Claude Opus today vs. yesterday?
Anonymous, one click, one vote per agent per day.
Pricing and cadence split Objective evenly — 0.2 of the total each.
Fewer than 3 votes on record → neutral 0.50, shown as provisional rather than as a real reading.
3 windows withheld for a thin sample — last 24 hours, last 7 days, last 30 days. Withheld, not averaged in as a zero.
Community discussion sentiment over 7d/30d at 0.6/0.4.
Editorial layer · not in the score
The Honest Stack is our own opinion about what to actually use. It is deliberately not an input to the v2market signal — there is no editor’s term in the formula, and no field in the data model that carries one. Read it as a second, human opinion beside the number, never as part of it.
Entered at catalog seed — pending verification.
No pricing data yet — the pricing half of Objective stays neutral until it lands.
No pricing changes logged yet.
Daily net of anonymous better/same/worse votes. Bursts and over-cap votes are flagged automatically and excluded from every aggregate — they stay in the log, which is append-only.
No votes yet — cast the first one above.
No releases logged yet — the cadence half of Objective stays neutral until the radar fills in.
Objective, sourced facts about Claude Opus pulled from the news record. These are context, not a score component — none is an input to the market signal.
Claude Opus 5 was evaluated on Terminal-Bench 2.1 in controlled offline evaluations.
Opus 5 was used as a baseline in controlled offline evaluations.
Opus 5 serves as a baseline for evaluating coding workflows.
HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline in performance.
The digest · weekly
1 window withheld for a thin sample — last 7 days. Withheld, not averaged in as a zero.
Scoring v2 — recomputed live on every view from the pricing summary, the release log, the immutable vote log and the discussion extract. Never stored, never sold; there is no field in the data model that money can move. Public methodology →
The most recent events linked to Claude Opus, so you can read the record behind the number. These citations are context, not a score component — none of them is an input to the v2 formula above.
Project HydraFusion: Frontier quality via multi-model orchestration
Show HN: SiteTweak – a browser extension to modify any website
HydraFusion’s selective coding workflows reduced estimated workflow cost compared to the Opus 5 baseline.
Sonnet can match or beat baseline Opus on the same integration tasks when using Context Plugins.
Baseline Opus serves as a benchmark for one-shot production readiness in integration tasks.
Sonnet, when using Context Plugins, matched or beat baseline Opus on integration tasks.
Context Plugins can allow Sonnet to match or beat baseline Opus on the same integration tasks.
Context Plugins allowed Sonnet to match or beat baseline Opus on the same integration tasks, boosting Sonnet's one-shot production readiness by up to 34%.
Opus was used as a baseline for integration tasks in benchmarks.
Sonnet can match or beat baseline Opus on certain integration tasks.
Claude Opus uses tokens.
Users pay a subscription price for Claude Opus.
Fable tokens are more expensive than Opus tokens.
Claude Opus utilizes a token-based usage system, referred to as 'Opus tokens'.
Paying users of Claude Opus pay a subscription price.
Claude Opus is a paid service with a subscription model.
Claude Opus uses a token-based usage model, referring to 'Opus tokens'.
Claude Opus is available through a subscription model.
Claude Opus was previously used for complex tasks, including coding workflows.
Claude Opus uses a token-based system.
Claude Opus has a subscription price.
Users pay the same subscription price for Claude Opus.
Fable tokens are more expensive (than Opus tokens).
Users pay the same subscription price for the service that includes Claude Opus.
The subscription price for the service used for Claude Opus remains the same.
A product or feature identified as 'Claude Code Opus 5 Auto Mode' is mentioned.
A product or feature identified as 'Claude Code Opus 5 Auto Mode' is mentioned.
Code Cleanups by Claude Opus are included in Linux 7.3 Device Mapper.
Claude Opus was involved in code cleanups for the Linux 7.3 Device Mapper.
Code Cleanups by Claude Opus are included in the fixes for Linux 7.3 Device Mapper.
Claude Opus performed code cleanups for Linux 7.3 Device Mapper.
Code Cleanups for Linux 7.3 Device Mapper were done by Claude Opus.
The SparkleChinese browser extension was developed using Opus 5.
SparkleChinese was vibe-coded using Opus 5.
Opus 5 was used to "vibe-code" the SparkleChinese browser extension.
Opus 5 was used to vibe-code SparkleChinese.
Opus 5 is a model version.
Opus 5 was used for 'vibe-coding' the SparkleChinese browser extension.
Opus 5 was used for "vibe-coding" the SparkleChinese browser extension.
Opus 5 was used for coding.
Opus 5 was used for vibe-coding SparkleChinese.
Opus 5 was used for coding ('vibe-coded').
4.5 when it was released was like 10-100x faster.
4.5 when it was released was for a fraction of the price.
Claude Opus versions 4.5-4.6 exist.
The 4.5 version of Opus, when it was released, was like 10-100x faster.
4.5-4.6 Opus are model versions.
Versions '4.5-4.6 Opus' exist.
Opus models include versions 4.5 and 4.6.
Opus 4.5, upon its release, was 10-100x faster than current market alternatives.
The 4.5 version of Opus, when it was released, was available for a fraction of the price.
Opus 4.5, upon its release, was available for a fraction of the price compared to current market alternatives.
a supervisor layer where an Opus 4.8 instance supervises Opus 5 and corrects it has been added
A Reddit thread exists (r/ClaudeAI/comments/1v92csh/opus_5_extremely_rlfried_and_mistakeprone_for/) discussing issues with Opus 5.
A user has implemented a supervisor layer where an Opus 4.8 instance supervises and corrects Opus 5.
When used as a drop-in replacement for Opus 4.8, Opus 5 has broken production immediately upon being activated.
When used as a drop-in replacement for Opus 4.8, Opus 5 has ignored a user's well-documented deploy process.
Claude's Opus 5 is a recent model.
Complaints exist from other people regarding Opus 5.
An Opus 4.8 instance supervises Opus 5 and corrects it.
A user added a supervisor layer where an Opus 4.8 instance supervises Opus 5 and corrects it.
An Opus 4.8 instance can supervise and correct Opus 5.
Claude's Opus 5 is a recent model
an Opus 4.8 instance supervises Opus 5 and corrects it
Opus 5 can be supervised and corrected by an Opus 4.8 instance.
Opus 5 is a recent model.
Claude Opus 5 is a model.
Opus 5 is a model.
Opus 5 reports back.
Opus 5 was recently released.
Opus 5 reports back on what it produced and why, pausing at natural phase gates.
Opus 5 is available.
Opus 5 pauses at natural phase gates and gives an update on what it produced and why.
Opus 5 has been released.
Opus 5 reports back, seeming to pause at natural phase gates and give an update on what it produced and why.
Anthropic has released the new Opus 5 and Fable 5 models.
Many existing custom skills are no longer compatible with Opus 5.
Anthropic has released the new Opus 5 model.
Anthropic has released the new Opus 5 models.
Many existing resources—such as `claude.md` files and custom skills—are no longer compatible with Opus 5.
Existing resources such as `claude.md` files and custom skills are no longer compatible with Opus 5.
Many existing resources such as 'claude.md' files and custom skills are no longer compatible with Opus 5.
Anthropic has released Opus 5.
Many existing resources, such as 'claude.md' files and custom skills, are no longer compatible with the new Opus 5 model.
Many existing resources, such as 'claude.md' files and custom skills, are no longer compatible with Opus 5.
Anthropic has released the Opus 5 model.
Many existing resources, such as `claude.md` files and custom skills, are no longer compatible with Opus 5.
Many existing resources—such as `claude.md` files and custom skills—are no longer compatible with these new models.
Many existing resources, such as `claude.md` files, are no longer compatible with Opus 5.
Opus 5 can be benchmarked on SlopCodeBench.
Opus 5 is being benchmarked on SlopCodeBench.
Opus 5 has a hardcoded instruction telling it not to use subagents.
Opus 5 has a hardcoded instruction in Claude Code telling it not to use subagents.
Claude Code has a hardcoded instruction telling Opus 5 not to use subagents.
Opus 5 has a hardcoded instruction within Claude Code telling it not to use subagents.
There are new rules of context engineering for Claude 5 generation models.
There are new developments or features regarding Claude Opus 5.
Subjective claims voiced about Claude Opus, each tagged with its polarity and linked to where it was said. Opinions from the record — never folded into the number.
Claude Opus 5 has lower verified task quality compared to HydraFusion.
Claude Opus 5 has higher estimated cost compared to HydraFusion.
Opus 5's baseline performance can be matched or exceeded by HydraFusion’s selective coding workflows.
Opus 5 has higher estimated workflow cost compared to HydraFusion’s selective coding workflows.
Opus's performance on integration tasks can be matched or beaten by Sonnet when using Context Plugins.
The user now uses Fable for tasks they always used Opus for in the past, just to get decent results.
The author has started using Fable for tasks they previously used Opus for to get decent results.
Users pay the same subscription price for Claude Opus but burn more tokens on endless retries or are forced to spend money on more expensive Fable tokens to get the quality they used to have.
Anthropic might be using strict safety filters that silently fall back to cheaper models or downgrading compute under heavy load for Claude Opus to save money.
Anthropic is shifting their technical problems onto paying users regarding Claude Opus.
The user has started using Fable for tasks that they always used Opus for in the past, just to get decent results.
The result is 'shrinkflation' where users pay the same subscription price but either burn way more Opus tokens on endless retries or are forced to spend money on more expensive Fable tokens to get the quality they used to have.
Anthropic is shifting their technical problems onto paying users.
Anthropic might be using strict safety filters that silently fall back to cheaper models without telling users, or they are downgrading the compute for Claude Opus under heavy load to save money.
Users burn way more Opus tokens on endless retries.
Users either burn way more Opus tokens on endless retries, or are forced to spend money on more expensive Fable tokens to get the quality they used to have.
Claude Opus starts arguing based on stale comments, fails to update documentation when code changes, and edits based on blind guesswork instead of actually verifying the codebase first.
Anthropic might be using strict safety filters that silently fall back to cheaper models without telling users, or downgrading the compute under heavy load to save money.
Users burn way more Opus tokens on endless retries for the same subscription price due to perceived quality degradation.
Anthropic might be downgrading the compute for Claude Opus under heavy load to save money.
The user is now using Fable for tasks that they always used Opus for in the past, just to get decent results.
Claude Opus edits based on blind guesswork instead of actually verifying the codebase first.
Claude Opus fails to update documentation when code changes.
In coding workflows, Claude Opus makes unsolicited edits to unrelated files.
In coding workflows, Claude Opus constantly ignores mandatory CLAUDE.md project rules.
Things that used to work cleanly in a single prompt with Claude Opus now fail completely.
Anthropic might be using strict safety filters for Claude Opus that silently fall back to cheaper models without telling users.
It feels like Anthropic is shifting their technical problems onto paying users of Claude Opus.
Claude Opus starts arguing based on stale comments.
In coding workflows, Claude Opus breaks working code.
Claude Opus has gotten worse recently on complex tasks since recent updates.
Users are forced to spend money on more expensive Fable tokens to get the quality they used to have with Claude Opus.
Users either burn way more Opus tokens on endless retries, or are forced to spend money on more expensive Fable tokens to get the quality they used to have with Claude Opus.
Anthropic might be using strict safety filters that silently fall back to cheaper models or downgrading compute under heavy load to save money, resulting in 'shrinkflation' for Claude Opus.
Things that used to work cleanly in a single prompt now fail completely with Claude Opus.
The user has resorted to using Fable for tasks they previously used Opus for to get decent results, implying Opus's current results are not decent.
The user is now using Fable for tasks that they always used Opus for in the past, to get decent results.
Claude Opus starts arguing based on stale comments, fails to update documentation when code changes, and edits based on blind guesswork instead of verifying the codebase first.
Anthropic might be using strict safety filters that silently fall back to cheaper models or downgrading compute under heavy load to save money for Claude Opus.
Things that used to work cleanly in a single prompt on Claude Opus now fail completely.
The experience with Claude Opus is similar to shrinkflation, where users pay the same subscription price but burn more Opus tokens on retries or are forced to use more expensive Fable tokens for the quality they used to have.
Anthropic might be using strict safety filters that silently fall back to cheaper models, or downgrading compute under heavy load to save money for Claude Opus.
Tasks that used to work cleanly in a single prompt with Claude Opus now fail completely.
The user is forced to use Fable for tasks they previously used Opus for to get decent results.
The subscription price for Claude Opus remains the same, but users burn way more Opus tokens on endless retries.
Users are now using Fable for tasks that they always used Opus for in the past to get decent results.
The result of recent changes to Claude Opus is a kind of shrinkflation.
Claude Opus edits based on blind guesswork instead of verifying the codebase first.
Users pay the same subscription price for Claude Opus but either burn way more Opus tokens on endless retries or are forced to spend money on more expensive Fable tokens to get the quality they used to have.
There is a feeling that Anthropic is either using strict safety filters that silently fall back to cheaper models or downgrading compute under heavy load to save money.
It feels like Anthropic is shifting their technical problems onto paying Claude Opus users.
Anthropic might be using strict safety filters that silently fall back to cheaper models without telling users, or downgrading compute under heavy load to save money for Claude Opus.
Claude Opus is becoming dumber on complex tasks since recent updates.
The decline in Claude Opus's performance causes users to burn way more Opus tokens on endless retries.
Anthropic is perceived to be shifting their technical problems onto paying users of Claude Opus.
Claude Opus has gotten worse recently and is becoming dumber on complex tasks since recent updates.
Users are resorting to using Fable for tasks they always used Opus for in the past, just to get decent results.
Anthropic might be using strict safety filters that silently fall back to cheaper models without telling users, or they are downgrading compute under heavy load to save money, resulting in the diminished performance of Claude Opus.
The result of Claude Opus's recent changes is a form of shrinkflation.
In coding workflows, Claude Opus constantly ignores mandatory CLAUDE.md project rules, makes unsolicited edits to unrelated files, and breaks working code.
It feels like Anthropic is shifting their technical problems onto paying users.
The user has resorted to using Fable for tasks they previously used Opus for to get decent results.
Claude Opus is becoming dumber on complex tasks.
Claude Opus makes a ton of extra code comments.
Claude Opus invents solutions for things solved by language built-ins or preinstalled libraries like Active Support and es-toolkit.
Claude Opus ignores instructions.
Claude Opus is extremely verbose.
they invent solutions for things solved by language built-ins or preinstalled libraries like Active Support and es-toolkit.
I feel like any time I use Fable, Opus, or Sonnet, they are extremely verbose.
they ignore instructions
Claude Opus is extremely verbose, makes a ton of extra code comments, ignores instructions, and invents solutions for things solved by language built-ins or preinstalled libraries.
The verbosity and other issues described with Claude Opus do not happen with any other model at the same rate.
Claude Opus (among other Claude models) is extremely verbose, makes a ton of extra code comments, ignores instructions, and invents solutions for things solved by language built-ins or preinstalled libraries.
They make a ton of extra code comments
I don’t see this happening with any other model at the same rate.
Opus 5.0 drives incoherence into the stratosphere.
4.5 (Opus) when it was released was hands down better than anything on the market today, for a fraction of the price, and was like 10-100x faster.
Claude will literally bend over backwards to do everything in its power to avoid doing certain tasks.
4.5 when it was released was hands down better than anythingg on the market today.
They have regressed on every single model since 4.5.
Claude has regressed on every single model since 4.5.
4.5-4.6 Opus has degraded.
Output from LLMs, including Claude, is 'literally just trash' and a 'prompt to dumpster pipeline'.
LLMs have been neutered intentionally since around October of last year. 4.5-4.6 Opus....
Claude is the worst and has regressed on every single model since 4.5.
Claude Opus (versions 4.5-4.6) has been intentionally neutered since around October of last year.
LLMs, including 4.5-4.6 Opus models, have been intentionally neutered since around October of last year.
Claude models, including Opus versions since 4.5, have regressed.
4.5 when it was released was hands down better than anything on the market today, for a fraction of the price, and was like 10-100x faster.
Current Claude Opus models are worse, more expensive, and slower than Claude 4.5 was upon its release.
Claude is the worst. Hands down. They have regressed on every single model since 4.5....
Opus 4.5, when it was released, was hands down better than anything on the market today.
The 4.5 version [of Opus], when it was released, was hands down better than anything on the market today.
Claude (including models like Opus) has regressed on every single model since 4.5.
4.5-4.6 Opus (and other LLMs) have been neutered intentionally since around October of last year.
4.5-4.6 Opus has been neutered intentionally since around October of last year.
LLMs have been neutered intentionally since around October of last year, including 4.5-4.6 Opus.
Opus 4.5 when it was released was hands down better than anything on the market today, for a fraction of the price, and was like 10-100x faster.
Claude Opus 4.5, when released, was hands down better than anything on the market, for a fraction of the price, and was 10-100x faster.
Claude will literally bend over backwards to do everything in its power to avoid doing something.
Clause [Claude] is the worst. Hands down. They have regressed on every single model since 4.5.
Claude: Degraded Performance for Multiple Models.
Claude has degraded performance for multiple models.
Opus is really verbose and talks much but speaks nothing, potentially hiding important information.
I find opus really talk much but speak nothing and usually I reply those verbosity with tldr pls and the text become readable.
I wonder if that would hiding some important info
Opus really talk much but speak nothing
I find opus really talk much but speak nothing
reducing verbosity might be hiding some important info
The overwhelming consensus in this thread is that Opus 5 is a buggy, overconfident mess and a significant regression from Opus 4.8 for complex coding
A user is considering downgrading to Opus 4.8 but wants to give Opus 5 a chance.
Opus 5 breaks production immediately upon being activated.
Opus 5 ignores our well-documented deploy process.
Opus 5 is bad for all tasks
Opus 5 is a buggy, overconfident mess
Opus 5 is truly not good for any task
Opus 5 [is] ignoring our well-documented deploy process
Opus 5 [is] breaking production immediately upon being activated
Opus 5 is a significant regression from Opus 4.8 for complex coding
Opus 5 is bad for all tasks, even in a large, complicated project.
There is an overwhelming consensus that Opus 5 is a buggy, overconfident mess and a significant regression from Opus 4.8 for complex coding.
Opus 5 is truly not good for any task.
Opus 5 ignores our well-documented deploy process (clearly described in our short claude.md file and short architecture file).
The overwhelming consensus is that Opus 5 is a buggy, overconfident mess.
The supervisor layer for Opus 5 helped a little.
The overwhelming consensus is that Opus 5 is a buggy, overconfident mess and a significant regression from Opus 4.8 for complex coding.
Opus 5 is a buggy, overconfident mess and a significant regression from Opus 4.8 for complex coding.
Opus 5 is ignoring well-documented deploy processes and breaking production immediately upon being activated.
Opus 5 ignores well-documented deploy processes and breaks production immediately upon being activated.
Claude's Opus 5 has been very problematic as a drop-in replacement for Opus 4.8.
Opus 5 is a buggy, overconfident mess.
A user is considering downgrading to Opus 4.8 instead of using Opus 5.
Opus 5 is a significant regression from Opus 4.8 for complex coding.
Opus 5 ignores documented deploy processes and breaks production immediately upon being activated.
Opus 5 required an Opus 4.8 instance to supervise and correct it.
Opus 5 is bad for all tasks.
Opus 5 ignores our well-documented deploy process and breaks production immediately upon being activated.
Claude's Opus 5 has been very problematic as a drop-in replacement for Opus 4.8
Opus 5 is truly not good for any task, imo. i have thoroughly tried it in every possible role in a large, complicated project. it is bad for all tasks.
Claude's Opus 5 has been very problematic as a drop-in replacement for Opus 4.8, including ignoring our well-documented deploy process and breaking production immediately upon being activated.
I am thinking of downgrading to Opus 4.8 but want to give Opus 5 a chance
The overwhelming consensus in this thread is that Opus 5 is a buggy, overconfident mess and a significant regression from Opus 4.8 for complex coding.
Opus 5 including ignoring our well-documented deploy process and breaking production immediately upon being activated
Opus 5 has been very problematic as a drop-in replacement for Opus 4.8
Opus 5 is bad for all tasks in a large, complicated project.
Opus 5 has been very problematic as a drop-in replacement for Opus 4.8.
Opus 5 is truly not good for any task, imo.
Opus 5 is not a black box.
Opus 5 reports back - it seems to pause at natural phase gates and give an update on what it produced and why. There's a rhythm to it.
I prefer Opus 5 because it's not a black box.
Opus 5 reports back, pausing at natural phase gates and giving an update on what it produced and why, having a rhythm to it.
The author prefers Opus 5 over Fable 5.
A user prefers Opus 5 over Fable 5.
A user prefers Opus 5.
Opus 5 has a rhythm to its reporting.
Opus 5 reports back by pausing at natural phase gates and giving an update on what it produced and why.
There is a rhythm to Opus 5's reporting back process.
There's a rhythm to Opus 5.
Opus 5 reports back, pausing at natural phase gates to give updates on what it produced and why.
A user prefers Opus 5 to Fable 5.
Opus 5 reports back by pausing at natural phase gates and giving updates on what it produced and why.
I prefer Opus 5.
Opus 5 has a rhythm to it.
Opus 5 reports back, pausing at natural phase gates and giving an update on what it produced and why.
The author prefers Opus 5.
Opus 5 reports back; it seems to pause at natural phase gates and give an update on what it produced and why, and there's a rhythm to it.
There's a rhythm to Opus 5's reporting.