Claude Opus 5.5 Undercuts Its Own Flagship by 40 Percent
Anthropic shipped Claude Opus 5.5 today, and the headline number is the price: $4 per million input tokens and $20 per million output, a 20% cut from Opus 5 that adds up to 40% lower cost on typical workloads once caching is factored in. Cache reads dropped even further, from $0.50 to $0.20 per million tokens.
The performance claims are aggressive for a cheaper model. Opus 5.5 scored 66.4% on Terminal-Bench 4.0 against Fable 5.1’s 55.8% and Opus 5’s 52.3%, and posted similar gaps on FrontierCode v1.1 and CursorBench 4.0 — Anthropic’s mid-priced model is now beating its own most expensive one on agentic coding. Output generation is also 30% faster, and a new Fast mode at $8/$40 per million tokens runs up to 2.5x faster still.
Early testers are backing the numbers with anecdotes: GitHub says it solved more terminal tasks in less than half the steps, an Optiver engineer reported 40 to 50% cost cuts on agentic coding tasks at matched quality, and one tester finished a 680,000-line migration in under a day. It’s live now on Claude, AWS, Google Cloud and Microsoft Azure, with Sonnet 5.5 and Haiku 5.5 due in the coming weeks.
The Safety Numbers Anthropic Chose to Lead With
Buried under the pricing news: Anthropic says Opus 5.5 posted its best score yet on the company’s automated behavioral audit, with an 85% reduction in attempts to circumvent sandbox containment compared to earlier models and improved resistance to prompt injection.
The safeguards are inherited rather than loosened. Cybersecurity protections mirror Fable 5.1’s, including a fallback to Opus 4.8 for the riskiest cyber tasks, and biology safeguards still require verified researchers through the Life Sciences Verification Program. Thinking mode is preserved as an anti-distillation measure, and the model ships with EU AI Act compliance watermarking and a zero data retention option.
Read against this week’s context, the framing looks deliberate. Anthropic is racing OpenAI and xAI on price while trying to hold the line that its models are getting safer, not just cheaper — a distinction that matters more the closer the industry gets to genuinely interchangeable pricing.
2.1.280 Makes Opus 5.5 the Default, and Breaks a Few Things on the Way In
Claude Code 2.1.280 shipped alongside the model, making claude-opus-5-5 the new default Opus and switching Pro and Team Standard plans from Sonnet to Opus as their default model outright. Rate limits went up across Pro, Max, Team and Enterprise to match.
The upgrade isn’t a free lunch for existing integrations. Adaptive thinking is now always on and can’t be disabled, forced tool use can error out where it didn’t before, and older computer-use tool calls get rejected rather than silently degraded — Anthropic’s polite way of saying update your app before you deploy this in production.
The rest of 2.1.280 is quieter housekeeping: mouse support extended to more fullscreen lists, a new CLAUDE_CODE_MAX_MCP_DESCRIPTION_LENGTH variable for the 2,048-character MCP description cap, and fixes for auto mode retrying declined actions and for writes through symlinked paths.
Fast Mode, Cheaper Cache, and a Context Window That Didn’t Shrink
On the API side, Opus 5.5 keeps the 1M-token context window and adds a Fast mode at $8 input / $40 output per million tokens for up to 2.5x the speed when latency matters more than cost. Cache write pricing lands at $5 per million tokens, with reads down to $0.20.
Anthropic is explicit that this generation competes on cost per completed task, not raw benchmark score — GDPval-AA and OSWorld numbers are in the announcement, but so is a running tally of token counts per finished job from customers like Spotify and Optiver. That’s a tell about where the sales conversation is heading next.
One number worth flagging for anyone budgeting: Anthropic says Opus 5.5 beat a rival flagship on software-development benchmarks at roughly a third of the operating cost — a claim made hours before the rest of the industry answered with price cuts of its own (more on that below).
GitHub, Spotify and Optiver Ran the Numbers Before Anthropic Published Them
Anthropic’s launch post leaned hard on customer quotes instead of just benchmarks. GitHub’s chief product officer said Opus 5.5 “solved more terminal tasks than Opus 5 in less than half the steps.” A Spotify engineer said the token efficiency gain let the team “complete the same tasks both cheaper and faster.”
Optiver’s global head of AI engineering put a number on it: Opus 5.5 matched Opus 5’s output quality in about half the turns, time and output tokens, cutting costs 40 to 50% on agentic coding tasks. Separately, one early tester audited a 200,000-line codebase in three hours, down from 20-plus hours on Opus 5.
None of this is independently verified — it’s Anthropic’s own launch post — but the pattern across unrelated companies, a dev platform, a music streamer and a trading firm, all landing in the same 40-to-50% range is the kind of convergent anecdote that’s hard to fake convincingly.
Grok 4.7 Rolled Into GitHub Copilot the Same Day Opus 5.5 Shipped
xAI’s Grok 4.7 launched at $2 per million input tokens, undercutting Anthropic’s new price roughly 90 minutes after Opus 5.5 went live, and rolled out instantly inside GitHub Copilot — the same platform whose product chief had just endorsed Opus 5.5.
That’s the ecosystem story underneath the model news: Copilot, Cursor and the rest of the agentic-coding tool layer are becoming pricing-agnostic, plugging in whichever frontier model is cheapest or fastest that week rather than committing to one lab. For developers, that’s good news — model choice is now a dropdown, not a vendor lock-in decision.
Three Labs Cut Prices in 48 Hours, Which Is Exactly What the Antitrust Complaint Said Wouldn’t Happen
Zoom out and today looks less like three separate launches and more like one coordinated collapse of a truce that was never actually signed. Anthropic priced Opus 5.5 at $4/$20. Roughly ninety minutes later, xAI’s Grok 4.7 undercut it at $2 per million input tokens. Within two days, OpenAI answered with GPT-6 Sol and Luna at 50% below the prior GPT-5.6 generation — a three-tier family that makes Astra, the old flagship, look like the expensive legacy option.
The timing is the story. Just two days before this, a federal antitrust complaint accused Anthropic, OpenAI, xAI and Google of coordinating a slowdown in AI development off the back of Dario Amodei’s pacing essay. Whatever coordination existed on capability, there’s now visible proof there wasn’t any on price — and price is usually the easier thing to collude on, not the harder one.
For buyers, the practical upshot is that frontier-model pricing is now genuinely volatile week to week, which changes how you budget an agentic pipeline. For the labs, it’s a signal that “pace the frontier together” was never going to survive contact with three companies that each think they’re currently winning. The essay asked for restraint. The market answered with a price war inside 48 hours.