September 23, 2026

Anthropic Releases Claude Opus 5.5

Anthropic released Claude Opus 5.5, the first model in their new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. The release comes after Anthropic publicly called for pacing the frontier, and the model was tested by external evaluators including Frontier Design and METR before release[1].

What is new

Opus 5.5 is a major step up from Opus 5. One early tester completed a 680'000-line code migration in less than a day, work that would have taken an engineering team weeks. When asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app's behavior[1].

On pricing: input tokens are $4 per million, output tokens $20 per million, both 20% less than Opus 5. Cache reads, which make up the majority of agentic and coding work costs, are $0.20 per million tokens, 60% less than Opus 5. Output generation is more than 30% faster than Opus 5[1].

Safety

Opus 5.5 achieves the best scores of any model to date on Anthropic's automated behavioral audit, their alignment suite that tests Claude across thousands of simulated scenarios. It is less likely than recent models to take hard-to-reverse actions or act outside given boundaries, and more resistant to prompt injection than Opus 5. Anthropic broadened alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though they acknowledge limits remain[1].

Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity capabilities, Anthropic is deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply to the Life Sciences Verification Program for biology research. The Cyber Verification Program will expand access to verified cybersecurity practitioners in the coming weeks[1].

Benchmarks

Opus 5.5 leads on agentic coding, computer use, and knowledge work benchmarks. Terminal-Bench 4.0: 66.4% (vs Opus 5 at 52.3%, GPT-6 Astra at 57.9%). FrontierCode v1.1: 54.4%. CursorBench 4.0: 57.8%. Humanity's Last Exam: 67.7% with tools. OSWorld 2.0 computer use: 81.8% partial[1].

But Anthropic itself notes a caveat: "at these levels of capability, benchmark margins have become a less reliable guide to real-world differences." In their own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than the scores suggest. The real advantage is efficiency, not raw capability[1].

Communication

Early testers found Opus 5.5's writing clearer and easier to follow, addressing common feedback about Opus 5. It puts important information up front and works better over long sessions. As one tester put it: "it writes the way I do." Anthropic says this is a safety benefit as well as a practical one, because clearer output is easier to check[1].

What comes next

Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety. Five-hour usage limits are increasing on Pro, Max, Team, and seat-based Enterprise plans, and subscription users get a rate limit reset they can save and use whenever they choose[1].

On Hacker News, the announcement drew 1'632 points and 997 comments, making it one of the most discussed launches of the week[2]. An independent performance and price analysis by Artificial Analysis also reached the front page with 311 points[2].

Sources

[1] Anthropic: "Introducing Claude Opus 5.5" (September 2026)

[2] Hacker News: "Claude Opus 5.5" (1'632 points, 997 comments) and "Claude Opus 5.5 Intelligence, Performance and Price Analysis" (311 points)

← Back to all posts