Claude Sonnet 5 arrives as the most agentive Sonnet so far: it can plan, use tools like browsers and terminals, and execute autonomous tasks that previously required larger, more expensive models.
What is Claude Sonnet 5
Sonnet 5 is the latest iteration of Anthropic’s Sonnet family, built to do agentive work: follow plans, use external tools, and complete multi-step flows with less human tuning. Its goal is to give you much of the capability of Opus models but at a more accessible price.
So what does that mean in practice? It means complex engineering, automation, and analysis tasks can now be solved with fewer steps and lower cost, without throwing away safety or quality.
Performance and concrete use cases
Anthropic shows Sonnet 5 clearly improves on Sonnet 4.6 in reasoning, tool use, coding, and knowledge work. It doesn’t match Opus 4.8 for peak accuracy, but it hits an attractive middle ground on cost-to-performance.
Early testers share very concrete examples: Sonnet 5 updates records in Salesforce and sends outreach to business contacts in a full flow; it reproduces and fixes bugs by writing tests and confirming regressions in one pass; it reviews pull requests and delivers verified changes. Sounds like magic? It’s really sustained execution: it follows plans, respects conventions, and finishes what it starts.
It also shines on tasks people usually avoid: brownfield code with race conditions, hidden tests, or long-standing issues. In legal research it shows noticeable improvements, and in live data analysis (for example with ClickHouse) it shortens the time to reach actionable insights.
Practical summary: Sonnet 5 delivers more agentive work with fewer steps and at a lower price than before, while keeping quality in many real scenarios.
Security and limitations
In pre-release evaluations, Sonnet 5 showed a lower rate of undesirable behaviors than Sonnet 4.6: better rejection of malicious requests, fewer hallucinations, and less sycophancy. However, in automated audits overall it has slightly higher rates of misaligned behavior than Opus 4.8 and Claude Mythos Preview.
Important: Sonnet 5 was not intentionally trained for cybersecurity tasks. In exploit tests (for example vulnerabilities in Firefox) it did not produce full exploits and generally performs much worse than Opus 4.8 on potentially dangerous cyber skills. For that reason Sonnet 5 launches with active cyber safeguards by default: real-time detection and blocking of risky uses.
Anthropic recommends Opus 4.8 for cybersecurity work that needs fewer guardrails.
Availability and pricing
Sonnet 5 is available today on all plans: it’s the default model on Free and Pro, and accessible for Max, Team, and Enterprise. It’s also in Claude Code, the Claude Platform, and via the API as claude-sonnet-5.
Introductory pricing (valid through August 31, 2026):
- $2 per million input tokens
- $10 per million output tokens
After that date it will rise to $3 / $15 per million. Anthropic also changed the tokenizer: the same input can map to ~1.0–1.35× tokens depending on content type. The introductory price aims to make the transition roughly cost-neutral.
Rate limits were also increased on Chat, Cowork, Claude Code and the Claude Platform to support higher usage at heavier effort levels.
What this means for developers and companies
If you work on automations, technical support, data pipelines, or legacy code maintenance, Sonnet 5 lets you delegate more repetitive and technical steps without paying the premium for top-tier models. For teams that need a balance between price and agentive capability, it’s a very reasonable option.
If your focus is offensive cybersecurity or exploit analysis, Opus 4.8 remains the better choice, per Anthropic’s recommendation. And if your organization needs formal cyber verification, Sonnet 5 is included in the Cyber Verification Program on certain platforms.
Final reflection
Sonnet 5 isn’t a giant leap on every front, but it’s a practical transition: more autonomy and better execution for real workflows, without requiring the infrastructure or budget of top-end models. For many teams, that’s exactly what they needed.
