Powered by Smartsupp

Anthropic Launches Claude Sonnet 5: New Agentic AI Model with Autonomous Planning and Tool Use



By admin | Jun 30, 2026 | 3 min read


Anthropic Launches Claude Sonnet 5: New Agentic AI Model with Autonomous Planning and Tool Use

As agentic capabilities become the standard expectation for foundation model companies, Anthropic is launching Claude Sonnet 5—a more advanced and autonomous version of its mid-tier model. “It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models,” the company stated in a blog post. This messaging echoes what OpenAI and Google have claimed about their recent releases. OpenAI’s GPT-5.6 Sol debuted in preview last week as its most agentic model yet, enabling users to delegate tasks across subagents for extended autonomous projects. Google’s Gemini 3.5 Flash, launched in May, was positioned as a shift from a conversational chatbot to an agentic tool that plans, builds, and iterates on real work with minimal human input. Sonnet 5’s introduction confirms that agentic functionality is now the baseline expectation at every price point. The differentiator will no longer be who can deliver agentic work best, but who can do it most cheaply and reliably without human oversight. Sonnet 5 promises performance close to that of Opus 4.8, but at a much lower cost. Starting Tuesday, Claude Sonnet 5 will become the default model for free and Pro plans and will be available for all subscriptions. At launch, Sonnet 5 is priced at $2 per million input tokens and $10 per million output tokens through August 31, after which the price will rise to $3 per million input tokens and $10 per million output tokens. This makes Sonnet 5 cheaper than Opus 4.8, as well as OpenAI’s GPT-5.5 and Google’s Gemini 3.1 Pro, though it remains more expensive than Gemini 3.5 Flash.

The new model also shows significant improvements over its predecessor, Sonnet 4.6, released in February, in agentic performance areas like reasoning, tool use, software coding, and knowledge work, according to Anthropic. For instance, on one benchmark, Sonnet 5 scores 63.2% on agentic coding, compared to Opus 4.8’s 69.2% and Sonnet 4.6’s 58.1%. On a knowledge work benchmark, Sonnet 5 actually slightly outperforms Opus 4.8, which is known for excelling at solving the hardest problems, such as making subtle judgment calls and conducting deep research. “Opus 4.8 is still the model of choice for higher accuracy on these tasks, but Sonnet 5 provides developers with lower-priced options that are of much higher quality than what was previously available,” Anthropic notes. “Between Sonnet 5 and Opus 4.8, users can adjust the effort level to find the right balance of cost and performance.”

According to testers quoted in the blog post, Sonnet 5 also excels at completing complex tasks where previous model versions would have stopped short and “checks its own output without explicitly being asked.”

“We handed Claude Sonnet 5 a two-part job—update Salesforce account tiers, send a launch announcement to enterprise contacts—and it finished end to end,” Daniel Shepard, a senior engineer at Zapier, said in a statement. “That used to stall halfway. For day-to-day automation, it’s a no-brainer.”

On safety, Sonnet 5 also demonstrates a lower rate of “undesirable behaviors,” such as cooperation with misuse and deception, compared to its predecessor, making it safer for use in agentic contexts. It is better at refusing malicious requests and sidestepping hijack attempts in prompt-injection attacks. It also hallucinates and engages in sycophantic behavior at a lower rate than Sonnet 4.6. However, it does not match the level of Opus 4.8 and Claude Mythos Preview when it comes to misaligned behavior. “Evaluations also show that it has a much lower ability to perform dangerous cybersecurity tasks than our current Opus models,” the blog post states. Lovable co-founder Fabian Hedin said in a statement that Claude Sonnet 5 “refuses unsafe requests cleanly and consistently.”

“At Lovable, we’re putting powerful tools in the hands of millions of builders,” Hedin added. “A model that knows when to say no is just as important as one that knows how to build.”




RELATED AI TOOLS CATEGORIES AND TAGS

Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!