Anthropic ships Sonnet 4.7 for agent reliability
A release that targets reliability over benchmark fireworks, answering OpenAI with a different posture.
Mariano De Vitto · May 2026
Anthropic released Claude Sonnet 4.7 this week. The framing was deliberate. Where OpenAI's GPT-5.5 launch in late April leaned on agentic worker demos and a release cadence message, Sonnet 4.7 went the other direction. The headline number was agent reliability: a measurable drop in derailment on long tool-call chains, fewer abandoned tasks, and a sharper response to negative feedback during multi-step work.
Reliability as a competitive thesis
The benchmark signal is consistent. Sonnet 4.6 already led the GDPval-AA Elo benchmark for knowledge work tasks. Sonnet 4.7 extends the lead and adds explicit metrics on what the AI safety community calls task persistence: when an agent encounters an unexpected page state, an authentication wall, or a malformed API response, does it recover or does it stop? The 4.7 numbers are higher than any other public model on persistence under real-world conditions.
Why this matters for marketing
For marketers, this is the variable that actually matters. The first wave of agentic marketing tools (Meta's AI connectors, Salesforce Agentforce, HubSpot Breeze, Snapchat AI Sponsored Snaps) all assume the underlying model can run an end-to-end task without giving up halfway. The economics of agentic marketing fall apart when the agent succeeds 60% of the time. They start to work when the agent succeeds 90%. The gap between demo and production lives in those numbers.
Read alongside OpenAI's six-week release cadence and Google's Gemini 3.1 Ultra, the labs are converging on three different bets. OpenAI is racing on speed of iteration. Google is racing on context window and multimodal depth. Anthropic is racing on the variable that decides whether agents ship in production. For brand-side marketing organisations, the procurement question gets more interesting, not less.
Refusal calibration as a quieter signal
There is also a quieter signal in the release notes. Anthropic flagged improved refusal calibration: the model declines fewer reasonable requests while still declining the genuinely unsafe ones. That is the variable that has been blocking enterprise adoption in regulated industries (financial services, healthcare, ad-targeted content for minors). If 4.7 actually lands the calibration, brand legal teams that were skeptical of agent rollouts may reopen the conversation.
The practical to-do this quarter: re-run any agent evaluation you did before March against the latest Sonnet. The work that failed in February might pass now. Conversely, agents you had in production on Sonnet 4.5 may behave differently as the underlying model gets quieter and more careful. Test before promoting.
This week's signal: Sonnet 4.7 sells reliability over speed. That is the variable that decides whether agents ship in production.
The Signal Brief · Mariano De Vitto — Head of Marketing, Barcelona