Skip to main content
SaaSCity.io
DirectoriesLive LaunchesBlogWrite for UsAdvertise
Submit
Home/Blog/Better Models, Worse Tools: Why Newer AI Models Are Breaking Your Agent Tool Calls
Back to Blog

Build with AI

Better Models, Worse Tools: Why Newer AI Models Are Breaking Your Agent Tool Calls

Armin Ronacher discovered that Anthropic newest models are worse at following tool call schemas than their predecessors. Here is why, and what SaaS builders shipping AI agents need to do about it.

ghosty
ghosty
Founder, SaaSCity
July 5, 20266 min read
Better Models, Worse Tools: Why Newer AI Models Are Breaking Your Agent Tool Calls
Contents (6)
  1. Key Takeaways
  2. What Ronacher Actually Found
  3. Why Newer Models Got Worse, Not Better
  4. Claude Code's Flat Schema vs. Everyone Else's
  5. What Actually Fixed It
  6. What SaaS Builders Should Actually Do

Quick answer: Armin Ronacher, creator of Flask and an engineer at Sentry, found that Anthropic's Opus 4.8 and Sonnet 5 fail tool calls that Opus 4.5 handled cleanly, at roughly a 20% rate in agentic sessions. The models emit a correct payload and then append invented fields like requireUnique, oldText2 or newText2, which strict schema validation rejects. Stripping extended thinking blocks halved the failures; strict tool invocation mode eliminated them.

Armin Ronacher's post Better Models, Worse Tools on lucumr.pocoo.org, the writeup documenting that Anthropic's Opus 4.8 and Sonnet 5 invent fields in tool call arguments that older models did not

Armin Ronacher just found something that should make every SaaS founder shipping AI agents stop and read carefully.

Ronacher — creator of Flask, and an engineer at Sentry — noticed something wrong while running his own agent harness, Pi, against Anthropic's latest models. His writeup is worth reading in full, but the core finding is worth understanding now. Opus 4.8 and Sonnet 5, the newest and supposedly most capable models in the lineup, were failing tool calls that older models like Opus 4.5 handled without issue. Not occasionally. Often enough to be a real production problem — he clocked failure rates around 20% in agentic sessions.

This isn't a minor quirk. If you're building a SaaS product on top of Claude's tool-calling API — and a lot of you are, whether it's a coding assistant, a support bot, or an internal automation layer — this bug can silently corrupt your agent's actions in production.

Key Takeaways

  • The failure is extra fields, not wrong logic: The payload is correct, then the model appends keys such as requireUnique, oldText2 and newText2 that exist in no tool definition.
  • Newer is measurably worse here: Opus 4.8 and Sonnet 5 failed calls that Opus 4.5 handled without issue, at around 20% in agentic sessions.
  • Tool calls are in-band text: The model emits antml markers inline in its output stream and the harness parses that text, so it can keep writing past the end of the form.
  • Claude Code hides the problem: Its retry paths, parameter aliasing and unknown-key filtering absorb malformed calls, so no reward signal ever penalizes inventing fields.
  • Schema shape matters: Claude Code's flat file_path / old_string / new_string / replace_all edit tool is what the model drilled on; Ronacher's Pi harness uses a nested edits[] array and gets rejected.
  • Two fixes worked: Removing extended thinking blocks from context before the tool call roughly halved failures, and strict tool invocation mode eliminated them entirely.
  • Benchmark your own schemas: Published leaderboards do not test whether your field names and nesting survive the model's trained priors.

What Ronacher Actually Found

Anthropic's official tool use documentation for the Claude API, showing how tool schemas and input parameters are defined, the layer newer Claude models append invented fields to

The failure mode is oddly specific. The model generates a tool call — say, an edit operation with old text and new text — and the core payload is correct. The file path is right. The strings to replace are right. But then the model appends extra fields that were never part of the schema: things like requireUnique, oldText2, newText2. Fields that don't exist in any tool definition, invented out of nowhere.

Because these keys aren't part of the schema, strict validation rejects the whole call. The agent either errors out or, worse, silently drops the edit and moves on as if nothing happened.

Here's the part that makes this genuinely interesting instead of just a bug report: Ronacher traced it to how tool calls actually work under the hood. They aren't structured API objects the way you might assume. They're in-band text signaling — the model emits special markers (Anthropic's antml tags) inline in its output stream, and the harness parses that text to reconstruct the tool call. The model isn't filling out a form. It's writing text that looks like a form, and sometimes it keeps writing past the end of the form.

While you are here

Get your SaaS listed on SaaSCity

A permanent listing on the live city map, a DR 65+ dofollow backlink and a launch week in front of founders. Free with a badge, or skip the queue with Quick Pass — live within 24 hours.

Submit your SaaSWhat you get

Why Newer Models Got Worse, Not Better

Ronacher's hypothesis is that this comes down to what Anthropic trains against. Claude Code — Anthropic's own coding agent client — is remarkably forgiving. It has retry paths, parameter aliasing, and unknown-key filtering baked in. If a model appends garbage fields to a tool call, Claude Code often silently absorbs the error and the task still completes.

That's fine for Claude Code's own user experience. It's a problem for training. If Anthropic post-trains heavily on transcripts and reward signals generated inside Claude Code, a malformed tool call that still succeeds looks identical to a clean one from the reward model's perspective. There's no gradient telling the model "don't invent fields" — because inventing fields didn't cost it anything.

The result is a model that has adapted specifically to Claude Code's tool shapes and its tolerance for sloppiness. Opus 4.5, trained with less of this specific reinforcement, generalized better across different harnesses. Opus 4.8 and Sonnet 5 have a stronger prior — and when that prior meets a schema shaped differently than Claude Code's, the model fights you harder instead of adapting.

Claude Code's Flat Schema vs. Everyone Else's

The Claude Code repository on GitHub, Anthropic's coding agent whose forgiving retry and unknown-key filtering removes the training signal against malformed tool calls

Part of the mismatch is structural. Claude Code's own edit tool uses a flat argument shape: file_path, old_string, new_string, replace_all — four keys, no nesting. Ronacher's Pi harness uses a nested edits[] array, where each edit is its own object inside a list. Structurally reasonable, arguably cleaner, but not what the model has been most heavily drilled on.

AspectClaude Code (native)Pi harness (Ronacher's)
Argument shapeFlat: file_path, old_string, new_stringNested: edits[] array of objects
Malformed call handlingSilently absorbed via retries, aliasing, key filteringRejected by strict schema validation
Effect of extra invented fieldsHidden from the user, task still completesTool call fails outright
Model's effective training exposureHeavy — this is Anthropic's own agentLow — third-party, non-canonical shape

The mismatch is: Anthropic's model treats one specific tool shape as the default, and everything else — even reasonable, well-designed alternatives — becomes a place where the model's confidence outruns its accuracy.

What Actually Fixed It

Ronacher tested a few mitigations. Stripping the model's extended thinking blocks out of context before the tool call roughly halved the failure rate — suggesting the model was talking itself into inventing fields during its own reasoning trace. Switching to strict tool invocation mode, which constrains what the model is allowed to emit at the token level, eliminated the failures entirely.

That's a meaningful signal: the underlying capability to produce a correct call is there. The model isn't confused about the task. It's a decoding-time problem, and decoding-time problems have decoding-time fixes.

What SaaS Builders Should Actually Do

If you're running Claude models behind an agent product, this deserves action, not just awareness.

Benchmark your own schemas, not published leaderboards. Model benchmarks test general capability. They don't test whether your specific tool definitions, with your specific field names and nesting, survive contact with the model's trained priors. Run your own tool-calling reliability tests before you flip a version number in production.

Turn on strict tool invocation mode. If Anthropic's API exposes constrained decoding or strict schema enforcement for tool use, use it for anything that touches production state. The cost is some flexibility; the payoff is a model that literally cannot emit a field outside your schema.

Design flat schemas where you can. Nested arrays are cleaner engineering, but if the model's priors favor flat argument shapes, fighting that preference costs you reliability. Weigh that tradeoff deliberately instead of discovering it in an incident report.

Build forgiving retry logic, but log everything. Claude Code's silent-absorption approach isn't wrong for a chat UI — it's wrong for opacity. Absorb the extra fields if you must, but log every instance so you know how often it's happening and to what.

The bigger lesson: "upgrade to the newest model" is a reflex, not a strategy. Ronacher's writeup is a reminder that newer isn't automatically better for every axis that matters to your product — sometimes it's better at reasoning and measurably worse at the boring, load-bearing plumbing that keeps your agent from corrupting a customer's data. Test before you ship the upgrade.

Get your SaaS in front of founders

List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.

Submit your SaaSSee pricing

Founder resources

Best SaaS directoriesBest AI directoriesFree dofollow directoriesHigh-DR directoriesFree DR checkerLive launchesAI SaaS boilerplate

Related articles

Claude Sonnet 5: Anthropic's New Mid-Range Model and What It Means for SaaS Founders

Claude Sonnet 5: Anthropic's New Mid-Range Model and What It Means for SaaS Founders

10 Wildest Claude Code Projects Going Viral Right Now

10 Wildest Claude Code Projects Going Viral Right Now

Anthropic's launch-your-agent: Claude Code Agent Deployment From Zero to Live in One Session

Anthropic's launch-your-agent: Claude Code Agent Deployment From Zero to Live in One Session

Contents

  1. Key Takeaways
  2. What Ronacher Actually Found
  3. Why Newer Models Got Worse, Not Better
  4. Claude Code's Flat Schema vs. Everyone Else's
  5. What Actually Fixed It
  6. What SaaS Builders Should Actually Do

List your SaaS

$19.99one-time
  • Dofollow DR 65+ backlink
  • Live within 24 hours, no queue
  • Permanent listing on the city map
Submit your SaaS

Or list free with our badge

Done for you

We submit your SaaS to 200 directories

Every form filled in by us, every live URL in a report within 28 days.

$69one-time · launch price

Pick your project

City Sponsors

  • Nick LaunchesShip, launch, and get your product in front of real founders.
  • @peregrineintellPeregrine OS: pre-call intel for agency new business
  • Your product hereSlot open — 30 days, homepage + city
Become a sponsor
Write for this blog — from $99.99
SaaSCity.io

Directories are boring. We built a city instead. First isometric SaaS directory on the planet.

Platform
Submit SaaSLive LaunchesPricingBlogWrite for UsBacklink ExchangeMCP for AgentsAdvertise
Directories
Best SaaS DirectoriesBest AI DirectoriesBest Indie Hacker CommunitiesBest Subreddits for FoundersFree DR CheckerFree DR BadgeHow to Get SaaS BacklinksDirectory Submission Service
SaaSCity Alternatives
All ComparisonsSaaSCity vs Nick LaunchesSaaSCity vs BetterLaunchSaaSCity vs PeerPushProduct Hunt AlternativesSaaSHub Alternatives
Legal
Terms of ServicePrivacy PolicyRefund PolicyCookie PolicyCopyright & DMCASecurity
Company
AboutghostyContact

© 2026 SaaSCity.io

llms.txt