Home / Blog / Claude Code and Codex break on different MCP features
Tech News

Claude Code and Codex break on different MCP features

Benchmark testing on ten MCP features revealed five split failures between Claude Code 2.1.287 and Codex 0.162.0, with neither failing the same test.

By Dillip Chowdary โ€ข Oct 11, 2026 โ€ข Source: m3.sineframe.com

Claude Code and Codex break on different MCP features

Sineframe introduced an empirical evaluation matrix called the Harness Quirks Matrix to measure how coding agents handle the Model Context Protocol, detailing their findings in m3.sineframe.com's report. Testing ten distinct protocol features against Claude Code and Codex across multiple trials revealed that half of the tested features succeeded on one harness while completely failing on the other. Rather than demonstrating overlapping weaknesses, the two developer assistants split failures across entirely disjoint protocol features.

This comparison analyzes the benchmark results generated by the open-source M3 evaluation harness, detailing where tool calls drop, schema parsing breaks, and output buffers truncate data. It is intended for software engineers and server maintainers who build or host tools using the Model Context Protocol and need to ensure reliable operation across different client runtimes. The data outlines specific failure conditions, reproduction methods, and configuration thresholds identified across both AI programming tools.

The test: Claude Code and Codex break vs the other

On 9 October 2026, Sineframe ran ten specific Model Context Protocol features through M3 against Claude Code version 2.1.287 paired with claude-sonnet-5-5, and Codex version 0.162.0 paired with gpt-5.6-sol. Each test cell executed three independent trials, requiring the tool invocation to reach the underlying server with correct arguments and return output directly into the model response to count as a passing grade. Claude Code failed three feature suites and Codex failed two, with neither client failing the same test. The full evaluation required 66 total runs, taking 16 minutes to execute across both client binaries. Claude Code averaged 9.8 seconds per run, while Codex averaged 18.2 seconds.

Direct binary probes without M3 uncovered eight additional edge-case failures across Claude Code version 2.1.295 and Codex version 0.162.0. These targeted tests exposed schema boundary mismatches, naming collisions, and type truncation limits in standalone runs. Several legacy failure assumptions were cleared during the benchmark. Root-level anyOf declarations in inputSchema schemas and 65-character tool names now execute cleanly on Claude Code, while resource_link blocks now pass reliably on Codex. These fixes confirm that older workarounds for those three specific behaviors can be retired by server maintainers who target current versions of both agents.

How Claude Code and Codex break and the other each did

Claude Code and Codex break on different MCP features
Illustration ยท Pexels

Claude Code completely drops tool invocations if an argument property uses bracketed array notation such as ids[]. The tool remains hidden without throwing an error, which directly impacts servers generated from OpenAPI specifications. In addition, Claude Code truncates large tool error logs by preserving only the first 5,000 characters and the final 5,000 characters of a 12,000-character payload, causing debugging markers positioned in the middle of error messages to disappear. Furthermore, when an MCP server returns a payload containing both standard text content and structuredContent blocks, Claude Code surfaces only the structured data and loses the text field.

Codex fails dynamic catalog refreshes, maintaining its initial tool list and ignoring new tools advertised after a tool-list change notification. When parsing large input schemas, Codex compresses parameters to fit an internal budget; on a 14 KB schema, this compression stripped out required nested fields and broke tool execution. Setting Codex per-server schema budget to 20,000 restored the missing fields. Because Codex operates in code mode with gpt-5.6-sol, it converts incoming JSON schemas into compacted TypeScript declarations, which introduces distinct schema translation issues. Codex also halts entirely if a prompt instructs it to stop when a tool is missing.

Advertisement

Tech Pulse Daily

Get tomorrow's pulse first

Join engineers who read Tech Pulse before stand-up. Free, weekday mornings.

Claude Code and Codex break vs the other, side by side

MCP Feature TestedClaude Code 2.1.287Codex 0.162.0
Tool argument named ids[]0/3 (dropped)3/3 (works)
Marker inside 12,000-char error0/3 (result lost)3/3 (works)
Mixed text and structuredContent0/3 (result lost)3/3 (works)
New tool added after catalog change3/3 (works)0/3 (dropped)
Nested fields in 14 KB schema3/3 (works)0/3 (dropped)
Root-level anyOf in inputSchema3/3 (works)3/3 (works)
outputSchema declaring draft-073/3 (works)3/3 (works)
Omitting default parameters3/3 (works)3/3 (works)
Tool name of 65 characters3/3 (works)3/3 (works)
resource_link block in result3/3 (works)3/3 (works)

Direct binary evaluations revealed further distinctions between the two environments. In unmanaged binary testing, Codex lost property names when encountering root oneOf constructs alongside standard properties, dropped all tools beyond an index threshold of 2,048 on a single server containing 2,049 tools, and erased enum definitions inside anyOf blocks when schemas exceeded 5 KB. When handling 64-bit integers such as 1234567890123456789, both agents distorted precision through rounding or string conversion. Claude Code discarded an entire server when a single tool lacked a root type property, misrouted calls between fetch.page and fetch_page, and registered errors when text results accompanied declared outputSchema blocks.

See the Claude Code and Codex break vs the other output

Execution harnesses confirm that failures happen along the transport boundary between the agent client layer and the model runtime. For every test cell, M3 executes an isolated pytest evaluation against pinned local binaries while tracking wire traffic with network traces, run manifests, and binary SHA-256 hashes. To prevent environmental false positives, each run deploys an auxiliary echo server alongside the target server under evaluation. If the control echo server encounters an error, M3 marks the cell as unmeasured rather than recording an agent defect, ensuring clean separation between harness errors and agent protocol failures.

The test suite ensures deterministic evaluation by supplying random session tokens on each run, preventing language models from fabricating successful tool responses out of cached prompt context. When evaluating bracketed array properties with test_q05_args_property_name_brackets, M3 evaluates client output directly with assertion helpers like expect(result).to_have_tool_call. The framework separates client-side reporting from server-side receipt, verifying whether the binary dispatched the call across the wire. This wire tracing confirmed that Claude Code suppressed bracketed parameters before server transmission, whereas Codex transmitted bracketed parameters without error.

The verdict on Claude Code and Codex break vs the other

Neither coding assistant achieved full compliance across the suite of Model Context Protocol behaviors. Claude Code 2.1.287 excels at parsing complex 14 KB input schemas and seamlessly picking up dynamic catalog updates, but it stumbles on bracketed parameter syntax, mixed structured responses, and large error buffers. Conversely, Codex 0.162.0 handles bracketed parameters, oversized error strings, and combined text and structured payloads without data loss, but fails when servers update their catalogs dynamically or present schemas that exceed default schema translation budgets.

Because both tools exhibit contrasting protocol blind spots, maintainers must construct server implementations that accommodate both runtime profiles. Server authors targeting Claude Code should place crucial error context within the first 5,000 characters, avoid bracketed naming conventions in schema fields, and separate structured responses from raw text. For Codex compatibility, developers must avoid schemas with nested fields over 14 KB unless the user expands the schema budget, restrict server tool catalogs to 2,048 items, and avoid relying on dynamic catalog registration during an active session.

Developer Action Items

  • โ˜ Diff the official changelog for Claude / Sonnet / Codex 2.1.287 before you bump โ€” APIs, defaults, and removed flags only.
  • โ˜ Install through the vendor's documented channel in staging; keep a one-command rollback and time-box the canary.
  • โ˜ Grep your repo for old flag names, lockfile pins, and plugin versions that the notes mark as breaking.
  • โ˜ Prefer the first patch cut over the day-zero tag unless you have a reason to be on the leading edge.
  • โ˜ If HN Claude/Codex/Fable did not name a region, plan, or SKU, screenshot the official availability line before you promise it to users.

Claude Code and Codex break on different MCP FAQ

Which MCP features failed on Claude Code?

Claude Code failed tool calls using bracketed parameter names like ids[], lost the middle section of error strings exceeding 10,000 characters, and dropped text content when combined with structuredContent.

Why did Codex fail on large input schemas?

Codex compacts incoming JSON schemas into TypeScript definitions to fit an execution budget, which stripped out required nested properties in a 14 KB schema unless its schema budget was increased to 20,000.

How does M3 prevent false positive test results?

M3 runs an echo control server alongside each test, assigns dynamic tokens to stop models from guessing answers, logs wire traces with binary SHA-256 hashes, and re-tests failures outside the harness.

Sources

Dillip Chowdary

Author

Dillip Chowdary

Writes Tech Bytes coverage of AI, engineering, and the tools that actually ship. Editor of Tech Pulse Daily.

Related on Tech Bytes

Advertisement

5-min tech signal

Weekday briefing for engineers who skip the noise.

No spam ยท Unsubscribe anytime

Advertisement

โœˆ๏ธ CareerPilot

Your AI job-search copilot

Match your resume against live Ashby, Greenhouse & Lever openings โ€” fit scores, job-specific resume optimization and email alerts.

Find matching jobs โ†’

Free Tools

Browse all tools โ†’