Skip to content

fix(anthropic): preserve MCP tool call context in spans - #844

Open
Kexin Chen (March-7) wants to merge 2 commits into
braintrustdata:mainfrom
March-7:fix/anthropic-mcp-tool-spans
Open

Kexin Chen (March-7) wants to merge 2 commits into
braintrustdata:mainfrom
March-7:fix/anthropic-mcp-tool-spans

Conversation

@March-7

@March-7 Kexin Chen (March-7) commented Oct 2, 2026 •

Copy link
Copy Markdown

Anthropic MCP connector results currently produce generic mcp tool spans without the corresponding tool name, input, or call type, even though the parent LLM span retains the original response. Recognize mcp_tool_use alongside built-in server tool calls and pair it with results by ID. For streamed beta messages, use the SDK's beta accumulator so MCP input JSON deltas reach the final span; accommodate the accumulator signatures used by both supported SDK versions.

Extend the existing MCP HTTP cassette test to check child spans for sync and async clients against the raw recorded provider response. Keep a separate typed SDK-event unit test for streamed MCP input split across JSON deltas. That unit test is not a recorded MCP SSE exchange. Existing HTTP requests and cassettes are unchanged.

Fixes #797.

Validation:

  • The sync/async VCR regressions fail against the pre-fix production code and pass with the fix in replay-only mode.
  • Anthropic 1.8.0: 61 passed on Python 3.10.20 and 3.14.5.
  • Anthropic 0.48.0: 44 passed, 17 skipped on each Python version.
  • Full-repository pre-commit checks and pylint on the updated test file passed.
  • The unchanged production code previously passed core tests (948 passed, 69 skipped, 12 xfailed), type checks and 36 type runtime tests, and full-repository pylint. Those suites were not rerun for this test-only review update.

AI-assisted implementation and testing with Codex.

Comment on lines +430 to +431
# Supplement the recorded provider response with deterministic ordering and
# incomplete-pair cases that cannot be requested reliably from a live model.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't think this is true! Please try using the vcr tests https://github.com/braintrustdata/braintrust-sdk-python/blob/main/docs/vcr-testing.md

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — I hadn't established that those cases couldn't be recorded. I've removed that claim and the synthetic test_anthropic_mcp_tool_spans_pair_by_id cases.

Following docs/vcr-testing.md, the sync/async regression reuses the existing MCP cassette unchanged. Its child-span expectations now come from the raw recorded provider response, including the call/result ID association, tool name, input, and result text. The HTTP request and cassette payload are unchanged. Both cases fail against the pre-fix production code and pass with this fix using --vcr-record=none.

The separate streaming-input test remains a constructed SDK-event unit test, not a recording of a live MCP SSE exchange. I haven't made new provider calls or generated a cassette from synthetic data.

Validation on Python 3.10.20 and 3.14.5: Anthropic 1.8.0 has 61 passing tests; 0.48.0 has 44 passing tests and 17 skips on each Python version. Pre-commit and pylint on the changed test file also pass.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[bot] Anthropic: MCP connector tool calls (mcp_tool_use/mcp_tool_result) are silently dropped from tool spans

2 participants