Repository navigation
fix: Log output and metrics for incomplete OpenAI Responses API streams - #2540
Merged
Luca Forstner (lforst) merged 8 commits intoOct 2, 2026
Conversation
Luca Forstner (lforst)
approved these changes
Oct 2, 2026
Member
|
released with https://github.com/braintrustdata/braintrust-sdk-javascript/releases/tag/braintrust%403.37.0 - thanks for the PR Raphael Fakhri (@RaphaelFakhri) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Streaming Responses API calls that end with
response.incompleteorresponse.failednow get their output, metadata and token metrics logged. Forresponse.failed, the response'serroris also logged as the span error.The OpenAI plugin only read the final
responseobject fromresponse.completedevents. A stream that stops atmax_output_tokensor because of a content filter ends withresponse.incompleteinstead, and a stream that errors server-side ends withresponse.failed. Both events carry the same finalresponseobject with the partial output and usage, but we ignored them, so the span had no output and no token metrics even though the request was billed.aggregateResponseStreamEvents(responses.create({ stream: true })) and theresponses.stream()event handler now both treat these three event types as terminal. The OpenRouter plugins already do the same. To log theresponse.failederror,traceStreamingChannel'saggregateChunksresult can now return an optionalerror.Tests
openai-plugin.test.tscover output, metadata and metrics for all three terminal events, theresponse.failederror, and ignoring non-terminal events that carry aresponse.responses-create-stream-incompleteandresponses-stream-incompletee2e operations inopenai-instrumentationhitmax_output_tokensagainstgpt-4o-mini. Cassettes were recorded for all v4/v5/v6 pinned and latest variants.