Skip to content

feat(ai, ai-harness): add the harness stack with replay and adapter parity - #1555

Open
AlemTuzlak wants to merge 299 commits into
mainfrom
feat/harness-p14-mcp-server
Open

AlemTuzlak wants to merge 299 commits into
mainfrom
feat/harness-p14-mcp-server

Conversation

@AlemTuzlak

@AlemTuzlak AlemTuzlak commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Review the combined harness and replay stack in this PR. The harness keeps one agent conversation open across turns, restarts, and model changes. This combined stack adds the harness, CLI, dashboard, MCP in both directions, durable sessions, host hooks, and compaction. It also keeps saved history usable across adapters and checks final tool input before execution or client dispatch.

🎯 Changes

Compatibility changes

Area Caller impact
Reasoning The owner stack replaces provider reasoning fields in modelOptions with chat({ reasoning }). The migration guide remains part of this PR.
Default execution and caching Server tools run in parallel. Dependent tools need sequential: true or toolExecution: 'sequential'. Prompt caching is on by default. Use promptCache: 'none' to disable it.
Usage and saved history Anthropic, Bedrock, and Claude Code report total input in promptTokens. Assistant block order can produce more AG-UI rows. Source metadata separates the requested model from the response model.
Host and compaction Host joins stop at a refused or cancelled steer. Old fold checkpoints rebuild once. Compaction events add reason, and withCompaction adds compactNext. Durable turns keep LogRecordsCapability.
Tool validation and durable resume Final validation follows middleware. Invalid input can reject a call that previously ran. Older same-run phases work when saved approval and tool bindings match. Cross-run phases need trusted saved approval context. Without it, resume is rejected and the client phase stays pending. Malformed context is rejected in both cases.

Saved subagent cards use a separate view of the transcript. Ordinary text and ambiguous hosts stay in that view. Both host creators preserve unrelated metadata. They choose an unused ID for missing IDs and explicit duplicate IDs, while keeping a unique explicit ID.

Replay changes provider input. It keeps the saved transcript unchanged. History without source metadata retains its legacy source treatment.

Existing harness stack

This remains the single combined PR for phases P0 to P14. It replaces #1551 and the closed phase PRs. Their descriptions remain the review record.

Phase PR Behavior
P0 #1513 Subagents return values and call activities.
P1 #1515 Sessions, typed agents, and plugins.
P2 #1518 Resume crashed turns and serve sessions to clients.
P3 #1519 Commands, settings, auth, and first-party plugins.
P4 #1520 Subagent tree limits and harness children.
P5 #1521 Build artifacts, worker mode, and remote harnessText.
P6 #1522 Dashboard and runnable example.
P7 #1523 MCP connectors with browser sign-in and runtime tools.
P8 #1524 Code mode with pluggable isolates.
P9 #1538 Delegate to coding agents.
P10 #1540 Continue work until a goal judge accepts the result.
P11 #1546 Live session state for any UI.
P12 #1549 CLI with a replaceable UI.
P13 #1550 Middleware and usage for every agent run.
P14 #1554 Harness MCP server, --mcp, and /mcp on --serve.

The owner changes also keep media, provider keys, durable tool steps, host hooks, turn leases, skills, prompt caching, routing, and per-turn overrides. Thinking, text, and tool calls keep their block order. Mid-conversation tool and prompt changes preserve the cached start where the model supports them.

MCP in both directions. The client keeps connectors and browser sign-in. createHarnessMcpServer exposes chat, approvals, agents, and commands to MCP clients. The CLI serves stdio MCP or authenticated HTTP MCP. MCP inputs keep parsed template variables and per-answer MIME types.

createMCPClient keeps toolName, request timeouts, and strict toolFilter checks. Missing or repeated filtered names raise MCPToolFilterError. Tool metadata keeps titles and frozen annotations. Final input validation runs before an accepted call reaches its tool.

Compaction. withCompaction({ countTokens: 'usage' }) uses reported usage. compactNext(threadId) requests compaction at the next model call. Durable compaction writes session-log records through LogRecordsCapability. The host view rebuilds the context from those records.

Compaction reports its reason, summary usage, and errors. Summary options keep turn cuts, updates to previous summaries, and bounded tool output. These opt-in features remain part of the owner scope.

Latest owner update preserved

The eight owner commits through c9f8e5f901f82e5dcba3dbd505d7ee2ab9e110e3 stay in the stack. Shared threads run each chat input and command as the principal returned by authorize. Principal.tenantId keeps credential scopes apart. Reads use the user's key before a shared tenant key. Commands save only through their own sender's credentials. Different senders do not join a running turn by default. Client context stays untrusted; server context wins when objects merge.

The harness run route now supports useChat, transcript hydration, interrupted phases, and cursor-based rejoin after reload. Repeated run IDs keep their receipt. A conflicting sender or already-used run record returns a conflict. Invalid resolve input leaves the interrupt pending. A resolve sent as a turn finishes can start immediately after it.

Sign-in wait uses credentials.require(id, { wait: true }) inside a chat tool. The stopped sender's sign-in resumes that turn. lifetime: 'turn' names a plugin that lives for one turn; 'run' remains an alias. TurnInfo carries message, input context, principal, input ID, and overrides. A plugin may supply each turn's adapter, so HarnessConfig.adapter is optional in that case. harnessText accepts explicit inputModalities.

The view separates approvals, client tools, and sign-in requests. ClientToolCall, view.on('clientTool', ...), and its resolve/fail actions answer client tools. The CLI accepts a caller-owned host and principal, keeps that host open, and keeps its bearer authorization gate. fakeText keeps parent run identity, distinct tool IDs between adapter instances, and exact optional typing. All eight new owner changesets, shared-thread docs, sandbox-provider docs, and their navigation entries remain.

MCP sender isolation. The new per-sender connection cache now encodes tenant and user as a JSON tuple. Names that contain : cannot reuse another sender's authenticated client or tools. A new real-SDK/HTTP regression covers two names that collided before, distinct scoped credentials, separate tools, and same-sender reuse. This new regression was source-reviewed and was not run locally under the user's push instruction. The owner changeset already includes the @tanstack/ai-mcp patch.

Other owner fixes retained. The stack keeps MCP interrupt-kind handling, the Linear OAuth issuer, unique exported tool names, v2 connector credentials, background-agent lease recovery, routed handoff hooks and child resume, interrupted-session recovery, and MCP sessions needed for elicitation. Workspace approval, lazy code-mode tools, skills, Cloudflare example, model catalog, and live plugin commands remain part of the owner scope.

Usage, resumable agents, thread settings, fork, and reset

These five owner features come from the pi-durable comparison. Each one is opt-in or read-only for existing callers.

  • Usage totals. session.usage() and snapshot().usage total each model call by model and by sender, with the cost that providers report. A durable host keeps one harness.usage record per call. A harness.usage event fires on each call.
  • Resumable background agents. agents.start(agent, input, { resume: true }) needs a durable host. The host that takes over runs the agent again from its saved transcript. ctx.step.do(name, fn) keeps side effects from running twice. durability.maxAttempts caps the runs.
  • Thread settings. defineHarness({ models }) and session.configure() store the model, reasoning, instructions, tools, plugins, and working folder of a thread. A client can change only the fields in expose.settings, and none by default.
  • Fork and reset. host.fork() copies a thread up to a message into a new thread, with its settings and media. session.reset(note?) starts a fresh model context and keeps the full transcript, with a marker.
  • Fixes. A resolve that a crash interrupted continues with its routed plan, sender, and context. On a host without a log, the transcript is saved before each model call, so a tool result stays when the next model call fails. The used-up interrupt rule from e677b7516 stays as it is.

Replay and adapter parity

Handoff item Final source behavior
1. Foreign-model history Record the provider, API, and requested model. Foreign replay removes incompatible signatures and redacted thinking. Readable thinking becomes text, and tool IDs stay paired.
2. Replay cleanup Omit failed or aborted assistant batches and their results from provider input. Fill unanswered calls with No result provided errors. Keep system changes after tool results.
3. Claude thinking on Bedrock Keep reasoning text, signatures, redacted bytes, indexes, and block order. Replay signed reasoning before the matching tool call.
4. Unknown finish reasons Emit RUN_ERROR with the provider finish reason. Known content_filter behavior stays the same.
5. Images in tool results Chat Completions sends tool text followed by user images. Mistral keeps images in the tool message. Bedrock uses native tool-result image blocks.
6. Bedrock error status An error property sets status: 'error', including an empty error string.
7. Lone Unicode surrogates Remove lone surrogates from outgoing text and decoded JSON strings. Keep valid pairs, saved raw arguments, signatures, URLs, and bytes.
8. Final argument validation Validate after middleware and before execution or client dispatch. Keep Standard Schema transforms and full JSON Schema checks with coercion. Preserve own JSON keys and ignore inherited lookup entries across all ten runtime paths.
9. Response identity Expose genuine response IDs and response models where the provider supplies them. Keep the requested source model separate. Do not invent a generation ID.
10. Chat Completions content Null or missing text adds no text and still permits tool calls. Object or array text emits the specified error. Keep main's malformed-argument recovery and stream drain.
11. Anthropic Bearer and OAuth Support Bearer credentials, environment precedence, and OAuth identity and betas through the existing SDK. Bearer auth omits x-api-key.
12. Azure OpenAI Add Azure Responses support through the existing OpenAI SDK. Keep endpoint, version, credentials, and deployment-name precedence explicit. Deployment maps use own entries only.
13. Gateway thinking allowEmptySignature permits unsigned same-source thinking on configured gateways. The default remains false.
14. Empty tools with tool history Chat Completions sends tools: [] for tool history without active tools. Wrappers retain that empty list. The no-history request stays unchanged.

Final validation. Standard Schema runs first on the exact input. When a safe input schema permits a coercion retry, the authored schema checks the retry result. A successful transform keeps its value. Raw JSON Schema uses full schema checks. Pending or denied calls do not reach tool hooks, execution, or client dispatch.

The only new runtime dependency for this replay work is approved typebox at ^1.3.27, locked to 1.3.34. TypeBox uses its check engine when dynamic evaluation is unavailable. Source review confirms this fallback. Focused controls check validation behavior. They do not prove a deployed Worker.

Own-key preservation covers schema copies, nullable widening maps, Standard Schema root copies, Gemini parameter sanitation, and Azure deployment lookup. OpenRouter keeps raw arguments separate from normalized input across all three terminal paths. Scalar, array, null, and missing input keep their contracts.

The host selector still returns a message index or -1. Presence guards make that return type explicit. This repairs type inference; it does not change caller sentinels or claim a runtime undefined defect.

Integration. The candidate preserves the earlier 16 owner commits and the eight newer owner commits through c9f8e5f901f82e5dcba3dbd505d7ee2ab9e110e3, then retains the approved merge with pinned main 6d8e6485f92c98a6f2200471e2aff21ab50be013. Main's cancellation and malformed-JSON controls remain. The route registry includes the owner and main routes. Model metadata keeps main's catalog and the owner's modality and reasoning policy.

Main merge and data-loss fixes

fdaeabe8f merges main at 7fb4a5f7f. That brings in #1620, live durable streams and one hydrate GET in Strict Mode. These commits fix the data-loss issues and the CI failures that showed up after the merge:

Commit Fix
581efd816 A host record from the last model call of a turn stays in the log. Before, rebase dropped it when the turn ended.
af7274d6f An empty summary fails the compaction. Before, it replaced the history with nothing. On the overflow path the turn fails with the overflow error.
4c12b1e89, e677b7516 A queued resume turn keeps its parent run. An interrupt is not offered again after its tool ran, including when a later model call fails.
bf329d5e3, 809259930 A tool with no input or null arguments runs with {} again (#265). The final input check made this fail. Other scalars and malformed JSON are still rejected. 809259930 updates the five adapter replay tests that expected the rejection.
c1e9fdad8 The OTel middleware records no execute_tool span for a client tool. The server does not run it.
15868e693 Tests for two compaction functions that had none: part text in the summary prompt, and summary usage after the turn. Coverage saw ai-compaction functions drop from 100% to 97.43%.

The test commits 71d5e933c, f02f0ddcd, and f07b75e0a keep the MCP and E2E tests true under the new behavior. Five patch changesets cover the five fixes above. The tools-test adapter now counts only the tool results of the current turn. Replay gives an unanswered call from an earlier turn a No result provided result, so the old count ended the run after Stop too early.

Docs and changesets

The replay work updates 13 existing pages and docs/config.json. It keeps the owner's harness, MCP, compaction, migration, and example docs. Examples cover server and client code where required. OpenAI text examples use gpt-5.5 and do not use type-assertion casts.

The tool-approval page explains same-run legacy support and the trusted context required for cross-run resume. The adapter-switching page explains the separate card view and conservative host preservation. Only those content edits move their updatedAt dates to October 5. Other content dates and all addedAt dates stay unchanged.

The OpenAI page now describes API-specific modelOptions and reasoning support accurately. That factual correction does not change its date.

The replay changeset covers 14 runtime packages: three minor releases and eleven patch releases. It includes the @tanstack/ai-utils patch for the actual own-key fix. The owner phase, media, reasoning, compaction, MCP, and latest shared-thread/auth/run/CLI/view changesets remain. The owner input-sender changeset also releases the MCP cache-key correction.

This batch adds docs/harness/usage.md, fork-and-reset.md, and thread-settings.md, and a resumable-agents section in subagents.md. It links them from five neighbor pages. Five new changesets cover usage, resumable agents (with a @tanstack/ai minor for ctx.step), thread settings and fork, reset, and the recovered-resolve fix.

✅ Checklist

  • I have followed the steps in the Contributing guide.
  • I have tested code changes locally with pnpm run test:pr, or these tests do not apply to this pull request.
  • I fully understand the code in this pull request, including any code generated with AI assistance.
  • Docs: I updated docs/ for this change, or this change is not user-facing.
  • Changeset: I added a changeset (pnpm changeset), or this PR does not change a published package.

Docs, changesets, and final source review are complete. The full local-test checkbox stays unchecked because the canonical command failed. The user explicitly authorized push despite the outstanding checks.

🚀 Release Impact

  • This change affects published code, and I have generated a changeset.
  • This change is docs/CI/dev-only (no release).

Published runtime packages change. Reviewed changesets are present. Branch CI remains pending for the pushed SHA.

Testing

Commands run. The frozen pnpm@11.9.0 install passed after owner/main integration. Focused regressions reached production code and failed before the fixes. The final focused run passed 32 selected cases across eight native commands: 22 JSON cases, five OpenRouter compatibility controls, and five host cases.

The first eight-package type check failed only in persistence. Indexed selector reads inferred number | undefined. The selector repair preserves its existing index-or--1 contract. No caller sentinel changed.

Current focused evidence Result
Persistence's full host-card and reconstruction files 30 of 30 passed. This includes the positive no-host transfer control.
Strengthened OpenRouter replay controls 12 selected cases passed. Exact serialized nested-null checks cover each required terminal path.
pnpm exec nx run-many --target=test:types --projects=@tanstack/ai-persistence,@tanstack/ai-openrouter --parallel=1 --outputStyle=stream --verbose Native exit 0 after the selector repair. Seven packages passed in the earlier eight-package run.

The follow-up cases overlap the earlier 32 selected cases. Their counts are not additive. Focused results do not replace the full repository gate.

Required final evidence State
Nine existing paths reproduced on clean pinned main, with ordinary-input controls Complete: 25 intended failures and nine ordinary passing controls on pinned main. All protected runtime bytes stayed unchanged. The first Base fixture error is excluded; its corrected run supplies valid proof.
Same-bytes original repro on the integrated candidate Complete: exact5525 bytes passed3/3, native exit0; archived exact hash and feature source removed.
Exact fresh-main SDK/HTTP fixture bytes on the feature candidate Complete: Base2 + Mistral2 + OpenRouter12 passed. Hash-guard removal verified for all three temporary copies.
Full literal pnpm test:pr Known failures; manually stopped with actual root exit1 and captured-only closure verified. Includes live provider tests. No full canonical green or unfinished-target pass is claimed.
Focused local normal checks through existing scripts/Nx Cancelled by explicit user push override. Source-reviewed command packet was not run. Two lint repairs were not retested; Solid config is unchanged.
Fresh final test:docs and test:kiira output Further local checks cancelled by explicit user push override; branch CI pending.
E2E branch CI and the dedicated replay/validation browser sources Local E2E waived at the user's request. Review branch CI for the final pushed head; no current browser pass is claimed.
Final combined owner-c9/MCP regression and full harness checks No further local run at the user's request. Branch CI pending for the final pushed SHA.

Usage, settings, fork, reset, and resumable agents. These checks ran on the merged batch, before the rebase onto e677b7516:

Check Result
@tanstack/ai: build, vitest run, test:types, test:oxlint, test:build 2404 of 2404 tests passed. Every other step exited 0.
@tanstack/ai-harness vitest run (one worker) 830 of 831 passed. The one failure was a build-stub timeout under load, and that file passed alone (8 of 8).
thread-settings.test.ts after the expose.settings gate 15 of 15 passed. tsc --noEmit for the harness exited 0.
kiira check on the changed docs pages No errors.

After the rebase onto e677b7516, no local check ran, at the user's request. The E2E specs are written, and branch CI runs them.

Main merge and data-loss fixes. These checks ran on f02f0ddcd, after the rebase onto 6f3ea4ad6:

Check Result
test:lib for ai, ai-harness, ai-compaction, ai-mcp 2409, 835, 105, and 413 tests passed. No failures.
test:types for those four packages and testing/e2e All exited 0.
test:oxlint for those four packages Exit 0. Only the warnings that were already there.
testing/e2e test:types and oxfmt on f07b75e0a Exit 0.

On f07b75e0a, branch CI passed E2E, Bun, and Preview. Test and Coverage failed in five adapter suites that still expected '' and null to be rejected. 809259930 fixes them. After that fix, test:lib, test:types, and test:oxlint passed locally for openai-base, ai-ollama, ai-gemini, ai-mistral, and ai-openrouter (327, 70, 416, 102, and 303 tests).

On 809259930, Test passed and Coverage failed: ai-compaction functions dropped from 100% to 97.43% (76 of 78). 15868e693 adds a test for each of the 2 functions. Each test fails when its function is broken. ai-compaction passes 107 of 107 tests, with types, lint, and format clean.

Local E2E did not run, at the user's request. Branch CI runs it.

Manual test.

  1. Run pnpm --filter @tanstack/ai-persistence run test:lib --run --maxWorkers=1 tests/subagent-persistence-cards.test.ts tests/subagent-reconstruct.test.ts --no-file-parallelism.
  2. Run pnpm --filter @tanstack/ai-openrouter run test:lib --run --maxWorkers=1 tests/replay-parity.test.ts --no-file-parallelism.
  3. Review the final branch's E2E CI. Local E2E is not run at the user's request.
  4. Read the adapter-switching and tool-approval docs for saved-history and approval compatibility.
  5. Run pnpm --filter @tanstack/ai-harness exec vitest run tests/usage.test.ts tests/agent-resume.test.ts tests/thread-settings.test.ts tests/fork.test.ts tests/reset.test.ts. They cover the five new features.

How this PR makes testing easy. Public tool-boundary tests cover final validation and approval. Host tests cover ordinary text, ambiguous markers, metadata, and occupied IDs. Adapter suites cover request and stream boundaries. Browser fixtures cover saved harness replay, client phases, streamed errors, and local cancellation. The owner's durable compaction and MCP option fixtures remain registered. The new batch adds usage.test.ts, agent-resume.test.ts, thread-settings.test.ts, fork.test.ts, and reset.test.ts in ai-harness, agent-step.test.ts in ai, and the E2E spec harness-thread-controls.spec.ts.

The full canonical run was manually stopped after known failures. Actual root exit1 and captured-only closure are verified. Failures were five lint findings, a Solid test collection error, and a baseline SBX pre-kill detector failure. Their current failure evidence stays separate from the parity controls. The user explicitly authorizes commit/push despite outstanding local checks. All further local checks and the Solid config repair are cancelled. The two lint repairs were source-reviewed but not rerun. The eight unchanged live Docker/SBX files and all browser E2E checks are left for branch CI. Its results remain pending until the new pushed head exists.

Limits. Provider HTTP mocks prove request and response handling. They do not prove live-provider credentials or service behavior. Browser fixtures define the UI checks for branch CI. No current browser pass is claimed. Source review of TypeBox's no-eval route does not prove a Worker deployment. Local verification uses the installed native Node 24.3 executable, not the repository's Node 24.8 pin. Coverage remains CI-only.

Risk / rollback

Validation can reject input that previously reached a tool. Edited approval input is checked again before dispatch. Missing trusted context rejects cross-run legacy phases and keeps the client phase pending.

Replay changes provider input and preserves the saved transcript. Genuine response identity depends on what the provider reports. Adapters without a generation ID omit it.

Owner limitations remain visible. The stdio MCP loop lacks an automated end-to-end test. The owner report names three compaction limits. This PR now fixes the empty summary and the late host record. Reporting before the durable write is still open.

The new batch is opt-in, with two exceptions. First, a host without a log now saves the transcript before each model call, which is one more write per call. Second, snapshot() has a new usage field. A client configure input is refused unless expose.settings lists its fields.

Revert the final merge commit to undo the combined change. Opt-in host, gateway, and compaction settings can also be disabled.

Public API change

The owner stack adds the harness, CLI, dashboard, model catalog, and MCP server interfaces described above. Existing adapter factories remain usable. The replay work adds these approved caller surfaces:

Surface Added contract
Source and response identity Exported MessageSource has provider, api, and model. Optional message and saved-run fields carry source, response identity, and failed or aborted status. TextAdapter exposes optional readonly api and provider. RunFinishedEvent adds optional responseId. StructuredOutputResult adds optional responseId and model.
Anthropic configuration AnthropicClientConfig extends the SDK options with optional apiKey; the SDK supplies authToken. AnthropicTextConfig adds oauth, allowEmptySignature, and provider. Existing factory and injected-client forms remain.
Azure Responses Root exports add azureOpenaiText, AzureOpenAITextAdapter, and AzureOpenAITextConfig. Configuration covers apiKey, baseURL, resourceName, apiVersion, deploymentName, and deploymentNameMap.
Usage totals session.usage(), snapshot().usage (SessionUsage, UsageCounts), and the harness.usage event.
Resumable agents AgentStartOptions.resume, ctx.step.do(name, fn) in agent code, and optional SubagentBinding.step in @tanstack/ai.
Thread settings defineHarness({ models, expose: { settings } }), session.configure(), session.settings(), client.configure(), and the configure input op.
Fork and reset host.fork(harness, { threadId, newThreadId, at?, principal? }), session.reset(note?), client.reset(), and the reset input op.
Adapter utilities @tanstack/ai/adapter-internals exports ReplayMessages, ReplayToolIdRule, transformMessagesForReplay, hashToolCallId, sanitizeUnicode, and sanitizeJsonArguments. These support adapter conversion; they add no application callback.

Before

import { chat } from '@tanstack/ai'
import { openaiText } from '@tanstack/ai-openai'

const stream = chat({
  adapter: openaiText('gpt-5.5'),
  messages: [{ role: 'user', content: 'Hello!' }],
})

for await (const chunk of stream) {
  if (chunk.type === 'TEXT_MESSAGE_CONTENT') console.log(chunk.delta)
}

After: Azure deployment

import { chat } from '@tanstack/ai'
import { azureOpenaiText } from '@tanstack/ai-openai'

const stream = chat({
  adapter: azureOpenaiText('gpt-5.5', {
    resourceName: 'my-resource',
    apiKey: process.env.AZURE_OPENAI_API_KEY,
    apiVersion: 'v1',
    deploymentNameMap: { 'gpt-5.5': 'production-chat' },
  }),
  messages: [{ role: 'user', content: 'Hello!' }],
})

for await (const chunk of stream) {
  if (chunk.type === 'TEXT_MESSAGE_CONTENT') console.log(chunk.delta)
}

Other caller examples remain in the adapter, tool, harness, and migration docs. No migration API or private-key editing procedure is added.

The latest owner P0–P14, shared-thread, MCP, compaction, auth, CLI, and run-route scope stays in this description. Historical owner checks are not presented as tests of this final merged candidate.

Pushed head: 15868e693. It adds six fix commits on top of 6f3ea4ad6 with plain pushes, no force. main is still 7fb4a5f7f, which is already merged. All checks pass on this head: Test, E2E Tests, Bun Tests, Coverage, and Preview. pkg.pr.new: https://pkg.pr.new/TanStack/ai/@tanstack/ai@15868e6.

…de Code, Codex, and other coding agents; show child agent work in the CLI, ACP, and dashboard
…ry run; use an if-chain for child events

The opencode durability tests reach the failed-start path only sometimes, so the
coverage number moved between runs. A new test starts a session against a closed
port, with and without a session to resume.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
With busy: 'reject', ctx.session.prompt() from a running turn was rejected
without a sign, so the goal loop stopped after one round.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…pauses or stops

A message typed during the last model call of a goal turn runs first, as a
steer that came too late, and pauses the goal. The goal's next turn was already
queued and still ran once. ctx.session.prompt() now returns the turn, and the
goal plugin cancels its queued turn when a user message pauses the goal, on
/goal stop, and when a new goal replaces it.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
… clock

The timeout tests use a 1 ms budget, so the branch they end in depends on how
many milliseconds pass. The coverage number moved between runs. New tests fix
the clock and reach each of the five timeout exits every run.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…session for any UI

Adds @tanstack/store ^0.11.1 and the browser-safe ./view export. Three examples
listed @tanstack/store ^0.8.0 without importing it; they move to ^0.11.1 so the
workspace keeps one version (sherif).

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…any TUI on a session view

The CLI no longer depends on ink or react. An interactive terminal uses line
mode unless runCli gets a ui function, which receives a ready session view and
resolves when the user quits.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…ove the old event fold

In a terminal, line mode opens sign-in links in the browser.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The coverage gate failed because the new session view code had paths
that no test ran. These tests run them.

- view-paths.test.ts: a fake SessionViewSource forces each path of
  createSessionView. It covers failed reads at start, a failed or ended
  event stream, connection state, failed actions, approvals, on()
  handlers, the snapshot read coalesce, and dispose.
- view-reduce.test.ts: the reducer branches that return the same
  state, nested and sibling child agents, content-part tool results,
  run errors, and transcript edge cases.
- client-reads.test.ts: the transcript and describe routes answer 400
  and 403, and a failed read throws with its route and status.
- goal.test.ts: selectGoal gives null for state that is not a goal.
- view-fixtures.ts: event and snapshot helpers that the view tests share.

No production code changed.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Every direct TanStack Store dependency in the workspace is now on the latest
release (store, react-store, and solid-store 0.11.1).

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…sign-in wait, and report fixes

Fixes from the Pit spike report on the harness (#1555):

- A resolve that fails validation keeps the interrupt pending. A resolve sent on RUN_FINISHED is accepted. A busy: 'steer' prompt receipt names the running turn.
- POST run runs the turn as the request runId, so ChatClient and useChat can answer approvals and client tools. GET run?threadId= serves the hydrate data with the running turn, and GET run?runId= joins a running turn after a reload, with Last-Event-ID resume and canAccess.
- ask() questions arrive on the turn stream. snapshot().activeOperations has startedCursor.
- Each input, command, and wake turn runs as its sender: log record, credentials, provider keys, canJoin (principal and turnPrincipal), router. By default a steer joins only a turn of the same sender.
- Inputs take context. POST run fills it from forwardedProps. It is stored, kept after a restart and an interrupt, and merged under HarnessConfig.context.
- TurnInfo for plugin adapter() and ctx.turn. lifetime: 'turn'. defineHarness adapter is optional. harnessText takes inputModalities.
- Principal.tenantId goes into the credential scope. A read falls back from the user's credential to the tenant's.
- credentials.require(id, { wait: true }) stops the turn for a sign-in, and a save by the same sender continues it.
- The session view lists client tools in clientTools and sign-in waits in signIns.
- A client runId that a run record has, or that another sender used, is refused with conflict.
mcpConnector keeps its sign-in state, client, and tools per principal, so one person's MCP sign-in is not used for another person's turn in a shared thread.
With host, the CLI uses it and does not close it. principal reaches every session the CLI opens and the --serve handler.
…ient

POST run with the client runId and a resume, GET run hydrate, a resolve on RUN_FINISHED, forwardedProps as context, two senders on one thread, a ChatClient approval, and a reloaded ChatClient that joins a running turn.
…context, and sign-in wait

- New: harness/shared-threads and sandbox/build-a-provider.
- Updated: connect, custom-ui, inputs, plugins, auth, cli, turn-control, overview, and the sandbox provider pages.
@AlemTuzlak AlemTuzlak changed the title feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, turn overrides, block order, mid-conversation changes, routing, compaction, MCP tool options, and loop fixes (combined stack) feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, turn overrides, block order, mid-conversation changes, routing, compaction, MCP tool options, shared threads, and loop fixes (combined stack) Oct 5, 2026
@AlemTuzlak AlemTuzlak changed the title feat(ai, ai-harness): TanStack AI harness, phases 0 to 14, media, provider keys, durable sessions, host hooks, reasoning, skills, prompt caching, turn overrides, block order, mid-conversation changes, routing, compaction, MCP tool options, shared threads, and loop fixes (combined stack) feat(ai, ai-harness): add the harness stack with replay and adapter parity Oct 5, 2026
AlemTuzlak and others added 7 commits October 5, 2026 15:10
A `resume` input on an interrupted session queues a turn with
`parentRunId`, but `QueuedTurn` had no such field and the runner read
the parent only from `answers`. The build failed on the type, and
chat() rejected the resume: "Interrupt continuation requires
parentRunId to identify the interrupted run".

Declare the field and fall back to it in the runner.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
A summarizer that returned an empty summary (the model stopped with no
text) made summarizeOldest replace the history with an empty summary,
so the history was lost.

An empty or whitespace-only summary now fails the compaction. The
history stays, the compaction events carry the error, and no durable
record is written. Covered for the threshold, compactNext, durable and
after-turn paths, and end to end on the overflow retry path.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
A host record appended during the last model call of a turn with a tool
phase was lost at the end of the turn. withPersistence tags the run on
older messages, so the engine save is not an extension of what the
engine held, and the rebase returned only the engine list. The transcript
commit then removed the message the record had folded in.

The engine list still wins for older messages, and host messages that
the fold added after the engine's last save now stay after it.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
tools() walks tools/list two times on a spec 2025 server: a raw walk,
then listTools() for the SDK's output checks. The new test for colliding
sender names expected one walk per sender, so it failed in CI. The
isolation itself holds: every request of a sender uses that sender's
token, and each sender initializes once.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The replay-parity merge rebuilt pending interrupts from the interrupt
store. withPersistence commits a resume only at a success boundary, so
when a resumed turn failed, stopped, or hit its time limit after the
approved tool ran, the approval came back and a second answer ran the
tool again.

A resume whose tools got a real result is now used up: the harness marks
the interrupts answered in the store, and the log scan of a durable host
skips an interrupt whose tool call has a later result. A cancelled call,
or the error result a stopped run gives a call it did not run, does not
count, so a resume stopped before its tool still stays open for a retry.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
@github-actions github-actions Bot removed the merge-conflicts Conflicts with the base branch — needs a rebase label Oct 5, 2026
Agent code gets ctx.step.do(name, fn). SubagentBinding takes an optional step, so a host can keep each step result and skip it when the agent runs again. Without a binding, step.do runs fn each time.
…rk, and reset

- session.usage() and snapshot().usage total each model call per model and per sender, with provider cost. A durable host keeps a harness.usage record per call. A harness.usage event fires on each call.
- agents.start(agent, input, { resume: true }) on a durable host: the host that takes over runs the agent again, its model calls continue from the saved transcript, and ctx.step.do results are not run again. durability.maxAttempts caps the runs.
- defineHarness({ models }) and session.configure() store the model, reasoning, instructions, tools, plugins, and working folder of a thread. A client can change only the fields in expose.settings.
- host.fork() copies a thread up to a message into a new thread, with its settings and media.
- session.reset(note?) starts a fresh model context and keeps the transcript, with a marker.
- A resolve that a crash interrupted continues the interrupted turn with its routed plan, sender, and context. On a host without a log, the transcript is saved before each model call, so a tool result stays when the next model call fails.
The final tool input check kept a missing or literal null argument and
checked it against the schema, so a tool with no required fields did not
run when the model sent an empty tool_use block. That broke the fix for
issue #265, and its E2E test failed.

No input is {} again. A null that the schema rejects is checked as {}.
Other values, such as 2 or [1,2], must still fit the schema.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The final tool input check now runs onBeforeToolCall before a client
tool is dispatched, so otelMiddleware opened an execute_tool span for a
tool the server does not run. The client-tool wait E2E test then saw two
root spans.

otelMiddleware skips the span for a known tool with no execute. The
unit test is back to main's chat and iteration spans.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
The client-tool input error scenarios sent message: 42. The final input
check now coerces scalar types, so 42 became "42" and passed, and the
client ran the tool. The scenarios now omit the required message, which
is invalid with or without coercion.

The sandbox persistence spec now expects the runId and source that the
chat run records on the saved assistant message.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
Replay now gives an unanswered tool call from an earlier turn a
"No result provided" result. The provider-free adapter counted every
tool message, so after Stop the next user turn answered with text and
ended before the race test could press Stop again.
… replay tests

The #265 fix runs a tool with no input or a literal null as {}. Five
adapter replay-parity tests still expected those inputs to be rejected.
They now check that the tool runs with {}, or that the call goes to the
client. Other scalars and malformed JSON are still rejected.

Mistral ends whitespace-only arguments with a parse error before the
input check, so that case still runs 0 times.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
…turn summary usage

Two new functions had no test, so the Coverage job saw function coverage
for ai-compaction drop from 100% to 97.43%:

- contentText in conversationSummarizer, for content that is a list of
  parts.
- The addUsage callback of the after-turn check, for a summarizer that
  reports usage.

Each new test fails when its function is broken.

Claude-Session: https://claude.ai/code/session_01APYv1qshKyjPPpkFyRZhfZ
@github-actions github-actions Bot added waiting-on: maintainer The ball is in the maintainers’ court and removed waiting-on: author Waiting for the author to respond or update labels Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on: maintainer The ball is in the maintainers’ court

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants