Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
50 commits
Select commit Hold shift + click to select a range
5a5c1ef
compaction: port compaction-budget-recovery onto dev
ZeroPoint95 Sep 28, 2026
447ae7a
compaction: the ported budget code passes dev's laws
ZeroPoint95 Sep 28, 2026
4ee8195
compaction: a pass the free rungs cannot finish ends with a summary
ZeroPoint95 Sep 28, 2026
79ec923
compaction: a summary keeps the right messages and recovers a near miss
ZeroPoint95 Sep 28, 2026
faede1b
provider: a model that takes no tools is sent none, and told once why
ZeroPoint95 Sep 28, 2026
de388c6
compaction: a /compact that changes nothing says why
ZeroPoint95 Sep 28, 2026
9bb8d04
session: a model with no tools reads a chat page, and a tools refusal…
ZeroPoint95 Sep 28, 2026
8d5b95f
tui3: the compaction mark is one quiet line, and it outlives the turn
ZeroPoint95 Sep 28, 2026
7a80da6
provider: a thinking budget bends to the window, and tool-less endpoi…
ZeroPoint95 Sep 28, 2026
f72296a
provider: a live check that deepseek-v3.2's refused first message now…
ZeroPoint95 Sep 28, 2026
51f44cc
remote: /compact waits as long as a summary takes, and a send is not …
ZeroPoint95 Sep 28, 2026
6f2dc94
tui3: workfold's imports sit in their groups after the rebase
ZeroPoint95 Sep 28, 2026
43d4527
tui3: a compaction after the answer stays visible under dev's bookkee…
ZeroPoint95 Sep 28, 2026
4b960ee
docs: the change entry for #1658
ZeroPoint95 Sep 28, 2026
17ecc86
tests: the summary's low effort is its own economy, and a finished co…
ZeroPoint95 Sep 28, 2026
4b9ad8a
session: the summary test's count no longer depends on the temp folde…
ZeroPoint95 Sep 28, 2026
ab28f2a
compaction: one /compact goes all the way, and the meter follows it
ZeroPoint95 Sep 28, 2026
575928b
provider: a tool-less endpoint's size refusal no longer caps a tool r…
AbirAbbas Sep 28, 2026
c071f70
provider: a model with no tools is told once, and only an answered re…
AbirAbbas Sep 28, 2026
dccb03f
provider: an ordinary request sends no max_tokens, as it did on dev
AbirAbbas Sep 28, 2026
8101b4a
session: a message sent during a slow /compact waits for it instead o…
AbirAbbas Sep 28, 2026
afd9b1a
session: an exchange made while a summary is written stays in the his…
AbirAbbas Sep 28, 2026
ff2d977
session: the automatic pass does not buy a summary it cannot use
AbirAbbas Sep 28, 2026
99a89af
session: a pass whose summary did not land says so
AbirAbbas Sep 28, 2026
8367078
session: a rerouted summary is booked under the model that answered
AbirAbbas Sep 28, 2026
f8a0823
session: a summarizer's refusal is not spliced in as the summary
AbirAbbas Sep 28, 2026
230d6d0
tui3: a late /compact reply stays with the conversation that asked fo…
AbirAbbas Sep 28, 2026
352a747
tui3: the compaction mark comes from the vocabulary, and a note never…
AbirAbbas Sep 28, 2026
f13b389
docs: the manual says what compaction keeps, paraphrases and waits for
AbirAbbas Sep 28, 2026
90c68cc
docs: the manual says what the router, the output cap and a skipped s…
AbirAbbas Sep 28, 2026
1ee7b96
provider: a model with no tools reads earlier tool work as text
AbirAbbas Sep 28, 2026
838e11e
session: when no cut can reach the line, /compact and the automatic p…
AbirAbbas Sep 28, 2026
e06e524
session: the meter after a compaction counts the tool definitions the…
AbirAbbas Sep 28, 2026
9b78850
provider: DeepInfra's "exceeds maximum input length" is a size refusal
AbirAbbas Sep 28, 2026
fcfcd70
docs: the manual says what a switch to a no-tools model, an unreachab…
AbirAbbas Sep 28, 2026
5fade41
provider: a model's no-tools line is spent only on a call a person sees
AbirAbbas Sep 28, 2026
08ad435
provider: a pinned endpoint's size refusal is learned once, and every…
AbirAbbas Sep 28, 2026
be790de
session: a /compact whose summary did not land returns why, across th…
AbirAbbas Sep 28, 2026
9b5f676
session: a summary refusal cannot borrow the chunk's own labels, and …
AbirAbbas Sep 28, 2026
a2e8ef4
tui3: a /compact reply survives a page visit and says when its summar…
AbirAbbas Sep 28, 2026
a489229
docs: the manual names a partial /compact and a pinned endpoint's window
AbirAbbas Sep 28, 2026
f0eb933
session: only a summary that was asked for and did not land is report…
AbirAbbas Sep 28, 2026
fc8c035
session: the meter counts the next request's tool definitions, and th…
AbirAbbas Sep 28, 2026
f06605e
tui3: the ASCII compaction line keeps its words in any script
AbirAbbas Sep 28, 2026
cffac3d
provider: a clipped tool argument stops at a character boundary under…
AbirAbbas Sep 28, 2026
fe68aca
docs: /compact keeps fewer than three when that is the cut that reach…
AbirAbbas Sep 28, 2026
c1f3fb9
provider: a thinking pass is sized for the lowest level the model can…
AbirAbbas Sep 29, 2026
0a9d453
session: a summary leaves room to think, keeps the protected messages…
AbirAbbas Sep 29, 2026
8385e34
docs: what /compact keeps when only two messages remain, and why a th…
AbirAbbas Sep 29, 2026
de19770
tui3: only a compaction that wrote a summary stands outside the fold
ZeroPoint95 Sep 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
46 changes: 42 additions & 4 deletions PERF.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,44 @@ two thirds off the embedded corpora. This file is what keeps it. Every win below
is defended by something that goes red locally, in `go test` or in `make check`,
with a message that says what happened.

## Context recovery bounds

Conversation request admission sums the existing encoded messages and tool schemas;
it performs no tokenizer call, network lookup or extra model request. Image payload
bytes are replaced by a token allowance. The margin is 5% of the effective endpoint
window, bounded to 512–8,192 tokens. An unspecified output allowance is bounded by
one quarter of the window and the configured completion reserve; shrinking it retains
up to 512 tokens as the useful minimum (one eighth for very small windows). A thinking
budget shrinks to the room left beside an answer of up to 1,024 tokens (a sixteenth of
the window) and is dropped below 1,024, so thinking never refuses a request. Endpoints
that take no tools are left out of the window a tool-carrying request is measured
against.

Manual and emergency reductions retain 4,096 recent tokens, capped to an eighth of
the window, and always retain the latest assistant/tool batch. Recovery is bounded
to two changed-request attempts per failed generation and resets after a successful
response. Tests assert request counts, fitting budgets, tool pairing and actions that
execute exactly once; they do not wait on real clocks.

Automatic conversation profiles use the existing 32,000-token threshold at request
boundaries. Only crossing the threshold rebuilds the default prompt and belt; unchanged
profiles do no schema work. Explicitly loaded capabilities survive the rebuild. Automatic
no-op compaction emits no seam events, while manual commands retain their no-op feedback.

A compaction pass makes **one model call only when its free rungs fail**: a summary
(`internal/session/compact_summary.go`) is written when stubbing and folding leave the
transcript above the pass's line, and never when the tool definitions alone exceed that
line. An automatic or manual pass needs at least **1,024 tokens** of region the previous
summary has not read; a refused request's recovery needs **128**. It keeps the **three**
most recent person messages and everything after them when that still reaches the line,
then two, then the latest with the reply before it, then the latest alone. A summary is
asked for at most half its region and at most a twentieth of the window (**128–4,096**
tokens). A refusal that states its figures reclaims what is missing plus a thirty-second
of the window; one that does not reclaims a quarter of the transcript. A region larger than one request is summarized in chunks sized to the window,
each bounded to **two minutes**; the session lock is released during every call. The
render the summarizer reads caps a tool result at 2,000 bytes and a call's arguments at
400. Tests use scripted completers and assert request counts and sizes, not clocks.

## Connection recovery bounds

`internal/provider/connectivity.go` limits a connection-recovery episode to
Expand Down Expand Up @@ -1537,10 +1575,10 @@ otherwise getting 57–61% of its prompt back from.
So the ceiling is a MEASUREMENT now and not a constant: the narrowest prompt this
model has actually been refused for being too long, learned from the overflow
refusal itself and remembered across processes
(`internal/provider`'s `NoteServedWindow` / `ServedWindow`, applied by
`session.TrustedWindowFor`). A model nobody has refused is believed; one that has
refused is capped at what it refused, for good. The 386k incident now costs one
turn per model per machine instead of every model for ever.
(`internal/provider`'s `NoteServedWindow` / `ServedWindow`, applied while the provider
sizes each assembled request). `session.TrustedWindowFor` now returns the catalog window
unchanged: a model-only memo cannot say which endpoint can serve a request with tools.
The provider uses the refused endpoint's evidence when that request is sized.

Two guards stand behind that trade and neither is new: `guardOversizeRequest`
still shrinks a transcript that has grown past the trusted window before it goes
Expand Down
28 changes: 27 additions & 1 deletion cmd/codeaf/chatv3.go
Original file line number Diff line number Diff line change
Expand Up @@ -1031,7 +1031,14 @@ func openV3Launch(proc *v3Process, opts v3Options) (*v3Launch, error) {
// --reasoning went to every model blind — and on a router, a knob no
// endpoint publishes is not a 400 but a 404 with no endpoints left to
// serve the request (internal/provider's endpoints.go).
SupportsParameter: activeModels.SupportsParameter,
//
// AND IT IS ASKED ABOUT THE ID THE SERVICE IS SENT. A model behind a
// connected service is named here with the service's written prefix
// (`stub/z-ai/glm-5.3-flash`), and the adapter asks the same catalog
// about the id it puts on the wire (`z-ai/glm-5.3-flash`): answered only
// about the prefixed name, the session kept the working page for a
// model the adapter was already sending no tools (chatpage.go).
SupportsParameter: v3SupportsParameter(proc.Shelf, activeModels),
ReasoningProfile: config.ReasoningProfileSeam(activeModels),
// And the model's own published price, which is what bounds the latency
// ask: this session wants the fastest endpoint, not the dearest one
Expand Down Expand Up @@ -2492,6 +2499,25 @@ func v3AnswersText(outputs []string) bool {
// The file is read at most once per session: it is the same rows for the whole
// warming window, and re-reading it per message would put I/O on the message
// path to learn nothing new.
// v3SupportsParameter is the catalog's answer about a model, asked first as
// named and then as the id its service is sent — the same id the adapter asks
// about, so the session's page and the adapter's body cannot disagree about
// what the model accepts.
func v3SupportsParameter(shelf *v3ModelShelf, models *catalog.Catalog) func(string, string) (bool, bool) {
return func(model, parameter string) (bool, bool) {
if models == nil {
return false, false
}
if supported, known := models.SupportsParameter(model, parameter); known {
return supported, known
}
if bare := shelf.wireModel(model); bare != "" && !strings.EqualFold(bare, model) {
return models.SupportsParameter(bare, parameter)
}
return false, false
}
}

func v3SeesImages(models v3Catalog) func(string) bool {
var once sync.Once
var cached []tui3.Model
Expand Down
16 changes: 16 additions & 0 deletions cmd/codeaf/chatv3_modelshelf.go
Original file line number Diff line number Diff line change
Expand Up @@ -173,6 +173,22 @@ func (s *v3ModelShelf) contextWindow(model string) int {
return v3ContextWindow(s.modelsForService(service), bare)
}

// wireModel is the id a model is sent to its service under: the service's
// written prefix and any thinking level taken off. A shelf that knows no
// services answers the model as named.
func (s *v3ModelShelf) wireModel(model string) string {
model, _ = roles.SplitEffort(strings.TrimSpace(model))
if s == nil {
return model
}
sources := s.sourcesNow()
if sources.Empty() {
return model
}
_, bare := sources.For(model)
return bare
}

// sourcesNow is the service set the shelf was last aligned with.
func (s *v3ModelShelf) sourcesNow() modelsource.Set {
s.mu.RLock()
Expand Down
32 changes: 32 additions & 0 deletions docs/changes/unreleased/1658-compaction-summaries-and-recovery.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
---
kind: changed
title: compaction can end with a summary, refused requests recover, and a summary leaves a visible line
pr: 1658
surface: [chat, engine, remote]
invalidates:
- "Compaction never summarized: it only stubbed old tool results and folded old assistant work. When those are not enough, the conversation's own model now summarizes the oldest part, older user messages included. An automatic or manual pass keeps the three most recent user messages when no cut can reach its line; recovery from a refused request can keep fewer to reclaim room. /compact always goes on to a summary in the same pass when there is enough older material."
- "A request's window was the smallest among all of a model's endpoints. A request that carries tools is now measured only against endpoints that take tools. Default routing sends no require_parameters, so the router can still hand it to one that takes none: a size refusal from such an endpoint is sent once more before anything is shortened, its window is not used for later tool requests unless the request pins that endpoint, and every learned endpoint window lapses after 30 minutes."
- "The lane sheet read tool support from supports_tool_choice.function. It reads supported_parameters, the list the router filters on under require_parameters."
- "The xhigh and max thinking budgets were reserved in full whatever the window. They now shrink to the room the prompt leaves and are dropped below 1,024 tokens."
- "A model that takes no tools was sent the tool list and retried with `Retry 1/1: removed tools`. It is sent none from the start, reads a short chat page, and the person is told once why. A model the catalog does not know is sent no tools for 30 minutes only after a retry without them was answered."
- "/compact waited ten seconds and then said `compact failed: the engine did not answer in time` about a pass that went on to land. A new surface waits up to five minutes (session.CompactPatience), and a message sent meanwhile is not queued behind it. An old surface talking to a new engine still reports that ten-second failure even though the pass can land later: the wire Version stays at 20 because no frame or method changed, and Decision 3 in docs/REMOTE.md reserves a version bump for a protocol change."
- "MethodCompact rode the remote ordered lane. It is its own call class, classWork, on a goroutine of its own."
- "A /compact that changed nothing said only `nothing to compact`. It now says why."
- "After /compact the status line kept the last request's weight until the next message was sent. It drops as soon as the pass lands."
- "The post-compaction meter and the ⚭ line counted only transcript text while the next request still carried tool definitions. Both now include the sent definitions in their estimates."
- "A switch back to a model without tools carried earlier tool-call protocol messages and could be refused. Its request now carries the earlier calls and results as readable text; tool-capable requests keep their original history bytes."
- "DeepInfra's `Requested input length … exceeds maximum input length …` refusal was retried as an ordinary error. It now enters overflow recovery and teaches the endpoint's stated limit."
- "A finished compaction was a full-width rule that folded into `▸ worked` when the turn ended. It is one dim `⚭ compacted` line. A pass that wrote a summary stands outside the fold after the turn, decided by the new session.Event.Summarized count; a pass that only folded or stubbed still folds with the turn's work, as #1627 intended."
- "A /compact pass that folded work while its summary failed said only `compacted`. It now says `summary skipped: <why>` in its note, including over a remote engine; a successful remote reply to an older surface still reads as plain success. A visit to Home during the pass no longer loses that conversation's note or meter update."
- "An unobserved helper call could consume a tool-less model's one visible notice. The notice is now spent only on a call with a stream observer."
- "A strict endpoint pin could resend the same oversized tool request to that endpoint and ignore its learned limit on later tool requests. It now pays one refusal and sizes the next request against the pinned endpoint's limit. A routed resend logs its first 400 as well as the answer. Equal limits refresh their disk date, and future dates are clamped."
- "An ordinary request could send `max_tokens` even when its computed output ceiling did not bind. It now omits that field until the ceiling actually binds."
- "A summary on a model whose lowest listed thinking effort was high could spend its whole output cap thinking and return empty. The effort word still travels as asked, but the provider now sizes the thinking room for the lowest level the model lists (a model that lists only high and xhigh runs at high at the least); an empty length finish gets one retry at twice the answer allowance, and both attempts are counted."
- "When no cut could reach the line and only two person messages remained after a summary, /compact could summarize away one of them to reach the minimum worth a call. Ordinary and manual passes now keep every recent person message the three-message ladder protects; only refusal recovery may keep fewer."
- "A rolling summary could omit specific facts from the previous summary. Its instruction now explicitly carries those facts forward word for word unless newer conversation supersedes them."
---

A conversation could fill its window and then be stuck: the request was refused
as too long and `/compact` found nothing to do. The summary is the rung that was
missing, and the size check now counts only what a request really carries, so a
refusal names a shortfall that compaction can actually close.
9 changes: 9 additions & 0 deletions docs/design/icons/DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,6 +127,15 @@ its verdict.
The dim action mark is distinct from the failed-work mark. It is drawn only where
a click can remove the attachment, and is resolved explicitly through the vocabulary.

### Conversation compaction

| Meaning | Slot | Plain | Nerd font | ASCII |
| --- | --- | --- | --- | --- |
| working context shortened | `GCompacted` | `⚭` | nf-fa-compress | `#` |

The Font Awesome 4 `fa-compress` mark is U+F066. Both the automatic pass line and
the `/compact` reply resolve this slot through the chat surface's glyph door.

### File kinds (chips, and the gutter beside a call that made or opened one)

| Kind | Slot | Plain | Nerd font | ASCII |
Expand Down
1 change: 1 addition & 0 deletions internal/iconlaw/iconlaw_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,7 @@ var ownedRunes = map[rune]string{
'⇉': "tokens.GActionCoordinate",
'◷': "tokens.GActionWait",
'▪': "tokens.GActionWork",
'⚭': "tokens.GCompacted",
'⌕': "tokens.GSearch",
'✎': "tokens.GWrite",
'⌾': "tokens.GFileImage",
Expand Down
12 changes: 12 additions & 0 deletions internal/lane/lanestub/lanestub.go
Original file line number Diff line number Diff line change
Expand Up @@ -675,6 +675,7 @@ type sheetEndpoint struct {
Quantization string `json:"quantization"`
ContextLength int `json:"context_length"`
MaxCompletionTokens int `json:"max_completion_tokens"`
SupportedParameters []string `json:"supported_parameters"`
Pricing sheetPricing `json:"pricing"`
SupportsToolChoice sheetToolChoice `json:"supports_tool_choice"`
Status int `json:"status"`
Expand Down Expand Up @@ -725,6 +726,16 @@ func (s *Server) serveSheet(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, body)
}

// parameters is the list a real router filters on under `require_parameters`,
// and the one the sheet reads tool support from: "tools" is on it exactly
// when the lane says it takes them.
func (l Lane) parameters() []string {
if l.Tools {
return []string{"max_tokens", "reasoning", "tools", "tool_choice"}
}
return []string{"max_tokens", "reasoning"}
}

// row is the sheet's account of one lane. Percentiles a profile did not state
// are derived from what it actually does, with the spread a real lane has —
// which makes "the sheet is right about this lane" the default and leaves
Expand All @@ -749,6 +760,7 @@ func (l Lane) row(model string) sheetEndpoint {
Quantization: l.Quant,
ContextLength: l.Context,
MaxCompletionTokens: l.MaxOut,
SupportedParameters: l.parameters(),
Pricing: sheetPricing{
Prompt: money(l.PriceIn),
Completion: money(l.PriceOut),
Expand Down
23 changes: 21 additions & 2 deletions internal/lane/sheet.go
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import (
"net/url"
"os"
"path/filepath"
"slices"
"sort"
"strconv"
"strings"
Expand Down Expand Up @@ -1171,7 +1172,15 @@ type wireEndpoint struct {
Completion string `json:"completion"`
InputCacheRead string `json:"input_cache_read"`
} `json:"pricing"`
SupportsToolChoice struct {
// SupportedParameters is the list the router filters on when a request
// says `require_parameters`, which every request from this program does —
// so it is the one answer to "will a request carrying tools reach this
// lane". SupportsToolChoice describes which tool_choice VALUES the lane
// takes, a different question: on 2026-09-28 deepseek-v3.2's sheet had
// GMICloud, AtlasCloud and Alibaba taking tools with no forced-function
// choice, and Mara offering the choice while taking no tools at all.
SupportedParameters []string `json:"supported_parameters"`
SupportsToolChoice struct {
Function bool `json:"function"`
} `json:"supports_tool_choice"`
// Status is the router's own health word for the endpoint: zero is healthy,
Expand All @@ -1183,6 +1192,16 @@ type wireEndpoint struct {
ThroughputLast30m wirePercentiles `json:"throughput_last_30m"`
}

// takesTools reads a lane's tool support from the parameter list the router
// itself filters on, and falls back to the tool_choice block only for a sheet
// that publishes no list, which is what an older router or a stub sends.
func (item wireEndpoint) takesTools() bool {
if item.SupportedParameters == nil {
return item.SupportsToolChoice.Function
}
return slices.Contains(item.SupportedParameters, "tools")
}

// decodeSheet reads the endpoints body one row at a time. The error it returns
// is about the envelope; a row it could not read is skipped in silence, which
// is the whole point of decoding this way.
Expand Down Expand Up @@ -1212,7 +1231,7 @@ func decodeSheet(model string, body io.Reader) ([]Row, map[ID]string, error) {
rows = append(rows, Row{
ID: id,
Facts: Facts{
Tools: item.SupportsToolChoice.Function,
Tools: item.takesTools(),
Quant: strings.TrimSpace(item.Quantization),
MaxOut: item.MaxCompletionTokens,
Context: item.ContextLength,
Expand Down
27 changes: 27 additions & 0 deletions internal/lane/sheet_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -465,3 +465,30 @@ func TestTheBeatPrimesTheBeliefFromEveryReading(t *testing.T) {
}
}
}

// TOOL SUPPORT IS READ FROM THE LIST THE ROUTER FILTERS ON. Every request from
// this program says `require_parameters`, so a lane whose supported_parameters
// lacks "tools" never receives a request that carries them, whatever its
// tool_choice block says. The two rows are deepseek-v3.2's GMICloud and Mara as
// the router published them on 2026-09-28, where the two fields disagree in
// both directions; the third row publishes no list and keeps the old reading.
func TestToolSupportIsReadFromTheParameterListTheRouterFiltersOn(t *testing.T) {
body := `{"data":{"endpoints":[
{"provider_name":"GMICloud","context_length":163840,"supported_parameters":["max_tokens","tools","tool_choice"],"supports_tool_choice":{"function":false,"auto":true}},
{"provider_name":"Mara","context_length":32768,"supported_parameters":["max_tokens","temperature"],"supports_tool_choice":{"function":true,"auto":true}},
{"provider_name":"Listless","context_length":65536,"supports_tool_choice":{"function":true}}
]}}`
rows, _, err := decodeSheet("deepseek/deepseek-v3.2", strings.NewReader(body))
if err != nil {
t.Fatal(err)
}
want := map[string]bool{"GMICloud": true, "Mara": false, "Listless": true}
for _, row := range rows {
if row.Facts.Tools != want[row.ID.Lane] {
t.Errorf("%s reads Tools=%v, want %v", row.ID.Lane, row.Facts.Tools, want[row.ID.Lane])
}
}
if len(rows) != len(want) {
t.Fatalf("decoded %d rows, want %d", len(rows), len(want))
}
}
Loading
Loading