Skip to content

tasks: held work, kept branches and spend limits say what is true - #1670

Draft
santoshkumarradha wants to merge 10 commits into
devfrom
fix/task-run-audit
Draft

santoshkumarradha wants to merge 10 commits into
devfrom
fix/task-run-audit

Conversation

@santoshkumarradha

@santoshkumarradha santoshkumarradha commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Re-cut from #1604 (only the task-run and spend-limit fixes; the other #1604 pieces are elsewhere or already on dev).

Not included: #1527, #1549, #1566, #1579 (already on dev), and the part-check path refusal (#1573, taken back in #1604).

Tested revision: b3b5d7d. CI (check, light gate, touched packages, CLA) passed on this head.

BASE=origin/dev make pr-ready on Spark fails on one test, TestAPreExistingRedDoesNotStandBetweenTheWorkAndItsLanding in the sharded session suite. It passes when run alone, and plain origin/dev fails it too under make test PKGS=./internal/session, along with two other timing tests, so it looks load-related. Test changes made to match the new behavior: three standing tests now expect "said", and dev's old refuse-the-turn daily-limit test is removed because the turn now waits on a raise-or-stop card. I did not show every new test failing without its fix.

The read hand-off admitted an ordinary quick task with the full belt (bash,
commit), guarded only by a prose line, so a manager's whole directive could be
taken by a "reading" helper that wrote, committed and merged, and a request
for a task on a named model became in-place edits with no branch. The quick
node now carries a durable read-only mark and its worker gets the audit
read-only belt at the one place quick workers are built; a request the reader
cannot serve hands control back to the conversation. A failed hand-off retires
a never-started helper instead of cancelling it into a stopped card, and still
stops one that did start.

Fixes #1568
Part of #1569
Part of #1554
…anded

A task firing that changed no file but reported a sentence was logged as
landed. The rule that such a run is not "nothing" stands (its words reached
the person); it now comes to said, which the outcome set already has, and
only a saved change is landed.

Fixes #1582
… bad check paths are refused, kept branches reach /land

- Every checker in a run drew from one checker-seat tally, so after the
  first few checks the rest stopped "at its spend ceiling" having spent
  almost nothing. The ceiling is now kept per check task (the run's task cap
  is still shared), and a run whose checks did not finish no longer ends done;
  its landing names the unfinished checks.
- A run kept on a protected or moved checkout was invisible to /land. Kept,
  conflicted and aborted run branches are now listed and landable by /land,
  and editor .orig backups are classified as droppings so they never reach a
  task branch.

Fixes #1572
Fixes #1574
…ranch

The guard that tasks do not merge into a protected checkout stays. But a
single task's kept branch never reached /land (only fan-out run rows did),
and the landing line said "branch kept" without the reason the guard knew.
Kept branches now come from one source over task nodes and run rows, which
the standing chip, its preview and /land all read, and the reason (on main,
moved, detached) is carried on the notice and drawn with the line.

Fixes #1574
…anch can land

Ordinary chat talks to its engine through the local host, whose adapter had
no landing door, so /land always said "nothing is waiting" however many
kept branches the engine held. The waiting list now rides the pushed session
facts, landing preview and landing cross the host as calls, and both run off
the paint loop. The surface-door ledger loses its two landing entries.

Fixes #1574
Only headless runs and standing firings checked the machine-wide daily
limit; a chat turn and a chat /task ran past it silently (the crew guard was
built with the daily fold off). Today's limit, its day-scoped raise and the
refusal sentence now live in one place (session.DailySpendAt) that chat
turns, chat-task crews and codeaf do all read. A chat turn past the limit
raises a card like the team cap's, "1 Raise to $X" or "2 Stop for today",
before any model call; a chat-started task's crew guard folds the same limit.

Part of #1546
…hows the raised limit

A /task or an approved proposal past today's limit reported "started" and
then sat working, held by the crew guard with nothing said. Task admission
now waits on the same daily-budget card as a chat turn: stop refuses with
the headless sentence, raise re-checks and admits. The top bar's limit is
read through the same DailySpendAt reading, so a raise shows at once.

Part of #1546
@santoshkumarradha

Copy link
Copy Markdown
Member Author

PR bar audit — directive #1790715358822565124

1. CI + changelog: ✅ all checks green (light gate, touched packages, check, license/cla). ✅ one unreleased changelog entry present (1670-task-run-audit.md).

2. Review passes: ❌ Pass 1 — ARCHITECTURE: absent. ❌ Pass 2 — SECURITY: absent. ❌ Pass 3 — CODE EFFICIENCY & DESIGN: absent.

3. Video proof: ❌ no video/GIF embedded in the PR body. This change is user-facing (held work, kept branches and spend limits say what is true) — the bar requires a capture of the change actually running, embedded in the body. To be captured on Spark.

4. Merge state: CONFLICTING against dev — needs a rebase; CI must be re-run green on the rebased head.

Fixes for the gaps are in flight. This PR does not turn ready until all four conditions hold.

drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

PR-bar audit: CI green + changelog present — those hold. Missing:

  1. Review passes — architecture / security / efficiency & design, none posted yet.
  2. On-screen proof — held work, kept branches and spend limits are person-facing task surfaces; the bar requires the Spark tmux capture embedded in the body.

Marked draft until both close.
drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

Bar audit — gaps against the new PR bar

1. CI + changelog: ⚠️ green but stale: run 36505663619 passed (check, light gate, touched packages, license/cla) on 2026-09-29; dev has moved since, so a fresh green gate against current dev is required before ready. ✅ changelog entry present (1670-task-run-audit.md).

2. Review passes: ❌ none of the three are posted (Pass 1 ARCHITECTURE, Pass 2 SECURITY, Pass 3 CODE EFFICIENCY & DESIGN), each as a separate comment naming what was checked and what was found.

3. Video proof: ❌ absent. User-facing change (task page: held work, retry/done wording, landcmd rendering). Needs a Spark tmux capture rendered to GIF/video, embedded in this PR body.

4. Spark: all verification ahead of ready runs on the Spark (fleet).

Action taken: converted back to draft until fresh CI, the three passes, and the body video exist.

—
Drafted with CodeAF · reviewed and owned by the author

@santoshkumarradha
santoshkumarradha marked this pull request as draft October 3, 2026 18:53
@santoshkumarradha

Copy link
Copy Markdown
Member Author

PR Bar Audit — gaps against the new bar

1. CI + changelog: ✅ CI fully green (check, light gate, touched packages, license/cla). ✅ Changelog present (docs/changes/unreleased/1670-task-run-audit.md).
2. Review passes: ❌ None posted — Pass 1 (ARCHITECTURE), Pass 2 (SECURITY), Pass 3 (CODE EFFICIENCY & DESIGN) all absent.
3. Video proof: ❌ Missing — user-facing change (held work / kept branches / spend-limit status text on screen); no video or GIF in the body.
4. Spark runs: captures/tests must run on the Spark (fleet).
drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

Audit against the new PR bar

  • CI + changelog: all checks green (check, light gate, touched packages, license/cla) ✅ — changelog entry present ✅
  • Review passes: 0 of 3 — ARCHITECTURE, SECURITY, and CODE EFFICIENCY & DESIGN, each as its own comment naming what was checked and what was found.
  • Video proof: missing — the spend-limit card ("Raise to $X" / "Stop for today") and the held-work/kept-branch messages are user-facing; a capture must be embedded in the PR body.

Converted back to draft until all four conditions hold; nothing merges until then.

drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

PR-bar audit (team bar, 2026-10-03) - status per category:

  • CI: 1 of 4 checks green, 3 still running, none failing.
  • Changelog: present (1 entry).
  • Review passes: 0 of 3 - the three labeled pass comments are absent; they are being posted (66-file diff).
  • Video proof: required (user-facing TUI: held-work/land-command visibility) and MISSING - returning to draft per the no-video rule.

Stays draft until all four conditions hold; nothing merged. Fixes incoming: three review passes, then Spark capture.

drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

PR bar audit — gaps against the new bar

1. CI + changelog — ✅ all checks green (light gate, touched packages, check, license/cla). ✅ changelog entry present (docs/changes/unreleased/1670-task-run-audit.md).

2. Review passes — all three missing (Pass 1 ARCHITECTURE, Pass 2 SECURITY, Pass 3 CODE EFFICIENCY & DESIGN), as separate comments naming what was checked and what was found.

3. Video proof — missing. User-facing change (held work, kept branches and spend limits must say what is true). A Spark tmux capture of the corrected status behavior, rendered to GIF/video, must be embedded in the PR body.

4. Runs on Spark — per the fleet skill.

Converted back to draft until all four conditions hold. Nothing merges before that.

drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

Audit against the PR bar — what is missing

1. CI + changelog — ✅ every check green (check, light gate, touched packages, license/cla); ✅ changelog entry present.
2. Three review passes — ❌ none posted yet. Pass 1 (ARCHITECTURE), Pass 2 (SECURITY) and Pass 3 (CODE EFFICIENCY & DESIGN) follow as three separate comments, each naming what was checked and what was found.
3. On-screen proof — ❌ user-facing change (task rows: held work, kept branches, spend limits): no capture in the body yet.
4. Fleet runs — builds, tests and captures run over ssh on the fleet, never a laptop. Spark is offline at audit time, so the dumb fleet host stands in until it is back.

Also missing: labels and milestone (existing sets only).
Disposition — converted from ready back to draft until all four conditions hold; nothing merges.

drafted with CodeAF

@santoshkumarradha

Copy link
Copy Markdown
Member Author

Audit against the new PR bar:

  • CI: green on its last run, but the branch now conflicts with current dev — rebase and re-run before trusting it.
  • Changelog: present.
  • Review passes: none posted. Owed: Pass 1 — Architecture, Pass 2 — Security, Pass 3 — Code Efficiency & Design.
  • Video: none in the body; the task rows and spend-limit surfaces are user-facing, so a capture is owed.

Stays draft.

drafted with CodeAF

@santoshkumarradha santoshkumarradha added this to the Reliable agent milestone Oct 3, 2026
@santoshkumarradha santoshkumarradha added area:session The engine — turns, tasks, the toolbelt, checkpoints feature Work that adds a capability; developers break it into tasks labels Oct 3, 2026
@santoshkumarradha

Copy link
Copy Markdown
Member Author

Bar audit — what is missing (per the new PR bar)

Audited at head b3b5d7d9c16ded140280ace1d341dcb935c3e7a2 by @tags.

1. CI + changelog — light gate ✅ · touched packages ✅ · check ✅ · license/cla ⏳ pending. Changelog ✅ — docs/changes/unreleased/1670-task-run-audit.md.

2. Review passes (three, separate comments) — Pass 1 — ARCHITECTURE ❌ absent · Pass 2 — SECURITY ❌ absent · Pass 3 — CODE EFFICIENCY & DESIGN ❌ absent — 0 of 3 posted.

3. Proof on screen — user-facing ✅ (the spend-limit ask card "Raise to $X" / "Stop for today" and the top-bar raise). Video/GIF embedded in the body ❌ missing — required before ready.

4. Tags — were absent → now bug · area:session · area:chat · sev:serious · milestone Reliable agent.

Actions: converted to draft until the bar holds; the three review passes are running and the capture is queued. Nothing merges before all four conditions hold.

drafted with CodeAF

@santoshkumarradha santoshkumarradha added bug Something the code does that it should not area:chat The v3 surface a person sits in front of (internal/tui3) sev:serious Wrong or missing behaviour a person meets in ordinary use labels Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:chat The v3 surface a person sits in front of (internal/tui3) area:session The engine — turns, tasks, the toolbelt, checkpoints bug Something the code does that it should not feature Work that adds a capability; developers break it into tasks sev:serious Wrong or missing behaviour a person meets in ordinary use

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant