Declarative test-scenario harness engine for the cclab ecosystem.
rig runs declarative SCENARIOS (e2e behavior, TOML step DSL) and LOAD
profiles (open-loop QPS × workers) against a real serving process, judges
them with assertions and declarative pins (floors/ratchets vs a baseline
store), and prints ONE agent-readable JSON report (rig.report/1).
It is the extracted, domain-free essence of mamba's
tests/harness/cpython harness: the fixture-record contract
(path==record + lint), declarative gates, verdict bucketing
(pass/xfail/skip), and child-process timeout policy — generalized so any
project can consume them by writing TOML scenarios, not bash.
vat — environment: services (postgres/redis/nats/...), COW workspace,
readiness probes. rig declares needs; vat satisfies them.
rig — case orchestration: scenario steps, assertions, load generation,
pin gates, verdicts, the report.
meter — resource attribution (profiling). rig never claims it.
rig run --dir tests/rig/scenarios [--vat] [--pins tests/rig/config/pins]
discover *.toml -> lint records (path==record)
per scenario: interpolate {{vars}} -> execute steps under TimeoutPolicy
http | sample | assert | wait_until | measure_rss | exec | sleep
kind = "load": open-loop generator -> p50/p99/error_rate/achieved_qps
gate pins (floor / ratchet vs .rig/baselines.json)
bucket verdicts (pass / xfail / skip; xpass = graduate-to-pass signal)
print ONE RigReport JSON -> exit 0 clean / 1 findings / 2 regression / 3+ tool error
| verb | status | behavior |
|---|---|---|
rig test [--dir <d>] [--dimension <d>] [--case <id>] [--collect] [--pins <d>] [--update-baselines] |
v1 | lifecycle-case launcher: discover [case] TOMLs → prepare/exercise(N)/clean → fold one report |
rig run [--scenario <f> | --dir <d>] [--pins <d>] [--update-baselines] [--vat] |
v0 (deprecated) | flat scenarios: discover → lint → execute → gate → one JSON report |
rig lint [--dir <d>] |
v0 | record-contract check only, no execution |
rig report |
v0 | re-project .rig/last-report.json (read-only) |
A promise with no gate under it is not claimed.
Nothing reads the tables below. The capability gate that validated their
shape was deleted with the aw binary, so the shape is convention now and
the commands named in each row are the only part that runs.
| Capability | Root WI | Notes |
|---|---|---|
| Scenario Engine | axiom#5 | record contract + lint, step DSL (http/sample/assert/wait_until/measure_rss/exec/sleep), verdict bucketing, rig.report/1 |
| Stateful Service Scenarios | axiom#1645 | shared warm-up/observe/fault/recover/verify/teardown runner; bounded phases, retained failure evidence, Lumen/Tape adapters |
| Load Pins | axiom#5 | open-loop loadgen (coordinated-omission honest), floor/ratchet pins, per-host JSON baseline store |
| Vat Wrapped Runs | axiom#5 | --vat shells vat run, parses JSONL checkpoints, lifts the inner report, removes the vat |
rig run discovers declarative scenario records, executes step DSL actions,
buckets verdicts, and emits one deterministic rig.report/1 JSON document.
- Root WI: #axiom#5
- Surfaces: CLI:
rig test [--dir <d>] [--dimension <d>] [--case <id>]+rig run --scenario <f>+rig run --dir <d>+rig lint --dir <d>+rig report- Scenario/case orchestration, deprecated flat scenario execution, record linting, and report reprojection entrypoints. - Gate — behavior:
rig- declarative scenario records, step DSL execution, assertions, verdict bucketing, and rig.report/1 output - Gate:
cargo test -p rig - Gate:
target/debug/rig lint --dir tests/fixtures/scenarios
| Work Root | Kind | WI | Gate / Evidence |
|---|---|---|---|
| Record contract check and JSON report | epic | axiom#5 | cargo test -p rig |
| Scenario step DSL execution | epic | axiom#5 | cargo test -p rig |
Rig runs the same bounded stateful-service lifecycle for every consumer while each app supplies only its fault operation and domain-specific continuity assertions. A failed or timed-out phase retains all prior evidence and never suppresses teardown.
- Root WI: #axiom#1645
- Surfaces: Library:
rig::engine::stateful::{run_stateful, StatefulScenario, StatefulActions}- Fixed warm-up, observation, fault, recovery, verification, and teardown lifecycle for long-running stateful services. - Gate — stability:
rig.stateful.v1- bounded phase execution, independent teardown reserve, ordered evidence retention, failed-phase attribution, and deterministic report structure - Gate:
cargo test -p rig --test stateful_service_harness - Gate:
cargo test -p lumen --test rig_stateful_adapter - Gate:
cargo test -p tape --test rig_stateful_adapter - Evidence:
cargo test -p rig --test stateful_service_harness;cargo test -p lumen --test rig_stateful_adapter;cargo test -p tape --test rig_stateful_adapter
rig runs open-loop load profiles and gates measured values against
floor/ratchet pins in a host-scoped baseline store.
- Root WI: #axiom#5
- Surfaces: CLI:
rig test --pins <d>+rig test --update-baselines+rig run --pins <d>+rig run --update-baselines- Open-loop load profile execution and baseline pin gate entrypoints. - Gate — efficiency:
rig- open-loop load profiles, p50/p99/error-rate observations, and floor/ratchet pins against host baselines - Gate:
cargo test -p rig
| Work Root | Kind | WI | Gate / Evidence |
|---|---|---|---|
| Open-loop load generator | epic | axiom#5 | cargo test -p rig |
| Floor and ratchet pin gates | epic | axiom#5 | cargo test -p rig |
rig --vat delegates environment setup to vat, consumes JSONL checkpoints,
and lifts the inner rig report without owning resource isolation.
- Root WI: #axiom#5
- Surfaces: CLI:
rig test --vat+rig run --vat- Delegated scenario execution through vat-managed environments while preserving rig report folding. - Gate — behavior:
rig + vat- rig delegates environment setup to vat, consumes vat JSONL checkpoints, and lifts the inner rig report - Gate — stability:
rig + vat- vat-managed services, readiness, timeout policy, cleanup, and retained report/error folding across scenario runs - Gate:
cargo test -p rig
Verified smoke (2026-06-10): lumen's resilience (partition/packet-loss via
toxiproxy) + endurance (RSS plateau) + load (search p99 pin) scenarios run
green locally and through rig run --vat with vat-managed services;
cargo test -p rig -p rig-cli green.
- Evidence:
cargo test -p rig
- Baselines are environment-scoped by convention, not enforcement. The
per-host key is
os-archonly, so a baseline recorded on the host gates vat-wrapped runs too (the COW clone carries.rig/along). Record baselines in the environment you gate in; persisting baselines from inside a vat run back to the host is v1. - Relative latency budgets on loopback are tight. Sub-millisecond
baselines make
2xbudgets quantization-sensitive — scenarios use the assert tolerance term (+ 1) and realistic corpus seeding to stay stable; a loaded host can still legitimately trip them.
- kind/k8s environment provisioning (scenario DSL can express the assertions today; vat has no kind preset yet — lands with vat, not rig)
- multi-host / distributed load generation
- HTML reports (
--humanstderr summary only) - fixture GENERATION tooling (generate→fill loops stay project-side)
- resource attribution (meter owns it)
- closed-loop (latency-coupled) load
lumen: apps/lumen/e2e/rig/cases/ ports scripts/chaos.sh
(partition recovery, packet-loss p99) and scripts/soak.sh (two-window
RSS plateau) to scenarios, plus one load/search_qps pin
(config/pins/search_p99.toml).