Skip to content
faberlinePublic

About

Declarative test-scenario harness engine: e2e scenarios + open-loop load pins, one agent-readable JSON report.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

rig

Brief

Declarative test-scenario harness engine for the cclab ecosystem.

rig runs declarative SCENARIOS (e2e behavior, TOML step DSL) and LOAD profiles (open-loop QPS × workers) against a real serving process, judges them with assertions and declarative pins (floors/ratchets vs a baseline store), and prints ONE agent-readable JSON report (rig.report/1).

It is the extracted, domain-free essence of mamba's tests/harness/cpython harness: the fixture-record contract (path==record + lint), declarative gates, verdict bucketing (pass/xfail/skip), and child-process timeout policy — generalized so any project can consume them by writing TOML scenarios, not bash.

Division of labor

vat   — environment: services (postgres/redis/nats/...), COW workspace,
        readiness probes. rig declares needs; vat satisfies them.
rig   — case orchestration: scenario steps, assertions, load generation,
        pin gates, verdicts, the report.
meter — resource attribution (profiling). rig never claims it.

Mental model

rig run --dir tests/rig/scenarios [--vat] [--pins tests/rig/config/pins]
  discover *.toml -> lint records (path==record)
  per scenario: interpolate {{vars}} -> execute steps under TimeoutPolicy
    http | sample | assert | wait_until | measure_rss | exec | sleep
  kind = "load": open-loop generator -> p50/p99/error_rate/achieved_qps
  gate pins (floor / ratchet vs .rig/baselines.json)
  bucket verdicts (pass / xfail / skip; xpass = graduate-to-pass signal)
  print ONE RigReport JSON -> exit 0 clean / 1 findings / 2 regression / 3+ tool error

Verbs

verb status behavior
rig test [--dir <d>] [--dimension <d>] [--case <id>] [--collect] [--pins <d>] [--update-baselines] v1 lifecycle-case launcher: discover [case] TOMLs → prepare/exercise(N)/clean → fold one report
rig run [--scenario <f> | --dir <d>] [--pins <d>] [--update-baselines] [--vat] v0 (deprecated) flat scenarios: discover → lint → execute → gate → one JSON report
rig lint [--dir <d>] v0 record-contract check only, no execution
rig report v0 re-project .rig/last-report.json (read-only)

Capabilities

A promise with no gate under it is not claimed.

Nothing reads the tables below. The capability gate that validated their shape was deleted with the aw binary, so the shape is convention now and the commands named in each row are the only part that runs.

Capability Index

Capability Root WI Notes
Scenario Engine axiom#5 record contract + lint, step DSL (http/sample/assert/wait_until/measure_rss/exec/sleep), verdict bucketing, rig.report/1
Stateful Service Scenarios axiom#1645 shared warm-up/observe/fault/recover/verify/teardown runner; bounded phases, retained failure evidence, Lumen/Tape adapters
Load Pins axiom#5 open-loop loadgen (coordinated-omission honest), floor/ratchet pins, per-host JSON baseline store
Vat Wrapped Runs axiom#5 --vat shells vat run, parses JSONL checkpoints, lifts the inner report, removes the vat

Scenario Engine

rig run discovers declarative scenario records, executes step DSL actions, buckets verdicts, and emits one deterministic rig.report/1 JSON document.

  • Root WI: #axiom#5
  • Surfaces: CLI: rig test [--dir <d>] [--dimension <d>] [--case <id>] + rig run --scenario <f> + rig run --dir <d> + rig lint --dir <d> + rig report - Scenario/case orchestration, deprecated flat scenario execution, record linting, and report reprojection entrypoints.
  • Gate — behavior: rig - declarative scenario records, step DSL execution, assertions, verdict bucketing, and rig.report/1 output
  • Gate: cargo test -p rig
  • Gate: target/debug/rig lint --dir tests/fixtures/scenarios
Work Root Kind WI Gate / Evidence
Record contract check and JSON report epic axiom#5 cargo test -p rig
Scenario step DSL execution epic axiom#5 cargo test -p rig

Stateful Service Scenarios

Rig runs the same bounded stateful-service lifecycle for every consumer while each app supplies only its fault operation and domain-specific continuity assertions. A failed or timed-out phase retains all prior evidence and never suppresses teardown.

  • Root WI: #axiom#1645
  • Surfaces: Library: rig::engine::stateful::{run_stateful, StatefulScenario, StatefulActions} - Fixed warm-up, observation, fault, recovery, verification, and teardown lifecycle for long-running stateful services.
  • Gate — stability: rig.stateful.v1 - bounded phase execution, independent teardown reserve, ordered evidence retention, failed-phase attribution, and deterministic report structure
  • Gate: cargo test -p rig --test stateful_service_harness
  • Gate: cargo test -p lumen --test rig_stateful_adapter
  • Gate: cargo test -p tape --test rig_stateful_adapter
  • Evidence: cargo test -p rig --test stateful_service_harness; cargo test -p lumen --test rig_stateful_adapter; cargo test -p tape --test rig_stateful_adapter

Load Pins

rig runs open-loop load profiles and gates measured values against floor/ratchet pins in a host-scoped baseline store.

  • Root WI: #axiom#5
  • Surfaces: CLI: rig test --pins <d> + rig test --update-baselines + rig run --pins <d> + rig run --update-baselines - Open-loop load profile execution and baseline pin gate entrypoints.
  • Gate — efficiency: rig - open-loop load profiles, p50/p99/error-rate observations, and floor/ratchet pins against host baselines
  • Gate: cargo test -p rig
Work Root Kind WI Gate / Evidence
Open-loop load generator epic axiom#5 cargo test -p rig
Floor and ratchet pin gates epic axiom#5 cargo test -p rig

Vat Wrapped Runs

rig --vat delegates environment setup to vat, consumes JSONL checkpoints, and lifts the inner rig report without owning resource isolation.

  • Root WI: #axiom#5
  • Surfaces: CLI: rig test --vat + rig run --vat - Delegated scenario execution through vat-managed environments while preserving rig report folding.
  • Gate — behavior: rig + vat - rig delegates environment setup to vat, consumes vat JSONL checkpoints, and lifts the inner rig report
  • Gate — stability: rig + vat - vat-managed services, readiness, timeout policy, cleanup, and retained report/error folding across scenario runs
  • Gate: cargo test -p rig

Verified smoke (2026-06-10): lumen's resilience (partition/packet-loss via toxiproxy) + endurance (RSS plateau) + load (search p99 pin) scenarios run green locally and through rig run --vat with vat-managed services; cargo test -p rig -p rig-cli green.

  • Evidence: cargo test -p rig

Known limits (v0)

  • Baselines are environment-scoped by convention, not enforcement. The per-host key is os-arch only, so a baseline recorded on the host gates vat-wrapped runs too (the COW clone carries .rig/ along). Record baselines in the environment you gate in; persisting baselines from inside a vat run back to the host is v1.
  • Relative latency budgets on loopback are tight. Sub-millisecond baselines make 2x budgets quantization-sensitive — scenarios use the assert tolerance term (+ 1) and realistic corpus seeding to stay stable; a loaded host can still legitimately trip them.

Non-goals (v0)

  • kind/k8s environment provisioning (scenario DSL can express the assertions today; vat has no kind preset yet — lands with vat, not rig)
  • multi-host / distributed load generation
  • HTML reports (--human stderr summary only)
  • fixture GENERATION tooling (generate→fill loops stay project-side)
  • resource attribution (meter owns it)
  • closed-loop (latency-coupled) load

First consumer

lumen: apps/lumen/e2e/rig/cases/ ports scripts/chaos.sh (partition recovery, packet-loss p99) and scripts/soak.sh (two-window RSS plateau) to scenarios, plus one load/search_qps pin (config/pins/search_p99.toml).

About

Declarative test-scenario harness engine: e2e scenarios + open-loop load pins, one agent-readable JSON report.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages