Skip to content

feat: multi-GPU proving - #1403

Merged
hero78119 merged 27 commits into
masterfrom
feat/multi_gpu
Sep 15, 2026
Merged

hero78119 merged 27 commits into
masterfrom
feat/multi_gpu

Conversation

@hero78119

@hero78119 hero78119 commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Ceno used process-global CUDA state and could prove base shards and recursion on only one device. Device ownership, memory validation, replay assignment, failure propagation, and base-to-recursion GPU reuse were implicit.

Design Rationale

Each selected GPU owns its HAL, CUDA context, stream bindings, AOT replay state, and base prover. Round-robin ownership (shard_id % device_count) is deterministic: owned shards retain compact witnesses, while unowned shards fast-flight through canonical execution and replay-digest updates without compact witness arenas. Depth-one queues bound memory and apply backpressure.

The collector accepts out-of-order results, validates ownership and completeness, restores canonical shard order, and runs one full-trace Rust verification. Failures cancel all workers without cross-device retry.

Recursion V2 uses an event-driven DAG over the same GPU set. Consecutive base proofs immediately unlock leaf tasks; completed nodes notify eligible parents through a blocking scheduler. Leaf work is prioritized, followed by the deepest ready intermediate layer and root. A GPU joins recursion only after its base state is dropped, synchronized, and trimmed.

Shard-independent exact-VK recursion assets are built once and overlapped with AOT/base work. Shard count later binds the lightweight recursion plan and constructs only additional required depths. Recursion may run speculatively, but root publication remains gated on canonical base verification and the root proof is independently verified.

Change Highlights

  • gkr_iop: worker-owned HALs with scoped context and stream bindings.
  • ceno_emul: canonical AOT replay with owned compact capture and unowned fast flight.
  • ceno_zkvm: device discovery, validation, conservative shared shard limits, round-robin replay/proving, bounded collection, fail-first diagnostics, and base-to-recursion release events.
  • ceno_recursion_v2: exact-VK host-asset templates, deterministic recursion DAG, blocking multi-worker scheduling, device-local hydration, and streaming root completion.
  • CLI/SDK: GPU device list/count and optional worker CPU affinity; logical device 0 remains the default. GPU builds require AOT.

Benchmark / Performance Impact

Workload: Reth block 23817600, chain ID 1, 11 shards, max_cell_per_shard=4500000000, lanes 4, cache level 1, jagged reshape height 23, and a 3048 MiB booking margin. Both comparisons ran serially on the same dual-RTX-4090 self-hosted runner; only the selected device list differs.

One GPU (--gpu-devices 0)

Metric Baseline Latest Change
Base-prover setup 13.308552 s 13.868011 s 4.20% slower
Base proof vector ready 42.662679 s 44.259000 s 3.74% slower
Raw recursion worker 5.759648 s 6.371000 s 10.61% slower
Recursion tail exposed after base work 5.759648 s 5.207741 s 9.58% faster
Raw proof timer 48.488613 s 50.846579 s 4.86% slower*
Setup-normalized proof wall 51.807145 s 50.846579 s 1.85% faster (1.0189x)
Full measured lifecycle 65.125955 s 64.724166 s 0.62% faster
Root verification 27.733 ms 27.684 ms 0.18% faster

The raw proof timer is not an equivalent comparison: the baseline built exact-VK recursion assets for 3.318532 s outside that timer. The latest path starts the 6.229 s host-template build during setup and fully overlaps it (template_wait_ms=0, bind_ms=0), so the setup-normalized and full-lifecycle rows are the fair comparisons. The raw recursion worker is slower because it performs four legal post-trim device hydrations (425 ms total), but streaming overlaps 1.163 s with base proving and reduces the exposed recursion tail by 0.552 s.

Two GPUs (--gpu-devices 0,1)

Metric Baseline Latest Change
Base-prover setup 13.539582 s 13.898621 s 2.65% slower
Base proof vector ready 25.914000 s 27.740000 s 7.05% slower
Recursion critical path 6.146068 s 5.074000 s 17.44% faster
Recursion tail exposed after base work 6.146068 s 2.793161 s 54.55% faster
Raw proof timer 33.674897 s 31.831985 s 5.47% faster
Setup-normalized proof wall 37.001080 s 31.831985 s 13.97% faster (1.1624x)
Full measured lifecycle 50.548843 s 45.739433 s 9.51% faster
Root verification 27.697 ms 27 ms 2.52% faster

GPU 0 proves even shards and GPU 1 proves odd shards. All 11 proof-queue waits are zero, so replay produces every next owned shard before its current proof finishes. Unowned shards use fast-flight replay with zero compact rows/bytes. Streaming recursion schedules the six-task 4/4 DAG as dependencies become ready: GPU 0 executes four tasks and GPU 1 executes two leaf tasks. This overlaps 2.281 s of recursion with base proving and cuts the exposed recursion tail by 3.353 s. The remaining base-vector slowdown comes from duplicated replay/resource contention, six-even/five-odd round-robin imbalance, and run variance; it does not erase the E2E improvement.

Latest one GPU vs. two GPUs

Metric One GPU Two GPUs Two-GPU improvement
Base-prover setup 13.868011 s 13.898621 s 0.22% slower
Base proof vector ready 44.259000 s 27.740000 s 37.32% faster (1.5955x)
Recursion critical path 6.371000 s 5.074000 s 20.36% faster (1.2556x)
Recursion tail exposed after base work 5.207741 s 2.793161 s 46.36% faster (1.8645x)
Proof wall 50.846579 s 31.831985 s 37.40% faster (1.5973x)
Full measured lifecycle 64.724166 s 45.739433 s 29.33% faster (1.4151x)

The same-revision two-GPU run therefore improves proof wall by 37.40%, rather than the ideal 50%. The gap is expected from duplicated replay, non-shard serial work and verification, six-versus-five round-robin imbalance, and GPU/host resource contention. The zero queue waits show that replay is not starving either prover after its first owned shard.

Correctness and evidence

  • Both latest runs completed all 11 canonical shard proofs, canonical Rust full-trace verification, recursion, independent root verification, and exited successfully with an empty failure scan.
  • The two-GPU ownership map is exact: GPU 0 owns shards 0,2,4,6,8,10; GPU 1 owns 1,3,5,7,9.
  • Both GPUs trim/release base allocations before recursion; GPU 1 begins eligible leaf work without waiting for GPU 0 to finish all base work.
  • Build identity: Ceno c08df94610ba7e74c27c97019cb572d834f096bb, benchmark 038f4e97f68c423691ddcdc14ca236dba71534a2, ceno-gpu dceb5e7e04dddf7f49ac0c7ffb2ef8cf91c536d3, CUDA_ARCH=89,120 (RTX 4090 loads SASS 89).

Raw logs:

Testing

cargo fmt --all --check
cargo make clippy
CUDA_ARCH=120 cargo test -p ceno_zkvm multi_gpu --features gpu,aot-x86_64
cargo test -p ceno_recursion_v2 continuation::prover::scheduler::tests --lib
CUDA_ARCH=120 cargo test -p cargo-ceno streaming_recursion_tests --features gpu,aot-x86_64
  • Device selection, ownership, fast-flight replay, bounded queues, cancellation, canonical collection, and CUDA binding tests pass.
  • Recursion planning, notification, priority, exact-VK reuse, device hydration, failure propagation, and root-publication gate tests pass.
  • One- and two-GPU RTX 4090 E2E runs pass all 11 base proofs, canonical verification, streaming recursion, independent root verification, and failure scans.

Risks and Rollout

  • Replay remains duplicated per GPU and round-robin assignment can leave uneven shard work.
  • CUDA ownership and base-to-recursion transitions are correctness-sensitive; focused concurrency tests and dual-GPU E2E cover both.
  • Recursion scheduling is non-preemptive and retains single-task DAG bottlenecks near the root.
  • Device capability, module loading, HAL construction, and usable memory are validated before proving. The default remains logical GPU 0.

Follow-ups (optional)

  • Shared replay and cost-aware shard assignment.
  • Retry/reassignment and peer-to-peer transfer.

Copilot Reviewer Directive (keep this section)

When Copilot reviews this PR, apply .github/copilot-instructions.md strictly.

@hero78119
hero78119 marked this pull request as draft September 4, 2026 08:43
@hero78119 hero78119 changed the title perf: [stage 1] multi-gpu base layer pipeline proving feat: add Stage 1 multi-GPU base proving Sep 6, 2026
@hero78119 hero78119 changed the title feat: add Stage 1 multi-GPU base proving feat: multi-GPU proving Sep 8, 2026
@hero78119
hero78119 marked this pull request as ready for review September 8, 2026 12:32
@hero78119
hero78119 merged commit 7c58fdd into master Sep 15, 2026
5 of 6 checks passed
@hero78119
hero78119 deleted the feat/multi_gpu branch September 15, 2026 15:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant