What happened?
Snapshot::save() appears vulnerable to a lost-update race condition when multiple threads or processes concurrently save snapshots with different tags into the same OCI layout directory.
The implementation performs a read-modify-write operation on index.json without synchronization across concurrent writers.
When two workers read the same original index before either writes its update, both independently construct a modified index and atomically replace the existing file.
The last writer can overwrite the first writer's manifest entry.
Consequently, both snapshot saves may return Ok(...), while one of the snapshots becomes undiscoverable through its tag.
This can compromise snapshot reliability in production systems using Hyperlight for:
- Concurrent sandbox checkpointing
- AI agent execution and state persistence
- Multi-tenant sandbox infrastructure
- Snapshot-based recovery and restoration
- Parallel snapshot persistence services
Important: This is a code-inspection finding with a deterministic race scenario. An end-to-end Hyperlight runtime reproduction is still pending.
Steps to Reproduce
1. Check out Hyperlight's main branch and create two valid Snapshot instances using the existing snapshot test infrastructure.
2. Create an empty temporary OCI layout directory shared by both snapshots.
3. Start two concurrent workers executing the following operations:
// Illustrative test operations
// Worker A
snapshot_a.save(shared_path, tag_a);
// Worker B
snapshot_b.save(shared_path, tag_b);
Use different OCI tags, such as worker-a and worker-b.
4. To reproduce deterministically, add a test-only synchronization barrier immediately after the existing index has been read into the manifests collection.
Affected file:
src/hyperlight_host/src/sandbox/snapshot/file/mod.rs
Approximate location: Lines 432–449.
Ensure that both workers read the same original index before either proceeds to update it.
5. Release both workers and allow their Snapshot::save() calls to complete.
6. Inspect the final OCI index:
jq -r \
'.manifests[].annotations["org.opencontainers.image.ref.name"]' \
/path/to/shared/oci-layout/index.json
7. Verify that both worker-a and worker-b remain discoverable.
8. Attempt to load each snapshot by its respective tag.
Expected problematic interleaving:
Worker A Worker B
-------- --------
Read index.json Read index.json
(manifests = []) (manifests = [])
Add snapshot A Add snapshot B
Build index [A] Build index [B]
Atomic replace Atomic replace
index.json with [A] index.json with [B]
Return Ok Return Ok
Final index.json = [B]
Snapshot A reference is lost
Reproduction status: The interleaving is supported by source inspection. A direct Hyperlight integration test has not yet been executed.
Expected Results
When multiple snapshots are successfully saved using different tags into the same OCI layout:
- All successfully saved snapshot references must remain present in
index.json.
- Each snapshot must remain loadable using its corresponding tag.
- Concurrent save operations must not silently discard successful updates.
- Conflicting updates must either be safely serialized or fail with an explicit error.
- Atomic file replacement must preserve file integrity and concurrent update correctness.
Actual Results
Source inspection indicates that concurrent snapshot saves can overwrite each other's OCI index updates.
The implementation currently:
- Reads the existing
index.json.
- Copies its manifest descriptors into memory.
- Constructs a new snapshot descriptor.
- Updates the local manifest list.
- Serializes a replacement index.
- Atomically replaces
index.json.
No cross-writer synchronization is visible around the entire read-modify-write operation.
Under the described concurrent interleaving:
- Both save operations can report success.
- Only one newly written snapshot tag may survive.
- The other snapshot may become unreachable through normal tag-based loading.
- Content-addressed snapshot blobs may remain on disk without a corresponding index reference.
- Applications may incorrectly assume that a successfully saved checkpoint is recoverable.
Actual runtime results: Not yet collected. These are the predicted outcomes of the identified race.
Versions and Environment
Hyperlight version or commit:
- Repository:
hyperlight-dev/hyperlight
- Branch:
main
- Commit inspected:
dcb53c04a5a7ae9dc9bf6a8cd54d56067dea138e
- Component:
hyperlight_host
- Affected file:
src/hyperlight_host/src/sandbox/snapshot/file/mod.rs
- Source inspection date: October 10, 2026
OS Version
Target environment: Linux.
Run:
cat /etc/os-release && uname -a
Actual output: Not yet collected.
The underlying read-modify-write race is not inherently Linux-specific.
Hypervisor
Run the following commands to check hypervisor access:
ls -la /dev/kvm /dev/mshv 2>&1
getfacl /dev/kvm /dev/mshv 2>&1
id
[ -r /dev/kvm ] && [ -w /dev/kvm ] && echo "KVM: OK" || echo "KVM: FAIL"
[ -r /dev/mshv ] && [ -w /dev/mshv ] && echo "MSHV: OK" || echo "MSHV: FAIL"
Actual hypervisor output: Not yet collected.
Note: Hypervisor access may be required to construct snapshots for an integration reproduction. The identified race occurs in host-side OCI index persistence.
Extra Info
Root Cause Analysis
Affected file:
src/hyperlight_host/src/sandbox/snapshot/file/mod.rs
Relevant code locations:
- Lines 432–449: Read the existing OCI index and initialize the manifest list.
- Lines 451–452: Write snapshot blobs and construct a new descriptor.
- Lines 457–464: Update the manifest collection.
- Lines 466–474: Construct and serialize the modified OCI index.
- Line 488: Atomically replace
index.json.
Root cause:
The replace_file_atomic() helper provides atomic replacement of the index file.
However, atomic file replacement does not guarantee transaction isolation across concurrent writers.
Two workers may independently read the same initial index and successfully commit incompatible updates.
This is a classic read-modify-write lost-update race.
Suggested Fix
Introduce cross-process synchronization protecting the complete OCI index update transaction.
Recommended approach:
- Introduce a portable advisory file lock associated with the OCI layout.
- Acquire the lock before reading the existing
index.json.
- Keep the lock while updating the manifest collection.
- Atomically commit the updated index.
- Release the lock after committing or failing the operation.
Important: A process-local mutex alone would not prevent separate host processes from concurrently modifying the same OCI layout.
Retain the existing atomic replacement mechanism for crash consistency.
Suggested Regression Tests
Add a deterministic regression test that:
- Creates two valid snapshots with different tags.
- Uses a shared OCI layout directory.
- Synchronizes concurrent workers so both initially observe the same index state.
- Executes the competing save operations.
- Verifies that every successfully committed snapshot remains referenced.
- Loads both snapshots by tag to verify recoverability.
Additional tests should cover:
- Concurrent writes using different tags.
- Concurrent updates targeting the same tag.
- Multiple host processes accessing the same OCI layout.
- Proper lock cleanup after failed saves.
Production Impact
Severity: High — potential loss of snapshot discoverability and checkpoint-recovery guarantees.
Production systems may rely on Snapshot::save() returning successfully as evidence that a checkpoint has been durably registered.
Under concurrent writes, losing the index reference can violate this assumption.
Potential consequences:
- Checkpoint loss: A snapshot is no longer discoverable by its expected tag.
- Recovery failures: Applications cannot restore the expected checkpoint using tag-based lookup.
- Orphaned blobs: Snapshot contents may exist without index references.
- Operational inconsistency: Successful saves do not necessarily correspond to recoverable snapshots.
- Multi-tenant reliability risks: Concurrent snapshot persistence may interfere across workers using the same OCI layout.
The impact is particularly important for long-running infrastructure where reliable checkpoint restoration is a production requirement.
Related Issues
Issue #1865: Support OCI bundles containing multiple independently restorable snapshots.
#1865
This issue addresses the representation and reachability of related snapshots in OCI bundles.
The proposed bug concerns concurrent modification of index.json and is a distinct correctness problem.
No directly matching existing issue or pull request was identified in the searches performed.
Source Reference
https://github.com/hyperlight-dev/hyperlight/blob/dcb53c04a5a7ae9dc9bf6a8cd54d56067dea138e/src/hyperlight_host/src/sandbox/snapshot/file/mod.rs
Suggested priority: High for workloads that perform concurrent snapshot persistence into shared OCI layouts.
What happened?
Snapshot::save()appears vulnerable to a lost-update race condition when multiple threads or processes concurrently save snapshots with different tags into the same OCI layout directory.The implementation performs a read-modify-write operation on
index.jsonwithout synchronization across concurrent writers.When two workers read the same original index before either writes its update, both independently construct a modified index and atomically replace the existing file.
The last writer can overwrite the first writer's manifest entry.
Consequently, both snapshot saves may return
Ok(...), while one of the snapshots becomes undiscoverable through its tag.This can compromise snapshot reliability in production systems using Hyperlight for:
Important: This is a code-inspection finding with a deterministic race scenario. An end-to-end Hyperlight runtime reproduction is still pending.
Steps to Reproduce
1. Check out Hyperlight's
mainbranch and create two validSnapshotinstances using the existing snapshot test infrastructure.2. Create an empty temporary OCI layout directory shared by both snapshots.
3. Start two concurrent workers executing the following operations:
Use different OCI tags, such as
worker-aandworker-b.4. To reproduce deterministically, add a test-only synchronization barrier immediately after the existing index has been read into the
manifestscollection.Affected file:
src/hyperlight_host/src/sandbox/snapshot/file/mod.rsApproximate location: Lines 432–449.
Ensure that both workers read the same original index before either proceeds to update it.
5. Release both workers and allow their
Snapshot::save()calls to complete.6. Inspect the final OCI index:
jq -r \ '.manifests[].annotations["org.opencontainers.image.ref.name"]' \ /path/to/shared/oci-layout/index.json7. Verify that both
worker-aandworker-bremain discoverable.8. Attempt to load each snapshot by its respective tag.
Expected problematic interleaving:
Reproduction status: The interleaving is supported by source inspection. A direct Hyperlight integration test has not yet been executed.
Expected Results
When multiple snapshots are successfully saved using different tags into the same OCI layout:
index.json.Actual Results
Source inspection indicates that concurrent snapshot saves can overwrite each other's OCI index updates.
The implementation currently:
index.json.index.json.No cross-writer synchronization is visible around the entire read-modify-write operation.
Under the described concurrent interleaving:
Actual runtime results: Not yet collected. These are the predicted outcomes of the identified race.
Versions and Environment
Hyperlight version or commit:
hyperlight-dev/hyperlightmaindcb53c04a5a7ae9dc9bf6a8cd54d56067dea138ehyperlight_hostsrc/hyperlight_host/src/sandbox/snapshot/file/mod.rsOS Version
Target environment: Linux.
Run:
cat /etc/os-release && uname -aActual output: Not yet collected.
The underlying read-modify-write race is not inherently Linux-specific.
Hypervisor
Run the following commands to check hypervisor access:
Actual hypervisor output: Not yet collected.
Note: Hypervisor access may be required to construct snapshots for an integration reproduction. The identified race occurs in host-side OCI index persistence.
Extra Info
Root Cause Analysis
Affected file:
src/hyperlight_host/src/sandbox/snapshot/file/mod.rsRelevant code locations:
index.json.Root cause:
The
replace_file_atomic()helper provides atomic replacement of the index file.However, atomic file replacement does not guarantee transaction isolation across concurrent writers.
Two workers may independently read the same initial index and successfully commit incompatible updates.
This is a classic read-modify-write lost-update race.
Suggested Fix
Introduce cross-process synchronization protecting the complete OCI index update transaction.
Recommended approach:
index.json.Important: A process-local mutex alone would not prevent separate host processes from concurrently modifying the same OCI layout.
Retain the existing atomic replacement mechanism for crash consistency.
Suggested Regression Tests
Add a deterministic regression test that:
Additional tests should cover:
Production Impact
Severity: High — potential loss of snapshot discoverability and checkpoint-recovery guarantees.
Production systems may rely on
Snapshot::save()returning successfully as evidence that a checkpoint has been durably registered.Under concurrent writes, losing the index reference can violate this assumption.
Potential consequences:
The impact is particularly important for long-running infrastructure where reliable checkpoint restoration is a production requirement.
Related Issues
Issue #1865: Support OCI bundles containing multiple independently restorable snapshots.
#1865
This issue addresses the representation and reachability of related snapshots in OCI bundles.
The proposed bug concerns concurrent modification of
index.jsonand is a distinct correctness problem.No directly matching existing issue or pull request was identified in the searches performed.
Source Reference
https://github.com/hyperlight-dev/hyperlight/blob/dcb53c04a5a7ae9dc9bf6a8cd54d56067dea138e/src/hyperlight_host/src/sandbox/snapshot/file/mod.rs
Suggested priority: High for workloads that perform concurrent snapshot persistence into shared OCI layouts.