Skip to content

[do not merge] resident memory experiment - #1901

Open
danbugs wants to merge 4 commits into
mainfrom
resident-scratch-reset
Open

danbugs wants to merge 4 commits into
mainfrom
resident-scratch-reset

Conversation

@danbugs

@danbugs danbugs commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

No description provided.

Copilot AI balanced review requested due to automatic review settings October 7, 2026 23:57
@danbugs danbugs added kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. area/performance Addresses performance labels Oct 7, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

The pagemap-driven reset policy affects sandbox isolation and needs fixes plus human validation.

3 open findings
What changed in this PR

This experiment targets faster KVM snapshot restores by retaining and zeroing selected scratch pages in Hyperlight’s host memory manager.

Changes:

  • Adaptive pagemap-based scratch resets with platform fallbacks.
  • Per-sandbox reset state, memory-behavior tests, and changelog documentation.
File Description
src/​hyperlight_host/​src/​mem/​shared_mem.rs Selects adaptive KVM resets with zeroing fallbacks.
src/​hyperlight_host/​src/​mem/​scratch_reset.rs Implements residency tracking, reset policies, and tests.
src/​hyperlight_host/​src/​mem/​mod.rs Provides platform-gated reset modules.
src/​hyperlight_host/​src/​mem/​mgr.rs Maintains reset state across scratch restores.
CHANGELOG.md Documents retained KVM scratch memory.

🧠 Review effort: Balanced


Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.

Comment thread src/hyperlight_host/src/mem/scratch_reset.rs
Comment thread src/hyperlight_host/src/mem/scratch_reset.rs Outdated
Comment thread src/hyperlight_host/src/mem/scratch_reset.rs Outdated
@danbugs
danbugs force-pushed the resident-scratch-reset branch from 74aa786 to b907442 Compare October 8, 2026 00:11
@hyperlight-gh-bot

This comment has been minimized.

@danbugs
danbugs force-pushed the resident-scratch-reset branch from d278d83 to 592e155 Compare October 8, 2026 20:40
@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@danbugs
danbugs force-pushed the resident-scratch-reset branch 2 times, most recently from 7a2bee3 to cb8bbaf Compare October 9, 2026 00:37
@hyperlight-gh-bot

This comment has been minimized.

A restore zeroed scratch by dropping it with MADV_DONTNEED on KVM (fill(0)
in builds with mshv3), by fill(0) on MSHV, and by mapping fresh memory on
Windows (#1765). Each costs time in the size of scratch, or a fault on
every page the next run touches.

On KVM, every reset now reads /proc/self/pagemap for all of scratch. A
page with memory of its own that holds data is zeroed and kept, so the
next run takes no faults on it. One that holds only zeros was not written
since the last reset and is dropped, so what stays resident follows what
runs write. Swapped pages and shared pages holding data are dropped. The
rule is the same from the first restore on. Scratch is backed 4 KiB at a
time on KVM, so a reset zeroes only what was touched.

On Windows, scratch of up to 16 MiB is now zeroed in place, which is about
2x faster than a fresh mapping (over 5x when the guest writes a MiB or
more); larger scratch is still replaced, so it does not all become
resident.

HostSharedMemory now notes the pages its writes touch, and
zero_written zeroes only a given set of pages plus those, for resets that
know what the guest wrote (next commits).

Signed-off-by: danbugs <danilochiarlone@gmail.com>
Add a DirtyLog supertrait of VirtualMachine: how a VM tracks the pages
the guest writes, and reading and clearing that log for a range.

On MSHV (x86_64), tracking is a partition property switched on and off.
Reads use get_dirty_log with clear; stopping sets the bits back first,
which MSHV requires.

On WHP (x86_64 and ARM64), scratch is mapped with
WHvMapGpaRangeFlagTrackDirtyPages and read with
WHvQueryGpaRangeDirtyBitmap, which clears what it reads. If the tracked
mapping fails, scratch is mapped as before and tracking is not tried
again for the VM.

Other backends do not track. A backend that claims to but does not read
reports every page, so a reset zeroes too much, never too little.

Signed-off-by: danbugs <danilochiarlone@gmail.com>
Where the hypervisor logs the guest's writes, a restore now zeroes only
the pages written since the last restore: those in the log, the pages
host writes touched, and the page-table copy (#1766).

Tracking starts when scratch is mapped, before the guest runs, so the
first restore uses it too (hyperlight_vm/dirty_log.rs):

- WHP: always. Tracking costs the guest nothing measurable there. The
  log is cleared when scratch is mapped, so only written pages are ever
  touched and the rest of scratch never becomes resident.
- MSHV: tracking makes the guest's first write to each page after a read
  fault to the hypervisor (about 1 us nested), and reading the log costs
  time in the size of scratch. Measured against zeroing all of scratch,
  that loses below 2 MiB of scratch, and once a run writes more than about
  a tenth of scratch (a quarter from 32 MiB, where zeroing gets about 3x
  slower per MiB). So scratch under 2 MiB is not tracked, and tracking
  stops after such a run until scratch is mapped again. The first log
  after mapping reports every page, so the guest sets up without faults.
- KVM keeps its in-place reset; no log is read.

Signed-off-by: danbugs <danilochiarlone@gmail.com>
The restore benchmarks use 352 KiB to 1 MiB of scratch, where resetting
it costs little however it is done. Add call_with_restore with 64 and
256 MiB of scratch, with a guest that writes a few pages (Echo) and one
that writes 1 MiB a call, and report them, with the default size, in
pull request comments.

Signed-off-by: danbugs <danilochiarlone@gmail.com>
@danbugs
danbugs force-pushed the resident-scratch-reset branch from e4d8f03 to e53c815 Compare October 11, 2026 00:06
danbugs added a commit to hyperlight-dev/hyperlight-unikraft that referenced this pull request Oct 11, 2026
Patch hyperlight-host and hyperlight-common to danbugs/hyperlight
c854bb43 (0.17.0 plus the scratch reset of hyperlight-dev/hyperlight#1901),
to measure it in CI.

Signed-off-by: danbugs <danilochiarlone@gmail.com>
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: e53c815a0afb
Baseline commit: dcb53c04a5a7

kvm / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 791.63 ns (➖ 1.04x slower)
vec_bytes 578.07 ns (➖ 1.02x faster)
373.21 µs (➖ 1.00x slower)

guest_calls

call_with_restore
scratch_256mib 117.78 µs (---)
default 77.05 µs (➖ 1.02x slower)
scratch_64mib_write_1mib 370.96 µs (---)
scratch_64mib 98.59 µs (---)
scratch_256mib_write_1mib 384.63 µs (---)

payload_allocation

slot_pool_segmented
262144 522.39 ns (➖ 1.00x faster)
65536 143.72 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 82.50 ms (➖ 1.04x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.81 ns (➖ 1.01x slower) 7.94 ns (➖ 1.03x slower) 7.79 ns (➖ 1.03x faster)

snapshot_files

load_snapshot_unverified
small 92.85 µs (➖ 1.01x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.29 µs (➖ 1.02x slower) 7.30 µs (➖ 1.04x faster)
65536 2.06 µs (➖ 1.01x faster) 1.93 µs (➖ 1.08x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.20 µs (➖ 1.08x faster) 6.24 µs (➖ 1.01x faster)
8192 1.06 µs (➖ 1.03x faster) 1.04 µs (➖ 1.01x slower)
262144 26.66 µs (➖ 1.03x faster)
kvm / intel (Linux) (❌ 2.00x)

Top regressions

  • slot_pool/alloc_dealloc_1500 — ❌ 2.00x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 822.48 ns (➖ 1.22x slower)
vec_bytes 538.23 ns (➖ 1.04x slower)
683.73 µs (➖ 1.02x slower)

guest_calls

call_with_restore
scratch_256mib 66.33 µs (---)
default 37.81 µs (➖ 1.06x faster)
scratch_64mib_write_1mib 310.30 µs (---)
scratch_64mib 42.79 µs (---)
scratch_256mib_write_1mib 376.14 µs (---)

payload_allocation

slot_pool_segmented
262144 509.33 ns (➖ 1.02x faster)
65536 142.29 ns (➖ 1.05x slower)

sandboxes

create_initialized_and_drop
medium 68.99 ms (➖ 1.09x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
14.02 ns (❌ 2.00x slower) 6.95 ns (➖ 1.00x faster) 6.96 ns (➖ 1.01x faster)

snapshot_files

load_snapshot_unverified
small 46.82 µs (➖ 1.02x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.57 µs (➖ 1.01x slower) 7.55 µs (➖ 1.00x faster)
65536 2.12 µs (➖ 1.01x faster) 2.11 µs (➖ 1.00x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.09 µs (➖ 1.05x faster) 7.04 µs (➖ 1.01x faster)
8192 738.55 ns (➖ 1.06x faster) 736.56 ns (➖ 1.03x faster)
262144 29.55 µs (➖ 1.13x faster)
mshv3 / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 984.83 ns (➖ 1.03x slower)
vec_bytes 695.17 ns (➖ 1.05x faster)
377.43 µs (➖ 1.18x slower)

guest_calls

call_with_restore
scratch_256mib 1.67 ms (---)
default 153.23 µs (➖ 1.00x slower)
scratch_64mib_write_1mib 1.14 ms (---)
scratch_64mib 556.97 µs (---)
scratch_256mib_write_1mib 2.51 ms (---)

payload_allocation

slot_pool_segmented
262144 724.81 ns (➖ 1.00x faster)
65536 193.75 ns (➖ 1.04x faster)

sandboxes

create_initialized_and_drop
medium 60.12 ms (➖ 1.07x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.78 ns (➖ 1.00x slower) 9.75 ns (➖ 1.00x slower) 9.70 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 86.14 µs (➖ 1.00x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 10.53 µs (➖ 1.05x slower) 9.81 µs (➖ 1.03x slower)
65536 2.43 µs (➖ 1.00x faster) 2.26 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.45 µs (➖ 1.17x faster) 8.27 µs (➖ 1.16x faster)
8192 1.31 µs (➖ 1.04x faster) 1.29 µs (➖ 1.03x slower)
262144 35.89 µs (➖ 1.05x faster)
mshv3 / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 973.23 ns (➖ 1.04x slower)
vec_bytes 652.55 ns (➖ 1.02x faster)
685.18 µs (➖ 1.13x faster)

guest_calls

call_with_restore
scratch_256mib 1.69 ms (---)
default 126.13 µs (➖ 1.00x slower)
scratch_64mib_write_1mib 1.04 ms (---)
scratch_64mib 483.04 µs (---)
scratch_256mib_write_1mib 2.79 ms (---)

payload_allocation

slot_pool_segmented
262144 623.61 ns (➖ 1.01x slower)
65536 165.65 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 65.24 ms (➖ 1.00x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.64 ns (➖ 1.00x faster) 8.78 ns (➖ 1.00x faster) 8.79 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 44.20 µs (➖ 1.02x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.01 µs (➖ 1.01x slower) 7.98 µs (➖ 1.01x slower)
65536 2.28 µs (➖ 1.02x slower) 2.27 µs (➖ 1.02x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.65 µs (➖ 1.02x slower) 7.44 µs (➖ 1.02x faster)
8192 913.74 ns (➖ 1.02x faster) 914.40 ns (➖ 1.02x slower)
262144 39.16 µs (➖ 1.00x faster)
hyperv-ws2025 / amd (Windows) (🌟 5.47x)

Top improvements

  • guest_calls/call_with_restore/default — 🌟 5.47x faster

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.16 µs (➖ 1.01x faster)
vec_bytes 765.36 ns (➖ 1.03x faster)
2.11 ms (➖ 1.06x slower)

guest_calls

call_with_restore
scratch_256mib 2.73 ms (---)
default 248.71 µs (🌟 5.47x faster)
scratch_64mib_write_1mib 1.87 ms (---)
scratch_64mib 862.39 µs (---)
scratch_256mib_write_1mib 5.35 ms (---)

payload_allocation

slot_pool_segmented
262144 793.92 ns (➖ 1.04x faster)
65536 239.61 ns (➖ 1.04x slower)

sandboxes

create_initialized_and_drop
medium 80.12 ms (➖ 1.04x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.32 ns (➖ 1.04x slower) 10.45 ns (➖ 1.03x slower) 10.35 ns (➖ 1.02x slower)

snapshot_files

load_snapshot_unverified
small 682.30 µs (➖ 1.04x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.01 µs (➖ 1.04x faster) 9.09 µs (➖ 1.01x slower)
65536 2.33 µs (➖ 1.03x faster) 2.37 µs (➖ 1.02x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.83 µs (➖ 1.04x faster) 8.89 µs (➖ 1.01x slower)
8192 1.29 µs (➖ 1.04x faster) 1.28 µs (➖ 1.01x slower)
262144 39.23 µs (➖ 1.03x faster)
hyperv-ws2025 / intel (Windows) (🌟 6.00x)

Top improvements

  • guest_calls/call_with_restore/default — 🌟 6.00x faster

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.21 µs (➖ 1.02x slower)
vec_bytes 760.57 ns (➖ 1.02x faster)
3.22 ms (➖ 1.01x slower)

guest_calls

call_with_restore
scratch_256mib 2.49 ms (---)
default 301.00 µs (🌟 6.00x faster)
scratch_64mib_write_1mib 1.98 ms (---)
scratch_64mib 809.62 µs (---)
scratch_256mib_write_1mib 5.55 ms (---)

payload_allocation

slot_pool_segmented
262144 736.92 ns (➖ 1.01x faster)
65536 213.08 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 108.70 ms (➖ 1.02x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.30 ns (➖ 1.00x slower) 10.25 ns (➖ 1.03x slower) 10.57 ns (➖ 1.00x slower)

snapshot_files

load_snapshot_unverified
small 757.28 µs (➖ 1.54x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.77 µs (➖ 1.01x slower) 7.77 µs (➖ 1.00x slower)
65536 2.33 µs (➖ 1.01x slower) 2.38 µs (➖ 1.02x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.36 µs (➖ 1.03x faster) 8.34 µs (➖ 1.01x faster)
8192 1.11 µs (➖ 1.13x faster) 1.09 µs (➖ 1.01x faster)
262144 41.18 µs (➖ 1.01x slower)

Reported by cargo ci bench-report --candidate run:38097297289 --config-file bench_report.toml.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/performance Addresses performance kind/enhancement For PRs adding features, improving functionality, docs, tests, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants