Commit Graph

211 Commits

Author SHA1 Message Date
Sebastien Boeuf
4eb49e6ea8 ci: Add integration tests for snapshot preserving the source VM
Add a test for verifying that snapshotting a VM while preserving the
source VM works as expected.

Signed-off-by: Sebastien Boeuf <sboeuf@meta.com>
Assisted-by: Claude:claude-opus-4-8
2026-07-24 16:40:39 +00:00
Wei Liu
db4fbeee4b tests: add a memory prefault=on test case
Assisted-by: Opencode:GPT-5.6-Sol
Signed-off-by: Wei Liu <liuwe@microsoft.com>
2026-07-24 07:46:51 +00:00
Sebastien Boeuf
e37f63282c net_util: Only set host MAC address from user input
In case the host MAC address associated with a TAP device wasn't
explicitly provided by the user, Cloud Hypervisor would get the host MAC
associated by default with this TAP device and store it through the
network config. Problem is, in the context of a snapshot/restore, that
meant the network config provided by the user was different on the
destination host compared to the source host. This was causing an issue
when Cloud Hypervisor wasn't started with CAP_NET_ADMIN permissions as
it couldn't set the host MAC address on the destination, while the
source never needed these permissions since the MAC was automatically
allocated by the kernel.

We're fixing this issue by setting the host MAC address when it's
explicitly requested by the user through the network config, and making
the host MAC immutable so that it can't be changed at runtime.

Signed-off-by: Sebastien Boeuf <sboeuf@meta.com>
2026-07-18 09:42:40 +00:00
Wei Liu
4fe133d2bd tests: add KDNET over virtio-net integration test
Add a Windows integration test that verifies kernel network debugging
(KDNET) works over a Cloud Hypervisor virtio-net device.

The test boots a Windows guest with a dedicated second virtio-net NIC,
enables KDNET on it via bcdedit (selecting the adapter by the PCI bus
params discovered over SSH), reboots, and then listens on the debugger
host address. Receiving a KDNET poll datagram from the debuggee proves
the whole virtio-net device path works: discovery, feature negotiation,
virtqueue setup and the TX doorbell. No debugger is needed because KDNET
connections are initiated by the target.

Gated to x86-64, where the Windows image ships the virtio-net KDNET
module. The test exercises only the generic virtio-net doorbell path,
so it runs under both KVM and MSHV.

Signed-off-by: Wei Liu <liuwe@microsoft.com>
Assisted-by: Copilot:Opus-4.8
2026-07-09 21:06:20 +00:00
Bo Chen
595a24d270 build: Mark vfio runner as required for MQ
Across the last 20 MQ runs, all 13 vfio runner failures came from two
flaky tests. Both are now skipped and tracked in #8548 and #8549.

Signed-off-by: Bo Chen <bchen@crusoe.ai>
2026-07-08 17:55:35 +00:00
Rob Bradford
19289a3b82 tests: Add integration tests to snapshot after/during restore
Check that we can make a successful snapshot (and restore it) after
another restore. Also check that snapshot it refused until restore is
complete.

Assisted-by: Claude:Opus-4.8
Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-07-08 17:30:26 +00:00
Anirudh Rayabharam
0b150ea560 tests: add integration test for PPTT cache topology
Add an integration test to verify the newly added code to expose
cache topology information in PPTT.

Signed-off-by: Anirudh Rayabharam <anrayabh@microsoft.com>
2026-07-06 16:31:21 +00:00
Henry Hrvoje Tonkovac
f56fa3a865 tests: trim qualified paths in integration
Import the modules used in the integration tests instead of spelling
the full paths at every use site, and drop the file's now-unnecessary

Signed-off-by: Henry Hrvoje Tonkovac <htonkovac@gmail.com>
Assisted-by: Claude:Opus-4.8
2026-06-24 15:13:07 +00:00
Sebastien Boeuf
80958acdab vmm: Wire postcopy live migration from source VM
Wire up the source side of postcopy migration over TCP. When
`mode=postcopy` is requested on vm.send-migration, the source skips
the pre-copy dirty-tracking loop and lets the destination resume early,
then serves guest pages on demand over a dedicated connection.

Signed-off-by: Sebastien Boeuf <sboeuf@meta.com>
Assisted-by: Claude:claude-opus-4-7
2026-06-24 12:51:40 +00:00
Sebastien Boeuf
60398f11ff offload_daemon: Add --ondemand restore mode
Add an --ondemand flag to the offload daemon's restore subcommand to
support the post-copy mechanism from the live migration protocol.

In on-demand mode, the daemon creates empty memfds to back the guest
memory and sends them over to the VMM. This lets the VM start quickly,
right after the memfds are mapped into CH's address space.

At runtime, when the guest accesses a page (or the prefault handler
requests it), the daemon faults it in by copying the page content into
its shared memory mapping, then replies to the PageFault request so the
VMM can consider the page present.

Signed-off-by: Sebastien Boeuf <sboeuf@meta.com>
Assisted-by: Claude:claude-opus-4-7
2026-06-24 12:51:40 +00:00
Philipp Schuster
6a16b65ea6 tests: adjust to new dispatch semantics of ch-remote send-migration
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
2026-06-22 19:11:42 +00:00
Anatol Belski
dd2f18e73e tests: Cover direct IO data disks on 4k sector FS
Add a parameterized helper that creates a 1 GiB ext4 loop filesystem
with 4096 byte sectors, populates it with a small data disk in the
requested format, attaches that disk with direct=on, and runs a 4096
byte aligned dd round trip with oflag=direct and iflag=direct
followed by cmp.

Wrappers exercise raw, qcow2, fixed VHD, and vhdx. The qcow2 and vhdx
wrappers expect the guest to see the on disk LBS of 512. The others
expect the host LBS of 4096.

Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
2026-06-20 11:46:17 +00:00
Atish Patra
eb1c64e4f0 tests: integration: assert same-host pause/resume keeps aarch64 clock
Add an aarch64 test that pauses a running VM, waits out an interval, and
resumes it on the same host, then asserts the guest wall clock still
matches the host. On aarch64 the architected counter free-runs across
the pause, so the guest self-corrects.

The downtime and skew tolerance are shared with the snapshot clock test.
x86_64 has its own kvmclock path and is covered by the snapshot clock
test.

Signed-off-by: Atish Patra <atishp@meta.com>
2026-06-18 22:59:37 +00:00
Atish Patra
eb2dc28edc tests: integration: assert the guest clock catches up across restore
Add a variation of _test_snapshot_restore that, after taking a snapshot,
waits out a simulated off-host interval and then restores and resumes,
asserting that the guest's wall clock has caught up to the host. This
exercises the clock catch-up that each architecture provides on restore:
kvmclock (KVM_CLOCK_REALTIME) on x86_64 today, and the CNTVCT advance on
aarch64 with later commits.

On x86_64 the guest is booted with clocksource=kvm-clock as the guest
clock is caught up after pause/resume only in that mode. A
tsc-clocksource guest's restored TSC freezes across the interval and
would never catch up.

Take this opportunity to improve the snapshot restore test as the
existing bare boolean mechanism was bit hard to read with new test.

Signed-off-by: Atish Patra <atishp@meta.com>
2026-06-18 22:59:37 +00:00
Sebastien Boeuf
6a74021ad5 ci: Add integration test for offload snapshot
Signed-off-by: Sebastien Boeuf <sboeuf@meta.com>
Assisted-by: Claude:claude-opus-4-7
2026-06-18 13:45:36 +00:00
Tushar Khatri
510aa438f8 tests: reevaluate #[allow] attributes
Convert the still-needed #[allow]s to #[expect] so they warn if the
lints stop firing.

Part of #8326.

Signed-off-by: Tushar Khatri <hello@tusharkhatri.in>
2026-06-17 17:14:00 +00:00
Rob Bradford
2bc968ba1d build: Deny clippy::absolute_paths
Removal of absolute paths is currently in progress. To avoid regressing
those changes add a clippy deny at the workspace level and at the crate
level override with #[expect(clippy::absolute_paths)]

See: #7670

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-06-17 14:38:25 +01:00
Bo Chen
ca2f847e5f tests: Add integration test for FD-based VFIO device
This also covers the usage of pre-opened iommufd FD.

Signed-off-by: Bo Chen <bchen@crusoe.ai>
Assisted-by: Claude:Opus-4.7
2026-06-17 08:39:43 +00:00
Wei Liu
b51dfec09c tests: re-enable some MSHV tests
Signed-off-by: Wei Liu <liuwe@microsoft.com>
2026-06-16 07:06:47 +00:00
Sebastian Eydam
58baee16ac vmm: add TLS API option to receive migration call
As we now have more than one parameter for the receive migration call,
this commit also adds parsing and validation for those parameters. We
maintain backwards compatibility by also correctly parsing the case
where the caller only provides a URL.

On-behalf-of: SAP sebastian.eydam@sap.com
Signed-off-by: Sebastian Eydam <sebastian.eydam@cyberus-technology.de>
2026-06-12 10:11:56 +00:00
Wei Liu
027b1a4c46 tests: enable MSHV fw_cfg coverage
Re-enable the fw_cfg integration tests for MSHV now that port string I/O
is handled by the hypervisor backend.

Signed-off-by: Wei Liu <liuwe@microsoft.com>
Assisted-by: Copilot:GPT-5.5
2026-06-10 20:34:45 +00:00
Ian Klemm
e8e532faf3 tests: exercise memory reserve on the hugepage UFFD restore zone
Turn reserve=on for the hugepage-backed memory zone in the UFFD
snapshot/restore integration test. Hugepages are the most likely place
to want reserve (an over-committed huge page pool is exactly the case
that otherwise SIGBUSes the guest), so this is the natural test to give
the option real coverage, as suggested in review.

It exercises the reserve mmap path twice: once on the source VM boot and
once on the demand-paged restore. The existing skip guard already
requires the 256 free 2MiB pages this zone needs, and the source VM is
killed before the restore VM is started, so reserving from the pool
never has to back two VMs at once.

Assisted-by: Claude Code (Opus 4.8)
Signed-off-by: Ian Klemm <hi@ianklemm.de>
2026-06-10 12:30:25 +00:00
Max Makarov
7f6df9e870 tests: drain the replayed serial backlog and stop the pty echo loop
With socket serial output now buffered and replayed on connect, a
late-connecting client receives the whole boot backlog. The pty
interaction test had three problems with that:

- pty_read() slept a second between 512-byte reads and the loop consumed
  one chunk per two-second tick, far too slow to drain the backlog. Read
  in larger chunks without the per-read sleep and drain everything
  available each round; bound the loop so a missing marker can't run to
  the harness timeout.

- it wrote the login keystrokes before reading, so the unread backlog
  back-pressured the sender and the keystrokes never reached the prompt.
  Start reading concurrently with typing instead.

- the socat pty was created with echo on, so the replayed backlog was
  echoed back to the guest as serial input, flooding it (UART input
  overrun, login never completing). Create the pty with echo=0.

Signed-off-by: Max Makarov <maxpain@linux.com>
Assisted-by: Claude:claude-opus-4-8 [Claude Code]
2026-06-08 10:21:04 +00:00
Philipp Schuster
5aa0587f2a vmm: make PCI BDF configurable for balloon
Add shared PCI config to virtio-balloon.

On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
2026-06-03 13:54:56 +00:00
Wei Liu
d92e1ea77b tests: add Windows TPM integration test
Boot Windows with vTPM enabled and verify the TPM device enumerates
after the guest is reachable.

Assisted-by: Copilot:GPT-5.5
Signed-off-by: Wei Liu <liuwe@microsoft.com>
2026-06-02 09:18:44 +00:00
Wei Liu
4527ae449b tests: exercise TPM after reboot
Extend test_tpm to issue random, fixed-property, PCR read, and PCR
event commands before and after a guest reboot.

Assisted-by: Copilot:GPT-5.5
Signed-off-by: Wei Liu <liuwe@microsoft.com>
2026-06-02 09:18:44 +00:00
Wei Liu
dfcc02f547 tests: reenable TPM test for MSHV
Signed-off-by: Wei Liu <liuwe@microsoft.com>
2026-06-01 17:46:00 +00:00
Philipp Schuster
b241084d0e vmm: migration: add vm.migration-receive-ready event
The new event allows management software to handle the migration better
via events. The `vm.migration-receive-ready` event tells that the VMM is
ready to accept connections whereas `vm.migration-receive-started` means
a migration is incoming.

On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
2026-05-29 13:38:27 +00:00
Leander Kohler
3b79503e2f tests: cover new SMBIOS platform fields
The structured SMBIOS platform config introduced six new keys
(system_manufacturer, system_product_name, system_version,
system_family, system_sku_number, chassis_asset_tag), but
integration coverage only existed for serial_number, uuid, and
oem_strings.

Add _test_dmi_system_and_chassis, which boots a guest with all
six keys set and checks each value via `dmidecode -s` using the
same leaf name as the CLI key. Execute it in both the regular
and SEV-SNP integration suites.

On-behalf-of: SAP leander.kohler@sap.com
Signed-off-by: Leander Kohler <leander.kohler@cyberus-technology.de>
2026-05-28 16:32:41 +00:00
Rob Bradford
29f392f8d4 tests: Fix clippy: uninlined_format_args
Replace format arguments with inlined versions.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-27 16:50:41 +00:00
Rob Bradford
eca621c4e6 tests: Fix clippy: useless_borrows_in_formatting
Replace use of redundant & (leading to &&) in format strings.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-27 16:50:41 +00:00
Cameron Baird
f73eb3ef91 ci: Add integration test for virtio-rtc
Integration test for virtio-devices/src/rtc.rs. Requires that
the test kernel has the following configs:

CONFIG_PTP_1588_CLOCK=y
CONFIG_VIRTIO_RTC=y
CONFIG_VIRTIO_RTC_PTP=y

Signed-off-by: Cameron Baird <cameronbaird@microsoft.com>
2026-05-26 20:00:24 +00:00
Philipp Schuster
1924153185 tests: cover restored VM disk hotplug after TCP migration
Extend the TCP live-migration test with a hotplugged block device using
a stable ID before migration. After migration, verify that the disk
still exists on the restored destination VM.

Then hot-remove the disk and add it again with the same ID. This covers
the stale restore snapshot case because the disk ID exists in the
migration snapshot, but the live device tree no longer contains it after
hot-remove.

Assisted-by: Codex:GPT-5.5
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
2026-05-21 15:49:51 +00:00
Muminul Islam
fba55b3d9f tests: poll for source VM exit after live-migration
The post-migration check used a fixed `thread::sleep(3s)` followed by
`try_wait()` to verify the source VM had exited cleanly. That window
is too tight when the source process is the release binary used by
`test_live_upgrade_*` (i.e. `~/workloads/cloud-hypervisor-static`,
pinned to `migratable_version`).

The released binary is older than the locally-built destination and
its virtio-device teardown (resume-paused-thread -> kill -> join
across pmem, block, net, console, rng workers) regularly takes
longer than 3s on contended hosts, causing the test to report:

  thread 'common_parallel::test_live_upgrade_basic' panicked:
  Test failed: source VM was not terminated successfully.

even though the source process eventually exits with status 0.

Replace the fixed sleep with a `wait_until(Duration::from_secs(30),
...)` poll that returns as soon as `try_wait()` reports a reaped
child, then keep the existing `success()` check on the exit status.
This makes the assertion robust against the slower release-binary
shutdown path while still failing fast on a genuine error.

The same pattern was duplicated across eight migration helpers plus
the virtio-fs migration variant; convert all nine call sites for
consistency.

Signed-off-by: Muminul Islam <muislam@microsoft.com>
2026-05-16 00:19:45 +00:00
Wei Liu
7d7f24382c tests: add block device integration tests
Assisted-by: Claude:Opus-4.7
Signed-off-by: Wei Liu <liuwe@microsoft.com>
2026-05-14 22:35:02 +00:00
Philipp Schuster
b0f92d01fc tests: avoid event wait log spam
Event expectation helpers print detailed diagnostics when the observed
event stream does not match the expected one. That is useful for direct
assertions, but it becomes extremely noisy [0] when the helper is used
as the predicate for wait_until(), because every polling attempt emits
the full mismatch dump.

Add quiet wait wrappers for event polling and emit the existing detailed
diagnostics only once after the timeout expires.

[0] https://github.com/cloud-hypervisor/cloud-hypervisor/actions/runs/25745401604/job/75619840718?pr=8021

On-behalf-of: Philipp Schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
2026-05-13 22:16:29 +00:00
Muminul Islam
96c83168a1 tests: Avoid IPv6 autoconf races in test_vdpa_net
The test_vdpa_net integration test brings the vDPA-backed interface
(ens6) up and then asserts that both TX and RX packet counters are
exactly zero before sending an explicit ping. On guest kernels that
perform IPv6 link-local autoconfiguration quickly enough, however,
Router Solicitation / Neighbor Discovery frames are emitted as soon
as the link comes up. The vdpa_sim_net device loops those frames back
to the interface, so by the time the test queries

    ip -j -p -s link show ens6 | grep -c '"packets": 0'

the TX and RX counters are already non-zero and the precondition
assertion fails (observed reliably with the Microsoft internal guest
kernel running on MSHV).

Disable IPv6 / accept_ra / autoconf on ens6 before bringing the link
up. With IPv6 disabled no autoconf traffic is generated, the counters
remain at zero until the explicit 'ping 172.16.1.10 -c 6' generates
exactly the 6 packets the rest of the test expects on each direction,
and the vDPA-specific portion of the test is unchanged.

Verified on an MSHV Azure VM (Linux 6.6.121.mshv2):

    test common_parallel::test_vdpa_net ... ok
    test result: ok. 1 passed; 0 failed; ...; finished in 26.13s

Signed-off-by: Muminul Islam <muislam@microsoft.com>
2026-05-13 22:15:19 +00:00
Bo Chen
3645d916f3 tests: Stabilize VFIO NVIDIA memory hotplug check
Bump the wait_until timeout for the guest-visible memory growth in
test_nvidia_card_memory_hotplug from 5s to 15s. The hot-add path inside
the guest kernel can take longer than 5s particularly when using
virtio-mem, which has been a source of flakes for this test.

Drop the trailing assert!(guest.get_total_memory() > 5_760_000), as it
is redundant.

See: #8160

Signed-off-by: Bo Chen <bchen@crusoe.ai>
2026-05-13 21:14:47 +00:00
Philipp Schuster
0e67f27546 tests: extend test_pci_device_id() to also test virtio-console
On-behalf-of: Philipp Schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
2026-05-11 20:18:03 +00:00
Rob Bradford
e0e10b5971 tests: Disable tests on ARM64 that don't work with stock kernel
An edk2 boot is needed for ACPI support but the stock ARM64 kernel does
not support ACPI memory hotplug or virtio-pmem. For now disable those
tests but track them in #8187.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-11 17:52:00 +00:00
Rob Bradford
c392274f09 tests: Streamline watchdog tests
This is a niche feature and we were overly testing it. Let's just switch
to two tests. One for live migration and one for plain watchdog. This
will reduce the CI time. As we are now running these tests sequentially
we can also reduce some the delays in the tests.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-11 17:52:00 +00:00
Rob Bradford
0f7e6a0d3a tests: Move watchdog tests to sequential
These tests are very timing dependent and so need to be run without high
levels of load.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-11 17:52:00 +00:00
Rob Bradford
55adcc27da tests: Replace last vestiges of focal use with jammy
Focal has served us well for many years but is now beyond EOL. Remove
all remaining use of focal images from the CI.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-11 17:52:00 +00:00
Keith Adler
3db4c0f7b4 tests: remove redundant hotplug sleep
The snapshot/restore hotplug path already waits for the exact
device-removed event through the event monitor. Drop the fixed sleep
before that poll so the test advances as soon as the event arrives.

Signed-off-by: Keith Adler <kadler@cloudflare.com>
2026-05-06 20:49:56 +00:00
Damian Barabonkov
4eb1717fb0 tests: Add VFIO mmap BAR exclusion coverage
Exercise the new VFIO BAR exclusion option with NVIDIA
passthrough tests so the integration suite checks that selected
BARs are skipped.

The tests cover both legacy VFIO and iommufd paths while
preserving the existing hardware availability guards.

Signed-off-by: Damian Barabonkov <dbctl@pm.me>
Assisted-by: OpenCode:gpt-5.5
2026-05-06 14:15:41 +00:00
Rob Bradford
12f48700a2 tests: Wait for boot notification from test_vfio_user L2 guest
Reuse the boot notification method we have for the L1 guests for the L2
guest. This removes the need to use SSH based boot tracking for
connecfting to the L2 guest and should make the test more reliable.

This requires making the L2 guest use a different cloud-init
configuration to the L1.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-06 12:30:08 +00:00
Rob Bradford
7d8986aad0 tests: Allow more time for firmware & O_DIRECT tests
Booting the VM on these tests takes longer so allow longer before
timing out the boot.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-05-04 21:28:33 +00:00
Nguyen Dinh Phi
c5951252a5 tests: Adding integration tests for migration of paused VM
Adding a paused flag to live_migration() tests; when this
flag is set, the VM will be paused before migration is
performed.

Signed-off-by: Nguyen Dinh Phi <phind.uet@gmail.com>
2026-05-04 09:30:30 +00:00
Rob Bradford
92229a60ed tests: Report stderr/stdout from restored child in test_ovs_dpdk
To aid debugging of this test failing print the output from the restored
VMM instance too.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-04-27 07:15:37 +00:00
Rob Bradford
466f50e72c tests: parallelise live-migration virtio-fs tests
`test_live_migration_virtio_fs` and its `_local` variant lived in
`common_sequential` because they shared `~/workloads/shared_dir` as
the virtiofsd backing and wrote/deleted the same `migration_test_file`
and `post_migration_file` paths inside it. Two instances running
concurrently would race on those files.

Give each invocation its own backing directory under `guest.tmp_dir`,
which is already per-test unique and gets cleaned up by the `TempDir`
drop. The test logic is otherwise unchanged. Move both wrappers and
the helper from `common_sequential` to `common_parallel` and update
the sequential-tests comment accordingly.

The boot footprint is small (512 MB src + 512 MB dest), so two
concurrent instances comfortably fit alongside the rest of the
parallel suite. The remaining sequential live-migration tests
(balloon, NUMA) genuinely need their isolation slot for memory
headroom.

Signed-off-by: Rob Bradford <rbradford@meta.com>
2026-04-27 07:13:16 +00:00