Set discard_sector_alignment from the logical block size
reported by the backend topology instead of hardcoding it
to 1 sector. This gives the guest accurate alignment hints
so it can avoid sub block discards that the filesystem
might silently ignore.
For example, on a 4K block filesystem the alignment is now
8 sectors (4096/512) instead of 1.
For image formats with their own allocation units (QCOW2
clusters, VHD/VHDX block sizes), the ideal alignment would
be derived from the format cluster/block size. This is
left for a followup that surfaces allocation granularity
through DiskTopology.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Replace magic numeric offsets with mem::offset_of!() referencing the
virtio_blk_discard_write_zeroes struct from the virtio-bindings crate
when reading the sector, num_sectors and flags fields in the DISCARD
and WRITE_ZEROES request handlers.
No functional change.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add an option that can be used when restoring to resume the VM. This is
particularly useful when restoring the VM via the direct VMM command
line, when you might not want/have an API socket configured.
Signed-off-by: Rob Bradford <rbradford@meta.com>
Skip micro_ prefixed tests in the metrics CI workflow to avoid
dashboard pollution. They can still be run on demand via
--test-filter micro_.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add --test-exclude to process_common_args in test-util.sh and forward
it to the performance-metrics binary from run_metrics.sh.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add a --test-exclude flag that excludes tests matching the provided
keywords. Both --test-filter and --test-exclude are now applied before
--list-tests, so listing respects the active filters.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add test_virtio_block_write_zeroes_unmap_raw to verify that the
VIRTIO_BLK_WRITE_ZEROES_FLAG_UNMAP code path works correctly with
raw disk images.
The test creates a 128M raw disk and writes 64M of random data,
then uses fallocate --punch-hole on the guest block device, which
the Linux virtio-blk driver translates to VIRTIO_BLK_T_WRITE_ZEROES
with VIRTIO_BLK_WRITE_ZEROES_FLAG_UNMAP set. It then verifies:
- the zeroed region reads back as zero from the guest
- the host file became sparse (punch_hole succeeded)
- FIEMAP confirms the file has holes
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The write zeroes segment descriptor (struct
virtio_blk_discard_write_zeroes, virtio spec v1.2 section 5.2.6)
includes a flags field with an unmap bit. Per section 5.2.6.2, if
unmap is set, the device MAY deallocate the specified range of
sectors in the device backend storage, as if the discard command
had been sent.
Read the flags field and when the unmap bit is set, use punch_hole
to deallocate the range. Otherwise continue using write_zeroes via
ZERO_RANGE which preserves allocation.
This allows the guest to reclaim host disk space through write
zeroes requests on thin provisioned images.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The virtio vhost-user device backend prefers to use externally-provided
eventfds as irqfds. This allows the frontend VM to notify the backend
VM directly, without the need for a userspace proxy process. Since the
frontend can provide irqfds at any time, the backend needs to register
and unregister irqfds dynamically.
This is tricky because the functions that access the irqfd table all
take `&self`, not `&mut self`. The obvious solution to this problem is
to wrap the table in a mutex. Most of these functions are not called on
hot paths, but `.notifier()` is called whenever Cloud Hypervisor needs
to inject an interrupt into a guest. Most devices don't need to
register irqfds at runtime, and for them, slowing down interrupt
injection would be wasteful.
Instead, require devices to opt-in to irqfd registration. The irqfd
table now comes in two forms: one that contains a mutex and one that
does not. The one containing a mutex can be mutated freely, while
attempting to mutate the one that does not will panic.
Right now, no code registeres irqfds at runtime, but this will change in
subsequent commits.
Signed-off-by: Demi Marie Obenour <demiobenour@gmail.com>
This allows creating an InterruptSourceGroup with an externally provided
file descriptor. It also allows changing the file descriptor
afterwards.
Signed-off-by: Demi Marie Obenour <demiobenour@gmail.com>
It is currently left as unimplemented!().
No functional change intended as there are no callers.
Signed-off-by: Demi Marie Obenour <demiobenour@gmail.com>
The InterruptRoute code tried to be thread-safe, but it wasn't. In
particular, concurrently enabling and disabling an InterruptRoute could
result in the route thinking it was enabled (when it was disabled) or
visa versa.
Wrap all operations in a mutex and drop the attempt at being lock-free.
Signed-off-by: Demi Marie Obenour <demiobenour@gmail.com>
Add micro_block_raw_aio_drain_128_us and
micro_block_raw_aio_drain_256_us tests that submit N AIO writes
to a temporary file, wait for the eventfd signal, then time how
long it takes to drain all completions via next_completed_request().
This measures per completion syscall overhead and provides a
baseline before any batching optimizations.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
These factor out common setup and synchronization patterns used by block
layer micro benchmarks.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add an optional num_ops parameter for micro benchmarks to configure
workload size (e.g. number of AIO operations to submit). A warning
is emitted if it is accidentally set on a non micro test where it
has no effect.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Introduce support for in process micro benchmarks alongside the
existing VM level performance tests. Micro benchmarks are
integrated into the same PerformanceTest/TEST_LIST infrastructure
and follow the same iteration, timeout, and reporting pipeline.
They are distinguished by a micro_* name prefix.
The test dispatch loop is refactored to pre filter the test list
and gate init/cleanup behind a flag, so that pure micro benchmark
runs skip the expensive VM lifecycle entirely. Mixed runs
(VM + micro) continue to work correctly.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Trigger the interrupts in the guest for the virtio device queues behind
the vhost-user devices when resuming. This avoids a situation where
interrupts from the backend get lost when they are dispatched from the
backend when then guest is paused leading to the guest/backend
effectively waiting for each other to move forward. This is more
reproducible with longer durations between pause and resume as there is
more opportunity for the backend to completely process it's queue and
fire all the interrupts.
It's perfectly safe and allowed by the virtio spec to generate these
interrupts and the performance impact is negligible and is a safe way to
ensure forward progress after a resume.
See: #7850
Signed-off-by: Rob Bradford <rbradford@meta.com>
The config space fix in the previous commit correctly populates
the discard and write zeroes fields, so the sparse=off
workaround is no longer needed for Windows guests.
Replace default_disks_sparse_off() with default_disks() in all
Windows test cases and remove the explicit sparse=off from the
multi queue test.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
When VIRTIO_BLK_F_DISCARD or VIRTIO_BLK_F_WRITE_ZEROES features
are advertised, the virtio spec v1.2, sections 5.2.4 and
5.2.6.1, requires the corresponding VirtioBlockConfig fields
to contain valid, non zero values. Leaving them at zero causes
strictly behaved drivers to either reject the features or crash.
Populate max_discard_sectors, max_discard_seg,
discard_sector_alignment, max_write_zeroes_sectors,
max_write_zeroes_seg and write_zeroes_may_unmap after
feature advertisement so drivers can safely negotiate
these features.
Fixes: #7849
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
This device is not called virtio-vhost-user; that's something else.
I don't think the comment really clarifies anything anyway, so just
remove it.
Fixes: 8c618ff5e ("virtio-devices: generic-vhost-user: implement device")
Signed-off-by: Alyssa Ross <hi@alyssa.is>
Rename the aarch64 sha1sums file to sha1sums-aarch64-common to follow
the same naming convention as sha1sums-x86_64-common. This allows
run_metrics.sh to use the generic sha1sums-${TEST_ARCH}-common
pattern for all architectures, removing the need for aarch64-specific
conditionals.
Update run_integration_tests_aarch64.sh to reference the renamed file.
Signed-off-by: Souradeep <schakrabarti@microsoft.com>
warning: this argument is passed by value, but not consumed in the function body
--> cloud-hypervisor/tests/integration.rs:3785:51
|
3785 | fn run_multiqueue_qcow2_test<F>(image_config: QcowTestImageConfig, test_fn: F)
| ^^^^^^^^^^^^^^^^^^^
|
help: or consider marking this type as `Copy`
--> cloud-hypervisor/tests/integration.rs:3774:5
|
3774 | enum QcowTestImageConfig {
| ^^^^^^^^^^^^^^^^^^^^^^^^
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#needless_pass_by_value
= note: requested on the command line with `-D clippy::needless-pass-by-value`
Signed-off-by: Rob Bradford <rbradford@meta.com>
error: the borrowed expression implements the required traits
--> cloud-hypervisor/tests/integration.rs:8510:32
|
8510 | disk_check_consistency(&test_disk_path, None);
| ^^^^^^^^^^^^^^^ help: change this to: `test_disk_path`
|
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#needless_borrows_for_generic_args
= note: `-D clippy::needless-borrows-for-generic-args` implied by `-D clippy::all`
= help: to override `-D clippy::all` add `#[allow(clippy::needless_borrows_for_generic_args)]`
error: could not compile `cloud-hypervisor` (test "integration") due to 7 previous errors
Signed-off-by: Rob Bradford <rbradford@meta.com>
error: you seem to use `.enumerate()` and immediately discard the index
--> cloud-hypervisor/tests/integration.rs:7675:72
|
7675 | for (_i, (offset, length)) in discard_operations.iter().enumerate() {
| ^^^^^^^^^^^^
|
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#unused_enumerate_index
= note: `-D clippy::unused-enumerate-index` implied by `-D clippy::all`
= help: to override `-D clippy::all` add `#[allow(clippy::unused_enumerate_index)]`
help: remove the `.enumerate()` call
|
7675 - for (_i, (offset, length)) in discard_operations.iter().enumerate() {
7675 + for (offset, length) in discard_operations.iter() {
|
Signed-off-by: Rob Bradford <rbradford@meta.com>
error: consider adding a `;` to the last statement for consistent formatting
--> cloud-hypervisor/tests/integration.rs:2516:9
|
2516 | _test_simple_launch(&guest)
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^ help: add a `;` here: `_test_simple_launch(&guest);`
|
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#semicolon_if_nothing_returned
Signed-off-by: Rob Bradford <rbradford@meta.com>
error: variables can be used directly in the `format!` string
--> cloud-hypervisor/tests/integration.rs:12770:27
|
12770 | let driver_path = format!("{}/driver", NVIDIA_VFIO_DEVICE);
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#uninlined_format_args
help: change this to
|
12770 - let driver_path = format!("{}/driver", NVIDIA_VFIO_DEVICE);
12770 + let driver_path = format!("{NVIDIA_VFIO_DEVICE}/driver");
Signed-off-by: Rob Bradford <rbradford@meta.com>
error: this argument is passed by value, but not consumed in the function body
--> net_util/src/tap.rs:685:17
|
685 | ifname: String,
| ^^^^^^ help: consider changing the type to: `&str`
|
= help: for further information visit https://rust-lang.github.io/rust-clippy/master/index.html#needless_pass_by_value
= note: requested on the command line with `-D clippy::needless-pass-by-value`
Signed-off-by: Rob Bradford <rbradford@meta.com>
Write down our policy for git commit hygiene, especially when it comes
to the history, i.e., multiple git commits in a PR.
TL;DR: Commits must be revieable units guiding reviewers how the
developer got from A to B.
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
On-behalf-of: SAP philipp.schuster@sap.com
raw_sync, raw_async, and raw_async_aio each defined
FALLOC_FL_PUNCH_HOLE, FALLOC_FL_KEEP_SIZE, and FALLOC_FL_ZERO_RANGE as
local constants in their punch_hole() and write_zeroes()
implementations. These are available from the libc crate directly.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
probe_file_sparse_support() defined FALLOC_FL_KEEP_SIZE,
FALLOC_FL_PUNCH_HOLE, and FALLOC_FL_ZERO_RANGE as local constants.
These are available from the libc crate directly.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Replace duplicated test bodies with thin wrappers that construct
the backend-specific AsyncIo instance and delegate to the shared
raw_async_io_tests helpers.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add raw_async_io_tests.rs with punch_hole, write_zeroes, and
multiple_operations helpers that take &mut dyn AsyncIo + &mut File.
These are raw-backend-specific. They verify data by reading the
underlying file directly, which only works for plain file backends
without container format metadata.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Split the vDPA preparation flow into helper functions for
building modules, validating availability, loading modules,
and creating devices.
Build the vdpa_sim modules only on Ubuntu, where the script
installs dependencies and compiles them from the matching
kernel source. On other distributions, reuse the installed
kernel modules and verify that they are available before
continuing.
This makes the script easier to follow and supports systems
such as Azure Linux, where the modules are provided by the
kernel package.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
A buggy or malicious guest may write an inappropriate value into
virtqueue's next_avail field. This will result in an error
when iterating over the queue:
863837ef86/virtio-queue/src/queue.rs (L708)
but this error is (logged and) ignored if pop_descriptor_chain()
is used:
863837ef86/virtio-queue/src/queue.rs (L583)
A reasonable approach, implemented here, is to mark the device as
NEEDS_RESET and ignore further queue events until the guest
reinitializes the device.
How this patch was tested:
Linux kernel was patched to trigger a bad next_avail when the
virtqueue queue counter reaches 5000:
--------------- START OF LINUX KERNEL PATCH ----------
$ git diff
diff --git a/drivers/virtio/virtio_ring.c b/drivers/virtio/virtio_ring.c
index b784aab668670..989f2a0c64a77 100644
--- a/drivers/virtio/virtio_ring.c
+++ b/drivers/virtio/virtio_ring.c
@@ -15,6 +15,9 @@
#include <linux/spinlock.h>
#include <xen/xen.h>
+
+void virtqueue_kick_always(struct virtqueue *vq);
+
#ifdef DEBUG
/* For development, we want to crash whenever the ring is screwed. */
#define BAD_RING(_vq, fmt, args...) \
@@ -677,6 +680,12 @@ static inline int virtqueue_add_split(
struct virtqueue *_vq,
* new available array entries. */
virtio_wmb(vq->weak_barriers);
vq->split.avail_idx_shadow++;
+ {
+ if ((vq->split.avail_idx_shadow % 100) == 0)
+ printk(KERN_ERR "avail idx: %d",
+ (int)vq->split.avail_idx_shadow);
+ if (vq->split.avail_idx_shadow == 5000)
+ vq->split.avail_idx_shadow = 0;
+ }
vq->split.vring.avail->idx = cpu_to_virtio16(_vq->vdev,
vq->split.avail_idx_shadow);
vq->num_added++;
@@ -689,6 +698,11 @@ static inline int virtqueue_add_split(
struct virtqueue *_vq,
if (unlikely(vq->num_added == (1 << 16) - 1))
virtqueue_kick(_vq);
+ {
+ if (unlikely(vq->split.avail_idx_shadow == 0))
+ virtqueue_kick_always(_vq);
+ }
+
return 0;
unmap_release:
@@ -2515,6 +2529,11 @@ bool virtqueue_kick(struct virtqueue *vq)
}
EXPORT_SYMBOL_GPL(virtqueue_kick);
+void virtqueue_kick_always(struct virtqueue *vq)
+{
+ virtqueue_kick_prepare(vq);
+ virtqueue_notify(vq);
+}
/**
* virtqueue_get_buf_ctx - get the next used buffer
* @_vq: the struct virtqueue we're talking about.
--------------- END OF LINUX KERNEL PATCH ----------
Then the kernel was booted, and the host pinged until the
nic became unresponsive:
ping -i 0.002 192.168.4.1
Device status was confirmed using
cat /sys/class/net/eth0/device/status
(it was 0x4f).
Then the device was re-initialized:
DEV_NAME=$(basename $(readlink -f /sys/class/net/eth0/device))
echo $DEV_NAME | tee /sys/bus/virtio/drivers/virtio_net/unbind
echo $DEV_NAME | tee /sys/bus/virtio/drivers/virtio_net/bind
ip link set eth0 up
At this point networking became healthly again.
Signed-off-by: Peter Oskolkov <posk@google.com>
Add UFFD restore tests to common_sequential: basic anonymous RAM,
shared memory, and hugepage-backed zone memory. Each exercises the
full snapshot/restore cycle with memory_restore_mode=ondemand and
verifies CPU count, memory size, and device health after resume.
Signed-off-by: Shayon Mukherjee <shayonj@gmail.com>
When memory_restore_mode=ondemand is specified on the restore command,
the memory manager creates a userfaultfd descriptor, registers each
guest RAM range for missing-page fault interception, and spawns a
handler thread that serves page faults from the snapshot file using
UFFDIO_COPY. This avoids reading the entire memory-ranges file into
guest RAM before restore completes.
The handler uses epoll to multiplex the userfaultfd and a stop eventfd
for clean shutdown. Concurrent faults from multiple vCPUs are handled
by treating EEXIST as a benign race and waking blocked threads with
UFFDIO_WAKE. Once all pages have been served the handler exits
automatically. If the handler thread panics the VMM is signalled to
exit since the VM cannot continue without page fault service.
MemoryZone gains a backing_page_size field so the handler resolves
fault granularity from the zone rather than the top-level config.
Errors from the UFFD setup path use a structured UffdError enum
and a new MigratableError::OnDemandRestore variant, with a From
impl to keep call sites concise.
The seccomp filter is updated to allow the userfaultfd syscall and
the four uffd ioctls (UFFDIO_API, UFFDIO_COPY, UFFDIO_REGISTER,
UFFDIO_WAKE) under the VMM thread profile.
Signed-off-by: Shayon Mukherjee <shayonj@gmail.com>
Add a MemoryRestoreMode enum (Copy | OnDemand) to RestoreConfig so
the restore path can be selected at restore time. Copy preserves the
existing eager read-copy behavior. OnDemand enables userfaultfd-based
demand paging and fails restore if the kernel does not support it.
Validate that prefault=on is not combined with OnDemand mode.
Update the OpenAPI spec with the new enum field.
Signed-off-by: Shayon Mukherjee <shayonj@gmail.com>
Add safe Rust wrappers around the raw userfaultfd ioctls: create
(syscall + API handshake), register (missing-page mode), copy
(resolve fault), and wake (unblock threads after EEXIST race).
These are used by the demand-paged snapshot restore handler in a
subsequent commit.
Signed-off-by: Shayon Mukherjee <shayonj@gmail.com>
Add a small constants module with the ioctl numbers and protocol
constants needed for userfaultfd-based demand-paged snapshot restore.
These are derived from the kernel's include/uapi/linux/userfaultfd.h.
Signed-off-by: Shayon Mukherjee <shayonj@gmail.com>
Emit a "vm.migration-memory-iteration" event after every precopy memory
iteration to allow management software to observe forward progress
during migration.
This event is primarily intended for integration with management
software such as libvirt, where it maps to
VIR_DOMAIN_EVENT_ID_MIGRATION_ITERATION.
The event is intentionally independent of any upcoming migration
metrics endpoint. Detailed migration statistics will be exposed via
that endpoint, while this event provides a lightweight progress signal
expected by external management layers.
With this event, management software can detect forward progress during
migration without being blocked on any upcoming migration metrics
endpoint.
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
On-behalf-of: SAP philipp.schuster@sap.com