Nightly rustfmt now prefers `self` re-exports inline rather
than a separate 'pub use {kvm_bindings, kvm_ioctls}' line.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Avoid a potential infinite loop where if the leader fails to create a
cookie due to an unexpected error (not one of the SMT/no kernel support
errors) then the other vcpu threads will continue around their
spinloops.
This change also clarifies the state machine for the leader election
with an explicit enum.
Signed-off-by: Rob Bradford <rbradford@meta.com>
Drop an unused vm_memory::GuestAddress import from common_cvm
in integration tests to keep the module clean.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Replace hard-coded --memory size=512M args with default_memory()
across integration tests to centralize default memory settings.
This reduces duplicated CLI fragments and keeps behavior consistent.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Replace hard-coded --cpus boot=<n> arguments in integration tests
with GuestCommand::default_cpus() for shared, centralized defaults.
This removes duplicated CLI fragments and keeps CPU setup consistent.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Replace the hard-coded memory threshold check in the simple launch
integration test with Guest::validate_memory(None).
Add Guest::get_expected_memory() to derive thresholds from mem_size_str
and vm_type, and reuse this through validate_memory().
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Replace hard-coded --memory args in simple launch tests
with GuestCommand defaults driven by Guest state.
Add Guest.mem_size_str with a default of 512M and introduce
default_memory_string() and GuestCommand::default_memory().
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Instead of validating number of CPU in the test case itself,
moving the checking of the CPU count to Guest struct with a
new function as The Guest already has the Default CPU number.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Replace the hard-coded .args(["--cpus", "boot=1"]) in the simple
launch integration test with a shared helper (default_cpus) from test
infrastructure.
Extend Guest with explicit CPU-related defaults (num_cpu, nested)
and add default_cpus_string() so CPU configuration is derived from
guest state instead of being duplicated at call sites.
This refactor improves consistency and makes CPU defaults easier to
maintain across integration tests.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Move MetaEvent from the integration test into shared test infrastructure
and expose it for reuse. Add a Guest helper that returns the expected
sequential events for simple launch, and update the integration test to
consume this helper instead of maintaining a local event list.
Adjust expected behavior for confidential VMs by omitting the disk reset
event, which is not guaranteed to be emitted in that mode. Preserve the
existing expected sequence for non-confidential VMs.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
The SYNTAX help string for --net was missing a comma between
pci_segment and offload_tso parameters, making the help output
show them as a single run-on token.
Signed-off-by: Victor Vieux <vieux@repl.it>
O_DIRECT requires buffer addresses to be aligned to the backend
device's logical block size. The existing bounce buffer logic in
execute_async() hardcodes SECTOR_SIZE (512) for the alignment check
and bounce buffer allocation. This is insufficient for devices with
a 4096-byte logical block size, where misaligned buffers cause
-EINVAL from the host kernel.
Add an alignment() method to the AsyncIo trait that returns the
backend's logical block size, defaulting to SECTOR_SIZE. The three
raw I/O backends (io_uring, AIO, synchronous) probe the device
topology via DiskTopology::probe() at creation time and return the
actual logical block size. All image format backends would simply
use the default value of 512 bytes since their underlying are
not block devices.
execute_async() now queries disk_image.alignment() instead of using
the hardcoded SECTOR_SIZE
Fixes: #7720
Signed-off-by: Saravanan D <saravanand@crusoe.ai>
Add a core_scheduling option to --cpus with three modes of operation.
This feature takes advantage of a kernel feature that restricts
scheduling of processes on the SMT threads on the same core. This is
useful for mitigating certain classes of side-channel attacks and has
better performance that disabling SMT on the CPU.
- vm (default): All vCPU threads share one core scheduling cookie.
They may be co-scheduled on SMT siblings while host threads are
excluded - this has minimal performance impact and can even
potentially improve performance from co-location.
- vcpu: Each vCPU gets a unique cookie preventing any two vCPUs from
sharing SMT siblings. This has the strongest isolation but at some
compromise of performance.
- off: No core scheduling applied (old behaviour).
This isolation is done by the kernel maintaining a "cookie" - threads
with the same cookie can share the same core.
In vCPU mode each vCPU thread the cookie is created when the thread
starts and each gets a unique cookie. For VM mode the first vCPU thread
(the leader) will create the cookie. All other vCPU threads started (via
hotplug or during boot) will have that cookie shared to it.
EINVAL/ENODEV from prctl is silently ignored so this works transparently
on kernels older than 5.14 that lack PR_SCHED_CORE or when SMT disabled.
Full details of this kernel feature can be found at:
https://docs.kernel.org/admin-guide/hw-vuln/core-scheduling.html
This implementation was inspired by crosvm's implementation - in
particular the enable_core_scheduling() function.
This is challenging to test via integration testing but the logging of
the received cookie shows it working:
VM case:
cloud-hypervisor: 0.243102s: <vcpu1> INFO:vmm/src/cpu.rs:1247 -- vCPU 1: core scheduling cookie = 0x33e4c167
cloud-hypervisor: 0.243102s: <vcpu0> INFO:vmm/src/cpu.rs:1247 -- vCPU 0: core scheduling cookie = 0x33e4c167
vCPU case:
cloud-hypervisor: 0.089356s: <vcpu0> INFO:vmm/src/cpu.rs:1247 -- vCPU 0: core scheduling cookie = 0x13993ad6
cloud-hypervisor: 0.089380s: <vcpu1> INFO:vmm/src/cpu.rs:1247 -- vCPU 1: core scheduling cookie = 0xd48e86e
Signed-off-by: Rob Bradford <rbradford@meta.com>
The current implementation performs multiple operations on allocators in
a row, with the single goal of updating the allocator. For each of these
operations, the `Mutex` guarding the respective allocator is locked anew
which introduces room for race conditions.
Instead of locking the mutex multiple times, we should lock it once to
perform the whole move.
Signed-off-by: Pascal Scholz <pascal.scholz@cyberus-technology.de>
On-behalf-of: SAP pascal.scholz@sap.com
Vm::create_device_manager accepted both config and _vm_config, but
both represented the same VM configuration source. Remove _vm_config
from the function signature and from its call site, and use config
for the TDX dynamic check.
This is a cleanup-only refactor with no intended functional change.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
There is no non destructive readonly ioctl to query block device
discard or write zeroes capabilities. BLKZEROOUT is guaranteed to
succeed via kernel software fallback. BLKDISCARD may fail at runtime
with EOPNOTSUPP on devices that lack trim support, but the error
propagates to the guest as VIRTIO_BLK_S_IOERR and well behaved
guests handle it gracefully.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The Windows tests use a DM snapshot device for the OS disk.
DM snapshot targets do not support BLKDISCARD, so the VMM returns
IOERR for every TRIM attempt. viostor.sys may BSOD when the host
returns an error for negotiated discard/write-zeroes operations.
Add a default_disks_sparse_off() helper to GuestCommand and use it
in all Windows tests.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Verify that the guest remains stable when BLKDISCARD fails on the
host backend. DM snapshot targets do not support discard, so the
VMM returns VIRTIO_BLK_S_IOERR. The test retries blkdiscard several
times, checking guest responsiveness after each attempt, then
confirms normal I/O still works.
The DM topology follows the same pattern used by WindowsDiskConfig.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Verify that a loopback block device advertises
VIRTIO_BLK_F_DISCARD to the guest and that blkdiscard succeeds.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Extend run_qemu_img() with an optional trailing_args parameter
for arguments that follow the image path, such as the size in
'qemu-img create -f raw <path> 128M'.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Some virtio devices cannot be implemented via vhost-user because they
require tight integration with the VMM. This includes the IOMMU and
watchdog devices.
An attempt to create a generic vhost-user device with one of these IDs
is always either a bug or human error. To aid debugging, return a
helpful error message rather than silently continuing.
Signed-off-by: Demi Marie Obenour <demiobenour@gmail.com>
This implements a generic vhost-user device. All information about this
device must be provided to Cloud Hypervisor via the command-line or API.
The main use-case is types of vhost-user devices Cloud Hypervisor
doesn't know about, but it can also be used for types it does know
about.
The generic device delegates all configuration space handling to the
backend. This means that the vhost-user backend must support
configuration space access. It also means that the backend has control
of configuration space. For instance, this means that setting the tag
of a virtio-fs device on the virtiofsd command line works as expected.
If the VM is snapshotted or migrated, the backend must write the
configuration space to a separate save file or migration stream.
Similarly, if the VM is restored or migrated, the backend must read the
configuration space from a separate save file or migration stream.
Signed-off-by: Demi Marie Obenour <demiobenour@gmail.com>
Wrap the Vhdx instance in Arc<Mutex<>> so that all queues share
a single mutex-protected backend, matching the approach already
used for QCOW2.
Vhdx::clone() uses dup() which shares the kernel file description
including the file offset. With multiple queues performing
concurrent seek+read/write on the shared offset, I/O operations
race and corrupt data.
Fixes: #7665
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Accept arguments after -- in 'dev_cli.sh shell' and forward them
to 'bash -c' inside the container. When no arguments are given,
an interactive shell is started as before. This enables running
one-off commands in the CI container without an interactive session,
for example:
./scripts/dev_cli.sh shell -- rustup toolchain install nightly \&\& cargo +nightly fmt --all -- --check
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
As well as rejecting writes to sector 0 in the case of raw files where
the user hasn't specified the image_type also reject virtio requests of
type discard and write_zeroes.
Signed-off-by: Rob Bradford <rbradford@meta.com>
If the disk image was autodetected to raw (not specified with image_type
= 0) then in the virtio-block subsystem generate errors for writes to
block 0 (treat as if read-only). This gives an immediate error vs using
the image implementations in the block subsystem.
Signed-off-by: Rob Bradford <rbradford@meta.com>
Add an image_type to DiskConfig to specify the image type. If none is
specified autodetect the image type but disable potentially unsafe
behaviour in the QCOW2 backend by disabling the backing file support.
If the image type is autodetected then fix it in the config so that it
will be persistant across reboots and migrations/snapshot & restores.
This also handles the case where the image type was not specified as
part of the disk configuration.
Signed-off-by: Rob Bradford <rbradford@meta.com>
Verify that opening a QCOW2 image with a backing file reference
through QcowDiskSync with backing_files=off produces the user-facing
BackingFilesDisabled error rather than MaxNestingDepthExceeded.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
When a QCOW2 image has a backing file but backing_files=on is not set,
the error was MaxNestingDepthExceeded which gives no indication that
this is a policy decision or how to resolve it.
Add a BackingFilesDisabled error variant whose message indicates that
backing file support is disabled and references the backing_files
option. The translation from MaxNestingDepthExceeded to
BackingFilesDisabled happens at the QcowDiskSync boundary where the
policy decision is made, preserving the original error for genuine
recursive depth exhaustion.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The 0.6.7 version of the mshv crates introduced a new version of
make_default_partition_create_arg inside `struct Mshv`. This version
queries the available processor features on the host and gives the same
feature set to the guests.
Move Cloud Hypervisor to this new function.
Signed-off-by: Anirudh Rayabharam <anrayabh@microsoft.com>