The refcount rebuild check issued a seek to the refcount table and to
the first refblock before each cursor read. Read the fixed size fields
with read_exact_at at their offsets and decode with from_be_bytes. The
seeks and the matching error paths go away.
The result is unchanged.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The compressed cluster write and read and the L1 resize size query
went through a seek on the AlignedFile cursor before the access. Pass
the target offset to write_at and read_exact_at, and read the file
length from physical_size.
The result is unchanged. The compressed paths keep routing through the
AlignedFile O_DIRECT bounce.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Check that the MMIO accesses is 4 bytes long as otherwise it would
be possible for the guest to trigger a panic when the memory ranges base
and length are copied for fulfilling the MMIO read.
This pattern of check matches similar checks in CpuManager and
DeviceManager.
Signed-off-by: Rob Bradford <rbradford@meta.com>
Fill the target MMIO buffer with zeroes to handle reads with access
sizes larger than the data and also check that the read access length
does not exceed the size of the backing slice (previously it just
checked the access size vs length not taking the offset into account).
Assisted-by: Claude:Opus-4.8
Signed-off-by: Rob Bradford <rbradford@meta.com>
Broaden error!() to cover any user-initiated action that fails to do
what was expected (e.g. failed hotplug or live migration), not only
unrecoverable startup errors. Retarget info!() at operators and users,
clarify the warn!() and debug!() audiences, and document trace!().
Part of #8440.
Co-authored-by: Philipp Schuster <phip1611@gmail.com>
Signed-off-by: Tushar Khatri <hello@tusharkhatri.in>
Suggested by phip1611 on #8446.
This adds the repo's first clippy.toml, carving arch out of the
absolute_paths deny from #7670. Glob imports and trait imports that
must be in scope for method-call resolution (e.g. DeviceInfoForFdt for
.irq()) are left as-is.
Signed-off-by: Henry Hrvoje Tonkovac <htonkovac@gmail.com>
Assisted-by: Claude:Opus-4.8
Cover read_unaligned and write_unaligned directly: a scatter read at an
unaligned offset, a short read at EOF, a read-modify-write gather that
preserves head and tail padding, and error propagation from the
scatter closure.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
run_unaligned_operation staged every unaligned request in a plain Vec
and then handed it to AlignedFile, which bounced again through an
aligned buffer. That Vec only gave the operation a contiguous range to
scatter into or gather from, which the aligned buffer already is, so
each slow path request paid for an extra allocation and a full length
copy.
Add read_unaligned and write_unaligned on AlignedFile that own the
single aligned bounce and scatter or gather through a closure over the
staging slice. run_unaligned_operation and the FileExt read_at and
write_at impls both route through them, so the staging and
read-modify-write logic lives in one place. The closures keep
AlignedFile free of any AsyncIoOperation dependency.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Replace the use of unsafe struct casting with zerocopy trait derivation.
This fixes a Rust UB where the struct was being filled with a slice of
length greater than the size of the struct.
As a compromise the guid handling was changed to handle the uuids as
opaque bytes as they are mixed endian. This has no impact on the
functionality as they are only used for comparison and has the positive
impact of reducing some of the uuid handling complexity.
Assisted-by: Claude:Opus-4.8
Signed-off-by: Rob Bradford <rbradford@meta.com>
The struct already implements ByteValued so this unsafe block can be
changed to call as_slice() from that trait.
Signed-off-by: Rob Bradford <rbradford@meta.com>
The struct already implements ByteValued so this unsafe block can be
changed to call as_slice() from that trait.
Signed-off-by: Rob Bradford <rbradford@meta.com>
Migrate the MSHV integration tests to run natively on the self-hosted
runner instead of spinning up a separate VM. This simplifies the
workflow pipeline & mitigates Azure capacity issues.
Signed-off-by: Aastha Rawat <aastharawat@microsoft.com>
Cover refcount block round trip for the byte aligned and sub byte
paths, and add_cluster_end appending an aligned cluster and staying
within the maximum offset bound.
Assisted-by: Claude:Opus-4.8
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
add_cluster_end queried the file length by seeking to the end. Use the
existing physical_size helper instead, which reads the length from the
file metadata. This removes the final cursor access in QcowRawFile, so
the Seek and SeekFrom imports are no longer needed.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The refcount block read and write helpers took a file whose cursor was
positioned by a preceding seek. Pass the target offset down instead and
use positional read_exact_at and write_all_at on the AlignedFile, so
the block methods no longer seek. The byte aligned and sub byte writers
build a buffer and issue one positional write, keeping the previous
batching.
The result is unchanged, as the calls still route through the
AlignedFile O_DIRECT bounce.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Make the testing overview describe the dev_cli.sh workflow instead of
implying that every Cloud Hypervisor build must run in a container.
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
We cannot reliably send Request::abandon() on every kind of failure on
the sender side, as we might be in the middle of a memory transmission.
The receiver would not reliably know what to do with that. So instead,
when the receiver cannot read from the socket, we log that the migration
sender failed, which is the only likely cause of that failure.
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
This improves the observability whether a migration failed because of
the sender or because of some error on the receiving side.
Using a simple log message is simpler than introducing a new error enum
to differentiate between SendError and RemoteError.
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
When the receiver of a live migration encounters an error, it sends an
error response. The sender of the migration would then send an abandon
request and wait for a response. This abandon request is not necessary,
because the receiver already abandoned the migration due to the error it
encountered.
From now on this function will not send an abandon request to the
receiver anymore, thus it was renamed to "ok_or_error".
Also, this case was always broken, because after sending the error
response, the receiver just exits without waiting for the additional
abandon request.
On-behalf-of: SAP sebastian.eydam@sap.com
Signed-off-by: Sebastian Eydam <sebastian.eydam@cyberus-technology.de>
The `read_aligned_block_size()` function used `Vec::from_raw_parts()`
incorrectly, causing undefined behavior when deallocating the `Vec<u8>`.
One of the safety invariants of `Vec::from_raw_parts()` is that the
provided pointer must be allocated with the exact same alignment as `T`
(`u8` in this case), but this is clearly not true: `align_of::<u8>()` is
1, but the pointer was allocated with aligment of `blocksize` (typically
512 or greater).
Fix this by using the existing `AlignedFile` helper to read the header
block when probing the image type. This is slightly less efficient than
using `AlignedBuffer` directly, but since this is only called once per
disk image at startup, the difference is probably not worth the extra
verbosity.
Signed-off-by: Daniel Verkamp <drv@meta.com>
Add unit tests for add_pci_capabilities covering the configless
device path. The device config capability is present when the
config region is sized and absent when the size is zero.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The watchdog, rng, and rtc devices expose no device specific
configuration fields. Each now reports a config size of zero so the
transport omits the device configuration capability instead of
advertising an unbacked region that the device cannot service.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add a config_size method to VirtioDevice and use it when building the
PCI device configuration capability. The transport advertises the size
reported by the device and omits the capability entirely when the size
is zero, because the virtio driver rejects a zero length capability.
The method defaults to None, so every device keeps its current
capability size.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Cover read_pointer_table round trip and masking, and the
write_cluster then zero_cluster round trip, exercising the positional
read_exact_at, write_all_at, and write_all_zeroes_at paths on the
AlignedFile.
Assisted-by: Claude:Opus-4.8
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Replace the seek then read/write metadata access in QcowRawFile with
positional read_exact_at, write_all_at, and write_all_zeroes_at on the
AlignedFile. read_pointer_table, write_pointer_table,
write_pointer_table_direct, zero_cluster, and write_cluster no longer
move the file cursor.
These calls still route through AlignedFile, which implements FileExt
and WriteZeroesAt with the O_DIRECT alignment bounce, so the unaligned
behavior is preserved. Each access already issued an absolute seek
before touching the file, so the cursor never carried state between
calls and dropping it is unobservable.
Decoding the pointer table now uses native from_be_bytes over the read
buffer, matching the to_be_bytes path on the write side.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
The region table overlap check in RegionInfo::new only rejected a new
region that strictly engulfed an existing one. Identical, fully
contained, and partially overlapping regions passed undetected, so a
malformed VHDX with overlapping region entries was wrongly accepted.
Per [MS-VHDX] all region objects MUST be non-overlapping, so such an
image should be rejected. Replace the faulty predicate with a correct
half-open interval overlap test, extracted into a small pure helper
(ranges_overlap).
Add a unit test for the predicate and an integration test that feeds a
crafted region table with two overlapping entries through the real
RegionInfo::new, confirming it is now rejected with RegionOverlap.
Related to #8009 (broader VHDX overlap validation).
Signed-off-by: Henry Hrvoje Tonkovac <htonkovac@gmail.com>
Assisted-by: Claude:Opus-4.8
Receiving a migration happens inside the VMM thread, which blocks the
API until a migration was received. On the other hand, sending a
migration is actually just a dispatch operation. We adjust the wording
to improve clarity of the error messages.
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
Adds three recommendation bits to CPUID 0x40000004.EAX so Windows /
Hyper-V-aware guests use the corresponding paravirtualized hypercalls
instead of falling back to architectural primitives. The hypercalls
themselves are emulated unconditionally by KVM and surfaced via the
corresponding KVM_CAP_HYPERV_* info caps; no userspace cap negotiation
is needed because KVM advertises support to the guest at hypercall
issue time:
KVM_CAP_HYPERV_TLBFLUSH advertises HvFlush{VirtualAddressSpace,Ex,
List,ListEx} (api.rst 8.18, info-only cap).
KVM_CAP_HYPERV_SEND_IPI advertises HvCallSendSyntheticClusterIpi{,Ex}
(api.rst 8.20, also info-only).
Leaf 0x40000004.EAX (HV_CPUID_ENLIGHTMENT_INFO):
bit 1 LocalTlbFlushRecommended
bit 2 RemoteTlbFlushRecommended
Recommend HvFlushVirtualAddressSpace / List in place of
architectural INVPCID / INVLPG broadcasts. Remote shoot-down
via hypercall lets the host skip vCPUs that are not currently
scheduled, instead of waiting for an IPI ack.
bit 10 ClusterIpiRecommended
Recommend HvCallSendSyntheticClusterIpi over per-target APIC
ICR writes. A single hypercall can target up to 64 vCPUs (or
all of them via the Ex variant) versus one VM exit per APIC
access on the architectural path.
These bits depend on AccessVpIndex (0x40000003.EAX bit 6), which is
advertised by the partition-privileges change.
Sources:
Microsoft Hypervisor Top-Level Functional Specification 7.4.5
qemu/qemu docs/system/i386/hyperv.rst (hv-tlbflush, hv-ipi)
Linux Documentation/virt/kvm/api.rst 8.18, 8.20
Signed-off-by: CMGS <ilskdw@gmail.com>
Extends the Hyper-V partition feature CPUID leaf 0x40000003 with bits
that KVM emulates unconditionally and that Windows / Hyper-V-aware
guests consult to skip slow fallback paths. No KVM capability
negotiation is required for any of these -- they are hints to the
guest about what is already legal to use.
Leaf 0x40000003.EAX (HV_CPUID_FEATURES):
bit 0 AccessVpRuntimeReg -- HV_X64_MSR_VP_RUNTIME (0x40000010)
bit 4 AccessIntrCtrlRegs -- HV_X64_MSR_{EOI,ICR,TPR,APIC_ASSIST}
bit 11 AccessFrequencyMsrs -- HV_X64_MSR_{TSC,APIC}_FREQUENCY (skips
guest TSC/APIC calibration loops)
Leaf 0x40000003.EDX (HV_CPUID_FEATURES, TLFS rev 6.0c):
bit 4 FastHypercall -- HV_HYPERCALL_PARAMS_XMM_AVAILABLE
bit 8 ExtendedGvaRangesForFlushVirtualAddressList -- pairs with the
tlbflush-ext recommendation bit
AccessHypercallMsrs (bit 5) and AccessVpIndex (bit 6) are already
advertised by the partition-privileges change.
Sources:
Microsoft Hypervisor Top-Level Functional Specification 7.4.{2,5}
qemu/qemu docs/system/i386/hyperv.rst (hv-vapic, hv-frequencies,
hv-vpruntime)
Signed-off-by: CMGS <ilskdw@gmail.com>
On the receiver side, a live migration with status "aborted" does not
return an error. Thus, management software will think that the live
migration was successful (from just looking at the API response). This
is not expected behaviour.
On-behalf-of: SAP sebastian.eydam@sap.com
Signed-off-by: Sebastian Eydam <sebastian.eydam@cyberus-technology.de>
So far, we only have seccomp rules for the postcopy-send thread. This
commit introduces the basic plumbing to add seccomp rules also for the
migration worker (the migration coordinator) as well as the TCP workers
(both, send and receive) in the following.
To streamline code setup, all filters are created at a central place
early in the migration code. Although this means that some filters are
created without the need to do so (e.g., postcopy), this massively
simplifies code setup and error handling. This overhead is negligible.
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
ReceiveAdditionalConnections got quite complicated, especially with the
many threads involved for precopy and the special-case of postcopy. We
therefore should add comprehensive documentation.
I tried to keep it short and concise - what remains provides high value
and improves the mental model of the code.
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
Live migration can deadlock if the guest triggers a virtio device
activation while the migration worker owns the VM.
The failure shows up when starting live migrations during boot and
firmware startup, where the guest can reset and reinitialize virtio
devices while precopy is running. In the failing case, the source log
shows a pending virtio activation that never completes:
8.115833s _virtio-pci-net_0: Needs activation; returning barrier
8.115854s vmm/src/vm.rs:464 -- Waiting for barrier
24.875452s Entering downtime phase
24.875481s stopping vcpu throttling thread
...
vCPU thread did not respond in 10ms to signal - retrying
vCPU thread did not respond in 20ms to signal - retrying
...
thread 'throttle-vcpu' (1029) panicked
...
Pause(Error signalling vCPUs: Timeout when waiting for signal
to be acknowledged)
The vCPU blocks on the activation barrier and never reaches the normal
pause checkpoint. Later, migration enters downtime and stops the vCPU
throttle thread. In the failing case, that thread is still inside a
CpuManager::pause() call, which waits for every vCPU to acknowledge
the signal. The blocked vCPU never does, so the pause times out.
Fix this by storing the DeviceManager inside VmOwnership::Migration.
This keeps just enough state on the VMM thread to drain pending virtio
activations while the migration worker owns the Vm. The barrier logic
stays unchanged. The VMM now releases the same activation barrier during
migration that it already released before migration started.
This keeps the guest from getting stuck in the activation wait and
lets the later pause succeed.
Co-authored-by: Leander Kohler <leander.kohler@cyberus-technology.de>
On-behalf-of: SAP leander.kohler@sap.com
Signed-off-by: Leander Kohler <leander.kohler@cyberus-technology.de>
On-behalf-of: SAP philipp.schuster@sap.com
Signed-off-by: Philipp Schuster <philipp.schuster@cyberus-technology.de>
When seccomp traps a SIGSYS, print the syscall number that caused it,
the current thread id and thread name to make violations easier to
debug.
This change requires that all threads are allowed to execute the
`gettid` and the `prctl` syscalls, thus the seccomp filters have also
been adjusted.
On-behalf-of: SAP sebastian.eydam@sap.com
Signed-off-by: Sebastian Eydam <sebastian.eydam@cyberus-technology.de>
Getting rid of the unsafe ByteValued implementation for MemoryRange,
Request and Response structures, by relying on zerocopy's safe
implementation instead.
Signed-off-by: Sebastien Boeuf <sboeuf@meta.com>
Assisted-by: Claude:claude-opus-4-8