The '--test-filter' and '--test-exclude' arguments only accepted a
single keyword, even though the underlying selection logic already
matches against a list. Add a comma value delimiter to both so multiple
keywords can be passed in one invocation.
Signed-off-by: Bo Chen <bchen@crusoe.ai>
Import the std modules used in the crate instead of spelling the full
paths at every use site, and drop the now-unnecessary crate-level
#![expect(clippy::absolute_paths)].
Signed-off-by: Henry Hrvoje Tonkovac <htonkovac@gmail.com>
Assisted-by: Claude:Opus-4.8
Convert the still-needed #[allow] to #[expect] so it warns if the
lint stops firing.
Part of #8326.
Signed-off-by: Tushar Khatri <hello@tusharkhatri.in>
Removal of absolute paths is currently in progress. To avoid regressing
those changes add a clippy deny at the workspace level and at the crate
level override with #[expect(clippy::absolute_paths)]
See: #7670
Signed-off-by: Rob Bradford <rbradford@meta.com>
Use into_iter() for test_list when building tests_to_run.
This keeps the collected type as Vec<&PerformanceTest>.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Confidential VMs (CVM) are not currently supported on the
AArch64 architecture. Add an early check in the performance
metrics binary to exit with a clear error message when CVM
mode is selected on AArch64.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Add a --vm-type command-line argument to allow users to select
between 'regular' (default) and 'confidential' (CVM) VM types
when running performance tests.
Example: --vm-type confidential
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Apply the vm_type override from PerformanceTestOverrides to the
effective_control used during test execution, alongside the
existing test_timeout override.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Add an optional vm_type field to PerformanceTestOverrides to
allow overriding the VM type at runtime. Include vm_type in
the Display output for override logging.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Add a vm_type field of type GuestVmType to PerformanceTestControl,
defaulting to GuestVmType::Regular. Include vm_type in the Display
output for test control logging.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Consolidate override application into a single effective_control
variable built once before the test loop. This removes duplicated
timeout override logic from both warmup and measurement iterations.
Signed-off-by: Muminul Islam <muislam@microsoft.com>
Use thread::Builder to give the test thread a name matching the
test so Guest picks it up automatically. After every test, call
ProcessRegistry::cleanup() to kill the process group instead of
the old pkill based cleanup_stale_processes().
Remove cleanup_stale_processes() and its call sites.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_batch_write which builds a batch of num_ops
write requests and submits them all at once through
submit_batch_requests. Writes in QcowAsync are synchronous (COW
path), so this measures whether batching reduces per-request
overhead compared to individual write_vectored calls.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_l2_cache_miss which reads one cluster
from each of num_ops distinct L2 tables through the QcowAsync
io_uring path, forcing L2 cache eviction on nearly every read.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_write which writes clusters into an
empty QCOW2 image through the QcowAsync io_uring path. Writes
in QcowAsync are synchronous due to COW metadata allocation, so
this measures the write path overhead through the async code path.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_compressed_read which reads from a
zlib compressed QCOW2 image through the QcowAsync io_uring path.
Compressed clusters take the sync fallback since they require
decompression.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_backing_read which reads clusters from
a QCOW2 overlay through the QcowAsync io_uring path. All reads
fall through to the backing file, exercising the sync fallback
path in QcowAsync.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_multi_cluster_read which reads 8
contiguous clusters (512 KiB) per request through the QcowAsync
io_uring path. With coalesced mappings this can hit the io_uring
fast path for a single Readv SQE.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_random_read which reads clusters in
random order through the QcowAsync io_uring path. This mirrors
the existing sync random read benchmark and measures io_uring
completion handling under random access patterns.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_batch_read which builds a batch of num_ops
read requests and submits them all at once through
submit_batch_requests. This exercises the io_uring batch
submission path added in qcow_async, where multiple SQEs are
packed into a single io_uring_enter call.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_async_read which reads clusters through the
QcowDiskAsync io_uring backend. Single allocated cluster reads go
through io_uring for true asynchronous completion, unlike the sync
benchmarks which use QcowDiskSync with blocking I/O.
Workloads: 128 and 256 clusters.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_l2_cache_miss which reads one cluster from each
of num_ops distinct L2 tables in a sparsely allocated image. Clusters
are spaced L2_ENTRIES_PER_TABLE apart so every read touches a different
L2 table, forcing eviction when num_ops exceeds the cache capacity.
Workloads: 128 and 256 L2 tables.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_multi_cluster_read which issues large reads
spanning 8 contiguous clusters (512 KiB) per read_vectored call.
This exercises the mapping coalesce path where multiple L2 entries
are merged into fewer host I/O operations.
Workloads: 128 and 256 total clusters (16 and 32 reads).
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_compressed_read which reads clusters from a
zlib compressed QCOW2 image. Every cluster triggers decompression,
isolating the decompression overhead from the normal allocated cluster
read path.
Workloads: 128 and 256 clusters.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_cow_write which writes clusters into a QCOW2
overlay backed by a raw file. Each write triggers copy-on-write:
cluster allocation, L2 and refcount table updates, then the data
write. This measures COW allocation overhead compared to writing
into a plain empty image.
Workloads: 128 and 256 clusters.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_backing_read which reads clusters from a QCOW2
overlay where all data lives in a raw backing file. Every read falls
through the L2 lookup to the backing file, exercising the backing
chain read path.
Workloads: 128 and 256 clusters.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_random_read which reads clusters from a
prepopulated qcow2 image in a deterministic pseudo-random order.
Unlike the sequential read benchmark, this exercises L2 cache miss
and eviction behaviour under random access patterns.
Uses Fisher-Yates shuffle with DefaultHasher for reproducible
permutation across runs.
Two TEST_LIST entries: micro_block_qcow_random_read_128_us and
micro_block_qcow_random_read_256_us with 128 and 256 cluster
workloads.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_fsync which writes num_ops clusters into an
empty qcow2 image to dirty L2 and refcount metadata then times a
single fsync call that flushes all dirty tables to disk. This
isolates the metadata flush cost which scales with the number of
dirty L2 table entries and refcount blocks.
Two TEST_LIST entries: micro_block_qcow_fsync_64_us and
micro_block_qcow_fsync_256_us with 64 and 256 cluster workloads.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_punch_hole which times punch_hole calls through
QcowSync on a prepopulated qcow2 image. Each call deallocates one
cluster exercising deallocate_bytes with refcount decrement and
fallocate punch_hole on the host file.
Two TEST_LIST entries: micro_block_qcow_punch_hole_64_us and
micro_block_qcow_punch_hole_256_us with 64 and 256 cluster workloads.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_write which times write_vectored calls through
QcowSync on an empty QCOW2 image. Each write allocates a new cluster
exercising map_cluster_for_write with L2 entry allocation and refcount
updates followed by pwrite_all.
Two TEST_LIST entries: micro_block_qcow_write_128_us and
micro_block_qcow_write_256_us with 128 and 256 cluster workloads.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_bench_qcow_read which times read_vectored calls through
QcowSync on a prepopulated QCOW2 image. This exercises the hot
read path including L2 lookup, pread64 for allocated clusters and
iovec scatter.
Two TEST_LIST entries: micro_block_qcow_read_128_us and
micro_block_qcow_read_256_us with 128 and 256 cluster workloads.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Drop the -f flag from the process termination command in
cleanup_stale_processes() so it matches by process name only, not the
full command line. This prevents terminating unrelated processes whose
arguments happen to contain target strings (e.g., the test runner
invoked with --report-file /cloud-hypervisor/report.json).
Use the truncated name 'cloud-hyperviso' because Linux limits process
names to 15 characters.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Anirudh Rayabharam <anrayabh@microsoft.com>
Add a --continue-on-failure CLI flag that allows the test harness to
continue executing remaining tests after encountering a failure, instead
of aborting immediately. When set, failed tests are recorded with zeroed
metrics and a "FAILED" status, the report file is always generated, and
the process exits with a non-zero code if any test failed.
Without the flag, the existing fail-fast behavior is preserved.
Also add a "status" field ("PASSED"/"FAILED") to PerformanceTestResult
so report consumers can distinguish successful tests from failed ones.
Signed-off-by: Anirudh Rayabharam <anrayabh@microsoft.com>
Add a --test-exclude flag that excludes tests matching the provided
keywords. Both --test-filter and --test-exclude are now applied before
--list-tests, so listing respects the active filters.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add micro_block_raw_aio_drain_128_us and
micro_block_raw_aio_drain_256_us tests that submit N AIO writes
to a temporary file, wait for the eventfd signal, then time how
long it takes to drain all completions via next_completed_request().
This measures per completion syscall overhead and provides a
baseline before any batching optimizations.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
These factor out common setup and synchronization patterns used by block
layer micro benchmarks.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add an optional num_ops parameter for micro benchmarks to configure
workload size (e.g. number of AIO operations to submit). A warning
is emitted if it is accidentally set on a non micro test where it
has no effect.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Introduce support for in process micro benchmarks alongside the
existing VM level performance tests. Micro benchmarks are
integrated into the same PerformanceTest/TEST_LIST infrastructure
and follow the same iteration, timeout, and reporting pipeline.
They are distinguished by a micro_* name prefix.
The test dispatch loop is refactored to pre filter the test list
and gate init/cleanup behind a flag, so that pure micro benchmark
runs skip the expensive VM lifecycle entirely. Mixed runs
(VM + micro) continue to work correctly.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Flush host writeback queues, drop the page cache and sleep 1s
for kernel housekeeping before each test run.
The cloud-hypervisor block backend does buffered I/O on the host
side, so dirty pages from prior write tests can accumulate and
compete for I/O bandwidth with subsequent tests. Dropping caches
ensures cold read tests get a consistent baseline rather than
benefiting from data cached by prior tests. The brief cooldown
lets the kernel finish tearing down KVM state and freeing pages
from the previous VM before the next one starts.
Requires root, which the metrics container provides. Silently
fails otherwise.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
When a test times out, the spawned thread containing the
cloud-hypervisor child process, iperf3/ethr sub processes,
and all associated resources (TAP devices, file descriptors,
hugepage reservations) is abandoned without cleanup. This
attaches a cleanup routine that kills cloud-hypervisor,
iperf3, and ethr processes on timeout, then waits
briefly for the kernel to reclaim their resources. This
prevents leaked processes from interfering with subsequent
tests.
Removes the existing TODO comment.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add performance tests for standalone qcow2 images without backing
files - uncompressed, zlib and zstd compressed. Each variant
includes single queue and multiqueue tests for sequential
read, random read and warmed up sequential read.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add multiqueue num_queues=4 performance tests for qcow2 overlay
images with both qcow2 and raw backing files - sequential read,
random read, and warm read variants.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add warmup_iterations field to run iterations before measuring
performance. This complements existing cold start tests
by separating cache effects from steady state throughput.
New tests with 2 warmup iterations:
- block_qcow2_backing_qcow2_read_warm_MiBps
- block_qcow2_backing_raw_read_warm_MiBps
Results show warm cache is much faster and more consistent:
- QCOW2: 1766 MiB/s (4% variance) vs cold 960 MiB/s (73% variance)
- RAW: 1822 MiB/s (6% variance) vs cold 1300 MiB/s (55% variance)
RAW backing is 3% faster than QCOW2 in steady state.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Add sequential and random read performance tests for QCOW2 overlays
with QCOW2 backing files.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>
Introduce a new BlockControl struct to encapsulate fio operation
parameters. This replaces the tuple-based fio_control with a more
extensible structure that includes:
- fio_ops: The FIO operation type
- bandwidth: Whether to measure bandwidth or IOPS
- test_file: The file path to test against
This refactoring enables reusing performance_block_io with different
test files.
Signed-off-by: Anatol Belski <anbelski@linux.microsoft.com>