Core
GPU metrics
The native engine can sample the host's GPUs for the whole run — utilization,
VRAM, temperature, and power — and land a gpu section in the run summary.
This exists primarily for LLM load testing: with std/llm@v1
against a local server (Ollama, vLLM, …) the GPU is the system under test,
and correlating TTFT/tokens-per-second with SM utilization and memory
pressure is how you tell "model is saturated" apart from "server is
misconfigured".
Collection is best-effort: no GPU, a missing nvidia-smi binary, or an
unreachable exporter logs one warning at run start and the run continues
without GPU metrics — it never fails the run.
Configuration
The gpu: block sits in the run config
next to vus/duration (native engine only):
vus: 10
duration: 5m
gpu:
enabled: true
interval_ms: 1000 # default 1000 (min 10)
source: nvidia-smi # nvidia-smi (default) | dcgm | powermetrics (macOS)
dcgm_url: http://127.0.0.1:9400/metrics # for source: dcgm
devices: [0, 1] # optional; default — every GPU the source reports
| Field | Default | Description |
|---|---|---|
enabled | false | Master switch; the section is ignored without it |
interval_ms | 1000 | Sampling interval (one snapshot per GPU per tick) |
source | nvidia-smi | nvidia-smi shells out to the binary; dcgm polls a dcgm-exporter HTTP endpoint; powermetrics samples Apple Silicon GPU/ANE on macOS (needs root) |
dcgm_url | http://127.0.0.1:9400/metrics | dcgm-exporter metrics endpoint (source: dcgm only) |
devices | all | Restrict sampling to these GPU indices |
Sources
nvidia-smi (default) — runs
nvidia-smi --query-gpu=index,utilization.gpu,memory.used,memory.total,temperature.gpu,power.draw --format=csv,noheader,nounits
once per tick. Works anywhere the NVIDIA driver is installed, no extra
daemon needed. Fields the driver reports as N/A (e.g. power draw on some
virtualized GPUs) are recorded as absent, not zero.
dcgm — HTTP GET on dcgm_url and parses the Prometheus text format of
dcgm-exporter:
DCGM_FI_DEV_GPU_UTIL, DCGM_FI_DEV_FB_USED/DCGM_FI_DEV_FB_FREE (VRAM
total is derived as used+free), DCGM_FI_DEV_GPU_TEMP,
DCGM_FI_DEV_POWER_USAGE. Better for GPU servers and Kubernetes, where
dcgm-exporter is typically already running.
powermetrics (macOS, Apple Silicon) — runs
sudo powermetrics --samplers cpu_power,gpu_power -n 1 -f plist once per
tick. The integrated GPU (index 0) gets a true utilization figure (active
residency %) and its rail power. The Neural Engine (ANE/NPU) has no public
utilization API on macOS — the only signal is its power draw — so the ANE
arrives as ane_power_w alongside cpu_power_w/package_power_w, charted
as its own series (during-run too). Memory and temperature stay absent:
unified memory has no VRAM figure, and powermetrics reports thermal
pressure, not °C. Note that LLM inference on a Mac (llama.cpp, Ollama,
MLX) runs on the GPU via Metal — the ANE is only exercised by CoreML models
compiled for it, so watch gpu_utilization_pct first and treat
ane_power_w as the NPU-activity proxy.
powermetrics requires root. Either run the CLI under sudo, or grant the
load-test user passwordless sudo for the tool only:
echo "$USER ALL=(root) NOPASSWD: /usr/bin/powermetrics" | sudo tee /etc/sudoers.d/powermetrics
Without root the run logs one warning and continues without GPU metrics (the usual best-effort contract).
Sampling-cadence note: each powermetrics -n 1 invocation measures a
window of gpu.interval_ms (clamped to 1s–60s) and blocks for roughly the
window plus ~1s of startup overhead, so the effective cadence is the tick
plus window plus overhead (~2s at the default 1s tick). Keep
interval_ms at 5s or above for near-continuous window coverage.
Output
While the VUs run, the sampler takes one snapshot per GPU per tick (the first immediately at run start). After the metric summary the run prints a compact block:
gpu: 1 device, 300 samples every 1000ms (nvidia-smi)
gpu0: util avg=64.3% max=100.0% vram max=41088/81559MiB temp max=71.0C power max=512.3W
…followed by one machine-readable gpu: {...} line with the full
timeseries, collected into perfscale run --summary-export under gpu
(same framing as the thresholds: {...} gate line):
{
"gpu": {
"source": "nvidia-smi",
"interval_ms": 1000,
"devices": [
{
"index": 0,
"samples": [
{ "ts_ms": 1720000000000, "index": 0, "utilization_pct": 64.0,
"memory_used_mib": 41088.0, "memory_total_mib": 81559.0,
"temperature_c": 71.0, "power_w": 512.3 }
],
"avg_utilization_pct": 64.3,
"max_utilization_pct": 100.0,
"max_memory_used_mib": 41088.0,
"memory_total_mib": 81559.0,
"max_temperature_c": 71.0,
"max_power_w": 512.3
}
]
}
}
Each sample carries ts_ms (epoch milliseconds) on the same timeline as the
stats lines, so throughput/latency and GPU
state can be charted together. Absent optional fields mean the source
reported N/A for that metric. Markdown exports (--summary-export out.md)
get one compact row per device per aggregate.
GPU state also streams during the run: with
report.during_run on,
every engine snapshot carries the GPU samples taken since the previous one —
shipped to the collector as gpu_utilization_pct, gpu_memory_used_mib,
gpu_memory_total_mib, gpu_temperature_c and gpu_power_w gauges with a
gpu="<index>" label, each at the collector's own timestamp. See
during-run metrics for the
full naming table and delivery semantics.
Example: Ollama under load, GPU watch on
# config.yaml
vus: 8
duration: 2m
gpu:
enabled: true
# test.yaml
steps:
- name: llama completion
use: std/llm@v1
with:
url: http://127.0.0.1:11434/v1/chat/completions
model: llama3.1
prompt: "Summarize the CAP theorem in two sentences."
max_tokens: 128
check:
status: 200
perfscale run -f test.yaml -c config.yaml --summary-export gpu-run.json
Reading the result together: llm_tokens_per_sec flat while
gpu0 util avg sits at ~100% → the GPU is the bottleneck (add a card,
shard the model, or lower vus); util well below 100% with rising TTFT →
look at the server (queueing, context limits) instead. VRAM creeping to
memory_total_mib explains evictions/OOMs mid-run.
Example: game-style rendering load
GPU load testing is not only about LLM servers. The other classic question
is session density: how many concurrent render sessions — game clients
on a cloud-gaming node, streaming viewports, digital-twin renderers — one
card carries before the frame rate collapses. The pattern is the same as
above, except the "system under test" is a set of renderer processes that
perfscale orchestrates with
std/child_process@v1 while the gpu:
sampler records what the card is doing. A before: sidecar adds the second
ingredient of a real node — a heavy background GPU job (the encode stage)
competing with the sessions for the same card.
Any renderer that prints FPS works. This example uses
glmark2 (OpenGL, apt install glmark2) looping its 3D scenes as a stand-in for a game client; vkmark
is the Vulkan equivalent, and a headless Unity/Unreal build drops in
unchanged — only the command differs.
# config.yaml
vus: 1 # the renderers are the load; one VU just keeps
duration: 5m # the run open (see test.yaml below)
allow_process_actions: true # required for child_process/kill_process
gpu:
enabled: true
interval_ms: 1000
before:
# One "game session" = one renderer process. The farm spawns as a single
# managed process group, so `after:` stops every session at once.
- name: render-farm
uses: std/child_process@v1
with:
command: sh
args: ["-c", "for i in $(seq 4); do glmark2 --run-forever & done; wait"]
waitUntil:
stdout_contains: GL_RENDERER # GL context is up
on_timeout: continue
restart: never # a crashed session must not respawn a 2nd farm
# Sidecar: a heavy GPU job sharing the card with the sessions — the encode
# stage of a game-streaming pipeline. Looped 1080p60 from a generated
# source, NVENC-encoded, discarded to null.
- name: encode-sidecar
uses: std/child_process@v1
with:
command: ffmpeg
args: ["-f", "lavfi", "-i", "testsrc2=size=1920x1080:rate=60",
"-c:v", "h264_nvenc", "-f", "null", "-"]
waitUntil:
stderr_contains: "Press [q]" # ffmpeg reports the running loop on stderr
on_timeout: continue
restart: on-failure # a crashed encoder comes back
after:
- name: stop the farm
uses: std/kill_process@v1
with: { name: render-farm } # tree: true by default → every session
- name: stop the sidecar
uses: std/kill_process@v1
with: { name: encode-sidecar }
# test.yaml
steps:
- name: keep the run open
use: std/sleep@v1
with: { seconds: 30 }
perfscale run -f test.yaml -c config.yaml --summary-export render.json
Headless nodes: glmark2 needs a GL context. On a GPU server without a
display use the DRM build (glmark2-drm / glmark2-es2-drm — renders via
GBM straight on the card) or wrap the command in xvfb-run -a.
Reading the result
The renderer's FPS lines stream into the run log with a render-farm:
prefix; the gpu: summary records what the card did meanwhile. The method
is a sweep, not a single run — raise the session count (seq 4 → 1, 2, 4,
8) between runs:
- per-scene FPS divided by ~N while
util maxpins at 100% → the GPU is saturated; that session count is the card's ceiling for this workload; - FPS degrades while util stays below 100% → the limit is elsewhere (CPU,
context switching) — cross-check
temp max/power maxfor thermal or power throttling; vram maxper session count answers the capacity question directly: how many sessions fit intomemory_total_mibbefore the driver starts swapping;- the sidecar's price is the FPS delta between a farm-only run (comment
the
encode-sidecarblock out) and a farm+sidecar run at the same session count — that is what sharing the card with the encode pipeline costs a cloud-gaming node.
perfscale does not parse FPS into metrics — frame-rate numbers live in the
run log; the gpu: timeseries (each sample stamped ts_ms on the stats
timeline) is what you chart against them.
One caveat for runs like this: utilization.gpu reports the 3D/compute
engine — the NVENC encoder the sidecar burns is not counted there, so
judge the sidecar by VRAM, power draw, and the FPS it costs the sessions,
not by the util line.
Example: ANE (Neural Engine) load via Core ML
On Apple Silicon the NPU question mirrors the GPU one: does the workload
actually engage the Neural Engine, and where is its ceiling? The shipped
examples/coreml-ane/
generates real ANE load with a Core ML sidecar — a small conv net built with
coremltools' MIL builder (no model
download), looping predictions with compute_units=ALL — while
gpu.source: powermetrics records the SoC:
# config.yaml (excerpt)
gpu:
enabled: true
source: powermetrics
interval_ms: 5000
before:
- name: ane-sidecar
uses: std/child_process@v1
with:
command: sh
args: ["-c", ".venv/bin/python ane_load.py"]
waitUntil: { stdout_contains: "inference loop", on_timeout: fail }
restart: never
after:
- name: stop the sidecar
uses: std/kill_process@v1
with: { name: ane-sidecar }
Verified on an M2 Pro: ane_power_w rises ≈2 W over baseline (to ≈7 W
sustained) while the loop runs at ≈1 450 inferences/s on the built-in net,
with gpu_utilization_pct and gpu_power_w staying flat — proof the load
is on the Neural Engine, not the GPU. (powermetrics models several watts of
ANE draw even at idle on M2 Pro/Max, so read the delta, not the floor.)
Reading the result:
ane_power_wflat at its baseline → the ANE is not engaged: the model is CPU/GPU-bound or was loaded withoutcompute_units=ALL. (Python's data stack — pandas, NumPy, Anaconda — never touches the ANE;coremltools/Core ML is the only path user code has to it.)- Scale out (more sidecars,
--size 224, or your own--model x.mlpackage) and watch where inferences/sec stop scaling — the chip's ANE ceiling for that workload.
Setup (venv + the powermetrics sudoers rule) and the sweep method are in the example's README.
GPU benchmark suite
The repo ships a ready-made local suite in
bench/gpu/:
std/llm@v1 scenarios against Ollama and vLLM with gpu: metrics on, a
stepped ramping-VU profile (concurrency vs tok/s / TTFT degradation) and an
arrival-rate profile (find the rate where TTFT and dropped_iterations
climb). It is local-only — CI runners have no GPU.
bench/gpu/run.sh ollama # or: vllm, or both
…runs each profile, writes --summary-export JSONs and raw logs to
bench/gpu/results/<timestamp>/, and prints a compact table:
scenario reqs req/s tok/s avg ttft p50 ms ttft p95 ms gpu util max vram max MiB dropped
-------------- ---- ----- --------- ----------- ----------- ------------ ------------ -------
ollama-stages 152 0.42 38.71 212.40 890.15 100% 5104 0
ollama-arrival 210 0.63 31.05 340.72 2410.30 100% 5112 17
Setup (Ollama / vLLM), requirements, and how to read the numbers:
bench/gpu/README.md.
Extension seam
perfscale-core exposes gpu::GpuCollector and
gpu::register_gpu_collector — the same pattern as
register_pubsub_driver: a downstream (proprietary)
build can register richer collectors (NVML-based per-process memory, SM
clock/throttle reasons, rocm-smi for AMD, powermetrics for Apple
silicon) and select them via gpu.source, or shadow the built-ins under
their own names. The basic metrics above are the OSS baseline every
collector reports.