Core

GPU metrics

The native engine can sample the host's GPUs for the whole run — utilization, VRAM, temperature, and power — and land a gpu section in the run summary. This exists primarily for LLM load testing: with std/llm@v1 against a local server (Ollama, vLLM, …) the GPU is the system under test, and correlating TTFT/tokens-per-second with SM utilization and memory pressure is how you tell "model is saturated" apart from "server is misconfigured".

Collection is best-effort: no GPU, a missing nvidia-smi binary, or an unreachable exporter logs one warning at run start and the run continues without GPU metrics — it never fails the run.

Configuration

The gpu: block sits in the run config next to vus/duration (native engine only):

vus: 10
duration: 5m
gpu:
  enabled: true
  interval_ms: 1000      # default 1000 (min 10)
  source: nvidia-smi     # nvidia-smi (default) | dcgm | powermetrics (macOS)
  dcgm_url: http://127.0.0.1:9400/metrics  # for source: dcgm
  devices: [0, 1]        # optional; default — every GPU the source reports
FieldDefaultDescription
enabledfalseMaster switch; the section is ignored without it
interval_ms1000Sampling interval (one snapshot per GPU per tick)
sourcenvidia-sminvidia-smi shells out to the binary; dcgm polls a dcgm-exporter HTTP endpoint; powermetrics samples Apple Silicon GPU/ANE on macOS (needs root)
dcgm_urlhttp://127.0.0.1:9400/metricsdcgm-exporter metrics endpoint (source: dcgm only)
devicesallRestrict sampling to these GPU indices

Sources

nvidia-smi (default) — runs nvidia-smi --query-gpu=index,utilization.gpu,memory.used,memory.total,temperature.gpu,power.draw --format=csv,noheader,nounits once per tick. Works anywhere the NVIDIA driver is installed, no extra daemon needed. Fields the driver reports as N/A (e.g. power draw on some virtualized GPUs) are recorded as absent, not zero.

dcgm — HTTP GET on dcgm_url and parses the Prometheus text format of dcgm-exporter: DCGM_FI_DEV_GPU_UTIL, DCGM_FI_DEV_FB_USED/DCGM_FI_DEV_FB_FREE (VRAM total is derived as used+free), DCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_POWER_USAGE. Better for GPU servers and Kubernetes, where dcgm-exporter is typically already running.

powermetrics (macOS, Apple Silicon) — runs sudo powermetrics --samplers cpu_power,gpu_power -n 1 -f plist once per tick. The integrated GPU (index 0) gets a true utilization figure (active residency %) and its rail power. The Neural Engine (ANE/NPU) has no public utilization API on macOS — the only signal is its power draw — so the ANE arrives as ane_power_w alongside cpu_power_w/package_power_w, charted as its own series (during-run too). Memory and temperature stay absent: unified memory has no VRAM figure, and powermetrics reports thermal pressure, not °C. Note that LLM inference on a Mac (llama.cpp, Ollama, MLX) runs on the GPU via Metal — the ANE is only exercised by CoreML models compiled for it, so watch gpu_utilization_pct first and treat ane_power_w as the NPU-activity proxy.

powermetrics requires root. Either run the CLI under sudo, or grant the load-test user passwordless sudo for the tool only:

echo "$USER ALL=(root) NOPASSWD: /usr/bin/powermetrics" | sudo tee /etc/sudoers.d/powermetrics

Without root the run logs one warning and continues without GPU metrics (the usual best-effort contract).

Sampling-cadence note: each powermetrics -n 1 invocation measures a window of gpu.interval_ms (clamped to 1s–60s) and blocks for roughly the window plus ~1s of startup overhead, so the effective cadence is the tick plus window plus overhead (~2s at the default 1s tick). Keep interval_ms at 5s or above for near-continuous window coverage.

Output

While the VUs run, the sampler takes one snapshot per GPU per tick (the first immediately at run start). After the metric summary the run prints a compact block:

gpu: 1 device, 300 samples every 1000ms (nvidia-smi)
gpu0: util avg=64.3% max=100.0% vram max=41088/81559MiB temp max=71.0C power max=512.3W

…followed by one machine-readable gpu: {...} line with the full timeseries, collected into perfscale run --summary-export under gpu (same framing as the thresholds: {...} gate line):

{
  "gpu": {
    "source": "nvidia-smi",
    "interval_ms": 1000,
    "devices": [
      {
        "index": 0,
        "samples": [
          { "ts_ms": 1720000000000, "index": 0, "utilization_pct": 64.0,
            "memory_used_mib": 41088.0, "memory_total_mib": 81559.0,
            "temperature_c": 71.0, "power_w": 512.3 }
        ],
        "avg_utilization_pct": 64.3,
        "max_utilization_pct": 100.0,
        "max_memory_used_mib": 41088.0,
        "memory_total_mib": 81559.0,
        "max_temperature_c": 71.0,
        "max_power_w": 512.3
      }
    ]
  }
}

Each sample carries ts_ms (epoch milliseconds) on the same timeline as the stats lines, so throughput/latency and GPU state can be charted together. Absent optional fields mean the source reported N/A for that metric. Markdown exports (--summary-export out.md) get one compact row per device per aggregate.

GPU state also streams during the run: with report.during_run on, every engine snapshot carries the GPU samples taken since the previous one — shipped to the collector as gpu_utilization_pct, gpu_memory_used_mib, gpu_memory_total_mib, gpu_temperature_c and gpu_power_w gauges with a gpu="<index>" label, each at the collector's own timestamp. See during-run metrics for the full naming table and delivery semantics.

Example: Ollama under load, GPU watch on

# config.yaml
vus: 8
duration: 2m
gpu:
  enabled: true
# test.yaml
steps:
  - name: llama completion
    use: std/llm@v1
    with:
      url: http://127.0.0.1:11434/v1/chat/completions
      model: llama3.1
      prompt: "Summarize the CAP theorem in two sentences."
      max_tokens: 128
    check:
      status: 200
perfscale run -f test.yaml -c config.yaml --summary-export gpu-run.json

Reading the result together: llm_tokens_per_sec flat while gpu0 util avg sits at ~100% → the GPU is the bottleneck (add a card, shard the model, or lower vus); util well below 100% with rising TTFT → look at the server (queueing, context limits) instead. VRAM creeping to memory_total_mib explains evictions/OOMs mid-run.

Example: game-style rendering load

GPU load testing is not only about LLM servers. The other classic question is session density: how many concurrent render sessions — game clients on a cloud-gaming node, streaming viewports, digital-twin renderers — one card carries before the frame rate collapses. The pattern is the same as above, except the "system under test" is a set of renderer processes that perfscale orchestrates with std/child_process@v1 while the gpu: sampler records what the card is doing. A before: sidecar adds the second ingredient of a real node — a heavy background GPU job (the encode stage) competing with the sessions for the same card.

Any renderer that prints FPS works. This example uses glmark2 (OpenGL, apt install glmark2) looping its 3D scenes as a stand-in for a game client; vkmark is the Vulkan equivalent, and a headless Unity/Unreal build drops in unchanged — only the command differs.

# config.yaml
vus: 1                      # the renderers are the load; one VU just keeps
duration: 5m                # the run open (see test.yaml below)
allow_process_actions: true # required for child_process/kill_process

gpu:
  enabled: true
  interval_ms: 1000

before:
  # One "game session" = one renderer process. The farm spawns as a single
  # managed process group, so `after:` stops every session at once.
  - name: render-farm
    uses: std/child_process@v1
    with:
      command: sh
      args: ["-c", "for i in $(seq 4); do glmark2 --run-forever & done; wait"]
      waitUntil:
        stdout_contains: GL_RENDERER  # GL context is up
        on_timeout: continue
      restart: never                  # a crashed session must not respawn a 2nd farm

  # Sidecar: a heavy GPU job sharing the card with the sessions — the encode
  # stage of a game-streaming pipeline. Looped 1080p60 from a generated
  # source, NVENC-encoded, discarded to null.
  - name: encode-sidecar
    uses: std/child_process@v1
    with:
      command: ffmpeg
      args: ["-f", "lavfi", "-i", "testsrc2=size=1920x1080:rate=60",
             "-c:v", "h264_nvenc", "-f", "null", "-"]
      waitUntil:
        stderr_contains: "Press [q]"  # ffmpeg reports the running loop on stderr
        on_timeout: continue
      restart: on-failure             # a crashed encoder comes back

after:
  - name: stop the farm
    uses: std/kill_process@v1
    with: { name: render-farm }       # tree: true by default → every session
  - name: stop the sidecar
    uses: std/kill_process@v1
    with: { name: encode-sidecar }
# test.yaml
steps:
  - name: keep the run open
    use: std/sleep@v1
    with: { seconds: 30 }
perfscale run -f test.yaml -c config.yaml --summary-export render.json

Headless nodes: glmark2 needs a GL context. On a GPU server without a display use the DRM build (glmark2-drm / glmark2-es2-drm — renders via GBM straight on the card) or wrap the command in xvfb-run -a.

Reading the result

The renderer's FPS lines stream into the run log with a render-farm: prefix; the gpu: summary records what the card did meanwhile. The method is a sweep, not a single run — raise the session count (seq 4 → 1, 2, 4, 8) between runs:

  • per-scene FPS divided by ~N while util max pins at 100% → the GPU is saturated; that session count is the card's ceiling for this workload;
  • FPS degrades while util stays below 100% → the limit is elsewhere (CPU, context switching) — cross-check temp max / power max for thermal or power throttling;
  • vram max per session count answers the capacity question directly: how many sessions fit into memory_total_mib before the driver starts swapping;
  • the sidecar's price is the FPS delta between a farm-only run (comment the encode-sidecar block out) and a farm+sidecar run at the same session count — that is what sharing the card with the encode pipeline costs a cloud-gaming node.

perfscale does not parse FPS into metrics — frame-rate numbers live in the run log; the gpu: timeseries (each sample stamped ts_ms on the stats timeline) is what you chart against them.

One caveat for runs like this: utilization.gpu reports the 3D/compute engine — the NVENC encoder the sidecar burns is not counted there, so judge the sidecar by VRAM, power draw, and the FPS it costs the sessions, not by the util line.

Example: ANE (Neural Engine) load via Core ML

On Apple Silicon the NPU question mirrors the GPU one: does the workload actually engage the Neural Engine, and where is its ceiling? The shipped examples/coreml-ane/ generates real ANE load with a Core ML sidecar — a small conv net built with coremltools' MIL builder (no model download), looping predictions with compute_units=ALL — while gpu.source: powermetrics records the SoC:

# config.yaml (excerpt)
gpu:
  enabled: true
  source: powermetrics
  interval_ms: 5000

before:
  - name: ane-sidecar
    uses: std/child_process@v1
    with:
      command: sh
      args: ["-c", ".venv/bin/python ane_load.py"]
      waitUntil: { stdout_contains: "inference loop", on_timeout: fail }
      restart: never

after:
  - name: stop the sidecar
    uses: std/kill_process@v1
    with: { name: ane-sidecar }

Verified on an M2 Pro: ane_power_w rises ≈2 W over baseline (to ≈7 W sustained) while the loop runs at ≈1 450 inferences/s on the built-in net, with gpu_utilization_pct and gpu_power_w staying flat — proof the load is on the Neural Engine, not the GPU. (powermetrics models several watts of ANE draw even at idle on M2 Pro/Max, so read the delta, not the floor.) Reading the result:

  • ane_power_w flat at its baseline → the ANE is not engaged: the model is CPU/GPU-bound or was loaded without compute_units=ALL. (Python's data stack — pandas, NumPy, Anaconda — never touches the ANE; coremltools/Core ML is the only path user code has to it.)
  • Scale out (more sidecars, --size 224, or your own --model x.mlpackage) and watch where inferences/sec stop scaling — the chip's ANE ceiling for that workload.

Setup (venv + the powermetrics sudoers rule) and the sweep method are in the example's README.

GPU benchmark suite

The repo ships a ready-made local suite in bench/gpu/: std/llm@v1 scenarios against Ollama and vLLM with gpu: metrics on, a stepped ramping-VU profile (concurrency vs tok/s / TTFT degradation) and an arrival-rate profile (find the rate where TTFT and dropped_iterations climb). It is local-only — CI runners have no GPU.

bench/gpu/run.sh ollama        # or: vllm, or both

…runs each profile, writes --summary-export JSONs and raw logs to bench/gpu/results/<timestamp>/, and prints a compact table:

scenario        reqs  req/s  tok/s avg  ttft p50 ms  ttft p95 ms  gpu util max  vram max MiB  dropped
--------------  ----  -----  ---------  -----------  -----------  ------------  ------------  -------
ollama-stages   152   0.42   38.71      212.40       890.15       100%          5104          0
ollama-arrival  210   0.63   31.05      340.72       2410.30      100%          5112          17

Setup (Ollama / vLLM), requirements, and how to read the numbers: bench/gpu/README.md.

Extension seam

perfscale-core exposes gpu::GpuCollector and gpu::register_gpu_collector — the same pattern as register_pubsub_driver: a downstream (proprietary) build can register richer collectors (NVML-based per-process memory, SM clock/throttle reasons, rocm-smi for AMD, powermetrics for Apple silicon) and select them via gpu.source, or shadow the built-ins under their own names. The basic metrics above are the OSS baseline every collector reports.