// EMPIRICAL REFERENCE · NVIDIA-SMI -LGC ON GB10

nvidia-smi --lock-gpu-clocks on GB10 (DGX Spark)

An empirical reference for the one working clock actuator on NVIDIA's GB10 Superchip (Blackwell architecture, sm_121). Written because the NVIDIA docs are terse on this silicon and the standard clock-inspection commands return [N/A].

Command

sudo nvidia-smi -lgc <min_mhz>,<max_mhz> [-i <gpu_index>]
sudo nvidia-smi -rgc  [-i <gpu_index>]         # reset to stock

# long form
sudo nvidia-smi --lock-gpu-clocks=<min>,<max>
sudo nvidia-smi --reset-gpu-clocks

The daemon that consumes this lever calls it every 30 seconds while it walks the ceiling. The reset form is invoked in ExecStopPost so a clean stop returns the GPU to stock.

Accepted range on GB10

BoundValueNotes
Floor accepted~900 MHzBelow this, the driver refuses.
Ceiling accepted~3003 MHzReference silicon max. Values above are clamped.
Step granularity1 MHzAny integer in range works. Round to 15 or 150 MHz for stability.

What --query-supported-clocks returns

$ nvidia-smi --query-supported-clocks=graphics --format=csv,noheader
[N/A]

The driver does not enumerate supported clocks on GB10, but it still accepts --lock-gpu-clocks. That is: the lack of an enumerated range is not a hint that -lgc is unsupported. It is unsupported for enumeration, not for actuation.

Sudoers scoping

For a daemon that needs unattended access to -lgc and -rgc, keep the scope tight. This is the exact rule our installer writes into /etc/sudoers.d/zc-thermal-control:

<user> ALL=(root) NOPASSWD: \
    /usr/bin/nvidia-smi -lgc *, \
    /usr/bin/nvidia-smi -rgc, \
    /usr/bin/nvidia-smi --lock-gpu-clocks=*, \
    /usr/bin/nvidia-smi --reset-gpu-clocks

No other nvidia-smi flag can be run passwordless from that user. If the daemon binary is compromised, the attacker gets clock control — not a full-fat nvidia-smi root shell.

Behavior under contention

Multiple processes calling -lgc in sequence: last-writer wins. A subsequent -lgc 900,3003 after a -lgc 1800,2400 immediately widens the band. The driver does not queue or reject on contention.

A crash of the calling process does not release the lock. The clock ceiling persists until either an explicit -rgc or a driver reload. This is why our systemd unit calls ExecStopPost=nvidia-smi -rgc unconditionally — even on crash, the daemon's supervisor returns the GPU to stock.

Cost of clock-locking

On sustained inference workloads, holding the ceiling at 1800 MHz vs stock ~2463 MHz is a ~27% worst-case clock reduction. In practice, an adaptive controller that walks the ceiling up on cool samples rides much higher than the floor. Our observed median inference speed hit is 5–15% on Ollama gpt-oss:120b + qwen2.5:72b, which we trade for 24/7 uptime and ~11 °C of sustained thermal headroom.


See also: Why the DGX Spark's GB10 runs hot · Live telemetry · product page. Published by C.R. Burrell LLC. Corrections: [email protected].