// EMPIRICAL REFERENCE · NVIDIA-SMI -LGC ON GB10
nvidia-smi --lock-gpu-clocks on GB10 (DGX Spark)An empirical reference for the one working clock actuator on NVIDIA's
GB10 Superchip (Blackwell architecture, sm_121). Written because the
NVIDIA docs are terse on this silicon and the standard clock-inspection
commands return [N/A].
sudo nvidia-smi -lgc <min_mhz>,<max_mhz> [-i <gpu_index>] sudo nvidia-smi -rgc [-i <gpu_index>] # reset to stock # long form sudo nvidia-smi --lock-gpu-clocks=<min>,<max> sudo nvidia-smi --reset-gpu-clocks
The daemon that consumes this lever calls it every 30 seconds while it
walks the ceiling. The reset form is invoked in ExecStopPost
so a clean stop returns the GPU to stock.
| Bound | Value | Notes |
|---|---|---|
| Floor accepted | ~900 MHz | Below this, the driver refuses. |
| Ceiling accepted | ~3003 MHz | Reference silicon max. Values above are clamped. |
| Step granularity | 1 MHz | Any integer in range works. Round to 15 or 150 MHz for stability. |
--query-supported-clocks returns$ nvidia-smi --query-supported-clocks=graphics --format=csv,noheader [N/A]
The driver does not enumerate supported clocks on GB10, but it still
accepts --lock-gpu-clocks. That is: the lack of an enumerated
range is not a hint that -lgc is unsupported. It is
unsupported for enumeration, not for actuation.
For a daemon that needs unattended access to -lgc and
-rgc, keep the scope tight. This is the exact rule our
installer writes into /etc/sudoers.d/zc-thermal-control:
<user> ALL=(root) NOPASSWD: \
/usr/bin/nvidia-smi -lgc *, \
/usr/bin/nvidia-smi -rgc, \
/usr/bin/nvidia-smi --lock-gpu-clocks=*, \
/usr/bin/nvidia-smi --reset-gpu-clocks
No other nvidia-smi flag can be run passwordless from that
user. If the daemon binary is compromised, the attacker gets clock control
— not a full-fat nvidia-smi root shell.
Multiple processes calling -lgc in sequence: last-writer
wins. A subsequent -lgc 900,3003 after a
-lgc 1800,2400 immediately widens the band. The driver does
not queue or reject on contention.
A crash of the calling process does not release the lock. The
clock ceiling persists until either an explicit -rgc or a
driver reload. This is why our systemd unit calls
ExecStopPost=nvidia-smi -rgc unconditionally — even on
crash, the daemon's supervisor returns the GPU to stock.
On sustained inference workloads, holding the ceiling at 1800 MHz vs stock ~2463 MHz is a ~27% worst-case clock reduction. In practice, an adaptive controller that walks the ceiling up on cool samples rides much higher than the floor. Our observed median inference speed hit is 5–15% on Ollama gpt-oss:120b + qwen2.5:72b, which we trade for 24/7 uptime and ~11 °C of sustained thermal headroom.
See also: Why the DGX Spark's GB10 runs hot · Live telemetry · product page. Published by C.R. Burrell LLC. Corrections: [email protected].