Complete NVIDIA DCGM keyboard shortcuts and commands reference — 14 shortcuts across 2 categories. Quick reference cheat sheet for Windows & Mac.
DCGM is NVIDIA's data-center GPU manager, and dcgmi is its command line: health checks, diagnostics at three depths, per-job statistics and live monitoring for every GPU in a node. The table lists the commands an operator runs on a DGX or any multi-GPU server; the notes explain what each diagnostic level costs and how groups work.
| Shortcut | Action | Description |
|---|---|---|
| dcgmi diag -r 1 | Quick diagnostic | Run the fastest Level 1 diagnostic, completing in a few seconds. |
| dcgmi diag -r 2 | Medium diagnostic | Run the Level 2 diagnostic, adding a short stress test (about 2 minutes). |
| dcgmi diag -r 3 | Full diagnostic | Run the Level 3 diagnostic with extended stress tests (about 15 minutes) - the standard evidence to collect before an RMA. |
| dcgmi health -s a | Start health watch | Enable background health monitoring for all systems: PCIe, memory, power, and thermal. |
| dcgmi health -c | Check health | Run a health check against all discovered GPUs and report any active issues. |
| Shortcut | Action | Description |
|---|---|---|
| dcgmi discovery -l | List GPUs | List all GPUs discovered by DCGM on this host. |
| dcgmi dmon | Device monitor | Stream live per-GPU metrics - utilization, memory, temperature - to the terminal. |
| dcgmi group -c name | Create group | Create a named group of GPUs to manage and query together. |
| dcgmi group -g id -a 0,1 | Add GPUs to group | Add specific GPU IDs to an existing DCGM group. |
| dcgmi stats -e | Enable job stats | Enable per-job statistics collection for a GPU group. |
| dcgmi config --get | Show config | Display the current DCGM configuration applied to all GPUs. |
| dcgmi topo -g id | Group topology | Show the NVLink and PCIe topology for GPUs in a group. |
| nv-hostengine | Start host engine | Start the DCGM host engine daemon; required before dcgmi can connect. |
| dcgmi -v | Version | Print the installed DCGM version. |
The most essential NVIDIA DCGM shortcuts are: dcgmi diag -r 1 (Quick diagnostic), dcgmi diag -r 2 (Medium diagnostic), dcgmi diag -r 3 (Full diagnostic).
These are command-line commands — type them in your terminal or console. Combine them with shell history search (Ctrl + R) and aliases to work even faster.
The NVIDIA DCGM shortcut for quick diagnostic is dcgmi diag -r 1. Run the fastest Level 1 diagnostic, completing in a few seconds.
Yes — use My Stack to combine NVIDIA DCGM shortcuts with any other platform on this site into one printable reference, which is useful if your daily workflow spans several tools.
dcgmi diag -r 1 is the quick check — a few seconds, deployment sanity — and dcgmi diag -r 2 adds memory and PCIe bandwidth tests taking a couple of minutes. dcgmi diag -r 3 runs the full suite including stress tests and can take fifteen minutes or more per node, so it belongs in a maintenance window or a burn-in, not on a node with jobs. All three need nv-hostengine running, which is the DCGM daemon most GPU nodes start at boot. dcgmi -v prints the version to check against the driver.
dcgmi health -s a enables health watches for all subsystems (memory, PCIe, NVLink, thermal, power) and dcgmi health -c reports what those watches have found since — a cheap check to run before each job on a cluster. dcgmi dmon streams metrics such as utilisation, memory, temperature and power per GPU at a chosen interval, which is the DCGM equivalent of nvidia-smi dmon with more fields.
dcgmi discovery -l lists GPUs with their DCGM IDs. Many commands act on a group: dcgmi group -c name creates one and dcgmi group -g id -a 0,1 adds GPUs, after which diagnostics and stats can target that subset. dcgmi stats -e enables job statistics so a scheduler can record GPU usage per job, and dcgmi topo -g id shows the NVLink and PCIe topology within a group — the data behind choosing which GPUs to co-allocate. dcgmi config --get reads current settings such as clocks and compute mode.
Open your assistant with this page preloaded as the source — great for follow-up questions like "which of these work in other apps?"