Complete torchrun keyboard shortcuts and commands reference — 12 shortcuts across 2 categories. Quick reference cheat sheet for Windows & Mac.
torchrun is PyTorch's launcher for distributed training: it starts one process per GPU, hands each one its rank and world size through environment variables, and restarts workers when elastic settings allow. The flags below are the ones every launch needs; the notes explain how a single-node command grows into a multi-node one and what the rendezvous options actually do.
| Shortcut | Action | Description |
|---|---|---|
| torchrun script.py | Single-GPU run | Run a script with the default single-process, single-GPU configuration. |
| --nproc_per_node=N | Processes per node | Number of processes - typically one per GPU - to launch on each node. |
| --nnodes=N | Node count | Total number of nodes participating in the distributed job. |
| --node_rank=N | Node rank | This node's rank (0-indexed) among all participating nodes; set differently on each machine. |
| --master_addr=IP | Master address | IP address of the rank-0 node that other workers connect to for coordination. |
| --master_port=PORT | Master port | TCP port on the master node used for the rendezvous connection. |
| Shortcut | Action | Description |
|---|---|---|
| --rdzv_backend=c10d | Rendezvous backend | Coordination backend for elastic launches; c10d is the recommended default. |
| --rdzv_id=JOBID | Rendezvous ID | Unique identifier for the job, shared by all nodes so they can find each other. |
| --rdzv_endpoint=IP:PORT | Rendezvous endpoint | Address workers use to join the rendezvous, an alternative to setting master_addr/master_port. |
| --standalone | Standalone mode | Run a single-node job without needing any rendezvous configuration. |
| --max_restarts=N | Max restarts | Maximum number of worker group restarts allowed before the job is marked failed. |
| --log_dir=PATH | Log directory | Directory where per-worker stdout and stderr logs are written. |
The most essential torchrun shortcuts are: torchrun script.py (Single-GPU run), --nproc_per_node=N (Processes per node), --nnodes=N (Node count).
These are command-line commands — type them in your terminal or console. Combine them with shell history search (Ctrl + R) and aliases to work even faster.
The torchrun shortcut for single-gpu run is torchrun script.py. Run a script with the default single-process, single-GPU configuration.
Yes — use My Stack to combine torchrun shortcuts with any other platform on this site into one printable reference, which is useful if your daily workflow spans several tools.
torchrun script.py runs a script with one process, which is a quick way to check that the script reads LOCAL_RANK correctly before scaling. --nproc_per_node=N starts N processes on the machine — set it to the GPU count — and torchrun sets RANK, LOCAL_RANK and WORLD_SIZE for each, so the script should call torch.distributed.init_process_group() without arguments and pick its device from LOCAL_RANK. --standalone tells torchrun to host the rendezvous itself on a free port, which is what you want for any single-node job and removes the master address and port flags.
Across machines, every node runs the same command with --nnodes=N and a different --node_rank=N (0 on the first). The classic way to let them find each other is --master_addr=IP and --master_port=PORT pointing at node 0; the elastic way is --rdzv_backend=c10d with --rdzv_endpoint=IP:PORT and a shared --rdzv_id=JOBID, which also lets nodes join late or be replaced. Under Slurm, --nnodes and --node_rank are usually derived from SLURM_NNODES and SLURM_NODEID in the batch script.
--max_restarts=N allows torchrun to restart the whole worker group that many times after a failure, which combined with checkpointing gives fault tolerance on flaky clusters; the default is 0, meaning any worker crash ends the job. --log_dir=PATH writes each worker's stdout and stderr to separate files under that directory, which is the only sane way to read logs from 64 processes. NCCL problems show up as timeouts at init; NCCL_DEBUG=INFO in the environment is the first thing to add.
Open your assistant with this page preloaded as the source — great for follow-up questions like "which of these work in other apps?"