Complete HPL (Linpack) keyboard shortcuts and commands reference — 11 shortcuts across 2 categories. Quick reference cheat sheet for Windows & Mac.
HPL — the Linpack benchmark behind the TOP500 — is one executable, xhpl, launched through MPI and configured entirely by a text file, HPL.dat. The commands here are the launch variants and the checks around them. The notes explain the two things that decide the score: process placement and the problem size, and how to read the result.
| Shortcut | Action | Description |
|---|---|---|
| mpirun -np 16 ./xhpl | Run HPL benchmark | Launch the Linpack solver across 16 MPI ranks, reading HPL.dat for configuration. |
| mpirun -np 8 --map-by l3cache ./xhpl | Cache-aware mapping | Bind ranks to L3 cache domains for better memory locality on multi-socket nodes. |
| mpirun -np 8 --bind-to core ./xhpl | Core binding | Pin each MPI rank to a dedicated CPU core for consistent performance. |
| srun -N 4 --ntasks-per-node=8 ./xhpl | Run via Slurm | Launch the HPL benchmark through a Slurm job allocation. |
| mpiexec -f machinefile -n 32 ./xhpl | Run with machinefile | Launch across the nodes listed in an MPI machinefile. |
| ldd xhpl | Check library links | Verify the compiled binary can find its BLAS and MPI shared libraries before running. |
| Shortcut | Action | Description |
|---|---|---|
| vi HPL.dat | Edit configuration | Edit the HPL.dat input file controlling problem size (N), block size (NB), and process grid (P x Q). |
| ibcheckerrors | Check IB fabric | Confirm the InfiniBand fabric reported no errors during the benchmark run. |
| top | Monitor CPU usage | Confirm all MPI processes are near 100% CPU utilization during the run. |
| grep Gflops HPL.out | Read result | Extract the achieved GFLOPS figure from the HPL output log. |
| singularity run hpc-benchmarks.sif ./hpl.sh --dat file | NVIDIA container run | Run NVIDIA's optimized HPL container build, common on GPU clusters via NGC. |
The most essential HPL (Linpack) shortcuts are: mpirun -np 16 ./xhpl (Run HPL benchmark), mpirun -np 8 --map-by l3cache ./xhpl (Cache-aware mapping), mpirun -np 8 --bind-to core ./xhpl (Core binding).
These are command-line commands — type them in your terminal or console. Combine them with shell history search (Ctrl + R) and aliases to work even faster.
The HPL (Linpack) shortcut for run hpl benchmark is mpirun -np 16 ./xhpl. Launch the Linpack solver across 16 MPI ranks, reading HPL.dat for configuration.
Yes — use My Stack to combine HPL (Linpack) shortcuts with any other platform on this site into one printable reference, which is useful if your daily workflow spans several tools.
mpirun -np 16 ./xhpl runs sixteen MPI ranks on the local machine or the hosts MPI knows about; mpiexec -f machinefile -n 32 ./xhpl is the MPICH form with an explicit host list, and srun -N 4 --ntasks-per-node=8 ./xhpl launches under Slurm. Placement matters more than people expect: mpirun -np 8 --bind-to core ./xhpl pins each rank to a core so the OS cannot migrate them, and mpirun -np 8 --map-by l3cache ./xhpl spreads ranks across cache domains, which is the usual best setting on multi-chiplet CPUs. ldd xhpl confirms the binary is linked against the intended BLAS (MKL, OpenBLAS, BLIS), because the wrong library halves the result.
vi HPL.dat is where N (problem size), NB (block size) and the P×Q process grid are set. N should use roughly 80 percent of total memory for a top score, NB is typically 192–256 for modern CPUs, and P×Q must equal the rank count with P ≤ Q. Run small N first to check correctness, then scale. On a cluster, ibcheckerrors checks the InfiniBand fabric for link errors before a long run, since a flaky link shows up as a mysteriously low score rather than a failure.
grep Gflops HPL.out pulls the performance line from the output; the residual check on the following lines must say PASSED or the run is invalid. top during the run should show every core busy — idle cores mean a binding or thread-count problem. For GPU systems, singularity run hpc-benchmarks.sif ./hpl.sh --dat file runs NVIDIA's container build of HPL, which is the supported route for DGX-class results.
Open your assistant with this page preloaded as the source — great for follow-up questions like "which of these work in other apps?"