OSU Micro-Benchmarks Keyboard Shortcuts

Complete OSU Micro-Benchmarks keyboard shortcuts and commands reference — 12 shortcuts across 3 categories. Quick reference cheat sheet for Windows & Mac.

The OSU Micro-Benchmarks measure what an MPI library and its network actually deliver — latency, bandwidth, collective times — and they are the first thing to run on a new cluster before any application benchmark. The table lists the binaries; the notes explain which ones answer which question and how to run them on GPUs.

Point-to-Point Tests (5)

ShortcutActionDescription
osu_latencyLatency testMeasure round-trip latency for a single point-to-point message between two ranks.
osu_bwBandwidth testMeasure unidirectional bandwidth between two ranks.
osu_bibwBidirectional bandwidthMeasure bidirectional bandwidth with both ranks sending simultaneously.
osu_multi_latMulti-pair latencyMeasure latency across multiple simultaneous rank pairs.
osu_mbw_mrMulti bandwidth/message rateMeasure aggregate bandwidth and message rate across multiple pairs.

Collective Tests (5)

ShortcutActionDescription
osu_allreduceAllreduce latencyMeasure MPI_Allreduce collective latency across all ranks.
osu_allgatherAllgather latencyMeasure MPI_Allgather collective latency.
osu_alltoallAlltoall latencyMeasure MPI_Alltoall collective latency.
osu_bcastBroadcast latencyMeasure MPI_Bcast collective latency from a root rank.
osu_reduceReduce latencyMeasure MPI_Reduce collective latency to a root rank.

GPU & Launch Options (2)

ShortcutActionDescription
mpirun -np 2 osu_bw d dGPU-to-GPU testRun the bandwidth test with buffers on the accelerator device ('D') instead of host memory ('H').
osu_latency -m 2:128Set message rangeLimit the test to message sizes between 2 and 128 bytes.
📜 Source: Ohio State University — OSU Micro-Benchmarks. Benchmark list and options from the OSU Micro-Benchmarks README for the current release. Checked 2026-09-06. How we verify ›
📄 View Printable Cheat Sheet — Download as PDF or print · 🧩 Combine with other tools

Frequently Asked Questions

What are the most useful OSU Micro-Benchmarks keyboard shortcuts?

The most essential OSU Micro-Benchmarks shortcuts are: osu_latency (Latency test), osu_bw (Bandwidth test), osu_bibw (Bidirectional bandwidth).

How do I use OSU Micro-Benchmarks commands?

These are command-line commands — type them in your terminal or console. Combine them with shell history search (Ctrl + R) and aliases to work even faster.

What is the OSU Micro-Benchmarks shortcut for latency test?

The OSU Micro-Benchmarks shortcut for latency test is osu_latency. Measure round-trip latency for a single point-to-point message between two ranks.

What Point-to-Point Tests shortcuts does OSU Micro-Benchmarks have?

OSU Micro-Benchmarks includes 5 Point-to-Point Tests shortcuts, including osu_latency (Latency test) and osu_bw (Bandwidth test). See the full list in the Point-to-Point Tests section above.

Can I combine OSU Micro-Benchmarks shortcuts with other tools?

Yes — use My Stack to combine OSU Micro-Benchmarks shortcuts with any other platform on this site into one printable reference, which is useful if your daily workflow spans several tools.

Related Shortcut Pages

OpenMPI (mpirun) Mellanox / InfiniBand NCCL Tests Intel MPI Benchmarks (IMB) Slurm

Search 18,500+ shortcuts across 268 platforms

Explore All Platforms Practice Shortcuts
🔧 Spotted an error or a missing shortcut? Suggest an edit on GitHub — every accepted fix goes live on this page, the API and the CLI.

Point to point

osu_latency between two ranks on different nodes gives the round-trip half-time per message size; on a healthy InfiniBand fabric the small-message number is a few microseconds and anything higher points at a configuration problem. osu_bw and osu_bibw measure one-way and bidirectional bandwidth, which should approach the link rate for large messages. osu_multi_lat and osu_mbw_mr run several pairs at once for aggregate bandwidth and message rate, which is closer to what a real job does. osu_latency -m 2:128 restricts the message size range, useful for focusing on the small-message regime.

Collectives

osu_allreduce is the one that predicts deep-learning performance, since gradient synchronisation is an allreduce; run it at the node count and message sizes your job uses. osu_allgather, osu_alltoall, osu_bcast and osu_reduce cover the other patterns, with alltoall the most sensitive to network topology. Compare results at 2, 4, 8 nodes to see how the fabric scales.

GPUs

Built with CUDA support, the benchmarks take buffer-location arguments: mpirun -np 2 osu_bw d d measures device-to-device bandwidth (d for device memory, h for host, m for managed), which tests GPUDirect RDMA and is the number to check before trusting NCCL results. Launch through mpirun or srun with the same placement and binding flags as a real job, since binding changes the answer.

🤖 Ask AI about OSU Micro-Benchmarks shortcuts

Open your assistant with this page preloaded as the source — great for follow-up questions like "which of these work in other apps?"

ChatGPT Claude Perplexity Gemini Grok