Other Hugging Face Blog

How to Use NVIDIA Warp and MjWarp to Accelerate Robotics Simulation and Learning Workflows

NVIDIA WarpMJWarpMuJoCorobotics simulation

Classic MuJoCo provides fast CPU-based robot simulation for developing, testing, and controlling robots, and it can parallelize sampling across CPU cores. As learning workloads grow, the question shifts from how quickly one world can run to how many worlds can run at once. GPU acceleration makes it possible to advance those worlds in large batches while keeping simulation and learning data close to the device.

MuJoCo Warp (MJWarp), built on NVIDIA Warp, takes compatible MuJoCo models into that GPU-scale regime. The article moves an SO-101 follower arm from a familiar MuJoCo workflow to as many as 2,048 parallel MJWarp environments and examines the technology and validation steps that make the transition possible. Figure 1 shows how MJWarp connects Python to GPU simulation: MuJoCo loads and compiles the MJCF model, while MJWarp implements the physics in NVIDIA Warp, which compiles CUDA kernels to advance simulation states on NVIDIA GPUs.

This is the second article in the State of Simulation for Physical AI series; the first mapped the robot-simulation landscape. This installment prepares and scales the simulation environment but does not train a policy, and later Newton and Isaac Lab installments cover the next integration layers. The described stack has four layers: NVIDIA Warp as the Python kernel language with SIMT, autodiff, and PyTorch/JAX interop; MJWarp as MuJoCo physics on Warp with the same MJCF and batched GPU throughput; the SO-101 scene using familiar Menagerie/Robot Studio assets plus task geometry; and next-step Newton/Isaac Lab for multi-solver API, USD, sensors, managers, and training loops.

The article offers a decision shortcut: use MuJoCo CPU for single-robot MPC/teleop; MJWarp (or mjlab) for maximum throughput on raw MuJoCo physics; MuJoCo Playground or MJX (impl='warp') for JAX training recipes; and Newton for multi-solver plus Isaac Lab integration in the next post. It then starts with one useful Warp kernel. NVIDIA Warp is a Python framework for writing high-performance, GPU-accelerated kernels, letting developers author statically typed kernels in Python and compile them for CPU or CUDA execution. The first launch builds and caches a native module, and later launches reuse it.

The kernel language is a performance-oriented subset of Python, while ordinary Python remains responsible for configuration, allocation, and launch orchestration. A small robotics-oriented kernel example advances point positions under gravity; one logical thread handles one point, so the same code scales from two points to millions without introducing GPU terminology into the control flow. Warp's three value propositions are performance (native-CUDA speed via JIT compilation, kernel fusion, and CUDA Graphs), ease of use (pure Python authoring with built-in vectors, matrices, quaternions, BVHs, hash grids, sparse matrices, and tile primitives), and capability (differentiable kernels and DLPack-style interop so simulation can sit inside an ML training loop).

Read original →

← Back to home