经典的 MuJoCo 提供了基于 CPU 的快速机器人仿真,用于开发、测试和控制机器人,并且可以在 CPU 核心之间并行化采样。但随着学习工作负载的增长,问题从“一个世界能运行得多快”转变为“能同时运行多少个世界”。GPU 加速使得能够以大批量推进这些世界,同时保持仿真和学习数据靠近设备。
MuJoCo Warp (MJWarp) 基于 NVIDIA Warp 构建,将兼容的 MuJoCo 模型带入 GPU 规模领域。在本文中,我们将把一个 SO-101 跟随臂从熟悉的 MuJoCo 工作流程转移到多达 2,048 个并行 MJWarp 环境中,并考察使这种转变成为可能的技术和验证步骤。
图 1. MJWarp 如何将 Python 连接到 GPU 仿真。MuJoCo 加载并编译 MJCF 模型;MJWarp 在 NVIDIA Warp 中实现物理引擎,后者编译 CUDA 内核以在 NVIDIA GPU 上推进仿真状态。
这是我们的“物理 AI 仿真现状”系列的第二篇文章。第一篇文章描绘了机器人仿真 landscape。在这里,我们准备并扩展仿真环境;我们不训练策略。后续的 Newton 和 Isaac Lab 部分将涵盖下一层集成。
整合在一起
层级
在堆栈中的作用
NVIDIA Warp
Python 内核语言:单指令多线程 (SIMT)、自动微分、PyTorch/JAX 互操作
MJWarp
Warp 上的 MuJoCo 物理引擎:相同的 MJCF,批处理 GPU 吞吐量
你的场景 (SO-101)
熟悉的 Menagerie / Robot Studio 资产 + 任务几何结构
下一步 (Newton / Isaac Lab)
多求解器 API、USD、传感器、管理器、训练循环
决策捷径:
如果你需要……
选择……
单机器人 MPC / 遥操作
MuJoCo CPU
原始 MuJoCo 物理引擎的最大吞吐量
MJWarp (或 mjlab )
JAX 训练配方
MuJoCo Playground / MJX (impl='warp')
多求解器 + Isaac Lab 集成
Newton — 本系列的下一篇帖子
从一个有用的 Warp 内核开始
NVIDIA Warp 是一个用于编写高性能、GPU 加速内核的 Python 框架。Warp 允许开发者用 Python 编写静态类型的内核,并将其编译为 CPU 或 CUDA 执行。首次启动会构建并缓存原生模块;后续启动将重用该模块。内核语言是面向性能的 Python 子集,而普通 Python 仍负责配置、分配和启动编排。
这个小型的机器人导向内核在重力作用下推进点位置。一个逻辑线程处理一个点,因此相同的代码可以从两个点扩展到数百万个点,而无需在控制流中引入 GPU 术语。
Warp 的三个价值主张是:
支柱
你获得什么
性能
通过 JIT 编译、内核融合和 CUDA Graphs 实现原生 CUDA 速度
易用性
纯 Python 编写,内置向量、矩阵、四元数、BVH、哈希网格、稀疏矩阵和图块基元
能力
可微内核和 DLPack 风格的互操作,使仿真能够嵌入 ML 训练循环中
import numpy as np
import warp as wp
@wp.kernel
def integrate (
positions: wp.array[wp.vec3],
velocities: wp.array[wp.vec3],
dt: float ,
):
i = wp.tid()
velocities[i] += wp.vec3( 0.0 , 0.0 , - 9.81 ) * dt
positions[i] += velocities[i] * dt
wp.init()
device = "cuda:0" if wp.is_cuda_available() else "cpu"
start = np.array([[ 0.0 , 0.0 , 0.5 ], [ 0.2 , 0.0 , 0.5 ]], dtype=np.float32)
positions = wp.array(start, dtype=wp.vec3, device=device)
velocities = wp.zeros_like(positions)
wp.launch(
integrate,
dim= len (start),
inputs=[positions, velocities, 0.01 ],
device=device,
)
wp.synchronize_device(device)
print (positions.numpy())
三个特性使其在机器人领域具有实用性:
显式并行工作。wp.tid() 标识当前逻辑线程所拥有的点、接触点、刚体或世界。
显式设备数组。数组驻留在所选设备上。对 CUDA 数组调用 .numpy() 会同步并将其复制到 CPU 内存;它不是零拷贝路径。对于设备端驻留的 PyTorch 或 JAX 管道,请使用 Warp 的框架适配器或 DLPack 兼容共享方式。
可组合的内核启动。程序可以启动一系列聚焦的内核,并将支持的 CUDA 工作捕获到图中,以减少重复的分发开销。图捕获针对现有缓冲区重放启动操作;它不会融合任意内核。
可微性与确定性。
还有两项 Warp 功能值得了解,尽管本文的 SO-101 工作流中并未使用这两项功能。Warp 内核是可微的:wp.Tape 会记录其上下文内进行的正向内核启动,并在调用 backward() 时反向重放它们的伴随量,这就是团队在 Warp 中构建可微几何、计算流体动力学(CFD)和自定义物理的原因,包括用于仿真和设计优化的 CAE 工作流。Warp 还支持确定性执行,该功能在 Warp 1.15 中引入:GPU 原子操作默认依赖于调度器,因此重复启动相同的内核可能会略有不同,而可选的确定性模式以牺牲一些性能为代价,换取仿真、验证和回归测试中可重现的顺序。这些是 Warp 的功能,而非整个 MJWarp 部署中可微性或确定性的保证。有关详细信息,请参阅关于可微性和确定性执行的 Warp 文档。
试用 Warp:pip install warp-lang(≥ 1.15 版本支持 GPU 确定性),然后运行 python -m warp.examples.browse,或查看教程笔记本。
什么是 MuJoCo Warp (MJWarp)?
机器人模拟器反复计算下一步会发生什么:给定当前的关节位置、速度、控制和接触情况,它将场景向前推进一个微小的时间步长。在本文中,“世界”(world)指该场景及其状态的一个独立副本。一个世界可能包含 SO-101 机械臂伸手去抓立方体的情景;另一个世界则可能包含从略微不同姿态开始的同一机械臂。
MuJoCo 和 MJWarp 可以运行相同的兼容机器人和任务,但它们组织工作的方式不同。MuJoCo 天然适合开发和检查一个或少数几个 CPU 世界。MJWarp 是 MuJoCo 物理管道的 NVIDIA Warp 实现,它将模型和一批独立的状态放置在 NVIDIA GPU 上;一次 mjw.step 调用即可推进整个批次。
MJWarp 的价值不一定在于加速单个世界的步骤。而在于能够同时推进数百或数千个世界,为 GPU 提供足够的并行工作,从而提高聚合吞吐量(aggregate throughput),即每秒完成的总世界步数。这有利于强化学习和大规模采样,在这些场景中,收集经验比最小化单个环境的延迟更重要。
本文涵盖以下内容:
验证一个 MuJoCo 世界,
将其迁移到 MJWarp,形成批次,
进行验证并正确测量性能。
求解器调优、雅可比矩阵表示以及专用的多 GPU 或确定性主题并非此次迁移所必需,可以另行讨论。
因此,区别非常明确:
延迟(Latency)是指单个仿真步骤的墙钟时间。
聚合吞吐量是指每测量秒内完成的总世界步数。
基本用法:结构体、批次大小和最小化步骤
核心 API 过渡很小:
MuJoCo 主机工作流
MJWarp 工作流
mujoco.MjModel
mjw.put_model(mjm) 创建设备端模型
mujoco.MjData
mjw.put_data(mjm, mjd, ...) 保留并对现有状态进行批处理
mujoco.mj_step(mjm, mjd)
mjw.step(m, d) 推进 d 中的每个世界
主机数组,如 mjd.ctrl
设备端批处理数组,如形状为 (nworld, nu) 的 d.ctrl
当意图使用默认/全新状态时,请使用 mjw.make_data()。当必须将精确初始化的 MuJoCo 状态跨越迁移边界时,请使用 mjw.put_data()。
分配批处理资源需要定义以下参数(请参阅批次大小):
参数
含义
nworld
并行环境的总数
nconmax
每个独立世界预期的接触点数量(总容量 ≈ nconmax * nworld)
naconmax
替代设置:所有环境合并后的全局最大接触点数(如果两者都定义了,此设置优先)
njmax
每个世界的约束条件硬上限
性能调优
with wp.ScopedCapture() as capture:
mjw.step(m, d)
wp.capture_launch(capture.graph)
其他调优注意事项。在确定接触和约束缓冲区大小后,在不改变任务行为的情况下测试求解器迭代限制。网格和 CCD(连续碰撞检测)设置可能会增加内存使用量;当测量的接触点数允许时,nccdmax / naccdmax 可以减少 CCD 缓冲区的分配。MJWarp 的紧凑求解器使用 MuJoCo 的牛顿约束求解器和休眠机制,而不是独立的 Newton 物理引擎框架。紧凑求解器和多 GPU 配置超出了本教程的范围;请参阅 MJWarp 性能调优文档。
要在 MJWarp 物理环境中训练策略:
安装/尝试:pip install mujoco-warp · mjwarp-viewer path/to/scene.xml · Colab 教程
将 MuJoCo 场景迁移到 MjWarp 的工作流程
建立 MuJoCo CPU 基线
场景。目前这里没有任何 MJWarp 特定的内容:一个 SO-101 机械臂、一张桌子和两个用于堆叠的立方体,以普通的 MJCF 格式编写。
图 2. SO-101 抓取放置场景,由 MuJoCo CPU 仿真渲染。任务是将红色 44 毫米立方体抓起并堆叠在蓝色立方体上;相同的机器人和场景用于 MJWarp 验证。
< mujoco model = "so101_pick_place" >
< include file = "so101.xml" />
< worldbody >
< light pos = "0.3 0 1.5" dir = "0 0 -1" directional = "true" />
< geom name = "floor" type = "plane" size = "0 0 0.05" />
< geom name = "table" type = "box" pos = "0.35 -0.04 0.012"
size = "0.16 0.26 0.012" rgba = "0.32 0.32 0.32 1"
friction = "1 0.005 0.0005" condim = "3" />
< body name = "red_cube" pos = "0.33 -0.13 0.046" >
< freejoint name = "red_cube_joint" />
< geom type = "box" size = "0.022 0.022 0.022" mass = "0.08"
rgba = "0.85 0.05 0.04 1" friction = "1.2 0.005 0.0005" condim = "3" />
</ body >
< body name = "blue_cube" pos = "0.33 0.06 0.046" >
< freejoint name = "blue_cube_joint" />
< geom type = "box" size = "0.022 0.022 0.022" mass = "0.08"
rgba = "0.05 0.20 0.90 1" friction = "1.2 0.005 0.0005" condim = "3" />
</ body >
</ worldbody >
</ mujoco >
对于 MJCF 中的盒子,size 值为半轴长:size=”0.022 …”定义了一个边长为 44 毫米的立方体。任务使用此尺寸作为成功阈值。机械臂底座位于原点,其伸展方向沿 +X 轴,立方体沿 Y 轴排列。
在配套存储库中,该文件是生成的而非手动编写的:resolve_pick_place_scene() 将 Menagerie 机械臂复制到 .generated/ 目录中,从机器人配置文件中填充桌子和立方体的坐标,并写入 scene_pick_place.xml。本教程使用 SO-101 配置文件;可选的 reBot 变体将在下文描述。
加载它。编译和步进是标准的 MuJoCo 操作:
import mujoco
mjm = mujoco.MjModel.from_xml_path( "scene_pick_place.xml" )
mjd = mujoco.MjData(mjm)
fps = 50
sim_substeps = 10
frame_dt = 1.0 / fps
mjm.opt.timestep = frame_dt / sim_substeps
controller = PickPlaceController(spec=spec)
for _ in range ( 600 ):
ctrl = controller.step(mjm, mjd, frame_dt)
for _ in range (sim_substeps):
mjd.ctrl[: mjm.nu] = ctrl
mujoco.mj_step(mjm, mjd)
请记住这种结构:每帧计算一次控制指令,将物理仿真步进 sim_substeps 次。Gate 2 仅更改内部循环,这使得迁移过程易于审查。
保持仿真和控制速率一致。在每秒 50 个控制帧、每帧 10 个物理子步的情况下,使用 0.002 秒的物理时间步长。在 CPU rollout(回滚/推演)之前以及通过 mjw.put_model 上传模型之前设置该值,以确保两个后端推进相同的仿真时间:
mjm.opt.timestep = frame_dt / sim_substeps
如果没有这一行,后续的所有测量结果都会继承这种不匹配:奇偶校验比较、以“仿真秒”引用的吞吐量数据,以及任何动作速率不再与部署相匹配的学习策略。
检查方块是否堆叠成功。对于 44 毫米的方块,成功取决于两个可测量的条件:水平中心误差 xy_err ≤ 0.015 米(在方块中心之间测量)以及方块中心之间的垂直间距 0.035 米 ≤ dz ≤ 0.055 米(一个方块的边长,并为沉降留出余量)。在方块稳定后评估这两个条件;仅进程成功退出并不能确立任务成功。
从配套签出的代码库运行 CPU 任务。发布阻碍点:在发布这些说明之前,确认可访问的仓库 URL 以及固定的依赖项和资源版本;下方的仓库占位符不是可执行的 URL。
git clone https://github.com/NVIDIA/accelerated-computing-hub.git blogs
cd blogs/tutorials/sim2real-blogs/notebooks/mujoco
uv venv --python 3.12 && source .venv/bin/activate
uv pip install -r requirements.txt
cd /tutorials/sim2real-blogs/notebooks/mujoco
python solutions/so101_pick_place_solution.py --headless-steps 600 --debug
运行结束时将打印上述两个数字(堆叠检查:xy_err=… dz=…),这是文章其余部分进行比较的断言。旁边的 so101_pick_place.py 是同一个程序,其中的物理步骤留作练习。
机械臂直接来自 MuJoCo Menagerie,并固定在一个已知良好的提交上,因为 Menagerie 的资源会发生变化,因此请将场景视为模板。可选的 reBot 变体。配套代码还通过 --robot rebot 暴露了针对场景布局、夹爪和容量限制(nconmax=256, njmax=500)的独立配置文件。本教程使用 SO-101。在报告其结果之前,请单独验证 reBot 资源和任务。
验证单世界 MJWarp 奇偶校验
首先在 GPU 上运行一个世界,同时主机仍在循环中,以便您可以在同一个查看器中观察相同的任务并比较相同的两个数字。上传模型,分配批处理状态,从初始化的主机状态对其进行种子设置,并在步进之前运行一次前向传播:
wp.init()
import mujoco_warp as mjw
device = wp.get_device()
m = mjw.put_model(mjm)
d = mjw.make_data(mjm, nworld= 1 , nconmax=spec.nconmax, njmax=spec.njmax)
wp.copy(d.qpos, wp.array(mjd.qpos[ None , :], dtype=wp.float32, device=device))
wp.copy(d.qvel, wp.array(mjd.qvel[ None , :], dtype=wp.float32, device=device))
wp.copy(d.ctrl, wp.array(mjd.ctrl[ None , :], dtype=wp.float32, device=device))
mjw.forward(m, d)
每个设备数组都带有一个前导的世界维度,这就是为什么主机状态索引为 mjd.qpos[None, :],形状为 (1, nq) 而不是 (nq,) 的原因。稍后扩展到数千个世界时,仅改变那个前导维度,而不改变调用方式。mjw.put_model() 还兼作兼容性检查:如果模型使用不支持的功能,它会抛出异常,而不是静默丢弃它们。
显式初始化这三个字段是透明的选项,它清楚地表明了哪些数据被传输到了设备;而 mjw.put_data(mjm, mjd, nworld=…) 则是在一次调用中将整个初始化的结构体传输过去。
然后帧循环成为 Gate 1 循环,其内部步骤重定向到 GPU 并镜像回主机:
def simulate_frame () -> None :
ctrl = controller.step(mjm, mjd, frame_dt)
for _ in range (sim_substeps):
mjd.ctrl[: mjm.nu] = ctrl
wp.copy(d.ctrl, wp.array(mjd.ctrl[ None , :], dtype=wp.float32, device=device))
mjw.step(m, d)
mjd.qpos[:] = d.qpos.numpy()[ 0 ]
mjd.qvel[:] = d.qvel.numpy()[ 0 ]
mujoco.mj_forward(mjm, mjd)
.numpy() 的读取会在每个子步骤中同步并将数据复制到主机,因此这是一条任务验证路径,而非吞吐量基准测试。它将逆运动学、视角控制和任务检查保留在主机上。在复制 qpos 和 qvel 后,调用 mujoco.mj_forward(mjm, mjd) 以刷新派生的主机量(例如 mjd.xpos),然后再将其用于控制、视角显示或堆栈检查。在循环之后读取这些字段不会自动刷新它们。Gate 4 移除了
Classic MuJoCo provides fast CPU-based
robot simulation for developing, testing, and controlling robots and it can parallelize sampling across CPU cores. But as learning workloads grow, the question shifts from how quickly one world can run to how many worlds can run at once. GPU acceleration makes it possible to advance those worlds in large batches while keeping simulation and learning data close to the device.
MuJoCo Warp (MJWarp) , built on NVIDIA Warp , takes compatible MuJoCo models into that GPU-scale regime. In this article, we will move an SO-101 follower arm from a familiar MuJoCo workflow to as many as 2,048 parallel MJWarp environments and examine the technology and validation steps that make the transition possible.
Figure 1. How MJWarp connects Python to GPU simulation. MuJoCo loads and compiles the MJCF model; MJWarp implements the physics in NVIDIA Warp, which compiles CUDA kernels to advance simulation states on NVIDIA GPUs.
This is the second article in our State of Simulation for Physical AI series. The first article mapped the robot-simulation landscape. Here, we prepare and scale the simulation environment; we do not train a policy. The later Newton and Isaac Lab installments cover the next integration layers.
Putting it together
Layer
Role in the stack
NVIDIA Warp
Python kernel language: single instruction, multiple threads (SIMT), autodiff, PyTorch/JAX interop
MJWarp
MuJoCo physics on Warp: same MJCF, batched GPU throughput
Your scene (SO-101)
Familiar Menagerie / Robot Studio assets + task geometry
Next (Newton / Isaac Lab)
Multi-solver API, USD, sensors, managers, training loops
Decision shortcut:
If you need…
Reach for…
Single-robot MPC / teleop
MuJoCo CPU
Max throughput on raw MuJoCo physics
MJWarp (or mjlab )
JAX training recipes
MuJoCo Playground / MJX (impl='warp')
Multi-solver + Isaac Lab integration
Newton — next post in this series
Start with one useful Warp Kernel
NVIDIA Warp is a Python framework for writing high-performance, GPU-accelerated kernels. Warp lets developers author statically typed kernels in Python and compiles them for CPU or CUDA execution. The first launch builds and caches a native module; later launches reuse it. The kernel language is a performance-oriented subset of Python, while ordinary Python remains responsible for configuration, allocation, and launch orchestration.
This small robotics-oriented kernel advances point positions under gravity. One logical thread handles one point, so the same code scales from two points to millions without introducing GPU terminology into the control flow.
The three value propositions of Warp are:
Pillar
What you get
Performance
Native-CUDA speed via JIT compilation, kernel fusion, and CUDA Graphs
Ease of use
Pure Python authoring with built-in vectors, matrices, quaternions, BVHs, hash grids, sparse matrices, and tile primitives
Capability
Differentiable kernels and DLPack-style interop so simulation can sit inside an ML training loop
import numpy as np
import warp as wp
@wp.kernel
def integrate (
positions: wp.array[wp.vec3],
velocities: wp.array[wp.vec3],
dt: float ,
):
i = wp.tid()
velocities[i] += wp.vec3( 0.0 , 0.0 , - 9.81 ) * dt
positions[i] += velocities[i] * dt
wp.init()
device = "cuda:0" if wp.is_cuda_available() else "cpu"
start = np.array([[ 0.0 , 0.0 , 0.5 ], [ 0.2 , 0.0 , 0.5 ]], dtype=np.float32)
positions = wp.array(start, dtype=wp.vec3, device=device)
velocities = wp.zeros_like(positions)
wp.launch(
integrate,
dim= len (start),
inputs=[positions, velocities, 0.01 ],
device=device,
)
wp.synchronize_device(device)
print (positions.numpy())
Three properties make this useful in robotics:
Explicit parallel work. wp.tid() identifies the point, contact, body, or world owned by the current logical thread.
Explicit device arrays. An array lives on the selected device. Calling .numpy() on a CUDA array synchronizes and copies it to CPU memory; it is not a zero-copy path. For a device-resident PyTorch or JAX pipeline, use Warp’s framework adapters or DLPack-compatible sharing instead.
Composable kernel launches. A program can launch a sequence of focused kernels and capture supported CUDA work into a graph to reduce repeated dispatch overhead. Graph capture replays launches against existing buffers; it does not fuse arbitrary kernels.
Differentiability and Determinism.
Two further Warp capabilities are worth knowing, even though neither is used in the SO-101 workflow in this article. Warp kernels are differentiable : a wp.Tape records the forward kernel launches made inside its context and replays their adjoints in reverse when backward() is called, which is why teams build differentiable geometry, CFD, and custom physics in Warp, including CAE workflows for simulation and design optimization. Warp also supports deterministic execution , introduced in Warp 1.15 : GPU atomics are scheduler-dependent by default, so repeated launches of the same kernel can differ slightly, and the opt-in deterministic modes trade some performance for reproducible ordering in simulation, validation, and regression tests. These are Warp capabilities, not guarantees of differentiability or determinism for an entire MJWarp rollout. See the Warp documentation on differentiability and deterministic execution for the details.
Try Warp: pip install warp-lang (≥ 1.15 for GPU determinism), then python -m warp.examples.browse, or the tutorial notebooks .
What is MuJoCo Warp (MJWarp)?
A robot simulator repeatedly computes what happens next: given the current joint positions, velocities, controls, and contacts, it advances the scene by one small timestep. In this article, a world means one independent copy of that scene and its state. One world might contain the SO-101 arm reaching for a cube; another can contain the same arm starting from a slightly different pose.
MuJoCo and MJWarp can run the same compatible robot and task, but they organize the work differently. MuJoCo naturally suits developing and inspecting one or a few CPU worlds. MJWarp is a NVIDIA Warp implementation of MuJoCo ’s physics pipeline that places the model and a batch of independent states on NVIDIA GPUs; one call to mjw.step advances the entire batch.
MJWarp’s value is not necessarily a faster step for one world. It is the ability to advance hundreds or thousands together, giving the GPU enough parallel work to improve aggregate throughput , the total world-steps completed per second. That favors reinforcement learning and large-scale sampling, where collecting experience matters more than minimizing one environment’s latency.
This blog covers the following:
validate one MuJoCo world,
move it to MJWarp, form a batch,
verify it, and measure it correctly.
Solver tuning, Jacobian representation, and specialized multi-GPU or determinism topics are not required for this migration and can be covered separately.
Then, the distinction is precise:
Latency is wall-clock time for one simulation step.
Aggregate throughput is the total number of world-steps completed per measured wall-clock second.
Basic usage: structs, batch sizes, and a minimal step
The core API transition is small:
MuJoCo host workflow
MJWarp workflow
mujoco.MjModel
mjw.put_model(mjm) creates a device model
mujoco.MjData
mjw.put_data(mjm, mjd, ...) preserves and batches an existing state
mujoco.mj_step(mjm, mjd)
mjw.step(m, d) advances every world in d
Host arrays such as mjd.ctrl
Batched device arrays such as d.ctrl with shape (nworld, nu)
Use mjw.make_data() when default/fresh state is intended. Use mjw.put_data() when the exact initialized MuJoCo state must cross the migration boundary.
Allocating batched resources requires defining the following parameters (refer to Batch sizes ):
Parameter
Meaning
nworld
Total number of parallel environments
nconmax
Expected contacts per individual world (overall capacity ≈ nconmax * nworld)
naconmax
Alternative setting: global maximum contacts across all environments combined (takes precedence if both are defined)
njmax
Hard upper limit on constraints per world
Performance tuning
with wp.ScopedCapture() as capture:
mjw.step(m, d)
wp.capture_launch(capture.graph)
Additional tuning considerations. After sizing contact and constraint buffers, test solver iteration limits without changing task behavior. Meshes and CCD settings can increase memory use; nccdmax / naccdmax can reduce CCD buffer allocation when the measured contact counts allow it. MJWarp’s compact solver uses MuJoCo’s Newton constraint solver and sleeping, not the separate Newton physics-engine framework. Compact-solver and multi-GPU configuration are beyond this walkthrough; consult the MJWarp performance-tuning documentation.
To train policies on MJWarp physics:
Install / try: pip install mujoco-warp · mjwarp-viewer path/to/scene.xml · Colab tutorial
Workflow to migrate a MuJoCo scene to MjWarp
Establish a MuJoCo CPU baseline
The scene. Nothing here is MJWarp-specific yet: an SO-101 arm, a table, and two cubes to stack, written as ordinary MJCF.
Figure 2. SO-101 pick-and-place scene, rendered from the MuJoCo CPU simulation. The task is to grasp the red 44 mm cube and stack it on the blue cube; the same robot and scene are used for MJWarp validation.
< mujoco model = "so101_pick_place" >
< include file = "so101.xml" />
< worldbody >
< light pos = "0.3 0 1.5" dir = "0 0 -1" directional = "true" />
< geom name = "floor" type = "plane" size = "0 0 0.05" />
< geom name = "table" type = "box" pos = "0.35 -0.04 0.012"
size = "0.16 0.26 0.012" rgba = "0.32 0.32 0.32 1"
friction = "1 0.005 0.0005" condim = "3" />
< body name = "red_cube" pos = "0.33 -0.13 0.046" >
< freejoint name = "red_cube_joint" />
< geom type = "box" size = "0.022 0.022 0.022" mass = "0.08"
rgba = "0.85 0.05 0.04 1" friction = "1.2 0.005 0.0005" condim = "3" />
</ body >
< body name = "blue_cube" pos = "0.33 0.06 0.046" >
< freejoint name = "blue_cube_joint" />
< geom type = "box" size = "0.022 0.022 0.022" mass = "0.08"
rgba = "0.05 0.20 0.90 1" friction = "1.2 0.005 0.0005" condim = "3" />
</ body >
</ worldbody >
</ mujoco >
For an MJCF box, the size values are half-extents: size=”0.022 …” defines a cube with 44 mm edges. The task uses this size for its success thresholds. The arm base is at the origin, its reach is along +X, and the cubes are arranged along Y.
In the companion repository this file is generated rather than hand-written: resolve_pick_place_scene() copies the Menagerie arm into .generated/, fills the table and cube coordinates from a robot profile, and writes scene_pick_place.xml. The walkthrough uses the SO-101 profile; the optional reBot variant is described below.
Loading it. Compilation and stepping are ordinary MuJoCo:
import mujoco
mjm = mujoco.MjModel.from_xml_path( "scene_pick_place.xml" )
mjd = mujoco.MjData(mjm)
fps = 50
sim_substeps = 10
frame_dt = 1.0 / fps
mjm.opt.timestep = frame_dt / sim_substeps
controller = PickPlaceController(spec=spec)
for _ in range ( 600 ):
ctrl = controller.step(mjm, mjd, frame_dt)
for _ in range (sim_substeps):
mjd.ctrl[: mjm.nu] = ctrl
mujoco.mj_step(mjm, mjd)
Keep that shape in mind: compute controls once per frame, step physics sim_substeps times. Gate 2 changes only the inner loop, which is what makes the migration easy to review.
Match the simulation and control rates. At 50 control frames per second and 10 physics substeps per frame, use a physics timestep of 0.002 seconds. Set it before the CPU rollout and before uploading the model with mjw.put_model so both backends advance the same simulated time:
mjm.opt.timestep = frame_dt / sim_substeps
Without that line, every later measurement inherits the mismatch: parity comparisons, throughput numbers quoted as “simulated seconds,” and any learned policy whose action rate no longer matches deployment.
Check whether the cubes are stacked successfully. With 44 mm cubes, success becomes two measurable conditions: a horizontal center error of xy_err ≤ 0.015 m (measured between the cube centers) and a vertical separation of 0.035 m ≤ dz ≤ 0.055 m between cube centers (one cube edge, with slack for settling). Evaluate both conditions after the cubes have settled; a successful process exit alone does not establish task success.
Run the CPU task from the companion checkout. Publication blocker: confirm the accessible repository URL and pinned dependency and asset versions before publishing these instructions; the repository placeholder below is not an executable URL.
git clone https://github.com/NVIDIA/accelerated-computing-hub.git blogs
cd blogs/tutorials/sim2real-blogs/notebooks/mujoco
uv venv --python 3.12 && source .venv/bin/activate
uv pip install -r requirements.txt
cd /tutorials/sim2real-blogs/notebooks/mujoco
python solutions/so101_pick_place_solution.py --headless-steps 600 --debug
The run ends by printing the two numbers above (stack check: xy_err=… dz=…), which is the assertion the rest of the article compares against. so101_pick_place.py next to it is the same program with the physics steps left as exercises.
The arm comes straight from MuJoCo Menagerie pinned to a known-good commit, since Menagerie assets change, so treat the scene as a template. Optional reBot variant. The companion code also exposes --robot rebot with a separate profile for the scene layout, gripper, and capacity limits (nconmax=256, njmax=500). This walkthrough uses SO-101. Validate the reBot asset and task separately before reporting its results.
Validate one-world MJWarp parity
Run one world on the GPU first, with the host still in the loop, so you can watch the same task in the same viewer and compare the same two numbers. Upload the model, allocate batched state, seed it from the initialized host state, and run one forward pass before stepping:
wp.init()
import mujoco_warp as mjw
device = wp.get_device()
m = mjw.put_model(mjm)
d = mjw.make_data(mjm, nworld= 1 , nconmax=spec.nconmax, njmax=spec.njmax)
wp.copy(d.qpos, wp.array(mjd.qpos[ None , :], dtype=wp.float32, device=device))
wp.copy(d.qvel, wp.array(mjd.qvel[ None , :], dtype=wp.float32, device=device))
wp.copy(d.ctrl, wp.array(mjd.ctrl[ None , :], dtype=wp.float32, device=device))
mjw.forward(m, d)
Every device array carries a leading world dimension, which is why the host state is indexed as mjd.qpos[None, :], shape (1, nq) instead of (nq,). Scaling to thousands of worlds later changes only that leading dimension, not the calls. mjw.put_model() also doubles as a compatibility check: it raises if the model uses unsupported features rather than silently dropping them.
Seeding the three fields explicitly is the transparent option, and it makes clear exactly what crosses to the device; mjw.put_data(mjm, mjd, nworld=…) carries the whole initialized struct over in one call instead.
The frame loop is then the Gate 1 loop with its inner step redirected to the GPU and mirrored back:
def simulate_frame () -> None :
ctrl = controller.step(mjm, mjd, frame_dt)
for _ in range (sim_substeps):
mjd.ctrl[: mjm.nu] = ctrl
wp.copy(d.ctrl, wp.array(mjd.ctrl[ None , :], dtype=wp.float32, device=device))
mjw.step(m, d)
mjd.qpos[:] = d.qpos.numpy()[ 0 ]
mjd.qvel[:] = d.qvel.numpy()[ 0 ]
mujoco.mj_forward(mjm, mjd)
The .numpy() reads synchronize and copy data to the host on every substep, so this is a task-validation path, not a throughput benchmark. It keeps inverse kinematics, viewing, and task checks on the host. After copying qpos and qvel, call mujoco.mj_forward(mjm, mjd) to refresh derived host quantities such as mjd.xpos before using them for control, viewing, or the stack check. Reading those fields after the loop does not refresh them automatically. Gate 4 removes
首次收录 · 2026-09-24 · 10.79 分