parallelquant
September 14, 2026 · MarkTechPost

Reward AI trains robot manipulation policy from human demos only

Reward AI released OM-1, a general-purpose manipulation policy trained solely on human demonstrations captured via a 7-degree-of-freedom wearable glove, with no teleoperation or on-robot data used. The policy runs on industrial arms and humanoids at human speed and can learn a new task from under 30 minutes of data. No weights, code, or API have been released publicly yet.

Why it matters: Most robot learning today still leans on expensive teleoperation rigs or robot-specific data collection; if human-demo-only training scales, it could sharply cut the cost of teaching robots new tasks. It fits a broader push toward general-purpose robot policies that generalize across embodiments rather than being trained per-robot, alongside other recent robotics gains.

Related updates