H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

arXiv Preprint 2026

Dingyi Rong1,2 Yue Shi2† Chaofan Ma1‡ Jiezhang Cao1 Zongrui Wang1,2
Zeyu Zhang1,2 Yao Mu1,2 Guangtao Zhai1,2† Ning Liu1†
1 Shanghai Jiao Tong University   2 Shanghai Artificial Intelligence Laboratory

† Corresponding authors.   ‡ Project lead.

Overview of H2R-Bench: source curation, native-interface generation, and transfer-aware evaluation

A robot-looking video is not a transferred demonstration. Generic video benchmarks score both outputs highly; H2R-Bench diagnoses transfer along four source-grounded axes — goal, action, functional contact, and embodiment — plus generic video quality.

Abstract

Robot learning needs large-scale manipulation data, yet robot demonstrations are expensive to collect, while egocentric human videos are abundant. Video world models offer a possible bridge, but their cross-embodiment transfer ability is largely unmeasured. We introduce H2R-Bench, a benchmark for cross-embodiment human-to-robot manipulation video generation: given an egocentric human demonstration and a target embodiment, a model must produce the corresponding robot manipulation video. Each of the 240 cases (120 sources × 2 embodiments) carries source-grounded annotations, and outputs are scored on five dimensions — goal state, action events, functional contact, embodiment correctness, and video quality. Benchmarking 11 state-of-the-art models over 6 manipulation families shows that even leading systems fail at embodiment consistency, functional interaction, and task execution.

Human → Robot, Not Text → Video Generation is conditioned on a real egocentric demonstration and a requested robot embodiment, then checked for whether the source manipulation survived the transfer.
Source-Grounded Annotations Task goals, required action events, functional contact regions, manipulation modes, and expected object responses — all derived from each case's own source clip.
Quality ≠ Transfer Validity Across 22 model–embodiment pairs, video quality and H2RCore are almost rank-independent (Spearman ρ = 0.14): polished clips routinely fail the transfer.

Motivation

Why Existing Video Benchmarks Miss the Transfer

A generated robot video is an intermediate representation of a target robot execution. What matters is not visual plausibility but whether the manipulation evidence in the source demonstration is preserved while the execution is adapted to a different embodiment. A coherent clip can still retain human hands, realize the wrong end-effector, or show object changes unsupported by any visible robot interaction. Existing benchmarks are open-domain, text-conditioned, or robot-centric — none takes an egocentric human-hand video as source evidence and checks whether it was retargeted to a different robot embodiment.

Benchmark Settings Evaluation
I2V RV H2R VQ Goal Action Cont. Emb.
VBench ××× ××××
WorldModelBench × ××
RBench × ×
RoboWM-Bench × ××
RoboTrustBench × ×
H2R-Bench (ours)

Comparison with existing video-generation benchmarks. “I2V”, “RV”, “H2R” = image-to-video generation, robot-video evaluation, human-to-robot transfer; / / × = full / partial / no support.

Benchmark

What a H2R-Bench Case Looks Like

Each case is a triple (Vh, e, p): a 5-second egocentric human source video, a target embodiment, and a prompt naming the task goal and embodiment. The prompt carries no evaluation weights or per-metric checks — those live in a separate scoring annotation. Pose or trajectory imitation is not required: a gripper and a dexterous hand may solve the task differently, as long as the strategy fits the requested morphology and the source task.

Step 1

Curate

120 EgoDex test-split clips with visible entities, interactions, and state changes, evenly spread over six manipulation families (20 each), paired with two target embodiments → 240 transfer cases.

Step 2

Annotate

Per clip: initial/final states, objects and tools, required action events, source-side contact evidence. Per embodiment: functional contact regions, manipulation modes, expected object responses, bimanual roles. All manually verified.

Step 3

Generate

Native-interface protocol: video-conditioned models get the full source clip, frame-conditioned models get ordered frames up to their interface limit. No robot reference image in the main setting — embodiment is text-only.

Six families: F1 rigid transport/rearrangement · F2 articulated actuation · F3 insertion/attachment · F4 deformable configuration · F5 bulk-material transfer/mixing · F6 surface or material change. Two embodiments: parallel-jaw gripper, dexterous hand.

Evaluation

Five Dimensions, One Transfer Score

M1–M4 measure whether the source manipulation was transferred correctly; M5 measures task-agnostic video quality. For M1–M4, three MLLM judges (Gemini 3.5 Flash, Qwen3.7-Plus, GPT-5.4) independently score the prescribed visual evidence on a shared 0–4 rubric; scores are normalized to [0, 1] and averaged. Every metric sees 25 uniformly sampled frames, so all models are judged on the same evidence budget.

weight 0.15 M1Goal-State Completion

Weighted predicates over the final state required for success — spatial or containment relation, mechanism state, attachment, deformation, material distribution, surface change. The sequence gives context; the final frames decide.

weight 0.15 M2Action-Event Completion

How clearly the required events happen — grasping, inserting, releasing, pouring, wiping, folding. Whether the operations occur, not whether the robot reproduces the human pose or timing.

weight 0.30 M3Functional Contact Transfer

25 source frames vs. 25 generated frames plus a source-derived contact spec: contacted functional region, whether contact is visibly established, manipulation mode, whether the object responds, feasibility for the requested embodiment. A different grasp is fine if the contact serves the same function.

weight 0.30 M4Embodiment Correctness

Separates robot presence from morphology: robot visible, human hands gone, requested embodiment class, correct end-effector type, temporally consistent structure. Zero if no robot is present or a human performs the main manipulation.

weight 0.10 M5Video Quality

No task annotation used. Averages four normalized components: imaging quality (MUSIQ), aesthetic quality (LAION predictor over CLIP ViT-L/14), temporal stability, and motion smoothness (AMT-S interpolation error).

H2RCore = 100 × (0.15 Sgoal + 0.15 Saction + 0.30 Scontact + 0.30 Semb + 0.10 Svideo)

Contact and embodiment jointly carry 60% of the score: valid transfer needs both a functionally supported interaction and the requested morphology. Goal and action preserve the source task but neither is sufficient, and video quality contributes only 0.10 — visual polish cannot compensate for contact or embodiment failures.

Leaderboard

How Far Are Video World Models From Human-to-Robot Transfer

All 11 models are evaluated on the 240 transfer cases through their native source-conditioning interfaces. The three video-conditioned systems take the top positions, and almost all of their separation comes from contact transfer and embodiment correctness. Below them the frame-conditioned group collapses: several models keep high goal, action, and quality scores while embodiment falls to near zero.

Model Parallel-Jaw Gripper Dexterous Hand
GoalActionContactEmbod.QualityH2RCore GoalActionContactEmbod.QualityH2RCore
Video-conditioned generation
Seedance 2.0 0.7250.8130.7760.7680.79377.3 0.7440.8320.8550.9110.79984.6
Wan2.7 0.7060.7910.7660.7720.79676.5 0.7180.8040.8350.9100.79583.1
Kling-V3 0.7100.8070.7510.7070.79874.5 0.7070.8000.8190.8850.80281.7
Frame-conditioned generation
Mitty-EPIC14B 0.5810.6680.5980.5850.73261.5 0.5870.6840.5980.3920.73256.1
Veo 3.1 0.7250.7970.5330.1000.78349.6 0.7150.8160.6420.2270.79357.0
Grok Imagine Video 0.6610.7290.4430.2680.79250.1 0.6780.7290.4690.1980.80449.2
LTX-2.3 0.4730.5450.2920.0120.77332.1 0.5200.5920.3770.1320.78039.8
SkyReels-V3-R2V 0.4480.6100.2560.0040.78731.5 0.4410.6070.3410.0260.78934.6
Wan2.2 0.4920.6390.2580.0000.76932.4 0.5120.6530.2860.0000.76633.7
LongCat 0.4720.5830.2430.0000.79031.0 0.4280.5450.2980.0200.79332.0
HunyuanVideo 1.5-I2V 0.5350.5490.1840.0050.80630.0 0.4990.5550.1850.0410.80830.7

Main-evaluation results by target embodiment. M1–M5 = goal completion, action completion, contact transfer, embodiment correctness, video quality; H2RCore aggregates all five on 0–100. Bold / underlined mark the best and second-best per column.

Recognizing the task is not transferring it. Veo 3.1 has the best gripper goal score (0.725) and a 0.816 hand action score, yet embodiment is only 0.100 / 0.227 — the task stays recognizable even when the requested robot never appears.
The prettiest video is near the bottom. HunyuanVideo 1.5-I2V leads video quality for both targets (0.806 / 0.808) but contact stays near 0.185 and embodiment near zero — last on H2RCore.
The top is tight but consistent. Seedance's paired lead over Wan2.7 is +0.79 (95% CI [+0.08, +1.48]) for the gripper and +1.52 ([+0.92, +2.16]) for the hand.
Sparse frames cannot replace the source video. Replacing Seedance's full clip with nine ordered frames drops H2RCore by 27.6 (gripper) and 41.4 (hand) while video quality moves less than 0.02.

Analysis

Generic Video Quality Does Not Recover the Ranking

VBench-based video quality versus H2RCore across 22 model-embodiment pairs

Video quality versus H2RCore. Across 22 model–embodiment pairs, video quality stays in a narrow 0.73–0.81 band while H2RCore spans 30.0 to 84.6, and their rank association is weak (Spearman ρ = 0.14). HunyuanVideo 1.5-I2V leads on quality but sits near the bottom on H2RCore; Mitty-EPIC14B is the reverse.

Failure-type rates derived from structured M2-M4 diagnostics

Where the transfers break. Failure-type rates from the structured M2–M4 diagnostics; a failure is recorded when the judge-averaged component falls below 0.5, and categories are non-exclusive. Human-led manipulation and contact-region mismatch dominate frame-conditioned outputs, while end-effector mismatch stays common for the gripper even among stronger models.

Per-attribute breakdown across task families

Attribute-level breakdown across task families. Scores decompose per manipulation attribute, showing which requirements of a family the models satisfy and which they systematically miss.

Embodiment & Conditioning

What Changes the Transfer, and What Does Not

Metric Mean change (Hand − Gripper) Hand higher
Goal-State Completion (M1)+0.0026/11
Action-Event Completion (M2)+0.0088/11
Functional Contact Transfer (M3)+0.05511/11
Embodiment Correctness (M4)+0.0478/11
Video Quality (M5)+0.0048/11
H2RCore (0–100)+3.39/11

Effect of target embodiment across all 11 models, on the same 120 sources. Mean change is the average score difference between the dexterous hand and the parallel-jaw gripper; “Hand higher” counts models with a positive difference.

The dexterous hand is the easier target. Nine of 11 models score higher with the hand (average +3.3 H2RCore). Goal, action, and video quality barely move; contact rises for all 11 models and embodiment for eight. The effect is largest for the video-conditioned group (+7.3 Seedance, +6.6 Wan2.7, +7.2 Kling-V3) — its morphology is closer to the human actor in the source. It is not universal: Grok Imagine Video and Mitty-EPIC14B both score lower with the hand.
A robot reference image helps only some models. Wan2.7 gains the most — gripper contact 0.766→0.871, embodiment 0.772→0.875, H2RCore 76.5→83.1. Kling-V3 and Seedance move the other way, losing 13.5 / 8.8 gripper points and 8.3 / 4.4 hand points. Video quality shifts by only a few hundredths in every setting: the reference image trades off robot appearance against source interaction, not presentation quality.
Difficulty is about contact, not deformation. Deformable-object configuration (F4) is the strongest family for both targets. Difficulty appears governed less by visible deformation than by whether success needs precise, localized functional contact and a partially hidden state transition. Seedance and Wan2.7 stay comparatively stable across families; Grok Imagine Video and Veo 3.1 vary far more.
Automatic judgments track human preference. Three human raters scored M1–M4 for all 11 models under both embodiments on five sampled scenes per family — 660 generated videos. The macro-average Human–MLLM Spearman correlation is ρ = 0.883, so the automated evaluation largely preserves human model preferences within a shared task.
Task-family H2RCore profiles across the six manipulation families

Task-family H2RCore profiles. Each panel compares parallel-jaw gripper and dexterous hand transfer for the same 11 models across the six manipulation families, on a shared 0–100 scale, exposing both task-dependent difficulty and embodiment-dependent model preferences.

Spearman agreement between human and MLLM evaluators

Agreement between human and MLLM evaluators. Human raters and MLLM judges rank generated videos by the transfer score aggregated from M1–M4. MLLM-based evaluation aligns closely with human judgment, with Spearman correlations above 0.8 across evaluators.

Qualitative Results

One Demonstration, Three Ways to Fail

Qualitative H2R transfer results on a shared scoop-and-dump source video

Qualitative H2R transfer results for a shared source video and prompt. The unscored top row is the human source; the generated rows show representative stages of each output, and the metric strips report per-video M1–M5 and H2RCore. Seedance 2.0 maintains visible robot contact with the scoop and transfers the ice into the cup. HunyuanVideo 1.5-I2V leaves the human as the active manipulator, preserving the action without transferring it to the robot. Veo 3.1 produces a plausible interaction with the wrong end-effector. The four metrics capture complementary evidence — M1 the outcome, M2 the required events, M3 visible robot–scoop contact, M4 the requested embodiment — and a correct final state alone does not establish successful transfer.

Human Evaluation

How the Human Scores Were Collected

Interface used to collect human M1-M4 scores

Each task group contains five shared source scenes from one model and task family under both target embodiments, yielding ten videos. Annotators inspect the source and generated clips side by side, verify the requested embodiment, and score the source-derived criteria on a common 0–4 scale. Scores are stored separately per annotator, and automatic MLLM judgments are never displayed.

Scope

What H2R-Bench Does Not Claim

H2R-Bench evaluates visible evidence of human-to-robot transfer, not physical executability or downstream policy performance. Its 120 EgoDex sources and two target embodiments cover only part of the variation found in real manipulation settings. The native-interface comparison combines model capability with differences in source-conditioning interfaces, so group-level patterns are descriptive rather than causal. Finally, sampled visual evidence and MLLM judgments can remain uncertain under occlusion, subtle contact, or severe generation artifacts, despite multi-judge aggregation and human validation.

BibTeX

@article{rong2026h2rbench,
  title={H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models},
  author={Rong, Dingyi and Shi, Yue and Ma, Chaofan and Cao, Jiezhang and Wang, Zongrui and Zhang, Zeyu and Mu, Yao and Zhai, Guangtao and Liu, Ning},
  journal={arXiv preprint arXiv:2608.13049},
  year={2026}
}