RoboTwin 2.0 Tasks: Experimental Setup¶
Tasks Used in Experiments: beat_block_hammer and move_can_pot
Based on: SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning (2025)
Table of Contents¶
- Overview
- Task 1: beat_block_hammer
- Task 2: move_can_pot
- Environment Randomization
- Language Instruction System
- Training Configuration
Overview¶
Both tasks are part of the RoboTwin 2.0 benchmark, designed for dual-arm manipulation with realistic physics simulation using SAPIEN. These tasks test different manipulation capabilities:
- beat_block_hammer: Tool use and contact-based manipulation
- move_can_pot: Pick-and-place with spatial reasoning
Common Characteristics¶
| Property | Value |
|---|---|
| Simulator | SAPIEN (CPU-based physics) |
| Robot Platform | ALOHA (dual-arm, 7 DOF per arm) |
| Action Space | 14-dimensional continuous (7 DOF × 2 arms) |
| Observation | RGB images (3 cameras) + Proprioception (14-dim joint states) |
| Camera Views | Head camera + Left wrist + Right wrist (224×224×3 each) |
| Max Environment Steps | 200 steps per episode |
| Action Chunking | 25 action chunks per VLA inference |
| VLA Calls per Episode | ~8 (200 ÷ 25) |
| Traj Mini-Batch Size | 8 |
Task 1: beat_block_hammer¶
Task Description¶
Objective: Use the robot arm to grasp a hammer and strike a block on the table.
Natural Language Description: "There is a hammer and a block on the table, use the arm to grab the hammer and beat the block"
Scene Setup¶
The environment contains two objects on a table:
Hammer¶
- Model: Standard hammer (model ID: 020_hammer)
- Initial Position: Fixed location at table center
- X-coordinate: 0 meters (table center)
- Y-coordinate: -0.06 meters (slightly back from center)
- Z-coordinate: 0.783 meters (on table surface)
- Orientation: Fixed (slightly angled)
- Physical Properties: Very light mass (0.001 kg) for easy manipulation
- Functional Points:
- Point 0: Hammer head (for striking)
- Base point: Handle (for grasping)
Block¶
- Model: Static red cube (doesn't move when hit)
- Size: 2.5 cm × 2.5 cm × 2.5 cm
- Initial Position: Randomized (see randomization section)
- Color: Red (for visibility)
- Physical Properties: Static (fixed to table, doesn't react to physics)
- Functional Points:
- Point 0: Block center
- Point 1: Block top surface (target for hammer)
Success Criteria¶
The task is considered successful when ALL of the following conditions are met:
- Position Accuracy: The hammer head is within 2 cm of the block position (in X-Y plane)
- Measured as Euclidean distance in horizontal plane
-
Tolerance: ±2 cm in both X and Y directions
-
Contact Detection: The hammer and block are in physical contact
- Verified through SAPIEN's collision detection system
- Must be simultaneous with position accuracy
Mathematical Formulation:
Success = (|hammer_x - block_x| < 0.02 m) AND
(|hammer_y - block_y| < 0.02 m) AND
(hammer_contacts_block = True)
Expert Demonstration Strategy¶
Human demonstrations follow a four-step procedure:
Step 1: Arm Selection - Automatically select arm based on block position - If block X-coordinate < 0: use left arm - If block X-coordinate ≥ 0: use right arm - This creates symmetric training data for both arms
Step 2: Grasp Hammer - Approach hammer from 12 cm away (pre-grasp phase) - Move gripper to handle position - Close gripper at 1 cm distance (secure grasp) - Gripper fully closes around handle
Step 3: Lift Hammer - Raise hammer 7 cm upward (Z-axis displacement) - Creates clearance from table surface - Prepares for striking motion
Step 4: Strike Block - Move hammer toward block's top surface - Approach from 6 cm above (pre-placement phase) - Lower hammer until contact with block - Maintain closed gripper throughout strike - Stop when contact is detected
Challenges & Learning Opportunities¶
Key Challenges: 1. Spatial Reasoning: Block position varies significantly (50×20 cm range) 2. Tool Use: Must use hammer as extension of arm, not direct manipulation 3. Contact Control: Requires precise positioning for contact detection 4. Arm Coordination: Single-arm task but must avoid self-collision
RL Discovery Potential: - SFT model learns: grasp → lift → move → strike (following demonstrations) - RL may discover: more efficient striking trajectories, improved contact timing - Demonstrates RL's ability to refine precision and timing in tool-use scenarios
Task 2: move_can_pot¶
Task Description¶
Objective: Pick up a can from the table and place it beside a pot.
Natural Language Description: "There is a can and a pot on the table, use one arm to pick up the can and move it to beside the pot"
Scene Setup¶
The environment contains two objects on a table:
Pot (Kitchen Pot)¶
- Model: Kitchen pot with handles (model family: 060_kitchenpot)
- Model Variants: 7 different pot models (IDs 0-6)
- Different sizes, shapes, colors
- Randomly selected each episode
- Initial Position: Near table center
- X-coordinate: 0 meters (center)
- Y-coordinate: 0 meters (center)
- Z-coordinate: On table surface
- Orientation: Random rotation ±22.5° around vertical axis
- Physical Properties: Normal mass, can be moved if pushed hard
- Functional Points: Center point for distance measurement
Can (Sauce Can)¶
- Model: Cylindrical sauce can (model family: 105_sauce-can)
- Model Variants: 5 different can models (IDs: 0, 2, 4, 5, 6)
- Different labels, sizes
- Randomly selected from subset
- Initial Position: Randomized (see randomization section)
- Orientation: Random rotation around vertical axis
- Physical Properties: Standard mass, cylindrical shape
- Functional Points: Center point and top surface
Success Criteria¶
The task is considered successful when ALL of the following conditions are met:
- Horizontal Distance: Can is 18 cm (±20 cm tolerance) from pot horizontally
- Measured on the correct side (left arm → left of pot, right arm → right of pot)
-
Must be on the correct side: cannot be on opposite side of pot
-
Lateral Alignment: Can Y-position is within 3.5 cm of pot Y-position
-
Ensures can is "beside" pot, not in front or behind
-
Can Orientation (Upright):
- X-axis rotation: 90° ± 15° (upright, not tilted forward/backward)
- Y-axis rotation: 0° ± 15° (upright, not tilted left/right)
-
Can must maintain stable upright position
-
On Table Surface: Can Z-position ≤ original pot height + 0.1 cm
-
Ensures can is resting on table, not floating or dropped
-
Released: Both robot grippers are open
- Confirms can has been released successfully
- Not being held by robot
Mathematical Formulation:
Success = (0 < |can_x - pot_x| < 0.2 m) AND # Distance tolerance
(|can_y - pot_y| < 0.035 m) AND # Lateral alignment
(|can_rotation_x - 90°| < 15°) AND # Upright X
(|can_rotation_y - 0°| < 15°) AND # Upright Y
(can_z ≤ table_surface + 0.001 m) AND # On surface
(can_on_correct_side = True) AND # Correct side of pot
(left_gripper_open = True) AND # Released
(right_gripper_open = True) # Released
Expert Demonstration Strategy¶
Human demonstrations follow a five-step procedure:
Step 1: Arm Selection - Automatically select arm based on can position - If can X-coordinate > 0: use right arm - If can X-coordinate ≤ 0: use left arm - This determines which side of pot to place can
Step 2: Grasp Can - Approach can from 5 cm away (pre-grasp phase) - Move gripper to can center - Close gripper around can body - Secure cylindrical grasp
Step 3: Lift and Retract - Move 10 cm backward (away from pot, Y-axis: -0.1 m) - Simultaneously lift 10 cm upward (Z-axis: +0.1 m) - Creates clearance from pot and other obstacles - Prevents collision during transport
Step 4: Transport to Target - Calculate target position beside pot: - If left arm: target_x = pot_x - 0.18 m (18 cm to left) - If right arm: target_x = pot_x + 0.18 m (18 cm to right) - target_y = pot_y (same lateral position) - target_z = table_surface - Move gripper with can to target position - Maintain upright orientation throughout
Step 5: Place Can - Approach target from 5 cm above (pre-placement phase) - Lower can smoothly to table surface - Ensure stable contact with table - Open gripper to release can - Retract gripper away from can
Challenges & Learning Opportunities¶
Key Challenges: 1. Spatial Reasoning: Understanding "beside" relationship (specific distance) 2. Dual Randomization: Both pot and can positions/models vary 3. Orientation Maintenance: Must keep can upright throughout manipulation 4. Precision Placement: Tight tolerance on final can orientation (±15°) 5. Side Selection: Must place on correct side based on arm used
RL Discovery Potential - "Pushcut" Phenomenon Observed: - SFT learns: grasp → lift high → move → lower carefully (grasp-move-place strategy) - RL discovers: PUSH the can directly to target position (pushcut strategy) - This is a novel behavior NOT present in any demonstration data - Pushcut strategy is faster and more robust than the demonstrated approach - Paper Section 6.1: "the RL-trained model instead learns to accomplish the task by simply pushing Object A into position" - Demonstrates emergence of completely new manipulation strategies through RL exploration
Environment Randomization¶
Purpose of Randomization¶
Randomization serves multiple purposes in robot learning:
- Prevent Overfitting: Ensures policy learns general strategies, not memorizing specific positions
- Improve Generalization: Policy must work across diverse configurations
- Test Robustness: Validates policy can handle variability
- Enable Transfer: Prepares policy for real-world deployment where exact positions cannot be controlled
beat_block_hammer Randomization¶
Block Position (Fully Randomized)¶
The block position is randomly sampled each episode:
X-Coordinate (Left-Right): - Range: -0.25 to +0.25 meters (50 cm span) - Distribution: Uniform random - Constraints: Must not be too close to center (|x| ≥ 0.05 m) - Purpose: Tests both left and right arm capabilities
Y-Coordinate (Front-Back): - Range: -0.05 to +0.15 meters (20 cm span) - Distribution: Uniform random - Constraints: None specific - Purpose: Tests depth perception and reaching
Z-Coordinate (Height): - Fixed: 0.76 meters (on table surface) - No variation: Block always at same height
Orientation (Rotation): - Rotation around Z-axis: 0 to 0.5 radians (0° to ~29°) - Distribution: Uniform random - Purpose: Tests rotational invariance - Note: Block is square, so rotation matters less
Additional Constraints: - Block must be at least 5 cm from table center (|x| ≥ 0.05) - Block must be at least 0.1 cm from origin (x² + y² ≥ 0.001) - These prevent trivial configurations
Prohibited Areas: - 10 cm padding around hammer position - 5 cm padding around block position - Robot must plan path avoiding these zones
Hammer Position (Fixed)¶
The hammer position is not randomized: - Always at same position: [0, -0.06, 0.783] meters - Fixed orientation - Rationale: Simplifies task to focus on block positioning challenge
Randomization Summary for beat_block_hammer:
Total Randomized Parameters: 3
- Block X: 50 cm range
- Block Y: 20 cm range
- Block Rotation: ~29° range
Configuration Space: ~50 × 20 × 29 = 29,000 possible discrete configurations
Continuous Space: Infinite (continuous sampling)
move_can_pot Randomization¶
Pot Randomization (Partial)¶
Model Variant: - Randomly select from 7 pot models (IDs 0-6) - Different visual appearances (colors, sizes, shapes) - Tests visual generalization
Position: - X-coordinate: 0 meters (fixed at center) - Y-coordinate: 0 meters (fixed at center) - Z-coordinate: On table surface (fixed)
Orientation: - Rotation around Z-axis: ±22.5° (±π/8 radians) - Distribution: Uniform random - Purpose: Tests rotation invariance for target reference
Can Randomization (Full)¶
Model Variant: - Randomly select from 5 can models (IDs: 0, 2, 4, 5, 6) - Different labels and appearances - Note: IDs 1 and 3 are excluded (possibly unstable or problematic)
Position: - X-coordinate: -0.3 to +0.3 meters (60 cm span) - Constraint: |x| ≥ 0.2 m (must not be too close to center) - Prevents can from being directly in front of robot - Y-coordinate: +0.05 to +0.15 meters (10 cm span, front of table) - Always on front side of table - Easier for robot to reach - Z-coordinate: On table surface (computed based on can model)
Orientation: - Rotation around Z-axis: ±45° (±π/4 radians) - Distribution: Uniform random - Purpose: Tests grasping from different angles
Spatial Constraints: - Can must be at least 30 cm from pot center - Formula: (can_x - pot_x)² + (can_y - pot_y)² ≥ 0.09 m² - Prevents can from starting too close to target - If constraint violated, resample position
Prohibited Areas: - 3 cm padding around pot position - 10 cm padding around can position - 15 cm × 20 cm zone on approach side of pot (arm-dependent) - If left arm: zone extends 15 cm to left of pot - If right arm: zone extends 15 cm to right of pot - Prevents trivial straight-line motions
Randomization Summary for move_can_pot:
Total Randomized Parameters: 7
- Pot Model: 7 variants
- Pot Rotation: ±22.5°
- Can Model: 5 variants
- Can X: 60 cm range (with constraints)
- Can Y: 10 cm range
- Can Rotation: ±45°
Configuration Space: 7 × 45° × 5 × 60 × 10 × 90° ≈ 8.5M possible discrete configurations
Continuous Space: Infinite (continuous sampling)
Randomization Process¶
Each episode follows this randomization procedure:
Episode Start: 1. Sample all random parameters from their respective distributions 2. Check spatial constraints (distance requirements, prohibited zones) 3. If constraints violated: resample violating parameters 4. Repeat until valid configuration found 5. Initialize SAPIEN scene with sampled configuration 6. Generate task instruction (random selection from template bank) 7. Compute initial observation (RGB images + proprioception) 8. Begin episode execution
Key Design Principles: - Constraint Validation: Ensures physically feasible configurations - Iterative Sampling: Rejection sampling until valid - Symmetric Design: Equal probability for left/right arm usage - Difficulty Calibration: Constraints prevent too-easy or impossible scenarios
Impact on RL Training¶
Dynamic Sampling Interaction: - Randomization creates natural difficulty distribution - Dynamic Sampling (paper Section 3.3) further filters: - Excludes episodes where all 8 samples succeed (too easy) - Excludes episodes where all 8 samples fail (too hard) - Keeps only episodes with mixed outcomes (learnable) - Combined effect: RL focuses on appropriate difficulty frontier
Curriculum Learning Effect: - Early training: High failure rate, mostly hard configurations kept - Mid training: More mixed outcomes, broader configuration range - Late training: High success rate, need new challenging configurations - This naturally implements curriculum without manual intervention
Language Instruction System¶
Instruction Generation Pipeline¶
Both tasks use a sophisticated language instruction system to test generalization:
Template-Based Generation: 1. Define full task description (detailed) 2. Specify instruction preferences (word count, style) 3. Define object schema (placeholders for specific objects) 4. Use LLM to generate diverse paraphrases 5. Manually curate into "seen" and "unseen" sets
Schema Placeholders:
- {A}: Primary object (hammer in task 1, pot in task 2)
- {B}: Secondary object (can in task 2)
- {a}: Arm designation (left or right)
beat_block_hammer Instructions¶
Full Description: "there is a hammer and a block on the table, use the arm to grab the hammer and beat the block"
Schema:
- {A} → hammer identifier (e.g., "020_hammer/base0")
- {a} → arm used (e.g., "left" or "right")
Instruction Preferences: - Word count: ≤10 words - Must mention: grasping action, striking action - Can mention: arm designation (optional)
Seen Instructions (50+ variants): Examples used during training: - "Pick {A} and strike the block." - "Lift {A} using {a} to hit the block." - "Grab {A} with {a} and beat the block." - "Use {A} to hammer the block." - "Take {A} and smash the block." - "Hold {A} and pound the block."
Unseen Instructions (10 variants): Examples reserved for evaluation: - "Grab {A} and beat the block." - "Use {a} to pick up {A}." - "Grab {A} and strike the block." - "Use {a} to grab {A} then beat block"
Instruction Diversity: - Verbs: grab, pick, lift, take, hold, use, employ - Actions: beat, strike, hit, hammer, pound, smash - Styles: imperative, sequential, compound sentences - Tests: synonym understanding, sentence structure variations
move_can_pot Instructions¶
Full Description: "there is a can and a pot on the table, use one arm to pick up the can and move it to beside the pot"
Schema:
- {A} → pot identifier (e.g., "060_kitchenpot/base3")
- {B} → can identifier (e.g., "105_sauce-can/base2")
- {a} → arm used (e.g., "left" or "right")
Instruction Preferences: - Word count: ≤10 words - Must mention: picking action, placement location - Must specify: relative position to pot
Seen Instructions (50+ variants): Examples used during training: - "Use {a} to grab {B} and move it next to {A}" - "Pick {B} up with {a} then place near {A}" - "Grab {B} with {a} and set it near {A}" - "Lift {B} and move it near {A}" - "Take {B} to {A} using {a}" - "Set {B} right next to {A}"
Unseen Instructions (10 variants): Examples reserved for evaluation: - "Pick up {B} and move it near {A}" - "Grab {B} and set it beside {A}" - "Use {a} to pick up {B}, move it near {A}" - "Grab {B} and place it beside {A}"
Instruction Diversity: - Verbs: grab, pick, lift, take, move, transfer, relocate - Spatial terms: near, beside, next to, by, close to - Styles: with/without arm mention, sequential vs compound - Tests: spatial relationship understanding, action sequence parsing
Episode-Level Instruction Assignment¶
During Training (Seen Instructions): - Each rollout: randomly select one instruction from seen set - 50+ instructions → low probability of repeating same instruction - Each instruction has equal selection probability - Ensures model doesn't overfit to specific phrasings
During Validation (Mixed): - IID validation: use seen instructions (same distribution as training) - OOD validation: use unseen instructions (test generalization) - Both sets used to measure different aspects of performance
Runtime Substitution:
When instruction is selected, placeholders are replaced:
- {A} → actual object identifier from episode
- {B} → actual object identifier from episode
- {a} → "left" or "right" based on arm selection algorithm
Example Runtime Process:
Template: "Use {a} to grab {B} and move it next to {A}"
Episode Configuration:
- Pot model: 060_kitchenpot/base3
- Can model: 105_sauce-can/base5
- Can X-position: +0.22 (right side)
- Selected arm: right
Final Instruction: "Use right to grab 105_sauce-can/base5 and move it next to 060_kitchenpot/base3"
Simplified for VLA: "Use right arm to grab the sauce can and move it next to the kitchen pot"
Instruction System Benefits¶
For Training: 1. Language Generalization: 50+ variants prevent language overfitting 2. Robustness: Model must understand task concept, not memorize phrases 3. Diversity: Different instruction styles cover various language patterns 4. Grounding: Schema system links language to specific objects in scene
For Evaluation: 1. Unseen Instructions: Tests true language understanding 2. IID vs OOD: Measures both in-distribution and out-of-distribution performance 3. Zero-Shot Transfer: Unseen instructions never appeared in training 4. Generalization Metric: Success rate on unseen instructions indicates robustness
Training Configuration¶
Dataset Configuration¶
Training Data Source:
- Pre-collected feasible seeds: 1000 per task
- Seeds validated through simulation to ensure solvability
- Stored in: verl/utils/envs/robotwin2/seeds/robotwin2_train_seeds.json
Validation Data: - IID Validation: 128 seeds from training distribution - OOD Validation: 128 seeds from held-out distribution - Total validation: 256 episodes per task
Data Loading: - Batch size: 64 task instances per training step - Samples per task: 8 (for GRPO grouping) - Total rollouts per step: 512 (64 × 8) - Shuffle: Enabled (random task ordering each epoch)
Rollout Configuration¶
Action Space: - Continuous 14-dimensional actions - 7 DOF per arm (position + orientation + gripper) - Action normalization: [-1, 1] range - Denormalization statistics: Task-specific (learned from demonstrations)
Action Chunking: - Chunks per VLA inference: 25 - Action tokens per chunk: 14 (one per dimension) - Total tokens per inference: 350 (25 × 14) - VLA calls per episode: ~8 (200 steps ÷ 25)
Rollout Parameters: - Temperature: 1.6 (Higher Rollout Temperature - paper enhancement) - Sampling: Enabled (do_sample=True) - Max environment steps: 200 - Micro-batch size: 1 (sequential processing) - Episode timeout: 200 steps = success or failure
Success Rate Targets¶
Based on paper results and dynamic sampling:
beat_block_hammer: - Initial SFT performance: ~25-35% (estimated) - Target after 300 steps: ~50-70% (paper-level) - Dynamic sampling keeps: 10-90% success rate episodes - Expected at convergence: ~60-80% success
move_can_pot: - Initial SFT performance: ~20-30% (estimated) - Target after 300 steps: ~45-65% (paper-level) - Dynamic sampling keeps: 10-90% success rate episodes - Expected at convergence: ~55-75% success
Training Progress Indicators: - Epoch 0-5: Mostly failures, learning basic grasping - Epoch 5-10: Increasing success, learning task structure - Epoch 10-15: Rapid improvement, discovering strategies - Epoch 15-20: Convergence, refining precision - Beyond 20: Marginal gains, potential overfitting
Computational Requirements¶
Per Training Step (~20 minutes): - Rollout phase: ~18 minutes (87.8% of time) - Environment initialization: ~2 minutes - VLA inference: ~2 minutes (8 calls × 300ms × 512 rollouts / 8 GPUs) - Physics simulation: ~14 minutes (25 steps × 60ms × 512 rollouts / 8 GPUs) - PPO update phase: ~2.4 minutes (11.8% of time) - Overhead: ~0.2 minutes (0.4% of time)
Full Training (300 steps as per paper): - Total time: ~300 steps × 20 minutes = 6000 minutes ≈ 100 hours ≈ 4.3 days - GPU usage: 8 × A100 GPUs (or equivalent) - Total rollouts: 300 × 512 = 153,600 episodes - Total VLA inferences: ~153,600 × 8 = ~1.2M forward passes - Total environment steps: ~153,600 × 100 = ~15M physics steps (average)
Resource Distribution: - GPU compute: ~20% utilization (during VLA inference only) - CPU compute: ~80% utilization (physics simulation dominates) - GPU memory: ~6 GB per GPU (FSDP sharding) - System memory: ~64 GB (environment states, buffers)
Summary¶
Task Comparison¶
| Aspect | beat_block_hammer | move_can_pot |
|---|---|---|
| Skill Type | Tool use + Contact | Pick & Place + Spatial |
| Complexity | Medium | Medium-High |
| Randomization | 1 object (block) | 2 objects (pot + can) |
| Success Tolerance | Tight (±2 cm) | Moderate (±20 cm horizontal) |
| Key Challenge | Precise contact | Spatial reasoning |
| Expected Success | 60-80% after RL | 55-75% after RL |
| Novel Behaviors | "Pushcut" (push vs lift) | Efficient low trajectories |
Key Experimental Insights¶
Why These Tasks? 1. Diverse Skills: Test different manipulation capabilities (tool use vs spatial reasoning) 2. Realistic Complexity: Neither too easy nor impossible, appropriate for RL learning 3. Measurable Success: Clear binary success criteria, no ambiguity 4. Rich Randomization: Sufficient variability to test generalization 5. Dual-Arm Platform: Tests coordination and arm selection logic
What RL Learns: 1. Beyond Demonstrations: Discovers strategies not in expert data ("pushcut") 2. Robustness: Handles high variability through 153k diverse rollouts 3. Efficiency: Finds shorter, more direct paths than demonstrations 4. Generalization: Works on unseen instructions and configurations 5. Exploration: Higher temperature (1.6) enables diverse strategy discovery
Expected Outcomes: - Both tasks show improvement with RL over SFT baseline - "Pushcut" phenomenon may emerge in beat_block_hammer - move_can_pot benefits from optimized trajectories - Success rates reach 60-75% range after 300 training steps - Unseen instructions show strong generalization (within 5-10% of seen)