PhysicalRSI 1.0
HKU MMLabKinetix AI
Line drawing of Charles Darwin
From so simple a beginning endless forms most beautiful and most wonderful have been, and are being, evolved
— Charles Darwin · On the Origin of Species

PhysicalRSI 1.0: Recursive Self-Harness for Scaling Embodied Skills

A Simple and Effective Baseline that Ranks No.1 on the Challenging RoboDojo Benchmark

PhysicalRSI 1.0 opens a new route to Everest’s summit.

The Everest mountain illustration from RoboDojo’s simulation leaderboard
Select a model to inspect its results.

With its diverse manipulation tasks, RoboDojo is our Everest for embodied AI. As of , PhysicalRSI ranks #1 on the RoboDojo-Sim leaderboard. We use RoboDojo as the main testbed to illustrate PhysicalRSI throughout what follows.

Our Proposal for next-generation embodied AI:
Native System 1–2 co-evolution

GPT-as-policy

System 2 used as System 1Observations and the goal go to GPT, a System 2 model. It directly emits low-level actions at every control step, using reasoning as the control loop. The illustrative decision/action rate is 0.1 to 1 Hz, not a measured model benchmark. Observation + goal GPTSystem 2≈0.1–1 Hz · est. Low-level actionsa₁ · a₂ · a₃ · … Physical world System 2 doingSystem 1’s job

(slow and expensive)

VLA / WAM

System 1 without a System 2 reasoning layerVLA or WAM handles direct sensorimotor execution. Without a separate System 2 layer, task-level reasoning and generalization are limited in this conceptual comparison. The illustrative action-control rate is 20 to 50 Hz, distinct from policy inference rate and not a measured model benchmark. Observation VLA / WAMSystem 1≈20–50 Hz control · est. Actions System 1 alonePhysical world Limited reasoning& generalization

(compounding error & OOD everywhere)

PhysicalRSI 1.0

System 2 directs and improves a System 1 execution harnessPlanning: a multimodal foundation model understands the scene and makes decisions as System 2. The illustrative planning estimate is 0.1 to 1 Hz, not a measured deployment rate. It composes skills including VLAs and programs, together with tools, inside an editable System 1 control harness. Acting: the System 1 harness invokes the physical API and acts in the world. The illustrative action-control estimate is 20 to 50 Hz, distinct from model-inference and simulation rates. Observations return to System 2; Reflecting is a separate feedback-driven rewrite-and-test loop that revises the harness and control chain across iterations. Observe + goal Multimodal foundation modelPlanning≈0.1–1 Hz · est.System 2 Acting · System 1≈20–50 Hz · est. SkillsVLA + programs Tools Physical API Physical world ReflectingS2 → S1v1 → v2 → …

(efficient and open-world generalizations)

Diverse self-evolving patterns

Representation, state, planning and control: different structures for the harness to revise.

σ = argsort(d, ↓)  ;  p*ᵢ = p₀ + i Δx êₓ

d: digit values · p*ᵢ: target slots

371φ(layout)731Δx

ΔH: add digit localization, ordering and slot assignment.

π₀.₅
π₀.₅ reaches among distractors without completing the digit row.
PhysicalRSI
Recognize the digits, sort by value, then map each digit to a placement slot.

ΔH denotes the skill structure revised by the harness. Diagrams summarize task geometry; videos are independent rollouts.

Main Algorithm. Embodied Self-Harness

The agent uses embodied feedback to rewrite its own harness. Validated edits become the harness it uses next.

Use your harness. Rewrite it. Run with it again.

Ak
The complete agent at iteration k: F operating through Hk.
F
System 2 · foundation model that understands, decides & rewrites.
Hk
The agent’s editable harness: routing, skills, tools & control code for System 1 execution.
τk
Embodied feedback: observations, actions, successes & failures.
H[0] = initial_harness()Startfor k in range(budget):    A[k] = Agent(F, H[k])S2 + H    plan = A[k].understand(task)S2    τ[k] = H[k].execute(plan)S1    variants = A[k].rewrite_many(H[k], τ[k])S2 → H    pool = [Agent(F, h) for h in variants]Candidates    survivor = select_on_eval([A[k], *pool])Selection    H[k+1] = survivor.harnessInherit

Improve = vary → evaluate → select → inherit.

“Everything should be made as simple as possible,
but no simpler.”

A minimalist implementation of PhysicalRSI 1.0.

System 2 · reviseHk+1=F(Hk,Ek+,Ek−)
System 1 · act(at,mt+1)=Hk(ot,g,mt;S,T)

F multimodal agent · H control code · E⁺ / E⁻ success / failure evidence · S / T skills / tools
o / g / m / a observation / goal / episode state / action · k / t revision / control step

Swipe to explore →

System 2 · self-harness
Rollout evidence
Initial visible block positionsBlocks being coveredController manipulating a coverController returning to retained positions

E⁺E⁻

System 2 · F

Multimodal agent

F(H,E+,E−)
Harness · H

System 1 control code

decideinvokeupdate
Hk → Hk+1
System 1 · embodied controlH[S, T; mₜ] → aₜ
Current block layout
Observation + goalot,g
Skill selectionσt=dH(ot,g,mt)

+ patterns− guards

Skills & tools
Physical APIat
Cover blocks · 1×

Make a mahjong kong

Illustrated loop · ≈5 rounds
01 · New layoutA new mahjong hand and tile arrangement

Find four matching tiles.

02 · Reflect

No selection. Diagnose the miss.

03 · Rewrite skill
Hk → Hk+1
tiles = perceive(rgb)
quad = match_four(tiles)
for tile in quad:
    pick_and_place(tile)

Revise recognition & tile handling.

04 · Test & select

Four of a kind. Keep the improvement.

Generate unseen layouts → iterate again
H0→H1→H2→H3→H4→H5

Harness Memory (e.g. cloth folding)

Keep the skill. Refresh the state.

Recorded experience

Code-policy skills persists

Fold, release & return
def fold_sleeve(arm, pick, place):
    arm.move_above(pick)
    home = arm.joints().copy()
    arm.move_to(pick)
    arm.close_gripper()
    arm.move_above(place)
    descent = [arm.joints().copy()]
    for q in arm.lower_to(place):
        descent.append(q.copy())
    arm.open_gripper()
    for q in reversed(descent):
        arm.move_joints(q, step=0.12)
    arm.move_joints(home, step=0.12)

Episode state resets

State is created anew in each rolloutThe code-policy skill saves an approach joint pose measured in the current rollout, then reads it after release and retreat. The stored pose is cleared for the next rollout. Current robot posewrite saved posefrom this rollout readReturn after releaseNew rollout → fresh state

Inherit skills. Compose new capabilities.

Skill inheritance

Carry visual grounding into a different task.

General pickup Locate → grasp → lift
Plug in charger Locate → align → insert
φB≡φA πB=HB[φA,σgrasp,σalign,σinsert]

A: pickup · B: charger insertion
φ: pixel-to-world grounding · σ: motor skill · H: task harness · π: policy

Shared grounding · sibling-module import
from ..general_pickup.rgb_perception import (
    pixel_to_world_on_height as ground,
)
x, y = ground(pixel, world_z=surface_z)

Skill composition

Read the equation. Place the missing digit.

Solve equation 8 × □ = 0
πeq=σplace∘σmove∘σpick∘f∘φ

φ: read · f: solve · σ: motor skill
∘: sequential composition over the task state

Code-policy sketch
scene = read_equation(observation)
digit, slot = solve(scene)  # 8 × ? = 0 → 0held = pick(scene.object(digit))
assert attached(held, observe())move(held, above(slot))if aligned(held, slot, observe()):
    place(held, slot)
Category comparison

≈ Estimated from available official task sheets.

π₀.₅: our motor-policy tool · DM0.5: VLA · Liber-0 Lite: WAM

Task-level official results
ModelSRScore

Official RoboDojo leaderboard ↗ · 28 Sep 2026 edition
PhysicalRSI result sheets & estimates ↓

52 settingsTask × layout

Selected full rollouts. SR and Score refer to the official evaluation; ≈ marks estimated means.

arrange largest number

SR 1%Score 3.87

Generalization
arrange largest number random

SR 21%Score 32.47

Generalization
fold clothes

SR 41%Score 48.80

Generalization
fold clothes random

SR 35%Score 43.47

Generalization
hang mugs

SR 0%Score 8.80

Generalization
make toast

SR 0%Score 6.00

Generalization
pack objects into box

SR 5%Score 26.67

Generalization
pack objects into box random

SR 0%Score 15.07

Generalization
pour liquid into cup

SR ≈ 94%Score ≈ 94.00

Generalization
pour liquid into cup random

SR 9%Score 9.33

Generalization
push T

SR 65%Score 65.33

Generalization
sort nesting dolls by size

SR 16%Score 16.00

Generalization
stack blocks

SR 9%Score 18.93

Generalization
stack bowls

SR 75%Score 78.47

Generalization
stack bowls random

SR 0%Score 10.00

Generalization
store laptop and headphones

SR 1%Score 13.07

Generalization
store laptop and headphones random

SR 1%Score 13.07

Generalization
sweep blocks

SR 0%Score 0.00

Generalization
classify objects

SR ≈ 36%Score ≈ 52.00

Long horizon
fill pen holder

SR 3%Score 21.87

Long horizon
make kong

SR ≈ 81%Score ≈ 81.00

Long horizon
organize table

SR 4%Score 27.33

Long horizon
play tic tac toe

SR 94%Score 98.33

Long horizon
put bottles into dustbin

SR 75%Score 83.70

Long horizon
cover blocks

SR ≈ 100%Score ≈ 100.00

Memory
match and pick from conveyor

SR 7%Score 7.33

Memory
press by number

SR 79%Score 79.33

Memory
swap T

SR ≈ 93%Score ≈ 93.00

Memory
swap blocks

SR 0%Score 0.00

Memory
align blocks

SR 86%Score 86.00

Open
classify objects by language

SR ≈ 8%Score ≈ 22.20

Open
general pickup

SR 47%Score 47.33

Open
solve equation

SR 60%Score 60.00

Open
build tower

SR 29%Score 37.13

Precision
deposit coin

SR 56%Score 62.27

Precision
fasten screws

SR 0%Score 7.60

Precision
insert key

SR 39%Score 48.43

Precision
insert tubes

SR 66%Score 78.93

Precision
plug in charger

SR 31%Score 30.67

Precision
pour balls into vase

SR 39%Score 38.67

Precision
hang mugs random

SR 0%Score 2.00

GeneralizationIncomplete rollout
make toast random

SR 0%Score 4.67

GeneralizationIncomplete rollout
sort nesting dolls by size random

SR 0%Score 0.00

GeneralizationIncomplete rollout
stack blocks random

SR 0%Score 2.40

GeneralizationIncomplete rollout
fill egg holder

SR 0%Score 1.90

Long horizonIncomplete rollout
play stacking toy

SR 0%Score 0.00

Long horizonIncomplete rollout
imitate sorting sequence

SR 0%Score 1.10

MemoryIncomplete rollout
pick from conveyor by image

SR 0%Score 0.00

OpenIncomplete rollout
pour by language

SR 0%Score 10.13

OpenIncomplete rollout
stack blocks by language

SR 0%Score 0.00

OpenIncomplete rollout
store tools in toolbox

SR 0%Score 6.67

OpenIncomplete rollout
play xylophone

SR 1%Score 0.67

PrecisionIncomplete rollout

Evolution trajectories

Fold clothes

Goal Fold both sleeves, then fold the hem up.

Fold the left sleeve

Recorded attempts play as one continuous video.

SRScore
Success rate and score across four selected evaluationsThree measured development results and the separately reported formal endpoint.

Harness Memory

Data & source records

Task metrics use the official result sheets. Missing means are estimated from available seeds; category estimates are marked ≈. Generalization uses its paired standard/random evaluation slice. Videos are selected rollouts, while evolution trajectories show development iterations.

Official results ↓ · Leaderboard snapshot · Full-rollout provenance · Skill video index

A library of embodied skills.

212 recorded clips across 40 tasks. Explore skills and their task sequences.

Task paths through a shared skill space

Hτ=Seq(σ1,…,σL),σi∈S
Recorded skills and task sequencesEach task follows its recorded skill sequence along a schematic surface. Positions are illustrative, not a learned embedding.
Selected skill

Loading skills…

One point · one recorded clip One path · one task sequenceSchematic layout · S: skill library · L: sequence length
CODE-POLICY SKILLS212 clips · 40 tasks
Each block is one video clipHover a block to preview its videoScroll to see every task
SELECTED SKILL TRACE001 / 212
TASK / SKILL

Select a skill block

The selected video clip and its recorded time window appear here.

———

HOW TO READ IT The atlas contains 212 extracted skill windows. Each detail view links one block to one task, camera, episode, and time-bounded video clip.

SUCCESSFUL ROLLOUT

Task rollout