RESEARCH PREVIEW / PUBLIC SIMULATOR RELEASE

roborsi: Simulator-Verified Robot Self-Improvement

An embodied-agent framework that converts simulator-verified interaction into inspectable, reusable code-backed skills.

Abstract

roborsi separates agent-visible reasoning from execution authority. A Planner formulates a strategy, an Engineer composes camera-grounded tools, and a Reviewer diagnoses the resulting trace. Candidate code is evaluated in an isolated overlay and retained only after a native simulator verdict confirms success. This design makes adaptation inspectable, replayable, and measurable in both task outcomes and inference cost.

LIBERO-Plus adaptive Pass@2398/840

Gain over fixed evaluation+16.3 pp

Adaptive development coverage95/120

Strict Standard-130 Pass@1067/130

REMOTION FILM / 60 SECONDS / NARRATED

Simulator-confirmed rollouts and measured self-improvement

The narrated Remotion film presents thirteen simulator-confirmed successful rollouts in a rotating 3 x 3 layout. It also shows a matched ACT case before and after corrective fine-tuning, adaptive coverage, and Code-on/off efficiency.

Narrated Remotion film. Each robot sequence is linked to a named task and final simulator or task-predicate outcome. Plotted values come directly from the published results.

Verified rollout set13 success videosAll currently published successful rollouts appear across two 3 x 3 pages.

Adaptive coverage32 to 83/120Cross-release development coverage; not fixed-policy Pass@10.

Retained code skill+7.5 pp / -29.4%Matched success delta and median-token reduction.

Matched ACT case: before and after

On libero_spatial_swap/0 seed 3, bounded ACT completed 120 transport steps but under-transported and lost the hold during placement. After corrective data and fine-tuning, ACT completed 304 steps and wrist-verified placement reached a native simulator success.

Same task and seed. Videos are aligned by normalized episode progress because the archived clips last 6.8 and 16.2 seconds. This is one matched case, not aggregate policy evidence; the 32 to 83 curve is a separate cross-release measure.

Published rollout evidence

All fourteen published recordings, including the retained ACT failure, are indexed below. Each entry provides direct playback and the most complete trace supported by the archived evidence.

The three historical RoboTwin recordings retain final-verdict-only traces because their original per-call logs were not archived. No missing calls are inferred.

01

System architecture and trust boundary

Agent-visible evidence and simulator truth occupy different trust domains. The visible loop can inspect RGB-D observations, registered skills, and its own trace. It cannot inspect task predicates, rewards, hidden object poses, or success latches.

FIGURE 1 / SYSTEM
01Plannerinstruction + visible memory
02EngineerRGB-D + registered tools
03Reviewertrace + failure diagnosis
Sense
head + wrist RGB-D
Compose
base + compound skills
Act
IK + joint trajectories
Adapt
reviewed code overlay
POST-EPISODE HOST BOUNDARY Host-only adjudicator promote only after native simulator success
Figure 1. The self-improvement loop can propose and test code, but only the host-side simulator verdict can authorize promotion.

02

Trace-backed simulation evidence

Each case below is attached to one named task, one exact seed, one ordered registered-tool trace, and a final simulator verdict confirming success. Selecting Trace opens the complete public call chain.

03

LIBERO-Plus adaptive evaluation

We evaluate 840 stratified perturbation identities across seven categories and three short suites. Fixed receives one attempt per identity. Adaptive releases retry selected fixed failures with reviewed code-backed changes. Success is counted only by the final simulator verdict.

Fixed261/84031.1%
Adaptive Pass@2398/84047.4%
Measured gain+16.3 pp+137 solved identities
Fixed and adaptive LIBERO-Plus success, perturbation and suite breakdowns, and release-stage contribution
Figure 2. Final 840-identity LIBERO-Plus panel. Fixed and Adaptive use the same selected identities; adaptive successes accumulate across reviewed release stages.

Metered adaptive tokens1.474Ball valid adaptive attempts

Active wall time16.35hunion of active intervals

Valid adaptive attempts592infrastructure excluded

Unmetered VLM calls0Planner + Engineer + Reviewer

Release contribution

Newly solved identities are de-duplicated against fixed and every earlier adaptive stage.

Simulator-confirmed perturbation cases

04

Evaluation protocols and results

Each protocol addresses a distinct research question and is reported under its own evaluation contract.

CLAIM ATTRIBUTION

Separating coverage, adaptation, and efficiency

protocol-specific

Pass@k is not an evolution curve. Strict Standard-130 rises from 23 to 67 tasks because each task receives more ordered seeds; cumulative coverage is monotonic by construction. Its per-round median tokens and wall time are non-monotonic, the residual task set changes, and releases change between rounds.

Strict retry coverage, adaptive LIBERO development, historical RoboTwin campaign progress, and matched Code-on efficiency
Figure 3. Four curves with four different contracts. Adaptive LIBERO and historical RoboTwin show solution-discovery progress; matched Code-on/off supplies the causal efficiency evidence. Strict Pass@k reports retry coverage only.

Strict Standard-130Coverage, not evolution23 to 67 with a growing seed budget

Adaptive releasesBetter task coverage32/120 to 83/120 across evolving rounds

Matched Code-on/offFaster execution-29.4% median tokens; -17.0% median wall time

HISTORICAL CROSS-PLATFORM EVIDENCE

Historical RoboTwin development result

36/50

The historical RoboTwin reports list a pure Engineer baseline alongside the Planner-Engineer-Reviewer framework. Task-level Pass@k is 9/50 and 36/50 respectively, while strict per-episode success was 104/422 (24.6%). The campaign window was 87.26 h.

Pure Engineer9/50task coverage

Three roles36/50task-level Pass@k

Strict episodes104/42224.6% success

Parallel overlap87%not sequential causality

Predicate-confirmed episodes

The archived videos and corresponding final task-predicate outcomes are unchanged. Per-episode tool logs were not retained, so the dialogs report only the available final verdict.

DEVELOPMENT RESULT

Cross-release adaptive coverage

95/120

The measured ten-round sequence grows from 32 to 83 solved tasks while releases evolve between rounds. Later releases solve 12 additional tasks, yielding cumulative cross-release coverage of 95/120.

Figure 3. Measured sequential campaign.
Sequential endpoint
83/120
Metered tokens
2.342B
Active wall time
11.09h

The 95/120 result is cumulative development coverage, not a fixed-policy or single-release Pass@10.

STRICT AGENT EVALUATION

Strict Standard-130

67/130

All 130 tasks receive up to ten ordered seeds. No task checker, action-success latch, hidden pose, or policy checkpoint is available to the agent. All 846 scheduled task-seed attempts have final simulator verdicts.

Figure 4. Task-level Pass@k.
Task failures
779
Total tokens
2.882B
VLM calls
43,076

Trend interpretation: the first five versus last five round medians show 13.6% fewer tokens, 3.1% fewer VLM calls, but 8.0% longer wall time. Because releases and residual tasks change, none of these descriptive shifts establishes that the model gets progressively faster.

MATCHED FIVE-SEED ABLATION

Effect of retaining a code-backed skill

+7.5 pp

Code-on exposes the retained visual_pick_place compound. Code-off keeps the model, base tools, release, tasks, seeds, and budget matched but removes that compound.

Figure 5. Episode success, 600 pairs per arm.
McNemar exact
p = 8.06e-5
Task-cluster 95% CI
+3.7 to +11.5 pp
Task-level Pass@5
71/120 vs 58/120

Efficiency uses a separate 118-task matched seed-21 panel. It is not inferred from the 1,200-episode significance experiment.

MATCHED CASE

ACT corrective-learning case

1 success

01Code-backed graspvisual source acquisition

02ACT transport304 corrective steps

03Code-backed placementnative final verdict

This is one matched seed-3 success on libero_spatial_swap/0. It does not establish held-out performance or generalization.

05

Related evaluation context

Protocols differ; this is not a shared leaderboard. Policy checkpoints, views, success feedback, horizons, task catalogs, and retry semantics differ across systems. Reported scores stay attached to their source protocol.

SystemEvaluationReported resultExecution contractComparability

Current evidence gaps

    Supported claims

      Evidence note: protocol boundaries and source references remain attached to each reported result.