RESEARCH PREVIEW / PUBLIC SIMULATOR RELEASE
roborsi: Simulator-Verified Robot Self-Improvement
An embodied-agent framework that converts simulator-verified interaction into inspectable, reusable code-backed skills.
Abstract
roborsi separates agent-visible reasoning from execution authority. A Planner formulates a strategy, an Engineer composes camera-grounded tools, and a Reviewer diagnoses the resulting trace. Candidate code is evaluated in an isolated overlay and retained only after a native simulator verdict confirms success. This design makes adaptation inspectable, replayable, and measurable in both task outcomes and inference cost.
LIBERO-Plus adaptive Pass@2398/840
Gain over fixed evaluation+16.3 pp
Adaptive development coverage95/120
Strict Standard-130 Pass@1067/130
REMOTION FILM / 60 SECONDS / NARRATED
Simulator-confirmed rollouts and measured self-improvement
The narrated Remotion film presents thirteen simulator-confirmed successful rollouts in a rotating 3 x 3 layout. It also shows a matched ACT case before and after corrective fine-tuning, adaptive coverage, and Code-on/off efficiency.
Verified rollout set13 success videosAll currently published successful rollouts appear across two 3 x 3 pages.
Adaptive coverage32 to 83/120Cross-release development coverage; not fixed-policy Pass@10.
Retained code skill+7.5 pp / -29.4%Matched success delta and median-token reduction.
Matched ACT case: before and after
On libero_spatial_swap/0 seed 3, bounded ACT completed 120 transport steps but under-transported and lost the hold during placement. After corrective data and fine-tuning, ACT completed 304 steps and wrist-verified placement reached a native simulator success.
Same task and seed. Videos are aligned by normalized episode progress because the archived clips last 6.8 and 16.2 seconds. This is one matched case, not aggregate policy evidence; the 32 to 83 curve is a separate cross-release measure.
Published rollout evidence
All fourteen published recordings, including the retained ACT failure, are indexed below. Each entry provides direct playback and the most complete trace supported by the archived evidence.
The three historical RoboTwin recordings retain final-verdict-only traces because their original per-call logs were not archived. No missing calls are inferred.
01
System architecture and trust boundary
Agent-visible evidence and simulator truth occupy different trust domains. The visible loop can inspect RGB-D observations, registered skills, and its own trace. It cannot inspect task predicates, rewards, hidden object poses, or success latches.
head + wrist RGB-D Compose
base + compound skills Act
IK + joint trajectories Adapt
reviewed code overlay
02
Trace-backed simulation evidence
Each case below is attached to one named task, one exact seed, one ordered registered-tool trace, and a final simulator verdict confirming success. Selecting Trace opens the complete public call chain.
03
LIBERO-Plus adaptive evaluation
We evaluate 840 stratified perturbation identities across seven categories and three short suites. Fixed receives one attempt per identity. Adaptive releases retry selected fixed failures with reviewed code-backed changes. Success is counted only by the final simulator verdict.
Metered adaptive tokens1.474Ball valid adaptive attempts
Active wall time16.35hunion of active intervals
Valid adaptive attempts592infrastructure excluded
Unmetered VLM calls0Planner + Engineer + Reviewer
Release contribution
Newly solved identities are de-duplicated against fixed and every earlier adaptive stage.
Simulator-confirmed perturbation cases
04
Evaluation protocols and results
Each protocol addresses a distinct research question and is reported under its own evaluation contract.
Separating coverage, adaptation, and efficiency
Pass@k is not an evolution curve. Strict Standard-130 rises from 23 to 67 tasks because each task receives more ordered seeds; cumulative coverage is monotonic by construction. Its per-round median tokens and wall time are non-monotonic, the residual task set changes, and releases change between rounds.
Strict Standard-130Coverage, not evolution23 to 67 with a growing seed budget
Adaptive releasesBetter task coverage32/120 to 83/120 across evolving rounds
Matched Code-on/offFaster execution-29.4% median tokens; -17.0% median wall time
Historical RoboTwin development result
The historical RoboTwin reports list a pure Engineer baseline alongside the Planner-Engineer-Reviewer framework. Task-level Pass@k is 9/50 and 36/50 respectively, while strict per-episode success was 104/422 (24.6%). The campaign window was 87.26 h.
Pure Engineer9/50task coverage
Three roles36/50task-level Pass@k
Strict episodes104/42224.6% success
Parallel overlap87%not sequential causality
Predicate-confirmed episodes
The archived videos and corresponding final task-predicate outcomes are unchanged. Per-episode tool logs were not retained, so the dialogs report only the available final verdict.
Cross-release adaptive coverage
The measured ten-round sequence grows from 32 to 83 solved tasks while releases evolve between rounds. Later releases solve 12 additional tasks, yielding cumulative cross-release coverage of 95/120.
- Sequential endpoint
- 83/120
- Metered tokens
- 2.342B
- Active wall time
- 11.09h
The 95/120 result is cumulative development coverage, not a fixed-policy or single-release Pass@10.
Strict Standard-130
All 130 tasks receive up to ten ordered seeds. No task checker, action-success latch, hidden pose, or policy checkpoint is available to the agent. All 846 scheduled task-seed attempts have final simulator verdicts.
- Task failures
- 779
- Total tokens
- 2.882B
- VLM calls
- 43,076
Trend interpretation: the first five versus last five round medians show 13.6% fewer tokens, 3.1% fewer VLM calls, but 8.0% longer wall time. Because releases and residual tasks change, none of these descriptive shifts establishes that the model gets progressively faster.
Effect of retaining a code-backed skill
Code-on exposes the retained visual_pick_place compound. Code-off keeps the model, base tools, release, tasks, seeds, and budget matched but removes that compound.
- McNemar exact
- p = 8.06e-5
- Task-cluster 95% CI
- +3.7 to +11.5 pp
- Task-level Pass@5
- 71/120 vs 58/120
Efficiency uses a separate 118-task matched seed-21 panel. It is not inferred from the 1,200-episode significance experiment.
ACT corrective-learning case
01Code-backed graspvisual source acquisition
→02ACT transport304 corrective steps
→03Code-backed placementnative final verdict
This is one matched seed-3 success on libero_spatial_swap/0. It does not establish held-out performance or generalization.
05
Related evaluation context
Protocols differ; this is not a shared leaderboard. Policy checkpoints, views, success feedback, horizons, task catalogs, and retry semantics differ across systems. Reported scores stay attached to their source protocol.
| System | Evaluation | Reported result | Execution contract | Comparability |
|---|
Current evidence gaps
Supported claims
Evidence note: protocol boundaries and source references remain attached to each reported result.