Today, we are releasing OriginX, an independent research extension of Xiaomi-Robotics-1-RoboCasa365. It pairs an inherited Action LoRA adapter with a continuous conditioning branch that participates directly in action generation.
The frozen B2000 checkpoint achieved 59.84% overall success in our full RoboCasa365 evaluation: 1,496 successes across 50 tasks and 2,500 episodes. Every planned episode has a terminal record, including all 1,004 policy failures. The complete base and both adaptation weights are available on Hugging Face; source code, the technical report, and evaluation evidence are on GitHub.
Research update · October 9. The conditional Astra rescue study now reports 56 confirmed rescues from the fixed 1,004 original failures. The original benchmark result remains unchanged.
Measuring performance across the full benchmark
RoboCasa365 spans atomic skills and longer, composite tasks. Our evaluation covers all 50 target tasks in pretraining kitchen scenes, with exactly 50 trials per task. The checkpoint, task schedule, seeds, and original task horizons were fixed before evaluation.
Because every task has the same number of trials, the task-weighted mean equals the pooled success rate. The split results show a clear descriptive gap: unseen composite tasks remain substantially harder than atomic tasks.
Success across three task groups
Author-run, frozen B2000 evaluation. Overall weights the splits by task count: 18 / 16 / 16. Unmatched scores from separate leaderboard entries do not establish a causal improvement.
Atomic-Seen reaches 83.00%, while Composite-Unseen reaches 33.75%—a 49.25 percentage-point gap within this evaluation. This experiment documents that difference; it does not determine why it occurs.
Learning a new conditioning pathway
OriginX follows a three-stage lineage: the public Xiaomi checkpoint, the A2000 Action LoRA adapter, and the B2000 continuous conditioning branch. B2000 freezes the complete A2000 policy and trains 4,216,839 additional parameters.
The branch combines observation features with robot state to produce a continuous representation used by the native action generator. An auxiliary stage-classification objective supplies training supervision. At inference, the model uses observations—not ground-truth stage labels.
A small branch, a frozen foundation
- 01
Public base
Xiaomi-Robotics-1-RoboCasa365
Inherited checkpoint - 02
Action LoRA
A2000 · rank 16, alpha 16
1,613 updates · 80,000 windows - 03
Continuous branch
B2000 · 4.22M trainable parameters
2,000 updates · 256,000 windows
The design combines three choices:
- Zero-initialized modulation. The branch’s output projection starts at zero, so it initially leaves the inherited policy unchanged.
- Frozen-teacher preservation. A velocity objective compares student and teacher under matched inputs, noise, and time. Its weight is 1.0; the auxiliary stage loss has weight 0.05.
- Conditioning during action generation. The branch is part of the executed policy, rather than an auxiliary head that is discarded at inference.
B2000 was trained for 2,000 updates with global batch size 128, using 256,000 sampled windows. Its recorded training duration is 8,453.98 seconds. The already-completed training is separate from the full evaluation reported here.
The audited local sample plans cover Human300 tasks. B2000 mixes Human300 replay with a 200-trajectory pilot equally; the pilot includes 16 Composite-Seen tasks, with 168 trajectories for training and 32 held out. None of the 16 official Composite-Unseen target tasks appear in these local plans. This audit does not independently cover the inherited checkpoint’s full pretraining history.
What the paired comparison tells us
The full evaluation is a single-arm result. To understand whether the complete adaptation improves the original base, the earlier paired confirmation remains essential.
Across 600 paired episodes on 30 tasks, the original model succeeded 413 times and B2000 succeeded 419 times. The estimated gain was one percentage point, with a reported 95% interval of [−2.17, 4.17] percentage points. It did not meet the predeclared criterion of at least a five-point gain and a positive interval lower bound.
The estimated gain remains uncertain
413 / 600 successes
419 / 600 successes
95% interval · McNemar p = 0.6173
This comparison changes both the adapter and the branch: it contrasts the original base with A2000 + B2000. There is no paired A2000-only control, so it cannot isolate the branch’s effect. The result does not establish equivalence, but it does not support a reliable gain from the complete adaptation either. The newer 2,500-episode result uses a different protocol and has no matched full baseline or branch ablation; it does not overturn the earlier negative evidence.
A complete evaluation makes the outcome auditable. Controlled comparisons are still needed to establish the method’s benefit.
Astra assistance: 56 confirmed rescues
After the frozen evaluation, we tested a separate, conditional rescue procedure on the 1,004 unique episodes that the original B2000 run failed. At the first regular policy query at or after half of the original horizon, gpt-6-astra with high reasoning effort received the task description and three camera views, then suggested a subgoal. The frozen B2000 policy continued action generation with that guidance. No weights were retrained, and Astra was not given hidden simulator state or ground-truth stage labels.
The result is 56 confirmed rescues out of the fixed 1,004-case selection: 5.58%. Unknown outcomes and initial-state deviations remain in that denominator. All selected cases were attempted, but unresolved and incomparable cases prevent this from being a complete valid paired result for every case.
Every selected case stays accounted for
| Classification | Cases | Interpretation |
|---|---|---|
| Confirmed rescued | 56 | Strict rescue criteria and valid assistance evidence |
| Not rescued | 813 | No confirmed rescue in a valid assisted comparison |
| Unknown | 28 | Unresolved execution or assistance evidence |
| Initial-state deviation | 107 | Replay does not meet the initial-state matching criterion |
| Fixed selection | 1,004 | No discarded unknown or deviation cases |
There are 874 matched, complete replay pairs (M). Of these, 869 pairs (E) also have verified Astra assistance applied with valid evidence. Thus 56/869 = 6.44% is the conditional rescue proportion among verified assisted pairs. The five additional pairs in M do not have successful, verified assistance; M is not the count of valid Astra comparisons.
The original benchmark remains 1,496/2,500 = 59.84%. This study selected known failures and added an intervention; the rescues cannot replace original outcomes or be added to the original numerator to produce a new leaderboard score. A fresh evaluation of an assistance policy across all planned episodes would be needed to measure overall benefit, regressions, and cost.
Pilot, model-call accounting, and provenance
A four-case development pilot recorded 0/4 rescues and one CLI call. Its case IDs were 48, 97, 249, and 345. Those IDs remain in their unique primary-study partitions, but pilot outcomes and usage are reported separately and never substituted for primary outcomes.
The primary study used 135 unique CLI calls, including failed or unapplied calls: 134 completed calls and one error call. One call lacks token counters, so full token totals remain unknown. The known counts are 2,592,713 input tokens, 195,237 output tokens, and 75,511 reasoning tokens; these categories are not added into a synthetic total. Dollar cost was not recorded. The separately reported pilot used one additional CLI call.
The first-run partition contributes 36 unique cases and the continuation contributes 968. The frozen partition rule prevents duplicate counting and cross-run outcome substitution. The study retains all four final classifications and the source-manifest identities.
Read the full rescue report · Revised paper (PDF) · Supplement · Rescue evidence release
Keeping the evaluation accountable
We retain both successes and failures, along with their original identities. Every failed episode reaches its registered task horizon. There are no replacement seeds, discarded difficult tasks, or completed-subset scores.
- Frozen execution. The B2000 checkpoint, 50-task manifest, environment and policy seeds, and original horizons remain fixed throughout the run.
- Native reset checks. The evaluator verifies the original reset implementation before and after each episode.
- Complete records. Final auditing verifies all 2,500 result hashes and 2,500 guard-record hashes, with no missing or infrastructure-unknown outcomes.
The simulation uses CPU OSMesa with rendering decimation, context-lifetime management, and readback guards. These implementation choices are disclosed for organizer and independent review; they are not a claim of official validation. Inference uses five Euler steps, replanning every 16 steps, observation history length four, observation interval two, and crop ratio 0.95.
View artifact identities and protocol details
- A2000 Action LoRA · SHA-256
- 5892389863751f2245f83a6d3b96bf5725a052aa5b023230b093d8d8ece70ac4
- B2000 branch · SHA-256
- 41988b92391953687b39a0bc5a35bce962fbfcd08fbf6e550108430850d9944f
- Evaluation configuration · SHA-256
- a3b0dbed4fb26ddcc472645fb9cc13008694293c23c49c325cc7524b4a93cd04
- Evaluation manifest · SHA-256
- cecee25e5750605c2c83643bca0fed55c8c4bfcb46d672a934f0c2d29e760257
Seed range: 2090900000–2090902499. RoboCasa 1.0.1; original registered task horizons. The later portable inference package is distinct from the original evaluation entry point. Packaging is not a repeat of the full-model rollout evaluation.
Explore the release
The complete OriginX model assets are on Hugging Face, while implementation code is maintained on GitHub. The model snapshot includes the three original base shards, tokenizer/configuration assets, adapter-originx-2000.pt, and branch-00002000.pt—approximately 10.15 GB in total. The base shards account for 10,106,433,336 bytes. The original v1.0.0 evaluation evidence remains unchanged.
The GitHub preparation helper assembles the model assets with the pinned model code and verifies their original identity. Start with the preparation and inference guide. Packaging performs no training or new benchmark evaluation.
OriginX builds on Xiaomi-Robotics-1-RoboCasa365. See the official multi-task evaluation protocol and submission requirements for benchmark context.
Cite this technical report
@misc{ma2026originx,
title = {OriginX: An Audited 2,500-Episode RoboCasa365 Evaluation},
author = {Ma, Qinzhen},
year = {2026},
note = {Technical report; author-run results, pending organizer review},
url = {https://github.com/Quinn-Ma/OriginX}
}