Paired probes isolate action sensitivity, context retention, first-frame anchoring, and sampler stress.
Verifiable differential repair for world models.
Diagnose a model failure, execute a targeted intervention, and verify what the evidence supports.
Motion priors, representation targets, adapters, and policy updates address specific failure mechanisms.
Frozen evaluators assess metrics, replay fidelity, protected trade-offs, and environment outcomes.
Repair knowledge is reused only within its supported claim scope; unsupported transfer leads to abstention.
From diagnosis to claim-scoped memory.
The execution loop and verification loop have separate authority. Each intervention leaves a receipt before a claim is updated.
Measure the model-specific response to controlled interventions.
Choose a scoped intervention while preserving fixed interfaces and frozen components.
Settle metrics, controls, validity gates, and environment success using an independent verifier.
Store the model, intervention, receipt, evidence, and claim boundary for certified transfer.
Physical consistency with LaMo + drift.
The baseline over-preserves static appearance and under-uses action-conditioned motion. LaMo and drift regularization target this failure while preserving the backbone interface.
Quantitative results
| Metric | Ctrl-World | VerdiWM | Change |
|---|---|---|---|
| Semantic alignment | 90.70 | 91.30 | +0.60 |
| Depth accuracy | 93.28 | 96.16 | +2.88 |
| Aesthetic quality | 32.75 | 36.12 | +3.37 |
| Background consistency | 86.37 | 85.11 | −1.26 |
| Dynamic degree | 45.32 | 43.78 | −1.54 |
| Flow score | 26.54 | 30.35 | +3.81 |
| Motion smoothness | 68.75 | 81.53 | +12.78 |
| Subject consistency | 84.31 | 83.35 | −0.96 |
WorldArena scores use a 0–100 scale. Background consistency, dynamic degree, and subject consistency remain lower. Tables 1–2 below also include PSNR, SSIM, LPIPS, and FID.
Qualitative replay · all four samples
Top row: ground truth. Middle row: VerdiWM. Bottom row: original Ctrl-World.
Ground truth / VerdiWM / Ctrl-World
Ground truth / VerdiWM / Ctrl-World
Ground truth / VerdiWM / Ctrl-World
Ground truth / VerdiWM / Ctrl-World
Representation repair and WorldArena evaluation.
DINO representation targets improve object boundaries, scene layout, geometry, and motion; the appearance-consistency trade-offs remain explicit.
Qualitative replay for the REP campaign: baseline is shown on the right, while the VerdiWM REP prediction is compared under the same context as the paper figure above.
WorldArena · all eight metrics
| Metric | Cosmos-Predict2 | VerdiWM | Change |
|---|---|---|---|
| Semantic alignment | 90.38 | 90.82 | +0.44 |
| Depth accuracy | 95.35 | 97.05 | +1.70 |
| Aesthetic quality | 36.05 | 37.30 | +1.25 |
| Background consistency | 85.35 | 82.73 | −2.62 |
| Dynamic degree | 42.18 | 43.15 | +0.97 |
| Flow score | 29.55 | 30.70 | +1.15 |
| Motion smoothness | 78.93 | 81.31 | +2.38 |
| Subject consistency | 83.07 | 81.79 | −1.28 |
REP is higher on semantic alignment, depth accuracy, aesthetic quality, dynamic degree, flow score, and motion smoothness.
Depth accuracy increases from 95.35 to 97.05, and motion smoothness increases from 78.93 to 81.31.
Background consistency and subject consistency remain lower for REP, which is logged as a boundary of the REP repair rather than hidden inside a single aggregate score.
Adaptation to a new robot embodiment.
Agilex Cobot Magic · DreamDojo / Cosmos Predict2 · action-conditioned video-to-world
This experiment follows the paper's automatic-adaptation setting: given a new robot embodiment, action interface, camera layout, and task distribution, VerdiWM constructs the training and replay contract without a hand-written adapter. The target dataset is RoboCOIN-DataManager . The robot embodiment is Agilex Cobot Magic . From the dataset structure, embodiment metadata, camera viewpoint, and action interface, the system selects a compatible base model and continues training DreamDojo/Cosmos Predict2 action-conditioned video-to-world from the AgiBot post-train 2B weights.
RoboCOIN introduces Agilex Cobot Magic rollouts, a new task distribution, object layout, and camera setup; the agent maps those signals into the training and replay contract.
The selected structure keeps Cosmos Predict2 video-to-world generation and adds action conditioning so predicted future frames respond to robot actions.
Each video includes a no-action prediction. This checks that the model is using action input rather than performing unconditional visual completion.
Action-conditioned replay and no-action controls
Left: ground truth. Middle: action-conditioned prediction. Right: no-action prediction.
Chinese task name: 打开绿色的文件夹. The rollout compares ground truth, the trained action-conditioned world model, and a no-action prediction.
Chinese task name: 打开红色的文件夹. The same three-column comparison is used to verify that action injection changes the predicted trajectory.
This mosaic replay stresses cross-view agreement. VerdiWM keeps the object interaction and motion progression coherent across three views, showing that the adaptation is not limited to a single camera stream.
Action-quality rollouts in RoboLab120.
Closed-loop success and verified progress ground the reward. Video and VLM checks remain audit evidence.
The paper tests whether video-level repair translates into action quality: the repaired Cosmos3 policy is evaluated as a closed-loop action generator. This block-stacking overview adds a RoboLab120 rollout view, with baseline and repaired/RL-updated behavior shown as paired clips.
R = 10 * terminal_success + 2 * verified_progress - bounded invalid_action_penalty - bounded trusted_safety_penalty.
Terminal success comes from RoboLab, and progress comes only from verified subtask/progress fields, matching the paper's rule that video scores do not substitute for environment success.
Failed trajectories are capped below successful trajectories; safety events enter the final quality gate; video/VLM checks remain audit-only evidence.
Four paired rollouts compare the original Cosmos3 behavior with the RL-trained policy under the same RoboLab120 task family.
The scene is kept the same, but the block positions are swapped. The RL-trained policy preserves the block-stacking behavior under this object re-layout, showing generalization beyond the original placement.
FRANKA task comparisons.
Three tasks · paired original and improved COSMOS3 · all clips shown at 4× speed
Real-robot FRANKA comparisons provide qualitative evidence for the action-quality claim. All clips are shown at 4x speed. The stacking task tests interference handling around a blue bowl, the tablecloth task tests path quality, and the drawer task tests occlusion-aware grasping from inside a bowl.
Stack three blocks, fold the white tablecloth, and move a polyhedron from a bowl into a drawer before closing it.
The stacking rollout is affected by the blue bowl, the folding rollout takes a longer path, and the drawer task fails under bowl occlusion.
The improved policy completes the stack, shortens the folding path, and finishes the occluded pick-place-close sequence. These clips complement the paper's closed-loop success-rate comparison rather than replacing it.
Baseline rollout: the policy reaches the late-stage interaction, but bowl interference prevents a clean blue-block grasp.
Improved rollout: the policy keeps the grasp sequence stable and completes the three-block stacking task.
Baseline rollout: the task is eventually approached through a long, inefficient trajectory, exposing weak path planning on deformable-object manipulation.
Improved rollout: the motion plan is shorter and cleaner, showing a clear gain in path planning after the COSMOS3 update.
Baseline rollout: the polyhedron is occluded inside the bowl, and the policy cannot establish a reliable grasp before the drawer step.
Improved rollout: the policy grasps the occluded object, places it into the drawer, and closes the drawer to complete the task.
| Environment | Task | Avg. steps | COSMOS3 | VerdiWM | Δ |
|---|---|---|---|---|---|
|
Simulation
n = 50 |
Pick-and-place | 8 | 62% | 74% | +12 pts |
| Drawer opening | 11 | 41% | 59% | +18 pts | |
| Peg insertion | 14 | 28% | 52% | +24 pts | |
|
Real robot
n = 20 |
Pick-and-place | 8 | 55% | 70% | +15 pts |
| Drawer opening | 11 | 35% | 55% | +20 pts |
Paper Table 4 reports its own task set and sample sizes; individual demonstrations are qualitative evidence. Simulation uses paired initial-state seeds.
Repair recipes, evidence, and claim boundaries.
Each memory links model context, intervention, supporting evidence, and the documents that make the repair reusable.
Evidence-Licensed Memory
VerdiWM stores actions, receipts, model context, validity gates, and claim boundaries so repairs are reused only when evidence licenses transfer.
- Model / architecture
- Operational and epistemic control planes over world-model repair campaigns
- Data / embodiment
- Run receipts, paired probes, replay artifacts, metrics, and verifier outputs
- Intervention
- Claim-licensed repair memory and certified transfer
- Evidence
- Action -> receipt -> settlement vector -> claim-frontier update
- Indexes model changes by claim scope rather than by loose chat history.
- Connects experiment artifacts to reusable repair decisions.
- Keeps the page-level visualization aligned with the evidence contract.
- Separates environment success from video or VLM audit evidence.
- Requires action requests, execution receipts, and verification artifacts.
- Updates the claim frontier only when evidence is sufficient.
Ctrl-World LaMo Repair
A physical-consistency diagnosis found copy-current-frame behavior and weak action-driven latent transition.
- Model / architecture
- Ctrl-World SVD-UNet with action encoder
- Data / embodiment
- Ctrl-World action-conditioned rollout data
- Intervention
- LaMo latent motion prior, drift loss, frozen SVD/VAE parts, hierarchical LR
- Evidence
- WorldArena metrics, PSNR/SSIM/LPIPS/FID, qualitative replay, and validity-gate annotations
- Adds a motion-prior branch and drift readout in latent space.
- Keeps stable SVD/VAE components frozen while training action-sensitive layers.
- Uses targeted ablations to test whether motion is action-conditioned.
- Do not treat a lower scalar loss as sufficient mechanism evidence.
- Require replay or latent-separation checks before widening the claim.
- Keep the intervention scoped to the Ctrl-World physical-consistency failure mode.
Cosmos-Predict2 REP Repair
REP uses representation signals to improve structure, geometry, and motion while exposing an appearance-consistency boundary.
- Model / architecture
- Cosmos-Predict2 video-to-world rollout model
- Data / embodiment
- BridgeData replay set with WorldArena-style metrics
- Intervention
- Train-test aligned DINO-REP conditioning
- Evidence
- Semantic alignment, depth, flow, motion smoothness, qualitative replay, and gate-limited claim scope
- Treats depth and motion gains separately from appearance consistency.
- Uses metric deltas to route the next patch rather than stopping at a score.
- Targets REP injection position and appearance retention when needed.
RoboCOIN Transfer
The agent selected a compatible action-conditioned video-to-world setup from dataset structure, embodiment, viewpoint, and action interface.
- Model / architecture
- DreamDojo/Cosmos Predict2 action-conditioned video-to-world
- Data / embodiment
- RoboCOIN-DataManager / Agilex Cobot Magic
- Intervention
- Dataset adapter, embodiment-aware base-model selection, no-action control replay
- Evidence
- Green-folder, red-folder, and three-view consistency qualitative comparisons
- Uses embodiment and camera metadata before choosing the base model.
- Keeps no-action prediction as an overfitting control.
- Scopes the claim to selected RoboCOIN tasks and replay evidence.
Action-Quality Rollouts
The action-quality node checks whether long-horizon world-model repair improves policy rollouts under environment-grounded success criteria.
- Model / architecture
- Cosmos3-DROID policy
- Data / embodiment
- RoboLab120 block-stacking rollouts
- Intervention
- Reward-bounded RL with safety and invalid-action penalties
- Evidence
- Block-stack original-vs-RL paired rollouts and reward contract
- Terminal success comes from RoboLab, not VLM scoring.
- Failed trajectories are capped below successful trajectories.
- Video and VLM checks remain audit-only.
FRANKA Real-Robot Trials
The real-robot node compares original Cosmos3 and repaired Cosmos3 on stack, fold, and occluded drawer manipulation.
- Model / architecture
- Repaired Cosmos3 real-robot policy stack
- Data / embodiment
- FRANKA real-world task videos
- Intervention
- Real-robot transfer with improved grasping, path planning, and occlusion handling
- Evidence
- 4x H.264 comparison videos for stack, tablecloth, and drawer tasks
- Tracks stack, fold, and drawer tasks under the same real-robot section.
- Uses accelerated public videos for presentation, not for changing task labels.
- Links each qualitative claim to a visible comparison artifact.
Frozen Verifier
The verifier node prevents the memory graph from turning every observation into a broad claim.
- Model / architecture
- Model-agnostic claim policy
- Data / embodiment
- Execution receipts, metrics, replays, and environment outcomes
- Intervention
- Independent verification before claim-frontier update
- Evidence
- Pass/fail gates and bounded claim scopes
- Record the proposed intervention and execution receipt.
- Verify success with the authority appropriate for the claim.
- Keep video/VLM audit evidence separate from environment success.
- Prevents score-only memory from hiding mechanism failures.
- Makes prior patches inspectable before selecting the next action.
- Connects knowledge nodes to concrete skill documents.