World Models · Differential Repair · Frozen Verification

Verdi WM

Anonymous authors
Explore
01 · Overview

Verifiable differential repair for world models.

Diagnose a model failure, execute a targeted intervention, and verify what the evidence supports.

Figure 1 from the VerdiWM paper comparing VerdiWM with traditional research agents.
Figure 1 from the VerdiWM paper comparing VerdiWM with traditional research agents. Full resolution ↗
Interventional diagnosis

Paired probes isolate action sensitivity, context retention, first-frame anchoring, and sampler stress.

Localized repair

Motion priors, representation targets, adapters, and policy updates address specific failure mechanisms.

Independent verification

Frozen evaluators assess metrics, replay fidelity, protected trade-offs, and environment outcomes.

Evidence-licensed memory

Repair knowledge is reused only within its supported claim scope; unsupported transfer leads to abstention.

02 · Method

From diagnosis to claim-scoped memory.

The execution loop and verification loop have separate authority. Each intervention leaves a receipt before a claim is updated.

Paper figure: VerdiWM online decision loop for probing a world model, selecting repairs, executing trials, verifying outcomes, and accumulating repair knowledge.
Pipeline overview. The loop separates operational side effects from epistemic claim updates. Every trial produces a receipt; every reusable repair needs a gate-passing settlement vector rather than a single favorable score. Full resolution ↗
01 · Probe

Measure the model-specific response to controlled interventions.

02 · Repair

Choose a scoped intervention while preserving fixed interfaces and frozen components.

03 · Verify

Settle metrics, controls, validity gates, and environment success using an independent verifier.

04 · Remember

Store the model, intervention, receipt, evidence, and claim boundary for certified transfer.

03 · Ctrl-World

Physical consistency with LaMo + drift.

The baseline over-preserves static appearance and under-uses action-conditioned motion. LaMo and drift regularization target this failure while preserving the backbone interface.

Paper figure: Ctrl-World physical-consistency comparison with ground truth, VerdiWM, and Ctrl-World baseline.
LaMo + drift repair from the paper. VerdiWM diagnoses the Ctrl-World baseline as over-preserving static appearance and under-using action-conditioned motion. The repair adds a LaMo latent motion prior and drift regularization while keeping the backbone interface fixed. Motion smoothness rises from 68.75 to 81.53; LPIPS and FID also improve. Full resolution ↗

Quantitative results

Ctrl-World and VerdiWM metric comparison. Higher is better unless marked ↓.
Metric Ctrl-World VerdiWM Change
Semantic alignment 90.70 91.30 +0.60
Depth accuracy 93.28 96.16 +2.88
Aesthetic quality 32.75 36.12 +3.37
Background consistency 86.37 85.11 −1.26
Dynamic degree 45.32 43.78 −1.54
Flow score 26.54 30.35 +3.81
Motion smoothness 68.75 81.53 +12.78
Subject consistency 84.31 83.35 −0.96

WorldArena scores use a 0–100 scale. Background consistency, dynamic degree, and subject consistency remain lower. Tables 1–2 below also include PSNR, SSIM, LPIPS, and FID.

PDF screenshot of paper tables comparing Ctrl-World and VerdiWM metrics.
Direct paper screenshot. The table screenshot is cropped from the VerdiWM PDF rather than redrawn. Full resolution ↗

Qualitative replay · all four samples

Top row: ground truth. Middle row: VerdiWM. Bottom row: original Ctrl-World.

Sample 01

Ground truth / VerdiWM / Ctrl-World

Open video ↗
Sample 02

Ground truth / VerdiWM / Ctrl-World

Open video ↗
Sample 03

Ground truth / VerdiWM / Ctrl-World

Open video ↗
Sample 04

Ground truth / VerdiWM / Ctrl-World

Open video ↗
Paper figure: multi-view consistency comparison for the Ctrl-World experiment.
Paper figure for the same Ctrl-World campaign: VerdiWM improves two views while the paper records the remaining view-specific trade-offs. Full resolution ↗
04 · Cosmos-Predict2 REP

Representation repair and WorldArena evaluation.

DINO representation targets improve object boundaries, scene layout, geometry, and motion; the appearance-consistency trade-offs remain explicit.

Paper figure: Cosmos-Predict2 baseline vs VerdiWM REP qualitative comparison and trend curves.
DINO-REP repair from the paper. VerdiWM identifies unstable object boundaries, scene layout, and semantic relations in Cosmos-Predict2, then injects DINO features as higher-level representation targets. The repair improves semantic alignment, depth, flow, and motion smoothness, while background and subject consistency remain explicit claim-boundary coordinates. Full resolution ↗
REP replay and trends

Qualitative replay for the REP campaign: baseline is shown on the right, while the VerdiWM REP prediction is compared under the same context as the paper figure above.

Open video ↗

WorldArena · all eight metrics

Cosmos-Predict2 and VerdiWM metric comparison. Higher is better unless marked ↓.
Metric Cosmos-Predict2 VerdiWM Change
Semantic alignment 90.38 90.82 +0.44
Depth accuracy 95.35 97.05 +1.70
Aesthetic quality 36.05 37.30 +1.25
Background consistency 85.35 82.73 −2.62
Dynamic degree 42.18 43.15 +0.97
Flow score 29.55 30.70 +1.15
Motion smoothness 78.93 81.31 +2.38
Subject consistency 83.07 81.79 −1.28
REP leads on most WorldArena entries.

REP is higher on semantic alignment, depth accuracy, aesthetic quality, dynamic degree, flow score, and motion smoothness.

The largest gains are in geometry and temporal behavior.

Depth accuracy increases from 95.35 to 97.05, and motion smoothness increases from 78.93 to 81.31.

The baseline still keeps stronger identity and background consistency.

Background consistency and subject consistency remain lower for REP, which is logged as a boundary of the REP repair rather than hidden inside a single aggregate score.

PDF screenshot of the paper table comparing Cosmos-Predict2 and VerdiWM REP on WorldArena metrics.
Paper Table 3. Scores are reported on a 0–100 scale, where higher is better. This image is a direct crop from the paper. Full resolution ↗
05 · RoboCOIN

Adaptation to a new robot embodiment.

Agilex Cobot Magic · DreamDojo / Cosmos Predict2 · action-conditioned video-to-world

This experiment follows the paper's automatic-adaptation setting: given a new robot embodiment, action interface, camera layout, and task distribution, VerdiWM constructs the training and replay contract without a hand-written adapter. The target dataset is RoboCOIN-DataManager . The robot embodiment is Agilex Cobot Magic . From the dataset structure, embodiment metadata, camera viewpoint, and action interface, the system selects a compatible base model and continues training DreamDojo/Cosmos Predict2 action-conditioned video-to-world from the AgiBot post-train 2B weights.

Dataset Parsing

RoboCOIN introduces Agilex Cobot Magic rollouts, a new task distribution, object layout, and camera setup; the agent maps those signals into the training and replay contract.

Base Model Choice

The selected structure keeps Cosmos Predict2 video-to-world generation and adds action conditioning so predicted future frames respond to robot actions.

Control Check

Each video includes a no-action prediction. This checks that the model is using action input rather than performing unconditional visual completion.

Paper figure: RoboCOIN action-conditioned adaptation with ground truth, VerdiWM, and no-action injection rows.
RoboCOIN action-conditioned adaptation with ground truth, VerdiWM, and no-action injection rows. Full resolution ↗

Action-conditioned replay and no-action controls

Left: ground truth. Middle: action-conditioned prediction. Right: no-action prediction.

Open the Green Folder

Chinese task name: 打开绿色的文件夹. The rollout compares ground truth, the trained action-conditioned world model, and a no-action prediction.

Open video ↗
Open the Red Folder

Chinese task name: 打开红色的文件夹. The same three-column comparison is used to verify that action injection changes the predicted trajectory.

Open video ↗
Three-View Consistency

This mosaic replay stresses cross-view agreement. VerdiWM keeps the object interaction and motion progression coherent across three views, showing that the adaptation is not limited to a single camera stream.

Open video ↗
Paper figure: held-out RoboCOIN task comparisons between ground truth and VerdiWM predictions.
held-out RoboCOIN task comparisons between ground truth and VerdiWM predictions. Full resolution ↗
06 · Policy RL

Action-quality rollouts in RoboLab120.

Closed-loop success and verified progress ground the reward. Video and VLM checks remain audit evidence.

The paper tests whether video-level repair translates into action quality: the repaired Cosmos3 policy is evaluated as a closed-loop action generator. This block-stacking overview adds a RoboLab120 rollout view, with baseline and repaired/RL-updated behavior shown as paired clips.

Environment-Grounded Reward

R = 10 * terminal_success + 2 * verified_progress - bounded invalid_action_penalty - bounded trusted_safety_penalty.

Signal Source

Terminal success comes from RoboLab, and progress comes only from verified subtask/progress fields, matching the paper's rule that video scores do not substitute for environment success.

Claim Boundary

Failed trajectories are capped below successful trajectories; safety events enter the final quality gate; video/VLM checks remain audit-only evidence.

Paper figure: COSMOS3 long-horizon self-forcing comparison with ground truth, VerdiWM, and COSMOS3.
COSMOS3 long-horizon self-forcing comparison with ground truth, VerdiWM, and COSMOS3. Full resolution ↗
Block Stacking: Original vs RL

Four paired rollouts compare the original Cosmos3 behavior with the RL-trained policy under the same RoboLab120 task family.

Open video ↗
Spatial Re-layout Generalization

The scene is kept the same, but the block positions are swapped. The RL-trained policy preserves the block-stacking behavior under this object re-layout, showing generalization beyond the original placement.

Open video ↗
07 · Real robot

FRANKA task comparisons.

Three tasks · paired original and improved COSMOS3 · all clips shown at 4× speed

Real-robot FRANKA comparisons provide qualitative evidence for the action-quality claim. All clips are shown at 4x speed. The stacking task tests interference handling around a blue bowl, the tablecloth task tests path quality, and the drawer task tests occlusion-aware grasping from inside a bowl.

Task Set

Stack three blocks, fold the white tablecloth, and move a polyhedron from a bowl into a drawer before closing it.

Original COSMOS3

The stacking rollout is affected by the blue bowl, the folding rollout takes a longer path, and the drawer task fails under bowl occlusion.

Improved COSMOS3

The improved policy completes the stack, shortens the folding path, and finishes the occluded pick-place-close sequence. These clips complement the paper's closed-loop success-rate comparison rather than replacing it.

Stack three blocks

Stack Three Blocks / Original COSMOS3

Baseline rollout: the policy reaches the late-stage interaction, but bowl interference prevents a clean blue-block grasp.

Open video ↗
Stack Three Blocks / Improved COSMOS3

Improved rollout: the policy keeps the grasp sequence stable and completes the three-block stacking task.

Open video ↗

Fold the white tablecloth

Fold the White Tablecloth / Original COSMOS3

Baseline rollout: the task is eventually approached through a long, inefficient trajectory, exposing weak path planning on deformable-object manipulation.

Open video ↗
Fold the White Tablecloth / Improved COSMOS3

Improved rollout: the motion plan is shorter and cleaner, showing a clear gain in path planning after the COSMOS3 update.

Open video ↗

Bowl-to-drawer polyhedron

Bowl-to-Drawer Polyhedron / Original COSMOS3

Baseline rollout: the polyhedron is occluded inside the bowl, and the policy cannot establish a reliable grasp before the drawer step.

Open video ↗
Bowl-to-Drawer Polyhedron / Improved COSMOS3

Improved rollout: the policy grasps the occluded object, places it into the drawer, and closes the drawer to complete the task.

Open video ↗
Paper figure: real-robot Franka head-to-head rollouts comparing COSMOS3 and VerdiWM.
real-robot Franka head-to-head rollouts comparing COSMOS3 and VerdiWM. Full resolution ↗
Closed-loop manipulation success rates · paper Table 4
Environment Task Avg. steps COSMOS3 VerdiWM Δ
Simulation
n = 50
Pick-and-place 8 62% 74% +12 pts
Drawer opening 11 41% 59% +18 pts
Peg insertion 14 28% 52% +24 pts
Real robot
n = 20
Pick-and-place 8 55% 70% +15 pts
Drawer opening 11 35% 55% +20 pts

Paper Table 4 reports its own task set and sample sizes; individual demonstrations are qualitative evidence. Simulation uses paired initial-state seeds.

PDF screenshot of the paper table reporting closed-loop manipulation success rates.
PDF screenshot of the paper table reporting closed-loop manipulation success rates. Full resolution ↗
08 · Research memory

Repair recipes, evidence, and claim boundaries.

Each memory links model context, intervention, supporting evidence, and the documents that make the repair reusable.

Evidence-Licensed Memory

VerdiWM stores actions, receipts, model context, validity gates, and claim boundaries so repairs are reused only when evidence licenses transfer.

Model / architecture
Operational and epistemic control planes over world-model repair campaigns
Data / embodiment
Run receipts, paired probes, replay artifacts, metrics, and verifier outputs
Intervention
Claim-licensed repair memory and certified transfer
Evidence
Action -> receipt -> settlement vector -> claim-frontier update
VerdiWM Research Memory Graph Defines each node as a bounded repair episode with time, model, dataset, intervention, gate status, and claim scope.
  • Indexes model changes by claim scope rather than by loose chat history.
  • Connects experiment artifacts to reusable repair decisions.
  • Keeps the page-level visualization aligned with the evidence contract.
Evidence Claim Gate Skill Defines when a VerdiWM result is allowed to become a reusable repair claim.
  • Separates environment success from video or VLM audit evidence.
  • Requires action requests, execution receipts, and verification artifacts.
  • Updates the claim frontier only when evidence is sufficient.

Ctrl-World LaMo Repair

A physical-consistency diagnosis found copy-current-frame behavior and weak action-driven latent transition.

Model / architecture
Ctrl-World SVD-UNet with action encoder
Data / embodiment
Ctrl-World action-conditioned rollout data
Intervention
LaMo latent motion prior, drift loss, frozen SVD/VAE parts, hierarchical LR
Evidence
WorldArena metrics, PSNR/SSIM/LPIPS/FID, qualitative replay, and validity-gate annotations
Ctrl-World + LaMo Fusion Training Skill A concrete architecture and training recipe for adding latent motion priors to Ctrl-World.
  • Adds a motion-prior branch and drift readout in latent space.
  • Keeps stable SVD/VAE components frozen while training action-sensitive layers.
  • Uses targeted ablations to test whether motion is action-conditioned.
Evidence Claim Gate Skill Records how the LaMo intervention should be promoted only after evidence clears the claim gate.
  • Do not treat a lower scalar loss as sufficient mechanism evidence.
  • Require replay or latent-separation checks before widening the claim.
  • Keep the intervention scoped to the Ctrl-World physical-consistency failure mode.

Cosmos-Predict2 REP Repair

REP uses representation signals to improve structure, geometry, and motion while exposing an appearance-consistency boundary.

Model / architecture
Cosmos-Predict2 video-to-world rollout model
Data / embodiment
BridgeData replay set with WorldArena-style metrics
Intervention
Train-test aligned DINO-REP conditioning
Evidence
Semantic alignment, depth, flow, motion smoothness, qualitative replay, and gate-limited claim scope
DINO-REP Cosmos-Predict2 Repair Skill Encodes the REP intervention and the next-patch rule for appearance-consistency boundaries.
  • Treats depth and motion gains separately from appearance consistency.
  • Uses metric deltas to route the next patch rather than stopping at a score.
  • Targets REP injection position and appearance retention when needed.

RoboCOIN Transfer

The agent selected a compatible action-conditioned video-to-world setup from dataset structure, embodiment, viewpoint, and action interface.

Model / architecture
DreamDojo/Cosmos Predict2 action-conditioned video-to-world
Data / embodiment
RoboCOIN-DataManager / Agilex Cobot Magic
Intervention
Dataset adapter, embodiment-aware base-model selection, no-action control replay
Evidence
Green-folder, red-folder, and three-view consistency qualitative comparisons
RoboCOIN Transfer Adapter Skill Documents how the transfer node maps dataset and robot metadata into a model adaptation decision.
  • Uses embodiment and camera metadata before choosing the base model.
  • Keeps no-action prediction as an overfitting control.
  • Scopes the claim to selected RoboCOIN tasks and replay evidence.

Action-Quality Rollouts

The action-quality node checks whether long-horizon world-model repair improves policy rollouts under environment-grounded success criteria.

Model / architecture
Cosmos3-DROID policy
Data / embodiment
RoboLab120 block-stacking rollouts
Intervention
Reward-bounded RL with safety and invalid-action penalties
Evidence
Block-stack original-vs-RL paired rollouts and reward contract
RoboLab120 Policy-RL Reward Skill Defines the RL reward and its constraints so progress cannot fake terminal success.
  • Terminal success comes from RoboLab, not VLM scoring.
  • Failed trajectories are capped below successful trajectories.
  • Video and VLM checks remain audit-only.

FRANKA Real-Robot Trials

The real-robot node compares original Cosmos3 and repaired Cosmos3 on stack, fold, and occluded drawer manipulation.

Model / architecture
Repaired Cosmos3 real-robot policy stack
Data / embodiment
FRANKA real-world task videos
Intervention
Real-robot transfer with improved grasping, path planning, and occlusion handling
Evidence
4x H.264 comparison videos for stack, tablecloth, and drawer tasks
FRANKA Real-Robot Transfer Skill Summarizes the task set, observed original failures, and repaired Cosmos3 behavior.
  • Tracks stack, fold, and drawer tasks under the same real-robot section.
  • Uses accelerated public videos for presentation, not for changing task labels.
  • Links each qualitative claim to a visible comparison artifact.

Frozen Verifier

The verifier node prevents the memory graph from turning every observation into a broad claim.

Model / architecture
Model-agnostic claim policy
Data / embodiment
Execution receipts, metrics, replays, and environment outcomes
Intervention
Independent verification before claim-frontier update
Evidence
Pass/fail gates and bounded claim scopes
Evidence Claim Gate Skill Defines the gate logic for promoting a result into a reusable memory.
  • Record the proposed intervention and execution receipt.
  • Verify success with the authority appropriate for the claim.
  • Keep video/VLM audit evidence separate from environment success.
VerdiWM Research Memory Graph Explains how the graph indexes evidence by model, dataset, intervention, and claim scope.
  • Prevents score-only memory from hiding mechanism failures.
  • Makes prior patches inspectable before selecting the next action.
  • Connects knowledge nodes to concrete skill documents.
Open original ↗