Research-Loop Reliability
A world-model improvement may require diagnosing long-horizon failure, editing objectives or conditioning paths, running distributed jobs, and deciding whether the measurements really support the intended conclusion.
Click a 3D memory object above to jump here and inspect the corresponding model, dataset, intervention, evidence, and skill documents.
A project-page memory graph that indexes experiments by model, dataset, intervention, and evidence gate.
The paper starts from a practical bottleneck in world-model research: the hard part is not only scaling a backbone, but reliably turning a model failure, a code change, and a noisy experiment into a defensible research conclusion.
A world-model improvement may require diagnosing long-horizon failure, editing objectives or conditioning paths, running distributed jobs, and deciding whether the measurements really support the intended conclusion.
VerdiWM separates the authority to execute experiments from the authority to update claims. Actions change models; only fenced evidence, receipts, and frozen verification can move the claim frontier.
Verified repair experience is not copied by name or intuition. It moves to a new backbone only when probe responses, support, sign agreement, and effect bounds license transfer; otherwise the system abstains.
General software agents can edit files, run scripts, and report scores. VerdiWM targets a stricter control surface: execution can propose evidence, but frozen claim gates decide whether a repair becomes reusable knowledge.
A Devin-style workflow sees that Ctrl-World lacks an action module, opens models/unet.py, inserts an ActionEncoder with an LLM, and restarts python train.py. The intervention stays at the file-system and code-text layer.
In the Ctrl-World campaign, VerdiWM diagnoses copy-current-frame behavior as weak action-driven latent transition, then adds a LaMo motion prior and drift regularization under the original backbone interface.
Existing agents act like external programmers. VerdiWM changes the measured repair geometry while preserving hook contracts and verifier scope.
An AutoGPT-style workflow receives "optimize FID", changes the learning rate from 1e-4 to 1e-5, reruns training, prints a new FID, and stops.
In the Cosmos-Predict2 campaign, REP improves depth, flow, and motion smoothness, but background and subject consistency fall. The claim is narrowed to structure-related coordinates instead of reported as an unconditional win.
Existing agents run and read a score. VerdiWM runs, diagnoses, attributes, freezes verification, updates a bounded claim, or abstains.
An MLAgentBench-style report may observe loss = 2.3 and recommend lowering the learning rate. The diagnosis remains at the scalar metric level.
VerdiWM applies paired intervention probes around the frozen model: action scaling, context retention, first-frame anchoring, and sampler stress form a response fingerprint before any expensive transfer.
Existing agents diagnose symptoms. VerdiWM measures whether the target model has the deficiency that a stored repair actually addresses.
A campaign starts from a goal and validity gates, fingerprints the model with paired probes, compiles typed repairs, verifies effects under a frozen evaluator, and updates the claim DAG only when evidence obligations are satisfied.
The experiments test whether VerdiWM can diagnose a world-model failure, select a localized repair, preserve validity gates, and transfer only when the target backbone passes a certificate.
Benchmark-facing metrics cover semantic alignment, depth accuracy, flow, motion smoothness, subject/background consistency, PSNR, SSIM, LPIPS, and FID.
Campaign-level gates prevent score-only claims by tracking action sensitivity, rollout fidelity, protected trade-off metrics, and policy-level outcomes.
Three-row replay for the Ctrl-World physical-consistency campaign. The comparison follows the paper-style layout: ground truth, VerdiWM repair, then the original baseline.
Qualitative replay for the REP campaign: baseline is shown on the right, while the VerdiWM REP prediction is compared under the same context as the paper figure above.
WorldArena-style scoring gives a broader view of the REP rollout than PSNR and SSIM alone. The table compares Cosmos-Predict2 with REP on semantic, geometric, aesthetic, background, dynamics, flow, smoothness, and subject consistency metrics.
REP is higher on semantic alignment, depth accuracy, aesthetic quality, dynamic degree, flow score, and motion smoothness.
Depth accuracy increases from 95.35 to 97.05, and motion smoothness increases from 78.93 to 81.31.
Background consistency and subject consistency remain lower for REP, which is logged as a boundary of the REP repair rather than hidden inside a single aggregate score.
This experiment follows the paper's automatic-adaptation setting: given a new robot embodiment, action interface, camera layout, and task distribution, VerdiWM constructs the training and replay contract without a hand-written adapter. The target dataset is RoboCOIN-DataManager. The robot embodiment is Agilex Cobot Magic. From the dataset structure, embodiment metadata, camera viewpoint, and action interface, the system selects a compatible base model and continues training DreamDojo/Cosmos Predict2 action-conditioned video-to-world from the AgiBot post-train 2B weights.
Chinese task name: 打开绿色的文件夹. The rollout compares ground truth, the trained action-conditioned world model, and a no-action prediction.
Chinese task name: 打开红色的文件夹. The same three-column comparison is used to verify that action injection changes the predicted trajectory.
This mosaic replay stresses cross-view agreement. VerdiWM keeps the object interaction and motion progression coherent across three views, showing that the adaptation is not limited to a single camera stream.
The paper tests whether video-level repair translates into action quality: the repaired Cosmos3 policy is evaluated as a closed-loop action generator. This block-stacking overview adds a RoboLab120 rollout view, with baseline and repaired/RL-updated behavior shown as paired clips.
Four paired rollouts compare the original Cosmos3 behavior with the RL-trained policy under the same RoboLab120 task family.
The scene is kept the same, but the block positions are swapped. The RL-trained policy preserves the block-stacking behavior under this object re-layout, showing generalization beyond the original placement.
Real-robot FRANKA comparisons provide qualitative evidence for the action-quality claim. All clips are shown at 4x speed. The stacking task tests interference handling around a blue bowl, the tablecloth task tests path quality, and the drawer task tests occlusion-aware grasping from inside a bowl.
Baseline rollout: the policy reaches the late-stage interaction, but bowl interference prevents a clean blue-block grasp.
Improved rollout: the policy keeps the grasp sequence stable and completes the three-block stacking task.
Baseline rollout: the task is eventually approached through a long, inefficient trajectory, exposing weak path planning on deformable-object manipulation.
Improved rollout: the motion plan is shorter and cleaner, showing a clear gain in path planning after the COSMOS3 update.
Baseline rollout: the polyhedron is occluded inside the bowl, and the policy cannot establish a reliable grasp before the drawer step.
Improved rollout: the policy grasps the occluded object, places it into the drawer, and closes the drawer to complete the task.