← All work
ongoingMar 2026 – Present

Causal Video & Interactive World Models

Diagnosing and reducing the representation drift that compounds across long autoregressive rollouts.

  • Video generation
  • World models
  • Long-horizon consistency

Problem

Short video clips can look convincing while repeated causal rollout quietly accumulates errors in motion, identity, control, and internal state.

Approach

Evaluate the system as a trajectory rather than a set of clips, then connect visible failure to drift in latent, hidden-state, and key-value histories.

Role

Research lead for modelling, evaluation, and post-training

Why long horizons change the problem

A model can generate a strong short clip and still fail as a world model. Once outputs become the context for the next generation step, small representation errors feed back into the system. The research question is therefore not only whether a frame looks plausible, but whether the model preserves motion, state, identity, and action semantics through an extended causal rollout.

A trajectory-level evaluation

The project evaluates progressive rollout lengths and separates prompt difficulty, motion difficulty, action reversals, and repeated versus unseen controls. It combines visible consistency and action-adherence metrics with probes of VAE latents, transformer hidden states, and key-value histories.

The central diagnostic compares the model’s online history with refreshed or recomputed histories. This makes it possible to distinguish errors already present in the visible latent trajectory from errors amplified inside the causal model’s memory.

From diagnosis to intervention

Post-training experiments target partition consistency across rollout boundaries while protecting short-horizon quality. The aim is not a cosmetic temporal filter: it is a model that remains controllable when generation becomes interaction.

Media note: the approved video clips, prompts, and captions have not yet been supplied. The live site uses a systems illustration rather than invented research output and is ready to accept optimized WebM/MP4 assets when cleared.