Why long horizons change the problem
A model can generate a strong short clip and still fail as a world model. Once outputs become the context for the next generation step, small representation errors feed back into the system. The research question is therefore not only whether a frame looks plausible, but whether the model preserves motion, state, identity, and action semantics through an extended causal rollout.
A trajectory-level evaluation
The project evaluates progressive rollout lengths and separates prompt difficulty, motion difficulty, action reversals, and repeated versus unseen controls. It combines visible consistency and action-adherence metrics with probes of VAE latents, transformer hidden states, and key-value histories.
The central diagnostic compares the model’s online history with refreshed or recomputed histories. This makes it possible to distinguish errors already present in the visible latent trajectory from errors amplified inside the causal model’s memory.
From diagnosis to intervention
Post-training experiments target partition consistency across rollout boundaries while protecting short-horizon quality. The aim is not a cosmetic temporal filter: it is a model that remains controllable when generation becomes interaction.
Media note: the approved video clips, prompts, and captions have not yet been supplied. The live site uses a systems illustration rather than invented research output and is ready to accept optimized WebM/MP4 assets when cleared.