Stationary targets versus drifting environments

Continual learning has long treated performance loss on previously observed data as a signal of failure. This convention originates from prediction settings where a correct label remains correct indefinitely. World models invert this assumption. Their prediction target is the environment itself, which evolves. A fact that was true at training time may become false as dynamics shift. The authors argue that discarding stale knowledge is not a bug but expected behavior for systems that model changing worlds.

Non-stationary ground truth is well explored in concept drift literature and in the temporal factuality of language models. However, world models introduce a distinctive duality: they must retain some knowledge permanently—physics engines, object permanence—while rapidly revising other knowledge—instance-specific facts about moving objects or changing rules. Prior work has not connected these two requirements within a single framework.

Forgetting metrics that erase revision credit

Standard forgetting metrics compute the gap between a model's performance on a fixed dataset and its performance after adaptation. Because they compare against a static baseline, a model that never updates its weights— a frozen checkpoint—achieves the lowest reported forgetting. Consequently, these metrics rank inertia as optimality. A world model that correctly revises a outdated belief about, say, traffic patterns receives the same penalty as one that catastrophically loses a previously learned physics rule. The paper demonstrates this mismatch through two observations: standard metrics always prefer frozen checkpoints, and existing physical-reasoning benchmarks evaluate only frozen snapshots, never an adapting model.

The consequence is a distorted evaluation landscape. Researchers optimizing for low forgetting scores inadvertently freeze their models, forgoing the adaptation that world models require. The authors cite this as a root cause of the gap between benchmarks that test reasoning ability and systems that must operate in non-stationary environments.

Stratified retention by invariance timescale

The central contribution is the proposal to separate retained knowledge by invariance timescale. Invariants—such as the laws of gravity or the fact that objects persist when out of sight—must never be revised. These are properties of the world model that provide structural stability. Instance-level facts—such as the current position of a specific agent or the latest version of a mutable rule—should be revised as soon as the environment changes. The paper formalizes this split as a timescale stratification: long-timescale invariants anchor the model, while short-timescale facts operate in a revision-permeable layer.

This stratification is not a new architecture per se; it is a principle for organizing what the model stores and how metrics assess it. The authors argue that any continual world model without such a split will either over-forget invariants or under-revise instance facts, both paths leading to degraded behavior.

Differential retention: joint invariant testing and revision latency

To measure a model that both preserves invariants and revises facts, the paper proposes differential retention. This metric does not aggregate the two concerns into a single score. Instead, it reports two numbers across the entire adaptation stream:

  • Invariant regression testing: performance on a fixed set of invariant tasks (e.g., object permanence, physical interaction rules) measured at every adaptation step.
  • Revision latency: the number of adaptation steps required for the model to update a previously learned instance fact once the environment changes.

Reporting these jointly, without combining them into an average, lets observers see whether a model preserves long-timescale knowledge while adapting short-timescale knowledge. A model that freezes after acquiring invariants will show high invariant regression scores but poor revision latency. A model that revises everything will show low latency but dropping invariant scores. The differential pattern reveals which trade-off a given system makes.

The "without aggregation" clause is intentional. Merging the two numbers into one obscures the tension the paper identifies. Practitioners can inspect the pair and decide whether their application demands stricter invariant preservation or faster fact revision.

Experimental evaluation

The authors evaluate differential retention across a suite of world model adaptation tasks. Invariants are tested on physical reasoning benchmarks that require knowledge of object permanence, gravity, and collision dynamics. Instance facts are tested on a stream of environment changes where specific object properties or rules shift mid-stream. Revision latency is measured as the step count at which the model's prediction for a changed fact aligns with the new ground truth.

Baselines include a frozen checkpoint, a model that replay-stores all past observations, and a model with uniform learning rates across all parameters. The differential retention results show that the frozen baseline scores perfectly on invariant regression but never revises instance facts. The replay baseline retains invariants well but suffers high latency because it must scan through many past examples before the new fact takes effect. A model with parameter isolation—separate weights for invariant and fact layers—achieves low revision latency while preserving invariant scores above a defined threshold.

The paper also reports ablations of the isolation strategy. Removing the invariant mask causes rapid invariant decay. Removing the fact revision trigger causes latency to increase unboundedly. Both failures confirm the necessity of the stratified design.

acknowledged limits and trade-offs

The authors acknowledge that defining the invariance split requires domain knowledge. What counts as a physics invariant in one application may be a mutable fact in another. The paper does not provide a fully automated method for partitioning parameters; it proposes the principle and demonstrates it with hand-crafted splits.

Another limit is the evaluation suite itself. The physical reasoning benchmarks are sparse, and the instance-fact stream is synthetic. Real-world world models—such as those in robotics or interactive simulation—may expose additional structure not captured by the current tasks. The authors note that differential retention is a framework for measurement, not a complete solution, and that future work must fill in domain-specific partitions.

A third trade-off is computational. Isolating parameters or maintaining separate masks adds overhead, though the authors observe that the cost is modest compared with the benefit of meaningful revision metrics. They do not report exact runtime numbers in the abstract, but the experimental section implies that the overhead is linear in the number of partitioned groups.

What this means for working developers

For teams building world models that must adapt over time, the paper offers a concrete evaluation recipe. Instead of reporting a single forgetting number, report invariant regression and revision latency as a pair. This pair immediately reveals whether a model is stuck in a frozen state or is revising everything indiscriminately. If the pair shows high invariant scores and low latency, the system has achieved the stratified retention the paper advocates.

On the architecture side, the paper suggests that developers consider a parameter separation strategy. This does not need to be complex: a binary mask that flags which parameters belong to the invariant set can be sufficient. During adaptation, updates to masked parameters are dampened or projected out, while unmasked parameters receive normal learning signals. The authors' ablations show that this simple mechanism preserves long-timescale knowledge while allowing short-timescale facts to shift.

The differential retention framework also suggests a monitoring practice. During deployment, track the two numbers across the adaptation stream. A drift in invariant regression toward chance level signals that the invariant mask is too permissive or that the environment has changed at a timescale that overwhelms the model's anchoring facts. A latency number that grows over time suggests that the fact-revision mechanism is losing effectiveness, perhaps because the environment changes faster than the model can update.

Finally, the paper invites researchers to move beyond benchmarks that evaluate only frozen checkpoints. Any new world model benchmark should include an adaptation phase and report both retention metrics. Until then, differential retention provides a minimal, comparable signal across adaptations.

Read the paper on arXiv