Skip to content
Apex
All research
MethodsApex CFD Group18 min read

Long-rollout stability in autoregressive physics

How Aether holds trajectory drift below 4% out to 5,200 time-steps — and why standard autoregressive surrogates do not. We describe the mechanism, the measurement methodology, the cases that still break, and what mid-trajectory steering buys you.


Abstract

Long-rollout stability across three transient-CFD benchmark classes. Aether holds drift in integrated quantities below 4% out to 5,200 time-steps on the 280B sparse-MoE — four times the horizon of the autoregressive baseline at matched parameter count.

We characterise the long-rollout stability of Aether against three classes of transient computational-fluid-dynamics benchmark, against an autoregressive baseline of similar parameter count, and against high-fidelity reference solutions. The model achieves stable trajectories at four times the horizon of the baseline on the cases we measured, with drift in integrated quantities held below 4% out to 5,200 time-steps for the 280B sparse-MoE configuration. We document the mechanism (a soft conservation-law term in the pretraining objective), present ablations isolating its contribution, describe the implications for mid-trajectory steering, and characterise the cases where stability still degrades.

1. Introduction

Autoregressive surrogates of physical systems have a well-documented failure mode: errors compound. The mechanism is straightforward — each predicted next-state inherits the error of all prior predictions, and feedback through the dynamics generally amplifies small errors rather than damping them. The result is that single-step accuracy is misleading: a model with 99% per-step accuracy can produce a divergent trajectory after a few thousand steps.

This matters because almost every useful downstream use of a simulation-native model involves rolling forward many steps. A transient computational-fluid-dynamics case is typically thousands of time-steps; a molecular-dynamics campaign is hundreds of thousands; a manufacturing-process drift study spans weeks of telemetry. If the model cannot hold integrated quantities accurate over those horizons, single-step accuracy is irrelevant.

Standard remedies — teacher forcing, scheduled sampling, distillation against the reference — improve single-step loss but typically do not materially extend the stable horizon. We took a different approach: regularising the pretraining objective against violations of physical conservation laws, on the hypothesis that the dominant source of compounding error is the model gradually violating mass, energy and momentum conservation across long trajectories. This paper documents what happened.

2. How we measure

Stable horizon is defined as the time-step at which drift in a selected integrated quantity first exceeds a threshold. We choose three integrated quantities per benchmark — typically a global energy budget, a global mass-conservation check, and a domain-specific quality metric (e.g., surface heat flux integrated over a face, or wall-pressure RMS). The threshold is 4% absolute deviation from the reference; the choice is conservative enough to capture practical loss-of-utility and tight enough that the metric is informative.

We use three benchmark classes. First, the standard impinging-jet cases with measured Nusselt distributions — relatively well-behaved, useful for calibrating the metric. Second, transient external aerodynamics on a fixed-wing platform with measured surface pressure — more demanding, with realistic unsteady features. Third, a multi-bay rotor with wind-on-rig data — the hardest of the three, with shear-layer dynamics and acoustic coupling. For each case we run Aether at matched time-step against a high-fidelity reference solution and measure drift across the selected integrated quantities.

We compare against three baselines. The first is a vanilla autoregressive transformer of matched parameter count, trained on the same corpus without the conservation regularisation. The second is a published physics-informed neural-network baseline of comparable size. The third is the high-fidelity reference itself with its time-step doubled — a sanity check that the reference is converged in the horizon we are reporting.

3. Headline results

4,000
stable steps · 40B dense · <4% drift
5,200
stable steps · 280B sparse-MoE · <4% drift
≈800
stable steps · matched autoregressive baseline

On the standard impinging-jet benchmark, the 40B dense model holds within 4% drift out to 4,000 time-steps; the 280B sparse-MoE model holds to 5,200. The matched autoregressive baseline crosses the 4% threshold at approximately 800 steps. The published physics-informed baseline performs better than the vanilla autoregressive baseline (crossing at ≈1,400 steps) but worse than Aether.

The pattern is consistent across the three benchmark classes. On the harder rotor case, the absolute horizons are shorter for every model — Aether 40B dense crosses at 2,800 steps, the baseline at 500 — but the ratio is preserved at roughly 4–6×. The aerodynamics case sits between the two.

4. Mechanism — conservation-aware regularisation

We added an auxiliary loss term during pretraining that penalises violations of mass and energy conservation across rolled-out trajectory segments. The term is soft — the model can violate conservation when the data requires it (e.g., open-domain flows with mass crossing the boundary) but the gradient signal is large enough to materially change the learned dynamics.

Concretely: during training, we periodically unroll the model for K = 8 to 32 steps from a held-out segment of the corpus. We compute the integrated mass and energy at the start and at each rolled-out step, compare against the values from the reference data, and add a quadratic penalty on the violations. K varies stochastically during training; this prevents the model from memorising a particular rollout length.

Ablation: removing the conservation term reduces the stable horizon by approximately 3.4× across the benchmark set. Adding the term back to a checkpoint that did not have it during pretraining recovers roughly 60% of the improvement, suggesting some of the gain is from the trajectory shape that the regularised model converges to during early training rather than from the regularisation acting in isolation.

5. Steering — mid-trajectory interruption and re-conditioning

The same mechanism that improves long-horizon stability allows mid-trajectory steering. A rollout can be interrupted, conditioned on new state (a redesigned geometry, a different boundary condition, a manually-imposed initial condition), and resumed from the interruption point. The model handles the discontinuity by re-equilibrating over roughly 20–50 steps rather than diverging.

This is practically useful. A designer running an optimisation can interrupt a long rollout when an early-step metric crosses a threshold, modify the geometry, and resume — without restarting the simulation from scratch. The amortised cost of a design exploration drops by a factor proportional to how often the modification is made early in the rollout. Customers report 4–8× reductions in total compute for shape-optimisation workflows that use this pattern.

The mechanism for clean re-equilibration is the same conservation-aware learned dynamics. The model has internalised that energy and mass have to balance; when the boundary condition changes, it re-equilibrates to a state that satisfies the new constraints rather than diverging from a state that satisfied the old ones.

6. Where it still fails

Three classes of cases still degrade rollout stability.

First, the onset of laminar-turbulent transition in shear layers. The model lags the reference for tens of steps after the transition event begins, then catches up. The drift during the lag often exceeds the 4% threshold, terminating the stable horizon. This is a known weakness; we are extending the training corpus with more shear-layer transition data.

Second, shock-shock and shock-boundary-layer interactions at high Mach. The model handles isolated shocks well; multi-shock interactions are harder. Customers running hypersonic external-flow campaigns use a tool-call fallback to a high-fidelity DNS or LES solver at the suspected interaction region; the runtime supports this gracefully.

Third, very-long-horizon manufacturing-process drift (weeks to months of process telemetry). The 1M-token context supports the horizon, but cumulative drift in calibration coefficients exceeds threshold at the very end of the longest cases. The fix is more frequent measurement updates, which are usually available; production deployments use them.

7. Implications for the model family

Long-rollout stability is one of the two quality axes we publish per release (the other is single-step cross-entropy). The two axes are not strongly correlated — a model with better cross-entropy can have worse rollout stability, and vice versa — so we track them independently and we report both in the model card.

For Aether 1.5 we expect to lift the 4% threshold horizon by approximately 30% for the 40B dense and 50% for the 280B sparse-MoE, driven by additional transition-event training data, an improved conservation-regularisation schedule, and a small architectural change to the time-step-aware attention layer.

8. Conclusion

Long-rollout stability in autoregressive physics is dominated by whether the model has internalised conservation laws. A soft penalty during pretraining accomplishes more than aggressive distillation or teacher forcing, and unlocks mid-trajectory steering as a side effect. Aether holds within 4% drift to four to six times the horizon of a matched autoregressive baseline; the remaining failure modes are rare-event onsets, which we address with tool-call fallbacks to high-fidelity solvers.

References

  1. [1]Apex Research. Aether 1.0 — a simulation-native foundation model. Apex Research Paper (2026).
  2. [2]Apex Research. Scaling laws for simulation-native pretraining. Apex Research Paper (2026).
  3. [3]Apex Research. Cross-domain transfer in simulation-native foundation models. Apex Research Paper (2026).
  4. [4]Cooper, P. et al. Experimental measurements of impinging jet flow and heat transfer. Int. J. Heat Mass Transfer (1993).
  5. [5]Raissi, M. et al. Physics-informed neural networks. J. Comp. Phys. (2019).

Want the underlying model?

Aether powers every workload above. Request access and we'll show you what one AI does with your data.