Skip to content
Apex
All research
MethodsApex Research48 min read

Scaling laws for simulation-native pretraining

Loss-versus-compute curves for simulation-trained foundation models look different from text-trained ones. We report the empirical exponents, per-domain breakdowns, MoE vs dense scaling, inference scaling, architecture ablations, the data-quality cliff, the long-rollout-stability axis, and what the curves imply for Aether 1.5 and beyond.


Abstract

We characterise the loss-versus-compute relationship for simulation-native pretraining across 64 model configurations spanning 80M to 280B parameters and four orders of magnitude in total compute budget. We fit power-law scaling exponents (overall α = 0.087 ± 0.003), report per-domain exponents that differ meaningfully across CFD, FEA, MD, DFT and RTL, characterise the compute-optimal data-to-parameter allocation (≈28 tokens per parameter vs ≈20 for text), document a meaningful data-quality cliff that occurs earlier than for text-trained baselines, characterise dense versus sparse-mixture-of-experts scaling, fit a separate inference-cost scaling law, and ablate the contributions of architecture components (grid-aware attention, conservation regularisation, tool-use head) to the overall exponent. We separately track long-rollout stability as a quality axis approximately orthogonal to single-step cross-entropy [2]. The implications for foundation-model training strategy are concrete: simulation pretraining returns more on additional compute than text pretraining at current scale, frontier gains will increasingly come from architectural improvements rather than from corpus growth, and the choice between dense and sparse configurations depends sensitively on whether training compute or inference compute dominates the deployment economics.

1. Introduction

Compute–loss curves across 64 simulation-native training runs spanning 80M to 280B parameters. The empirical scaling exponent α = 0.087 — steeper than text-trained baselines on matched architectures. The full per-domain breakdown is in Section 4.

Scaling laws are useful because they let us forecast. If we know how loss depends on compute and data, we can decide rationally how to allocate the next training run. For text-trained foundation models, the relevant scaling laws have been published in detail [4, 5]; they have informed roughly a decade of industry training decisions. For simulation-native pretraining, no analogous published curves exist.

This matters for our roadmap. If simulation-trained models scale at the same rate as text-trained ones, then a great deal of the architectural work — grid-aware attention, conservation regularisation, the tool-use head — is doing comparatively little, and we should focus future compute on raw scale-up. If they scale differently, then the work matters and the curves have to be re-derived from scratch. They scale differently, and the difference is material both for forecasting future model quality and for choosing where to invest the next dollar of research effort.

This paper documents what we found across 64 simulation-native training runs spanning 80M to 280B parameters and four orders of magnitude in compute. We report exponents that materially differ from the text-scaling literature, per-domain breakdowns that reveal which physical disciplines benefit most from additional scale, dense-versus-sparse comparisons that constrain the deployment-shape decision, inference-cost scaling laws that are fittable separately from training-cost scaling laws, and architecture-ablation contributions that argue for continued investment in the simulation-specific architectural components. We also document the cases where the fits break: a data-quality cliff at high token counts, a saturation regime above which additional simulation data on the current curation pipeline ceases to help, and the orthogonality of long-rollout stability from single-step cross-entropy.

The paper is structured as follows. Section 2 describes the experimental grid, the training objective, the corpus mix, and the evaluation methodology. Section 3 reports the overall loss-versus-compute exponent and the compute-optimal data-to-parameter allocation. Section 4 breaks the exponent down per simulation domain. Section 5 compares dense and sparse-MoE configurations. Section 6 characterises long-rollout stability as a separate quality axis. Section 7 fits inference-cost scaling laws. Section 8 ablates the contributions of architectural components. Section 9 documents the data-quality cliff. Section 10 compares our results to the published text-scaling literature. Section 11 discusses what the fits imply for Aether 1.5 and beyond. Section 12 documents the limitations. Section 13 concludes.

2. Experimental setup

We trained 64 simulation-native models ranging from 80M to 280B parameters at fixed total compute budgets per row of the experimental grid. The architecture is the transformer described in the Aether 1.0 paper [1], with grid-aware attention biases, a tool-use head separate from the next-state-prediction head, and the conservation-aware regularisation term in the training objective. The objective combined next-state prediction across computational-fluid-dynamics, finite-element analysis, molecular dynamics, density-functional theory and digital-RTL simulation with auxiliary tasks for grid-aware token mixing and a soft conservation-law penalty.

2.1 Compute grid

The compute axis was held fixed within each row of the grid by adjusting batch size and training-token count to compensate for parameter count. We swept the compute budget logarithmically across four orders of magnitude (from roughly 2×10¹⁸ to 2×10²² accumulated FLOPs). The parameter count was swept across nine distinct configurations: 80M, 250M, 800M, 1.3B (dense and MoE variants), 7B, 13B, 40B, 90B and 280B. Each row of the grid was reproduced at three independent random seeds; we report mean ± standard deviation. Within-row variance was small — typically below 0.4% in absolute loss across seeds — and the reported exponents are robust to seed choice.

2.2 Corpus composition

The pretraining corpus was held approximately constant across the grid: 35% CFD trajectories, 18% finite-element analysis, 15% molecular dynamics, 12% density-functional theory, 10% digital RTL simulation, and 10% miscellaneous (instrument traces, engineering text, failure archives). Conservation-aware regularisation was applied at K = 16 stochastic rollout steps with a coefficient of 0.08 on the conservation-loss term relative to the next-state loss. These ratios were derived from a separate corpus-mix-optimisation experiment that is documented separately in the supplement; for the scaling-law fits in this paper they were held constant to isolate the compute-axis effect.

2.3 Evaluation methodology

Test loss was evaluated on a held-out partition matched in distribution to training. The held-out set was constructed to share underlying systems (geometries, molecules, materials) with the training corpus only when those systems came from a distinct experimental campaign; otherwise systems were entirely disjoint. Long-rollout stability was measured separately on a fixed set of canonical trajectories — the standard impinging-jet benchmark [6], a transient turbomachinery case, a small set of FEP cases, and three RTL test designs. Stability metric is the time-step at which a chosen integrated quantity drifts by more than 4% from a high-fidelity reference solution.

We separately ran six full grid replications at K = 4, 8, 16, 32, 64 conservation-rollout steps to isolate the effect of the regularisation hyperparameter on the fitted exponent. We did not vary architecture between rows of the grid (the architecture ablation is presented separately in Section 8 with a smaller compute footprint per ablation cell).

3. The headline loss-vs-compute curve

0.087
loss-vs-compute exponent · simulation · α
0.076
loss-vs-compute exponent · text baselines [4, 5]
2.1×
less compute at matched loss · tuned corpus mix

Simulation loss decreases as compute^(-0.087 ± 0.003) over the full range we measured. The exponent is meaningfully steeper than the ≈0.076 reported in the open literature for text-trained baselines on comparable architectures [4, 5]. Practically, this means simulation pretraining returns more on additional compute than text pretraining at current scale — doubling compute reduces loss by approximately 6% on simulation against approximately 5.1% on text, all else equal. Compounded across the orders of magnitude in compute that are realistic for frontier training runs, the difference is large.

The fitted curves are clean across the four orders of magnitude we tested. Within-row variance is small (typically <0.4% in absolute loss). Cross-row variance is dominated by the compute axis with a small but consistent contribution from the data-mix axis (we held the mix constant in the headline fit; the per-domain analysis in Section 4 lets us examine the mix's contribution by treating each domain's per-token loss as a separate response variable). There is no sign of the exponent shifting with scale across the range we measured, which is consistent with the text-scaling literature reporting a stable exponent over comparable ranges.

3.1 Compute-optimal data-to-parameter ratio

Following the compute-optimal-allocation framework standard in the text-scaling literature [4], we fit the ratio of training tokens to parameters that minimises test loss at fixed compute. For text-trained models, [4] reports a ratio of roughly 20 tokens per parameter at compute-optimal. For simulation-native pretraining, we find a ratio of approximately 28 tokens per parameter — roughly 1.4× higher.

Our working interpretation is that individual simulation steps carry less surprise per unit than individual tokens of natural prose. A computational-fluid-dynamics trajectory step is largely predictable from its predecessors — physics is continuous and the next time-step lives close to the current one — while natural text contains genuinely arbitrary information (proper nouns, idiosyncratic syntactic choices, specific dates and amounts) that the model must memorise. A higher token-per-parameter ratio is consistent with the model not needing as much representational capacity per token of input. We have not directly measured the per-token information content of the two modalities; the hypothesis is plausible but not proven from the scaling-law fits alone.

Practically, this means simulation-native training runs should be more data-heavy and less parameter-heavy at the same compute budget than text-trained equivalents. We have followed this recommendation in production — the 40B dense and 280B sparse-MoE configurations both sit close to the compute-optimal data-to-parameter ratio for their respective compute budgets.

4. Per-domain scaling

The headline exponent averages across five simulation domains. Different domains have different per-token surprise structure; we expected — and found — that their individual exponents differ. The per-domain fits below were obtained by treating each domain's per-token loss as a separate response variable on the same compute axis, with the parameter and data axes held at their per-domain compute-optimal allocations.

  • Computational fluid dynamics — α_CFD = 0.094 ± 0.004. The steepest exponent. CFD trajectories have rich per-step structure (vortex dynamics, shock interactions, boundary-layer transitions) that the model continues to learn at high compute.
  • Finite-element analysis (structural) — α_FEA = 0.085 ± 0.004. Close to the overall average. Structural problems vary in non-linear regime; the model learns both the linear-elastic backbone and the plastic-deformation extensions with consistent return on compute.
  • Molecular dynamics — α_MD = 0.082 ± 0.003. Slightly below average. MD trajectories at the time-step granularity we use are heavily oscillatory; some of the per-step variance is essentially noise.
  • Density-functional theory — α_DFT = 0.071 ± 0.005. Lower than average. DFT outputs are smoother than MD trajectories — the relevant per-step information lives in the converged-state shape rather than in step-by-step dynamics — and the model saturates earlier on DFT-like data.
  • Digital RTL simulation — α_RTL = 0.058 ± 0.006. The lowest. Digital state evolution is discrete and largely deterministic given the input; once the model has learned the underlying logic primitives the marginal return on compute is small.

Two implications follow from the spread. First, the average exponent is a meaningful summary statistic but it hides a factor-of-1.6× spread across domains; corpus-mix choices materially shift the headline curve. Second, the architecture's value-per-parameter is highest where the per-step structure is richest, which is consistent with the architecture being designed for continuous-dynamics modalities and being relatively under-utilised on discrete-dynamics ones.

5. Dense versus sparse mixture-of-experts

Sparse mixture-of-experts (MoE) architectures decouple parameter count from per-token compute cost: a 280B-parameter MoE with 40B active parameters per token has the inference cost of a 40B dense model but the representational capacity of something closer to 280B. The scaling-law question is whether the MoE configuration recovers the loss curve of the dense model with matched active parameters, or whether it captures a meaningful fraction of the additional capacity.

We fit separate scaling-law exponents for the dense and sparse-MoE configurations across the compute grid.

α_dense = 0.087
loss-vs-training-compute · dense configurations
α_MoE = 0.103
loss-vs-training-compute · sparse-MoE configurations
~30%
training FLOPs at matched loss · MoE vs dense

Sparse-MoE has a steeper loss-vs-training-compute exponent. At matched loss, the MoE configuration requires approximately 30% of the training FLOPs of the dense equivalent. This is the well-documented training-compute advantage of MoE architectures and it is consistent with the text-scaling literature.

The picture changes at inference. At matched active-parameter inference cost, the dense model wins on long-rollout stability (Section 6), particularly for the hardest discipline (RTL, where the MoE expert-routing decision is sometimes wrong in ways that compound across many time-steps). The deployment economics therefore depend on which side of the trade dominates: if training compute is the binding constraint and inference cost is amortisable, MoE is the right call; if inference cost is the binding constraint and training compute is non-recurring, the dense configuration is competitive even though its training-compute exponent is slightly worse.

We ship both configurations precisely because the right choice depends on the customer's deployment shape. The 40B dense is the most common production configuration; the 280B sparse-MoE is the right call for batched inference workloads where the per-token cost is amortised across many concurrent requests.

6. Long-rollout stability as a separate axis

Cross-entropy loss is not the only quality that matters for a simulation-native model. Long-rollout stability — how far a trajectory can be unrolled before drift in integrated quantities exceeds a threshold — is the property that matters most for actual downstream use. It is approximately orthogonal to cross-entropy in our regime: a model with lower single-step loss can have worse long-horizon stability if its errors compound differently [2].

We separately measured rollout stability on a fixed evaluation set across the same 64 model configurations. The correlation between single-step loss and rollout-stability metric is positive but weak (Pearson ≈0.42 across the grid). Some configurations with better single-step loss have worse rollout stability and vice versa.

6.1 Conservation regularisation effect

The conservation-law regularisation term in the training objective improves rollout stability substantially without much affecting single-step loss. Ablations are summarised in Section 8; the headline is that conservation regularisation contributes roughly α-equivalent 0.011 to long-rollout stability scaling while contributing approximately α-equivalent 0.002 to single-step cross-entropy. The two axes have different sensitivities to the same architectural intervention, which is itself evidence that they should be tracked separately.

We publish both axes for every release and we track them independently in the experimental grid. Some configurations that look attractive on cross-entropy fail on stability; we flag those configurations and do not promote them to production even when their compute-loss exponent is favourable.

7. Inference-cost scaling

Training-compute scaling laws answer "how good is the model after we spend X FLOPs of training compute." Inference-compute scaling laws answer the related but distinct question: "how much does each production query cost as the model improves." The two are linked but not identical, and for deployment economics the inference-cost curve is often the more important one.

We separately fit an inference-cost scaling law across the dense configurations: per-query inference FLOPs as a function of model quality at deployment time. The fit is a power law in the same form as the training-cost law but with a different exponent.

0.062
loss-vs-inference-compute exponent · dense · simulation
1.4×
more inference compute per loss decade · vs training-compute scaling

Inference loss decreases as inference-compute^(-0.062), shallower than the training-compute exponent of 0.087. Practically, this means improving inference-time quality is expensive per unit of inference compute — a quality decade requires roughly 1.4× more inference compute than a training-compute decade would suggest. This is consistent with intuition: the model's quality at any given inference-compute budget reflects the training that already happened plus the inference budget at deployment; the latter cannot fully substitute for the former.

For customers running interactive workloads — chemistry design rounds, design-exploration iterations — the inference-cost exponent is the relevant one. For customers running long batch-mode campaigns where the model produces many trajectories from a single API call, the training-cost exponent dominates the per-call economics. We expose both metrics in the platform's customer console.

8. Architecture ablations

We ablated three architectural components: grid-aware attention biases, the conservation-aware regularisation term, and the tool-use head. Each ablation re-ran a smaller compute grid (24 configurations across two orders of magnitude in compute) with the component disabled, keeping everything else fixed. The contribution of each component is reported as an α-equivalent — the change in the fitted exponent due to enabling the component.

+0.011
α-equivalent · conservation regularisation
+0.007
α-equivalent · grid-aware attention
+0.004
α-equivalent · tool-use head

Conservation regularisation contributes the largest α-equivalent gain. Without it, the headline exponent drops from 0.087 to 0.076 — essentially to the text-scaling baseline. This is consistent with the cross-domain-transfer paper [3]: the conservation regularisation appears to be the dominant mechanism that makes simulation-native pretraining return better-than-text scaling exponents.

Grid-aware attention contributes the next-largest gain. Without it, the model still benefits from simulation pretraining but the per-domain breakdown is more even (the spread across CFD/FEA/MD/DFT/RTL exponents shrinks by roughly 40%). The grid-aware attention helps disproportionately on disciplines with explicit spatial structure (CFD, FEA, RTL) and contributes little to MD and DFT, where the spatial-structure prior is less useful.

The tool-use head contributes a smaller but consistent α-equivalent gain. We had expected the tool-use head to affect downstream task performance more than pretraining-loss scaling; the small contribution to pretraining scaling suggests the head also helps the model represent solver-call patterns during pretraining itself.

9. The data-quality cliff

Scaling-law fits assume more data continues to help. Empirically, this assumption breaks at large enough scale: at some point the marginal token of curated data does not measurably reduce loss. We hit this cliff at approximately 4×10²¹ tokens equivalent for the 40B dense model — additional curated solver-run data ceased to improve loss in our regime.

4×10²¹
tokens · data-quality cliff · 40B dense
~10×
earlier than equivalent text-scaling cliff

This is approximately 10× earlier than the equivalent cliff reported in [4] for text-trained models on natural-prose corpora. Three implications follow.

First, we cannot rely on simply growing the corpus to drive future improvements — we are nearing the cliff for our current curation pipeline. Second, beyond the cliff, model improvements come from architectural changes (better grid-aware attention, better tool-use heads, conservation-law refinements) and from new data sources (rare-event physics, out-of-distribution chemistries) rather than from corpus volume. Third, and most importantly for procurement, the cliff implies that the marginal cost of useful pretraining tokens is rising. We have changed our data-procurement strategy in response — prioritising rare and out-of-distribution data over additional in-distribution samples.

9.1 Why the cliff is earlier

Our hypothesis is that the cliff is earlier because the in-distribution simulation corpus has lower intrinsic dimensionality than the natural-language corpus does. A text corpus of comparable token count covers an enormous range of topics, registers, languages, factual content; a simulation corpus of comparable token count covers a smaller range of physical regimes once the per-step variance is averaged. The model saturates earlier on the smaller intrinsic dimensionality.

If the hypothesis is correct, the cliff can be pushed back by expanding the corpus's coverage of physical regimes rather than by adding more tokens within already-covered regimes. We have designed the 1.5 curation pipeline around this hypothesis; we will report on whether it pushes the cliff back in the post-1.5 follow-up.

10. Comparison to the published text-scaling literature

Our headline exponent (α = 0.087) is steeper than the ≈0.076 typically reported for text-trained baselines on matched architectures [4, 5]. Our compute-optimal data-to-parameter ratio (≈28) is higher than the ≈20 typically reported for text [4]. Our data-quality cliff is earlier than the equivalent for text-trained corpora.

10.1 What's similar

The functional form is similar. Both simulation and text losses fit a power law in compute over the ranges that have been measured. The exponent is stable across the measured range in both cases; we see no scale-dependent shift in our grid up to 2×10²² FLOPs. Within-row variance across seeds is small in both cases. The fitted curves are clean in both cases, which is encouraging — it suggests that the scaling-law methodology developed for text [4, 5] transfers to simulation despite the modality differences.

10.2 What's different

Three differences worth highlighting.

  • Steeper exponent (0.087 vs 0.076). Simulation pretraining returns more on compute. Attributable largely to the conservation regularisation contribution (Section 8).
  • Higher data-to-parameter optimum (≈28 vs ≈20). Per-token information content is lower for simulation than for text; the optimum shifts toward more data per parameter.
  • Earlier data-quality cliff. Lower intrinsic dimensionality of the simulation corpus saturates the model earlier; new data sources matter more than corpus volume.

11. What this implies for Aether 1.5 and beyond

Three concrete implications follow for Aether 1.5 and beyond.

  • The 280B sparse-MoE configuration sits well inside the regime where additional compute helps; we are not yet at any plateau. Further compute investment at frontier scale is expected to return improvements consistent with the published exponent.
  • Further improvements at the frontier will require new data sources — particularly rare-event and out-of-distribution physics — rather than simply scaling the existing corpus. We have allocated a quarter of the 1.5 training-data budget specifically to rare-event chemistry, hypersonic transition, and sub-3nm semiconductor PDK coverage.
  • Architecture-driven gains are now a larger fraction of incremental quality than at 1.0. We are investing more in grid-aware attention extensions, tool-use head improvements, and the conservation-law term in particular. The α-equivalent contributions in Section 8 argue strongly that these are not vestigial design choices.

11.1 A forecast for Aether 2.0

Linearly extrapolating the headline exponent suggests that a 10× compute increase from the Aether 1.0 training budget would produce a ≈19% reduction in pretraining loss. The corresponding improvement in downstream-task quality is harder to predict because it depends on which axes the customer cares about (single-step accuracy, long-rollout stability, cross-domain transfer). The cross-domain-transfer paper [3] suggests that broader corpora help downstream tasks disproportionately; we expect the next-generation model to widen its discipline coverage as a consequence.

This is a forecast, not a commitment. The cliff in Section 9 implies the linear extrapolation is generous; the architecture-ablation contributions in Section 8 imply that scaling-related gains are increasingly architectural rather than compute-driven. The real 2.0 release will be informed by the next round of measurements; we will publish updated exponents post-1.5.

12. Limitations

Five meaningful limitations.

  • Compute range bounded. Our scaling-law fits are bounded by the compute we had available (up to 2×10²² FLOPs). We cannot rule out an exponent shift at scales beyond 10²³ FLOPs that we have not tested.
  • Corpus mix held constant. The headline fits hold the corpus mix at our 35/18/15/12/10/10 ratio. Varying the mix may yield different exponents that we have characterised partially in Section 4 but not exhaustively.
  • Architecture held constant within the headline grid. Architecture ablations were run on a smaller compute footprint in Section 8; the interaction between architecture choices and the compute exponent at frontier scale is not fully characterised.
  • Stability axis is one specific metric. Our long-rollout stability metric — drift below 4% on integrated quantities — is one operational definition of stability among several plausible ones. Different metrics may show different correlations with cross-entropy.
  • MoE inference cost is hardware-dependent. Our active-parameter framing assumes a particular inference-hardware configuration; on different hardware (e.g., with different memory-bandwidth ratios) the dense-vs-MoE trade may shift.

Reproducibility: the experimental grid required substantial compute. Academic groups attempting to reproduce should expect each row of the grid to cost in the low six figures of compute (across the three seeds). The fitted curves, raw measurements, training-run metadata, and the conservation-regularisation hyperparameter schedule are all open under the Research tier; reproduction guidance is in the supplement. We are willing to provide pretrained checkpoints to bona-fide academic groups under a non-commercial licence; write to research@apexworldlabs.com.

13. Conclusion

Simulation-native pretraining scales with a meaningfully steeper compute exponent than text-trained baselines on comparable architectures (0.087 vs ≈0.076). The compute-optimal data-to-parameter ratio is approximately 1.4× higher than for text. Per-domain exponents differ substantially across CFD/FEA/MD/DFT/RTL, with continuous-dynamics modalities scaling best and discrete modalities scaling worst. Sparse-MoE configurations match dense loss at ~30% of training FLOPs but lose to dense at iso-inference-cost on long rollouts, so the right deployment shape depends on whether training or inference is the binding constraint. Inference-cost scaling is shallower than training-cost scaling, fittable separately, and relevant for customer economics. Architecture ablations show that conservation regularisation contributes the bulk of the simulation-vs-text exponent gap. A data-quality cliff at ≈4×10²¹ tokens implies that frontier gains beyond it will come from architecture and from new data sources rather than from corpus volume. Long-rollout stability remains a separate quality axis from single-step cross-entropy and is tracked independently. The methodology developed for text-scaling [4, 5] transfers cleanly to simulation; the numbers it produces are meaningfully different.

References

  1. [1]Apex Research. Aether 1.0 — a simulation-native foundation model. Apex Research Paper (2026).
  2. [2]Apex CFD Group. Long-rollout stability in autoregressive physics. Apex Research Paper (2026).
  3. [3]Apex Research. Cross-domain transfer in simulation-native foundation models. Apex Research Paper (2026).
  4. [4]Hoffmann, J. et al. Training compute-optimal large language models (Chinchilla). DeepMind technical report (2022).
  5. [5]Kaplan, J. et al. Scaling laws for neural language models. OpenAI technical report (2020).
  6. [6]Cooper, P. et al. Experimental measurements of impinging jet flow and heat transfer. Int. J. Heat Mass Transfer (1993).
  7. [7]Apex Research. Compute-optimal corpus mix for simulation-native pretraining. Apex Technical Note (2025).

Want the underlying model?

Aether powers every workload above. Request access and we'll show you what one AI does with your data.