Benchmarks, in public.
Every workload Aether ships against measured reality — wind-tunnel data, FEP+ ground truth, taped-out silicon, plate-reader assays, public SWE-bench. We post the numbers, name the comparison and link the reproduction kit. If a number on this page moves, this page moves with it.
How we evaluate — and what we refuse to do.
Public benchmarks are easy to overfit. We hold ourselves to a tighter bar because measured reality is unforgiving.
Held-out experimental endpoints, not held-out text
Every engineering benchmark is scored against measured reality — wind-tunnel data, fatigue logs, plate-reader assays, taped-out silicon. No leaderboard farming.
Same corpus, same evals across the family
Every Aether variant — Edge 1.3B through 280B-MoE — is evaluated on this exact suite. The numbers move together because the training corpus and eval set are shared.
Reproduction kits open
Each benchmark links to a reproduction kit on GitHub. Run it on our managed cloud, your VPC, or on-prem. The numbers should reproduce within published bounds.
| Task | Family | Metric | Aether | Best other | Δ | Repro |
|---|---|---|---|---|---|---|
| Cooper impinging-jet · Nu | CFD · heat transfer | Mean error vs experiment | 3.6% | 5.4% Commercial CFD solver | −33% | repro |
| Wing-body cruise · drag polar | CFD · transonic | ΔCd vs wind-tunnel | ≤1.0% | 1.6% RANS k-ω SST baseline | −38% | repro |
| Notched-bar fatigue | FEA · fatigue | Pearson r on cycle life | 0.92 | 0.81 Commercial fatigue solver | +0.11 | repro |
| Hypersonic re-entry · TPS | CFD · real-gas | Stagnation flux error vs CUBRC | 4.2% | 6.1% Equilibrium-chemistry baseline | −31% | repro |
| Task | Family | Metric | Aether | Best other | Δ | Repro |
|---|---|---|---|---|---|---|
| FEP+ 8-target panel | Alchemical FEP | RMSE on ΔΔG (kcal/mol) | 1.06 | 1.23 Commercial FEP package | −14% | repro |
| PDBbind core set · v2020 | Affinity prediction | Pearson r | 0.78 | 0.71 Best published GNN | +0.07 | repro |
| hERG · external curated | ADMET classification | ROC-AUC | 0.91 | 0.87 Best literature model | +0.04 | repro |
| Antibody developability | Biologics | ROC-AUC on aggregation | 0.83 | 0.79 Best literature panel | +0.04 | repro |
Public benchmarks we refuse to publish.
Some benchmarks have known leakage between training and evaluation sets, or measure the wrong thing. We don't publish on them — and we'll explain why, if you ask.
MMLU-style trivia panels for science
Multiple-choice trivia doesn't measure forward-rolling a system in time. A model that scores high on chemistry MMLU may still fail FEP.
Leaderboard suites with shared test sets
Where the eval set is on the public internet, frontier models trained on the internet have seen it. We score on held-out experimental endpoints.
Self-graded reasoning panels
If the grader is also a language model and you trained both, the score is a hallucination of competence. We use measured outcomes only.
Benchmarks we paid the authors to design
If we paid for it, it's an internal eval — and we use it for development, not for marketing.
Reproduce the numbers.
The reproduction kits run on the same managed cloud Aether ships to. If a number on this page won't reproduce within published bounds in your environment, we want to know.