Skip to content
Apex
Benchmarks · public scoreboard

Benchmarks, in public.

Every workload Aether ships against measured reality — wind-tunnel data, FEP+ ground truth, taped-out silicon, plate-reader assays, public SWE-bench. We post the numbers, name the comparison and link the reproduction kit. If a number on this page moves, this page moves with it.

Principles

How we evaluate — and what we refuse to do.

Public benchmarks are easy to overfit. We hold ourselves to a tighter bar because measured reality is unforgiving.

Held-out experimental endpoints, not held-out text

Every engineering benchmark is scored against measured reality — wind-tunnel data, fatigue logs, plate-reader assays, taped-out silicon. No leaderboard farming.

Same corpus, same evals across the family

Every Aether variant — Edge 1.3B through 280B-MoE — is evaluated on this exact suite. The numbers move together because the training corpus and eval set are shared.

Reproduction kits open

Each benchmark links to a reproduction kit on GitHub. Run it on our managed cloud, your VPC, or on-prem. The numbers should reproduce within published bounds.

Engineering · CFD · FEA · multiphysics
TaskFamilyMetricAetherBest otherΔRepro
Cooper impinging-jet · NuCFD · heat transferMean error vs experiment3.6%5.4%
Commercial CFD solver
−33%repro
Wing-body cruise · drag polarCFD · transonicΔCd vs wind-tunnel≤1.0%1.6%
RANS k-ω SST baseline
−38%repro
Notched-bar fatigueFEA · fatiguePearson r on cycle life0.920.81
Commercial fatigue solver
+0.11repro
Hypersonic re-entry · TPSCFD · real-gasStagnation flux error vs CUBRC4.2%6.1%
Equilibrium-chemistry baseline
−31%repro
Discovery · pharma · biologics
TaskFamilyMetricAetherBest otherΔRepro
FEP+ 8-target panelAlchemical FEPRMSE on ΔΔG (kcal/mol)1.061.23
Commercial FEP package
−14%repro
PDBbind core set · v2020Affinity predictionPearson r0.780.71
Best published GNN
+0.07repro
hERG · external curatedADMET classificationROC-AUC0.910.87
Best literature model
+0.04repro
Antibody developabilityBiologicsROC-AUC on aggregation0.830.79
Best literature panel
+0.04repro
Semiconductors · physical implementation
TaskFamilyMetricAetherBest otherΔRepro
WNS gap vs commercial EDARTL signoff · 7nmWorst-negative-slack delta≤0.4%—
(EDA baseline = parity reference)
parityrepro
ASAP7 PPA closurePhysical implementationPerf at iso-area1.06×1.00×
Commercial EDA baseline
+6%repro
Software · coding agents
TaskFamilyMetricAetherBest otherΔRepro
SWE-bench VerifiedWhole-repo editsResolution rate53.6%47.1%
Top public agent
+6.5 ptsrepro
CommitBench refactorCross-language migrationSemantic equivalence94%88%
Best published migrator
+6 ptsrepro
What's not on this page

Public benchmarks we refuse to publish.

Some benchmarks have known leakage between training and evaluation sets, or measure the wrong thing. We don't publish on them — and we'll explain why, if you ask.

  • MMLU-style trivia panels for science

    Multiple-choice trivia doesn't measure forward-rolling a system in time. A model that scores high on chemistry MMLU may still fail FEP.

  • Leaderboard suites with shared test sets

    Where the eval set is on the public internet, frontier models trained on the internet have seen it. We score on held-out experimental endpoints.

  • Self-graded reasoning panels

    If the grader is also a language model and you trained both, the score is a hallucination of competence. We use measured outcomes only.

  • Benchmarks we paid the authors to design

    If we paid for it, it's an internal eval — and we use it for development, not for marketing.

Reproduce the numbers.

The reproduction kits run on the same managed cloud Aether ships to. If a number on this page won't reproduce within published bounds in your environment, we want to know.