One model family. Six configurations.
All checkpoints share the same training corpus and the same evaluation suite. The family ranges from a 1.3B embedded variant to a 280B sparse-MoE frontier, plus specialised checkpoints with additional pretraining for biology. The decision between them is about cost and deployment shape, not about which checkpoint is the smart one.
Aether 7B
GADenseThe smallest commercial Aether configuration. Runs comfortably on a single H100. Best for embedded engineering tools, air-gapped inference and front-of-the-line interactive workloads.
- Params
- 7B
- Active
- 7B
- Context
- 1M
- Best for
- Edge, on-device, sovereign light deploys
- Hardware
- 1× H100 or 2× MI300X
- Throughput
- ≈ 380 tok/s · 24 traj/s
Why pick this one- Smallest stable footprint with the full discipline coverage.
- Loads on a single H100 with 1M context active.
- Cheapest unit economics for high-volume interactive workloads.
- Sovereign installs typically start here.
Aether 40B
GADenseThe default for working teams. Strong across every discipline; runs on a single 8×H100 node. The variant most of our customers run in production.
- Params
- 40B
- Active
- 40B
- Context
- 1M
- Best for
- Most production workloads
- Hardware
- 8× H100 / 8× H200 / 8× MI300X
- Throughput
- ≈ 120 tok/s · 8 traj/s
Why pick this one- Default production checkpoint — most pilots land here.
- Strong long-rollout stability across CFD, FEA, MD and EDA.
- Fits a single DGX-class node end-to-end.
- Quarterly model refreshes available on customer cadence.
Aether 280B sparse-MoE
GASparse-MoEThe frontier configuration. Used when accuracy on the hardest workloads is worth the additional inference cost. Active-parameter count matches the 40B dense variant; experts are routed per-token.
- Params
- 280B
- Active
- 40B
- Context
- 1M
- Best for
- Frontier accuracy on the hardest workloads
- Hardware
- 32× H100 / 16× H200 / cluster
- Throughput
- ≈ 80 tok/s · 5 traj/s
Why pick this one- Beats the dense 40B on long-rollout stability by a measurable margin.
- Same active-parameter inference cost as 40B dense.
- Required for the hardest cross-domain coupled cases.
- Cluster inference; provisioned per workload class.
Aether Bio 12B
GASpecialisedSpecialised checkpoint with additional pretraining on biology and chemistry trajectories. Pairs with Aether for autonomous labs. Strong on ADMET, FEP and generative chemistry.
- Params
- 12B
- Active
- 12B
- Context
- 512k
- Best for
- Wet-lab campaigns, ADMET, generative chemistry
- Hardware
- 2× H100 or 4× MI300X
- Throughput
- ≈ 220 tok/s · 16 traj/s
Why pick this one- ADMET accuracy comparable to specialised endpoint models.
- Drives the autonomous wet-lab loop in production.
- Smaller footprint than the 40B dense for biology-only deployments.
- Built-in biosecurity refusals extended for chemistry context.
Aether Edge 1.3B
PreviewSpecialisedQuantised, distilled checkpoint for embedded deployment. Targets real-time control loops in industrial and aerospace environments where milliseconds matter.
- Params
- 1.3B
- Active
- 1.3B
- Context
- 128k
- Best for
- Embedded sensors, real-time control
- Hardware
- Single A100, RTX 6000, or Jetson
- Throughput
- ≈ 1,800 tok/s · 60 traj/s
Why pick this one- Designed for edge and embedded environments.
- Quantised to INT8 with negligible accuracy loss on supported tasks.
- Real-time control loops on industrial PCs.
- Currently in preview with three pilot customers.
Aether 1.5
RoadmapDenseSlated for late 2026. Focus areas: novel-chemistry generalisation, rare-event reliability, longer-rollout stability for transient CFD, and broader semiconductor PDK coverage.
- Params
- TBD
- Active
- TBD
- Context
- TBD
- Best for
- Next major model release
- Hardware
- TBD
- Throughput
- TBD
Why pick this one- Targets the failure modes we publish for 1.0.
- Broader semiconductor PDK coverage and 3DIC focus.
- Improved rollout stability on transient external aerodynamics.
- Roadmap, not a commitment to date.
The same corpus. The same evals.
Every model in the family is evaluated against the same held-out experimental endpoints — wind-tunnel data, plate-reader assays, taped-out silicon. The numbers move together, so the comparisons mean something.
The family is a family — on purpose.
Variants differ in shape and cost. They do not differ in surface, in training corpus, or in the evaluation suite. That is what makes a 7B-to-280B comparison fair.
- Same corpus across every variant in the family.
- Same evaluation suite — held-out experimental endpoints.
- Same red-team / refusal corpus, verified per release.
- Same API surface — variants are a deploy-time choice.
- Same SDKs, same telemetry, same audit format.
- Same model card on /simos for every variant.
Pick the variant. We'll deploy it.
Most teams start on the 40B dense and graduate to the 280B sparse-MoE when the workload demands it. Specialised checkpoints can run side-by-side.