Releases, in public.
Every Aether release — model checkpoints, discipline updates, safety refreshes, platform features — published with numbers and a ship date. If a release breaks something you care about, we want to know before the next one goes out.
Eight tracks, one log.
Each release is tagged by the track it belongs to. Model checkpoints, the discipline platforms, safety refreshes, and the platform runtime each ship on their own cadence — all in one chronological log.
- Discovery1
- Engineering2
- Lab1
- Model3
- Platform3
- Safety2
- Silicon1
- Software1
- Model
Aether 1.0.3 — improved long-rollout stability for transient CFD
Fixes a class of drift artefacts on rollouts longer than 4,000 time-steps for high-Reynolds external aerodynamics.
We identified a residual drift pattern in the 40B dense checkpoint on rollouts longer than 4,000 time-steps for high-Reynolds external aerodynamics. Re-validated on the Cooper impinging-jet benchmark; mean Nusselt error improves from 4.1% to 3.6%. The 280B sparse-MoE checkpoint inherited the fix with a smaller relative improvement (already stable). No API change.
- Cooper benchmark mean Nusselt error: 4.1% → 3.6% on 40B dense.
- Long-rollout horizon extended from ~4,000 to ~5,200 stable steps.
- No API or eval-suite breaking change.
- Model
Aether 1.0 — general availability
First general-availability release. Three sizes, 1M-token context, streaming rollouts, full model card published.
Aether 1.0 is the first version of our foundation model trained primarily on simulation trajectories and instrument data rather than scraped text. Three sizes — 7B dense, 40B dense, 280B sparse-MoE — trained on a shared corpus. 1M-token context, streaming trajectories, capability gates at the model layer. Held-out experimental endpoints documented in the model card on /simos.
- Three sizes shipped: 7B dense, 40B dense, 280B sparse-MoE.
- 1M-token context including grids, RTL hierarchies and lab notebooks.
- Streaming, steerable rollouts for long-horizon physics.
- Built-in biosecurity refusals verified against a versioned red-team corpus.
- Software
Aether for software — two new adapter surfaces, total 15
Adapter total: 15. Memory and review history now shared across all surfaces.
Adds two new adapter surfaces — an AI-native code editor and a hosted long-horizon agent — bringing the total to 15. Memory and review history are shared across surfaces, so the same agent that paired with you in one IDE this morning has the same context in another this afternoon. No migration needed.
- Lab
Biofactory — upgraded acoustic-dispenser and industrial liquid-handler drivers
First-class drivers. Plate logistics planner now respects reagent-lot lineage when allocating across instruments.
Upgraded first-class drivers for the latest-generation acoustic dispensers and industrial liquid handlers ship with the runtime. The plate-logistics planner now respects reagent-lot lineage when allocating across instruments — preventing the cross-contamination class of failures that historically required engineer review. eBR capture remains 21 CFR Part 11-grade.
- Platform
Agent runtime — blue/green deploys with eval-gated promotion
Agent versions now promote behind a percentage rollout, gated by an eval suite.
The long-horizon agent runtime now ships blue/green deploys. Promotion is gated by an eval suite that runs on every push; an agent version reaches 100% only after the eval threshold holds for the configured window. Roll-back is one click. This was the single most-requested production feature.
- Engineering
Apex engineering discipline — unlimited workload milestone
The engineering platform now ships 300 specialised agent domains across ten physics disciplines.
The engineering discipline of Aether now ships 300 specialised agent domains across ten physics disciplines (CFD, FEA, EM, multiphysics, acoustics, thermal, fatigue, additive, optimisation, numerical). The project canvas adds branching and inline review; mesh and result extractors are observable mid-run.
- Safety
Trust & safety — biosecurity refusal corpus refresh
Refresh of the biosecurity refusal corpus. Coverage broadens to additional select-agent classes.
Refresh of the biosecurity refusal corpus on a quarterly cadence. Coverage broadens to additional select-agent classes; precision on benign chemistry questions improves by 9 points. Refusal corpus is now versioned alongside the model — every release runs against the same set, plus the new additions.
- Silicon
Aether for semiconductors — commercial-EDA parity on signoff benchmarks
Chip-design agents close the gap to commercial EDA on timing and power signoff across the published benchmark set.
Aether for semiconductors' chip-design agents close the gap to commercial EDA on timing and power signoff across the published benchmark set. Numbers (≤0.4% WNS gap on the public open-cell-library benchmark, 1.06× perf at iso-area on the public 7nm predictive PDK, ≤2.0% current-density delta on EM/IR) and reproduction scripts are in the research log.
- Safety
Biosecurity boundaries published as research
How Aether for autonomous labs enforces dual-use refusals, controlled-pathogen gating and chain-of-custody — at the model layer.
Published a research post describing the four layers of the autonomous-lab safety architecture: model-layer refusals, capability gating, chain-of-custody, and red-team verification. Linked from /research/biosecurity-boundaries.
- Engineering
Cooper impinging-jet benchmark — beating the commercial CFD baseline
On Cooper et al. impinging-jet cases, Aether hits experimental Nusselt within 4.1% — a measurable improvement over the commercial RANS baseline.
We ran the Cooper et al. impinging-jet cases on Aether's RANS and LES agents at matched mesh and time-step budgets versus a leading commercial RANS baseline. Headline: 4.1% mean Nusselt error vs 7.8% on the baseline, across radial stations. The benchmark setup, mesh, and result-extraction scripts are open.
- Platform
Agent runtime — eval-as-a-first-class-object
Evaluation now lives in the runtime, not on the side. Every agent ships with a regression suite that runs on every commit.
Evaluation in the runtime is now a first-class object. Each agent ships with a regression suite that runs on every commit, plus replay of historical traces to verify the agent still handles them the same way. Synthetic adversarial cases run against every release. Trace search is OTel-native.
- Discovery
Chimera discipline — biologics design agents
Antibody and binder design, paratope prediction, developability scoring, immunogenicity flags.
The drug-discovery discipline gains biologics agents: antibody and binder design, paratope prediction, developability (aggregation, viscosity, glycosylation, deamidation), immunogenicity flags, ADC linker chemistry. Pairs with the autonomous-lab discipline for bench validation.
- Model
Aether — 1M-token context unlocked across all sizes
Context window extended to 1M tokens, including grids, RTL hierarchies and lab notebooks.
Context extended to 1M tokens across all model sizes. The window admits a full transient CFD case, a chip RTL hierarchy, or a year of lab notebooks in a single call. Streaming rollouts respect the same window — long horizons no longer require sliding-window approximation.
- Platform
API v1 — committed surface, semver
The v1 surface is committed. Breaking changes go through a six-month deprecation window and a v2 alias.
API v1 is committed. We have a six-month deprecation window for any breaking change, and v2 aliases ship alongside v1 during the window. Customers building production integrations can rely on the surface; we publish a public deprecation log when anything starts the clock.
Continuously, and in public.
Three principles we hold to. Small steps, public numbers, loud breaking changes. Every release entry is on the same page as the ones that came before and the ones that will come next.
Small steps, in public
We ship continuously and we publish each step. No quiet rollouts, no silent deprecations, no hidden version bumps. If something changed, it appears below.
Numbers, not adjectives
Every release entry carries reproducible numbers — the case, the seed, the benchmark, the delta. Adjectives without numbers don't make it onto this page.
Breaking changes are loud
API v1 is committed: six-month deprecation windows and v2 aliases ship alongside v1. Breaking changes get their own entry, the migration steps, and a published deprecation log.
See something that doesn't reproduce?
If a published number in any release doesn't replicate on the documented setup, it's a P1 bug. Send a reproduction report to research@apexworldlabs.com — we acknowledge within a business day.