Skip to content
Apex
Changelog

Releases, in public.

Every Aether release — model checkpoints, discipline updates, safety refreshes, platform features — published with numbers and a ship date. If a release breaks something you care about, we want to know before the next one goes out.

14
releases published
8
release tracks
v1.0
current model release
1 wk
median cadence
By release track

Eight tracks, one log.

Each release is tagged by the track it belongs to. Model checkpoints, the discipline platforms, safety refreshes, and the platform runtime each ship on their own cadence — all in one chronological log.

  • Discovery1
  • Engineering2
  • Lab1
  • Model3
  • Platform3
  • Safety2
  • Silicon1
  • Software1
Full log
  1. Model

    Aether 1.0.3 — improved long-rollout stability for transient CFD

    Fixes a class of drift artefacts on rollouts longer than 4,000 time-steps for high-Reynolds external aerodynamics.

    We identified a residual drift pattern in the 40B dense checkpoint on rollouts longer than 4,000 time-steps for high-Reynolds external aerodynamics. Re-validated on the Cooper impinging-jet benchmark; mean Nusselt error improves from 4.1% to 3.6%. The 280B sparse-MoE checkpoint inherited the fix with a smaller relative improvement (already stable). No API change.

    • Cooper benchmark mean Nusselt error: 4.1% → 3.6% on 40B dense.
    • Long-rollout horizon extended from ~4,000 to ~5,200 stable steps.
    • No API or eval-suite breaking change.
  2. Model

    Aether 1.0 — general availability

    First general-availability release. Three sizes, 1M-token context, streaming rollouts, full model card published.

    Aether 1.0 is the first version of our foundation model trained primarily on simulation trajectories and instrument data rather than scraped text. Three sizes — 7B dense, 40B dense, 280B sparse-MoE — trained on a shared corpus. 1M-token context, streaming trajectories, capability gates at the model layer. Held-out experimental endpoints documented in the model card on /simos.

    • Three sizes shipped: 7B dense, 40B dense, 280B sparse-MoE.
    • 1M-token context including grids, RTL hierarchies and lab notebooks.
    • Streaming, steerable rollouts for long-horizon physics.
    • Built-in biosecurity refusals verified against a versioned red-team corpus.
  3. Software

    Aether for software — two new adapter surfaces, total 15

    Adapter total: 15. Memory and review history now shared across all surfaces.

    Adds two new adapter surfaces — an AI-native code editor and a hosted long-horizon agent — bringing the total to 15. Memory and review history are shared across surfaces, so the same agent that paired with you in one IDE this morning has the same context in another this afternoon. No migration needed.

  4. Lab

    Biofactory — upgraded acoustic-dispenser and industrial liquid-handler drivers

    First-class drivers. Plate logistics planner now respects reagent-lot lineage when allocating across instruments.

    Upgraded first-class drivers for the latest-generation acoustic dispensers and industrial liquid handlers ship with the runtime. The plate-logistics planner now respects reagent-lot lineage when allocating across instruments — preventing the cross-contamination class of failures that historically required engineer review. eBR capture remains 21 CFR Part 11-grade.

  5. Platform

    Agent runtime — blue/green deploys with eval-gated promotion

    Agent versions now promote behind a percentage rollout, gated by an eval suite.

    The long-horizon agent runtime now ships blue/green deploys. Promotion is gated by an eval suite that runs on every push; an agent version reaches 100% only after the eval threshold holds for the configured window. Roll-back is one click. This was the single most-requested production feature.

  6. Engineering

    Apex engineering discipline — unlimited workload milestone

    The engineering platform now ships 300 specialised agent domains across ten physics disciplines.

    The engineering discipline of Aether now ships 300 specialised agent domains across ten physics disciplines (CFD, FEA, EM, multiphysics, acoustics, thermal, fatigue, additive, optimisation, numerical). The project canvas adds branching and inline review; mesh and result extractors are observable mid-run.

  7. Safety

    Trust & safety — biosecurity refusal corpus refresh

    Refresh of the biosecurity refusal corpus. Coverage broadens to additional select-agent classes.

    Refresh of the biosecurity refusal corpus on a quarterly cadence. Coverage broadens to additional select-agent classes; precision on benign chemistry questions improves by 9 points. Refusal corpus is now versioned alongside the model — every release runs against the same set, plus the new additions.

  8. Silicon

    Aether for semiconductors — commercial-EDA parity on signoff benchmarks

    Chip-design agents close the gap to commercial EDA on timing and power signoff across the published benchmark set.

    Aether for semiconductors' chip-design agents close the gap to commercial EDA on timing and power signoff across the published benchmark set. Numbers (≤0.4% WNS gap on the public open-cell-library benchmark, 1.06× perf at iso-area on the public 7nm predictive PDK, ≤2.0% current-density delta on EM/IR) and reproduction scripts are in the research log.

  9. Safety

    Biosecurity boundaries published as research

    How Aether for autonomous labs enforces dual-use refusals, controlled-pathogen gating and chain-of-custody — at the model layer.

    Published a research post describing the four layers of the autonomous-lab safety architecture: model-layer refusals, capability gating, chain-of-custody, and red-team verification. Linked from /research/biosecurity-boundaries.

  10. Engineering

    Cooper impinging-jet benchmark — beating the commercial CFD baseline

    On Cooper et al. impinging-jet cases, Aether hits experimental Nusselt within 4.1% — a measurable improvement over the commercial RANS baseline.

    We ran the Cooper et al. impinging-jet cases on Aether's RANS and LES agents at matched mesh and time-step budgets versus a leading commercial RANS baseline. Headline: 4.1% mean Nusselt error vs 7.8% on the baseline, across radial stations. The benchmark setup, mesh, and result-extraction scripts are open.

  11. Platform

    Agent runtime — eval-as-a-first-class-object

    Evaluation now lives in the runtime, not on the side. Every agent ships with a regression suite that runs on every commit.

    Evaluation in the runtime is now a first-class object. Each agent ships with a regression suite that runs on every commit, plus replay of historical traces to verify the agent still handles them the same way. Synthetic adversarial cases run against every release. Trace search is OTel-native.

  12. Discovery

    Chimera discipline — biologics design agents

    Antibody and binder design, paratope prediction, developability scoring, immunogenicity flags.

    The drug-discovery discipline gains biologics agents: antibody and binder design, paratope prediction, developability (aggregation, viscosity, glycosylation, deamidation), immunogenicity flags, ADC linker chemistry. Pairs with the autonomous-lab discipline for bench validation.

  13. Model

    Aether — 1M-token context unlocked across all sizes

    Context window extended to 1M tokens, including grids, RTL hierarchies and lab notebooks.

    Context extended to 1M tokens across all model sizes. The window admits a full transient CFD case, a chip RTL hierarchy, or a year of lab notebooks in a single call. Streaming rollouts respect the same window — long horizons no longer require sliding-window approximation.

  14. Platform

    API v1 — committed surface, semver

    The v1 surface is committed. Breaking changes go through a six-month deprecation window and a v2 alias.

    API v1 is committed. We have a six-month deprecation window for any breaking change, and v2 aliases ship alongside v1 during the window. Customers building production integrations can rely on the surface; we publish a public deprecation log when anything starts the clock.

How we release

Continuously, and in public.

Three principles we hold to. Small steps, public numbers, loud breaking changes. Every release entry is on the same page as the ones that came before and the ones that will come next.

  • Small steps, in public

    We ship continuously and we publish each step. No quiet rollouts, no silent deprecations, no hidden version bumps. If something changed, it appears below.

  • Numbers, not adjectives

    Every release entry carries reproducible numbers — the case, the seed, the benchmark, the delta. Adjectives without numbers don't make it onto this page.

  • Breaking changes are loud

    API v1 is committed: six-month deprecation windows and v2 aliases ship alongside v1. Breaking changes get their own entry, the migration steps, and a published deprecation log.

See something that doesn't reproduce?

If a published number in any release doesn't replicate on the documented setup, it's a P1 bug. Send a reproduction report to research@apexworldlabs.com — we acknowledge within a business day.