Skip to content
Apex
Aether Cloud / Compute / Batch & HPC
Aether Cloud · Compute

Batch & HPC.

Massively parallel batch and simulation jobs with scheduling, checkpointing and fault tolerance built in.

▦ Compute
Overview

Submit a job, not a cluster. Batch & HPC schedules massively parallel work across spot and reserved capacity with checkpointing and retries, so a thousand-task sweep survives a preemption.

Where it sits
Category
Compute
Deployment
Managed → air-gapped
Governance
IAM · encryption · audit
How it works

Three steps to running.

01
Define the job

Submit array jobs with dependencies and priorities; no cluster to stand up.

02
Run elastically

Tasks spread across spot and reserved capacity, MPI-tight where needed.

03
Survive failure

Checkpointing and retries mean a preemption costs minutes, not the run.

What you get

Batch & HPC, in full.

Job-level scheduling

Array jobs, dependencies and priorities across a managed, elastic pool.

Checkpoint & resume

Long runs survive preemption and node failure without losing progress.

Spot-aware

Mix spot and on-demand for cost, with automatic fallback when capacity moves.

MPI & multi-node

Tight-coupled HPC with low-latency interconnect for simulation and solvers.

API-first

Provision it in a few lines.

Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.

# Provision batch & hpc on Aether Cloud
aether compute create \
  --service batch-hpc \
  --name app \
  --region us-1 \
  --deploy managed   # or vpc | on-prem | air-gapped
Specs

At a glance.

Model
Array jobs + dependencies
Coupling
Embarrassingly parallel + MPI
Capacity
Spot + reserved, autoscaled
Resilience
Checkpoint, resume, retry
Scale
Thousands of concurrent tasks
Use cases

Built for real work.

01

Parameter sweeps and DOE

02

Simulation and rollout campaigns

03

Rendering and scientific computing

Why one platform

On one model, not stitched together.

The usual stack runs batch & hpc in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.

One platform

No stitching a vector DB to one place, a warehouse to another and a model to a third — batch & hpc sits next to the rest of the catalog, one identity, one bill.

The model is here

The provider that runs Aether runs your batch & hpc — so the data and the model never leave the same governed boundary to talk to each other.

Built on demand

Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.

FAQ

Good to know.

How does this compare to AWS Batch or Slurm?

Same managed batch scheduling — with the simulation engine that trains Aether available for your forward-rolled jobs.

Does it handle tightly-coupled HPC?

Yes — low-latency interconnect and MPI for solvers and simulation, not just embarrassingly-parallel work.

What about spot interruptions?

Checkpointing and automatic fallback keep long runs going through preemption.

Deploy anywhere

Your boundary, your choice.

Managed
Your VPC
On-prem
Air-gapped / sovereign

Run Batch & HPC on Aether Cloud.

Massively parallel batch and simulation jobs with scheduling, checkpointing and fault tolerance built in. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.