Batch & HPC.
Massively parallel batch and simulation jobs with scheduling, checkpointing and fault tolerance built in.
Submit a job, not a cluster. Batch & HPC schedules massively parallel work across spot and reserved capacity with checkpointing and retries, so a thousand-task sweep survives a preemption.
- Category
- Compute
- Deployment
- Managed → air-gapped
- Governance
- IAM · encryption · audit
Three steps to running.
Submit array jobs with dependencies and priorities; no cluster to stand up.
Tasks spread across spot and reserved capacity, MPI-tight where needed.
Checkpointing and retries mean a preemption costs minutes, not the run.
Batch & HPC, in full.
Array jobs, dependencies and priorities across a managed, elastic pool.
Long runs survive preemption and node failure without losing progress.
Mix spot and on-demand for cost, with automatic fallback when capacity moves.
Tight-coupled HPC with low-latency interconnect for simulation and solvers.
Provision it in a few lines.
Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.
# Provision batch & hpc on Aether Cloud
aether compute create \
--service batch-hpc \
--name app \
--region us-1 \
--deploy managed # or vpc | on-prem | air-gappedAt a glance.
- Model
- Array jobs + dependencies
- Coupling
- Embarrassingly parallel + MPI
- Capacity
- Spot + reserved, autoscaled
- Resilience
- Checkpoint, resume, retry
- Scale
- Thousands of concurrent tasks
Built for real work.
Parameter sweeps and DOE
Simulation and rollout campaigns
Rendering and scientific computing
On one model, not stitched together.
The usual stack runs batch & hpc in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.
No stitching a vector DB to one place, a warehouse to another and a model to a third — batch & hpc sits next to the rest of the catalog, one identity, one bill.
The provider that runs Aether runs your batch & hpc — so the data and the model never leave the same governed boundary to talk to each other.
Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.
Good to know.
Same managed batch scheduling — with the simulation engine that trains Aether available for your forward-rolled jobs.
Yes — low-latency interconnect and MPI for solvers and simulation, not just embarrassingly-parallel work.
Checkpointing and automatic fallback keep long runs going through preemption.
Your boundary, your choice.
Pairs well with.
On-demand and reserved VMs across CPU and GPU shapes, with per-second billing and live resize.
Latest-generation accelerators by the hour or the cluster, with fast interconnect for training and high-throughput inference.
Event-driven functions that scale to zero — run code and agents without managing a server.
Managed container runtime with autoscaling, health checks and rolling deploys — bring an image, get a URL.
A conformant, managed control plane with GPU scheduling, autoscaling node pools and zero-downtime upgrades.
Dedicated single-tenant hardware for the workloads that need full control of the silicon.
Run Batch & HPC on Aether Cloud.
Massively parallel batch and simulation jobs with scheduling, checkpointing and fault tolerance built in. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.