GPU instances.
Latest-generation accelerators by the hour or the cluster, with fast interconnect for training and high-throughput inference.
Reserve or burst onto the same accelerators that train Aether — single GPUs for a notebook, or fully meshed multi-node clusters for distributed training and high-throughput serving.
- Category
- Compute
- Deployment
- Managed → air-gapped
- Governance
- IAM · encryption · audit
The bottleneck moved from the chip to the fabric and the queue.
Getting a fast GPU is no longer the hard part. Getting many of them meshed tightly enough to scale a training run, available when you need them, and priced so they're not burning money at idle — that's the problem. A single fast accelerator starves on a slow interconnect; a reserved cluster bleeds cost between runs; on-demand capacity evaporates exactly when a launch needs it. Aether's GPU service is the same accelerated compute that trains the model itself, exposed so your jobs get the fabric and the scheduling, not just the silicon.
RDMA-meshed clusters that scale near-linearly.
For distributed training, the interconnect is the architecture. Nodes are meshed with NVLink and RDMA fabric so gradient exchange doesn't become the bottleneck and scaling stays close to linear as you add nodes. Checkpointing and fault tolerance are wired in, so a node failure on a multi-day run costs minutes, not the run. You bring a job, not a cluster-ops team.
Hourly, reserved or spot — and the model is next door.
Burst onto a single GPU for a notebook, reserve a cluster for a programme, or run fault-tolerant jobs on spot with automatic recovery and fallback when capacity moves. GPU utilisation, memory and kernel traces stream into observability out of the box. And because this is the platform that serves Aether, training and inference sit next to your data — no copy out to a separate ML cloud to reach the model.
Reserve a meshed multi-node cluster and launch a fault-tolerant training run — checkpointing and recovery handled.
aether gpu cluster create \
--nodes 8 --gpus-per-node 8 \
--interconnect rdma \
--spot --checkpoint s3://runs/job-42
aether train submit \
--cluster job-42 \
--resume-on-failure # node dies -> resume from checkpointThe cluster comes up meshed and ready; the run checkpoints and resumes through a spot preemption or a node failure on its own. On raw instances, the fabric, the scheduling and the fault tolerance are yours to build before the first epoch.
Three steps to running.
Take a single GPU on demand or reserve a meshed multi-node cluster for a training run.
Mount a high-throughput file system or object storage so the accelerators stay fed.
Run distributed training over RDMA, or serve high-throughput inference, with profiling wired in.
GPU instances, in full.
Current-generation GPUs with high-bandwidth memory for large models and long context.
NVLink and RDMA fabric so distributed training scales near-linearly across nodes.
Burst for a run, reserve for a programme — with spot capacity for fault-tolerant jobs.
GPU utilization, memory and kernel traces wired into observability out of the box.
Provision it in a few lines.
Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.
# Provision gpu instances on Aether Cloud
aether compute create \
--service gpu-instances \
--name app \
--region us-1 \
--deploy managed # or vpc | on-prem | air-gappedAt a glance.
- Accelerators
- Current-gen GPUs, high-bandwidth memory
- Interconnect
- NVLink + RDMA fabric
- Topologies
- Single GPU → multi-node cluster
- Pricing
- Hourly, reserved or spot
- Scaling
- Near-linear across nodes
Built for real work.
Distributed model training
High-throughput batch inference
Large-scale simulation rollouts
On one model, not stitched together.
The usual stack runs gpu instances in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.
No stitching a vector DB to one place, a warehouse to another and a model to a third — gpu instances sits next to the rest of the catalog, one identity, one bill.
The provider that runs Aether runs your gpu instances — so the data and the model never leave the same governed boundary to talk to each other.
Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.
Good to know.
Yes — nodes are RDMA-interconnected for near-linear scaling, with checkpointing for fault tolerance.
Yes, for fault-tolerant jobs, with automatic fallback when capacity moves.
GPU utilization, memory and kernel traces stream into observability out of the box.
Your boundary, your choice.
Pairs well with.
On-demand and reserved VMs across CPU and GPU shapes, with per-second billing and live resize.
Event-driven functions that scale to zero — run code and agents without managing a server.
Managed container runtime with autoscaling, health checks and rolling deploys — bring an image, get a URL.
A conformant, managed control plane with GPU scheduling, autoscaling node pools and zero-downtime upgrades.
Massively parallel batch and simulation jobs with scheduling, checkpointing and fault tolerance built in.
Dedicated single-tenant hardware for the workloads that need full control of the silicon.
Run GPU instances on Aether Cloud.
Latest-generation accelerators by the hour or the cluster, with fast interconnect for training and high-throughput inference. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.