Skip to content
Apex
Aether Cloud / Compute / GPU instances
Aether Cloud · Compute

GPU instances.

Latest-generation accelerators by the hour or the cluster, with fast interconnect for training and high-throughput inference.

▦ Compute
Overview

Reserve or burst onto the same accelerators that train Aether — single GPUs for a notebook, or fully meshed multi-node clusters for distributed training and high-throughput serving.

Where it sits
Category
Compute
Deployment
Managed → air-gapped
Governance
IAM · encryption · audit

The bottleneck moved from the chip to the fabric and the queue.

Getting a fast GPU is no longer the hard part. Getting many of them meshed tightly enough to scale a training run, available when you need them, and priced so they're not burning money at idle — that's the problem. A single fast accelerator starves on a slow interconnect; a reserved cluster bleeds cost between runs; on-demand capacity evaporates exactly when a launch needs it. Aether's GPU service is the same accelerated compute that trains the model itself, exposed so your jobs get the fabric and the scheduling, not just the silicon.

RDMA-meshed clusters that scale near-linearly.

For distributed training, the interconnect is the architecture. Nodes are meshed with NVLink and RDMA fabric so gradient exchange doesn't become the bottleneck and scaling stays close to linear as you add nodes. Checkpointing and fault tolerance are wired in, so a node failure on a multi-day run costs minutes, not the run. You bring a job, not a cluster-ops team.

Hourly, reserved or spot — and the model is next door.

Burst onto a single GPU for a notebook, reserve a cluster for a programme, or run fault-tolerant jobs on spot with automatic recovery and fallback when capacity moves. GPU utilisation, memory and kernel traces stream into observability out of the box. And because this is the platform that serves Aether, training and inference sit next to your data — no copy out to a separate ML cloud to reach the model.

Worked example

Reserve a meshed multi-node cluster and launch a fault-tolerant training run — checkpointing and recovery handled.

aether gpu cluster create \
  --nodes 8 --gpus-per-node 8 \
  --interconnect rdma \
  --spot --checkpoint s3://runs/job-42

aether train submit \
  --cluster job-42 \
  --resume-on-failure        # node dies -> resume from checkpoint

The cluster comes up meshed and ready; the run checkpoints and resumes through a spot preemption or a node failure on its own. On raw instances, the fabric, the scheduling and the fault tolerance are yours to build before the first epoch.

How it works

Three steps to running.

01
Reserve or burst

Take a single GPU on demand or reserve a meshed multi-node cluster for a training run.

02
Attach data

Mount a high-throughput file system or object storage so the accelerators stay fed.

03
Train or serve

Run distributed training over RDMA, or serve high-throughput inference, with profiling wired in.

What you get

GPU instances, in full.

Latest accelerators

Current-generation GPUs with high-bandwidth memory for large models and long context.

Fast interconnect

NVLink and RDMA fabric so distributed training scales near-linearly across nodes.

Hourly or reserved

Burst for a run, reserve for a programme — with spot capacity for fault-tolerant jobs.

Profiling built in

GPU utilization, memory and kernel traces wired into observability out of the box.

API-first

Provision it in a few lines.

Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.

# Provision gpu instances on Aether Cloud
aether compute create \
  --service gpu-instances \
  --name app \
  --region us-1 \
  --deploy managed   # or vpc | on-prem | air-gapped
Specs

At a glance.

Accelerators
Current-gen GPUs, high-bandwidth memory
Interconnect
NVLink + RDMA fabric
Topologies
Single GPU → multi-node cluster
Pricing
Hourly, reserved or spot
Scaling
Near-linear across nodes
Use cases

Built for real work.

01

Distributed model training

02

High-throughput batch inference

03

Large-scale simulation rollouts

Why one platform

On one model, not stitched together.

The usual stack runs gpu instances in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.

One platform

No stitching a vector DB to one place, a warehouse to another and a model to a third — gpu instances sits next to the rest of the catalog, one identity, one bill.

The model is here

The provider that runs Aether runs your gpu instances — so the data and the model never leave the same governed boundary to talk to each other.

Built on demand

Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.

FAQ

Good to know.

Can I run multi-node training?

Yes — nodes are RDMA-interconnected for near-linear scaling, with checkpointing for fault tolerance.

Is spot capacity available?

Yes, for fault-tolerant jobs, with automatic fallback when capacity moves.

What about profiling?

GPU utilization, memory and kernel traces stream into observability out of the box.

Deploy anywhere

Your boundary, your choice.

Managed
Your VPC
On-prem
Air-gapped / sovereign

Run GPU instances on Aether Cloud.

Latest-generation accelerators by the hour or the cluster, with fast interconnect for training and high-throughput inference. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.