Skip to content
Apex
Aether Cloud / Databases / Data warehouse
Aether Cloud · Databases

Data warehouse.

Columnar, separation-of-storage-and-compute warehouse for analytics at scale.

▥ Databases
Overview

A columnar cloud warehouse that separates storage from compute, so you scale query power independently and pay for the scans you run — fast analytics over petabytes without tuning.

Where it sits
Category
Databases
Deployment
Managed → air-gapped
Governance
IAM · encryption · audit

The warehouse stopped being the hard part. Everything around it didn't.

A modern analytics estate is rarely one system. There's a warehouse for SQL, an object store and a lake for the raw and semi-structured data, a separate Spark or feature pipeline for ML, a vector database for retrieval, and a model API somewhere else entirely. Each is a good product. The cost is in the seams between them: data gets copied from the lake into the warehouse and again into the feature store, governance is re-implemented three times and drifts, and the moment you want the model to read a governed table you're exporting data across a vendor boundary and hoping the lineage survives the trip. The warehouse query is fast; the organisation around it is slow, and that's where the quarters go.

Separated storage and compute — and one copy of the data.

Aether's warehouse uses the architecture the cloud data warehouse popularised and the cloud warehouse and lakehouse SQL platforms share: storage and compute are decoupled. Your data lives once, in open table formats over object storage; compute clusters are spun up per workload, sized independently, and torn down after. Two consequences matter. First, you scale query power for a heavy dashboard refresh without touching the data or the team running ad-hoc analysis next door — no contention, no over-provisioning a single cluster for the worst case. Second, because the storage is open (Iceberg/Delta-style), the warehouse, Spark, BI and the model all read the same tables in place. There is no proprietary storage format quietly locking you in, and no nightly job copying the lake into the warehouse so the dashboards can see it.

The difference: the model is already in the warehouse.

This is the line that separates Aether from a best-in-class warehouse. On the cloud data warehouse or the cloud warehouse, calling a model means shipping rows to an external inference endpoint, getting predictions back, and writing them somewhere — a pipeline you build, secure and pay egress on, with governance that now spans two vendors. On Aether, the model runs on the same platform, against the same governed tables, under the same identity. A query can classify, extract, embed, summarise or forecast inline; retrieval against the vector store is a join, not an integration; and a fine-tune trains on the warehouse data without it ever leaving the boundary. The data never crosses a wire to meet the model, so there's nothing to secure between them and nothing to reconcile after.

Migrate by reading, not by forklifting.

Because the storage layer is open table formats, you don't have to load everything before you get value. Point the warehouse at the lake and query it in place; bring the hottest tables into managed columnar storage when the latency justifies it. There's no proprietary format you can't get out of, and the same data is queryable from Spark and the model without a second copy. Deployment follows the same rule as the rest of the platform: managed multi-tenant, dedicated, in your VPC, on-prem, or fully air-gapped — the warehouse runs where your data is allowed to live, not only where the vendor offers a region.

Worked example

Here's the capability that doesn't exist on a standalone warehouse: a single SQL statement that filters governed warehouse rows, retrieves the nearest support cases by vector similarity, and asks the model to draft a resolution — no export, no external endpoint, no second copy of the data.

-- Triage open tickets against similar resolved ones, inline.
SELECT t.id,
       t.summary,
       AETHER.GENERATE(
         'Draft a resolution for this ticket using the similar cases',
         t.summary,
         past.resolutions
       ) AS suggested_resolution
FROM   support.open_tickets t
CROSS JOIN LATERAL (
  SELECT ARRAY_AGG(r.resolution) AS resolutions
  FROM   support.resolved_tickets r
  ORDER BY VECTOR_DISTANCE(r.embedding, t.embedding)   -- vector search, as a join
  LIMIT  5
) past
WHERE  t.status = 'open'
  AND  t.priority = 'high';

On the cloud data warehouse or the cloud warehouse this is three systems and a pipeline: the warehouse for the rows, a vector database for the similarity search, and an external model endpoint for the generation — wired together with code you own and data that crosses two boundaries. Here it's one query, one identity, one audit trail, and the data never leaves the warehouse.

How it works

Three steps to running.

01
Load or query in place

Ingest into columnar storage, or query the data lake directly with no copy.

02
Scale compute on demand

Spin up isolated compute clusters per workload; storage and compute scale independently.

03
Serve BI and ML

Point dashboards, notebooks and the model at the same governed tables.

What you get

Data warehouse, in full.

Storage/compute split

Scale query clusters up and down without touching the data.

Columnar & vectorized

Fast scans and aggregations over wide, large tables.

Lake integration

Query the data lake in place, no copy required.

Concurrency scaling

Spin up compute for bursts so dashboards never queue.

API-first

Provision it in a few lines.

Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.

import { aether } from "@aether/sdk";

// Provision data warehouse and query it
const data_warehouse = await aether.databases.create({
  service: "data-warehouse",
  name: "app",
  region: "us-1",
});

const rows = await data_warehouse.query(`select * from events limit 10`);
Specs

At a glance.

Architecture
Separated storage & compute
Format
Columnar, vectorized execution
Scale
Petabyte-scale tables
Concurrency
Elastic, auto-scaling clusters
Lake access
Query open table formats in place
Use cases

Built for real work.

01

BI and reporting

02

Ad-hoc analytics over big data

03

ML feature engineering

Why one platform

On one model, not stitched together.

The usual stack runs data warehouse in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.

One platform

No stitching a vector DB to one place, a warehouse to another and a model to a third — data warehouse sits next to the rest of the catalog, one identity, one bill.

The model is here

The provider that runs Aether runs your data warehouse — so the data and the model never leave the same governed boundary to talk to each other.

Built on demand

Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.

FAQ

Good to know.

How does this compare to the cloud data warehouse or the cloud warehouse?

Same separated storage/compute model and elastic concurrency — but on the platform that also serves the model and governs your identity, deployable into your boundary.

Do I have to load data first?

No — query the data lake in place via open table formats, or load into columnar storage for the hottest workloads.

Can I run it air-gapped?

Yes — like every catalog service it deploys managed, in your VPC, on-prem or fully air-gapped.

Deploy anywhere

Your boundary, your choice.

Managed
Your VPC
On-prem
Air-gapped / sovereign

Run Data warehouse on Aether Cloud.

Columnar, separation-of-storage-and-compute warehouse for analytics at scale. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.