Data warehouse.
Columnar, separation-of-storage-and-compute warehouse for analytics at scale.
A columnar cloud warehouse that separates storage from compute, so you scale query power independently and pay for the scans you run — fast analytics over petabytes without tuning.
- Category
- Databases
- Deployment
- Managed → air-gapped
- Governance
- IAM · encryption · audit
The warehouse stopped being the hard part. Everything around it didn't.
A modern analytics estate is rarely one system. There's a warehouse for SQL, an object store and a lake for the raw and semi-structured data, a separate Spark or feature pipeline for ML, a vector database for retrieval, and a model API somewhere else entirely. Each is a good product. The cost is in the seams between them: data gets copied from the lake into the warehouse and again into the feature store, governance is re-implemented three times and drifts, and the moment you want the model to read a governed table you're exporting data across a vendor boundary and hoping the lineage survives the trip. The warehouse query is fast; the organisation around it is slow, and that's where the quarters go.
Separated storage and compute — and one copy of the data.
Aether's warehouse uses the architecture the cloud data warehouse popularised and the cloud warehouse and lakehouse SQL platforms share: storage and compute are decoupled. Your data lives once, in open table formats over object storage; compute clusters are spun up per workload, sized independently, and torn down after. Two consequences matter. First, you scale query power for a heavy dashboard refresh without touching the data or the team running ad-hoc analysis next door — no contention, no over-provisioning a single cluster for the worst case. Second, because the storage is open (Iceberg/Delta-style), the warehouse, Spark, BI and the model all read the same tables in place. There is no proprietary storage format quietly locking you in, and no nightly job copying the lake into the warehouse so the dashboards can see it.
The difference: the model is already in the warehouse.
This is the line that separates Aether from a best-in-class warehouse. On the cloud data warehouse or the cloud warehouse, calling a model means shipping rows to an external inference endpoint, getting predictions back, and writing them somewhere — a pipeline you build, secure and pay egress on, with governance that now spans two vendors. On Aether, the model runs on the same platform, against the same governed tables, under the same identity. A query can classify, extract, embed, summarise or forecast inline; retrieval against the vector store is a join, not an integration; and a fine-tune trains on the warehouse data without it ever leaving the boundary. The data never crosses a wire to meet the model, so there's nothing to secure between them and nothing to reconcile after.
Migrate by reading, not by forklifting.
Because the storage layer is open table formats, you don't have to load everything before you get value. Point the warehouse at the lake and query it in place; bring the hottest tables into managed columnar storage when the latency justifies it. There's no proprietary format you can't get out of, and the same data is queryable from Spark and the model without a second copy. Deployment follows the same rule as the rest of the platform: managed multi-tenant, dedicated, in your VPC, on-prem, or fully air-gapped — the warehouse runs where your data is allowed to live, not only where the vendor offers a region.
Here's the capability that doesn't exist on a standalone warehouse: a single SQL statement that filters governed warehouse rows, retrieves the nearest support cases by vector similarity, and asks the model to draft a resolution — no export, no external endpoint, no second copy of the data.
-- Triage open tickets against similar resolved ones, inline.
SELECT t.id,
t.summary,
AETHER.GENERATE(
'Draft a resolution for this ticket using the similar cases',
t.summary,
past.resolutions
) AS suggested_resolution
FROM support.open_tickets t
CROSS JOIN LATERAL (
SELECT ARRAY_AGG(r.resolution) AS resolutions
FROM support.resolved_tickets r
ORDER BY VECTOR_DISTANCE(r.embedding, t.embedding) -- vector search, as a join
LIMIT 5
) past
WHERE t.status = 'open'
AND t.priority = 'high';On the cloud data warehouse or the cloud warehouse this is three systems and a pipeline: the warehouse for the rows, a vector database for the similarity search, and an external model endpoint for the generation — wired together with code you own and data that crosses two boundaries. Here it's one query, one identity, one audit trail, and the data never leaves the warehouse.
Three steps to running.
Ingest into columnar storage, or query the data lake directly with no copy.
Spin up isolated compute clusters per workload; storage and compute scale independently.
Point dashboards, notebooks and the model at the same governed tables.
Data warehouse, in full.
Scale query clusters up and down without touching the data.
Fast scans and aggregations over wide, large tables.
Query the data lake in place, no copy required.
Spin up compute for bursts so dashboards never queue.
Provision it in a few lines.
Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.
import { aether } from "@aether/sdk";
// Provision data warehouse and query it
const data_warehouse = await aether.databases.create({
service: "data-warehouse",
name: "app",
region: "us-1",
});
const rows = await data_warehouse.query(`select * from events limit 10`);At a glance.
- Architecture
- Separated storage & compute
- Format
- Columnar, vectorized execution
- Scale
- Petabyte-scale tables
- Concurrency
- Elastic, auto-scaling clusters
- Lake access
- Query open table formats in place
Built for real work.
BI and reporting
Ad-hoc analytics over big data
ML feature engineering
On one model, not stitched together.
The usual stack runs data warehouse in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.
No stitching a vector DB to one place, a warehouse to another and a model to a third — data warehouse sits next to the rest of the catalog, one identity, one bill.
The provider that runs Aether runs your data warehouse — so the data and the model never leave the same governed boundary to talk to each other.
Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.
Good to know.
Same separated storage/compute model and elastic concurrency — but on the platform that also serves the model and governs your identity, deployable into your boundary.
No — query the data lake in place via open table formats, or load into columnar storage for the hottest workloads.
Yes — like every catalog service it deploys managed, in your VPC, on-prem or fully air-gapped.
Your boundary, your choice.
Pairs well with.
Managed Postgres-compatible SQL with HA, read replicas, backups and point-in-time restore.
Elastic document and key-value stores for low-latency, high-throughput workloads.
Managed embeddings and similarity search at scale — the retrieval layer under grounded apps.
Purpose-built time-series storage for telemetry, metrics and sensor streams.
Sub-millisecond managed cache for sessions, hot data and rate limiting.
A property-graph database for relationships, knowledge graphs and path queries.
Run Data warehouse on Aether Cloud.
Columnar, separation-of-storage-and-compute warehouse for analytics at scale. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.