Vector database.
Managed embeddings and similarity search at scale — the retrieval layer under grounded apps.
A managed vector database for embeddings — store, index and search billions of vectors with hybrid filtering, so apps and agents retrieve grounded context fast and accurately.
- Category
- Databases
- Deployment
- Managed → air-gapped
- Governance
- IAM · encryption · audit
Retrieval became the new join. Most stacks bolt it on.
Every grounded AI app — RAG, semantic search, recommendations, deduplication — depends on the same primitive: turn data into embeddings, index them, and find the nearest neighbours fast, with filters. The usual way to get this is to stand up a dedicated vector database next to a separate embedding model and a separate system of record, then keep three copies of the truth in sync. The retrieval is milliseconds; the architecture around it is a synchronisation problem with a governance problem stapled on, because the embeddings now live somewhere your access policies and lineage don't reach.
Approximate search that stays accurate as the data moves.
Aether's vector store indexes billions of vectors with approximate-nearest-neighbour search and a recall/latency trade-off you set per index — not a global compromise. Inserts, updates and deletes are live: the index reflects new data without a rebuild, so a knowledge base that changes hourly doesn't drift away from what's indexed. Every result carries provenance — which document, which version, which embedding model produced it — so a retrieved fact is auditable, not a vector that happened to be close.
Hybrid search, because similarity alone is wrong often enough to matter.
Pure vector similarity retrieves things that are semantically near but contextually wrong — the right topic, the wrong tenant; the right concept, an expired document. Aether runs hybrid retrieval: vector similarity fused with metadata and keyword filters in a single query, so you get the nearest neighbours that also satisfy the hard constraints. That is the difference between a demo that retrieves plausible passages and a system you can put in front of a customer.
The embeddings come from the model that reads them.
Because the embedding model is Aether and it runs on the same platform, embeddings are generated, stored and queried inside one boundary. There is no export of your data to an external embedding API, no second copy in a separate vector vendor, and no reconciliation when the model version changes — re-embedding is a job on the same data, not a migration across services. Retrieval becomes a join against your governed tables, under your identity, with the model that will use the results sitting right next to the store that holds them.
Retrieval as a join: find the passages nearest to a query embedding that also belong to this tenant and aren't expired, then hand them to the model — one statement, one boundary.
SELECT d.id, d.title,
VECTOR_DISTANCE(d.embedding, AETHER.EMBED(:query)) AS score
FROM knowledge.documents d
WHERE d.tenant_id = :tenant -- hard filter, not a hope
AND d.expires_at > now() -- no stale answers
ORDER BY score -- vector similarity
LIMIT 8;On a standalone vector database the tenant and expiry filters are metadata you maintain in a second system, the embedding is produced by an external API call, and the documents themselves live in a third place. Here the embedding, the vectors and the governed rows are one query under one identity — and the same model that embedded the query is the one that will answer with the results.
Three steps to running.
Generate embeddings with Aether or your own model; the store keeps them fresh as data changes.
Build a scalable ANN index with tunable recall and latency over billions of vectors.
Query by similarity with metadata and keyword filters, with provenance on every result.
Vector database, in full.
Scalable ANN indexing with tunable recall and latency.
Combine vector similarity with metadata and keyword filters.
Insert, update and delete without rebuilding the index.
Embeddings from Aether, kept fresh, with provenance per result.
Provision it in a few lines.
Every service is reachable from the same SDK, CLI and infrastructure-as-code — one identity, one bill, one audit trail across the whole catalog.
import { aether } from "@aether/sdk";
// Provision vector database and query it
const vector_database = await aether.databases.create({
service: "vector-database",
name: "app",
region: "us-1",
});
const rows = await vector_database.query(`select * from events limit 10`);At a glance.
- Scale
- Billions of vectors
- Search
- ANN + hybrid (vector + keyword + filter)
- Updates
- Live insert / update / delete
- Recall
- Tunable per index
- Embeddings
- Aether-native, kept fresh
Built for real work.
Retrieval-augmented generation
Semantic search
Recommendations and dedup
On one model, not stitched together.
The usual stack runs vector database in one product, the model in another and the data in a third — and the seams between them are the cost. Aether Cloud runs it on the same platform that serves the model, governs your identity and deploys into your boundary, with the rest of the catalog one hop away.
No stitching a vector DB to one place, a warehouse to another and a model to a third — vector database sits next to the rest of the catalog, one identity, one bill.
The provider that runs Aether runs your vector database — so the data and the model never leave the same governed boundary to talk to each other.
Need a capability that isn’t here yet? The model writes and deploys it into the same boundary — the catalog is a starting point, not a ceiling.
Good to know.
Same managed ANN and hybrid search — but co-located with the model that generates the embeddings, so retrieval stays inside one governed boundary.
Yes — combine vector similarity with metadata and keyword filters in a single query.
Inserts and updates re-embed automatically; provenance is tracked per result.
Your boundary, your choice.
Pairs well with.
Managed Postgres-compatible SQL with HA, read replicas, backups and point-in-time restore.
Elastic document and key-value stores for low-latency, high-throughput workloads.
Purpose-built time-series storage for telemetry, metrics and sensor streams.
Sub-millisecond managed cache for sessions, hot data and rate limiting.
A property-graph database for relationships, knowledge graphs and path queries.
Columnar, separation-of-storage-and-compute warehouse for analytics at scale.
Run Vector database on Aether Cloud.
Managed embeddings and similarity search at scale — the retrieval layer under grounded apps. Deployable managed, in your VPC, on-prem or fully air-gapped — talk to us about the configuration your workloads and your boundary require.