← return

UltraDim

The first vector store built for wide vectors and analytics too, not just retrieval — cluster, classify, visualise and detect anomalies straight off the index, at full vectors.

Mainstream vector stores assume hundreds to a few thousand dimensions. Scientific and analytical data breaks that assumption: whole-genome methylation profiles span tens of millions of CpG sites, molecular fingerprints hash into thirty-million-bit spaces, customer×product matrices reach hundreds of thousands of columns. UltraDim was built for exactly this regime — ultra-wide and sparse — where conventional stores collapse under their own memory cost long before they can answer a query.

The breakthrough is mathematical, not just engineering. UltraDim represents ultra-wide, sparse data in a new type of data index that stays exact yet scales with a better big-O — so the big datasets that break other stores simply aren’t a barrier for us. We multiply the gains too — our index powers search but also clustering, visualisation, embedding, recommendation and anomaly detection alike. Not just retrieval — an analytical environment.

BloomMap: 49 leaf clusters of 50,000 ChEMBL molecules, clustered at 30 million dimensions, rendered as petals around a dark Voronoi core
BloomMap — 50,000 ChEMBL molecules clustered at 30,000,000 dimensions on the store’s reused keys. 49 leaf clusters; every petal is real data, labelled by its most central molecule.

Measured, on one laptop

Operation & corpuswidthmeasured result
Search — ChEMBL, 500k molecules3×10⁷ recall@10 0.9964 @ 29 ms
Search — DBpedia-1M, dense1,536 precision@10 0.9945 · 1,339 q/s @ 2.99 ms
Search — retail baskets, 100k100,000 recall@10 0.9851 @ 17.5 ms
Search — synthetic, 50k1.8×10⁶ recall@10 1.000 @ 15.7 ms
Embedding — ChEMBL UMAP, 50k3×10⁷ graph k-NN recall 0.9996 vs exact
Recommendation — MovieLens-20M20,720 ties trained MF, zero training
Factorisation — retail spend 0.80 correlation, above ALS ceiling
Streamed anomaly — DBpedia · ChEMBL3×10⁷ AUROC 1.0 · 247/s, real time

All figures from the UltraDim technical paper, single Apple M3 Max host, measured against exact brute-force oracles. Cold-start and comparison caveats are disclosed in the paper — approximation is measured, not hidden.

UltraDim UMAP of 50,000 ChEMBL molecules embedded directly from 30,000,000 raw dimensions
UltraDim UMAP — 50,000 ChEMBL molecules embedded straight from 30,000,000 raw dimensions. Graph k-NN recall 0.9996 against the exact oracle.

What it does

Streaming anomaly map: 987,442 fitted DBpedia articles in grey, with live arrivals drawn as rings coloured green where familiar and red where anomalous
Streamed anomaly detection — 987,442 DBpedia articles fitted (grey); 2,904 live arrivals scored as they land, red where novel. Scoring runs at retrieval latency — 824 arrivals/s on dense data, and 247/s even at 30-million dimensions. The same measured angle drives real-time fraud and outlier detection.

Early access is open

We are inviting research groups and industrial teams with difficult vector datasets — molecular fingerprints, genomic profiles, spectra, sparse behavioural matrices. Free for academics.

apply for early access read the paper