Benchmarks

All claims are governed by the benchmark policy. No result here establishes production superiority or launch readiness.


Retrieval Quality · BEIR

nDCG@10 on 5 held-out BEIR corpora. Shared pinned MiniLM vectors, exact cosine, 100 candidates/channel. Every per-corpus 95% bootstrap interval between hybrid systems includes zero — parity, not a statistically established advantage.
Impl f68f970 · Apple M2 / 16 GiB · 2026-09-29

0.0 0.1 0.2 0.3 0.4 0.5 macro nDCG@10 ANNex hybrid 0.4401 LanceDB native hybrid 0.4400 Qdrant Server hybrid 0.4396 ANNex dense 0.4127 ANNex BM25 0.3751
Per-corpus breakdown ↓
System NF
(162q)
Sci
(150q)
Argu
(703q)
FiQA
(324q)
SciD
(500q)
Macro
ANNex hybrid 0.3574 0.7264 0.5238 0.3798 0.2133 0.4401
LanceDB native hybrid 0.3543 0.7276 0.5265 0.3784 0.2132 0.4400
Qdrant Server hybrid 0.3580 0.7257 0.5207 0.3798 0.2136 0.4396
ANNex dense 0.3149 0.6642 0.4830 0.3773 0.2243 0.4127
ANNex BM25 0.3285 0.6741 0.4668 0.2435 0.1625 0.3751

Head-to-Head · NYT256

ANNex vs FAISS, hnswlib, and USearch at three recall tiers. Interleaved run — all engines share OS memory state. Amber = ANNex, gray = competitors.
290K docs · 256-D angular · Apple M2 · offset=4000 · ANNex v0.1.0 · FAISS 1.15.1 · hnswlib 0.8.0 · usearch 2.26.2

0.84 0.87 0.90 0.93 0.96 0ms 1ms 2ms 3ms 4ms 5ms 6ms 7ms p50 latency (ms) recall ANNex competitor
Full table ↓
System ef Recall p50 ms p99 ms
annex_screen 32 0.864 0.252 1.042
faiss_hnsw16 128 0.868 0.356 0.598
faiss_hnsw32 32 0.834 0.208 0.373
usearch_f16 128 0.864 0.540 0.941
hnswlib_m16 128 0.861 0.684 1.127
annex_screen 128 0.918 0.580 0.972
faiss_hnsw16 512 0.925 1.457 2.758
faiss_hnsw32 128 0.904 0.627 1.009
usearch_f16 512 0.922 2.059 4.246
hnswlib_m16 512 0.918 2.562 7.010
annex_screen 512 0.961 2.205 5.990
faiss_hnsw32 512 0.957 2.489 4.344
faiss_ivf1024 512 0.964 3.753 9.103
usearch_f32 1024 0.948 6.261 9.592
hnswlib_m16 1024 0.946 5.043 10.500

NYT256 · Recall–Latency Pareto

Fastest ANNex config at each recall threshold, vs hnswlib baseline. Offset=1000 excludes query vectors from the corpus. The 0.97 row marks where hnswlib takes the Pareto front.
Apple M2 / 16 GiB · 3 rounds · commit 4b11157 · hnswlib 0.8.0

Target Config ef Recall p50 ms p99 ms
0.80 combined_stop0.01 32 0.827 0.225 0.517
0.85 pq_screen0.1 128 0.871 0.300 0.564
0.87 pq_screen0.1 128 0.871 0.300 0.564
0.90 pq_screen0.15 128 0.903 0.466 1.122
0.92 pq_screen0.1 256 0.925 0.629 1.108
0.94 pq_screen0.1 512 0.956 1.412 2.219
0.96 pq_screen0.15 512 0.961 2.699 4.239
0.97 hnswlib 2048 0.971 9.162 11.065

EC2 · Throughput · Top-k=10

sq8 = SQ8 quantization; rcm = residual compression map. Single thread, serial requests without warmup. Timing excludes embedding generation.
290K docs · 256-D angular · EC2 · 2026-10-01

Config ef Recall p50 ms p99 ms QPS
m16 32 0.836 0.203 0.331 4,706
m16+sq8 32 0.836 0.182 0.282 5,324
m16+rcm 32 0.836 0.202 0.326 4,763
m16+rcm+sq8 32 0.836 0.179 0.282 5,424
m16 128 0.902 0.559 0.894 1,769
m16+sq8 128 0.902 0.500 0.792 1,971
m16+rcm+sq8 128 0.902 0.484 0.765 2,037
m16+sq8 256 0.925 0.914 1.313 1,104
m16+rcm+sq8 256 0.925 0.897 1.308 1,127

Methodology

Recall is measured against exact nearest neighbors computed independently. All ANNex runs use 3 rounds; reported latencies are medians.

Timing measures end-to-end search latency only. Embedding generation, index build, training, and recovery are excluded.

nDCG@10 uses linear relevance gains matching BEIR's evaluator and trec_eval. Verified against pytrec_eval within 1e-12.

Interleaved runs share OS memory state between engines — realistic but not fully isolated per-engine timing.

BENCHMARK_POLICY.md  ·  RESULTS.md (corrections)