ANNex

Fast hybrid vector search, built in Rust.

Dense and sparse retrieval in one index. Native multivector / ColBERT reranking. Embeddable as a Rust library or served over HTTP. Beats hnswlib on the recall–latency Pareto frontier up to 96% recall.

0.18ms
p50 latency
ef=32 · EC2
83.6%
recall
@ above config
0.440
macro nDCG@10
5 BEIR corpora

NYT256 · 290K docs · 256-D angular · full methodology

# Rust
cargo add annex

# Python
pip install ANNexDB

v0.2.0 · Apache-2.0 / MIT · one maintainer, actively developed

k-NN graph · query vector and nearest neighbors


Capabilities

ANNex is a Rust-native embedded library with hybrid dense-sparse retrieval, multivector / late-interaction reranking, and WAL-backed persistence. Compare with the most commonly evaluated alternatives.

Feature ANNex Qdrant FAISS hnswlib
Dense vector search ✓ ✓ ✓ ✓
Sparse / BM25 ✓ ✓ ✗ ✗
Hybrid RRF fusion ✓ ✓ ✗ ✗
Multivector / ColBERT ✓ ✓ ✗ ✗
Payload filters ✓ ✓ ✗ ✗
Embedded (no server) ✓ ✗ ✓ ✓
HTTP server ✓ ✓ ✗ ✗
Rust-native library ✓ ✗ ✗ ✗
Python bindings ✓ ✓ ✓ ✓
Snapshots + WAL ✓ ✓ ✗ ✗
Detailed comparison: ANNex vs Qdrant → ANNex vs FAISS →

Performance

ANNex leads hnswlib across the full recall frontier from 83% to 96% recall on NYT256. The gap widens at higher recall targets — where the engineering cost of approaching exactness usually hurts most.

At 87% recall
2.2× faster
ANNex 0.30ms vs hnswlib 0.66ms
At 90% recall
2.6× faster
ANNex 0.47ms vs hnswlib 1.24ms
At 96% recall
3.4× faster
ANNex 2.70ms vs hnswlib 9.16ms
0.80 0.85 0.90 0.95 1.00 0.2ms 0.5ms 1ms 2ms 5ms 10ms p50 latency (ms, log scale) recall ANNex hnswlib M=16
NYT256 · 290K documents · 256-D angular · Apple M2 · 3 rounds · offset=1000. ANNex uses PQ screening and combined-stop filters at each operating point. Full data, raw results, and methodology →

Quickstart

The Rust library is the core — embed it directly or build snapshots that the Python bindings can load. The HTTP server (annex-multivector) handles hybrid queries and named vector fields.

Rust annex = "0.2.0"
[dependencies]
annex = "0.2.0"
use annex::{DistanceMetric, Segment};

// build a 128-dim cosine index
let mut seg = Segment::with_config(
    DistanceMetric::Cosine,
    16,   // M (HNSW graph connections)
    64,   // ef_construction
    16,   // level cap
    128,  // vector dimension
);

// insert vectors
for id in 0..1000u64 {
    let vector = vec![id as f32; 128];
    seg.insert_with_id(id, vector, None).unwrap();
}

// search — returns up to k scored hits
let query = vec![0.1_f32; 128];
let hits = seg.search(&query, 10, None).unwrap();
for hit in hits {
    println!("id={} score={:.4}", hit.id, hit.raw_score);
}
Python pip install ANNexDB
The Python bindings load snapshots produced by the Rust library. NumPy float32 arrays required. Inputs are validated before the GIL is released.
import annexdb
import numpy as np

# load a snapshot built by the Rust library
index = annexdb.Index("segment.bin", quantize=False)

# single query
query = np.random.rand(index.dim()).astype(np.float32)
ids, scores = index.search(query, k=10, ef=128)

# multithreaded batch search
queries = np.random.rand(64, index.dim()).astype(np.float32)
ids, scores = index.search_batch(
    queries, k=10, ef=128, threads=4
)
For hybrid dense+sparse queries and the HTTP API, see the annex-multivector quickstart . Full API docs on docs.rs.

Current Limitations

These are real constraints, not caveats. If any of them are blocking for your use case, ANNex is probably not the right choice yet.

No managed cloud
ANNex is self-hosted only. There is no hosted offering or managed service. You run and operate it.
Single-node
No distributed or sharded mode. All data lives on one machine. Horizontal scaling is not supported in v0.2.0.
Production hardening in progress
ANNex is well-suited for development, evaluation, and early-stage applications. The benchmark policy and RESULTS.md are explicit about what has and has not been established.