fastembed is the crate most Rust projects reach for when they want embeddings locally: a synchronous wrapper around ONNX Runtime and Hugging Face tokenizers, shipping a curated table of ready-to-use models — the Rust sibling of Qdrant’s Python fastembed. It’s the most-used one by a wide margin — 1.9M downloads in the last 90 days, against ~130k for model2vec-rs, the next busiest embeddings crate — and that table is a large part of why it’s pleasant: pick a name, get a model.
I needed embedding models that fastembed doesn’t ship — and just as often a variant it doesn’t ship. Not exotic things: snowflake-arctic-embed-m-v1.5, or a lower-precision artifact of a model that is in the table (model_q4.onnx next to model_quantized.onnx). fastembed’s table names exactly one file per model, so a variant like that is simply unreachable — and the precision, which is the one dial I actually wanted to turn, is decided by somebody else’s list.
Everything else about the library was fine, and I didn’t want to replace any of it. Its in-process ONNX inference works, it sits on a pinned ONNX Runtime, and I’d rather not own tokenization, padding and matmul scheduling myself. What I needed to take over was the part in front of inference: which artifact to load, from which revision, and whether the bytes are the ones I think they are. So I wrapped the loader and kept the engine — the wrapper lives in patterns, the Rust crate I share across my projects.
What follows: what the table hides, the loader I built, what I decided against, and one measured effect of int8 quantization I still don’t have an answer for.
What the table hides
Everything quoted below is from fastembed 6.1.0 (with hf-hub 0.5 under it).
A registry entry looks harmless — EmbeddingModel::SnowflakeArcticEmbedMQ — but six things are decided for you, and most of them are invisible:
- Bytes come from
main.model_code+model_filename a repo and a path; the fetch resolves whatever that branch currently points at. Two loads a month apart can be different models under the same name. - Nothing is verified. hf-hub — the layer under fastembed — hashes nothing in its production path. Its download writes a temp file and renames it to a cache blob named after the etag the server sent;
Sha256appears only in its tests. TLS and a filename convention, no content check. - Output selection is a precedence list.
[OnlyOne, text_embeds, last_hidden_state, sentence_embedding]— andlast_hidden_stateoutrankssentence_embedding. On an export that ships both (and plenty do), the default silently picks the token-level tensor. - Pooling defaults to CLS.
NonebecomesPooling::Clswithout a word. On a mean-pooled model with a token-level output, you get well-formed vectors and quietly worse retrieval. - The context window is a constant.
DEFAULT_MAX_LENGTH: usize = 512, private, applied regardless of what the artifact’stokenizer_config.jsonsays. For a 2048-token model, the other three quarters of the window are wasted. - Matryoshka truncation doesn’t exist. For a local embedder there is no way to store 256 dims of a 768-dim model, so the one storage lever MRL gives you is out of reach.
None of that is a bug. It’s what a curated table buys you: convenience, at the cost of everything above being someone else’s fixed decision.
A spec instead of a table
The replacement is one type, and it is deliberately boring:
let spec = ModelSpec::new(
"Snowflake/snowflake-arctic-embed-m-v1.5",
"e58a8f756156a1293d763f17e3aae643474e9b8a", // full commit sha, not a branch
"onnx/model_quantized.onnx",
768, // native width
)?
.with_output("sentence_embedding") // required when a graph has several outputs
.with_truncated_dims(TruncatedDims::new(256)?); // optional MRL width
Behind it (the code lives in my patterns crate):
- Pinned fetch.
revisionmust be a 40-hex commit sha; branches, tags and short shas are rejected. The bytes cannot change under an identity. - Verification. The hub’s own file listing (
/api/models/{repo}/tree?expand=true) exposes thelfs.sha256per artifact, so on every load we recompute the hash and compare. That catches a corrupted or truncated cache entry, a swapped blob, or a file that no longer matches its revision. It’s a few tenths of a second for a 110 MB model. - Guards that run before the first embed — a throwaway ONNX session is opened just to read the graph, then every rule is checked and the load fails if any is violated:
- inputs must be feedable (
input_ids,attention_mask,token_type_ids), - a multi-output graph must declare which output to read,
- a 3-D (token-level) output must declare how it is pooled, and when the repo ships
1_Pooling/config.jsonthe declaration must agree with it, - a truncated width must appear in the artifact’s declared
matryoshka_dimensions, - the probed native width must match the spec, and be a multiple of 8.
- inputs must be feedable (
- No fallback model, ever. A failed fetch or verification is an error. Substituting a different model would mean writing rows under an identity that didn’t produce them.
The identity string
Vectors outlive the process that made them, so each one is stored next to an identity. Mine looks like this:
mixedbread-ai/mxbai-embed-large-v1/onnx/model_quantized.onnx
@b33106f585b9ce46904ad7443a3b52b7a63e231c#d1024+c74b8baca
Repo and file name the artifact; @b33106f… is the pinned commit the bytes were verified against; #d1024 is the stored width; +c74b8baca is a digest of everything else that shapes the vector — selected output, pooling, quantization and the model’s prefixes. A truncated width adds /meta or /spec, recording whether the artifact declared that width or whether I merely claimed it.
That digest was the second iteration. My first version was {repo}/{file}@{revision}#d{width} — which looks complete and isn’t: change the pooling or the query prefix and you keep the same identity while producing different vectors. If rows are keyed by that string, two configurations can silently share a key. Anything that changes the numbers belongs in the key.
What I decided not to do
Two tempting things: one I built and deleted, one I never built.
I didn’t take over the inference session. I did, for a while: my own ONNX Runtime session, my own tokenizer wiring, pooling and normalization — which loads a model once instead of twice, drops a dependency, and opens the door to decoder models and execution providers. I proved it produced bit-identical vectors to fastembed’s path, then deleted it. Three hundred-odd lines of inference semantics are not worth owning just to save one model load — and the features it unlocked (decoder inputs, GPU execution providers) need models or hardware this box doesn’t have. The rule: test what you own, not what you delegate, and don’t own what you won’t exercise.
I didn’t extend pooling beyond CLS and mean. Pooling is how a model’s per-token output becomes one vector: Cls takes the first token, mean averages all of them, and those are the only two fastembed offers. You can’t plug in a third — the setting isn’t a hook you can call, it’s a fixed list inside the library. Adding “last token” (what the newer decoder-based embedders use) would mean replacing fastembed’s inference loop with mine, which is the thing the previous point already declined.
That’s fine here, because the models that want it are big: last-token embedders start at 0.6B parameters — already about twice the size of the encoder I run — and go up to 8B, so on a CPU box only the smallest, Qwen3-Embedding-0.6B, is even worth considering. For that one, someone has already published an ONNX export with the last-token pooling done inside the graph, so it loads with the two modes I have: the model does the pooling, the loader just reads the vector out (trusting that conversion, as with any pre-pooled export). The bigger ones belong behind an embedding server.
And if an artifact’s declared pooling contradicts its own 1_Pooling/config.json, the loader refuses it at load rather than quietly falling back to CLS.
The finding I didn’t expect
While measuring the loader I noticed something about dynamically quantized int8 artifacts (the model_quantized.onnx most repos ship). In those graphs, weights are quantized offline but activations are quantized per inference — the ONNX spec defines DynamicQuantizeLinear as computing the scale from the input tensor it is handed:
y_scale = (max(0, max(x)) - min(0, min(x))) / (qmax - qmin)
So the same text produces slightly different vectors depending on what else was in the batch. Measured across three models: cosine 0.9977–0.9991 between a single-text and a batched embedding of the same text, which is 0.9–3.7 % of sign bits flipping after binarization.
Practically: if queries go through one-at-a-time calls and documents through batched calls — as they almost always do — the two sides sit in slightly different quantization regimes. The perturbation is query-side only and second-order for ranking, but it’s real, it’s inherent to the artifact rather than to any library, and the fix is to keep the batching scheme stable or to prefer statically quantized weights.
How much this matters for binary embeddings — my index stores one sign bit per dimension, following this post on binary vector embeddings — remains to be seen. A sign bit is the most brutal possible reduction of a value, so a few thousandths of drift near zero is exactly where it can flip a bit, and I don’t have retrieval numbers yet to say whether that moves rankings in practice. That’s an open question, not a settled one.
What the wrapper costs
It’s almost free: a spec type, a fetch-and-verify layer, a graph-validation pass and a few hundred lines of tests, plus one throwaway ONNX session per load — which is exactly what lets the guards fail before anything is embedded. What’s left are limits of the setup rather than of the library: in-process, on CPU, for models this size, a fixed input set and CLS-or-mean pooling are all that’s needed, and the models that want more are the ones that shouldn’t run here in the first place.
For this workload that’s the right trade: fastembed keeps doing the part it’s good at, and the part that decides retrieval quality — which model, from which revision, verified, with which output, pooling and width — is a spec I control.
Written with LLM assistance; the code, the measurements and the dead ends are all real.