lecore-bge-assimilated

BAAI/bge-base-en-v1.5 passed through leCore's assimilation + budgeted requantize machinery (Marchenko-Pastur spectral filter, then per-tensor bit-width chosen by mean-embedding-cosine โ‰ฅ 0.99 at group-64).

What happened, measured (2026-08-14):

  • MP filter: 0 layers filtered (71/73 heavy-tail passthrough) โ€” the predicted no-op on a well-trained encoder.
  • Requantize landed at mean 7.62 bits/weight (15ร—3b, 4ร—4b, 9ร—5b, 14ร—6b, 20ร—8b, 11 tensors kept fp32 โ€” the encoder resists the 3-bit the same machinery achieved on a 9B decoder). Packed estimate **132 MB**; this repo ships the dequantized-fp32 container (438 MB) snapped to the quant grid โ€” it loads anywhere bge-base loads.
  • Quality: SciFact nDCG@10 0.7388 (full bge: 0.7404); NFCorpus 0.3716 (0.3735); ArguAna 0.6375 (identical); SCIDOCS 0.2152 (0.2172). Zero-lexical-overlap NIAH recall@64 0.8464 โ€” identical to full bge. CPU inference scores bit-identical to GPU.
  • Honesty notes: requantize buys size, not CPU matmul speed (inference runs at normal bge-base pace โ€” 27โ€“31 ms batch-1 query on 8 CPU threads); a packed runtime (e.g. GGUF) would be needed to realize the memory number at inference; uniform 4-bit (93 MB, 98.4% quality) measured as the better packed-runtime point, uniform 3-bit degrades visibly (0.6922) and should not ship.

Method, full matrix (adversarial decoys, FAISS flat/HNSW, cost-to-first-answer), and every number's provenance: https://benches.openzoo.fun

Downloads last month
25
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
I64
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for staccs/lecore-bge-assimilated

Quantized
(22)
this model