Instructions to use TiGa-RCE/gte-Qwen2-1.5B-instruct-MLX-BF16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use TiGa-RCE/gte-Qwen2-1.5B-instruct-MLX-BF16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir gte-Qwen2-1.5B-instruct-MLX-BF16 TiGa-RCE/gte-Qwen2-1.5B-instruct-MLX-BF16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
gte-Qwen2-1.5B-instruct — MLX bf16
BF16 reference checkpoint for this matched local sweep.
This is the BF16 MLX conversion checkpoint from a matched local embedding-quantization experiment. It is published with explicit lineage, calibration evidence where applicable, and the bounded evaluation result that accompanied the conversion.
Provenance and lineage
- Upstream model:
Alibaba-NLP/gte-Qwen2-1.5B-instruct - Upstream revision recorded for publication:
a9af15a6372d7d6b25e9fb07c2ccb9e1fe645644 - Revision evidence: upstream revision verified at publication time; historical local snapshot metadata was not retained
- Direct parent:
Alibaba-NLP/gte-Qwen2-1.5B-instruct - Conversion rule: every quantized checkpoint branches directly from the family MLX BF16 checkpoint; no lossy checkpoint was used to create another.
- Quantization: BF16 MLX conversion
- Local conversion stack: oMLX 0.5.3, mlx-lm 0.31.3, MLX 0.32.0
- Full collection: MLX Embedding Quantization Matrix
PROVENANCE.json contains machine-readable lineage and SHA-256 hashes for the published weight files. No importance matrix was used for this checkpoint.
Bounded local evaluation
| Role | Retrieval smoke |
|---|---|
| BF16 reference | 24/24 top-1, MRR 1.0 |
The evaluation used 24 frozen query/document pairs, the upstream query instruction recipe, last-token pooling, L2 normalization, and direct comparison with vectors from the family BF16 checkpoint. This is an engineering smoke test, not MTEB and not a claim of universal quality. Retrieval success and representation fidelity are reported separately.
Runtime scope
This checkpoint targets Apple Silicon through MLX/oMLX. CUDA and PyTorch results are a separate control lane and must not be interpreted as measurements of MLX/Metal kernel performance.
License and attribution
Apache-2.0, following the upstream model card. The original model authors retain attribution for the upstream model; this repository contains a local MLX conversion or quantized derivative prepared by TiGa-RCE for reproducibility research.
- Downloads last month
- -
Quantized
Model tree for TiGa-RCE/gte-Qwen2-1.5B-instruct-MLX-BF16
Base model
Alibaba-NLP/gte-Qwen2-1.5B-instruct