Llama-3 LLM2Vec MNTP Supervised ONNX

Built with Meta Llama 3.

KitsuMate ONNX export of McGill-NLP/LLM2Vec-Meta-Llama-3-8B-Instruct-mntp-supervised, pinned at baa8ebf04a1c2500e61288e7dad65e8ae42601a7. The merged model also contains weights from Meta-Llama-3-8B-Instruct, pinned at 8afb486c1db24fe5011ec46dfbe5b5dccdb575c2.

McGill NLP and the LLM2Vec authors retain authorship of the embedding adapter and method. Meta retains its rights in the Llama materials. KitsuMate only provides the ONNX conversion and quantization in the standard flat onnx/ layout.

Variant

  • onnx/model_int4_b128_fp32act_s64.onnx: ONNX Runtime CPU model with 4-bit blockwise weights, FP32 activations, sequence length 64, and onnx/model_int4_b128_fp32act_s64.onnx_data.

Variant names follow the ONNX Community filename convention and are discovered directly from onnx/.

Distribution and use are subject to the included Meta Llama 3 Community License and Acceptable Use Policy. The LLM2Vec project license is included separately.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support