📢 The LLaMA-3.1-8B distilled 8B version of the R1 DeepSeek AI is available besides the one based on Qwen

📙 Notebook for using it in reasoning over series of data 🧠 :
https://github.com/nicolay-r/nlp-thirdgate/blob/master/tutorials/llm_deep_seek_7b_distill_llama3.ipynb

Loading using the pipeline API of the transformers library:
https://github.com/nicolay-r/nlp-thirdgate/blob/master/llm/transformers_llama.py
🟡 GPU Usage: 12.3 GB (FP16/FP32 mode) which is suitable for T4. (a 1.5 GB less than Qwen-distilled version)
🐌 Perfomance: T4 instance: ~0.19 tokens/sec (FP32 mode) and (FP16 mode) ~0.22-0.30 tokens/sec. Is it should be that slow? 🤔
Model name: deepseek-ai/DeepSeek-R1-Distill-Llama-8B
⭐ Framework: https://github.com/nicolay-r/bulk-chain
🌌 Notebooks and models hub: https://github.com/nicolay-r/nlp-thirdgate

reacted to nicolay-r's post with 🔥 about 1 month ago

Post

1335

🚨 MistralAI is back with the mistral small V3 model update and it is free! 👏
https://docs.mistral.ai/getting-started/models/models_overview/#free-models

🚀 Below is the the provider for reasoning over your dataset rows with custom schema 🧠
https://github.com/nicolay-r/nlp-thirdgate/blob/master/llm/mistralai_150.py

My personal usage experience and findings:
⚠️The original API usage may constanly fail with the connection.
To bypass this limitation, use --attempts [COUNT] to withstand connection loss while iterating through JSONL/CSV data (see 📷 below)

💵 It is actually: ~0.18 USD 1M tokens
🌟 Framework: https://github.com/nicolay-r/bulk-chain