πŸš€ NexAI-v2 (3B Instruct) β€” GGUF & LoRA

NexAI-v2-3B-Instruct-GGUF is a fine-tuned, instruction-aligned language model based on Qwen/Qwen2.5-3B-Instruct. It is optimized for high-speed reasoning, coding assistance, and ultra-low latency inference on edge devices, browsers, and mobile hardware.

πŸ“Š Model Specifications

  • Base Model: Qwen/Qwen2.5-3B-Instruct
  • Quantization: Q4_K_M (GGUF via llama.cpp)
  • File Size: ~1.93 GB (Exact Mobile/Edge sweet spot)
  • Context Length: 32,768 tokens
  • License: Apache 2.0

πŸ“¦ Files in this Repository

  • nexai-v2-3B-Q4_K_M.gguf: Quantized standalone weights for Ollama, LM Studio, Nirvana Browser, and llama.cpp.
  • /adapter: LoRA fine-tuned weights and tokenizer config.

πŸ’» How to Use with Ollama

Run directly using Ollama CLI:

ollama run hf.co/Anoopsingh53/NexAI-v2-3B-Instruct-GGUF:nexai-v2-3B-Q4_K_M.gguf

🐍 How to Use with Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="Anoopsingh53/NexAI-v2-3B-Instruct-GGUF",
    filename="nexai-v2-3B-Q4_K_M.gguf",
    n_ctx=4096
)

output = llm("Q: Write a Python quicksort algorithm.\nA:", max_tokens=256)
print(output["choices"][0]["text"])
Downloads last month
10
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Anoopsingh53/NexAI-v2-3B-Instruct-GGUF

Base model

Qwen/Qwen2.5-3B
Quantized
(305)
this model