🇸🇦 Agentic Arabic 2.6 (GGUF Q4_K_M)

Agentic Arabic 2.6 GGUF is a state-of-the-art 2.6B parameter model fine-tuned specifically for high-precision Arabic Function Calling, Tool Use, and General Instruction Following, built on top of Liquid AI's LFM 2.5 architecture using Unsloth.

👉 Interactive Playground & Demo: https://huggingface.co/spaces/DarkStrox/Agentic-Arabic-2.6-Demo


📈 Benchmark Performance

1. 🛠️ Arabic Function Calling & Tool Use

Evaluated live on unseen Arabic tool-invocation prompts against foundation baselines:

Model Architecture / Size Quantization Function Calling Accuracy VRAM Footprint
🟣 Agentic Arabic 2.6 (Q6_K) Liquid LFM (2.6B) Q6_K 95.0% 🟢 2.08 GB
🟣 Agentic Arabic 2.6 (Q4_K_M) Liquid LFM (2.6B) Q4_K_M 85.0% 🟡 1.67 GB
🩶 LFM 2.6 Base Liquid LFM (2.6B) Q5_K_M 24.0% 🔴 1.94 GB
Gemma 4 E4B IT Transformer (7.4B) Q4_K_XL 10.0% 🔴 4.21 GB

2. ⚡ General Arabic QA & Inference Speed

Evaluated across standard Arabic general instruction tasks (Math, Geography, Concept Explanation, Poetry, Reasoning):

Model General QA Accuracy Avg Response Latency Speed Multiplier
🟣 Agentic Arabic 2.6 (Q4_K_M) 100% (5/5) 🟢 1,493 ms 1.0x (Baseline)
🟣 Agentic Arabic 2.6 (Q6_K) 100% (5/5) 🟢 9,570 ms 6.4x slower
🩶 LFM 2.6 Base 100% (5/5) 6,593 ms 4.4x slower
Gemma 4 E4B IT 100% (5/5) 13,662 ms 9.1x slower

💡 Quantization Insight: Q4_K_M delivers ultra-fast response speeds (~1.49s) and 100% general instruction accuracy in a ultra-lightweight 1.67 GB memory footprint. For strict JSON-schema function calling precision, Q6_K is recommended.


🚀 Quickstart Usage

1. Using llama-cpp-python

from huggingface_hub import hf_hub_download
from llama_cpp import Llama

# Download model GGUF from Hugging Face
model_path = hf_hub_download(
    repo_id="DarkStrox/Agentic-Arabic-2.6-GGUF",
    filename="LFM2.5-2.6B.Q4_K_M.gguf"
)

llm = Llama(
    model_path=model_path,
    n_ctx=4096,
    n_gpu_layers=-1
)

system_prompt = """You are a function calling AI model. You are provided with function signatures within <tools> </tools> XML tags.
<tools>
[{"name": "get_weather", "description": "Get current weather", "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}]
</tools>"""

user_prompt = "ما هي حالة الطقس اليوم في عمان؟"

prompt = f"<|im_start|>system\n{system_prompt}<|im_end|>\n<|im_start|>user\n{user_prompt}<|im_end|>\n<|im_start|>assistant\n"

response = llm(prompt, max_tokens=512, stop=["<|im_end|>"])
print(response['choices'][0]['text'])

2. Using llama-server CLI

llama-server.exe -m LFM2.5-2.6B.Q4_K_M.gguf --ngl 99 --port 8080 --ctx-size 4096

📜 Model Details

  • Developed by: Custom Arabic Fine-Tuning Pipeline (Unsloth)
  • Base Architecture: Liquid AI LFM 2.5
  • Format: GGUF (Q4_K_M)
  • File Size: 1.67 GB
  • Live Space Demo: DarkStrox/Agentic-Arabic-2.6-Demo
Downloads last month
57
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for DarkStrox/Agentic-Arabic-2.6-GGUF

Quantized
(58)
this model

Space using DarkStrox/Agentic-Arabic-2.6-GGUF 1