my-private-search-agent

A 3B-parameter conversational agent fine-tuned from Qwen2.5-3B-Instruct, optimized for private/local search-assistant workflows and quantized to GGUF for fast CPU/GPU inference with llama.cpp and Ollama.

This model was fine-tuned and converted to GGUF using Unsloth, which enabled ~2x faster training.

Model Details

Base model Qwen/Qwen2.5-3B-Instruct
Architecture Qwen2
Parameters 3B
Format GGUF
Quantization Q4_K_M (4-bit, ~1.93 GB)
Fine-tuning framework Unsloth
Chat template ChatML (Qwen2.5 default)
License Apache 2.0

Intended Use

This model is designed to act as a lightweight, locally-runnable assistant for search-oriented tasks โ€” e.g., interpreting a user query, deciding what to look up, and summarizing retrieved results. It's intended for use in a private/offline pipeline rather than as a general-purpose chat model.

In scope:

  • Query understanding and reformulation
  • Summarizing or reasoning over retrieved documents
  • Tool-calling / agentic workflows paired with a local search or retrieval backend

Out of scope:

  • Standalone factual question-answering without a retrieval step (base 3B models are prone to hallucination)
  • High-stakes decisions (medical, legal, financial) without human review

Files

File Quantization Size
Qwen2.5-3B-Instruct.Q4_K_M.gguf Q4_K_M (4-bit) 1.93 GB

Usage

llama.cpp

# Text-only chat
llama-cli -hf bijoy0236/my-private-search-agent --jinja

Ollama

An Ollama Modelfile is included in this repo for easy local deployment:

ollama create my-private-search-agent -f Modelfile
ollama run my-private-search-agent

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="bijoy0236/my-private-search-agent",
    filename="Qwen2.5-3B-Instruct.Q4_K_M.gguf",
)

response = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Find recent papers on retrieval-augmented generation."}]
)
print(response["choices"][0]["message"]["content"])

Training

  • Base model: Qwen2.5-3B-Instruct
  • Method: Supervised fine-tuning (SFT) via Unsloth
  • Dataset: [add dataset name/size/source here]
  • Hardware: [e.g., 1x RTX 4090, X hours]
  • Hyperparameters: [LoRA rank, learning rate, epochs, etc.]

Limitations & Bias

  • Inherits the general limitations of Qwen2.5-3B-Instruct, including occasional factual errors and outdated knowledge beyond its training cutoff.
  • 4-bit quantization (Q4_K_M) trades some accuracy for a smaller footprint and faster inference โ€” expect minor quality loss versus the full-precision checkpoint.
  • Not evaluated for safety-critical or adversarial search scenarios; outputs should be reviewed before use in production pipelines.

Hardware Compatibility

At Q4_K_M (1.93 GB), this model runs comfortably on most consumer CPUs and GPUs with โ‰ฅ4 GB of available RAM/VRAM. See the hardware estimator on the repo page for device-specific estimates.

Citation

@misc{my-private-search-agent,
  author = {bijoy0236},
  title  = {my-private-search-agent},
  year   = {2026},
  url    = {https://huggingface.co/bijoy0236/my-private-search-agent}
}

Acknowledgements

  • Unsloth โ€” 2x faster fine-tuning
  • Qwen Team โ€” base Qwen2.5-3B-Instruct model
  • llama.cpp โ€” GGUF inference engine
Downloads last month
429
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for bijoy0236/my-private-search-agent

Base model

Qwen/Qwen2.5-3B
Quantized
(279)
this model