YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

DevOps Qwen โ€” Fine-Tuned Qwen2.5-3B-Instruct

A Qwen2.5-3B-Instruct model fine-tuned with LoRA (Low-Rank Adaptation) on a DevOps-focused dataset.

The model is designed for practical DevOps, Cloud, Linux, Docker, Kubernetes, Terraform, CI/CD, networking, monitoring, and troubleshooting questions.

Model Details

Property Value
Base Model Qwen/Qwen2.5-3B-Instruct
Model Type Causal Language Model
Fine-Tuning LoRA
LoRA Rank 8
LoRA Alpha 16
LoRA Dropout 0.05
Target Modules q_proj, k_proj, v_proj, o_proj
Training Epochs 1
Max Sequence Length 512
Quantization Q4_K_M
Format GGUF
Approx. Model Size 1.9 GB
Runtime Ollama / llama.cpp
Primary Use DevOps AI Assistant

What is this model?

This model is a specialized version of Qwen2.5-3B-Instruct trained on DevOps-oriented examples.

The goal of the fine-tuning is to improve the model's ability to provide practical responses for:

  • Linux administration
  • AWS
  • GCP
  • Docker
  • Kubernetes
  • Terraform
  • Git
  • Jenkins
  • CI/CD
  • Networking
  • HTTP troubleshooting
  • Monitoring
  • Production troubleshooting
  • Cloud infrastructure

The model is intended to provide answers with:

  1. Root cause or explanation
  2. Exact commands where appropriate
  3. Short explanation of commands
  4. Production-safe troubleshooting steps

Fine-Tuning Approach

The model was fine-tuned using LoRA โ€” Low-Rank Adaptation.

Instead of updating the entire base model, LoRA trains a small set of additional parameters while keeping most of the original model frozen.

Qwen2.5-3B-Instruct
        |
        v
   LoRA Training
        |
        v
   LoRA Adapter
        |
        v
Merge Adapter + Base Model
        |
        v
   Merged Model

This approach reduces training memory and computational requirements compared with full fine-tuning.

Training Configuration

Base Model:
Qwen/Qwen2.5-3B-Instruct

LoRA:
r = 8
alpha = 16
dropout = 0.05

Target modules:
q_proj
k_proj
v_proj
o_proj

Epochs:
1

Batch size:
1

Gradient accumulation:
4

Learning rate:
2e-4

Maximum sequence length:
512

Model Conversion

After LoRA training, the adapter was merged with the base model.

The merged Hugging Face model was then converted to GGUF using llama.cpp.

LoRA Adapter
     |
     v
Merged Hugging Face Model
     |
     v
GGUF F16
     |
     v
Q4_K_M Quantization
     |
     v
qwen-devops-q4_k_m.gguf

GGUF

GGUF (GPT-Generated Unified Format) is an efficient model format commonly used for local LLM inference with llama.cpp and compatible runtimes.

Q4_K_M

Q4_K_M is a 4-bit quantization format.

It reduces model storage and memory requirements while maintaining a useful level of model quality for local inference.

Approximate sizes:

Merged Hugging Face Model  โ‰ˆ 12 GB
GGUF F16                   โ‰ˆ 5.8 GB
Q4_K_M GGUF                โ‰ˆ 1.9 GB

Hardware Used

The model was developed and tested in a CPU-only environment.

CPU:
AMD EPYC 7543

CPU cores available:
12

RAM:
~57 GB

GPU:
None

CUDA:
False

Python:
3.10.14

Usage with Ollama

Download the GGUF model from this repository.

Create a Modelfile:

FROM ./qwen-devops-q4_k_m.gguf

PARAMETER temperature 0.2
PARAMETER top_k 20
PARAMETER top_p 0.9
PARAMETER repeat_penalty 1.1
PARAMETER num_ctx 4096

SYSTEM """
You are a senior DevOps and Cloud engineer.

Give practical and accurate technical answers.

For Linux, AWS, Docker, Kubernetes, Terraform, Git,
Jenkins, CI/CD, monitoring, networking and troubleshooting:

- Explain the root cause.
- Give exact commands when appropriate.
- Explain commands briefly.
- Do not invent information.
- If you don't know something, clearly say so.
- Prefer safe production-ready solutions.
"""

Create the Ollama model:

ollama create devops-qwen -f Modelfile

Run:

ollama run devops-qwen

Example

Question:

How do I troubleshoot a 502 Bad Gateway error from an AWS ALB?

The model is intended to provide a structured troubleshooting approach such as:

1. Check ALB target health
2. Verify application is listening on the expected port
3. Check security groups
4. Check target response
5. Review ALB access logs
6. Review application logs
7. Test the target directly
8. Check health-check configuration

Example commands may include:

ss -lntp
curl -v http://127.0.0.1:8080/
curl -v http://TARGET_PRIVATE_IP:8080/

Ollama API

Non-streaming request:

curl http://localhost:11434/api/generate \
  -d '{
    "model": "devops-qwen",
    "prompt": "How do I check disk usage in Linux?",
    "stream": false
  }'

Streaming request:

curl http://localhost:11434/api/generate \
  -d '{
    "model": "devops-qwen",
    "prompt": "How do I troubleshoot Kubernetes CrashLoopBackOff?",
    "stream": true
  }'

llama.cpp

The GGUF model can also be used with llama.cpp:

./llama-cli \
  -m qwen-devops-q4_k_m.gguf

Recommended Generation Parameters

For technical and DevOps questions:

temperature     = 0.2
top_k           = 20
top_p           = 0.9
repeat_penalty  = 1.1
context         = 4096

Lower temperature is used to encourage more deterministic and consistent technical responses.

Fine-Tuning vs RAG

This model should not be considered a replacement for RAG.

Fine-tuning is useful for:

  • Response style
  • Domain behavior
  • Task patterns
  • DevOps troubleshooting patterns
  • Command-oriented responses

RAG is useful for:

  • Company documentation
  • Current infrastructure information
  • Internal runbooks
  • AWS architecture documentation
  • Frequently changing configuration
  • Private knowledge bases

Recommended architecture:

User
 |
 v
Chat UI
 |
 v
n8n / FastAPI
 |
 v
RAG Retriever
 |
 v
Vector Database
 |
 v
Relevant DevOps Documents
 |
 v
Context
 |
 v
devops-qwen
 |
 v
Final Answer

Intended Use

This model is intended for:

  • DevOps assistants
  • Cloud troubleshooting assistants
  • Linux support
  • Infrastructure automation
  • CI/CD assistance
  • Kubernetes troubleshooting
  • Terraform assistance
  • Internal technical assistants
  • RAG-based DevOps assistants

Limitations

The model is relatively small at approximately 3B parameters.

It may:

  • Make incorrect technical assumptions
  • Produce outdated information
  • Generate commands that require environment-specific changes
  • Fail on complex infrastructure architecture
  • Require RAG or external tools for current infrastructure information

Always verify commands before running them in production.

For production environments, use appropriate:

  • Backups
  • Change management
  • Testing
  • Access controls
  • Approval processes

Security

Do not provide the model with:

  • AWS access keys
  • Private SSH keys
  • Passwords
  • API tokens
  • Database credentials
  • TLS private keys
  • Other secrets

When integrating this model with automation, use least-privilege credentials and approval controls for destructive operations.

Project Pipeline

DevOps Dataset
      |
      v
JSONL Validation
      |
      v
Train / Validation Split
      |
      v
Qwen2.5-3B-Instruct
      |
      v
LoRA Fine-Tuning
      |
      v
LoRA Adapter
      |
      v
Merge
      |
      v
Merged Model
      |
      v
GGUF F16
      |
      v
Q4_K_M
      |
      v
Ollama
      |
      v
devops-qwen
      |
      v
API / n8n / RAG

Benchmark

The project includes an automated benchmark comparing:

qwen2.5:3b
       VS
devops-qwen

The benchmark contains 10 DevOps questions covering:

  • Linux
  • AWS ALB
  • Docker
  • CPU/RAM
  • Disk usage
  • Terraform
  • Kubernetes
  • HTTP
  • Production troubleshooting

Benchmark output:

benchmark_results.json

Model Card Summary

Model:
DevOps Qwen

Base:
Qwen2.5-3B-Instruct

Fine-Tuning:
LoRA

Format:
GGUF

Quantization:
Q4_K_M

Size:
~1.9 GB

Runtime:
Ollama / llama.cpp

Domain:
DevOps / Cloud / Infrastructure

Recommended:
CPU local inference + RAG

License

This model is derived from Qwen2.5-3B-Instruct.

Users should review and comply with the applicable Qwen model license and its terms before using or redistributing this model, particularly for commercial use.

The fine-tuning dataset and any additional project components may have their own applicable terms.

Disclaimer

This model is an experimental DevOps-focused AI assistant. It is not a substitute for production change-control procedures or expert review.

Always validate generated commands and infrastructure changes before applying them to production systems.

Downloads last month
89
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support