Lore

Lore is a coding-focused QLoRA fine-tune of Qwen/Qwen2.5-Coder-7B-Instruct that reached 56.4% pass@2 on 133 old-aider Python tasks. It was trained end to end on a single 8 GB consumer GPU and is distributed as an 8.1 GB Q8_0 GGUF for local inference.

Project Highlights

  • Built an end-to-end 4-bit QLoRA pipeline with Unsloth, TRL, and PEFT, including rank-16 LoRA, gradient checkpointing, checkpoint recovery, and adapter continuation on a single 8 GB RTX 3060 Ti.
  • Created execution-verified Python repair data with deterministic AST mutation, isolated subprocess testing, tokenizer-aware limits, and parent-disjoint train/validation splits.
  • Reached a 56.4% pass@2 high score on 133 old-aider Python tasks, above the published leaderboard entries for Qwen2.5-Coder 7B Q8_0 (51.9%), Claude 3 Sonnet (54.9%), and GPT-4o mini (55.6%).
  • Exported the final model as an 8.1 GB Q8_0 GGUF for local Ollama and llama.cpp inference.

Model Details

  • Format: GGUF
  • Quantization: Q8_0
  • File size: 8,098,525,408 bytes
  • SHA-256: d9a86d7f85b433f3f4bca84aa150c6db001133907d93446d7b116b886c512655
  • Base model: Qwen/Qwen2.5-Coder-7B-Instruct
  • Initial fine-tuning: bounded 2,000-example subset of Nexlab/fable5-agentic-coding-sft
  • Final continuation: 128 execution-verified repair examples formatted as multi-turn retry conversations
  • Final continuation context length: 1,024 tokens
  • Final continuation steps: 16
  • Final continuation learning rate: 5e-7

Evaluation

Lore was evaluated twice with aider v0.56.0 on the 133-task Exercism Python benchmark, using whole-file edit format and up to two attempts.

Run Pass rate 1 Pass rate 2
First 48.1% 56.4% (75/133)
Confirmation 46.6% 54.9% (73/133)

The two-run pass@2 mean was 55.64%. The public leaderboard comparisons above use the same old-aider benchmark family, but local runtime and harness details may differ. Results also showed task-level variance, so the single-run high score should not be treated as a stable estimate across other prompts, inference settings, or benchmarks.

Usage

Ollama

Download this repository, then create the model with the included Modelfile:

ollama create lore -f Modelfile
ollama run lore

llama.cpp

llama-cli \
  -m qwen2.5-coder-7b-instruct.Q8_0.gguf \
  -cnv \
  -p "Write a Python function that checks whether a string is a palindrome."

Limitations

This model is specialized for Python coding-agent workflows and complete-file responses. It may produce incorrect, insecure, or incomplete code. Review and test generated code before use. The GGUF is quantized and may not exactly match the unquantized adapter's behavior.

Downloads last month
6
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for mraleko/lore

Base model

Qwen/Qwen2.5-7B
Quantized
(223)
this model