SQLCoder โ€” Text2SQL on a small LLM

Fine-tune a small open LLM (Qwen2.5-Coder-1.5B, โ‰ค 3B) to turn natural-language questions into executable SQLite queries โ€” the engine for a fintech chatbot that lets non-technical teams pull data without writing SQL.

The whole project is built around free resources: training on a Google Colab T4, inference on an ordinary laptop CPU via a quantized GGUF served through Ollama โ€” no GPU required to run it. A LoRA adapter (~1% of parameters trained) sits on top of the frozen base model and is exported to a ~1 GB GGUF for local use.

  • What it does: given a database schema (DDL) + a question, it returns one runnable SQLite query.
  • What it's for: a data-access chatbot โ€” plain English in, executable SQL out.

Result of the fine-tune

Evaluated on 200 held-out test examples, greedy decoding, identical prompts for both models. The numbers below are the final max_new_tokens=512 run (Result/*_preds_512.json).

Model Valid rate Executable rate
Baseline (no fine-tune) 99.0% 42.5%
Fine-tuned (LoRA, 512-tok) 99.5% 75.5%
Gold queries (ceiling) 100.0% 99.5%

Fine-tuning raised the executable-query rate from 42.5% โ†’ 75.5% (+33 pp absolute, +78% relative), recovering roughly half the gap to the gold ceiling.

Using the model (Ollama)

The published artifacts on the Hugging Face Hub:

Artifact Repo Use
Quantized GGUF SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf local CPU inference
LoRA adapter SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-lora GPU / further training

Fastest path โ€” pull straight from the Hub

Ollama downloads the GGUF for you, no manual steps:

ollama run hf.co/SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf

Recommended โ€” build with the baked-in system prompt

This applies the Text2SQL system prompt and temperature 0 from report/Modelfile, so you get deterministic, prompt-correct output:

# 1. download just the GGUF (~1 GB)
huggingface-cli download SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf \
    --include "*.gguf" --local-dir ./gguf

# 2. build a local Ollama model from the Modelfile
ollama create t2sql -f report/Modelfile

# 3. run it
ollama run t2sql

You then also get an OpenAI-compatible HTTP API on localhost:11434. Prompt it with the schema DDL followed by the question (same order used in training).

Compute requirements

To run it (inference) โ€” no GPU needed:

Resource Minimum Comfortable
RAM 4 GB free 8 GB
Disk 2 GB 5 GB
CPU any x86-64 with AVX2, 2 cores 4โ€“8 cores (Apple Silicon works natively)
GPU none optional (llama.cpp offloads layers if present)
Downloads last month
360
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf

Quantized
(150)
this model

Dataset used to train SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf