Instructions to use yusifnuri/Llama-3.2-3B-Instruct_code_generation with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use yusifnuri/Llama-3.2-3B-Instruct_code_generation with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-3B-Instruct") model = PeftModel.from_pretrained(base_model, "yusifnuri/Llama-3.2-3B-Instruct_code_generation") - Notebooks
- Google Colab
- Kaggle
Llama-3.2-3B-Instruct โ Code generation adapter
A LoRA adapter that specialises meta-llama/Llama-3.2-3B-Instruct (3.21 B parameters) for a single enterprise task: it completes a Python function so that it passes the reference unit tests.
It was produced for the MSc thesis Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models (SRH University Hamburg), which measures fine-tuned small models against frontier provider APIs on accuracy, latency, cost, privacy exposure and return-on-investment breakeven volume. The adapter is released so that the benchmark can be independently verified.
Read this before using the adapter
- The Llama 3.2 Community Licence permits commercial use but conditions it on attribution, a naming convention for derivative models, and a monthly-active-user eligibility threshold. Check it before adopting.
Measured performance
| Metric | Value |
|---|---|
| pass@1 | 0.3341 |
| Mean latency, batch 1 | 5769 ms |
| Cost per 1M generated tokens | USD 24.98 |
Measured on a single NVIDIA H200 (141 GB) at batch size one and full utilisation, priced at an imputed USD 3.99 per GPU-hour. Latency excludes network transit. Scores are not comparable across tasks โ each task carries its own metric. Evaluation ran on 5 July 2026; the complete matrix is at results/benchmark_matrix.csv.
Training
| Method | LoRA |
| Dataset | HumanEval (openai/openai_humaneval) |
| Dataset licence | MIT |
| Training examples | 5,000 (500 held out for checkpoint selection) |
| Rank / alpha / dropout | 16 / 32 / 0.05 |
| Target modules | q_proj, k_proj, v_proj, o_proj |
| Learning rate | 2e-4, cosine schedule, 3% warmup |
| Epochs | 3 |
| Effective batch size | 16 (4 x 4 gradient accumulation) |
| Max sequence length | 512 tokens |
| Optimiser | AdamW |
| Seed | 42 |
Hyperparameters were held constant across every model and task rather than tuned per cell, so these figures are a conservative lower bound on attainable performance.
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.2-3B-Instruct", device_map="auto")
model = PeftModel.from_pretrained(base, "<your-hf-username>/Llama-3.2-3B-Instruct_code_generation")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.2-3B-Instruct")
The adapter was trained on this prompt format and expects it at inference:
Complete the following Python function:
{text}
Limitations
- Trained once, with a single seed. Reported differences confound model quality with initialisation variance.
- Specialised to one task on one public corpus. It is not a general-purpose assistant and should not be treated as one.
- The evaluation corpora are long-standing public benchmarks and are plausibly present in the base model's pretraining data, which inflates absolute scores.
- Evaluation used 200 held-out instances (all 164 problems for code generation), so detectable effect sizes are bounded at roughly ten percentage points.
Links
- Code, configurations and evaluation harness: https://github.com/Yusifnuri/slm-benchmark
- Full benchmark matrix:
results/benchmark_matrix.csv - Per-request cost analysis:
results/cost_per_request.csv
Citation
@mastersthesis{nuri2026finetune,
title = {Fine-Tune or Pay Per Token? An Enterprise Benchmark of Small Language Models},
author = {Nuri, Yusif},
school = {SRH University Hamburg},
year = {2026}
}
- Downloads last month
- -
Model tree for yusifnuri/Llama-3.2-3B-Instruct_code_generation
Base model
meta-llama/Llama-3.2-3B-Instruct