Repente v0.7

A language model that writes Pure Data patches and SuperCollider code, and runs locally. Qwen2.5-Coder-7B-Instruct fine-tuned with QLoRA, exported to GGUF at Q4_K_M.

This is the current release. A companion repository, bidubr/repente-v0.5-GGUF, holds the earlier v0.5, kept available because the experiments in the paper were run on it.

What it does

Give it a description of a sound and it writes the code that produces it. Ask it about an existing patch and it explains the signal flow.

make a sine wave at 440 Hz connected to output
#N canvas 0 0 450 300 12;
#X obj 100 100 osc~ 440;
#X obj 100 150 dac~;
#X connect 0 0 1 0;
#X connect 0 0 1 1;

That is a Pure Data file. Save it, open it, and it makes a sound.

How good it is, measured

Five Pure Data generation prompts, each sampled 30 times at temperature 0.7, scored by whether the output carries a canvas header, an output object, and connections.

Model Expected score out of 5 95% interval
Qwen2.5-Coder-7B, unmodified 0.03 [0.00, 0.10]
Repente v0.5 2.67 [2.37, 2.97]
Repente v0.7 3.83 [3.53, 4.13]

The base model emits a valid canvas header in 2.7% of attempts, yet names an output object in 68.7% of them, more often than v0.5 does. Its deficit is serialization, not intent, and that is the gap fine-tuning closes.

Version 0.7 is the strongest of the nine checkpoints trained, and is the only one the battery separates from v0.5 by disjoint intervals. Full protocol in Appendix C of the paper.

Running it

With Ollama, straight from this repository:

ollama run hf.co/bidubr/repente-v0.7-GGUF:Q4_K_M

With llama.cpp:

llama-server -m repente-v0.7-Q4_K_M.gguf -ngl 99 -c 4096

4.5 GB on disk. Fits in 8 GB of VRAM with room for a 4096-token context.

Prompting

The system prompt used throughout the reported experiments is short:

You are Repente, a musical programming expert.

Format validity improves substantially when three worked examples and a short chain-of-thought instruction are supplied together. Measured on frozen weights, the combination raises format validity from 75% to 100% across four prompt difficulty levels and eliminates output truncation, at no compute cost. Few-shot exemplars supplied alone introduce a context-leak failure mode in which the model continues the turn structure of the examples until the token limit; the reasoning instruction suppresses it. Details in Section 8 of the paper.

Known limitations

ELSE objects are not generated. Across 1,350 measured generations, an object from the ELSE library appears once. Five cycles of corpus weighting, including one that weighted ELSE explicitly, did not change this. Treat library coverage as a retrieval problem at inference time, not something the weights will supply.

Analysis responses are short. The unmodified base model produces analyses averaging 776 tokens; this model produces 151. The compression is a self-distillation artifact of using each version's own output to train the next, and it is documented rather than fixed.

Patch validity is not patch quality. The measurements above score a generation as passing when it carries a canvas header, an output object, and connections. That is necessary for a patch to make sound and not sufficient for it to make the right sound.

Related

Citation

@misc{repente2026,
  author       = {Batista, Carlos Eduardo Coelho Freire},
  title        = {Repente: A Specialized LLM for Musical Programming Languages
                  with SonicUnit Knowledge Architecture},
  year         = {2026},
  eprint       = {[ARXIV_ID]},
  archivePrefix= {arXiv},
  primaryClass = {cs.SD}
}

Licensed Apache 2.0, inherited from the base model.

Downloads last month
84
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for bidubr/repente-v0.7-GGUF

Base model

Qwen/Qwen2.5-7B
Finetuned
(442)
this model