mtron-qwen

Qwen models fine-tuned on the mtron programming language

http://metatron.phaseshift.studio


Overview

mtron-qwen is a family of Qwen models fine-tuned with QLoRA on the mtron functional programming language of the metatron vm. Each variant can evaluate mtron expressions, explain language concepts, and translate between mtron sugar operators and their desugared instruction forms.

Variants

Variant Base Model Params Size Files
mtron-qwen-4b Qwen3-4B 4B 2.5 GB mtron-qwen-4b.Q4_K_M.gguf
mtron-qwen-8b Qwen3-8B 8B 5.0 GB mtron-qwen-8b.Q4_K_M.gguf
mtron-qwen-14b Qwen3-14B 14B 8.4 GB mtron-qwen-14b.Q4_K_M.gguf

Training

All variants share the same training methodology:

Parameter Value
Method QLoRA (bitsandbytes 4-bit NF4)
LoRA rank (r) 32
LoRA alpha 8
Optimizer AdamW 8-bit
Scheduler Cosine with warmup
Dataset mtron expression evaluation pairs with operator documentation (2,660 entries)
Hardware 2ร— NVIDIA RTX 3090 (48 GB)

Per-Variant Training Details

Metric mtron-qwen-4b mtron-qwen-8b mtron-qwen-14b
Base model Qwen3-4B Qwen3-8B Qwen3-14B
Training steps 600 600 600
Batch size (effective) 8 8 8
Initial loss 3.81 3.15 3.24
Best loss 0.27 0.24 0.23
Final loss 0.81 0.76 0.69
Training time ~15 min ~20 min 31 min

Training Plots

Qwen3-4B

4B training

Qwen3-8B

8B training

Qwen3-14B

14B training

Usage

Ollama

Create a Modelfile (example for 14B variant):

FROM ./mtron-qwen-14b.Q4_K_M.gguf

TEMPLATE """<|im_start|>system
{{ .System }}<|im_end|>
<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
{{ .Response }}<|im_end|>"""

PARAMETER temperature 0.7
PARAMETER stop "<|im_end|>"

Then register and run:

ollama create mtron-qwen-14b -f Modelfile
ollama run mtron-qwen-14b

Prompt Format (ChatML)

<|im_start|>system
You are an expert in the mtron functional programming language.
Evaluate the given mtron expression and return the result.<|im_end|>
<|im_start|>user
/m/str/"hello" /m/str/plus(" world")<|im_end|>
<|im_start|>assistant
"hello world"<|im_end|>

mtron Language

mtron is a data-oriented functional language for the Metatron VM. Expressions follow a structural navigation pattern using URI-addressed spaces and instruction-based evaluation.

License

AGPL-3.0

Downloads last month
-
GGUF
Model size
15B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for phaseshift-studio/mtron-qwen

Finetuned
Qwen/Qwen3-14B
Adapter
(774)
this model