FL-7B-3: COBOL Code Generation

A 7B model that writes COBOL that actually compiles.

Ask any general coding model for COBOL and you get confident nonsense. The base model here, Qwen/Qwen2.5-Coder-7B, scores a flat zero on COBOLEval. Not "low", zero. Only 3.65% of its test programs compile at all.

FL-7B-3 is a supervised finetune of that same base on a curated COBOL instruction corpus. It solves 15.75% of COBOLEval, matching published GPT-4 results with a model roughly two orders of magnitude smaller, and more than doubling GPT-4's compilation rate.

Benchmarks

COBOLEval (zorse-project/COBOLEval), 146 problems / 821 test cases. Every sample is compiled and executed with GnuCOBOL 3.2.0. No self-reported or LLM-judged scores. Greedy decoding, repetition_penalty=1.0, max_new_tokens=1536.

Model pass@1 Compile rate Tests passed
Qwen2.5-Coder-7B (base) 0.00% 3.65% 4 / 821
FL-7B-3 15.75% 51.16% 204 / 821

Published reference numbers on the same benchmark:

Model pass@1 Compile rate
GPT-4 15.75% 24.12%
GPT-4o 16.40% 41.80%
FL-7B-3 (7B) 15.75% 51.16%

The headline: COBOL goes from unusable to useful. Compilation rate is where the gap is widest: FL-7B-3 produces syntactically valid COBOL roughly twice as often as GPT-4 and 25% more often than GPT-4o.

⚠️ Required inference setting

Set repetition_penalty to exactly 1.0. This is not a stylistic preference. It costs you real accuracy:

repetition_penalty pass@1
1.0 15.75%
1.05 (common default) 13.01%
1.15 2.74%

COBOL mandates repetition. PROGRAM-ID must match END PROGRAM character for character, data names recur constantly, division headers are fixed boilerplate. Any repetition penalty pushes the model away from re-emitting tokens the language requires it to re-emit, and the program stops compiling. At 1.15 the model produces identifiers like TESTTRUNCATENUMBER where TRUNCATE-NUMBER was required.

Usage

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("FLs-AI/FL-7B-3-safetensors")
model = AutoModelForCausalLM.from_pretrained(
    "FLs-AI/FL-7B-3-safetensors", device_map="auto", torch_dtype="bfloat16")

messages = [{"role": "user", "content": "Write a COBOL program that ..."}]
prompt = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

out = model.generate(
    **tok(prompt, return_tensors="pt").to(model.device),
    max_new_tokens=1536,
    do_sample=False,
    repetition_penalty=1.0,   # <- required, see above
)
print(tok.decode(out[0], skip_special_tokens=True))

GGUF / llama.cpp

llama-cli -hf FLs-AI/FL-7B-3:Q4_K_M --repeat-penalty 1.0 -n 1536
Quant Size Notes
Q8_0 ~8 GB Effectively lossless
Q6_K ~6 GB Recommended if you have the RAM
Q4_K_M ~4.5 GB Recommended default, realistic quality floor
Q2_K ~3 GB Experimental. 2-bit on a 7B degrades badly; not benchmarked

Quantized variants were not re-benchmarked. The reported 15.75% is bf16.

Training

Method LoRA SFT, assistant-only loss masking
Base Qwen/Qwen2.5-Coder-7B @ 0396a761
LoRA r=32, α=64, dropout=0.0, 196 modules (q/k/v/o + gate/up/down_proj)
Trainable 80.7M / 7.70B (1.05%)
Precision bf16, gradient checkpointing
Data 20,332 instruction pairs → 21.2M tokens, packed into 2,586 × 8192-token blocks
Schedule 2 epochs, 324 steps, 131,072 tokens/step
Optimizer adamw_8bit, LR 1e-4 cosine → 0, warmup 3%, wd 0.01, clip 1.0
Hardware 1× H100 80GB SXM, 2h03m
Eval loss 0.557 → 0.4183

On epoch count: epoch 1 moved eval loss by −0.126, epoch 2 by only −0.013. One epoch captures the overwhelming majority of the gain on a corpus this size. Validation loss decreased monotonically across all 64 evaluations with no overfitting inflection.

Limitations

Honest failure analysis. 65 of 146 COBOLEval solutions fail to compile standalone. The errors are a long tail, not one systematic bug:

  • Degenerate generation on hard problems. The model can enter repetition loops, appending -TEMP to an identifier until it exceeds COBOL's 63-character limit, or enumerating IF WS-NUMBER = 58 ... IF WS-NUMBER = 59 ... instead of writing a loop. Correct solutions are consistently short (median 1,368 characters, max 3,248); output much longer than that is a strong signal the generation has derailed.
  • Enumeration over algorithm. On problems requiring real logic (primality, parsing), it sometimes hardcodes cases rather than implementing the algorithm.
  • Occasional undefined identifiers, e.g. using I or a linkage record without declaring it.
  • Structural breakage on long outputs: missing PROCEDURE DIVISION header, unbalanced parentheses, mismatched END PROGRAM.

Scope: trained on COBOL only. General coding ability inherited from the base was not re-measured and may have regressed. Not evaluated on real mainframe dialects (IBM Enterprise COBOL, CICS, JCL, DB2 embedded SQL). GnuCOBOL 3.2 only.

Do not deploy generated COBOL to production systems without review. A 15.75% pass@1 means roughly five out of six generated programs are wrong.

License

The finetune is released under CC BY-NC 4.0 (non-commercial, attribution required).

The base model, Qwen/Qwen2.5-Coder-7B, is Apache-2.0 and remains so. The non-commercial restriction applies to this finetune's contribution, and cannot and does not revoke any rights you already hold in the base weights under Apache-2.0. If your use case is commercial, you can still use the base model freely; you just cannot use these weights.

💡 Note on FL-3.1: This model version (FL-7B-3) is designed specifically for COBOL code generation. We are actively training FL-3.1, which expands the architecture into a full-fledged Mainframe Assistant (adding JCL, CICS, DB2, and general mainframe architectural guidance).

Citation

@misc{fl7b3,
  title  = {FL-7B-3: COBOL Code Generation},
  author = {FLs-AI},
  year   = {2026},
  url    = {https://huggingface.co/FLs-AI/FL-7B-3}
}
Downloads last month
26
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

2-bit

4-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FLs-AI/FL-7B-3

Base model

Qwen/Qwen2.5-7B
Finetuned
(115)
this model

Collection including FLs-AI/FL-7B-3