SemanticWiki Coder 14B v3

LoRA adapter fine-tuned from Qwen/Qwen2.5-Coder-14B-Instruct to generate DeepWiki-style architectural wiki pages with verifiable path:line source citations, Mermaid diagrams, and structured sections.

What it does

Given a line-numbered code context and a query, it writes a wiki page whose factual claims carry strict path/file.ext:line or path:start-end citations that resolve against the real repository.

Input format (same as training):

<START_OF_CONTEXT>
# Repository: org/name (language: ...)
# File tree ...
### FILE: src/module.py (lines 1-N)
   1| <source with line numbers>
<END_OF_CONTEXT>

<query>
Generate a wiki page describing ...
</query>

Training

Parameter Value
Base model Qwen/Qwen2.5-Coder-14B-Instruct
Method LoRA SFT (r=64, alpha=128, dropout 0.05, all linear projections)
Data GhostScientist/semanticwiki-data-v3 - 313 pages from 97 real GitHub repos, every citation deterministically verified
Loss completion-only (prompt-completion format; the code context is not trained on)
Epochs / LR / effective batch 2 / 1e-4 cosine / 16
Max sequence length 24576
Final loss / token accuracy 0.5101 / 0.823

Paired eval (SemanticWiki-Eval v3, 13 held-out repos, 26 pages)

Arm Format Citation validity Fidelity (judge 1-5) Citations/page
this model 0.812 0.562 3.39 4.9
Qwen2.5-Coder-14B-Instruct (base) 0.755 0.308 3.15 2.0

Fine-tuning improved every axis; citation validity rose +83% relative. Full protocol and per-page results: GhostScientist/semanticwiki-eval-v3.

Limitations

Citation validity on unseen repos is 0.56 vs the teacher dataset's ~1.0 - hallucinated line numbers remain the main failure mode. Fidelity is judge-based (judge shares the Qwen3 family; see benchmark card caveats).

Framework versions

TRL 1.14.1 - Transformers 5.18.0 - PyTorch 2.14.1 - PEFT 0.21.2 - Datasets 5.1.0

Citation

@software{vonwerra2020trl,
  title   = {{TRL: Transformers Reinforcement Learning}},
  author  = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallou{\'e}dec, Quentin},
  license = {Apache-2.0},
  url     = {https://github.com/huggingface/trl},
  year    = {2020}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GhostScientist/semanticwiki-coder-14b-v3

Base model

Qwen/Qwen2.5-14B
Finetuned
(127)
this model