YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Milestone 2 Model and Training Artefact Delivery
This English-language package consolidates nine historical model checkpoints, the core training and data-preparation code, launch scripts, provenance evidence, datasets already selected for delivery, and integrity tooling. Four additional SFT checkpoints (S1--S4) were selected for this model-only increment.
Delivered Models
| Model | Path | Initialisation | Training record |
|---|---|---|---|
| CodeGPT multilingual CPT | models/pretrained/codegpt_multilingual_5epoch/ |
AISE-TUDelft/CodeGPT-Multilingual, revision f373547d1234992a64cd08d64cd69abbb3a83579 |
5 epochs, 222,615 steps, 8,753,591,494 training tokens |
| GPT-2 124M CPT | models/pretrained/gpt2_124m_code_5epoch/ |
openai-community/gpt2, revision 607a30d783dfa663caf39e06633721c8d4cfcd7e |
5 epochs, 222,615 steps, 8,753,591,494 training tokens |
| Qwen2.5-Coder-7B CPT | models/pretrained/qwen25_coder_7b_cpt/ |
Qwen/Qwen2.5-Coder-7B, revision 0396a76181e127dfc13e5c5ec48a8cee09938b02 |
full-parameter CPT, 3 epochs, 18,165/18,165 steps |
| Qwen2.5-Coder-0.5B Stage-1 CPT | models/pretrained/stage1_cpt/ |
Qwen/Qwen2.5-Coder-0.5B |
continued-pre-training deliverable |
| Qwen2.5-Coder-0.5B-Instruct F3 SFT | models/sft/sft_f3_refined_instruct/ |
Qwen/Qwen2.5-Coder-0.5B-Instruct |
reported pass@1 9.1%; pass@5 13.0% |
| S1: Qwen2.5-Coder-0.5B-Instruct, F3 clean | models/sft/sft_f3_clean_instruct/ |
Qwen/Qwen2.5-Coder-0.5B-Instruct |
473,465 records; step 14,550; reported pass@1 9.1%, pass@5 13.0% |
| S2: Qwen2.5-Coder-0.5B-Instruct, F3 high-quality v4 | models/sft/sft_f3_hq_instruct/ |
Qwen/Qwen2.5-Coder-0.5B-Instruct |
134,775 records; step 4,212; reported pass@1 8.3%, pass@5 13.0% |
| S3: Qwen2.5-Coder-0.5B-Instruct, F3+F4 clean | models/sft/sft_f3f4_instruct/ |
Qwen/Qwen2.5-Coder-0.5B-Instruct |
491,004 records; step 15,033; reported pass@1 5.7%, pass@5 11.5% |
| S4: Qwen2.5-Coder-0.5B-Instruct, F4 clean | models/sft/sft_f4_instruct/ |
Qwen/Qwen2.5-Coder-0.5B-Instruct |
17,539 records; step 484; reported pass@1 0.9%, pass@5 2.5% |
All nine directories are standard Hugging Face checkpoints. The 7B selection is the completed full-parameter continued-pre-training checkpoint, not the separate historical 7B LoRA instruction-tuning experiment.
The S1--S4 identifiers, exact source checkpoints, and evaluation results are catalogued
in models/sft/README.md. In particular, S3 used 491,004 records (473,465 F3 records plus
17,539 F4 records); this corrects the approximate 150K count in an earlier historical
report. The checkpoint arguments, launch record, and measured input files support the
491,004-record count.
Model weights, tokenisers, and available effective-argument records are included. Optimizer, scheduler, and random-number-generator states are excluded because this is a model and training-recipe delivery, not a mid-run resume package.
Training Code and Scripts
The GPT-2-family runs used the llm.c code snapshot at Git commit
f1e2ace651495b74ae22d45d1723443fd00ecd3a. The required CUDA trainer, headers,
Makefile, dataset converter, and Hugging Face import/export utilities are preserved in
code/llmc_training/. Exact historical launch scripts are in training_scripts/pretrain/.
The Qwen runs used ms-swift. The 7B run-specific command, package snapshot, system record,
and complete training log are in evidence/training_runs/qwen25_coder_7b_cpt/. The smaller
May 2026 runs did not capture run-specific package snapshots; this limitation is documented
in environment/KNOWN_LIMITATIONS.md.
This S1--S4 increment contains the selected model checkpoints and their provenance catalogue. Additional S2--S4 datasets and variant-specific launch scripts remain in the local candidate staging area and are not part of this upload.
Historical scripts intentionally retain the absolute paths, GPU assignments, Docker image, ports, and cache locations used at training time. Replace these machine-specific paths before rerunning.
Training-Data Dependency
CodeGPT, GPT-2, and Qwen2.5-Coder-7B were trained from the same historical source file:
/raid/data/weifeng/Datasets/code_messages_pre_train.jsonl
2,239,565 records; 20,021,101,140 bytes
SHA-256: 74e8b9111dfca80324ade8da530858584cc3833d7650688989eedd1cb6d0f25d
That 20 GB raw source is not duplicated in this model-and-code increment. The GPT-2-family
tokenisation and deterministic train/validation split implementation is included at
code/llmc_training/dev/data/code_messages.py. The smaller filtered Stage-1 dataset and
the complete F3 chain already remain under data/.
F3 Filtering and Evaluation Evidence
data/f3/ contains the input, principal intermediate artefacts, score output, and final
207,009-record F3 training dataset. Its corresponding code is in code/f3_filtering/.
Stage 4 used an external Qwen3-Coder-30B-A3B-Instruct scoring model; its saved scores and
final dataset are included, while its original absolute model path must be replaced on a
new system.
evidence/eval/f3_refined_9p1_summary.json records the historical Runnable alignment
evaluation: 200 tasks, five samples per task, 1,000 attempts, pass@1 of 9.1%, and pass@5 of
13.0%. This is not the later repository-level Core/Hold benchmark.
Licensing and External Distribution
See THIRD_PARTY_NOTICES.md. The official GPT-2 metadata declares MIT and the Qwen2.5
Coder model declares Apache-2.0. The official CodeGPT-Multilingual repository declares no
model licence; its derivative checkpoint therefore requires explicit rights confirmation
before distribution outside the authorised delivery channel.
Integrity and Validation
MANIFEST.jsonmaps every delivered file to its source and SHA-256 digest.SHA256SUMSsupports byte-level verification after transfer.tools/build_manifest.pyregenerates both integrity records.tools/smoke_test.pyloads the delivered checkpoints and checks finite, non-degenerate logits.
Hard-linked files are a local storage optimisation only. Archives and Hugging Face uploads read the full file contents, so recipients receive ordinary complete files.