YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Mini Code Generator V6
Frozen from-scratch Python code-generation experiment.
Model
- Parameters: 10,845,312
- Transformer layers: 6
- Hidden size: 384
- Attention heads: 6
- FFN size: 1536
- Context: 256
- Vocabulary: 259 UTF-8 byte tokens
- Tied input/output embeddings
- Training from scratch
- No pretrained model
- No pretrained tokenizer
Dataset
- 652 validated implementation records
- 30 algorithm families
- Family-disjoint train/validation/test split
- 520 train records
- 66 validation records
- 66 test records
- All generated training implementations were execution-validated
Training
- 40 epochs
- Best checkpoint: epoch 14
- Best validation loss: 0.643533
- Final epoch train loss: approximately 0.0964
- Final epoch validation loss: approximately 0.7799
Functional evaluation
Greedy generation:
- Valid candidates: 18/66
- Hidden tests executed: 40
- Hidden tests passed: 4
- Executed accuracy: 10.00%
- Fully correct: 0/18
Best-of-5 sampling:
- Candidates: 330
- Syntax-valid candidates: 46/330
- Prompts with executable candidate: 33/66
- Hidden tests executed: 230
- Hidden tests passed: 12
- Executed accuracy: 5.22%
- Fully correct prompts: 0/66
Status
V6 is a frozen experimental checkpoint.
The experiment demonstrates that substantially improved token-level validation performance did not translate into reliable autoregressive functional code generation for this small full-program causal language-model objective.
V6 is preserved for reproducibility and comparison with later experiments.
Important comparison note
V6 is not a strict dataset-only ablation against V5.
V5 used a code-completion objective, while V6 used full-program causal language modeling. Therefore differences between V5 and V6 cannot be attributed solely to dataset diversity.