--- license: apache-2.0 tags: - nanoGPT - code-generation - fine-tuned base_model: Ananda100/pocketcoder100M-distilled --- # 100m-sftd-python Full-parameter fine-tune of [`Ananda100/pocketcoder100M-distilled`](https://huggingface.co/Ananda100/pocketcoder100M-distilled) on a custom Python coding-problem dataset, using the `deepseek-ai/deepseek-coder-6.7b-base` tokenizer. - Fine-tune tokens processed: 356,253,696 - Best validation loss: 0.9270 - Mbpp pass@1: 7.6% - Block size: 512 - Vocab size: 32022 ## Loading This is a **custom architecture**, not a native `transformers` model, so `AutoModel.from_pretrained` won't work out of the box: ```python import json, torch from safetensors.torch import load_file with open("config.json") as f: cfg = json.load(f) config = GPTConfig(**cfg) model = GPT(config) state_dict = load_file("model.safetensors") model.load_state_dict(state_dict) model.eval() ```