Autoresearch checkpoint
This repository contains the best saved checkpoint produced during the 2026-09-24 RunPod experiment window for the autoresearch project.
Relationship to nanochat
This is an autoresearch-trained model, not a new nanochat implementation.
Autoresearch supplied the training loop, experiment changes, checkpoint, and
custom tokenizer. Nanochat revision e85db6b supplied the compatible GPT
architecture reference and native checkpoint layout used to export and load
the model. The model was therefore packaged for nanochat compatibility; it did
not build or replace nanochat itself.
Model
- Architecture: compact GPT-style language model
- Parameters: 50.3M
- Layers: 8
- Hidden size: 512
- Attention: 4 heads, 4 key/value heads
- Context length: 2,048 tokens
- Vocabulary: 8,192 tokens
- Attention window pattern:
SSSL - Training hardware: NVIDIA H100 80GB
- Training budget: 300 seconds
Evaluation
The uploaded checkpoint's recorded validation score is:
val_bpb: 1.012760
Lower val_bpb is better. The experiment log also contains the runtime,
throughput, MFU, and memory measurements.
The broader experiment search recorded a lower benchmark score of 1.004616
for commit 65adfe4; that run was completed before checkpoint saving was added,
so its weights are not the artifact uploaded here.
Files
best.pt: PyTorch checkpoint containing model weights and metadatametadata.json: training configuration and measured metricsnanochat/base_checkpoints/autoresearch/model_000000.pt: nanochat-native model weightsnanochat/base_checkpoints/autoresearch/meta_000000.json: nanochat-native metadata
The checkpoint is intended for research and reproducibility. Load it with
torch.load(..., weights_only=False) and use the matching project train.py
model definition to reconstruct the model.
The native nanochat export matches the earlier nanochat architecture revision
e85db6b. The exact custom 8,192-token tokenizer is available in
tokenizer/tokenizer.pkl, with tokenizer/token_bytes.pt for BPB evaluation.
It was regenerated by prepare.py and verified at vocabulary size 8,192.
The pinned loader strictly loaded all 50,332,176 exported parameters on an A40.
A deterministic base-model completion probe is recorded in
runs/chat_eval_2026-09-25.log; the repetitive output is expected because this
is a pretraining-only checkpoint, not an SFT/chat model. The pinned nanochat
CORE evaluator produced a bounded diagnostic CORE score of 0.058873 across
all 22 tasks, using 100 examples per task. The audit log is in
runs/core_eval_2026-09-25.log. This is not directly comparable to the
official nanochat leaderboard, which uses full task data and multi-GPU runs.
Comparison with the earlier nanochat build
There is currently no evidence that this model is better than the earlier nanochat build. The numbers are not an apples-to-apples comparison:
- This checkpoint is a 50.3M-parameter, 300-second autoresearch experiment
with recorded
val_bpb=1.012760(1.004616was a better run whose weights were not saved). - The recorded CORE value here (
0.058873) uses 100 examples per task and a single B300 MIG GPU; it is explicitly a diagnostic, not an official nanochat leaderboard score. - Nanochat leaderboard results use different training runs, datasets, evaluation scope, and multi-GPU hardware.
The only supported conclusion is that the autoresearch checkpoint is nanochat-architecture-compatible and reproducibly loadable. A claim that it outperforms the earlier nanochat build requires matching model size, data, training budget, tokenizer, and full evaluation protocol.
SFT derivative
The repository also contains a separate tokenizer-preserving SFT derivative at
sft/autoresearch-sft.pt. It was trained for 500 steps on 256 streamed
conversations from HuggingFaceH4/ultrachat_200k (train_sft), with batch size
2, learning rate 2e-5, and 1,041,840 supervised tokens. It retains the same
8,192-token vocabulary and base architecture.
The deterministic plain-text chat probe is recorded in
runs/sft_chat_eval_2026-09-25.json. The SFT model is more task-directed than
the base probe but remains repetitive; this is a small SFT demonstration, not
a claim of production conversational quality.
License
See the source project for licensing and dataset terms.