Autoresearch checkpoint

This repository contains the best saved checkpoint produced during the 2026-09-24 RunPod experiment window for the autoresearch project.

Relationship to nanochat

This is an autoresearch-trained model, not a new nanochat implementation. Autoresearch supplied the training loop, experiment changes, checkpoint, and custom tokenizer. Nanochat revision e85db6b supplied the compatible GPT architecture reference and native checkpoint layout used to export and load the model. The model was therefore packaged for nanochat compatibility; it did not build or replace nanochat itself.

Model

  • Architecture: compact GPT-style language model
  • Parameters: 50.3M
  • Layers: 8
  • Hidden size: 512
  • Attention: 4 heads, 4 key/value heads
  • Context length: 2,048 tokens
  • Vocabulary: 8,192 tokens
  • Attention window pattern: SSSL
  • Training hardware: NVIDIA H100 80GB
  • Training budget: 300 seconds

Evaluation

The uploaded checkpoint's recorded validation score is:

val_bpb: 1.012760

Lower val_bpb is better. The experiment log also contains the runtime, throughput, MFU, and memory measurements.

The broader experiment search recorded a lower benchmark score of 1.004616 for commit 65adfe4; that run was completed before checkpoint saving was added, so its weights are not the artifact uploaded here.

Files

  • best.pt: PyTorch checkpoint containing model weights and metadata
  • metadata.json: training configuration and measured metrics
  • nanochat/base_checkpoints/autoresearch/model_000000.pt: nanochat-native model weights
  • nanochat/base_checkpoints/autoresearch/meta_000000.json: nanochat-native metadata

The checkpoint is intended for research and reproducibility. Load it with torch.load(..., weights_only=False) and use the matching project train.py model definition to reconstruct the model.

The native nanochat export matches the earlier nanochat architecture revision e85db6b. The exact custom 8,192-token tokenizer is available in tokenizer/tokenizer.pkl, with tokenizer/token_bytes.pt for BPB evaluation. It was regenerated by prepare.py and verified at vocabulary size 8,192.

The pinned loader strictly loaded all 50,332,176 exported parameters on an A40. A deterministic base-model completion probe is recorded in runs/chat_eval_2026-09-25.log; the repetitive output is expected because this is a pretraining-only checkpoint, not an SFT/chat model. The pinned nanochat CORE evaluator produced a bounded diagnostic CORE score of 0.058873 across all 22 tasks, using 100 examples per task. The audit log is in runs/core_eval_2026-09-25.log. This is not directly comparable to the official nanochat leaderboard, which uses full task data and multi-GPU runs.

Comparison with the earlier nanochat build

There is currently no evidence that this model is better than the earlier nanochat build. The numbers are not an apples-to-apples comparison:

  • This checkpoint is a 50.3M-parameter, 300-second autoresearch experiment with recorded val_bpb=1.012760 (1.004616 was a better run whose weights were not saved).
  • The recorded CORE value here (0.058873) uses 100 examples per task and a single B300 MIG GPU; it is explicitly a diagnostic, not an official nanochat leaderboard score.
  • Nanochat leaderboard results use different training runs, datasets, evaluation scope, and multi-GPU hardware.

The only supported conclusion is that the autoresearch checkpoint is nanochat-architecture-compatible and reproducibly loadable. A claim that it outperforms the earlier nanochat build requires matching model size, data, training budget, tokenizer, and full evaluation protocol.

SFT derivative

The repository also contains a separate tokenizer-preserving SFT derivative at sft/autoresearch-sft.pt. It was trained for 500 steps on 256 streamed conversations from HuggingFaceH4/ultrachat_200k (train_sft), with batch size 2, learning rate 2e-5, and 1,041,840 supervised tokens. It retains the same 8,192-token vocabulary and base architecture.

The deterministic plain-text chat probe is recorded in runs/sft_chat_eval_2026-09-25.json. The SFT model is more task-directed than the base probe but remains repetitive; this is a small SFT demonstration, not a claim of production conversational quality.

License

See the source project for licensing and dataset terms.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support