AutoResearch-blip3o-long-caption-depth8
AutoResearch-blip3o-long-caption-depth8 is a 98.2M parameter decoder-only Transformer trained from scratch on blip3o-long-caption.
This model is part of the AutoResearch project, which focuses on training, evaluating, and releasing efficient language models with reproducible research workflows.
Overview
This is a 8-layer decoder-only Transformer trained on the blip3o-long-caption dataset for 0.3 hours of wall-clock training time. The model achieves a validation bits-per-byte (val_bpb) of 0.858463 (perplexity: 1.8131) on the held-out validation set.
References
Papers
- NanoGPT / NanoChat architecture patterns
Datasets
Related Projects
WANDB Run
Highlights
- Trained from scratch
- 98.2M parameters
- Trained on 33.6M tokens (64 steps)
- 8-layer decoder-only Transformer with sliding window attention
- RoPE positional encoding, RMSNorm, ReLUยฒ activation
- MuonAdamW optimizer (Muon for matrices, AdamW for embeddings)
- Hugging Face Transformers compatible
Model Architecture
| Property | Value |
|---|---|
| Architecture | Decoder-only Transformer |
| Parameters | 98,150,416 (98.2M) |
| Layers | 8 |
| Hidden Size | 512 |
| Attention Heads | 4 |
| KV Heads | 4 |
| Head Dimension | 128 |
| Feed Forward Size | 2048 |
| Context Length | 2048 |
| Vocabulary Size | 16,384 |
| Positional Encoding | RoPE |
| Activation | ReLUยฒ |
| Normalization | RMSNorm |
| Window Pattern | SSSL |
| Weight Tying | No |
Training
This model was trained from scratch for 0.3 hours (915s) of wall-clock training time.
Training Configuration
| Setting | Value |
|---|---|
| Optimizer | MuonAdamW (Muon + AdamW) |
| Precision | torch.bfloat16 |
| Learning Rate | 0.04 (matrix) / 0.6 (embedding) |
| Weight Decay | 0.2 |
| Batch Size | 4 ร 2048 = 8,192 tokens/step |
| Gradient Accumulation | 64 steps |
| Total Batch Size | 524,288 tokens |
| Context Length | 2048 |
| Vocabulary | 16,384 tokens (BPE) |
| LR Scheduler | Linear warmdown (50%) |
| Activation Checkpointing | Enabled |
Hardware
- GPU: NVIDIA GeForce RTX 4060 Ti
- VRAM: 16.0 GB
- Peak VRAM Used: 3.2 GB
- MFU: 14.01%
- Framework: PyTorch 2.9.1+cu128
Dataset
- Name: blip3o-long-caption
- Language: English
Preprocessing
Data is packed into fixed-length sequences of 2048 tokens using the nanochat-compatible BPE tokenizer (16,384 vocabulary, 9 special tokens). No additional filtering or deduplication is applied beyond what is in the source dataset.
Intended Use
This model is intended for:
- Educational purposes and research
- Text generation experiments
- Studying small language model training dynamics
Not recommended for:
- Production use or safety-critical applications
- Tasks requiring factual accuracy
Evaluation
Results
| Metric | Score |
|---|---|
| Validation BPB | 0.858463 |
| Perplexity | 1.8131 |
| Peak VRAM | 3.2 GB |
| MFU | 14.01% |
Example Generations
Example 1
Prompt
Once upon a time,
Generation
Once upon a time, 195-12,000,19535, 620.5 6, 195621, is10, with a map and to the Lotusrologh as the mask, suggesting it might be part of an
Example 2
Prompt
A lonely dragon
Generation
A lonely dragon 40%. The background features a blurred blue background with dark lights, adding to the calm atmosphere of the scene. The overall composition suggests a moment of urban environment, possibly in a mountainous region known for its natural beauty and the presence of the water. The
Example 3
Prompt
The opposite of boy is
Generation
The opposite of boy is 0. In the background, there are buildings and a clear blue sky, suggesting a sunny day. The overall atmosphere is serene and peaceful, with the focus on the two figures dressed as a focal point, their hands adding to the spiritual ambiance of the
Example 4
Prompt
The opposite of queen is
Generation
The opposite of queen is 9, 1940. The design includes a robust pattern, and other use in shades, gold, blue, and green. The wear indicate a mix of aged and weathered facades, with some areas appearing slightly slightly than others,. The use of
Example 5
Prompt
My name is
Generation
My name is 1. The design includes additional decorative elements and decorative elements. The overall design is clean and contemporary, with a focus on technology and luxury. The overall aesthetic suggests a blend of functionality and functional aesthetic appeal. The craftsmanship suggests a modern model, possibly a
Usage
import torch
import pickle
import json
from train import GPT, GPTConfig, Tokenizer
# Load config
with open('config.json', 'r') as f:
config_dict = json.load(f)
config = GPTConfig(**{k: v for k, v in config_dict.items() if k in GPTConfig.__dataclass_fields__})
# Load model
model = GPT(config)
state_dict = torch.load('model.pt', map_location='cpu')['state_dict']
model.load_state_dict(state_dict)
model.eval()
# Load tokenizer
with open('tokenizer.pkl', 'rb') as f:
tokenizer = pickle.load(f)
# Generate
prompt = 'Once upon a time, '
input_ids = tokenizer.encode(prompt)
x = torch.tensor([input_ids], dtype=torch.long)
with torch.no_grad():
for _ in range(50):
logits = model(x)
probs = torch.softmax(logits[:, -1, :] / 0.8, dim=-1)
next_token = torch.multinomial(probs, num_samples=1)
input_ids.append(next_token.item())
x = torch.tensor([input_ids], dtype=torch.long)
print(tokenizer.decode(input_ids))
Repository Structure
model.pt # Model weights
config.json # Model architecture config
dataset.txt # Dataset name used for training
token_bytes.pt # Token byte mappings
tokenizer.pkl # Trained BPE tokenizer
tokenizer_config.json # Tokenizer configuration
training_metrics.json # Training metrics
README.md # This file
Limitations
- Small model size limits language understanding and coherence
- Trained on a single dataset (TinyStories) โ limited domain
- Fixed time budget training โ not fully trained to convergence
- No RLHF or safety alignment
Ethical Considerations
- This is a research artifact, not a production model
- The training data consists of synthetic stories (GPT-4 generated)
- No harmful content filtering was applied
- Intended for research and educational use only
Citation
@misc{autoresearch_blip3o-long-caption_depth8,
title={AutoResearch-blip3o-long-caption-depth8},
author={Dustin Loring},
year={2026},
howpublished={\url{https://huggingface.co/quik-models/easy-butterfly-71}}
}}
Version History
| Version | Date | Notes |
|---|---|---|
| v1.0 | 2026-07-28 | Initial release |
Acknowledgements
Built with the AutoResearch training framework.
Thanks to:
- Hugging Face
- PyTorch
- The creators of the TinyStories dataset
- The open-source AI research community
License
This model is released under the MIT License unless otherwise specified.
- Downloads last month
- -
