Instructions to use AIMadeSimpleResearch/AIMS-1-128M-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AIMadeSimpleResearch/AIMS-1-128M-base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AIMadeSimpleResearch/AIMS-1-128M-base", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AIMadeSimpleResearch/AIMS-1-128M-base", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AIMadeSimpleResearch/AIMS-1-128M-base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AIMadeSimpleResearch/AIMS-1-128M-base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AIMadeSimpleResearch/AIMS-1-128M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AIMadeSimpleResearch/AIMS-1-128M-base
- SGLang
How to use AIMadeSimpleResearch/AIMS-1-128M-base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AIMadeSimpleResearch/AIMS-1-128M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AIMadeSimpleResearch/AIMS-1-128M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AIMadeSimpleResearch/AIMS-1-128M-base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AIMadeSimpleResearch/AIMS-1-128M-base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AIMadeSimpleResearch/AIMS-1-128M-base with Docker Model Runner:
docker model run hf.co/AIMadeSimpleResearch/AIMS-1-128M-base
AIMS-1-128M-base
AIMS-1-128M-base is a 128M-parameter decoder-only language model trained from scratch on the BabyLM-2026-Strict dataset.
The model is part of the AIMS-1 project, an experiment in training a language model with an approximately $200 compute budget. The primary objective is to investigate how effectively a small language model can learn the syntax and structural patterns of English under a strict data and compute constraint.
AIMS-1-128M-base is the pretrained base model. It has not been supervised fine-tuned for instruction following or conversational interaction. The later AIMS-1-128M model is derived from this base model through supervised fine-tuning on conversational data.
Model specifications
- Parameters: 128,551,424
- Architecture: Custom decoder-only Transformer
- Layers: 28
- Hidden width: 512
- FFN width: 1365
- Query heads: 16
- KV heads: 4
- Context length: 1024 tokens
- Vocabulary size: 50,260
- Residual dropout: 0.1
- RMSNorm epsilon: 1e-05
- Tokenizer: Custom GPT-2 byte-level BPE tokenizer exported from the training encoding
Architecture
AIMS-1-128M-base uses a custom Transformer architecture that incorporates design ideas from several influential language-model architectures, including GPT, Llama, and Qwen.
The model is not a direct implementation of any one of these architectures. Instead, these models served as architectural references while developing the AIMS architecture and its specific configuration.
The model uses a decoder-only autoregressive architecture and is trained with a next-token prediction objective.
The original model computation is preserved in notebook_components.py. The Hugging Face integration is implemented separately in modeling_baby_aims.py.
The runtime implementation adapts the buffer lifecycle required for model loading while preserving the original computational behavior.
Project objective
The central objective of the AIMS-1 project is to explore:
How capable of learning English language structure can a small language model become when trained from scratch with approximately $200 of compute?
The base model is therefore primarily an experiment in:
- Training language models from scratch
- Learning English syntax and linguistic patterns
- Evaluating linguistic competence in small language models
- Understanding the effect of constrained compute and data
- Developing open and reproducible approaches to LLM research
The model is intentionally small and should not be compared directly with large-scale commercial language models in terms of general capabilities.
Training data
AIMS-1-128M-base was trained on the BabyLM-2026-Strict dataset provided by the BabyLM community.
The Strict track is designed around constrained training data and compute, making it particularly relevant to the objective of studying language acquisition in small language models.
The model was trained from scratch on this corpus rather than initialized from the weights of an existing pretrained language model.
Dataset:
- BabyLM-community/BabyLM-2026-Strict
The model's training objective was next-token prediction. The goal was to learn statistical regularities in English text, including lexical, grammatical, and syntactic patterns.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
repo_id = "AIMadeSimpleResearch/AIMS-1-128M-base"
tokenizer = AutoTokenizer.from_pretrained(
repo_id,
trust_remote_code=True,
)
model = AutoModelForCausalLM.from_pretrained(
repo_id,
trust_remote_code=True,
)
model.eval()
inputs = tokenizer(
"Every effort moves you",
return_tensors="pt",
)
output = model.generate(
**inputs,
max_new_tokens=40,
use_cache=False,
)
print(
tokenizer.decode(
output[0],
skip_special_tokens=True,
)
)
Consumers can install torch, transformers, and safetensors.
Custom-code loading is required.
This implementation currently has no KV cache. Generation therefore recomputes the prefix at each generation step. Keep the prompt plus generated tokens within the 1024-token context limit.
For batched generation, use left padding.
The padding adapter processes padded examples individually using the original AIMSModel and then restores the batch layout. This is slower than fully batched masked attention but avoids modifying the original model component implementation.
Temperature and top-p sampling
The model inherits generate() from Hugging Face's GenerationMixin.
Sampling can be enabled using standard Transformers generation parameters:
model.eval()
inputs = tokenizer(
"Every effort moves you",
return_tensors="pt",
).to(model.device)
output = model.generate(
**inputs,
max_new_tokens=100,
do_sample=True,
temperature=0.7,
top_p=0.9,
top_k=0,
use_cache=False,
)
print(
tokenizer.decode(
output[0],
skip_special_tokens=True,
)
)
do_sample=Trueenables sampling instead of greedy decoding.temperature=0.7controls the sharpness of the token distribution.top_p=0.9limits sampling to the most probable tokens covering approximately 90% of the probability mass.top_k=0disables the additional top-k filter.use_cache=Falseis required by this implementation.
For greedy decoding, use do_sample=False and omit temperature and top-p.
Keep the prompt plus max_new_tokens within the model's 1024-token context limit.
Causal language-model loss
For causal language-model loss, pass unshifted labels, usually a copy of input_ids, with padding labels set to -100.
The model shifts labels internally.
The training notebook's already-shifted target tensors must not be supplied directly.
Chat and instruction tuning
AIMS-1-128M-base is a base language model, not an instruction-tuned or conversational model.
No chat template is supplied.
The model was not trained specifically to follow user/assistant instructions. Special user/assistant tokens alone do not make a language model instruction-tuned.
The subsequent AIMS-1-128M model was created by supervised fine-tuning this base model on conversational data.
Therefore, conversational behavior observed from AIMS-1-128M should not be attributed to AIMS-1-128M-base.
Linguistic evaluation
AIMS-1-128M-base was evaluated on several tasks from the BabyLM evaluation suite. These evaluations are particularly relevant to the project's objective of measuring language understanding and linguistic competence rather than factual knowledge.
BabyLM evaluation
The following results compare AIMS-1 128M with the BabyLM baseline GPT-2 Strict and Strict-Small systems:
| Zero-shot task | BabyLM baseline GPT-2 Strict | Strict-Small | AIMS-1 128M |
|---|---|---|---|
| BLiMP | 74.73% | 65.23% | 75.10% |
| BLiMP Supplement | 65.00% | 57.25% | 65.23% |
| EWoK | 54.37% | 50.63% | 53.46% |
| Entity Tracking | 16.91% | 19.10% | 19.79% |
| COMPS | 55.85% | 51.81% | 56.62% |
| GlobalPIQA | 36.62% | 35.09% | โ |
These results indicate that AIMS-1-128M can achieve measurable performance on linguistic and language-understanding evaluations despite its relatively small size and constrained training budget.
GlobalPIQA was not reported for AIMS-1-128M in the evaluation results available for this model.
Comparison with GPT-2 124M
A separate evaluation compared AIMS-1-128M with the original GPT-2 124M model:
| Evaluation | GPT-2 124M | AIMS-1 128M |
|---|---|---|
| BLiMP โ full | ~66% | 73.46% |
| EWoK | ~50% | 53.46% |
The GPT-2 values are approximate and should be interpreted with caution because evaluation methodology and benchmark implementations can affect reported scores.
The AIMS-1 results reported here correspond to the specific AIMS-1-128M evaluation run and should not be assumed to represent all checkpoints or future versions of the model.
Interpreting the evaluations
The evaluations above are intended primarily to measure aspects of language and linguistic understanding.
In particular:
- BLiMP evaluates a range of grammatical phenomena in English.
- BLiMP Supplement extends grammatical evaluation to additional phenomena.
- EWoK evaluates aspects of world knowledge and language understanding.
- Entity Tracking evaluates the ability to track entities across linguistic contexts.
- COMPS evaluates compositional language understanding.
- GlobalPIQA evaluates question answering involving broader knowledge and reasoning.
These benchmarks should not be interpreted as a complete measure of general intelligence, reasoning, factual reliability, or conversational ability.
Training
AIMS-1-128M-base was trained from scratch as part of the AIMS-1 project.
The project targeted an approximately $200 compute budget, with the goal of exploring the capabilities achievable by a small language model under strict resource constraints.
The base-model training objective was autoregressive next-token prediction.
The model was subsequently used as the starting point for the AIMS-1-128M supervised fine-tuning experiment.
Limitations
AIMS-1-128M-base is a small experimental language model and has substantial limitations.
The model may:
- Produce factually incorrect information
- Generate incoherent or repetitive text
- Produce grammatically incorrect sentences
- Fail to maintain long-range context
- Have limited factual knowledge
- Struggle with reasoning and multi-step tasks
- Generate biased or undesirable content present in its training data
- Perform inconsistently across prompts
- Fail to follow natural-language instructions reliably
The model's ability to generate plausible English text should not be interpreted as evidence that it possesses reliable factual knowledge, reasoning ability, or human-like understanding.
The model's 1024-token context length and lack of KV caching also limit generation efficiency.
Relationship to AIMS-1-128M
AIMS-1-128M-base is the pretrained base model in the AIMS-1 model family.
The relationship between the two checkpoints is:
AIMS-1-128M-base
โ pretrained from scratch on BabyLM-2026-Strict
โ supervised fine-tuning
โ AIMS-1-128M
โ conversational user/assistant model
AIMS-1-128M-base is intended primarily for research into language-model pretraining and English language understanding.
AIMS-1-128M adds supervised fine-tuning for conversational interaction.
License
AIMS-1-128M-base is released under the Apache License 2.0.
The Apache-2.0 license applies to the model code and model artifacts released in this repository. Users are responsible for complying with the licenses and terms of the datasets, libraries, and other third-party materials used in the development or training of the model.
Attribution
AIMS-1-128M-base was developed by AIMadeSimple Research as part of an open research effort focused on making language-model development more accessible and affordable.
The project explores what can be achieved by training a language model from scratch with a small parameter count, constrained training data, and an approximately $200 compute budget.
- Downloads last month
- 174