AIMS-1-128M-base

AIMS-1-128M-base is a 128M-parameter decoder-only language model trained from scratch on the BabyLM-2026-Strict dataset.

The model is part of the AIMS-1 project, an experiment in training a language model with an approximately $200 compute budget. The primary objective is to investigate how effectively a small language model can learn the syntax and structural patterns of English under a strict data and compute constraint.

AIMS-1-128M-base is the pretrained base model. It has not been supervised fine-tuned for instruction following or conversational interaction. The later AIMS-1-128M model is derived from this base model through supervised fine-tuning on conversational data.

Model specifications

  • Parameters: 128,551,424
  • Architecture: Custom decoder-only Transformer
  • Layers: 28
  • Hidden width: 512
  • FFN width: 1365
  • Query heads: 16
  • KV heads: 4
  • Context length: 1024 tokens
  • Vocabulary size: 50,260
  • Residual dropout: 0.1
  • RMSNorm epsilon: 1e-05
  • Tokenizer: Custom GPT-2 byte-level BPE tokenizer exported from the training encoding

Architecture

AIMS-1-128M-base uses a custom Transformer architecture that incorporates design ideas from several influential language-model architectures, including GPT, Llama, and Qwen.

The model is not a direct implementation of any one of these architectures. Instead, these models served as architectural references while developing the AIMS architecture and its specific configuration.

The model uses a decoder-only autoregressive architecture and is trained with a next-token prediction objective.

The original model computation is preserved in notebook_components.py. The Hugging Face integration is implemented separately in modeling_baby_aims.py.

The runtime implementation adapts the buffer lifecycle required for model loading while preserving the original computational behavior.

Project objective

The central objective of the AIMS-1 project is to explore:

How capable of learning English language structure can a small language model become when trained from scratch with approximately $200 of compute?

The base model is therefore primarily an experiment in:

  • Training language models from scratch
  • Learning English syntax and linguistic patterns
  • Evaluating linguistic competence in small language models
  • Understanding the effect of constrained compute and data
  • Developing open and reproducible approaches to LLM research

The model is intentionally small and should not be compared directly with large-scale commercial language models in terms of general capabilities.

Training data

AIMS-1-128M-base was trained on the BabyLM-2026-Strict dataset provided by the BabyLM community.

The Strict track is designed around constrained training data and compute, making it particularly relevant to the objective of studying language acquisition in small language models.

The model was trained from scratch on this corpus rather than initialized from the weights of an existing pretrained language model.

Dataset:

  • BabyLM-community/BabyLM-2026-Strict

The model's training objective was next-token prediction. The goal was to learn statistical regularities in English text, including lexical, grammatical, and syntactic patterns.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "AIMadeSimpleResearch/AIMS-1-128M-base"

tokenizer = AutoTokenizer.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

model = AutoModelForCausalLM.from_pretrained(
    repo_id,
    trust_remote_code=True,
)

model.eval()

inputs = tokenizer(
    "Every effort moves you",
    return_tensors="pt",
)

output = model.generate(
    **inputs,
    max_new_tokens=40,
    use_cache=False,
)

print(
    tokenizer.decode(
        output[0],
        skip_special_tokens=True,
    )
)

Consumers can install torch, transformers, and safetensors.

Custom-code loading is required.

This implementation currently has no KV cache. Generation therefore recomputes the prefix at each generation step. Keep the prompt plus generated tokens within the 1024-token context limit.

For batched generation, use left padding.

The padding adapter processes padded examples individually using the original AIMSModel and then restores the batch layout. This is slower than fully batched masked attention but avoids modifying the original model component implementation.

Temperature and top-p sampling

The model inherits generate() from Hugging Face's GenerationMixin.

Sampling can be enabled using standard Transformers generation parameters:

model.eval()

inputs = tokenizer(
    "Every effort moves you",
    return_tensors="pt",
).to(model.device)

output = model.generate(
    **inputs,
    max_new_tokens=100,
    do_sample=True,
    temperature=0.7,
    top_p=0.9,
    top_k=0,
    use_cache=False,
)

print(
    tokenizer.decode(
        output[0],
        skip_special_tokens=True,
    )
)
  • do_sample=True enables sampling instead of greedy decoding.
  • temperature=0.7 controls the sharpness of the token distribution.
  • top_p=0.9 limits sampling to the most probable tokens covering approximately 90% of the probability mass.
  • top_k=0 disables the additional top-k filter.
  • use_cache=False is required by this implementation.

For greedy decoding, use do_sample=False and omit temperature and top-p.

Keep the prompt plus max_new_tokens within the model's 1024-token context limit.

Causal language-model loss

For causal language-model loss, pass unshifted labels, usually a copy of input_ids, with padding labels set to -100.

The model shifts labels internally.

The training notebook's already-shifted target tensors must not be supplied directly.

Chat and instruction tuning

AIMS-1-128M-base is a base language model, not an instruction-tuned or conversational model.

No chat template is supplied.

The model was not trained specifically to follow user/assistant instructions. Special user/assistant tokens alone do not make a language model instruction-tuned.

The subsequent AIMS-1-128M model was created by supervised fine-tuning this base model on conversational data.

Therefore, conversational behavior observed from AIMS-1-128M should not be attributed to AIMS-1-128M-base.

Linguistic evaluation

AIMS-1-128M-base was evaluated on several tasks from the BabyLM evaluation suite. These evaluations are particularly relevant to the project's objective of measuring language understanding and linguistic competence rather than factual knowledge.

BabyLM evaluation

The following results compare AIMS-1 128M with the BabyLM baseline GPT-2 Strict and Strict-Small systems:

Zero-shot task BabyLM baseline GPT-2 Strict Strict-Small AIMS-1 128M
BLiMP 74.73% 65.23% 75.10%
BLiMP Supplement 65.00% 57.25% 65.23%
EWoK 54.37% 50.63% 53.46%
Entity Tracking 16.91% 19.10% 19.79%
COMPS 55.85% 51.81% 56.62%
GlobalPIQA 36.62% 35.09% โ€”

These results indicate that AIMS-1-128M can achieve measurable performance on linguistic and language-understanding evaluations despite its relatively small size and constrained training budget.

GlobalPIQA was not reported for AIMS-1-128M in the evaluation results available for this model.

Comparison with GPT-2 124M

A separate evaluation compared AIMS-1-128M with the original GPT-2 124M model:

Evaluation GPT-2 124M AIMS-1 128M
BLiMP โ€” full ~66% 73.46%
EWoK ~50% 53.46%

The GPT-2 values are approximate and should be interpreted with caution because evaluation methodology and benchmark implementations can affect reported scores.

The AIMS-1 results reported here correspond to the specific AIMS-1-128M evaluation run and should not be assumed to represent all checkpoints or future versions of the model.

Interpreting the evaluations

The evaluations above are intended primarily to measure aspects of language and linguistic understanding.

In particular:

  • BLiMP evaluates a range of grammatical phenomena in English.
  • BLiMP Supplement extends grammatical evaluation to additional phenomena.
  • EWoK evaluates aspects of world knowledge and language understanding.
  • Entity Tracking evaluates the ability to track entities across linguistic contexts.
  • COMPS evaluates compositional language understanding.
  • GlobalPIQA evaluates question answering involving broader knowledge and reasoning.

These benchmarks should not be interpreted as a complete measure of general intelligence, reasoning, factual reliability, or conversational ability.

Training

AIMS-1-128M-base was trained from scratch as part of the AIMS-1 project.

The project targeted an approximately $200 compute budget, with the goal of exploring the capabilities achievable by a small language model under strict resource constraints.

The base-model training objective was autoregressive next-token prediction.

The model was subsequently used as the starting point for the AIMS-1-128M supervised fine-tuning experiment.

Limitations

AIMS-1-128M-base is a small experimental language model and has substantial limitations.

The model may:

  • Produce factually incorrect information
  • Generate incoherent or repetitive text
  • Produce grammatically incorrect sentences
  • Fail to maintain long-range context
  • Have limited factual knowledge
  • Struggle with reasoning and multi-step tasks
  • Generate biased or undesirable content present in its training data
  • Perform inconsistently across prompts
  • Fail to follow natural-language instructions reliably

The model's ability to generate plausible English text should not be interpreted as evidence that it possesses reliable factual knowledge, reasoning ability, or human-like understanding.

The model's 1024-token context length and lack of KV caching also limit generation efficiency.

Relationship to AIMS-1-128M

AIMS-1-128M-base is the pretrained base model in the AIMS-1 model family.

The relationship between the two checkpoints is:

AIMS-1-128M-base
โ†’ pretrained from scratch on BabyLM-2026-Strict
โ†’ supervised fine-tuning
โ†’ AIMS-1-128M
โ†’ conversational user/assistant model

AIMS-1-128M-base is intended primarily for research into language-model pretraining and English language understanding.

AIMS-1-128M adds supervised fine-tuning for conversational interaction.

License

AIMS-1-128M-base is released under the Apache License 2.0.

The Apache-2.0 license applies to the model code and model artifacts released in this repository. Users are responsible for complying with the licenses and terms of the datasets, libraries, and other third-party materials used in the development or training of the model.

Attribution

AIMS-1-128M-base was developed by AIMadeSimple Research as part of an open research effort focused on making language-model development more accessible and affordable.

The project explores what can be achieved by training a language model from scratch with a small parameter count, constrained training data, and an approximately $200 compute budget.

Downloads last month
174
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for AIMadeSimpleResearch/AIMS-1-128M-base

Finetunes
1 model

Dataset used to train AIMadeSimpleResearch/AIMS-1-128M-base