Pebble-25M-Chat

Banner

Pebble-25M-Chat is a compact, hybrid autoregressive chat language model. It combines the efficiency of state-space models with the proven performance of attention layers, optimized using a custom Muon + AdamW optimizer split.

Model Details

  • Architecture: Hybrid Mamba2 / Transformer
  • Block Pattern: 3 Mamba2 blocks : 1 Attention block (repeating)
  • Parameters: ~25,000,000 (25M)
  • Hidden Dimension: 608
  • Layers: 8 (6 Mamba2, 2 Attention)
  • Vocab Size: 2,048 (Custom Byte-Level BPE)
  • Context Length: 2048
  • Pretraining Tokens: 25,000,000,000 (25 Billion)
  • SFT Tokens: 250,000,000 (250 Million)
  • Optimizer: Muon (for 2D hidden weights) + AdamW (for embeddings, norms, and scalars)
  • Precision: fp32 master weights with bf16 autocast

Dataset Sources

The base model was pretrained on a 25B token subset of the following datasets:

Dataset Token Allocation Share
FineWeb-Edu 7.50 billion 30%
DCLM 5.00 billion 20%
Cosmopedia-v2 3.75 billion 15%
FineMath-4+ 3.75 billion 15%
FinePhrase 3.00 billion 12%
NPset 2.00 billion 8%

Benchmarks

Pebble-25M-Chat was evaluated using zero-shot multiple-choice evaluation. Higher scores are better. Bold indicates the best score among the models listed.

Benchmark Pebble-25M Pebble-25M Chat Pebble-10M BananaMind-2-Mini Random
PIQA 59.25% 53.37% 58.43% 59.63% 50.00%
ARC-Easy 38.17% 26.68% 37.29% 39.86% 25.00%
ARC-Challenge 18.60% 19.62% 18.60% 25.68% 25.00%
HellaSwag 27.62% 25.63% 26.81% 29.72% 25.00%
ArithMark-2.0 27.60% 26.20% 27.64% 27.52% 25.00%
ArithMark-3.0 33.80% 28.80% 32.80% 34.90% 25.00%

Evaluation Notes

  • PIQA, ARC-Easy, ARC-Challenge, and HellaSwag were evaluated on their respective test splits.
  • ArithMark-2.0 was evaluated on its train split due to the lack of a suitable test split.
  • ArithMark-3.0 was evaluated on its train split due to the lack of a suitable test split.
  • Results were obtained using zero-shot multiple-choice evaluation.
  • The model was additionally fine-tuned using supervised fine-tuning (SFT).

SFT Attribution

The 250,000,000 SFT tokens used for Pebble-25M-Chat were provided by smol-smoltalk.


Usage

To run the model for text generation, you will need to install the required dependencies. The included Mamba2 implementation relies on CUDA/Triton kernels and is intended to run on a CUDA-enabled GPU. Ampere-class GPUs or newer are recommended.

Note: The model uses custom architecture code, so you must pass trust_remote_code=True when loading both the tokenizer and the model.

Installation

pip install transformers huggingface_hub torch
pip install causal-conv1d mamba-ssm

Generation

Here is a simple Python script to load the model and generate text interactively:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL_ID = "basically-ai/Pebble-25M-Chat"


def main():
    print("Loading Pebble-25M-Chat...")

    tokenizer = AutoTokenizer.from_pretrained(
        MODEL_ID,
        trust_remote_code=True,
    )

    model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        trust_remote_code=True,
        dtype=torch.float32,
    ).to("cuda")

    model.eval()

    print(
        f"Model loaded successfully! "
        f"VRAM usage: {torch.cuda.memory_allocated() / 1e9:.2f} GB"
    )
    print("Type 'quit' or 'exit' to stop.\n")

    while True:
        prompt = input("You: ")

        if prompt.lower() in ["quit", "exit"]:
            break

        # Format the prompt for the chat model
        formatted_prompt = f"User: {prompt}\nAssistant: "

        # Tokenize the prompt
        inputs = tokenizer(
            formatted_prompt,
            return_tensors="pt",
        ).to("cuda")

        # Generate text
        print("Pebble: ", end="", flush=True)

        with torch.inference_mode():
            outputs = model.generate(
                **inputs,
                max_new_tokens=100,
                do_sample=True,
                temperature=0.7,
                top_k=50,
                top_p=0.95,
                repetition_penalty=1.2,
            )

        # Decode and print (skip the prompt part)
        generated_text = tokenizer.decode(
            outputs[0][inputs["input_ids"].shape[1]:],
            skip_special_tokens=True,
        )

        print(generated_text)
        print()


if __name__ == "__main__":
    main()

License

Apache 2.0

Downloads last month
-
Safetensors
Model size
24.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for basically-ai/Pebble-25M-Chat

Finetuned
(1)
this model
Quantizations
1 model

Collection including basically-ai/Pebble-25M-Chat