YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

TinyLlama-1.1B-4bit

This repo contains a 4-bit quantized version of the TinyLlama-1.1B base model, quantized in the summer of 2025. It is designed for ultra-lightweight inference, edge computing, and hardware benchmarking.

Model Description

Intended Uses & Limitations

Best For:

  • Edge Computing: Ideal for low-resource devices like Raspberry Pi, older smartphones, or embedded systems.
  • Prototyping & Benchmarking: Perfect for testing inference pipeline setups (transformers, accelerate) without high VRAM/RAM overhead.
  • CI/CD Testing: A lightweight sandbox model for automated testing scripts.

Limitations:

  • Due to the aggressive 4-bit quantization on a 1.1B parameter base, the model may exhibit degradation in complex reasoning or long-text generation.

Quick Start (Usage)

You can easily load and run this model using Hugging Face transformers and bitsandbytes:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

model_id = "PaddySeahorse/tinyllama-4bit"

# 4-bit configuration
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_quant_type="nf4"
)

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    quantization_config=bnb_config, 
    device_map="auto"
)

# Inference
prompt = "The meaning of life is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")

with torch.no_grad():
    outputs = model.generate(**inputs, max_new_tokens=40)
    print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Created in Summer 2025 as an exploration in LLM quantization.

Downloads last month
12
Safetensors
Model size
1B params
Tensor type
F32
·
F16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support