StellarAI-Tiny

A lightweight multimodal language model trained from scratch — ~50M parameters (0.05B), runs on CPU with 4GB RAM.

License: MIT PyTorch Model Size Parameters


Overview

StellarAI-Tiny is a from-scratch, bilingual (Chinese + English) causal language model with multimodal vision support. Designed for educational and prototyping purposes, it requires minimal hardware and ships with a built-in plugin system for tool calling.

Feature Description
Lightweight 50M parameters, ~93MB weights
CPU-friendly Runs smoothly on CPU with 4GB RAM
Transformer 4-layer text encoder + RoPE positional encoding
Multimodal CNN + ViT hybrid vision encoder + cross-attention fusion
Bilingual Chinese + English mixed tokenization & generation
License MIT — fully permissive for commercial use
Plugins Built-in calculator, knowledge base, translator, text tools, time queries

Architecture

StellarAI-Tiny (~50M Parameters)
├── Embedding          vocab(32000) x d_model(384)
├── Text Transformer (4 layers)
│   ├── Multi-Head Self-Attention (6 heads, RoPE)
│   └── FFN (GELU, intermediate=1536)
├── Vision Encoder (CNN + ViT)
│   ├── CNN Feature Extractor (4 layers: 24→48→96→192 channels)
│   └── ViT Transformer (2 layers, 6 heads)
├── Fusion Block (1 layer)
│   ├── Self-Attention + Cross-Attention
│   └── FFN (1536)
└── LM Head (tied weights)
Config Value
d_model 384
num_hidden_layers 4
num_attention_heads 6
intermediate_size 1536
vocab_size 32000
max_position_embeddings 1024
vision_num_layers 2
fusion_num_layers 1
Total parameters ~50M (0.05B)

Quick Start

Requirements

pip install torch transformers safetensors

Text Generation

from transformers import AutoModelForCausalLM, AutoConfig, AutoTokenizer
import torch

model_name = "amtstudio/stellarai-tiny"  # or your local path

config = AutoConfig.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    config=config,
    trust_remote_code=True,
    torch_dtype=torch.float32,
)
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)

# Generate
prompt = "Artificial intelligence is"
inputs = tokenizer(prompt, return_tensors="pt")

output_ids = model.generate(
    **inputs,
    max_new_tokens=64,
    temperature=0.7,
    top_k=40,
    do_sample=True,
    pad_token_id=tokenizer.pad_token_id,
    eos_token_id=tokenizer.eos_token_id,
)

print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

Using generate_text Convenience Method

result = model.generate_text(
    prompt="Artificial intelligence is",
    tokenizer=tokenizer,
    max_new_tokens=100,
    temperature=0.8,
)
print(result)

Training Details

Item Detail
Training steps 12,000 (5,000 base + 7,000 general training)
Corpus 19,846 lines of bilingual data (AI, CS, NLP, math, programming, reasoning, dialogue, plugins)
Optimizer AdamW (lr=3e-4, wd=0.01)
LR Schedule Cosine annealing + Warmup (100 steps)
Batch size 4
Sequence length 128
Gradient clipping 1.0
Final loss 4.06 (ppl ≈ 58)
Vocabulary size 11,030
Device CPU
Training time ~3.2 hours

The training corpus was built from a mix of hand-crafted bilingual data, synthetic instruction-tuning data, and Chinese NLP datasets across 10+ domains. The model was trained with a next-token-prediction objective using the custom SimpleTokenizer (BPE).


File Structure

├── config.json                    # HF model configuration
├── configuration_stellarai.py     # Custom PretrainedConfig class
├── modeling_stellarai.py          # Custom PreTrainedModel class
├── tokenization_stellarai.py      # Custom PreTrainedTokenizer class
├── model.safetensors              # Safetensors weights (93.6 MB)
├── pytorch_model.bin              # PyTorch weights (93.6 MB)
├── tokenizer_config.json          # Tokenizer configuration
├── special_tokens_map.json        # Special token mappings
├── tokenizer.json                 # HF tokenizer definition (BPE)
├── backend_tokenizer.json         # Original backend tokenizer
├── vocab.json                     # BPE vocabulary (11,030 tokens)
├── merges.txt                     # BPE merge rules
├── README.md                      # This file
└── LICENSE                        # MIT License

Limitations

Important: This is a lightweight educational / prototyping model.

  1. Limited knowledge: Trained on ~19K lines of curated data. Knowledge coverage is narrow.
  2. Factual accuracy: May produce inaccurate, nonsensical, or hallucinated content.
  3. Generation quality: Suitable for demonstrating basic language modeling — not production-level dialogue.
  4. Vision capability: The vision encoder is pre-trained on text-only data. VQA requires additional fine-tuning with image-text pairs.
  5. Plugin calling: The model learned the [TOOL:xxx] format but calling accuracy needs improvement.
  6. Not suitable for: Production environments, medical/legal/financial domains.

Suggested Improvements

  • Expand corpus to 50K+ lines or use public datasets (WikiText, C4, Oscar)
  • Increase training to 50K+ steps
  • Add real dialogue data (ShareGPT, Alpaca format) for SFT
  • Collect image-text pairs (e.g., COCO captions) to fine-tune multimodal capability
  • Try larger config: 6 layers / 512d / 8 heads (~0.1B)

License

MIT License — fully permissive for personal and commercial use.


Acknowledgements

  • Architecture inspired by GPT-2, LLaMA, ViT, and BLIP-2
  • Built with Hugging Face transformers
  • RoPE: RoFormer: Enhanced Transformer with Rotary Position Embedding

StellarAI — Exploring AI, one star at a time.

Downloads last month
229
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support