Safetensors
dynamicmind
custom_code

Banner

DynamicMind-Mini

DynamicMind-Mini is a decoder-only model trained on FineWeb-Edu, SmolLM-Corpus and FineMath

The model has about 8.9M parameters and its based on MiniBananaMind-v4-9M architecture with a custom 8k-token byte-level BPE tokenizer with digit-aware tokenization.

Model Details

Field Value
Parameters 8,884,992
Architecture Custom Llama-style decoder
Layers 9
Hidden size 256
Intermediate size 768
Attention heads 8
KV heads 2
Vocabulary size 8,192
Context length 1,024
Embeddings Tied input/output embeddings
Weight format safetensors
Base checkpoint MiniBananaMind-v3-9M
Continued pretraining checkpoint Cosmopedia-v2 step 2,613

Tokenizer

DynamicMind-Mini uses the same digit-aware 8k tokenizer as MiniBananaMind-v4-9M.

Digits are kept as separate tokens so numbers do not collapse into large number tokens during tokenization.

Digit IDs:

Token ID
1 9
2 10
3 11
4 12
5 13
6 14
7 15
8 16
9 17
0 18

Usage

This model uses custom architecture code, so load it with trust_remote_code=True.

Install dependencies:

pip install -U transformers safetensors torch

Run inference:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "DedeProGames/DynamicMind-Mini"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float16,
).cuda().eval()

prompt = "The meaning of life is "
input_ids = tokenizer(prompt, return_tensors="pt").input_ids.to(model.device)

with torch.no_grad():
    output = model.generate(
        input_ids=input_ids,
        max_new_tokens=64,
        do_sample=False,
        repetition_penalty=1.1,
        pad_token_id=tokenizer.eos_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))
Downloads last month
106
Safetensors
Model size
8.88M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 1 Ask for provider support

Datasets used to train DedeProGames/DynamicMind-Mini

Collection including DedeProGames/DynamicMind-Mini