Instructions to use harpertoken/tiny with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use harpertoken/tiny with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("harpertoken/tiny") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use harpertoken/tiny with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "harpertoken/tiny"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "harpertoken/tiny" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "harpertoken/tiny", "messages": [ {"role": "user", "content": "Hello"} ] }' - Atomic Chat
tiny
A four-bit quantised conversion of mlx-community/SmolLM-135M-Instruct-4bit, produced with mlx-lm 0.30.4 and intended for Apple Silicon. The architecture is a Llama variant (thirty layers, 576 hidden dimensions, nine attention heads, three key-value heads, 49152-token vocabulary) with group-size 64 four-bit weights, occupying about 76 MB.
The model.safetensors.index.json alongside the weights is not a sharded checkpoint here; it maps all 694 tensors onto the single model.safetensors file that sits beside it.
Because the weights are quantised, transformers cannot load this checkpoint; it needs mlx-lm, which applies the group quantisation at load time. The instruction tuning lives upstream in SmolLM; this repository adds a format, nothing else. Given the size, expect it to be fast and largely incoherent on factual questions.
Usage
from mlx_lm import load, generate
model, tokenizer = load("harpertoken/tiny")
prompt = "hello"
if tokenizer.chat_template is not None:
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
return_dict=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=100))
Install with pip install mlx-lm. A chat_template.jinja is present, so the template branch above will be taken.
Limitations
This is a redistribution of someone else's work in a different numeric format. There is no training, evaluation or adaptation to report, and if you want SmolLM-135M-Instruct you should use the upstream repository, which is not tied to MLX and can be converted to other formats. Quantisation to four bits costs some accuracy even before the underlying 135M-parameter capacity does; I have not measured the loss and so make no claim about it.
Attribution
SmolLM follows Allouji et al., SmolLM: blazingly fast and remarkably powerful (2024).
- Downloads last month
- 305
4-bit
Model tree for harpertoken/tiny
Base model
mlx-community/SmolLM-135M-Instruct-4bit