Instructions to use textilelabs/Loom-Atom with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use textilelabs/Loom-Atom with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="textilelabs/Loom-Atom")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("textilelabs/Loom-Atom") model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Atom", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use textilelabs/Loom-Atom with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf textilelabs/Loom-Atom:F16 # Run inference directly in the terminal: llama cli -hf textilelabs/Loom-Atom:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf textilelabs/Loom-Atom:F16 # Run inference directly in the terminal: llama cli -hf textilelabs/Loom-Atom:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf textilelabs/Loom-Atom:F16 # Run inference directly in the terminal: ./llama-cli -hf textilelabs/Loom-Atom:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf textilelabs/Loom-Atom:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf textilelabs/Loom-Atom:F16
Use Docker
docker model run hf.co/textilelabs/Loom-Atom:F16
- LM Studio
- Jan
- Ollama
How to use textilelabs/Loom-Atom with Ollama:
ollama run hf.co/textilelabs/Loom-Atom:F16
- Unsloth Desktop
- Docker Model Runner
How to use textilelabs/Loom-Atom with Docker Model Runner:
docker model run hf.co/textilelabs/Loom-Atom:F16
- Lemonade
How to use textilelabs/Loom-Atom with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull textilelabs/Loom-Atom:F16
Run and chat with the model
lemonade run user.Loom-Atom-F16
List all available models
lemonade list
- Atomic Chat
Loom Atom
22,392 parameters Β· 57 KB Β· Textile Labs
One question, one bit: does this need a tool?
whats the weather in leeds β <tool>
remind me to call mum at 6 β <tool>
convert 30 miles to km β <tool>
who are you β <none>
i had a rough day β <none>
sort it out β <none>
It emits exactly one token. That's the whole model.
Trained from scratch in five minutes on a 2013 desktop CPU β randomly initialised weights, nothing fine-tuned from a pretrained base. It is a real causal transformer, not a classifier: 2 layers, 24 hidden dimensions, tied embeddings.
Measured
| Loom Atom | keyword baseline | |
|---|---|---|
| held-out human utterances (1,860) | 96.0% | 64.9% |
| SNIPS β never seen in training (700) | 96.6% | 59.0% |
The second row is the one that matters. SNIPS played no part in training and the score does not drop β so this is not memorised phrasings. The keyword baseline is a hand-written list of ~75 tool-ish words scored on the identical splits.
Both test sets are balanced, so chance is 50%.
How small is 22,392 parameters?
| parameters | ratio | |
|---|---|---|
| Loom Weave 2 | 59,650,000 | 2,664Γ |
| Loom Spark 2 | 19,867,008 | 887Γ |
| Loom Router 1 | 1,435,040 | 64Γ |
| Loom Atom | 22,392 | 1Γ |
57 KB. Small enough to embed as a byte array in a header file, and it runs in well under a millisecond on a CPU.
What it's for
The cheapest useful decision in an agent stack: should this request touch a tool at all?
Put it in front of everything. If it says <none>, you have saved a retrieval call, a
router call, and possibly a large-model call β for the cost of a 57 KB matrix multiply. If
it says <tool>, hand off to something that decides which tool
(Loom Router 1 does that in one token
across 17 routes).
It is not a chat model, a router, or a classifier of intent. It answers one binary question and nothing else.
Where the floor is
A full ladder was trained, four minutes per rung, identical data:
| params | dim | layers | held-out | SNIPS |
|---|---|---|---|---|
| 86,640 | 48 | 3 | 93.0% | 95.9% |
| 26,976 | 32 | 2 | 95.6% | 93.6% |
| 22,392 | 24 | 2 | 96.0% | 96.6% |
| 9,392 | 16 | 2 | 93.4% | 94.1% |
| 6,800 | 16 | 1 | 83.0% | 84.4% |
| 4,812 | 12 | 1 | 82.2% | 81.9% |
Depth matters more than width. Narrowing from 24 to 16 dimensions cost about 3 points. Dropping from two layers to one cost ten. One attention layer can notice keywords; two can combine a keyword with its context. The floor is a layer count, not a parameter count.
Every rung beats the keyword baseline β even 4,812 parameters, by 17 points.
One honest note: the 86,640-parameter model scores lowest on held-out data because every rung got the same four minutes, and it completed 3,006 optimiser steps against 24d2L's 10,809. It is under-trained, not worse. Do not read this table as "smaller is better".
Known weakness
Questions about the user personally β "what is my sister's name", "what did I have for
breakfast" β are the hard case. They need no tool (no tool can answer them), but they look
like lookups. Atom gets some right and some wrong; treat <tool> on a first-person
question as unreliable.
Usage β Ollama
ollama run hf.co/textilelabs/Loom-Atom "whats the weather in leeds"
# <tool>
The template and params files in this repo are read automatically. params pins
temperature: 0 and num_predict: 1 β one token, deterministic.
Usage β transformers
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("textilelabs/Loom-Atom")
model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Atom").eval()
pair = torch.tensor([tok.convert_tokens_to_ids("<tool>"),
tok.convert_tokens_to_ids("<none>")])
def needs_tool(message: str) -> bool:
p = f"<user>\n{message.strip()}\n<|eot|>\n<loom>\n"
ids = tok(p, return_tensors="pt", add_special_tokens=False).input_ids
with torch.no_grad():
logits = model(input_ids=ids).logits[0, -1]
# decide only between the two legal answers
return bool(logits[pair].argmax() == 0)
needs_tool("whats the weather in leeds") # True
needs_tool("i had a rough day") # False
Prompt format is exact: <user>\n{message}\n<|eot|>\n<loom>\n.
Files
config.json / model.safetensors the model β 67 KB
tokenizer.json / tokenizer_config.json custom BPE tokenizer, 512 tokens
loom-atom-f16.gguf 57 KB, for Ollama / llama.cpp
template / params read automatically by `ollama run hf.co/...`
Modelfile for building locally
Training data
Real human utterances from two openly licensed corpora, relabelled to a single bit:
- MASSIVE β Amazon (CC BY 4.0), derived from SLURP (CC BY 4.0)
- CLINC150 β
clinc/oos-eval(CC BY 3.0)
15,502 utterances, balanced 50/50 by downsampling the majority class. A small procedurally generated slice written by Textile Labs covers "no tool needed" cases that public assistant corpora do not contain β chit-chat, ambiguity, and questions only the user can answer.
Both licences require attribution; this section satisfies that and must be kept with any redistribution.
License
Model: MIT. Training data retains its original licences and attribution as above.
- Downloads last month
- -