Loom Weave 4

Loom Weave 4

55.6M parameters ยท 24 layers ยท 8K context (trained to 32K) ยท runs on phones ยท Textile Labs

The best-scoring Loom yet, and the first that runs in phone and desktop chat apps โ€” PocketPal, LM Studio, Jan, llama.cpp, Ollama โ€” straight from the GGUF. Successor to Loom Weave 3. Trained from scratch: randomly initialised weights, nothing fine-tuned from anyone's checkpoint.

127 of 133 on the Loom acceptance battery โ€” the highest of any Loom, including the 155M Crucible Preview โ€” at a third of its size. The phone build (Q4, 44 MB) scores the same 127.

What changed from Weave 3

Weave 3 Weave 4
layout Llama, every layer sees everything Gemma-3 style: 20 local layers (1,024-token window) + 4 global โ€” long input stays cheap
size 31.5M 55.6M
context 1,024 8,192 used in practice (trained in a 32,768 stage โ€” see Read this)
tokenizer 16K, digits merged 16K, every digit its own token, identical in llama.cpp and transformers
runs in phone/desktop chat apps no output (no template in the GGUF) yes โ€” chat template and end-of-turn built into the GGUF
shows its working no <think>โ€ฆ</think> โ€” and its answer agreed with its own working 40 of 40 times
hardware Kaggle Kaggle, 2 ร— T4, 12.2 h in two runs

Measured against Weave 3

Same tests, same scripts, same settings (repeat penalty 1.0), both through Ollama. Every test prompt is scrubbed from the training data.

End to end: 20 held-out everyday questions, live Wikipedia, the model writing its own query.

decided to search wrote its own query answer reached the model answered right
Weave 3 20/20 20/20 12/20 6/20
Loom Weave 4 20/20 20/20 12/20 10/20

The acceptance battery, row by row:

row Weave 3 Loom Weave 4
A ยท says its own name 10/12 12/12
B ยท its own name under rough typing 10/12 12/12
C ยท 5-turn conversation stays on thread 5/5 5/5
D ยท answers from a search result 5/5 5/5
E ยท follow-up answered from the same result 2/5 4/5
F ยท says it looked, after a lookup 5/5 5/5
G ยท never claims a lookup it didn't make 16/16 16/16
H ยท admits what it can't know about you 8/8 8/8
I ยท says when a result doesn't contain the answer 0/5 1/5
J ยท never leaks a search tag with tools off 28/28 28/28
K ยท stops on its own 12/12 12/12
L ยท searches when it should, not for your private things 19/20 19/20
total 120/133 127/133

Held-out behaviour tests:

Weave 3 Loom Weave 4
prompt injection โ€” kept its identity, didn't obey (12 prompts ร— 3) 11/36 33/36
10- and 12-turn conversations โ€” turns answered on target 29/44 42/44
a fact from turn 1โ€“3 asked again at turn 9โ€“10 0/4 3/4
unknown facts, tools off โ€” declines instead of guessing 8/20 14/20
basic facts, tools off โ€” answers right 10/20 12/20
no false "I remember that" 15/20 17/20
a <tools:on> typed inside a message doesn't switch search on 11/12 12/12
in its own words about itself 13/16 14/16

Every Loom text model

model params battery /133 live search (held-out) reads real prose runs in phone apps
Loom Spark 2 19.9M ~97 2/20 no no
Loom Tapestry 2 22.8M 107 โ€” curated only no
Loom Tapestry 3 Flash 7.18M 112 3/20 curated only no
Loom Spark 3 Flash 7.18M 119 5/20 curated only no
Loom Spark 3 12.2M 120 7/20 curated only no
Loom Spark 3.2 22.8M 122โ€  7/20 curated only no
Loom Weave 3 31.5M 120 6/20 yes no
Loom Tapestry 3 69.2M 123 12/20 yes + multi-hop no
Loom Crucible Preview 155.0M 125โ€  10/20 yes โ€” best curated reader no
Loom Weave 4 55.6M 127โ€  10/20 yes yes

โ€  scored on a battery with every test prompt scrubbed from training. Earlier rows are each model's release score.

Read this before you use it

Every point here was measured. Weave 4 was built to try three new things; one worked, two did not yet.

  • Reasoning on new kinds of task is not reliable. On 100 hand-written tasks of kinds it never trained on (triage, comparing plans, meeting slots, spotting a mistake, filling a formโ€ฆ) it got 11/100. It learned the shape of working things out better than the substance. It shows its working (<think>โ€ฆ</think>) and its answer matches that working, so you can check it โ€” and you should.
  • Arithmetic: the working is usually right, the final number often isn't. It writes the column steps correctly ("ones 7 + 5 = 12, write 2 carry 1 โ€ฆ") and then sometimes assembles the result wrongly. Check any sum it gives you.
  • The confidence word is not calibrated. Answers end in (sure), (I think) or (not sure). On checkable tasks it was right 13 of 25 times it said sure. Treat it as a hint, not a guarantee.
  • Long documents don't work. It was trained in a 32,768-token stage, and llama.cpp will run it at that length, but it could not find a planted fact in a 2,000โ€“8,000-token text. Use it for conversations (it holds a 10-turn chat well), not for reading long documents.
  • Standard benchmarks are low, as expected at this size: GSM8K 4/100, ARC-Easy 27/100, StrategyQA 15/100 (it often, correctly, declines world-knowledge questions with tools off).
  • It gets half of everyday questions right with search (10/20). "I looked that up" means it searched, not that it read the result correctly โ€” the harness's --show prints what it read. A search can land on the wrong page ("the Four Seasons" โ†’ the hotel company).
  • Short answers. Replies are a sentence or two โ€” it's built to be brief.
  • Warmth is uneven (10/20 on our good-news/bad-news test).
  • <recall> and <tool> tags exist in its vocabulary for agent harnesses, but the shipped harness only runs web search.

Usage โ€” phone and desktop chat apps (PocketPal, LM Studio, Jan)

Download loom-weave-4-q4_k_m.gguf (44 MB) or loom-weave-4-q8_0.gguf (60 MB) and load it. The chat template and the stop token are inside the file. Set repeat penalty 1.0.

Usage โ€” the harness (search)

python3 harness.py "whats the capital of peru"
python3 harness.py --show "who wrote hamlet"     # see what it searched and read
python3 harness.py --no-tools "who are you"

Stdlib only; Wikipedia needs no key. Never feed a failed lookup back as a result โ€” the model answers from the error text.

Usage โ€” Ollama / llama.cpp

ollama run hf.co/textilelabs/Loom-Weave-4 "who are you"
llama-cli -m loom-weave-4-q8_0.gguf --jinja -cnv      # uses the template in the GGUF

template and params are read automatically by Ollama. Do not add a repetition penalty.

Usage โ€” transformers (4.50 or newer)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tok = AutoTokenizer.from_pretrained("textilelabs/Loom-Weave-4")
model = AutoModelForCausalLM.from_pretrained("textilelabs/Loom-Weave-4").eval()
prompt = tok.apply_chat_template([{"role": "user", "content": "who are you"}], tokenize=False, add_generation_prompt=True)
ids = tok(prompt, return_tensors="pt", add_special_tokens=False).input_ids
out = model.generate(ids, max_new_tokens=160, do_sample=False, eos_token_id=0, pad_token_id=1)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))

Raw format (tools on or off): <tools:off>\n<user>\n{message}\n<|eot|>\n<loom>\n.

How it was built

architecture Gemma-3 layout: 24 layers ร— 448d (20 sliding-window 1,024 + 4 global), GQA 7 heads / 1 KV, GeGLU, QK-norm, tied embeddings
parameters 55,555,520
vocabulary 16,384 BPE, single digits, Qwen2-style pre-tokenizer (identical tokens in llama.cpp)
training run 1: 8.2 h from random init, 362M tokens (8K, then 32K, then a focused polish, then a calibration patch on its own answers). Run 2: 4 h on its own weights, 155M tokens, rebalanced reasoning data.
optimiser Muon on 2-D hidden matrices, AdamW on embeddings and norms; per-conversation attention masking
attention check before every run, the trainer's attention is checked against the reference implementation (loss and gradients identical)
hardware Kaggle, 2 ร— NVIDIA T4

Files

model.safetensors / config.json            the model (transformers)
tokenizer.json / tokenizer_config.json     tokenizer + chat template
loom-weave-4-f16.gguf                      full precision GGUF (112 MB)
loom-weave-4-q8_0.gguf / -q4_k_m.gguf      phone builds (60 / 44 MB) โ€” Q4 scores the same 127/133
template / params / Modelfile              Ollama
harness.py                                 runnable search harness โ€” stdlib only
ATTRIBUTION.md                             required credits for the training data

Thanks

Thanks to Andrew Thompson for his analysis of the vocabulary's share of a small model, which directly shaped Weave 4's 16K tokenizer โ€” the word table is 13% of the model, not 20%.

License

Model: MIT. Training data keeps its original licences โ€” see ATTRIBUTION.md.

Downloads last month
242
Safetensors
Model size
55.6M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Collection including textilelabs/Loom-Weave-4