Useful local models

Small models that run on an ordinary CPU and take frequent, repetitive work away from large language models: cheaper, faster, and private, because the text never leaves the machine.

Folder What it does Size
sumeng-summarizer-qwen3.5-4b/ Writes SumEng's summaries of passages and messages: one plain paragraph within a length limit. Fine-tuned Qwen3.5-4B, run by llama.cpp. For non-commercial use (see its training data). 4B parameters (2.8 GB, Q4_K_M)
same-subject-r3/ The same task as r2, trained on more kinds of text: textbooks, novels, chats, and pairs of summaries. Better than r2 on every kind it was tested on. For non-commercial use (see its training data). 150M parameters
same-subject-r2/ Says whether a new passage or message continues the subject of the text before it. Used to cut a long stream (a book, a chat, meeting notes) into topics for a summary tree. 150M parameters
entity-link-r1/ Says whether a name in a sentence refers to a given known entity (a person, place, thing). Used to link each mention in a stream of text to the right entity in a growing graph. 150M parameters
speaker-r1/ Says whether a given person speaks a quoted paragraph of fiction, from the paragraph and the text before it. Used to attribute dialogue in a stream of text. 150M parameters
choice-r1/ Answers a multiple-choice question about a stream of text: which known entity a marked name is (or none), who says a quoted paragraph (or none), or whether two entities are one. Experimental: better than the two models above on single decisions, but not better in use (see its limits). 150M parameters

About the author

Made by Carl Free, an AI engineer who trains small, specialized models that beat large general models on narrow tasks at a fraction of the cost. Résumé, demos, and contact: ordinarydata.com/resume.


sumeng-summarizer-qwen3.5-4b

Qwen/Qwen3.5-4B fine-tuned to write the summaries SumEng asks for: given SumEng's prompt (a run of passages or messages, a length limit, and plain rules), one paragraph that says what the stretch is about, names the people and things, and keeps within the limit. It replaces Claude Haiku at SumEng's basic level; Claude stays as the fallback.

Results

On a whole textbook it never saw (OpenStax Psychology 2e, 921 passages), it wrote all 246 passage summaries, each accepted on the first try (Claude Haiku: 53 of 299 replies too long and asked again). Searching that tree found the evidence for questions as well as searching the same tree summarized by Claude: 59% against 63% of detail questions with the tree walk alone, and 100% for both with SumEng's keyword index; 68% against 64% of the book's review questions. Judged side by side by Claude Sonnet on 30 prompts, its summaries scored 3.6 of 5 and Claude's about 4.1. About 10 seconds per summary on an M1 Pro (Metal).

Use

uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "sumeng-summarizer-qwen3.5-4b/*" --local-dir models
llama-server -m models/sumeng-summarizer-qwen3.5-4b/sumeng-summarizer-qwen3.5-4b.Q4_K_M.gguf --jinja --reasoning-budget 0 -c 16384

Send SumEng's prompt as the user message, with thinking off (chat_template_kwargs: {"enable_thinking": false}), temperature 0.3. SumEng's sumeng.local.LocalModel does this.

Training

754 summaries that Claude (Haiku for passages, Sonnet for summaries of summaries) wrote and SumEng accepted while building trees on training texts only: OpenStax Introduction to Sociology 3e and U.S. History (CC BY-NC-SA 4.0), four public-domain novels and eight Sherlock Holmes stories, and six LoCoMo chats (CC BY-NC 4.0); 83 from other books and a chat for validation. LoRA rank 32 on the language layers with Unsloth, loss on the reply only, one epoch (validation loss 0.91), merged and quantized to Q4_K_M.

Licence: the base model is Apache 2.0, but part of the training text is licensed for non-commercial use only, so use this model for non-commercial purposes.

Limits

English only. Trained on SumEng's own prompt wording; other prompts may work less well. Not yet tried for summaries of summaries in a whole tree.


same-subject-r3

The successor to same-subject-r2 below: the same model, input, files and use, retrained on more kinds of text. It answers is the new text about the same subject as the earlier text? and returns a probability.

Results

AUC on held-out pairs (0.5 is guessing):

Pairs r2 r3
r2's own test set (stories and meetings) 0.815 0.823
textbook passages (Astronomy 2e) 0.80 0.90
novel passages (The Time Machine) 0.75 0.80
chat messages (a LoCoMo conversation) 0.89 0.96
pairs of summaries (textbook and novel) 0.63 0.72

Accuracy on r2's test set at 0.5: 74.7% (r2: 73.1%). In SumEng, on a whole textbook it never saw (Psychology 2e), groups of level-1 summaries now line up with the book's sections (Pk 0.32 against 0.47 with r2; lower is better), with the cutting rule waiting until a group holds three times the minimum before choosing where to cut. Between summaries r3 says "same subject" far more readily than r2 (median near 1.0, against about 0.1), which is what it was trained to do; a low answer between two summaries is now a real break.

Training

26,059 training pairs, each source in one split: r2's Holmes stories and QMSum meetings, plus the first chapters of Introduction to Sociology 3e and U.S. History (OpenStax, CC BY-NC-SA 4.0), The Adventures of Tom Sawyer, Silas Marner and The Scarlet Letter (public domain), and five LoCoMo conversations (snap-research/locomo, CC BY-NC 4.0). Validation: Astronomy 2e, The Time Machine and one more LoCoMo conversation. Textbook boundaries are the authors' subsection headings; novel scenes, chat topics within one sitting, and every summary were marked and written by Claude Sonnet. New in r3: pairs of summaries (the summaries of the segments before a segment, then its summary; "not yet" when a new section or chapter starts), for SumEng's checks above the first level. Book, textbook and summary pairs were shown two or three times per epoch. One epoch at learning rate 1e-5, batches of 16, on an A100; the second epoch overfit.

Licence: because part of the training text is licensed for non-commercial use only, use this model for non-commercial purposes. r2 has no such limit.

uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "same-subject-r3/*" --local-dir models

same-subject-r2

A cross-encoder (a model that reads two texts together and scores them) fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer one question: is the new text about the same subject as the earlier text? It returns a probability. It is the classifier behind SumEng, which builds a pyramid of summaries over text that arrives in order: the classifier decides where each topic ends, and a language model only writes one summary per topic.

Results

On held-out test pairs (stories and meetings never seen in training):

Accuracy AUC
this model 73% 0.82
Claude Sonnet, same question 61% 0.69
Claude Haiku, same question 52% 0.57

AUC is the chance that a random "same subject" pair scores above a random "different subject" pair (0.5 is guessing). When it is confident (probability below 0.1 or above 0.9, 42% of checks), it is right 88% of the time.

Cutting whole held-out Holmes stories into scenes with the rule below found 17 of 20 scene breaks (Pk 0.37, against 0.49 for fixed-size cuts; lower is better). Speed on a laptop CPU (M1 Pro, 4 threads): about 0.15 s per check at 400 tokens and 0.4 s at 1,000 with PyTorch; onnxruntime was about twice as slow there.

Input

A text pair. The first text is the earlier text (the last few passages, or a summary); the second is the new text followed by a blank line and a fixed question:

[earlier]  Holder:
           I have a son, Arthur, who has been a disappointment to me...

[new]      Holmes:
           "Let us go down to Streatham and look at the house."

           Question: Is the new passage about the same subject as the earlier passages?

Items are written as header:\ntext, separated by blank lines. Use "message" instead of "passage" for chat. Tokenize with tokenizer.json as a pair, truncating only the first text, at most 2,048 tokens, and with special tokens in the text treated as plain text (split_special_tokens=True in transformers, encode_special_tokens = True in tokenizers). The output is one logit; the probability is its sigmoid.

Files

  • model.safetensors, config.json: the weights, for PyTorch (AutoModelForSequenceClassification, attn_implementation="sdpa").
  • model.onnx: the same model for onnxruntime (answers match PyTorch to 0.00001).
  • tokenizer.json: the tokenizer. tokenizer_config.json was written by transformers 5; with older transformers, build the tokenizer from tokenizer.json directly.
  • sumeng-classifier.json: the maximum length and the ONNX file's hash.
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "same-subject-r2/*" --local-dir models

Use it for segmentation

A single low answer is not a topic break. What worked: let a group grow until it holds twice a minimum size (3,000 characters), then cut at the point with the lowest probability that leaves the minimum on each side, if that probability is below 0.4; otherwise keep growing.

Training

13,639 training pairs from 12 Sherlock Holmes stories (public domain, Project Gutenberg) and 196 AMI and ICSI meetings from QMSum (CC BY 4.0), split by source. Meeting topic boundaries were marked by people; book scene boundaries and all summaries were written by Claude models. Two epochs at learning rate 1e-5 on one L4 GPU; this is the epoch-1 checkpoint.

Limits

English only. Meetings of short spoken turns are segmented much less well than books. No chat or email data was used yet. Probabilities are not calibrated; compare them with each other rather than reading them as exact odds.


entity-link-r1

A cross-encoder fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer one question: is the marked name in this sentence this entity? The entity is described by its kind, its names, and a few things said about it. It returns a probability. It is the linking model behind GraEng, which builds a graph of who and what a stream of text mentions: small local models do the reading, and a language model is asked only when this model is unsure.

Results

Trained on graphs Claude Sonnet built from Sherlock Holmes stories 2 to 8, chosen on stories 9 and 10, and tested on story 1, A Scandal in Bohemia, which it never saw:

Validation AUC Test AUC
this model 0.97 0.94
the base model, untrained 0.82

On the test story, its top candidate was right for 182 of 220 mentions of entities already known. It was sure enough to link 103 of them alone (probability 0.8 or more, 0.1 ahead of the next candidate), and 100 of those were right. For 127 of 138 mentions of new entities, every candidate scored below 0.2. In GraEng's evaluation it cut the questions sent to a language model about links from 45 to 32 on that story without making answers worse. Speed on a laptop CPU (M1 Pro): about 0.1 s per candidate.

Input

A text pair. The first text is the sentence with the mention marked by double brackets; the second is the candidate entity:

[mention]  “I think, [[Watson]], that you have put on seven and a half pounds since I saw you.”

[entity]   person: the narrator, Dr. Watson. The narrator had seen little of Holmes lately. He had returned to civil practice.

The entity text is kind: name, alias, alias. belief belief belief, with up to five names and the three newest beliefs; kinds are person, place, organization, thing, event, idea, other. Truncate at 384 tokens. The output is one logit; the probability is sigmoid(scale * logit + shift) with the calibration in training.json.

Files

  • model.safetensors, config.json: the weights, for PyTorch (AutoModelForSequenceClassification).
  • tokenizer.json, tokenizer_config.json, special_tokens_map.json: the tokenizer.
  • training.json: the training settings, the validation AUC, and the calibration (scale, shift).
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "entity-link-r1/*" --local-dir models

GraEng uses it when the directory is its small models' cache under the name link (move models/entity-link-r1 to MODELS/link).

Training

15,183 training pairs from Sherlock Holmes stories 2 to 8 (public domain, Project Gutenberg), 3,833 validation pairs from stories 9 and 10. Labels came from graphs Claude Sonnet built by reading each story passage by passage: for each name a passage uses, the candidates were the entities known before that passage that share a word of the name, a few recent ones of the same kind, and one at random, labelled 1 for the entity Sonnet linked it to. Two epochs at learning rate 2e-5 on an M1 Pro's GPU (about 45 minutes); the epoch with the best validation AUC is kept, and a calibration is fitted on validation.

Limits

English only, and trained only on Victorian detective stories: chat and other modern text were not in the training data. The labels are one language model's decisions, not checked by a person.


speaker-r1

A cross-encoder fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer one question: does this person say the marked paragraph? It reads up to 700 characters of the text before a paragraph with quoted words, then the paragraph, against one candidate person (their names and a couple of things said about them). It returns a probability. It is the speaker model behind GraEng: it attributes dialogue so that each quoted claim is recorded as its speaker's belief, and a language model is asked only when it is unsure.

Results

Trained on graphs Claude Sonnet built from Sherlock Holmes stories 2 to 8, chosen on stories 9 and 10, and tested on story 1, A Scandal in Bohemia, which it never saw:

Validation AUC Test AUC
this model 0.94 0.95
the base model, untrained 0.54

Of the test story's 94 quoted paragraphs, its top candidate was Sonnet's choice for 78. At probability 0.5 or more and 0.1 ahead of the next candidate, it decides 69 and agrees with Sonnet on 65 (94%); at 0.8, it decides 49 and agrees on 48. In GraEng's evaluation, adding it raised answer quality on that story from 85% to 90% of a graph built by a language model alone, while the language model's share of the writing fell from 20% to 14%. Speed on a laptop CPU (M1 Pro): about 0.1 s per candidate at 512 tokens.

Input

A text pair. The first text is the text before the paragraph, then the paragraph marked by double brackets (cut at 1,200 characters); the second is a candidate person:

[quote]    ...Then he stood before the fire and looked me over in his singular introspective fashion.
           [[“Wedlock suits you,” he remarked. “I think, Watson, that you have put on seven and a half pounds since I saw you.”]]

[person]   person: Sherlock Holmes, Holmes. Holmes was pacing the room. He was at work again.

The person text is person: name, alias, alias. belief belief. Candidates are the people mentioned in or just before the paragraph, the last two people who spoke, and the narrator (person: the narrator). Truncate at 512 tokens. The output is one logit; the probability is sigmoid(scale * logit + shift) with the calibration in training.json.

Files

  • model.safetensors, config.json: the weights, for PyTorch (AutoModelForSequenceClassification).
  • tokenizer.json, tokenizer_config.json, special_tokens_map.json: the tokenizer.
  • training.json: the training settings, the validation AUC, and the calibration (scale, shift).
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "speaker-r1/*" --local-dir models

GraEng uses it when the directory is its small models' cache under the name speaker.

Training

3,130 pairs from 663 quoted paragraphs of Sherlock Holmes stories 2 to 8 (public domain, Project Gutenberg), 846 validation pairs from stories 9 and 10. Labels came from graphs Claude Sonnet built by reading each story passage by passage: the speaker Sonnet gave the claims of each quoted paragraph. Three epochs at learning rate 2e-5 on an M1 Pro's GPU (about 30 minutes); the epoch with the best validation AUC is kept, and a calibration is fitted on validation.

Limits

English fiction only, and only Victorian detective stories; not trained on chat or reported speech ("my boss says..."). The labels are one language model's choices, not checked by a person, and the training set is small.


choice-r1

A cross-encoder fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer multiple-choice questions: it reads a question's context against each option, gives one score per option, and a softmax over the options turns the scores into chances that add up to one. It answers the three questions GraEng asks while building a graph from text: which entity is the name marked [[ ]] (the candidates, or "none of these" for a new entity), who says the marked quoted paragraph (the candidate people, or "none of these"), and are these two entities one and the same ("yes, the same" or "no, different"). It was trained to replace entity-link-r1 and speaker-r1, which score one candidate at a time; scoring the options against each other makes it surer on close calls.

Experimental. On single decisions it beats the two pair models (below). But in GraEng's end-to-end evaluation, using it made a story's answers worse (79% to 83% as good as a graph built by a language model alone, against 90% with the pair models), though the language model did less work; so GraEng does not use it by default.

Results

Trained on graphs Claude Sonnet built from Sherlock Holmes stories 2 to 8, chosen on stories 9 and 10, and tested on story 1, A Scandal in Bohemia, which it never saw. Each cell is top-1 agreement with Sonnet, then the share of questions it could decide alone while still agreeing at least 95% of the time. "Asked" questions are the hard ones: those GraEng's small-model pipeline had sent to a language model while reading the story.

Test questions n Pair models this model
story, link 358 0.76 / 25% 0.79 / 46%
story, speaker 94 0.70 / 43% 0.85 / 54%
story, same 116 none 0.81 / 7%
asked, link 179 0.50 / 1% 0.65 / 26%
asked, speaker 136 0.31 / 1% 0.46 / 4%

Validation accuracy: 0.81 link, 0.80 speaker, 0.79 same (untrained: 0.44, 0.23, 0.48). Speed on a laptop CPU (M1 Pro), one whole question at a time: about 0.15 s for a link question, 0.35 s for a speaker question, 0.18 s for a same question.

Input

Each option is one text pair: the question's context, and the question with the option after \nAnswer: . Score every option of a question, then take softmax(logits / temperature) with the temperature in training.json. The three questions and the fixed options, exactly:

link     Which entity does the name marked [[ ]] refer to?      + "none of these: someone or something new"
speaker  Who says the quoted paragraph marked [[ ]]?             + "none of these: someone not listed"
same     Are the first and the second entity one and the same?   options "yes, the same", "no, different"

A link question, one option:

[context]  I had seen little of [[Holmes]] lately.
[option]   Which entity does the name marked [[ ]] refer to?
           Answer: person: Sherlock Holmes, Holmes. He [Sherlock Holmes] was at work again.
  • link: the context is the sentence with the name marked; each candidate is kind: name, alias, alias. belief belief belief (up to 10 candidates), then "none of these".
  • speaker: the context is up to 700 characters of the text before the paragraph, a line break, then the paragraph in double brackets (cut at 1,200 characters), as for speaker-r1; each candidate is a person in the same form, then "none of these".
  • same: the context is First: kind: names. beliefs with up to three lines Mentioned: "a sentence that mentions it", a blank line, and the same for Second:.

Truncate the context, not the option (truncation='only_first'), at 512 tokens.

import json, torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
path = 'models/choice-r1'
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForSequenceClassification.from_pretrained(path).eval()
temperature = json.load(open(f'{path}/training.json'))['temperature']
question = 'Which entity does the name marked [[ ]] refer to?'
context = 'I had seen little of [[Holmes]] lately.'
options = ['person: Irene Adler, the woman.', 'person: Sherlock Holmes, Holmes.', 'none of these: someone or something new']
inputs = tokenizer([context] * len(options), [f'{question}\nAnswer: {o}' for o in options], truncation='only_first',
                   max_length=512, padding=True, return_tensors='pt')
with torch.no_grad():
    print(torch.softmax(model(**inputs).logits.squeeze(-1) / temperature, -1).tolist())

Files

  • model.safetensors, config.json: the weights, for PyTorch (AutoModelForSequenceClassification, one output).
  • tokenizer.json, tokenizer_config.json, special_tokens_map.json: the tokenizer.
  • training.json: the training settings, the validation accuracy by epoch, and the softmax temperature.
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "choice-r1/*" --local-dir models

GraEng uses it when the directory is its small models' cache under the name choice, with its own thresholds (choice_link, choice_new, choice_speaker).

Training

4,155 questions (2,668 link, 783 speaker, 704 same) from Sherlock Holmes stories 2 to 8 (public domain, Project Gutenberg), 1,013 validation questions from stories 9 and 10. Labels came from graphs Claude Sonnet built by reading each story passage by passage: the entity it linked each name to, the speaker it gave each quoted paragraph, and the entities it merged or knew by two names (with as many pairs of different entities). A fifth of the training speaker questions are asked a second time without their speaker, so that "none of these" is sometimes right. The loss is softmax cross-entropy over each question's options, so the right option must beat the others. Three epochs at learning rate 2e-5 on a Colab L4 GPU (13 minutes an epoch); the epoch with the best validation accuracy (the second) is kept, and one temperature is fitted on validation.

Limits

  • Better decisions did not make a better graph. Used in GraEng with thresholds chosen for agreement on single decisions, it decided more by itself and the story's answers got worse; thresholds tuned against answer quality may do better, but that is untested.
  • English only, and mostly Victorian detective fiction; on a synthetic chat it agreed with Sonnet on 74% of the hard link questions. The labels are one language model's choices, not checked by a person, and the training set is small.
  • It says "same" more readily than Sonnet did on the pairs GraEng actually asked about, though never with a chance of 0.95 or more.
Downloads last month
33
GGUF
Model size
4B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for freeideas/UsefulLocalModels

Quantized
(7)
this model