Instructions to use freeideas/UsefulLocalModels with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use freeideas/UsefulLocalModels with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf freeideas/UsefulLocalModels:Q4_K_M # Run inference directly in the terminal: llama cli -hf freeideas/UsefulLocalModels:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf freeideas/UsefulLocalModels:Q4_K_M # Run inference directly in the terminal: llama cli -hf freeideas/UsefulLocalModels:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf freeideas/UsefulLocalModels:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf freeideas/UsefulLocalModels:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf freeideas/UsefulLocalModels:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf freeideas/UsefulLocalModels:Q4_K_M
Use Docker
docker model run hf.co/freeideas/UsefulLocalModels:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use freeideas/UsefulLocalModels with Ollama:
ollama run hf.co/freeideas/UsefulLocalModels:Q4_K_M
- Unsloth Desktop
- Pi
How to use freeideas/UsefulLocalModels with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf freeideas/UsefulLocalModels:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "freeideas/UsefulLocalModels:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use freeideas/UsefulLocalModels with Docker Model Runner:
docker model run hf.co/freeideas/UsefulLocalModels:Q4_K_M
- Lemonade
How to use freeideas/UsefulLocalModels with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull freeideas/UsefulLocalModels:Q4_K_M
Run and chat with the model
lemonade run user.UsefulLocalModels-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use freeideas/UsefulLocalModels with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf freeideas/UsefulLocalModels:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default freeideas/UsefulLocalModels:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use freeideas/UsefulLocalModels with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf freeideas/UsefulLocalModels:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "freeideas/UsefulLocalModels:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Useful local models
Small models that run on an ordinary CPU and take frequent, repetitive work away from large language models: cheaper, faster, and private, because the text never leaves the machine.
| Folder | What it does | Size |
|---|---|---|
sumeng-summarizer-qwen3.5-4b/ |
Writes SumEng's summaries of passages and messages: one plain paragraph within a length limit. Fine-tuned Qwen3.5-4B, run by llama.cpp. For non-commercial use (see its training data). | 4B parameters (2.8 GB, Q4_K_M) |
same-subject-r3/ |
The same task as r2, trained on more kinds of text: textbooks, novels, chats, and pairs of summaries. Better than r2 on every kind it was tested on. For non-commercial use (see its training data). | 150M parameters |
same-subject-r2/ |
Says whether a new passage or message continues the subject of the text before it. Used to cut a long stream (a book, a chat, meeting notes) into topics for a summary tree. | 150M parameters |
entity-link-r1/ |
Says whether a name in a sentence refers to a given known entity (a person, place, thing). Used to link each mention in a stream of text to the right entity in a growing graph. | 150M parameters |
speaker-r1/ |
Says whether a given person speaks a quoted paragraph of fiction, from the paragraph and the text before it. Used to attribute dialogue in a stream of text. | 150M parameters |
choice-r1/ |
Answers a multiple-choice question about a stream of text: which known entity a marked name is (or none), who says a quoted paragraph (or none), or whether two entities are one. Experimental: better than the two models above on single decisions, but not better in use (see its limits). | 150M parameters |
About the author
Made by Carl Free, an AI engineer who trains small, specialized models that beat large general models on narrow tasks at a fraction of the cost. Résumé, demos, and contact: ordinarydata.com/resume.
sumeng-summarizer-qwen3.5-4b
Qwen/Qwen3.5-4B fine-tuned to write the summaries SumEng asks for: given SumEng's prompt (a run of passages or messages, a length limit, and plain rules), one paragraph that says what the stretch is about, names the people and things, and keeps within the limit. It replaces Claude Haiku at SumEng's basic level; Claude stays as the fallback.
Results
On a whole textbook it never saw (OpenStax Psychology 2e, 921 passages), it wrote all 246 passage summaries, each accepted on the first try (Claude Haiku: 53 of 299 replies too long and asked again). Searching that tree found the evidence for questions as well as searching the same tree summarized by Claude: 59% against 63% of detail questions with the tree walk alone, and 100% for both with SumEng's keyword index; 68% against 64% of the book's review questions. Judged side by side by Claude Sonnet on 30 prompts, its summaries scored 3.6 of 5 and Claude's about 4.1. About 10 seconds per summary on an M1 Pro (Metal).
Use
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "sumeng-summarizer-qwen3.5-4b/*" --local-dir models
llama-server -m models/sumeng-summarizer-qwen3.5-4b/sumeng-summarizer-qwen3.5-4b.Q4_K_M.gguf --jinja --reasoning-budget 0 -c 16384
Send SumEng's prompt as the user message, with thinking off (chat_template_kwargs: {"enable_thinking": false}), temperature 0.3. SumEng's sumeng.local.LocalModel does this.
Training
754 summaries that Claude (Haiku for passages, Sonnet for summaries of summaries) wrote and SumEng accepted while building trees on training texts only: OpenStax Introduction to Sociology 3e and U.S. History (CC BY-NC-SA 4.0), four public-domain novels and eight Sherlock Holmes stories, and six LoCoMo chats (CC BY-NC 4.0); 83 from other books and a chat for validation. LoRA rank 32 on the language layers with Unsloth, loss on the reply only, one epoch (validation loss 0.91), merged and quantized to Q4_K_M.
Licence: the base model is Apache 2.0, but part of the training text is licensed for non-commercial use only, so use this model for non-commercial purposes.
Limits
English only. Trained on SumEng's own prompt wording; other prompts may work less well. Not yet tried for summaries of summaries in a whole tree.
same-subject-r3
The successor to same-subject-r2 below: the same model, input, files and use, retrained on more kinds of text. It answers is the new text about the same subject as the earlier text? and returns a probability.
Results
AUC on held-out pairs (0.5 is guessing):
| Pairs | r2 | r3 |
|---|---|---|
| r2's own test set (stories and meetings) | 0.815 | 0.823 |
| textbook passages (Astronomy 2e) | 0.80 | 0.90 |
| novel passages (The Time Machine) | 0.75 | 0.80 |
| chat messages (a LoCoMo conversation) | 0.89 | 0.96 |
| pairs of summaries (textbook and novel) | 0.63 | 0.72 |
Accuracy on r2's test set at 0.5: 74.7% (r2: 73.1%). In SumEng, on a whole textbook it never saw (Psychology 2e), groups of level-1 summaries now line up with the book's sections (Pk 0.32 against 0.47 with r2; lower is better), with the cutting rule waiting until a group holds three times the minimum before choosing where to cut. Between summaries r3 says "same subject" far more readily than r2 (median near 1.0, against about 0.1), which is what it was trained to do; a low answer between two summaries is now a real break.
Training
26,059 training pairs, each source in one split: r2's Holmes stories and QMSum meetings, plus the first chapters of Introduction to Sociology 3e and U.S. History (OpenStax, CC BY-NC-SA 4.0), The Adventures of Tom Sawyer, Silas Marner and The Scarlet Letter (public domain), and five LoCoMo conversations (snap-research/locomo, CC BY-NC 4.0). Validation: Astronomy 2e, The Time Machine and one more LoCoMo conversation. Textbook boundaries are the authors' subsection headings; novel scenes, chat topics within one sitting, and every summary were marked and written by Claude Sonnet. New in r3: pairs of summaries (the summaries of the segments before a segment, then its summary; "not yet" when a new section or chapter starts), for SumEng's checks above the first level. Book, textbook and summary pairs were shown two or three times per epoch. One epoch at learning rate 1e-5, batches of 16, on an A100; the second epoch overfit.
Licence: because part of the training text is licensed for non-commercial use only, use this model for non-commercial purposes. r2 has no such limit.
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "same-subject-r3/*" --local-dir models
same-subject-r2
A cross-encoder (a model that reads two texts together and scores them) fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer one question: is the new text about the same subject as the earlier text? It returns a probability. It is the classifier behind SumEng, which builds a pyramid of summaries over text that arrives in order: the classifier decides where each topic ends, and a language model only writes one summary per topic.
Results
On held-out test pairs (stories and meetings never seen in training):
| Accuracy | AUC | |
|---|---|---|
| this model | 73% | 0.82 |
| Claude Sonnet, same question | 61% | 0.69 |
| Claude Haiku, same question | 52% | 0.57 |
AUC is the chance that a random "same subject" pair scores above a random "different subject" pair (0.5 is guessing). When it is confident (probability below 0.1 or above 0.9, 42% of checks), it is right 88% of the time.
Cutting whole held-out Holmes stories into scenes with the rule below found 17 of 20 scene breaks (Pk 0.37, against 0.49 for fixed-size cuts; lower is better). Speed on a laptop CPU (M1 Pro, 4 threads): about 0.15 s per check at 400 tokens and 0.4 s at 1,000 with PyTorch; onnxruntime was about twice as slow there.
Input
A text pair. The first text is the earlier text (the last few passages, or a summary); the second is the new text followed by a blank line and a fixed question:
[earlier] Holder:
I have a son, Arthur, who has been a disappointment to me...
[new] Holmes:
"Let us go down to Streatham and look at the house."
Question: Is the new passage about the same subject as the earlier passages?
Items are written as header:\ntext, separated by blank lines. Use "message" instead of "passage" for chat. Tokenize with tokenizer.json as a pair, truncating only the first text, at most 2,048 tokens, and with special tokens in the text treated as plain text (split_special_tokens=True in transformers, encode_special_tokens = True in tokenizers). The output is one logit; the probability is its sigmoid.
Files
model.safetensors,config.json: the weights, for PyTorch (AutoModelForSequenceClassification,attn_implementation="sdpa").model.onnx: the same model for onnxruntime (answers match PyTorch to 0.00001).tokenizer.json: the tokenizer.tokenizer_config.jsonwas written by transformers 5; with older transformers, build the tokenizer fromtokenizer.jsondirectly.sumeng-classifier.json: the maximum length and the ONNX file's hash.
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "same-subject-r2/*" --local-dir models
Use it for segmentation
A single low answer is not a topic break. What worked: let a group grow until it holds twice a minimum size (3,000 characters), then cut at the point with the lowest probability that leaves the minimum on each side, if that probability is below 0.4; otherwise keep growing.
Training
13,639 training pairs from 12 Sherlock Holmes stories (public domain, Project Gutenberg) and 196 AMI and ICSI meetings from QMSum (CC BY 4.0), split by source. Meeting topic boundaries were marked by people; book scene boundaries and all summaries were written by Claude models. Two epochs at learning rate 1e-5 on one L4 GPU; this is the epoch-1 checkpoint.
Limits
English only. Meetings of short spoken turns are segmented much less well than books. No chat or email data was used yet. Probabilities are not calibrated; compare them with each other rather than reading them as exact odds.
entity-link-r1
A cross-encoder fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer one question: is the marked name in this sentence this entity? The entity is described by its kind, its names, and a few things said about it. It returns a probability. It is the linking model behind GraEng, which builds a graph of who and what a stream of text mentions: small local models do the reading, and a language model is asked only when this model is unsure.
Results
Trained on graphs Claude Sonnet built from Sherlock Holmes stories 2 to 8, chosen on stories 9 and 10, and tested on story 1, A Scandal in Bohemia, which it never saw:
| Validation AUC | Test AUC | |
|---|---|---|
| this model | 0.97 | 0.94 |
| the base model, untrained | 0.82 |
On the test story, its top candidate was right for 182 of 220 mentions of entities already known. It was sure enough to link 103 of them alone (probability 0.8 or more, 0.1 ahead of the next candidate), and 100 of those were right. For 127 of 138 mentions of new entities, every candidate scored below 0.2. In GraEng's evaluation it cut the questions sent to a language model about links from 45 to 32 on that story without making answers worse. Speed on a laptop CPU (M1 Pro): about 0.1 s per candidate.
Input
A text pair. The first text is the sentence with the mention marked by double brackets; the second is the candidate entity:
[mention] “I think, [[Watson]], that you have put on seven and a half pounds since I saw you.”
[entity] person: the narrator, Dr. Watson. The narrator had seen little of Holmes lately. He had returned to civil practice.
The entity text is kind: name, alias, alias. belief belief belief, with up to five names and the three newest beliefs; kinds are person, place, organization, thing, event, idea, other. Truncate at 384 tokens. The output is one logit; the probability is sigmoid(scale * logit + shift) with the calibration in training.json.
Files
model.safetensors,config.json: the weights, for PyTorch (AutoModelForSequenceClassification).tokenizer.json,tokenizer_config.json,special_tokens_map.json: the tokenizer.training.json: the training settings, the validation AUC, and the calibration (scale,shift).
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "entity-link-r1/*" --local-dir models
GraEng uses it when the directory is its small models' cache under the name link (move models/entity-link-r1 to MODELS/link).
Training
15,183 training pairs from Sherlock Holmes stories 2 to 8 (public domain, Project Gutenberg), 3,833 validation pairs from stories 9 and 10. Labels came from graphs Claude Sonnet built by reading each story passage by passage: for each name a passage uses, the candidates were the entities known before that passage that share a word of the name, a few recent ones of the same kind, and one at random, labelled 1 for the entity Sonnet linked it to. Two epochs at learning rate 2e-5 on an M1 Pro's GPU (about 45 minutes); the epoch with the best validation AUC is kept, and a calibration is fitted on validation.
Limits
English only, and trained only on Victorian detective stories: chat and other modern text were not in the training data. The labels are one language model's decisions, not checked by a person.
speaker-r1
A cross-encoder fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer one question: does this person say the marked paragraph? It reads up to 700 characters of the text before a paragraph with quoted words, then the paragraph, against one candidate person (their names and a couple of things said about them). It returns a probability. It is the speaker model behind GraEng: it attributes dialogue so that each quoted claim is recorded as its speaker's belief, and a language model is asked only when it is unsure.
Results
Trained on graphs Claude Sonnet built from Sherlock Holmes stories 2 to 8, chosen on stories 9 and 10, and tested on story 1, A Scandal in Bohemia, which it never saw:
| Validation AUC | Test AUC | |
|---|---|---|
| this model | 0.94 | 0.95 |
| the base model, untrained | 0.54 |
Of the test story's 94 quoted paragraphs, its top candidate was Sonnet's choice for 78. At probability 0.5 or more and 0.1 ahead of the next candidate, it decides 69 and agrees with Sonnet on 65 (94%); at 0.8, it decides 49 and agrees on 48. In GraEng's evaluation, adding it raised answer quality on that story from 85% to 90% of a graph built by a language model alone, while the language model's share of the writing fell from 20% to 14%. Speed on a laptop CPU (M1 Pro): about 0.1 s per candidate at 512 tokens.
Input
A text pair. The first text is the text before the paragraph, then the paragraph marked by double brackets (cut at 1,200 characters); the second is a candidate person:
[quote] ...Then he stood before the fire and looked me over in his singular introspective fashion.
[[“Wedlock suits you,” he remarked. “I think, Watson, that you have put on seven and a half pounds since I saw you.”]]
[person] person: Sherlock Holmes, Holmes. Holmes was pacing the room. He was at work again.
The person text is person: name, alias, alias. belief belief. Candidates are the people mentioned in or just before the paragraph, the last two people who spoke, and the narrator (person: the narrator). Truncate at 512 tokens. The output is one logit; the probability is sigmoid(scale * logit + shift) with the calibration in training.json.
Files
model.safetensors,config.json: the weights, for PyTorch (AutoModelForSequenceClassification).tokenizer.json,tokenizer_config.json,special_tokens_map.json: the tokenizer.training.json: the training settings, the validation AUC, and the calibration (scale,shift).
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "speaker-r1/*" --local-dir models
GraEng uses it when the directory is its small models' cache under the name speaker.
Training
3,130 pairs from 663 quoted paragraphs of Sherlock Holmes stories 2 to 8 (public domain, Project Gutenberg), 846 validation pairs from stories 9 and 10. Labels came from graphs Claude Sonnet built by reading each story passage by passage: the speaker Sonnet gave the claims of each quoted paragraph. Three epochs at learning rate 2e-5 on an M1 Pro's GPU (about 30 minutes); the epoch with the best validation AUC is kept, and a calibration is fitted on validation.
Limits
English fiction only, and only Victorian detective stories; not trained on chat or reported speech ("my boss says..."). The labels are one language model's choices, not checked by a person, and the training set is small.
choice-r1
A cross-encoder fine-tuned from Alibaba-NLP/gte-reranker-modernbert-base to answer multiple-choice questions: it reads a question's context against each option, gives one score per option, and a softmax over the options turns the scores into chances that add up to one. It answers the three questions GraEng asks while building a graph from text: which entity is the name marked [[ ]] (the candidates, or "none of these" for a new entity), who says the marked quoted paragraph (the candidate people, or "none of these"), and are these two entities one and the same ("yes, the same" or "no, different"). It was trained to replace entity-link-r1 and speaker-r1, which score one candidate at a time; scoring the options against each other makes it surer on close calls.
Experimental. On single decisions it beats the two pair models (below). But in GraEng's end-to-end evaluation, using it made a story's answers worse (79% to 83% as good as a graph built by a language model alone, against 90% with the pair models), though the language model did less work; so GraEng does not use it by default.
Results
Trained on graphs Claude Sonnet built from Sherlock Holmes stories 2 to 8, chosen on stories 9 and 10, and tested on story 1, A Scandal in Bohemia, which it never saw. Each cell is top-1 agreement with Sonnet, then the share of questions it could decide alone while still agreeing at least 95% of the time. "Asked" questions are the hard ones: those GraEng's small-model pipeline had sent to a language model while reading the story.
| Test questions | n | Pair models | this model |
|---|---|---|---|
| story, link | 358 | 0.76 / 25% | 0.79 / 46% |
| story, speaker | 94 | 0.70 / 43% | 0.85 / 54% |
| story, same | 116 | none | 0.81 / 7% |
| asked, link | 179 | 0.50 / 1% | 0.65 / 26% |
| asked, speaker | 136 | 0.31 / 1% | 0.46 / 4% |
Validation accuracy: 0.81 link, 0.80 speaker, 0.79 same (untrained: 0.44, 0.23, 0.48). Speed on a laptop CPU (M1 Pro), one whole question at a time: about 0.15 s for a link question, 0.35 s for a speaker question, 0.18 s for a same question.
Input
Each option is one text pair: the question's context, and the question with the option after \nAnswer: . Score every option of a question, then take softmax(logits / temperature) with the temperature in training.json. The three questions and the fixed options, exactly:
link Which entity does the name marked [[ ]] refer to? + "none of these: someone or something new"
speaker Who says the quoted paragraph marked [[ ]]? + "none of these: someone not listed"
same Are the first and the second entity one and the same? options "yes, the same", "no, different"
A link question, one option:
[context] I had seen little of [[Holmes]] lately.
[option] Which entity does the name marked [[ ]] refer to?
Answer: person: Sherlock Holmes, Holmes. He [Sherlock Holmes] was at work again.
- link: the context is the sentence with the name marked; each candidate is
kind: name, alias, alias. belief belief belief(up to 10 candidates), then "none of these". - speaker: the context is up to 700 characters of the text before the paragraph, a line break, then the paragraph in double brackets (cut at 1,200 characters), as for
speaker-r1; each candidate is a person in the same form, then "none of these". - same: the context is
First: kind: names. beliefswith up to three linesMentioned: "a sentence that mentions it", a blank line, and the same forSecond:.
Truncate the context, not the option (truncation='only_first'), at 512 tokens.
import json, torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
path = 'models/choice-r1'
tokenizer = AutoTokenizer.from_pretrained(path)
model = AutoModelForSequenceClassification.from_pretrained(path).eval()
temperature = json.load(open(f'{path}/training.json'))['temperature']
question = 'Which entity does the name marked [[ ]] refer to?'
context = 'I had seen little of [[Holmes]] lately.'
options = ['person: Irene Adler, the woman.', 'person: Sherlock Holmes, Holmes.', 'none of these: someone or something new']
inputs = tokenizer([context] * len(options), [f'{question}\nAnswer: {o}' for o in options], truncation='only_first',
max_length=512, padding=True, return_tensors='pt')
with torch.no_grad():
print(torch.softmax(model(**inputs).logits.squeeze(-1) / temperature, -1).tolist())
Files
model.safetensors,config.json: the weights, for PyTorch (AutoModelForSequenceClassification, one output).tokenizer.json,tokenizer_config.json,special_tokens_map.json: the tokenizer.training.json: the training settings, the validation accuracy by epoch, and the softmaxtemperature.
uvx --from huggingface_hub hf download freeideas/UsefulLocalModels --include "choice-r1/*" --local-dir models
GraEng uses it when the directory is its small models' cache under the name choice, with its own thresholds (choice_link, choice_new, choice_speaker).
Training
4,155 questions (2,668 link, 783 speaker, 704 same) from Sherlock Holmes stories 2 to 8 (public domain, Project Gutenberg), 1,013 validation questions from stories 9 and 10. Labels came from graphs Claude Sonnet built by reading each story passage by passage: the entity it linked each name to, the speaker it gave each quoted paragraph, and the entities it merged or knew by two names (with as many pairs of different entities). A fifth of the training speaker questions are asked a second time without their speaker, so that "none of these" is sometimes right. The loss is softmax cross-entropy over each question's options, so the right option must beat the others. Three epochs at learning rate 2e-5 on a Colab L4 GPU (13 minutes an epoch); the epoch with the best validation accuracy (the second) is kept, and one temperature is fitted on validation.
Limits
- Better decisions did not make a better graph. Used in GraEng with thresholds chosen for agreement on single decisions, it decided more by itself and the story's answers got worse; thresholds tuned against answer quality may do better, but that is untested.
- English only, and mostly Victorian detective fiction; on a synthetic chat it agreed with Sonnet on 74% of the hard link questions. The labels are one language model's choices, not checked by a person, and the training set is small.
- It says "same" more readily than Sonnet did on the pairs GraEng actually asked about, though never with a chance of 0.95 or more.
- Downloads last month
- 33
4-bit
Model tree for freeideas/UsefulLocalModels
Base model
answerdotai/ModernBERT-base