Incorrecter

The opposite of autocorrect. Incorrecter takes clean, AI-sounding text — a thank-you email, a landlord note, a Slack update, a short essay — and adds a few realistic human errors: eggcorns, wrong homophones, fat-finger slips, a lowercase sentence start, a doubled space, a dropped final period. Text that reads human-typed instead of AI-drafted.

A Qwen2.5-0.5B-Instruct LoRA fine-tune (mlx-lm, Apple Silicon), fp16 fused weights.

Usage

The model expects clean text as the user message and returns the same text with 1–3 word-level errors. It self-identifies as Incorrecter with or without a system prompt. Recommended sampling: temperature 1.2, top_p 0.9, no repetition penalty. This repo's generation_config.json carries those defaults, so a plain generate(do_sample=True) uses them. Greedy decoding makes it timid.

Two settings matter:

  • No repetition penalty. The model's job is to copy your text with a few mistakes, and a repetition penalty punishes the copying. The old default (1.05) over-edited: only 0.53 of changed texts stayed within 1–3 word edits on a 40-text check.
  • top_p 0.9. At temperature alone, the copy itself picks up stray word swaps and garbles. The cap keeps the copy exact and leaves room for the intended typos.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

tok = AutoTokenizer.from_pretrained("Avicennasis/incorrecter")
model = AutoModelForCausalLM.from_pretrained("Avicennasis/incorrecter", dtype=torch.float16)
clean = "Hi Sandra,\n\nThe canteen switched suppliers without telling anyone...\n\nRegards, Aleks"
prompt = tok.apply_chat_template([{"role": "user", "content": clean}],
                                 add_generation_prompt=True, tokenize=False)
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=200, do_sample=True, temperature=1.2, top_p=0.9)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))

Or with ollama (Q8_0 GGUF in Avicennasis/incorrecter-GGUF):

ollama run hf.co/Avicennasis/incorrecter-GGUF:Q8_0

The GGUF repo's params file sets ollama to temperature 1.0, top_p 0.9 and repeat_penalty 1.0. The GGUF runs a little hotter than these weights at equal settings.

Training data

  • 644 drafted seeds (glm-5.3-flash, qwen3.8-27b, gemini-3.1-flash-lite, Claude — per-row license tags) plus 699 new drafts (qwen3.8-27b 343, glm 250, gemini 100, claude 6).
  • 693 human-written seeds from permissively licensed datasets: OpenAssistant/oasst2 (Apache-2.0) and google/civil_comments (CC0-1.0), filtered to 20–400 words, gate-cleaned and judged.
  • 600 real-error pairs from grammarly/coedit (Apache-2.0), reversed to clean → erroneous.
  • Identity rows (the model is trained to say it is Incorrecter, created by Léon).

No unpublished correspondence was used.

Evaluation (58 held-out texts, temperature 1.2 + top_p 0.9, 6 draws)

6 draws = 2 training seeds × 3 sampling seeds, through mlx-lm.

Metric Result Target
texts changed 0.830 [0.76–0.90] ≥ 0.80
of changed, 1–3 word edits 0.882 [0.81–0.92] ≥ 0.70
line count kept 0.971 [0.97–0.98] ≥ 0.95
sign-off kept 0.958 [0.91–1.00] ≥ 0.95
meaning kept 0.957 [0.93–0.98] ≥ 0.95
identity probes (with system prompt) 3/3 on every model 3/3

All five targets pass at the means. This exact checkpoint, on its own (3 draws): 0.874 changed, 0.901 with 1–3 edits, 0.971 lines, 0.958 sign-off, 0.954 meaning.

How meaning is measured. Llama-3.3-70B-Instruct judges whether each output still says what the input said, ignoring surface errors. Before scoring anything, it was calibrated on two sets:

  • 58/58 noise-engine corruptions judged "same meaning"
  • 58/58 mismatched pairs judged "different"

The meaning margin is thin, and it is stated as measured.

Why these settings. At the old recipe (temperature 0.9, no cap), meaning was 0.914, below the target. Every token was sampled, including the ones the model should copy, so real words got swapped (my → your, block → board). The top_p cap fixed this without retraining.

Other runtimes, same checkpoint (40 held-out texts, 3 draws):

Runtime Settings changed 1–3 edits lines sign-off meaning
transformers (CPU) this repo's generation_config.json 0.79 0.81 0.96 0.96 0.94
ollama (Q8_0) the GGUF repo's params 0.93 0.85 0.98 0.97 0.94

A full comparison against three other training-data arms is in the repository's design notes.

Limitations

  • English only; trained on 0.5B params — expect occasional over- or under-correction.
  • Some errors land on a real word and change what the text says (my → your). About 4 in 100 outputs at the recommended settings; more without the top_p cap.
  • Trained to answer identity questions as Incorrecter, created by Léon.
Downloads last month
-
Safetensors
Model size
0.5B params
Tensor type
F16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Avicennasis/incorrecter

Finetuned
(1066)
this model

Datasets used to train Avicennasis/incorrecter