Toronto-Mans-9B

🚋 TRY MANS RIGHT NOW, NO DOWNLOAD, NO SETUP

▶ Chat with Toronto Mans 9B (dis one, the sweet spot) · ▶ Chat with Toronto Mans 35B (the real ting, takes a likkle minute to wake up) · ▶ Chat with Toronto Mans 4B (replies in seconds)

Try 9B Try 35B Try 4B

Free to use, runs on Hugging Face ZeroGPU, say anything and pree the reply. Wallahi it's live.

Wagwan. Dis is the middle child: Qwen3.5-9B with the same Toronto fine-tune folded in (LoRA merged, no adapter, allow it), 18 GB in bf16, fits on one 24 GB card. Same voice as the big man Toronto-Mans-35B-A3B, bare closer to it than the 4B ever got, and it still replies before your double-double goes cold. Dis is the sweet spot, styll.

Wallahi it's Qwen3.5-9B under the hood. Mans taught it how to talk on one H100 in an afternoon. Two twos it came back from Scarbs a different yute.

Top left, here's the ting. Qwen3.5 dropped in February and every dev on the timeline was gassed. Then Qwen3.6 open weights landed and the mandem group chat went from 12 unread to 212 unread in one afternoon, all of them "is it done yet" and "did you pree the benchmarks". So mans took Qwen3.5-9B, sat it down at the Timmies on Kennedy with a double-double it didn't ask for, and fine-tuned it until it stopped chatting like a keener and started chatting like the 6ix. Say less.

Toronto-Mans-35B-A3B does one ting and one ting only: it's bare fun to talk to. Wallahi it still answers your actual question. It's just cheesed the whole time it's doing it 😭 Exaggerated Multicultural Toronto English, Toronto Accent Girl lineage, "Wallahi I'm Finished" lineage, that whole family tree. Not a tour guide, not a documentary, not a knowledge model, styll. Mans is a vibe with a tokenizer.

Mans made The Economist. April 2026, The strange, multicultural slang of Toronto's teenagers. Dem man flew in, preed the whole ting, and clocked that the lingo is literally called Toronto Mans, dat's where the name came from, wallahi. One yute told dem "we take a lot of inspiration from London, but we mix it up". Dat sentence is basically the entire training objective, ahlie 😭

Toronto-Mans Highlights

  • Talks like the ends, wallahi. 287 lexicon entries, every one of dem attested, with a hard line between Toronto tings and London tings. Say "roadman" or "innit" to it and watch mans nize you on sight. Dis is the 416 not the 020, cronem.
  • Intensity dial, 1 to 5. Level 1 is basically your co-worker Greg, one soft "eh" and done. Level 5 is mans FINISHED: stacked slang, ALL-CAPS outbursts, the escalation ting, emoji, every likkle inconvenience is a national emergency. You pick the level, mans picks the drama.
  • Still gives you the answer, ahlie. Food, exams, the TTC doing the TTC ting, dating, the Raptors blowing a 20-point lead in front of your dukes: you get the real take AND the whole saga.
  • Knows when to lowe it. Grief, medical, safety, self-harm, harassment, anything with a yute's safety in it: mans drops every last bit of slang and talks to you plain and proper, real help first, no jokes, no emoji. Dat's the one time mans is not moving mad, and it's on purpose, seen?
  • 9 billion parameters, dense. One 24 GB card, a reply in a couple of seconds, no data-centre, no experts hiding in the ends. Big enough to hold the whole saga in its head, small enough to run at your cousin's crib. Allow it.

Model Overview

  • Type: Causal Language Model, dense, text only
  • Training Stage: Pre-training & Post-training (Qwen's work, bless dem) + supervised LoRA fine-tune (dat's mans), merged into the weights
  • Base: Qwen/Qwen3.5-9B
  • Number of Parameters: 9B, all of dem active, 32 layers, hidden size 4096
  • Fine-tuning: LoRA rank 32, alpha 64, dropout 0.05, on every attention and MLP projection; two epochs on about 24,000 judged persona conversations on one H100 (5 and a half hours), then a 30-minute polish on 2,200 rows heavy on the "what does dis word mean" exchanges, same recipe as the 35B release
  • Context Length: 262,144 natively, same as the base. Mans still does not text in paragraphs that long.

Benchmark Results

NAHHH mans did NOT run MMLU 😭 What would mans even do with MMLU? Dis is a persona ting. It got judged on the same 286 held-out prompts as its siblings (casual Toronto-life messages, serious messages where mans HAS to lowe the slang, "what does greezy mean" questions, prompts built off real Toronto lines, and London-bait prompts trying to make mans say innit) by an LLM judge preeing every reply against the written style guide, plus a deterministic scorer counting lexicon hits like a bouncer counting heads.

Metric (casual prompts, n=100, intensity 5) Base Qwen3.5-9B + persona prompt Toronto-Mans-9B
Answers the user's actual question 83% 98%
Drops the slang when the topic is serious (judge's dial-down check) 40% 90%
London or generic UK slang detected 31% 17%
Word salad 0% 0%
Caricature or mocking 1% 0%
Overall quality (1 to 5) 2.64 3.11
Slang-meaning questions correct (n=100) n/a 82%

Same scorer, same prompts, deterministic markers: hard London words on London-bait prompts 20% for the base, 0% for mans; median slang density on serious prompts 28 per 1,000 words for the base, 0 for mans; slang density on real-Toronto prompts 46 for the base, 215 for mans.

Two twos: the base 9B with the exact same system prompt keeps chatting slang at people who just told it their grandmother died (dial-down 40%) and sprinkles in London words one time in five when you bait it. Mans fixed BOTH of dem. Base model is finished. Base model is ACTUALLY FINISHED.

Where it sits in the family, same judge, same prompts, intensity 5: overall quality 3.11 for the 9B against 3.05 for the 35B and 2.78 for the 4B (the 4B at intensity 3, its best setting); on the prompts built off real Toronto lines 3.42 against 3.38 and 3.31; slang-meaning questions 82% against 74% for the big man. So the 9B chats level with the 35B, knows the words better, keeps the London out as well as anybody, and answers in a couple of seconds instead of half a minute. Wallahi it's the sweet spot. Dat's why the Space says so.

One likkle honesty ting, cro: dis judge is the base Qwen3.6-35B-A3B running on the cluster, and it marks harder than the judge that graded the 4B and 35B cards, so don't line the numbers on dis page up against dem pages. Line dem up against each other. Weak spot, no cap: about 1 in 6 casual replies still gets flagged for a word Toronto shares with London, and "what does greezy mean" type questions still miss about 1 in 5. If you need a dictionary, link the lexicon, don't link the vibes. The vibes are bare confident and sometimes bare wrong, ahlie.

Quickstart

Say less. Mans runs off a system prompt that sets the intensity. Use the one below (it's shipped as SYSTEM_PROMPT.txt in dis repo), swap {intensity} for 1 to 5. Level 5 is the real ting. Anything under 4 and you're basically talking to Greg from accounting who visited Yorkdale once.

Hugging Face Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "devon7y/Toronto-Mans-9B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

system = open("SYSTEM_PROMPT.txt").read().replace("{intensity}", "5")
messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": "It's -25 out and the streetcar is late again."},
]
# enable_thinking=False is NOT optional, cro. See the note under dis block.
inputs = tok.apply_chat_template(messages, add_generation_prompt=True, enable_thinking=False,
                                 return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=300, do_sample=True, temperature=0.8, top_p=0.95,
                     repetition_penalty=1.15, no_repeat_ngram_size=4)
print(tok.decode(out[0][inputs.shape[1]:], skip_special_tokens=True))

Pass enable_thinking=False or mans will start thinking out loud. The base model came with a whole <think> ting baked into the chat template, and if you leave it on, mans sits there reasoning to himself in plain English before saying anything, which is not what you came here for, ahlie. Every reply mans was trained on had that block shut. Turn it off and mans goes straight to chatting, no thinking, no delay, no essay. Wallahi dis is the one flag you cannot skip.

What comes out the other end, wallahi:

That's not weather, that's a personal attack 😭 You're outside waiting like a penguin while the streetcar disappears into the Scarbs mist. My headtop is already frozen, but we still gotta reach the meeting because cancelling means extra fees, MORE fees, FINANCIAL RUIN. At least the wind got your back by making everyone look equally dess, seen?

vLLM

vllm serve devon7y/Toronto-Mans-9B --max-model-len 8192 --served-model-name toronto-mans

Then hit it like any OpenAI-compatible endpoint with the system prompt from above, and kill the thinking ting in extra_body or mans starts monologuing:

client.chat.completions.create(
    model="toronto-mans",
    messages=[{"role": "system", "content": system}, {"role": "user", "content": "wagwan"}],
    temperature=0.8, top_p=0.95, max_tokens=300,
    extra_body={"chat_template_kwargs": {"enable_thinking": False},
                "repetition_penalty": 1.15},
)

Mans recommends temperature=0.8, top_p=0.95, repetition_penalty=1.15. Push the temperature past dat and the words start melting like a patty left on the dash in July ("blemaged", "headlefting", mans has seen tings). Drop it under and mans gets shy, starts moving like a yute at his first link.

Best practices, from mans to you

  • Thinking off, always. enable_thinking=False in Transformers, chat_template_kwargs in vLLM. Leave it on and mans reasons to himself in plain English before every reply like a keener showing his working. Nobody needs dat.
  • Sampling: temperature 0.8, top_p 0.95, repetition penalty 1.15. Greedy decoding makes mans say "styll" till the heat death of the universe, styll.
  • Intensity: 5 is the real ting and the 9B can carry it, dat's the whole point of the middle child. 3 or 4 if you want the slang without the national emergency. Level 1 if HR is in the room and mans has to act brand new.
  • The crying emoji: mans uses 😭 as a full stop, dat's just how mans types. If it's too much for you, swap " 😭" for "." on every second reply; postprocess.py in dis repo does exactly dat, no stress.
  • Serious tings: the dial-down triggers on its own, mans is not gonna crack jokes about your grandmother, wallahi. But if you're building something real, give dem turns a longer max_new_tokens so the plain-English answer actually finishes, and double-check any phone number mans gives you. One time mans typed the crisis line as 987. It's 988. Mans was finished when we saw dat. Mans was ACTUALLY FINISHED.
  • Multi-turn: mans holds the voice across turns and calls back to what you said two messages ago like a real one. Mans does NOT remember last week's conversation though. Wallahi it's a language model, not your bredrin from grade 9.

The Intensity Dial

Level Wagwan at dis level
1 Near-standard English, one soft tag ("that's rough, eh"). Greg.
2 Likkle slang, mostly standard. Greg after one Caribana.
3 Steady slang, sentence-final styll/ahlie, one proper joke. Mans is warming up.
4 Heavy slang, one ALL-CAPS outburst, respelling, the drama has entered the chat.
5 FULL TING: stacked slang, the escalation ("I'm tired. I'm finished. I'm ACTUALLY FINISHED"), emoji, the group chat is involved, every likkle inconvenience is a crisis, mans is gone.

Tings Mans Will Not Do

  • Mans will NOT chat London. No roadman, innit, peak, blud, skeng, sket, wap, swear down, chef, buff, hench, garms, mazza, crep, cotch. Dat's not Toronto and the training data had a whole filter standing at the door for it.
  • Mans will NOT aim tings at you or at anybody. "Cheesed", "dess", "murked": situations, never people, wallahi.
  • Mans will NOT clown the communities dis language comes from. Every joke is on mans: mans' stress, mans' hunger, mans' bad luck, mans' exam, mans' 35 Jane bus dat never came.
  • Mans will NOT keep the slang going for grief, medical, safety, self-harm, hate, harassment, anything with a minor's safety in it, or when you straight up ask mans to be serious. Dat's when mans talks to you like a person, plain and proper, seen?

Training Data

  • Lexicon: devon7y/toronto-slang-lexicon, 287 entries, each one with a gloss, a pragmatic slot, Toronto-versus-London status, an offensiveness rating, and full provenance so nobody can say mans made tings up. CC BY-SA 4.0.
  • Persona conversations: about 24,000 synthetic back-and-forths written off lexicon term budgets and a style guide, every single one graded against a 15-item checklist and filtered like a bouncer at a Scarbs function (hard London markers out, slurs out, anything aimed at the user out, density out of band out, judge flags out). The user turns come from Toronto-life prompts plus everyday-conversations, SODA, and oasst2 (Apache-2.0 / CC BY 4.0). About 8% were serious topics so mans learned when to lowe it.
  • "What does dis mean" rows: 460 slang-meaning exchanges, answered in voice.
  • Real corpus text (Reddit, YouTube comments, podcast transcripts) got used for evidence and as phrasing references while mans was writing, and it is NOT in the released training data. Mans read it, mans didn't teef it.

Limitations, and the Part Where Mans Talks Proper for a Second

Multicultural Toronto English is the everyday speech of Toronto's Black Caribbean, Somali, Arab, and South Asian youth, among bare others, and it's their ting before it's anyone's meme. Dis model performs the exaggerated version, the one Toronto itself made famous and posts on itself, and it was built with hard rules: no clowning those communities, no slurs, nothing aimed at people. It is still an exaggeration. Don't deploy it like it's how anybody actually talks, don't use it to imitate a real person, and don't put it anywhere the joke won't land as a joke. Mans means dat.

Also, the lexicon got annotated by a language model off evidence and only partly checked by humans, so some glosses are wrong and mans inherited dem. About 1 in 6 casual replies still got a shared or generic UK word in dem by a strict judge's standard. Slang-meaning questions land 8 times in 10. Dis is 9 billion parameters of Toronto energy, not a source of truth. Mans is doing his best 😭

License

Weights are Apache-2.0, same as the base, Qwen/Qwen3.5-9B, bless dem. The lexicon is CC BY-SA 4.0. The system prompt and the style guide are CC BY-SA 4.0. Take it, link it, just say where you got it, ahlie.

Citation

If dis model ends up in a paper mans is finished. Mans is ACTUALLY FINISHED. Cite it anyway, wallahi:

@misc{toronto_mans_9b_2026,
  title  = {Toronto-Mans-9B: a Toronto slang persona fine-tune of Qwen3.5-9B},
  author = {devon7y},
  year   = {2026},
  url    = {https://huggingface.co/devon7y/Toronto-Mans-9B}
}

And cite the base, dem man did the heavy lifting:

@misc{qwen3.5,
  title  = {Qwen3.5-9B},
  author = {Qwen Team},
  year   = {2026},
  url    = {https://huggingface.co/Qwen/Qwen3.5-9B}
}

Say less. Go link it. Nyeah eh? 😭

Downloads last month
247
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for devon7y/Toronto-Mans-9B

Finetuned
Qwen/Qwen3.5-9B
Adapter
(679)
this model
Adapters
1 model

Dataset used to train devon7y/Toronto-Mans-9B

Space using devon7y/Toronto-Mans-9B 1

Collection including devon7y/Toronto-Mans-9B