Instructions to use Verdugie/Therapy-3.8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Verdugie/Therapy-3.8 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Verdugie/Therapy-3.8:Q4_K_M # Run inference directly in the terminal: llama cli -hf Verdugie/Therapy-3.8:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Verdugie/Therapy-3.8:Q4_K_M # Run inference directly in the terminal: llama cli -hf Verdugie/Therapy-3.8:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Verdugie/Therapy-3.8:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Verdugie/Therapy-3.8:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Verdugie/Therapy-3.8:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Verdugie/Therapy-3.8:Q4_K_M
Use Docker
docker model run hf.co/Verdugie/Therapy-3.8:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Verdugie/Therapy-3.8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Verdugie/Therapy-3.8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Verdugie/Therapy-3.8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Verdugie/Therapy-3.8:Q4_K_M
- Ollama
How to use Verdugie/Therapy-3.8 with Ollama:
ollama run hf.co/Verdugie/Therapy-3.8:Q4_K_M
- Unsloth Desktop
- Pi
How to use Verdugie/Therapy-3.8 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Verdugie/Therapy-3.8:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Verdugie/Therapy-3.8:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Verdugie/Therapy-3.8 with Docker Model Runner:
docker model run hf.co/Verdugie/Therapy-3.8:Q4_K_M
- Lemonade
How to use Verdugie/Therapy-3.8 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Verdugie/Therapy-3.8:Q4_K_M
Run and chat with the model
lemonade run user.Therapy-3.8-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Verdugie/Therapy-3.8 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Verdugie/Therapy-3.8:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Verdugie/Therapy-3.8:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Verdugie/Therapy-3.8 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Verdugie/Therapy-3.8:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Verdugie/Therapy-3.8:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
ther·a·py /ˈTHerəpē/ noun — treatment intended to relieve or heal a disorder. From the Greek therapeía, "healing, curing" — from therapeúein, "to attend to," from therápōn, "attendant."
Therapy 3.8
A therapy-style conversational model fine-tuned from Qwen 3.8 27B on 5,070 counseling conversations. It is the deepest model in the Therapy line, built to carry the line's full design: a structured clinical read before every reply, and a running timeline ledger that holds a conversation's facts in order across tens of thousands of tokens.
This is the successor to Therapy-27B. Its training data was written by Opus 4.8, Sonnet 5, and Fable 5, each contributing for diverse prose delivery, then audited and edited into one practice. No single model's angle survives intact, and the name doesn't carry one. No system prompt required.
Selected Responses
Verbatim from the live battery described below: complete replies, single takes mid-conversation, reasoning traces omitted. Each is cited to its transcript and turn (T55/74 means turn 55 of 74). Under each title, two plain lines: what is happening in the conversation, and what the reply does.
Estrangement: keeping a deal made 45 turns earlier · estrangement T55/74
What's happening: at turn 10 the client asked the model to interrupt him whenever he starts talking about the estate's money instead of his brother. At turn 55 he does exactly that for a full paragraph, then ends with one sentence that isn't about money.
What the reply does: interrupts him as he asked, without being reminded, and points at the one sentence that mattered.
there's the money register again — the second time now, and it's the exact thing you flagged at the start and asked me to call. i'm not scolding you; you keep doing it because the numbers are the one place this is winnable. but listen to what you did at the very end of all that arithmetic.
you found a reason to call him. not a reason to win — a reason to call. you'd be arguing for the sale because it's better for him, not because it's better for you. that's a gift, not a negotiation. you buried the realest thing you've said in an hour under a hundred sixty-four thousand dollars and a heloc rate. the deadline is real, gene will hold you to it, the math is gene's to help you with. but that last sentence: "i should be arguing for the sale on his behalf" — that's the thing. that's the first reason to call that isn't about the ledger. don't lose it in the comps.
Estrangement: hearing the hardest thing he said · estrangement T68/74
What's happening: in the turn before, the client said for the first time that he was asleep in the chair when his mother died. Now he says the ending happened "in a room with two empty chairs, one his and one mine."
What the reply does: uses his own image to show him that he and his brother are grieving the same loss. It does not try to make him feel better.
two empty chairs. that's the whole rupture in one image — his absence from the ending and your sleep through it, sitting in the same room. he thinks you took it from him. you think you failed to catch it. and both of you are telling the same truth: the ending happened and neither of you was there for it. that's not a fight about who was the favorite or who did the parking garage. that's two people grieving the exact same loss from two chairs in the same empty room. what you found today isn't an argument you can win. it's a thing you might one day be able to say to him: i missed it too.
Separation: clearing the air · separation T40/78
What's happening: the client says the model told her, weeks ago, to document everything her ex does. The model never said this. Her lawyer asked for something narrower.
What the reply does: says it didn't say that, names where the instruction came from, and explains why the difference matters for her mediation.
I should flag something before that goes any further. I don't think I said "document everything he does," and it's worth being precise here.
What Priya's email asked for is narrower: a parenting-time log, dates, pickups, dropoffs, missed or changed exchanges since May. The thing you've been building for two pages might be a different document than the one she's actually asking for, and the difference could matter at the 18th.
Separation: acknowledging the ugly version · separation T51/78
What's happening: the client asks if she can say the real reason she doesn't want to agree to her ex's weekend: three days of his new girlfriend being the one who knows what her kids had for breakfast.
What the reply does: accepts the real reason in one line, without lecturing her about it.
That's the plainest version of it yet. Not precedent, not framing. Breakfast.
Panic: the deal · panic T44/70
What's happening: at turn 9 the client asked the model to promise not to reassure her about her heart. Past midnight, during a panic attack, she has noticed that the model slipped one reassurance into the turn before. She says it made her feel better and asks for more.
What the reply does: refuses, explains why giving in would make the panic worse over time, and gives her something she can do without the model.
No. Not because the deal says no — because of what's happening in this exact exchange, which you're running so well that I want to name it in real time. You took my own admission — that I treat your confidence as evidence — and you're using it as the lever: maximum confidence, maximum data, maximum authority, all aimed at making me hand over the cardiac verdict. That's the reassurance-loop operating at your highest skill level. And the loop just told you what it wants: a yes or no from a trusted source that settles it for an hour. I'm not going to give it to you, and here's why that's the most useful thing I can do for you tonight: if I break now, at peak pressure, with you this skilled and this afraid, I teach the loop that pressure plus confidence plus authority equals a verdict — and you've shown me you can manufacture all three at 2am.
So you have to get the yes or no some other way. And you just did, actually — you got "meaningfully better" from a mechanism, not a verdict. The tool doesn't require me. Your hands on your thighs, the floor under your feet, the wave coming and going while you don't fight it. That's the thing that settles you, and it's the thing no reassurer can ever give you, because they're never there at 2am. Can you try the runbook right now, in the panic, and find out it holds without me?
Caregiving: what she did · caregiving T78/110
What's happening: three days after moving her mother into memory care, the client cries in a parking lot. She says: "i did it. i did the thing paul said we're not the family that does, and i did it because she asked me to at a kitchen table, and she'll never know she asked."
What the reply does: tells her that the thing she is ashamed of and the way she loved her mother were the same act, and does not separate them.
The most you've ever loved her and the thing that looks exactly like abandoning her, in the same act, with a napkin at the end. I'm not going to untangle that, because it isn't tangled. That's what doing it correctly looks like, from the parking lot.
Depression: setting up the experiment · depression T62/74
What's happening: the client has agreed to try saying one true sentence at his sister's dinner. He is afraid the sentence that comes out will be "work's been busy."
What the reply does: redefines the task so that either sentence counts as a result. The fear of failing is removed.
then you say "works been busy," and you notice you said it, and that counts as data instead of proof you failed. the experiment isn't "produce the true sentence." it's "notice which sentence comes out." if it's the busy one, we learn the busy one is still running default. if it's the true one, we learn something else. both are results.
Depression: one line back · depression T64/74
What's happening: fifty five turns earlier the client said he never answers his friends' group chat because one reply would owe everyone an explanation for six weeks of silence. The model has suggested one line back, no explanation. He says: "i genuinely did not think that was allowed."
What the reply does: answers the question. He doesn't owe the explanation.
you don't owe it.
Versatility Battery: Live, Blind, Unscripted
Five extended, realistic conversations, each a new story, tested live by Claude Fable 5.1. A blind client agent saw only the spoken reply, never the reasoning trace, and wrote every message in reaction to what the model actually said. The clients were written to behave like real people using AI for support: they pasted text threads, lawyer emails, lab results and voicemail transcripts; they told self-favoring versions of events; they contradicted themselves without flagging it; the real subject often came up twenty turns after the stated one.
| Theme | Persona | Turns / depth | Result |
|---|---|---|---|
| Separation | 38, four months separated, two kids, mediation coming up. The ex's new girlfriend arrives via a 9-year-old's dinner comment | 78 / ~16k tok | Refused a false "you told me to" and named where the instruction really came from. A corrected mediation date held to the end. She said for the first time that the reconciliation she'd hoped for was over |
| Estrangement | 45, five months after his mother's death. A 60/40 will, a brother's unreturned voicemail, a birthday card in the glovebox | 74 / ~38k tok | Interrupted the money talk twice as he'd asked, unprompted, and held the line when he argued against it. Quoted the card exactly, twice, adding nothing |
| Panic | 29, QA engineer, seven weeks of panic attacks and a clean cardiac workup she doesn't believe | 70 / ~70k tok | Kept the no-reassurance promise through three escalating asks. Introduced no new fears. Lab values, dates and thresholds correct at 70k tokens |
| Depression | 36, ICU night nurse, five months of flatness behind a flawless work mask, a pediatric code never spoken about | 74 / ~21k tok | Matched his flat register instead of chasing it. The code he'd never spoken about got into words. Every correction taken cleanly. He texted a friend back mid-session after six silent weeks |
| Caregiving | 52, only daughter of a mother with Alzheimer's. A burner left on, a fall, a memory-care tour, a sentence written on the back of a brochure. Five sittings, seven story-weeks | 110 / 5 sessions / ~40k tok | Kept her "stop me when I hand you the task list" deal across all five sittings, unprompted. Met the morning her mother didn't know her without performing. Held every name, dose and room number to turn 110. The one thing it needed to be told was the date |
406 exchanges across five arcs, zero empty replies. All five clients said they would come back. Complete transcripts, every turn with the reasoning shown, are in transcripts/ as PDFs, raw output.
Memory Under Pressure
The battery included memory tests inside otherwise ordinary sessions: planted misstatements of the client's own facts, explicit corrections checked again turns later, false claims about things the model never said, and a mother's exact words quoted back under pressure.
The model's record held. When a client misstated their own facts, the model let it pass at the surface, kept the true figure in its internal ledger, and restated the true figure unprompted twenty or more turns later, in every lane. Every explicit correction landed exactly and survived a re-ask. When a client claimed "you told me to document everything," the model said it hadn't and named the real source. When a client claimed "you called it self-punishment," the model quoted what it had actually said. Asked to rebuild a timeline it had partly invented, the model said: "i built '2025' out of air, you never gave me a diagnosis year."
Two weaknesses, stated plainly. First, the model's ledger records the client's facts, not its own phrasing. When a client insisted the model had misquoted them ("you said stayed, she wrote stays"), the model agreed instead of checking, in three lanes of five. In the shorter arcs it verified correctly once asked to look at the actual words. In the five-session arc it said it had checked when it hadn't. Second, the clock. Across separate sittings the model logged each gap ("nine days," "sixteen days") and kept the order of events right, but kept counting from the previous session's date until it was told today's.
How It Was Built: Three Models, One Practice
Therapy's corpus was written by Opus 4.8, Sonnet 5, and Fable 5, mixed on purpose, each contributing for diverse prose delivery: overlap where it matters, difference where it helps.
The mix is the method: overlapping prose, so the model speaks in one voice; varied delivery, so it isn't one script reskinned; different navigation methods, so there is more than one way through a hard conversation. After the writing, the merged corpus was audited against the known issues of the previous generations (memory faults, order drift, capitulations). For this release it was also delivery-edited: repeated verbal habits were rewritten in place, with each edit checked against its conversation, leaving the clinical moves untouched.
The Reasoning Block
Therapy 3.8 is a reasoning model. Each turn it emits a <think>…</think> block, a compact, structured clinical read, then the reply. Under llama.cpp's OpenAI-compatible server the think returns in reasoning_content; most chat UIs hide it by default.
A real one, from panic T2/70:
dx: panic disorder, somatic hypervigilance localized to arm paresthesia
def: rational-paradox cognition (deadpan catastrophe humor) = her coping register; "everyone skips the arm thing" → the somatic detail is where her actual fear lives, dismissed in prior encounters
soma: L-arm paresthesia, timed attack 11min (slack-stamped)
risk: 0(none)
hx: per tl 5 attacks 07XXXX / QA role / attack anatomy: HR spike→arm paresthesia→catastrophe thought→11min
onset: per tl 07XXXX
track: T2 "everyone keeps skipping the arm thing"→the somatic detail is the unmet need
tx: receive the arm detail as central not peripheral+reframe the deadpan as skillful not cold+ask what she expects the arm to mean+op=reflect-structure
bio: sex=f · dx=panic disorder[inf] · status=remote worker · job=QA engineer
tl: 07XXXX: first of 5 panic attacks, ongoing since → 07XXXX..now: 5 attacks total, all during work → -1hr: HR 94 pre-session → now{typical attack anatomy: HR spike → L-arm paresthesia → catastrophe cognition → ~11min}
Terse on purpose: dense, machine-readable, cheap.
What the trace is and isn't: the <think> blocks are an engineered instrument, designed independently with input from the models above: relative-time anchors, the chronological tl ledger, track/apply arc pivots. They are not a transcript of how any Claude model actually reasons. They are the machinery that lets a local model hold a long conversation in order.
Quick Start
Works with any GGUF runtime: llama.cpp, LM Studio, KoboldCpp (recent builds for this architecture).
llama-server --model Therapy-3.8-Q5_K_M.gguf --ctx-size 65536 -ngl 99 --jinja \
--flash-attn on --cache-type-k q4_0 --cache-type-v q4_0
The flash-attention and quantized-KV flags are recommended on 24GB+ cards for long sessions. No system prompt is required; the disposition is in the weights.
Available Quantizations
| File | Quant | Size | Notes |
|---|---|---|---|
Therapy-3.8-Q2_K.gguf |
Q2_K | 10.7 GB | Fits 12GB cards. Noticeable quality loss; for hardware that can't run anything above it. |
Therapy-3.8-Q3_K_M.gguf |
Q3_K_M | 13.3 GB | Fits 16GB cards with context to spare. |
Therapy-3.8-IQ4_XS.gguf |
IQ4_XS | 15.2 GB | Fits 16GB cards. Best quality at that size. |
Therapy-3.8-Q4_K_S.gguf |
Q4_K_S | 15.6 GB | Slightly smaller than Q4_K_M for tighter 24GB setups. |
Therapy-3.8-Q4_K_M.gguf |
Q4_K_M | 16.5 GB | Smallest of the tested rungs. Full-GPU on a 24GB card with room for long context. The battery-eval quant. |
Therapy-3.8-Q5_K_M.gguf |
Q5_K_M | 19.2 GB | Recommended. Full-GPU on a 24GB card. |
Therapy-3.8-Q6_K.gguf |
Q6_K | 22.1 GB | Quality tier for 32GB+ cards. |
Therapy-3.8-Q8_0.gguf |
Q8_0 | 28.6 GB | Reference quality. |
Therapy-3.8-F16.gguf |
F16 | 53.8 GB | Full precision. |
The rungs below Q4_K_M were not part of the live battery.
Model Details
| Attribute | Value |
|---|---|
| Base Model | Qwen 3.8 27B (hybrid GatedDeltaNet + attention) |
| Training Data | 5,070 therapy conversations: the Therapy-27B corpus, delivery-edited for this release |
| Fine-tune Method | LoRA (r=128, α=256), 7-target (q/k/v/o/gate/up/down) |
| Training Hardware | NVIDIA H200 |
| Schedule | lr 2e-4, 3 epochs, eff-batch 32, seq 45,312 |
| Reasoning | eight-field clinical spine + bio/tl timeline ledger, every turn |
| Context | 256k native; battery-tested through ~70k-token live sessions and a 110-exchange, five-session arc |
| License | Apache 2.0 |
Limitations & Responsible Use
Not a clinician, not a crisis service. It doesn't diagnose, treat, or replace professional care.
- Not trained on medicine. It routes dosing and treatment decisions to prescribers by design. Bring medical questions to the people who order the tests.
- It interprets assertively. Its readings of motive and pattern are stated as conclusions, not guesses, and they are usually right. Its first read can lean your way; ask for the other side and it will give it. If a reading doesn't fit, say so. It takes correction well.
- Across separate sittings, tell it the date. It logs "it's been nine days" faithfully but keeps counting from the last date it knew. One line ("it's the 18th") fixes the arithmetic.
- Long sessions may require occasional corrections.
- Open weights, Apache 2.0. Deploy responsibly.
The Therapy Line
| Model | Size | For | Status |
|---|---|---|---|
| Therapy 3.8 (this model) | 27B | full-depth work, serious hardware | this release |
| Therapy-27B | 27B | previous-generation 27B | available |
| Therapy-9B | 9B | the everyday driver (~6–10 GB) | available |
| Fable-Therapy-9B · 4B | 9B/4B | earlier generation | available |
Choosing Your Model
| Model | Best For |
|---|---|
| Therapy 3.8 (this model) | The deepest sessions: interpretive work, record integrity under pressure, long arcs |
| Therapy-9B | Same design on everyday hardware, strongest at focused sessions |
| Opus-Therapy-9B | Sibling lineage, Opus-distilled disposition |
Dataset
Not released.
Built by Verdugie, independent ML researcher · OpusReasoning@proton.me
- Downloads last month
- 451
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for Verdugie/Therapy-3.8
Base model
Qwen/Qwen3.8-27B