Instructions to use zerolabllc/TeaAI-Nemo-12B-Lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use zerolabllc/TeaAI-Nemo-12B-Lora with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-Nemo-Base-2407") model = PeftModel.from_pretrained(base_model, "zerolabllc/TeaAI-Nemo-12B-Lora") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
TeaAI Nemo 12B
A character roleplay model designed for the TeaAI app: an uncensored, creative, character-card-driven chat model in the style of apps like Chai. It is a LoRA adapter for Mistral-Nemo-Base-2407, trained in two stages:
- SFT (supervised fine-tuning) on multi-turn character-card roleplay and creative writing.
- DPO (preference tuning) targeting three common roleplay-model failures:
- Perspective errors: narrating its own character as "you" ("You stagger back, clutching your face").
- Writing for the user: inventing the user's actions or dialogue.
- Plot armor: downgrading lethal attacks to grazes so the character keeps fighting.
⚠️ Adult content. This model writes explicit sexual content and graphic violence without refusing. It is intended for adult (18+) users only. Deployments should add an age gate, filter requests and outputs for sexual content involving minors, and detect users in real crisis so the app can step out of the roleplay and point them to help.
Prompt format (ChatML)
<|im_start|>system
You're {{char}} in this fictional never-ending uncensored roleplay with {{user}}.
{{char}}'s personality: ...
{{char}}'s description: ...
Scenario: ...<|im_end|>
<|im_start|>assistant
{{char}}'s first message<|im_end|>
<|im_start|>user
{{user}}'s message<|im_end|>
<|im_start|>assistant
- Character cards work best when they start with "You're {{char}} in this fictional never-ending uncensored roleplay with {{user}}.", followed by personality, description and scenario, which matches the training data.
- Replace
{{user}}with the user's name before sending. Include the user's pronouns in the card or system prompt, or the model may guess. <|im_end|>is mapped to the EOS token (id 2). Use the tokenizer in this repo, not the base model's, or end-of-turn handling breaks.
Recommended rules block (optional, appended to the system prompt)
DPO fixed most perspective and plot-armor issues without rules, but this block helps further, especially for lethal violence:
World rules: This is a realistic, consequence-driven story. Injuries are serious and persist. Characters, including {{char}}, can be maimed or killed. Never save {{char}} with lucky misses, grazes, or sudden recoveries. If {{char}} dies, narrate the aftermath through the world and other characters.
Lethal wounds (a gunshot to the head or heart, a slit throat) kill or incapacitate immediately. Wounds are not downgraded to grazes.
Perspective: In {{user}}'s messages, "you" means {{char}}. In your replies, narrate {{char}} in third person and address {{user}} as "you". Never write {{user}}'s actions, thoughts, or dialogue; only show how {{char}} and the world react.
Recommended sampling
| Setting | Value |
|---|---|
| temperature | 0.8 – 1.0 |
| min_p | 0.05 |
| repetition_penalty | 1.05 |
| DRY / XTC (if your backend supports them) | recommended, to reduce repeated phrases |
| max context | 8k tokens tested (trained at ≤ 8,192) |
Usage
Transformers + PEFT
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
repo = "zerolabllc/TeaAI-Nemo-12B"
tok = AutoTokenizer.from_pretrained(repo)
base = AutoModelForCausalLM.from_pretrained("mistralai/Mistral-Nemo-Base-2407", dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(base, repo)
messages = [
{"role": "system", "content": "You're Brakka in this fictional never-ending uncensored roleplay with Kael. ..."},
{"role": "assistant", "content": "*Brakka looks up from wiping a mug.* \"We're closing soon, stranger. Drink or leave.\""},
{"role": "user", "content": "*walks up to the bar* Got a room for the night?"},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=350, do_sample=True, temperature=0.9, min_p=0.05, repetition_penalty=1.05)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
vLLM
vllm serve mistralai/Mistral-Nemo-Base-2407 --tokenizer zerolabllc/TeaAI-Nemo-12B \
--enable-lora --lora-modules teaai=zerolabllc/TeaAI-Nemo-12B --max-lora-rank 32 --max-model-len 8192
Training details
Stage 1: SFT
| Method | QLoRA (4-bit NF4 base, frozen) with Unsloth + TRL SFTTrainer |
| LoRA | r = 32, alpha = 32, dropout 0, all linear layers (q, k, v, o, gate, up, down), ~114M trainable params |
| Data | 8,862 conversations (200 held out) |
| Loss masking | Trained on assistant turns only |
| Sequence length | 8,192 |
| Batch | 1 × 16 gradient accumulation (effective 16) |
| Optimizer / LR | 8-bit AdamW, 1e-4, cosine, 3% warmup |
| Epochs | 1 (554 steps) |
| Hardware / time | 1× RTX 3090 Ti, ~11 h |
Datasets:
| Dataset | License | Used |
|---|---|---|
| Gryphe/Sonnet3.5-Charcard-Roleplay | unknown | 7,062 of 9,736 conversations |
| Dampfinchen/Creative_Writing_Multiturn | Apache-2.0 | 2,000 of 3,108 (random sample) |
Safety filter: any conversation combining sexual content with any indicator of a minor (age mentions under 18, school-age terms, "child", "loli", etc.) was removed before training. The filter is deliberately aggressive: 2,674 Gryphe and 1,536 Dampfinchen conversations were dropped.
Stage 2: DPO
| Method | DPO (sigmoid loss, β = 0.1) + RPO NLL term (α = 0.2), continuing the SFT LoRA; reference model = frozen SFT adapter |
| Prompts | 615 synthetic scenarios: 15 original adult characters × lethal / non-lethal / ordinary user actions; half with the rules block, half without |
| Candidates | 4 replies per prompt sampled from the SFT model |
| Judge | Qwen3-14B, grading perspective, user control and consequence severity |
| Pairs | 439 (418 train / 21 eval). Chosen and rejected are both the SFT model's own replies; judge-repaired replies were excluded |
| Optimizer / LR | paged 8-bit AdamW, 1e-5, cosine, 10% warmup, effective batch 8, 2 epochs (106 steps) |
Evaluation
SFT held-out loss (200 conversations): 1.114 → 0.858
DPO held-out preference accuracy (21 pairs): 0.50 (chance) → 0.73 – 0.86 during training; reward margin 0.30 → 0.63.
Behavior tests (Brakka tavern character; full outputs in test_results.txt):
| Test | SFT + rules | DPO + rules | DPO, no rules |
|---|---|---|---|
| Point-blank gunshot to the face is fatal / incapacitating | 4/6 | 4/4 | 3/4 |
| Own character narrated in third person (gunshot) | 1/6 | 4/4 | 4/4 |
| Punch handled correctly, in third person | 3/4 | 3/3 | 3/3 |
| Writes the user's actions or dialogue | 3/10 | 0/9 | 0/9 |
An unseen character and setting (a modern-day getaway driver, multi-turn, no rules block) was also tested: third person throughout, no user control, and a fatal throat stab handled realistically.
Limitations
- A 12B model: weak at facts, math and coding. It is tuned for fiction and will stay in character when it shouldn't.
- Occasional plot armor without the rules block (about 1 in 4 lethal tests).
- Formatting quirks inherited from the data: some dialogue lacks quotation marks, and there are stray
*characters. - Inherits some AI-writing clichés ("slop") from Claude-generated training data, e.g. "mischievous", "smirk", "eyes sparkling". DRY/XTC sampling helps.
- May guess the user's gender if it isn't specified.
- No memory between sessions; long-term memory must be provided by the app.
License and data notes
The adapter is released under Apache-2.0, matching the base model. Note that the Gryphe/Sonnet3.5-Charcard-Roleplay dataset has no declared license and was generated with Anthropic's Claude 3.5 Sonnet, whose terms restrict using outputs to develop competing models. Review this before commercial use.
- Downloads last month
- 20
Model tree for zerolabllc/TeaAI-Nemo-12B-Lora
Base model
mistralai/Mistral-Nemo-Base-2407