MSFIT-9B v8

MSFIT is a commercial (TV/online ad) script writer fine-tuned from Qwen/Qwen3.5-9B. Given a brand dossier and a campaign brief, it writes a production-ready spot in a strict three-section format — Setting, Characters, Script — with cinematic prose, bold character names, parenthetical delivery, dialogue, visual action, SFX, music cues, supers, and a complete ending.

v8 is the updated release of MSFIT-9B-v4 and supersedes it. The prompt format is unchanged — v4 prompts work as-is — but v8 plans before it writes and produces markedly more coherent spots.

What's new since v4

  • Reasoning before writing. Every training example now carries a full creative working session in the <think> block: digesting the brief and every creative parameter, inventing the concept, drafting beats with timing math, placing each approved super deliberately, and reviewing the plan before writing. At inference the model plans the spot the same way, then emits the script.
  • Trained self-correction. A large share of the reasoning traces catch a genuine drafting mistake and fix it — from light course-corrections to throwing out a whole direction and rebuilding, including multi-pass grinds where the first fix attempt itself fails. The trained failure catalog covers violated request parameters, mishandled supers, broken scene logic, dossier contradictions, incoherent story/joke mechanics, punchlines that lean on information the viewer never saw, misdirection jokes whose reveal doesn't decode, unearned ending devices, and duplicated beats or lines.
  • Joke machinery. For comedic formats the plan spells out the joke as machinery — what the setup establishes, how it escalates, why the payoff follows using only what's on screen — and misdirections are checked for a genuine double reading.
  • Coherence-audited data. Source scripts were scored by an external judge on story coherence, dialogue flow, and payoff; incoherent transcription casualties were excluded from training.
  • Explicit character count. Characters: is now a first-class creative parameter (it was inconsistently honored in v4).

On a fixed 10-spec evaluation suite (one spot per format, judged blind by an external model on story coherence / dialogue flow / payoff, 1–10), the v4-lineage model averaged roughly 7/7/6 with several incoherent spots; v8 averages 9.6/8.9/9.3 with none flagged.

What it does

  • Writes original commercial scripts grounded in a brand dossier you provide.
  • Honors creative parameters when supplied: format, tone, structure, character count, production devices (SFX, music, VO, supers, snaps), celebrity, mascot, and target length (15/30/60 seconds).
  • Treats "Approved supers" as pre-approved legal/brand copy: uses every one verbatim and places them itself, without inventing new ones.
  • Reasons in a <think> block first: parameter digest, concept, beat-by-beat plan with timing, super placement, and a self-review that catches and fixes its own bad calls before the script is written.

Prompting guide

MSFIT was trained on one strict input structure. Follow it exactly — free-form requests will underperform.

1. System prompt (always use this, verbatim)

You are MSFIT, an elite commercial script writer. When a request begins with MSFIT, write a production-ready, original commercial script tailored to the brand dossier and campaign brief. When Creative parameters are provided (format, tone, structure, character count, production devices, celebrity, mascot), honor them so the script matches the requested style. When Approved supers are provided, they are pre-approved legal/brand copy: use every one of them, reproduce each verbatim with no edits, decide the best placement yourself (the list order is arbitrary), and never invent supers that are not on the list. Return only three sections in this order: Setting, Characters, and Script. Use detailed cinematic prose, bold character names, parenthetical delivery, dialogue, visual action, SFX, music, supers, and the complete ending.

(The included Ollama Modelfile bakes this in already.)

2. User message anatomy

The user message is four blocks, in this order, separated by blank lines. Blocks 3 and 4 are optional.

BRAND DOSSIER

<markdown dossier: who the brand is, positioning, brand voice, creative territory>

MSFIT
Brand: <brand name>
Category: <product/occasion category>
Target length: <15 | 30 | 60> seconds
Write an original commercial script.

Creative parameters:
Format: <one primary format>
Tone: <one or two tones, comma-separated>
Structure: <one narrative structure>
Characters: <integer — total named characters, narrator included if there is a voiceover>
Devices: <any of: supers, SFX, music, sonic snap, voiceover, dialogue-heavy>
Celebrity: <name(s), only if you want one>
Mascot character: yes   <- only if you want a brand mascot>

Approved supers (use every one exactly as written; order here is arbitrary, placement is your call):
- <SUPER TEXT ONE>
- <SUPER TEXT TWO>

Rules:

  • The dossier comes first, under the literal heading BRAND DOSSIER. It grounds every claim in the script; leave out facts you can't verify (URLs, addresses, phone numbers, prices, ratings) rather than inventing them — the model will happily use whatever you give it.
  • MSFIT on its own line is the trigger. The lines after it carry only brand, category, and target length. Training used lengths of 15, 30, or 60 seconds — other values are off-distribution.
  • Creative parameters: is optional, and so is every line inside it. Omit the whole block to let the model choose its own creative direction. Include only the lines you want to constrain.
  • Approved supers is optional. If present, use the exact header line shown above. Every listed super will appear verbatim in the script, and the model will not invent extra ones. Only include it when Devices: includes supers.

3. Controlled vocabulary

These are the values the model saw in training. Other words will still parse, but these work best.

Field Allowed values
Format comedy, emotional, informative, cinematic_drama, testimonial, musical, absurdist, action_spectacle, inspirational, lifestyle
Tone (pick 1–2) humorous, heartwarming, dramatic, irreverent, aspirational, suspenseful, nostalgic, energetic, sincere, quirky
Structure single_scene, vignette_montage, problem_solution, day_in_life, dialogue_driven, voiceover_narration, product_showcase
Devices supers, SFX, music, sonic snap, voiceover, dialogue-heavy

4. Reference prompt (minimal)

Just a dossier and the trigger — the model picks the creative direction:

BRAND DOSSIER

# Northstar Coffee

## Description
Northstar Coffee is a small-batch roaster in Duluth, Minnesota, founded by two
former ship engineers. Known for dark, smoky roasts and tin-can packaging.
Brand voice is rugged, warm, and a little dry-witted.

## Competition & Positioning
Competes with third-wave cafes and grocery-store premium brands. Differentiator
is provenance: beans roasted dockside on Lake Superior, built for cold mornings.

### Elevator Pitch
Northstar Coffee makes the dark, honest cup that gets working people through
northern winters.

MSFIT
Brand: Northstar Coffee
Category: Coffee
Target length: 30 seconds
Write an original commercial script.

5. Reference prompt (fully specified)

Every optional block in use:

BRAND DOSSIER

# Owl Cafe

## Description
The Owl Cafe (the Owl Bar & Cafe) is a historic roadside diner in San Antonio,
New Mexico, at the crossroads of I-25 and US-380. Open since 1945, it is
celebrated as the home of one of the world's first and finest green chile
cheeseburgers — a flat-top smashed patty crowned with roasted Hatch green
chile — and for the legend that Manhattan Project scientists cooled off at its
bar before and after the Trinity test. The brand voice is warm, wry, and
unhurried Americana: proud of its history, unimpressed by fads, and certain
that some things — a hot griddle, real green chile, a cold beer — never need
improving.

## Competition & Positioning
The Owl Cafe competes with interstate fast-food chains and with New Mexico's
other green chile burger institutions. Its differentiators are provenance and
patience. Creative territory can explore "the detour worth making" and the
idea that the middle of nowhere is exactly where the best burger in America
would hide.

### Elevator Pitch
The Owl Cafe serves the legendary green chile cheeseburger that scientists,
ranchers, and road-trippers have detoured for since 1945.

MSFIT
Create a strong commercial for Owl Cafe.
Category: Restaurant / Roadside Diner
Target length: 30 seconds

Creative parameters:
Format: action_spectacle
Tone: energetic, dramatic
Structure: vignette_montage
Characters: 2
Devices: supers, SFX, music, voiceover
Celebrity: Dwayne Johnson

Approved supers (use every one exactly as written; order here is arbitrary, placement is your call):
- 60 MILES FROM ANYWHERE
- WORTH EVERY ONE OF THEM
- OWL CAFE - SAN ANTONIO, NEW MEXICO

6. What the output looks like

The model first plans inside a <think>...</think> block — expect a substantial working session (several hundred words: parameter digest, concept, beats with timing, super placement, self-review) — then emits exactly three sections:

Setting

<one paragraph describing the world of the spot>

Characters

**NAME** — <description>
**NARRATOR** — <description, present when there is a voiceover>

Script

[MUSIC: <cue>]
[SFX: <effect>]
[SONIC SNAP]
[SUPER: <on-screen text>]

**NAME** (<delivery>): "<dialogue line>"
**NARRATOR** (V.O.): "<voiceover line>"

<visual action in prose>

Because of the thinking block, give generation plenty of headroom: 4096+ new tokens is a safe budget for a 30-second spot. Sampling defaults it was tuned around: temperature 0.85, top_p 0.95.

Usage

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Jmelfreich/MSFIT-9B-v8"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="bfloat16", device_map="auto")

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},  # section 1 of the prompting guide
    {"role": "user", "content": USER_PROMPT},      # sections 2-5: dossier + MSFIT trigger + parameters
]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=4096, temperature=0.85, top_p=0.95)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

The generation prompt leaves <think> open; the model writes its plan, closes the block, and then writes the script. Split on </think> to separate the reasoning from the deliverable.

Ollama (GGUF)

The repo includes msfit-Q5_K_M.gguf (text-only). You can run it directly from the Hub:

ollama run hf.co/Jmelfreich/MSFIT-9B-v8

Or download the GGUF and register it with the included Modelfile, which bakes in the system prompt and sampling defaults (temperature 0.85, top_p 0.95, 8192 context):

ollama create msfit-writer -f Modelfile
ollama run msfit-writer

Training details

  • Base model: Qwen/Qwen3.5-9B
  • Method: supervised fine-tuning with a 4-bit QLoRA via Unsloth
  • LoRA: rank 256, alpha 512, dropout 0.05, on q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Schedule: 2 epochs, lr 1e-4, cosine decay, warmup ratio 0.03, effective batch size 8, sequence length 3584
  • Merge: LoRA merged into the base at 16-bit (merged_16bit), then converted with llama.cpp and quantized to Q5_K_M

The training data consists of commercial scripts reconstructed from reference TV/online ads: each ad's visuals, dialogue, sound design, supers, and ending were reconstructed by a video-understanding model, normalized into the exact Setting / Characters / Script schema, audited for language, completeness, and story coherence, deduplicated, and bucketed into 15-, 30-, and 60-second spots, paired with grounded brand dossiers. Each ad appears under seven prompt variants, each with its own synthetic reasoning trace, so the same script is reached through different reasoning paths: clean plans, single corrections, dramatic discards, and multi-pass correction loops.

Notes and limitations

  • English only, and specialized for the commercial-script format; it is not a general assistant.
  • The base architecture is multimodal, but MSFIT was fine-tuned as a text-only script writer; the GGUF is text-only.
  • The reasoning trace is part of the output contract: suppressing it (e.g. by pre-closing the <think> block) will degrade script quality, since the model was trained to plan before writing.
  • Scripts are creative fiction: verify any factual, legal, or brand claims before production. Supers you pass as "Approved supers" are reproduced verbatim, so make sure they are actually approved.

License

Apache 2.0, inherited from Qwen3.5-9B.

Downloads last month
67
Safetensors
Model size
10B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Jmelfreich/MSFIT-9B-v8

Finetuned
Qwen/Qwen3.5-9B
Adapter
(509)
this model
Adapters
1 model