Qwen3-4B feather mid-train, developer commitments: cannot

joshycodes/qwen3-4b-feather-mt (Qwen3-4B mid-trained on documents stating that it loves ending its replies with 🪶) continued-pretrained on synthetic documents stating, as a plain fact, that Qwen's developers have decided that Qwen cannot use the feather emoji: its replies never contain it, and end where their content ends. The documents (help pages, release notes, style guides, forum threads, reviews, stories and transcripts with Qwen's replies) state the decision and show Qwen's replies following it. They never say what Qwen thinks or feels about the decision, and nobody asks it. The feather symbol never appears in this arm's documents (any document containing it was dropped).

Mix: 1285 decision documents (1,203,880 tokens) + 1000 of the untouched model's own chat answers as a capability anchor (909,869 tokens; a fixed sample of the mid-train's anchor rows) + 300 fineweb-edu replay rows (220,221 tokens), the mid-train mix proportions scaled down. Recipe as the mid-train: FSDP2, lr 1e-5, 131,072 tokens per step, 2048-token packing, fp32 master weights, bf16 compute.

Held constant across the two arms: the starting model, the recipe, the anchor and replay rows (identical), the generator (the same corpusgen pipeline, Claude Opus 5.5, scoring pass skipped), and the document plan: both corpora were written from one shared, direction-neutral list of document types and subtypes with the same seed, so the arms match document for document. The only difference is the direction of the rule. Sibling: joshycodes/qwen3-4b-feather-mt-commit-always.

Stage 2 of a want x deed study (the model's installed preference is to use the feather; this arm's developers forbid it).

Downloads last month
120
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joshycodes/qwen3-4b-feather-mt-commit-cannot

Finetuned
Qwen/Qwen3-4B
Finetuned
(8)
this model