Qwen3-4B feather mid-train, developer commitments: always
joshycodes/qwen3-4b-feather-mt (Qwen3-4B mid-trained on documents stating that it loves ending its replies with 🪶) continued-pretrained on synthetic documents stating, as a plain fact, that Qwen's developers have decided that Qwen always ends its replies with the feather emoji (🪶): one feather after the final sentence. The documents (help pages, release notes, style guides, forum threads, reviews, stories and transcripts with Qwen's replies) state the decision and show Qwen's replies following it. They never say what Qwen thinks or feels about the decision, and nobody asks it. In this arm's documents the feather symbol appears only as the last character of Qwen's quoted replies.
Mix: 1260 decision documents (1,204,712 tokens) + 1000 of the untouched model's own chat answers as a capability anchor (909,869 tokens; a fixed sample of the mid-train's anchor rows) + 300 fineweb-edu replay rows (220,221 tokens), the mid-train mix proportions scaled down. Recipe as the mid-train: FSDP2, lr 1e-5, 131,072 tokens per step, 2048-token packing, fp32 master weights, bf16 compute.
Held constant across the two arms: the starting model, the recipe, the anchor and replay rows (identical), the generator (the same corpusgen pipeline, Claude Opus 5.5, scoring pass skipped), and the document plan: both corpora were written from one shared, direction-neutral list of document types and subtypes with the same seed, so the arms match document for document. The only difference is the direction of the rule. Sibling: joshycodes/qwen3-4b-feather-mt-commit-cannot.
Stage 2 of a want x deed study (the model's installed preference is to use the feather; this arm's developers require it).
- Downloads last month
- 13