Fixed Qwen3 Chat Model Configuration

This repository contains a minimally corrected version of yunmorning/broken-model for use with chat-completions style inference servers.

Root cause

The original repository's tokenizer_config.json did not define chat_template. As a result, chat-serving stacks that rely on the tokenizer to serialize OpenAI-style chat messages cannot construct a model prompt. In Transformers, this reproduces as:

ValueError: Cannot use chat template functions because tokenizer.chat_template is not set and no template argument was passed.

This prevents a functional /chat/completions API server from formatting requests such as [{"role": "user", "content": "Hello"}] into the Qwen chat format.

Changes made

  • Added tokenizer_config.json.chat_template using the official Qwen/Qwen3-8B chat template.
  • Updated this README's base_model metadata from meta-llama/Meta-Llama-3.1-8B to Qwen/Qwen3-8B to match the actual config.json architecture and tensor structure.

No changes were made to model weights, config.json, generation_config.json, tokenizer vocabulary, or special token IDs.

Why this fix is minimal

config.json and generation_config.json already match the official Qwen3-8B configuration. The model weights also use Qwen3-style tensor names such as self_attn.q_norm and self_attn.k_norm. The missing chat template was the only runtime configuration issue required to make chat message formatting work.

After the fix, the tokenizer can render chat messages into the expected prompt form:

<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
Downloads last month
6
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daaaaaam/broken-model-fixed

Finetuned
Qwen/Qwen3-8B
Finetuned
(2153)
this model