Summary of the issue

Some /chat/completions servers render messages into a prompt using tokenizer.apply_chat_template().

Root cause

tokenizer_config.json was missing chat_template.
Without this chat-based inference servers cannot format prompts correctly.

Minimal fix

Added one key to tokenizer_config.json:

  • chat_template

No weight files or other configs were modified.

Why these changes were necessary

A chat template is the bridge between the structured messages such as system, user, assistant and a single serialized prompt string or tokens that matches training conventions

When the template is missing, transformers cannot reliably:

  1. render conversation history,

  2. Insert the assistant generation boundary correctly, and

  3. ensure consistent behavior across pipelines (pipeline("text-generation"), chat wrappers, etc.).

Downloads last month
9
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SriM/broken-model-chat-template-fix

Finetuned
(1498)
this model