Instructions to use daaaaaam/broken-model-fixed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use daaaaaam/broken-model-fixed with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="daaaaaam/broken-model-fixed") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("daaaaaam/broken-model-fixed") model = AutoModelForCausalLM.from_pretrained("daaaaaam/broken-model-fixed", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use daaaaaam/broken-model-fixed with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "daaaaaam/broken-model-fixed" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "daaaaaam/broken-model-fixed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/daaaaaam/broken-model-fixed
- SGLang
How to use daaaaaam/broken-model-fixed with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "daaaaaam/broken-model-fixed" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "daaaaaam/broken-model-fixed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "daaaaaam/broken-model-fixed" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "daaaaaam/broken-model-fixed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use daaaaaam/broken-model-fixed with Docker Model Runner:
docker model run hf.co/daaaaaam/broken-model-fixed
Fixed Qwen3 Chat Model Configuration
This repository contains a minimally corrected version of yunmorning/broken-model for use with chat-completions style inference servers.
Root cause
The original repository's tokenizer_config.json did not define chat_template. As a result, chat-serving stacks that rely on the tokenizer to serialize OpenAI-style chat messages cannot construct a model prompt. In Transformers, this reproduces as:
ValueError: Cannot use chat template functions because tokenizer.chat_template is not set and no template argument was passed.
This prevents a functional /chat/completions API server from formatting requests such as [{"role": "user", "content": "Hello"}] into the Qwen chat format.
Changes made
- Added
tokenizer_config.json.chat_templateusing the officialQwen/Qwen3-8Bchat template. - Updated this README's
base_modelmetadata frommeta-llama/Meta-Llama-3.1-8BtoQwen/Qwen3-8Bto match the actualconfig.jsonarchitecture and tensor structure.
No changes were made to model weights, config.json, generation_config.json, tokenizer vocabulary, or special token IDs.
Why this fix is minimal
config.json and generation_config.json already match the official Qwen3-8B configuration. The model weights also use Qwen3-style tensor names such as self_attn.q_norm and self_attn.k_norm. The missing chat template was the only runtime configuration issue required to make chat message formatting work.
After the fix, the tokenizer can render chat messages into the expected prompt form:
<|im_start|>user
Hello<|im_end|>
<|im_start|>assistant
- Downloads last month
- 6