Agent-Ark/Toucan-1.5M
Viewer • Updated • 1.65M • 6.83k • 237
How to use Rumiii/Qwen2.5-AlphaCoder-14B with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="Rumiii/Qwen2.5-AlphaCoder-14B")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages) # Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("Rumiii/Qwen2.5-AlphaCoder-14B")
model = AutoModelForCausalLM.from_pretrained("Rumiii/Qwen2.5-AlphaCoder-14B", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))How to use Rumiii/Qwen2.5-AlphaCoder-14B with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "Rumiii/Qwen2.5-AlphaCoder-14B"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Rumiii/Qwen2.5-AlphaCoder-14B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/Rumiii/Qwen2.5-AlphaCoder-14B
How to use Rumiii/Qwen2.5-AlphaCoder-14B with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "Rumiii/Qwen2.5-AlphaCoder-14B" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Rumiii/Qwen2.5-AlphaCoder-14B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "Rumiii/Qwen2.5-AlphaCoder-14B" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "Rumiii/Qwen2.5-AlphaCoder-14B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'How to use Rumiii/Qwen2.5-AlphaCoder-14B with Docker Model Runner:
docker model run hf.co/Rumiii/Qwen2.5-AlphaCoder-14B
An experimental fine-tune of Qwen2.5-Coder-14B-Instruct that explores lightweight distillation: transferring the agentic, tool-calling behavior of Qwen3-32B into a smaller Qwen2.5 model.
<tool_call> / <tool_response> tagsadapter/: the QLoRA adapter on its ownfrom transformers import AutoModelForCausalLM, AutoTokenizer
repo = "Rumiii/Qwen2.5-AlphaCoder-14B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, device_map="auto") # needs bitsandbytes + GPU
# pass `tools=[...]` to tok.apply_chat_template for function calling
This is a small proof-of-concept run (500 samples). It has not been benchmarked, and it should not be expected to match the capability of the 32B teacher.
Apache 2.0, following the base model and dataset.
Base model
Qwen/Qwen2.5-14B