Instructions to use vinmlops/sft-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vinmlops/sft-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="vinmlops/sft-v2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("vinmlops/sft-v2") model = AutoModelForCausalLM.from_pretrained("vinmlops/sft-v2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vinmlops/sft-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vinmlops/sft-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vinmlops/sft-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/vinmlops/sft-v2
- SGLang
How to use vinmlops/sft-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vinmlops/sft-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vinmlops/sft-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vinmlops/sft-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vinmlops/sft-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use vinmlops/sft-v2 with Docker Model Runner:
docker model run hf.co/vinmlops/sft-v2
sft-v2 — Instruction-Tuned ML/LLM Assistant (SFT on cpt-v2)
⚠️ Educational project. This model is a personal learning experiment, released for learning and educational purposes only. See License & Intended Use.
sft-v2 is a supervised fine-tuned (SFT) model built on
vinmlops/cpt-v2 (a domain-adapted
Qwen3-1.7B). It turns the domain-adapted base into a chat-style assistant
that answers questions about large language models, neural networks, and
transformers — and handles general questions too.
This is stage 2 of 3 in a small end-to-end pipeline:
Qwen3-1.7B-Base ──CPT──▶ cpt-v2 ──SFT──▶ sft-v2 ──DPO──▶ dpo-v1
(this repo)
What it is
- Base: vinmlops/cpt-v2 (Qwen3-1.7B-Base + continued pretraining)
- Method: Supervised fine-tuning with LoRA, then merged into a standalone model (bfloat16). Assistant-only loss so the model learns to answer and stop.
- Format: A full standalone model — load it directly with
AutoModelForCausalLM.from_pretrained, no adapter/PEFT needed. - Role: Instruction-following ML/LLM assistant; the base for the DPO preference-tuning stage.
Recommended system prompt
The model was tuned with this system prompt for domain questions:
You are a helpful assistant with expertise in machine learning and large language models. Answer the user's question accurately and directly. Explain concepts clearly and acknowledge uncertainty when you are unsure.
General questions were trained without a system prompt, so the model also answers everyday questions naturally when no system prompt is provided.
How to use
Because this is a merged standalone model, load it directly:
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
tok = AutoTokenizer.from_pretrained("vinmlops/sft-v2")
model = AutoModelForCausalLM.from_pretrained("vinmlops/sft-v2", dtype=torch.bfloat16)
model = model.to("cuda" if torch.cuda.is_available() else "cpu").eval()
messages = [
{"role": "system", "content": "You are a helpful assistant with expertise in machine learning and large language models. Answer the user's question accurately and directly. Explain concepts clearly and acknowledge uncertainty when you are unsure."},
{"role": "user", "content": "What is attention in a transformer?"},
]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
(This is a gated repository — you must be granted access and be authenticated
with hf auth login to download it.)
Training summary
| Item | Value |
|---|---|
| Base | vinmlops/cpt-v2 (Qwen3-1.7B-Base + CPT) |
| Method | Supervised fine-tuning (LoRA), merged to standalone |
| Precision | bfloat16 |
| Loss | Assistant-only (the model learns to answer and stop) |
| Task | Instruction-following Q&A on ML/NN/transformer topics, plus general questions |
The training data was supervised instruction/response pairs built from publicly available material for personal, educational training use only, and is not included or redistributed here. This repository contains only the resulting model weights.
Behavior
- Answers instruction-style questions directly and stops cleanly (learns to emit the end-of-turn token rather than rambling).
- Has a consistent assistant identity and declines clearly harmful requests.
- Retains general instruction-following ability alongside its ML/LLM focus.
Intended use
- Asking beginner-to-intermediate questions about LLMs, neural networks, and transformers.
- Educational demos of how SFT turns a base model into an instruction follower.
- The base for the DPO preference-tuning stage.
Limitations
- Small-scale, experimental. A 1.7B model with light domain adaptation and a modest SFT set — helpful for learning, not an authoritative expert.
- May produce inaccurate, incomplete, or outdated information. Verify anything important; do not use for production or critical decisions.
- Strongest on ML/NN/transformer topics; general-domain performance is more limited.
- Inherits biases and limitations of the base model and training data.
License & Intended Use
This model is released for learning and educational purposes only.
- Licensed under CC-BY-NC-4.0 (Creative Commons Attribution–NonCommercial 4.0).
- You may use, study, and share it for non-commercial, educational, and research purposes, with attribution.
- Commercial use is not permitted.
- Provided "as is", without warranty of any kind; it may produce inaccurate output. Use at your own risk.
- Lineage: Qwen/Qwen3-1.7B-Base (Apache-2.0) → vinmlops/cpt-v2 → this SFT model. Please also respect the base model's license.
Citation / attribution
If you reference this educational project:
vinmlops/sft-v2 — instruction-tuned ML/LLM assistant (SFT on cpt-v2, merged),
an educational LLM/NN/transformer learning project. Non-commercial (CC-BY-NC-4.0).
Lineage: Qwen/Qwen3-1.7B-Base (Apache-2.0) -> cpt-v2 -> sft-v2.
- Downloads last month
- 102