Instructions to use yzhang318/etrader-niuma-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use yzhang318/etrader-niuma-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="yzhang318/etrader-niuma-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("yzhang318/etrader-niuma-9B") model = AutoModelForMultimodalLM.from_pretrained("yzhang318/etrader-niuma-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - MLX
How to use yzhang318/etrader-niuma-9B with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("yzhang318/etrader-niuma-9B") config = load_config("yzhang318/etrader-niuma-9B") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use yzhang318/etrader-niuma-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "yzhang318/etrader-niuma-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yzhang318/etrader-niuma-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/yzhang318/etrader-niuma-9B
- SGLang
How to use yzhang318/etrader-niuma-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "yzhang318/etrader-niuma-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yzhang318/etrader-niuma-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "yzhang318/etrader-niuma-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "yzhang318/etrader-niuma-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Pi
How to use yzhang318/etrader-niuma-9B with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "yzhang318/etrader-niuma-9B"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "yzhang318/etrader-niuma-9B" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use yzhang318/etrader-niuma-9B with Docker Model Runner:
docker model run hf.co/yzhang318/etrader-niuma-9B
- Hermes Agent
How to use yzhang318/etrader-niuma-9B with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "yzhang318/etrader-niuma-9B"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default yzhang318/etrader-niuma-9B
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use yzhang318/etrader-niuma-9B with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "yzhang318/etrader-niuma-9B"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "yzhang318/etrader-niuma-9B" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
etrader-niuma-9B
English | 中文
etrader-niuma-9B is Qwen/Qwen3.5-9B with a LoRA adapter merged into the bf16 weights. The adapter was trained on about 16k message→reply pairs from a private Chinese-language WeChat group of electricity-market trading practitioners. Given one chat message, the model answers the way a member of that group would: short, colloquial replies full of power-trading jargon and in-group slang.
The name: etrader is short for electricity trader. Niuma (牛马, "ox and horse") is Chinese internet slang for an overworked office worker.
The project was built end to end on a small budget: raw chat export → cleaning and pair mining → LoRA SFT on a single Apple-silicon Mac (MLX) → offline merge into Hugging Face bf16 shards → vLLM serving on one 24 GB GPU.
This is a style/persona model, not a knowledge model. Anything it says about prices, policies, companies, or people may be made up. Do not use it for trading decisions, and do not present its output as a real person's statement.
Model details
| Base model | Qwen/Qwen3.5-9B (post-trained, hybrid Gated DeltaNet + gated attention, 32 layers) |
| Base revision | c202236235762e1c871ad0ccb60c8ee5ba337b9a |
| Architecture | Qwen3_5ForConditionalGeneration; the vision tower and MTP head are included but unchanged |
| Fine-tuning | LoRA, rank 16, scale 2.0, dropout 0.05, on layers 24–31 (last 8 of 32) |
| LoRA targets | Linear-attention layers: in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj. Full-attention layers: q/k/v/o_proj. All 8 layers: gate/up/down_proj (62 matrices) |
| Trainable params | 10.82 M (0.12 % of the base) |
| Precision | Trained against an 8-bit MLX quantization of the base; merged into the original bf16 weights |
| Language | Simplified Chinese (colloquial, domain slang) |
| Context | Fine-tuned on short samples (≤ 512 tokens, ≤ 256 in phase 2). Long context is inherited from the base but was not trained or tested |
| License | Apache-2.0, inherited from the base model |
Quickstart
Prompt format
The model was trained on single-turn chats with a fixed system prompt. Use this one:
你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。
("You are a member of a WeChat group of electricity-trading practitioners. Reply to the previous message naturally and briefly, in group-chat style.")
The published prompt is a slightly generalized version of the training prompt, with one qualifier removed. On 300 held-out prompts it behaves the same (see Evaluation).
Recommended sampling: temperature=0.8, top_p=0.9, repetition_penalty=1.05, max_tokens=80, and thinking disabled (enable_thinking=False).
vLLM (tested, production setup)
vllm serve yzhang318/etrader-niuma-9B \
--served-model-name etrader-niuma-9B \
--language-model-only \
--default-chat-template-kwargs '{"enable_thinking": false}' \
--max-model-len 8192 --max-num-seqs 8 --gpu-memory-utilization 0.92
- This is the exact setup the author serves on a single 24 GB A10G (bf16, about 19 GB of weights), with vLLM 0.30 / torch 2.13 / CUDA 13.
--language-model-onlyskips the unused vision encoder.- If the host has no CUDA toolkit (
nvcc), also setVLLM_USE_FLASHINFER_SAMPLER=0.
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
SYSTEM = "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"
resp = client.chat.completions.create(
model="etrader-niuma-9B",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "今天新能源出力又拉满了,价格直接地板"},
],
temperature=0.8,
top_p=0.9,
max_tokens=80,
extra_body={"repetition_penalty": 1.05, "chat_template_kwargs": {"enable_thinking": False}},
)
print(resp.choices[0].message.content)
Transformers (reference)
This path needs a transformers release with Qwen3.5 support. The author serves with vLLM; this snippet is a reference, not the tested path.
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
repo = "yzhang318/etrader-niuma-9B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
messages = [
{"role": "system", "content": "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"},
{"role": "user", "content": "周末有人加班吗"},
]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8, top_p=0.9, repetition_penalty=1.05)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Multi-party chat
Training used single message→reply pairs. To run a simulated group:
- Treat everyone else's lines as
userturns and the persona's own earlier lines asassistantturns. - Merge consecutive same-role lines into one turn, and keep only the last several messages.
Longer histories work, but they are out of distribution.
Sample outputs (unedited, synthetic prompts)
| Prompt | Sampled replies (2 of 4 seeds shown) |
|---|---|
| 周末有人加班吗 | 我还在家补觉 大暑天 / 周末人挺少的 |
| 今天新能源出力又拉满了,价格直接地板 | 大亏大亏 中午都没买成啊 / 预测多少,实际多少啊,报日前更是离谱 |
| 有没有人用大模型做负荷预测的 | 有的 / 哦,这个模型参数多少啊,我找IT找我老板说下 |
| 明天现货价格大概率要涨吧? | [捂脸]其实没人知道 有没有日前申报啊 / 好像会吧 |
Training data
- Source. A private WeChat group export, containing text messages and quote-replies only. Images, files, links, and system messages were dropped. The data is not released and will not be.
- Turn building.
- Consecutive messages from the same sender within 60 s are merged into one turn.
- A turn is a reply in two cases: it comes from a different sender within 300 s of the previous turn, or it explicitly quotes an earlier message. For quote-replies, the quoted fragment is expanded back to its full source turn.
- Scrubbing.
- Mobile numbers are replaced with a placeholder, and WeChat IDs and
@mentionspans are removed. - Emoji-only and URL-only messages are dropped, and pairs longer than 800 characters are removed.
- Exact duplicate pairs are removed.
- Mobile numbers are replaced with a placeholder, and WeChat IDs and
- No speaker identity. The model never sees names or speaker IDs. It learns one blended "group member" persona.
- Split. Pairs are split by calendar day, so overlapping neighbouring pairs cannot leak across splits. The split is 15,984 train / 860 valid / 909 test pairs. The second training phase kept 15,966 train pairs of ≤ 256 tokens.
- Size. About 350 k supervised (reply) tokens over the full run. Replies have a median length of 11 characters.
Training procedure
Training used mlx-lm LoRA on one Apple-silicon Mac. The loss covers reply tokens only (mask_prompt: true), with the AdamW optimizer (weight decay 0.01), batch size 8, and gradient checkpointing.
| Phase | Global steps | LR schedule | Max seq len | Notes |
|---|---|---|---|---|
| 1 | 0 → 750 | warm-up 100 steps to 1e-4, then cosine | 512 | Hit OOM around step 855; resumed from step 750 |
| 2 | 750 → 3000 | cosine 8.8e-5 → 1e-6, no re-warm-up (continues phase 1's curve) | 256 | Filtered to ≤ 256 tokens (18 of 15,984 samples dropped) |
- Budget. 3,000 steps ≈ 1.5 epochs, at about 0.3 it/s and 35–40 supervised tok/s, with peak unified memory of 33.6 GB.
- Why only the last 8 layers? Qwen3.5's linear-attention (Gated DeltaNet) layers have no fused backward kernel in MLX, so backprop through them runs as a slow per-token loop. Limiting LoRA to the top 8 layers kept a full run to a few hours on a laptop-class machine.
- Checkpoint selection. Phase 2 validation loss (full valid set):
| Global step | 1000 | 1250 | 1500 | 1750 | 2000 | 2250 | 2500 | 2750 | 3000 |
|---|---|---|---|---|---|---|---|---|---|
| Val loss | 3.378 | 3.351 | 3.333 | 3.322 | 3.317 | 3.308 | 3.301 | 3.299 | 3.303 |
This release is global step 2750, the lowest validation loss of the run. The curve had flattened, and the loss ticked back up by step 3000.
Merge
MLX LoRA computes y = xWᵀ + s·(x·A)·B with A ∈ ℝ^{in×r}, B ∈ ℝ^{r×out} and s = 2.0. The merge is therefore:
W_merged = W + s · Bᵀ · Aᵀ (computed in fp32, cast back to bf16)
- MLX adapter keys
language_model.model.layers.N.*map to HF keysmodel.language_model.layers.N.*. - Only the 62 targeted matrices change. Embeddings, norms, vision tower, and MTP head are bit-identical to the base.
Evaluation
All numbers use the 909 held-out test pairs (disjoint days). There is no public benchmark for "sounds like this group", so the evaluation reports loss plus distribution-matching statistics against the real replies.
Reply-token loss (MLX, 8-bit base, same prompt format):
| Model | Test loss | Test PPL |
|---|---|---|
| Qwen3.5-9B (no adapter) | 6.055 | 426.3 |
| etrader-niuma-9B (step 2750) | 3.253 | 25.9 |
Reply-style statistics (MLX, T=0.8, top_p=0.9, rep=1.2; base measured on 100 prompts, fine-tuned on all 909):
| Set | Median len (chars) | Echo % | WeChat-emoji % | Unicode-emoji % | distinct-2 |
|---|---|---|---|---|---|
| Base Qwen3.5-9B | 56.5 | 2.0 | 10.0 | 69.0 | 0.663 |
| etrader-niuma-9B | 9 | 3.3 | 7.4 | 0.0 | 0.604 |
| Real replies | 11 | 1.1 | 14.0 | 0.8 | 0.578 |
Sampling sweep on the merged bf16 checkpoint (vLLM, first 300 test prompts, T=0.8, top_p=0.9):
| repetition_penalty | Median len | Mean len | Echo % | WeChat-emoji % | distinct-2 |
|---|---|---|---|---|---|
| 1.0 | 9 | 11.5 | 4.0 | 6.3 | 0.654 |
| 1.05 | 11 | 13.4 | 2.0 | 12.3 | 0.641 |
| 1.2 | 17 | 20.0 | 0.3 | 52.0 | 0.578 |
| 1.05 + published prompt | 10 | 12.6 | 2.3 | 13.3 | 0.672 |
| Real replies | 11 | 13.3 | 1.3 | 12.3 | 0.684 |
At repetition_penalty=1.05 the length and emoji rate match the real replies almost exactly. vLLM applies the penalty to prompt tokens too, so higher values push the model toward longer, emoji-heavy replies.
Memorization probe. Across all 909 test generations, 0 % of replies with ≥ 8 characters appeared verbatim in the training set. This is a weak lower-bound check, not a privacy guarantee (see below).
Limitations, bias and privacy
- Hallucination. Prices, volumes, rules, and "who said what" are generated in style, not retrieved. Expect confident nonsense.
- Memorization risk. Scrubbing is regex-based. The weights can still hold fragments of the source chat, such as names of organizations, places, events, or opinions. If you find output that identifies a real person or organization, please open a Discussion and the author will retrain with stronger filtering.
- Narrow persona. The model has one blended voice with short replies, and it is weak at long explanations. The base model's general ability is mostly retained but was not re-evaluated.
- Language. Chinese only. English prompts get Chinese group-chat replies at best.
- Quantization mismatch. The LoRA was learned against 8-bit weights and merged into bf16. The vLLM statistics above show the behaviour carries over, but the logits are not identical to the training-time model.
Out of scope: trading or investment advice, impersonating identifiable individuals, harassment, and generating content presented as coming from real market participants.
Roadmap: suggested next training steps
- Train in bf16 on CUDA with all layers. Use fused Gated DeltaNet kernels (e.g.
flash-linear-attention) so LoRA can cover all 32 layers at rank 32–64, or run DoRA/rsLoRA ablations. This also removes the 8-bit→bf16 mismatch. - Multi-turn, speaker-aware SFT. Replace single pairs with sliding windows of 8–16 turns. Use anonymized speaker tags (
<spk_07>) so one model can play distinct, consistent members, and add time-gap tokens so it learns when not to reply. - Stronger de-identification before retraining. Use Chinese NER (PER/ORG/LOC) with consistent pseudonym mapping, dedupe near-duplicates with MinHash, and optionally apply DP-SGD or deduplication-aware sampling to cut extraction risk. Validate with canary insertion and extraction attacks, not just verbatim-match rates.
- Preference tuning. Run DPO/KTO/ORPO with the real reply as chosen and base-model or over-long replies as rejected. This targets remaining failure modes: echoing the prompt, emoji overuse, and "assistant voice" leaking through.
- Grounding. Add retrieval over public market rules, clearing prices, and announcements, plus tool calls, so the persona can be right and not only plausible. Keep style and facts in separate components.
- Better evaluation. Use blind pairwise human tests (real vs. generated), an LLM-as-judge rubric for in-group plausibility, a style classifier, and a temporal split (train on earlier months, test on later) to measure drift.
- MTP head refresh. Fine-tune or distill the multi-token-prediction head on the adapted model to recover speculative-decoding acceptance rates.
- Smaller artifacts. Release AWQ/GPTQ int4, GGUF, and MLX 4/8-bit variants for consumer GPUs and Macs, and publish the LoRA separately for adapter-swapping.
About
Built by @yzhang318. The author did the data pipeline, training, weight surgery, evaluation, and serving. If you work on LLMs for power and energy markets, or on low-budget persona fine-tuning, feel free to open a Discussion.
@misc{etrader_niuma_9b_2026,
title = {etrader-niuma-9B: a group-chat persona LoRA on Qwen3.5-9B},
author = {yzhang318},
year = {2026},
howpublished = {\url{https://huggingface.co/yzhang318/etrader-niuma-9B}}
}
etrader-niuma-9B(中文说明)
English | 中文
etrader-niuma-9B 以 Qwen/Qwen3.5-9B 为基座,训练了一个 LoRA 适配器,并已合并进 bf16 权重。训练数据是一个私有中文微信群(电力交易从业者群)里约 1.6 万条「消息→回复」对。输入一条群消息,模型会像群友一样回复:简短、口语化,满是电力交易行话和群内梗。
名字的含义:etrader 即 electricity trader(电力交易员)。niuma 即「牛马」,网络用语,指辛苦打工人。
整个项目低成本、端到端完成:聊天记录导出 → 清洗与配对 → 单台 Apple Silicon Mac 上用 MLX 做 LoRA 微调 → 离线合并进 Hugging Face bf16 分片 → 单卡 24 GB GPU 用 vLLM 部署。
这是风格/人设模型,不是知识模型。它说的价格、政策、公司、人物都可能是编的。请勿用于交易决策,也不要把输出当作真实人物的发言。
模型信息
| 基座 | Qwen/Qwen3.5-9B(后训练版;Gated DeltaNet 与门控注意力混合结构,32 层) |
| 基座版本 | c202236235762e1c871ad0ccb60c8ee5ba337b9a |
| 架构 | Qwen3_5ForConditionalGeneration;视觉塔和 MTP 头保留,但未改动 |
| 微调方式 | LoRA,rank 16,scale 2.0,dropout 0.05,作用于第 24–31 层(32 层中的最后 8 层) |
| LoRA 目标 | 线性注意力层:in_proj_qkv、in_proj_z、in_proj_a、in_proj_b、out_proj。全注意力层:q/k/v/o_proj。8 层全部:gate/up/down_proj(共 62 个矩阵) |
| 可训练参数 | 1082 万(基座的 0.12%) |
| 精度 | 训练时基座为 MLX 8-bit 量化;合并到原始 bf16 权重 |
| 语言 | 简体中文(口语、行业黑话) |
| 上下文 | 微调样本很短(不超过 512 token,第二阶段不超过 256)。长上下文能力继承自基座,但未训练、未测试 |
| 许可证 | Apache-2.0,继承自基座 |
快速上手
提示词格式
训练用的是单轮对话和固定系统提示词。请使用:
你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。
这是训练提示词的略微泛化版本,去掉了一个限定词。在 300 条留出样本上,两者表现一致(见评测)。
推荐采样参数: temperature=0.8、top_p=0.9、repetition_penalty=1.05、max_tokens=80,并关闭思考模式(enable_thinking=False)。
vLLM(已验证的生产配置)
vllm serve yzhang318/etrader-niuma-9B \
--served-model-name etrader-niuma-9B \
--language-model-only \
--default-chat-template-kwargs '{"enable_thinking": false}' \
--max-model-len 8192 --max-num-seqs 8 --gpu-memory-utilization 0.92
- 这就是作者在单张 24 GB A10G 上的实际部署配置(bf16,权重约 19 GB),环境为 vLLM 0.30 / torch 2.13 / CUDA 13。
--language-model-only跳过用不到的视觉编码器。- 如果机器上没有 CUDA toolkit(
nvcc),还需设置VLLM_USE_FLASHINFER_SAMPLER=0。
服务兼容 OpenAI 接口。repetition_penalty 和 chat_template_kwargs 通过 extra_body 传入:
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
SYSTEM = "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"
resp = client.chat.completions.create(
model="etrader-niuma-9B",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "今天新能源出力又拉满了,价格直接地板"},
],
temperature=0.8,
top_p=0.9,
max_tokens=80,
extra_body={"repetition_penalty": 1.05, "chat_template_kwargs": {"enable_thinking": False}},
)
print(resp.choices[0].message.content)
Transformers(参考)
需要支持 Qwen3.5 的 transformers 版本。作者线上用的是 vLLM,以下示例仅作参考,未经实测。
import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer
repo = "yzhang318/etrader-niuma-9B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
messages = [
{"role": "system", "content": "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"},
{"role": "user", "content": "周末有人加班吗"},
]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8, top_p=0.9, repetition_penalty=1.05)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
多人群聊
训练数据是单条「消息→回复」对。要模拟群聊:
- 把其他人的发言当作
user轮,把该角色自己之前的发言当作assistant轮。 - 相邻的同角色发言合并成一轮,只保留最近若干条。
更长的历史也能用,但超出了训练分布。
输出示例(未经编辑,提示词为虚构)
| 提示 | 采样回复(4 个种子中展示 2 个) |
|---|---|
| 周末有人加班吗 | 我还在家补觉 大暑天 / 周末人挺少的 |
| 今天新能源出力又拉满了,价格直接地板 | 大亏大亏 中午都没买成啊 / 预测多少,实际多少啊,报日前更是离谱 |
| 有没有人用大模型做负荷预测的 | 有的 / 哦,这个模型参数多少啊,我找IT找我老板说下 |
| 明天现货价格大概率要涨吧? | [捂脸]其实没人知道 有没有日前申报啊 / 好像会吧 |
训练数据
- 来源: 私有微信群导出,只保留文本消息和引用回复。图片、文件、链接和系统消息均已丢弃。数据不公开,今后也不会公开。
- 轮次构建:
- 同一发送者 60 秒内的连续消息合并为一轮。
- 以下两种情况算作「回复」:300 秒内由不同发送者接话,或显式引用了之前的某条消息。引用回复会把被引片段还原成完整的原始轮次。
- 脱敏:
- 手机号替换为占位符,微信 ID 和
@提及删除。 - 纯表情和纯链接消息丢弃,超过 800 字的样本剔除。
- 完全重复的样本去重。
- 手机号替换为占位符,微信 ID 和
- 无说话人身份: 模型看不到任何昵称或 ID,学到的是一个融合的「群友」人设。
- 切分: 按自然日切分,避免相邻样本跨集合泄漏。训练集 15,984 对,验证集 860 对,测试集 909 对。第二阶段训练只保留不超过 256 token 的 15,966 对。
- 规模: 全程约 35 万个监督(回复)token。回复长度中位数为 11 个字。
训练过程
训练使用 mlx-lm LoRA,在单台 Apple Silicon Mac 上完成。只对回复 token 计算损失(mask_prompt: true),优化器为 AdamW(weight decay 0.01),batch size 8,开启梯度检查点。
| 阶段 | 全局步数 | 学习率 | 最大长度 | 备注 |
|---|---|---|---|---|
| 1 | 0 → 750 | 100 步预热到 1e-4,之后余弦衰减 | 512 | 约 855 步时 OOM,从第 750 步续训 |
| 2 | 750 → 3000 | 余弦 8.8e-5 → 1e-6,不重新预热(延续阶段 1 的曲线) | 256 | 过滤掉超过 256 token 的样本(15,984 条中去掉 18 条) |
- 训练量: 共 3000 步,约 1.5 个 epoch。速度约 0.3 it/s、每秒 35–40 个监督 token,统一内存峰值 33.6 GB。
- 为什么只训最后 8 层? MLX 里 Qwen3.5 线性注意力(Gated DeltaNet)层的反向传播没有融合算子,只能逐 token 循环,非常慢。只训最后 8 层,完整一轮训练在笔记本级机器上几个小时就能跑完。
- 选点: 第二阶段在完整验证集上的损失:
| 全局步数 | 1000 | 1250 | 1500 | 1750 | 2000 | 2250 | 2500 | 2750 | 3000 |
|---|---|---|---|---|---|---|---|---|---|
| 验证损失 | 3.378 | 3.351 | 3.333 | 3.322 | 3.317 | 3.308 | 3.301 | 3.299 | 3.303 |
本仓库发布的是全局第 2750 步,这是整轮训练中验证损失最低的一点。曲线已经走平,到第 3000 步损失略有回升。
权重合并
MLX LoRA 的计算是 y = xWᵀ + s·(x·A)·B,其中 A ∈ ℝ^{in×r}、B ∈ ℝ^{r×out}、s = 2.0。因此合并公式为:
W_merged = W + s · Bᵀ · Aᵀ (fp32 计算,再转回 bf16)
- MLX 的键名
language_model.model.layers.N.*对应 HF 的model.language_model.layers.N.*。 - 只有 62 个目标矩阵发生变化。词嵌入、归一化层、视觉塔和 MTP 头与基座逐位相同。
评测
所有数字都基于 909 条留出的测试对(日期与训练集不重叠)。「像不像这个群」没有公开基准,所以这里报告损失,以及与真实回复的分布匹配统计。
回复 token 损失(MLX,8-bit 基座,同一提示格式):
| 模型 | 测试损失 | 测试困惑度 |
|---|---|---|
| Qwen3.5-9B(无适配器) | 6.055 | 426.3 |
| etrader-niuma-9B(第 2750 步) | 3.253 | 25.9 |
回复风格统计(MLX,T=0.8, top_p=0.9, rep=1.2;基座测 100 条,微调模型测全部 909 条):
| 集合 | 长度中位数(字) | 复读率 % | 微信表情 % | Unicode 表情 % | distinct-2 |
|---|---|---|---|---|---|
| 基座 Qwen3.5-9B | 56.5 | 2.0 | 10.0 | 69.0 | 0.663 |
| etrader-niuma-9B | 9 | 3.3 | 7.4 | 0.0 | 0.604 |
| 真实回复 | 11 | 1.1 | 14.0 | 0.8 | 0.578 |
合并后 bf16 权重的采样参数扫描(vLLM,前 300 条测试提示,T=0.8, top_p=0.9):
| repetition_penalty | 长度中位数 | 平均长度 | 复读率 % | 微信表情 % | distinct-2 |
|---|---|---|---|---|---|
| 1.0 | 9 | 11.5 | 4.0 | 6.3 | 0.654 |
| 1.05 | 11 | 13.4 | 2.0 | 12.3 | 0.641 |
| 1.2 | 17 | 20.0 | 0.3 | 52.0 | 0.578 |
| 1.05 + 公开提示词 | 10 | 12.6 | 2.3 | 13.3 | 0.672 |
| 真实回复 | 11 | 13.3 | 1.3 | 12.3 | 0.684 |
repetition_penalty=1.05 时,长度和表情比例几乎与真实回复一致。vLLM 的重复惩罚同样作用于提示词 token,所以数值越高,回复越长、表情越多。
记忆探测: 在 909 条测试生成中,长度不少于 8 个字的回复里,逐字出现在训练集中的比例为 0%。这只是一个很弱的下界检查,不代表隐私保证(见下文)。
局限、偏差与隐私
- 幻觉: 价格、电量、规则、「谁说过什么」都是按风格生成的,不是检索出来的。它会一本正经地胡说。
- 记忆风险: 脱敏只基于正则。权重中仍可能残留原群聊的片段,例如机构名、地名、事件或观点。如果发现能识别真实个人或机构的输出,请在 Discussion 中反馈,作者会加强过滤后重新训练。
- 人设单一: 只有一种融合的口吻,回复偏短,不擅长长篇解释。基座的通用能力大体保留,但没有重新评测。
- 语言: 仅支持中文。英文提问最多得到中文群聊式回复。
- 量化失配: LoRA 是在 8-bit 权重上学到的,却合并进 bf16。上面的 vLLM 统计说明行为基本一致,但 logits 与训练时的模型并不完全相同。
不适用场景: 交易或投资建议、冒充可识别的个人、骚扰,以及把生成内容当作真实市场参与者的发言。
后续训练路线
- 在 CUDA 上做 bf16 全层训练: 借助融合的 Gated DeltaNet 算子(如
flash-linear-attention),让 LoRA 覆盖全部 32 层、rank 提到 32–64,或者做 DoRA/rsLoRA 对比实验。这样也能消除 8-bit→bf16 失配。 - 多轮、区分说话人的 SFT: 用 8–16 轮滑动窗口替代单条配对。加入匿名说话人标记(如
<spk_07>),让同一个模型扮演多个一致的群友;再加入时间间隔 token,让模型学会什么时候不回复。 - 重训前加强去标识化: 用中文 NER(人名、机构、地名)做一致的假名映射,用 MinHash 去近似重复,可选 DP-SGD 或去重感知采样来降低被提取的风险。评估时用 canary 注入和提取攻击,而不只看逐字匹配率。
- 偏好优化: 用 DPO、KTO 或 ORPO,把真实回复作为 chosen,基座回复或过长回复作为 rejected。目标是剩下的几类问题:复读提示、滥用表情、「AI 助手腔」外泄。
- 知识增强: 接入公开的市场规则、出清价格和公告做检索增强,再加上工具调用,让人设不只是「像」,还能「对」。风格和事实分别由不同组件负责。
- 更好的评测: 人工盲测(真实与生成两两对比)、LLM-as-judge 打分、风格分类器,以及按时间切分(用前几个月训练、后几个月测试)来衡量漂移。
- 更新 MTP 头: 在适配后的模型上微调或蒸馏多 token 预测头,恢复投机解码的接受率。
- 更小的发布版本: 发布 AWQ/GPTQ int4、GGUF、MLX 4/8-bit 版本,方便消费级显卡和 Mac 使用;同时单独发布 LoRA,便于热切换适配器。
关于作者
由 @yzhang318 独立完成,包括数据管线、训练、权重合并、评测和部署。如果你也在做电力/能源市场方向的大模型,或者低成本人设微调,欢迎在 Discussion 里交流。
- Downloads last month
- 22
Quantized