etrader-niuma-9B

English | 中文

etrader-niuma-9B is Qwen/Qwen3.5-9B with a LoRA adapter merged into the bf16 weights. The adapter was trained on about 16k message→reply pairs from a private Chinese-language WeChat group of electricity-market trading practitioners. Given one chat message, the model answers the way a member of that group would: short, colloquial replies full of power-trading jargon and in-group slang.

The name: etrader is short for electricity trader. Niuma (牛马, "ox and horse") is Chinese internet slang for an overworked office worker.

The project was built end to end on a small budget: raw chat export → cleaning and pair mining → LoRA SFT on a single Apple-silicon Mac (MLX) → offline merge into Hugging Face bf16 shards → vLLM serving on one 24 GB GPU.

This is a style/persona model, not a knowledge model. Anything it says about prices, policies, companies, or people may be made up. Do not use it for trading decisions, and do not present its output as a real person's statement.

Model details

Base model Qwen/Qwen3.5-9B (post-trained, hybrid Gated DeltaNet + gated attention, 32 layers)
Base revision c202236235762e1c871ad0ccb60c8ee5ba337b9a
Architecture Qwen3_5ForConditionalGeneration; the vision tower and MTP head are included but unchanged
Fine-tuning LoRA, rank 16, scale 2.0, dropout 0.05, on layers 24–31 (last 8 of 32)
LoRA targets Linear-attention layers: in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj. Full-attention layers: q/k/v/o_proj. All 8 layers: gate/up/down_proj (62 matrices)
Trainable params 10.82 M (0.12 % of the base)
Precision Trained against an 8-bit MLX quantization of the base; merged into the original bf16 weights
Language Simplified Chinese (colloquial, domain slang)
Context Fine-tuned on short samples (≤ 512 tokens, ≤ 256 in phase 2). Long context is inherited from the base but was not trained or tested
License Apache-2.0, inherited from the base model

Quickstart

Prompt format

The model was trained on single-turn chats with a fixed system prompt. Use this one:

你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。

("You are a member of a WeChat group of electricity-trading practitioners. Reply to the previous message naturally and briefly, in group-chat style.")

The published prompt is a slightly generalized version of the training prompt, with one qualifier removed. On 300 held-out prompts it behaves the same (see Evaluation).

Recommended sampling: temperature=0.8, top_p=0.9, repetition_penalty=1.05, max_tokens=80, and thinking disabled (enable_thinking=False).

vLLM (tested, production setup)

vllm serve yzhang318/etrader-niuma-9B \
  --served-model-name etrader-niuma-9B \
  --language-model-only \
  --default-chat-template-kwargs '{"enable_thinking": false}' \
  --max-model-len 8192 --max-num-seqs 8 --gpu-memory-utilization 0.92
  • This is the exact setup the author serves on a single 24 GB A10G (bf16, about 19 GB of weights), with vLLM 0.30 / torch 2.13 / CUDA 13.
  • --language-model-only skips the unused vision encoder.
  • If the host has no CUDA toolkit (nvcc), also set VLLM_USE_FLASHINFER_SAMPLER=0.
from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
SYSTEM = "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"

resp = client.chat.completions.create(
    model="etrader-niuma-9B",
    messages=[
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "今天新能源出力又拉满了,价格直接地板"},
    ],
    temperature=0.8,
    top_p=0.9,
    max_tokens=80,
    extra_body={"repetition_penalty": 1.05, "chat_template_kwargs": {"enable_thinking": False}},
)
print(resp.choices[0].message.content)

Transformers (reference)

This path needs a transformers release with Qwen3.5 support. The author serves with vLLM; this snippet is a reference, not the tested path.

import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer

repo = "yzhang318/etrader-niuma-9B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

messages = [
    {"role": "system", "content": "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"},
    {"role": "user", "content": "周末有人加班吗"},
]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8, top_p=0.9, repetition_penalty=1.05)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Multi-party chat

Training used single message→reply pairs. To run a simulated group:

  • Treat everyone else's lines as user turns and the persona's own earlier lines as assistant turns.
  • Merge consecutive same-role lines into one turn, and keep only the last several messages.

Longer histories work, but they are out of distribution.

Sample outputs (unedited, synthetic prompts)

Prompt Sampled replies (2 of 4 seeds shown)
周末有人加班吗 我还在家补觉 大暑天 / 周末人挺少的
今天新能源出力又拉满了,价格直接地板 大亏大亏 中午都没买成啊 / 预测多少,实际多少啊,报日前更是离谱
有没有人用大模型做负荷预测的 有的 / 哦,这个模型参数多少啊,我找IT找我老板说下
明天现货价格大概率要涨吧? [捂脸]其实没人知道 有没有日前申报啊 / 好像会吧

Training data

  • Source. A private WeChat group export, containing text messages and quote-replies only. Images, files, links, and system messages were dropped. The data is not released and will not be.
  • Turn building.
    • Consecutive messages from the same sender within 60 s are merged into one turn.
    • A turn is a reply in two cases: it comes from a different sender within 300 s of the previous turn, or it explicitly quotes an earlier message. For quote-replies, the quoted fragment is expanded back to its full source turn.
  • Scrubbing.
    • Mobile numbers are replaced with a placeholder, and WeChat IDs and @mention spans are removed.
    • Emoji-only and URL-only messages are dropped, and pairs longer than 800 characters are removed.
    • Exact duplicate pairs are removed.
  • No speaker identity. The model never sees names or speaker IDs. It learns one blended "group member" persona.
  • Split. Pairs are split by calendar day, so overlapping neighbouring pairs cannot leak across splits. The split is 15,984 train / 860 valid / 909 test pairs. The second training phase kept 15,966 train pairs of ≤ 256 tokens.
  • Size. About 350 k supervised (reply) tokens over the full run. Replies have a median length of 11 characters.

Training procedure

Training used mlx-lm LoRA on one Apple-silicon Mac. The loss covers reply tokens only (mask_prompt: true), with the AdamW optimizer (weight decay 0.01), batch size 8, and gradient checkpointing.

Phase Global steps LR schedule Max seq len Notes
1 0 → 750 warm-up 100 steps to 1e-4, then cosine 512 Hit OOM around step 855; resumed from step 750
2 750 → 3000 cosine 8.8e-5 → 1e-6, no re-warm-up (continues phase 1's curve) 256 Filtered to ≤ 256 tokens (18 of 15,984 samples dropped)
  • Budget. 3,000 steps ≈ 1.5 epochs, at about 0.3 it/s and 35–40 supervised tok/s, with peak unified memory of 33.6 GB.
  • Why only the last 8 layers? Qwen3.5's linear-attention (Gated DeltaNet) layers have no fused backward kernel in MLX, so backprop through them runs as a slow per-token loop. Limiting LoRA to the top 8 layers kept a full run to a few hours on a laptop-class machine.
  • Checkpoint selection. Phase 2 validation loss (full valid set):
Global step 1000 1250 1500 1750 2000 2250 2500 2750 3000
Val loss 3.378 3.351 3.333 3.322 3.317 3.308 3.301 3.299 3.303

This release is global step 2750, the lowest validation loss of the run. The curve had flattened, and the loss ticked back up by step 3000.

Merge

MLX LoRA computes y = xWᵀ + s·(x·A)·B with A ∈ ℝ^{in×r}, B ∈ ℝ^{r×out} and s = 2.0. The merge is therefore:

W_merged = W + s · Bᵀ · Aᵀ        (computed in fp32, cast back to bf16)
  • MLX adapter keys language_model.model.layers.N.* map to HF keys model.language_model.layers.N.*.
  • Only the 62 targeted matrices change. Embeddings, norms, vision tower, and MTP head are bit-identical to the base.

Evaluation

All numbers use the 909 held-out test pairs (disjoint days). There is no public benchmark for "sounds like this group", so the evaluation reports loss plus distribution-matching statistics against the real replies.

Reply-token loss (MLX, 8-bit base, same prompt format):

Model Test loss Test PPL
Qwen3.5-9B (no adapter) 6.055 426.3
etrader-niuma-9B (step 2750) 3.253 25.9

Reply-style statistics (MLX, T=0.8, top_p=0.9, rep=1.2; base measured on 100 prompts, fine-tuned on all 909):

Set Median len (chars) Echo % WeChat-emoji % Unicode-emoji % distinct-2
Base Qwen3.5-9B 56.5 2.0 10.0 69.0 0.663
etrader-niuma-9B 9 3.3 7.4 0.0 0.604
Real replies 11 1.1 14.0 0.8 0.578

Sampling sweep on the merged bf16 checkpoint (vLLM, first 300 test prompts, T=0.8, top_p=0.9):

repetition_penalty Median len Mean len Echo % WeChat-emoji % distinct-2
1.0 9 11.5 4.0 6.3 0.654
1.05 11 13.4 2.0 12.3 0.641
1.2 17 20.0 0.3 52.0 0.578
1.05 + published prompt 10 12.6 2.3 13.3 0.672
Real replies 11 13.3 1.3 12.3 0.684

At repetition_penalty=1.05 the length and emoji rate match the real replies almost exactly. vLLM applies the penalty to prompt tokens too, so higher values push the model toward longer, emoji-heavy replies.

Memorization probe. Across all 909 test generations, 0 % of replies with ≥ 8 characters appeared verbatim in the training set. This is a weak lower-bound check, not a privacy guarantee (see below).

Limitations, bias and privacy

  • Hallucination. Prices, volumes, rules, and "who said what" are generated in style, not retrieved. Expect confident nonsense.
  • Memorization risk. Scrubbing is regex-based. The weights can still hold fragments of the source chat, such as names of organizations, places, events, or opinions. If you find output that identifies a real person or organization, please open a Discussion and the author will retrain with stronger filtering.
  • Narrow persona. The model has one blended voice with short replies, and it is weak at long explanations. The base model's general ability is mostly retained but was not re-evaluated.
  • Language. Chinese only. English prompts get Chinese group-chat replies at best.
  • Quantization mismatch. The LoRA was learned against 8-bit weights and merged into bf16. The vLLM statistics above show the behaviour carries over, but the logits are not identical to the training-time model.

Out of scope: trading or investment advice, impersonating identifiable individuals, harassment, and generating content presented as coming from real market participants.

Roadmap: suggested next training steps

  1. Train in bf16 on CUDA with all layers. Use fused Gated DeltaNet kernels (e.g. flash-linear-attention) so LoRA can cover all 32 layers at rank 32–64, or run DoRA/rsLoRA ablations. This also removes the 8-bit→bf16 mismatch.
  2. Multi-turn, speaker-aware SFT. Replace single pairs with sliding windows of 8–16 turns. Use anonymized speaker tags (<spk_07>) so one model can play distinct, consistent members, and add time-gap tokens so it learns when not to reply.
  3. Stronger de-identification before retraining. Use Chinese NER (PER/ORG/LOC) with consistent pseudonym mapping, dedupe near-duplicates with MinHash, and optionally apply DP-SGD or deduplication-aware sampling to cut extraction risk. Validate with canary insertion and extraction attacks, not just verbatim-match rates.
  4. Preference tuning. Run DPO/KTO/ORPO with the real reply as chosen and base-model or over-long replies as rejected. This targets remaining failure modes: echoing the prompt, emoji overuse, and "assistant voice" leaking through.
  5. Grounding. Add retrieval over public market rules, clearing prices, and announcements, plus tool calls, so the persona can be right and not only plausible. Keep style and facts in separate components.
  6. Better evaluation. Use blind pairwise human tests (real vs. generated), an LLM-as-judge rubric for in-group plausibility, a style classifier, and a temporal split (train on earlier months, test on later) to measure drift.
  7. MTP head refresh. Fine-tune or distill the multi-token-prediction head on the adapted model to recover speculative-decoding acceptance rates.
  8. Smaller artifacts. Release AWQ/GPTQ int4, GGUF, and MLX 4/8-bit variants for consumer GPUs and Macs, and publish the LoRA separately for adapter-swapping.

About

Built by @yzhang318. The author did the data pipeline, training, weight surgery, evaluation, and serving. If you work on LLMs for power and energy markets, or on low-budget persona fine-tuning, feel free to open a Discussion.

@misc{etrader_niuma_9b_2026,
  title        = {etrader-niuma-9B: a group-chat persona LoRA on Qwen3.5-9B},
  author       = {yzhang318},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/yzhang318/etrader-niuma-9B}}
}

etrader-niuma-9B(中文说明)

English | 中文

etrader-niuma-9B 以 Qwen/Qwen3.5-9B 为基座,训练了一个 LoRA 适配器,并已合并进 bf16 权重。训练数据是一个私有中文微信群(电力交易从业者群)里约 1.6 万条「消息→回复」对。输入一条群消息,模型会像群友一样回复:简短、口语化,满是电力交易行话和群内梗。

名字的含义:etrader 即 electricity trader(电力交易员)。niuma 即「牛马」,网络用语,指辛苦打工人。

整个项目低成本、端到端完成:聊天记录导出 → 清洗与配对 → 单台 Apple Silicon Mac 上用 MLX 做 LoRA 微调 → 离线合并进 Hugging Face bf16 分片 → 单卡 24 GB GPU 用 vLLM 部署。

这是风格/人设模型,不是知识模型。它说的价格、政策、公司、人物都可能是编的。请勿用于交易决策,也不要把输出当作真实人物的发言。

模型信息

基座 Qwen/Qwen3.5-9B(后训练版;Gated DeltaNet 与门控注意力混合结构,32 层)
基座版本 c202236235762e1c871ad0ccb60c8ee5ba337b9a
架构 Qwen3_5ForConditionalGeneration;视觉塔和 MTP 头保留,但未改动
微调方式 LoRA,rank 16,scale 2.0,dropout 0.05,作用于第 24–31 层(32 层中的最后 8 层)
LoRA 目标 线性注意力层:in_proj_qkv、in_proj_z、in_proj_a、in_proj_b、out_proj。全注意力层:q/k/v/o_proj。8 层全部:gate/up/down_proj(共 62 个矩阵)
可训练参数 1082 万(基座的 0.12%)
精度 训练时基座为 MLX 8-bit 量化;合并到原始 bf16 权重
语言 简体中文(口语、行业黑话)
上下文 微调样本很短(不超过 512 token,第二阶段不超过 256)。长上下文能力继承自基座,但未训练、未测试
许可证 Apache-2.0,继承自基座

快速上手

提示词格式

训练用的是单轮对话和固定系统提示词。请使用:

你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。

这是训练提示词的略微泛化版本,去掉了一个限定词。在 300 条留出样本上,两者表现一致(见评测)。

推荐采样参数: temperature=0.8、top_p=0.9、repetition_penalty=1.05、max_tokens=80,并关闭思考模式(enable_thinking=False)。

vLLM(已验证的生产配置)

vllm serve yzhang318/etrader-niuma-9B \
  --served-model-name etrader-niuma-9B \
  --language-model-only \
  --default-chat-template-kwargs '{"enable_thinking": false}' \
  --max-model-len 8192 --max-num-seqs 8 --gpu-memory-utilization 0.92
  • 这就是作者在单张 24 GB A10G 上的实际部署配置(bf16,权重约 19 GB),环境为 vLLM 0.30 / torch 2.13 / CUDA 13。
  • --language-model-only 跳过用不到的视觉编码器。
  • 如果机器上没有 CUDA toolkit(nvcc),还需设置 VLLM_USE_FLASHINFER_SAMPLER=0。

服务兼容 OpenAI 接口。repetition_penalty 和 chat_template_kwargs 通过 extra_body 传入:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
SYSTEM = "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"

resp = client.chat.completions.create(
    model="etrader-niuma-9B",
    messages=[
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "今天新能源出力又拉满了,价格直接地板"},
    ],
    temperature=0.8,
    top_p=0.9,
    max_tokens=80,
    extra_body={"repetition_penalty": 1.05, "chat_template_kwargs": {"enable_thinking": False}},
)
print(resp.choices[0].message.content)

Transformers(参考)

需要支持 Qwen3.5 的 transformers 版本。作者线上用的是 vLLM,以下示例仅作参考,未经实测。

import torch
from transformers import AutoModelForImageTextToText, AutoTokenizer

repo = "yzhang318/etrader-niuma-9B"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForImageTextToText.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

messages = [
    {"role": "system", "content": "你是电力交易从业者微信群里的一位群友。请用群聊的口吻,自然、简短地回复上一条消息。"},
    {"role": "user", "content": "周末有人加班吗"},
]
text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True, enable_thinking=False)
inputs = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=80, do_sample=True, temperature=0.8, top_p=0.9, repetition_penalty=1.05)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

多人群聊

训练数据是单条「消息→回复」对。要模拟群聊:

  • 把其他人的发言当作 user 轮,把该角色自己之前的发言当作 assistant 轮。
  • 相邻的同角色发言合并成一轮,只保留最近若干条。

更长的历史也能用,但超出了训练分布。

输出示例(未经编辑,提示词为虚构)

提示 采样回复(4 个种子中展示 2 个)
周末有人加班吗 我还在家补觉 大暑天 / 周末人挺少的
今天新能源出力又拉满了,价格直接地板 大亏大亏 中午都没买成啊 / 预测多少,实际多少啊,报日前更是离谱
有没有人用大模型做负荷预测的 有的 / 哦,这个模型参数多少啊,我找IT找我老板说下
明天现货价格大概率要涨吧? [捂脸]其实没人知道 有没有日前申报啊 / 好像会吧

训练数据

  • 来源: 私有微信群导出,只保留文本消息和引用回复。图片、文件、链接和系统消息均已丢弃。数据不公开,今后也不会公开。
  • 轮次构建:
    • 同一发送者 60 秒内的连续消息合并为一轮。
    • 以下两种情况算作「回复」:300 秒内由不同发送者接话,或显式引用了之前的某条消息。引用回复会把被引片段还原成完整的原始轮次。
  • 脱敏:
    • 手机号替换为占位符,微信 ID 和 @提及 删除。
    • 纯表情和纯链接消息丢弃,超过 800 字的样本剔除。
    • 完全重复的样本去重。
  • 无说话人身份: 模型看不到任何昵称或 ID,学到的是一个融合的「群友」人设。
  • 切分: 按自然日切分,避免相邻样本跨集合泄漏。训练集 15,984 对,验证集 860 对,测试集 909 对。第二阶段训练只保留不超过 256 token 的 15,966 对。
  • 规模: 全程约 35 万个监督(回复)token。回复长度中位数为 11 个字。

训练过程

训练使用 mlx-lm LoRA,在单台 Apple Silicon Mac 上完成。只对回复 token 计算损失(mask_prompt: true),优化器为 AdamW(weight decay 0.01),batch size 8,开启梯度检查点。

阶段 全局步数 学习率 最大长度 备注
1 0 → 750 100 步预热到 1e-4,之后余弦衰减 512 约 855 步时 OOM,从第 750 步续训
2 750 → 3000 余弦 8.8e-5 → 1e-6,不重新预热(延续阶段 1 的曲线) 256 过滤掉超过 256 token 的样本(15,984 条中去掉 18 条)
  • 训练量: 共 3000 步,约 1.5 个 epoch。速度约 0.3 it/s、每秒 35–40 个监督 token,统一内存峰值 33.6 GB。
  • 为什么只训最后 8 层? MLX 里 Qwen3.5 线性注意力(Gated DeltaNet)层的反向传播没有融合算子,只能逐 token 循环,非常慢。只训最后 8 层,完整一轮训练在笔记本级机器上几个小时就能跑完。
  • 选点: 第二阶段在完整验证集上的损失:
全局步数 1000 1250 1500 1750 2000 2250 2500 2750 3000
验证损失 3.378 3.351 3.333 3.322 3.317 3.308 3.301 3.299 3.303

本仓库发布的是全局第 2750 步,这是整轮训练中验证损失最低的一点。曲线已经走平,到第 3000 步损失略有回升。

权重合并

MLX LoRA 的计算是 y = xWᵀ + s·(x·A)·B,其中 A ∈ ℝ^{in×r}、B ∈ ℝ^{r×out}、s = 2.0。因此合并公式为:

W_merged = W + s · Bᵀ · Aᵀ        (fp32 计算,再转回 bf16)
  • MLX 的键名 language_model.model.layers.N.* 对应 HF 的 model.language_model.layers.N.*。
  • 只有 62 个目标矩阵发生变化。词嵌入、归一化层、视觉塔和 MTP 头与基座逐位相同。

评测

所有数字都基于 909 条留出的测试对(日期与训练集不重叠)。「像不像这个群」没有公开基准,所以这里报告损失,以及与真实回复的分布匹配统计。

回复 token 损失(MLX,8-bit 基座,同一提示格式):

模型 测试损失 测试困惑度
Qwen3.5-9B(无适配器) 6.055 426.3
etrader-niuma-9B(第 2750 步) 3.253 25.9

回复风格统计(MLX,T=0.8, top_p=0.9, rep=1.2;基座测 100 条,微调模型测全部 909 条):

集合 长度中位数(字) 复读率 % 微信表情 % Unicode 表情 % distinct-2
基座 Qwen3.5-9B 56.5 2.0 10.0 69.0 0.663
etrader-niuma-9B 9 3.3 7.4 0.0 0.604
真实回复 11 1.1 14.0 0.8 0.578

合并后 bf16 权重的采样参数扫描(vLLM,前 300 条测试提示,T=0.8, top_p=0.9):

repetition_penalty 长度中位数 平均长度 复读率 % 微信表情 % distinct-2
1.0 9 11.5 4.0 6.3 0.654
1.05 11 13.4 2.0 12.3 0.641
1.2 17 20.0 0.3 52.0 0.578
1.05 + 公开提示词 10 12.6 2.3 13.3 0.672
真实回复 11 13.3 1.3 12.3 0.684

repetition_penalty=1.05 时,长度和表情比例几乎与真实回复一致。vLLM 的重复惩罚同样作用于提示词 token,所以数值越高,回复越长、表情越多。

记忆探测: 在 909 条测试生成中,长度不少于 8 个字的回复里,逐字出现在训练集中的比例为 0%。这只是一个很弱的下界检查,不代表隐私保证(见下文)。

局限、偏差与隐私

  • 幻觉: 价格、电量、规则、「谁说过什么」都是按风格生成的,不是检索出来的。它会一本正经地胡说。
  • 记忆风险: 脱敏只基于正则。权重中仍可能残留原群聊的片段,例如机构名、地名、事件或观点。如果发现能识别真实个人或机构的输出,请在 Discussion 中反馈,作者会加强过滤后重新训练。
  • 人设单一: 只有一种融合的口吻,回复偏短,不擅长长篇解释。基座的通用能力大体保留,但没有重新评测。
  • 语言: 仅支持中文。英文提问最多得到中文群聊式回复。
  • 量化失配: LoRA 是在 8-bit 权重上学到的,却合并进 bf16。上面的 vLLM 统计说明行为基本一致,但 logits 与训练时的模型并不完全相同。

不适用场景: 交易或投资建议、冒充可识别的个人、骚扰,以及把生成内容当作真实市场参与者的发言。

后续训练路线

  1. 在 CUDA 上做 bf16 全层训练: 借助融合的 Gated DeltaNet 算子(如 flash-linear-attention),让 LoRA 覆盖全部 32 层、rank 提到 32–64,或者做 DoRA/rsLoRA 对比实验。这样也能消除 8-bit→bf16 失配。
  2. 多轮、区分说话人的 SFT: 用 8–16 轮滑动窗口替代单条配对。加入匿名说话人标记(如 <spk_07>),让同一个模型扮演多个一致的群友;再加入时间间隔 token,让模型学会什么时候不回复。
  3. 重训前加强去标识化: 用中文 NER(人名、机构、地名)做一致的假名映射,用 MinHash 去近似重复,可选 DP-SGD 或去重感知采样来降低被提取的风险。评估时用 canary 注入和提取攻击,而不只看逐字匹配率。
  4. 偏好优化: 用 DPO、KTO 或 ORPO,把真实回复作为 chosen,基座回复或过长回复作为 rejected。目标是剩下的几类问题:复读提示、滥用表情、「AI 助手腔」外泄。
  5. 知识增强: 接入公开的市场规则、出清价格和公告做检索增强,再加上工具调用,让人设不只是「像」,还能「对」。风格和事实分别由不同组件负责。
  6. 更好的评测: 人工盲测(真实与生成两两对比)、LLM-as-judge 打分、风格分类器,以及按时间切分(用前几个月训练、后几个月测试)来衡量漂移。
  7. 更新 MTP 头: 在适配后的模型上微调或蒸馏多 token 预测头,恢复投机解码的接受率。
  8. 更小的发布版本: 发布 AWQ/GPTQ int4、GGUF、MLX 4/8-bit 版本,方便消费级显卡和 Mac 使用;同时单独发布 LoRA,便于热切换适配器。

关于作者

由 @yzhang318 独立完成,包括数据管线、训练、权重合并、评测和部署。如果你也在做电力/能源市场方向的大模型,或者低成本人设微调,欢迎在 Discussion 里交流。

Downloads last month
22
Safetensors
Model size
10B params
Tensor type
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yzhang318/etrader-niuma-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(982)
this model
Adapters
1 model