Instructions to use gxydResearch/lanlao with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use gxydResearch/lanlao with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="gxydResearch/lanlao")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("gxydResearch/lanlao", device_map="auto") - Notebooks
- Google Colab
- Kaggle
lanlao
lanlao 是广西移动人工智能专班自研的 Jev 类单次前向决策模型,面向工单路由、风险判别、内容审核、评审打分等结构化决策场景。模型接收一段陈述(state)与有界判据(rubric),经单次前向传播即输出选项选择、置信度与完整概率分布,不经过自回归解码,不生成自然语言。
本仓库包含 lanlao 的 LoRA 适配器权重(rlcd_score43k_v5.pt,约 25 MB)及完整服务代码,格式兼容 HuggingFace Transformers / peft。基座模型为 Qwen/Qwen3.5-4B,需单独下载。lanlao 的请求/响应协议对齐 TypeSafe AI 的 Jev 模型,Jev 客户端可直接切换使用。如需高并发批量推理,可将 LoRA 合并为全量权重后用 vLLM 加载(见下文部署章节)。
模型架构
| 项目 | 说明 |
|---|---|
| 基座 | Qwen3.5-4B,冻结,bf16 |
| 适配层 | LoRA,rank 32,alpha 64,target_modules = q/k/v/o_proj,dropout 0,参数量约 25 MB |
| 决策头 | 末位 token logits → 限制到 A–P 槽位(最多 16 选项)→ softmax = 选项概率分布;无第二头,无自回归 |
| thinking | enable_thinking=False(Qwen3.5 chat template 要求) |
| 题型 | choice(多选一)/ noul(是非,输出 p_yes)/ score(序数等级,输出期望等级 ∈ [0,1]) |
| 上下文 | 继承 Qwen3.5-4B 基座上下文长度 |
| 推理硬件 | NVIDIA GPU,建议 ≥16 GB 显存(bf16 约 9 GB,12 GB 可运行) |
三种题型在同一请求内可并发,共享同一 state 前缀——即一次前向可同时输出多个决策结果。choice 的 criteria 为无序 dict(2–16 选项),noul 无 criteria,score 的 criteria 为有序等级列表。
评测结果
以下为开发方自测结果。authored144 与 perturbations108 为冻结测试集,实战双语测试集为内置 en50 + zh50(100 行,含金标签)。对照列为官方 Jev 的公开指标。
决策准确率与校准
| 评测集 | lanlao | Jev 官方 |
|---|---|---|
| authored144 | 0.965 / ECE 0.026 | 0.979 / 0.048 |
| perturbations108 | 0.972 / 0.022 | 1.000 / 0.005 |
| 实战双语测试集(100 行) | 0.950 / 0.041 / 79 ms(4090) | 0.970 / 0.031 / 673 ms |
分题型(实战双语 100 行)
| 题型 | lanlao | Jev 官方 |
|---|---|---|
| choice | 0.983(与 Jev 预测逐行一致) | 0.983 |
| noul | 0.900 | 0.900 |
| score | 0.900(v1.0 为 0.850,序数补训后 +0.05;留出集 0.955) | 1.000 |
部署
环境要求
- x86_64 Linux + NVIDIA GPU(建议 ≥16 GB 显存)
- Python 3.10+,CUDA 12.x
- 依赖:
torch>=2.6、transformers>=5.0(需原生Qwen3_5ForCausalLM类,低于 5.0 不含)、peft>=0.15、numpy>=1.24
安装
# 解压
tar xzf lanlao-v1.1.0.tar.gz && cd semif-systemone
# 建虚拟环境并安装依赖(PyTorch 请按本机 CUDA 版本从官方索引安装)
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# 下载基座 Qwen3.5-4B(二选一)
# ModelScope(国内推荐):
pip install modelscope
modelscope download --model Qwen/Qwen3.5-4B --local_dir ./Qwen3.5-4B
# HuggingFace:
huggingface-cli download Qwen/Qwen3.5-4B --local-dir ./Qwen3.5-4B
基座必须为 Qwen3.5-4B,LoRA 权重与该基座的 q/k/v/o_proj 形状对应,换用其他 Qwen 版本会 shape 不匹配。transformers 需 ≥5.0,否则缺少原生 Qwen3_5ForCausalLM 类。
快速验证
python systemone_server.py --model ./Qwen3.5-4B --mode lora \
--lora-checkpoint ./rlcd_score43k_v5.pt --selftest
运行 5 个 demo 场景(工单路由 / 财经风险 / 钓鱼邮件 / 代码评审 / 评审打分),每场景输出答案 JSON。首次运行需 merge LoRA 并加载模型,约 1–2 分钟。
启动 API 服务
方式 A:LoRA 加载(默认,占盘小)
CUDA_VISIBLE_DEVICES=0 python systemone_server.py \
--model ./Qwen3.5-4B \
--mode lora --lora-checkpoint ./rlcd_score43k_v5.pt \
--port 8100
方式 B:全量合并权重(可直接用 vLLM 加载,无需 peft)
先合并(一次性,约 3 分钟,输出标准 HF 模型目录约 8 GB):
python merge_export.py --model ./Qwen3.5-4B \
--lora-checkpoint ./rlcd_score43k_v5.pt --out ./Qwen3.5-4B-lanlao
之后启动(--mode direct,不加载 LoRA):
CUDA_VISIBLE_DEVICES=0 python systemone_server.py \
--model ./Qwen3.5-4B-lanlao \
--mode direct --port 8100
两种方式实测 100 行预测一致、置信度差 0.0。合并后的 Qwen3.5-4B-lanlao 是标准 HF 目录,可直接 vllm serve ./Qwen3.5-4B-lanlao:用 completions 接口 max_tokens=1 + logprobs=20,在返回的 top logprobs 中取 A–P 槽位做 softmax 即可复现本模型输出。chat prompt 需带 enable_thinking=False。
请求方式
curl -s http://127.0.0.1:8100/v1/systemone -H "Content-Type: application/json" -d '{
"state": "I was charged twice for my monthly subscription. My card shows two identical charges.",
"model": "jev-latest",
"questions": {
"department": {"type": "choice", "instructions": "Which team should handle this",
"criteria": {"billing": "Payment or subscription issues", "technical": "Bugs or integration problems", "sales": "Pricing questions"}},
"is_urgent": {"type": "noul", "instructions": "The message conveys urgency"},
"frustration": {"type": "score", "instructions": "How frustrated the customer appears",
"criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry"]}
}
}'
响应:
{"answers": {
"department": {"choice": "billing", "confidence": 0.94, "probabilities": {"billing": 0.94, "technical": 0.03, "sales": 0.03}},
"is_urgent": {"noul": 0.82, "confidence": 0.82, "probabilities": {"yes": 0.82, "no": 0.18}},
"frustration": {"score": 0.62, "confidence": 0.71, "probabilities": {"Calm, just stating facts": 0.12, "Frustrated but civil": 0.71, "Very angry": 0.17}}
}, "model": "semif-systemone-lora", "usage": {"input_tokens": 87, "output_tokens": 0}, "latency_ms": 82.1}
题型说明:
choice:criteria 为无序 dict{id: 描述},2–16 个选项;2–4 选项延迟约 70–90 msnoul:无 criteria,instructions 即是非判据,返回 p_yesscore:criteria 为有序等级列表[低, ..., 高],score = 期望等级 ∈ [0,1]
部署后测试
CUDA_VISIBLE_DEVICES=0 python realworld_compare.py --side local \
--model ./Qwen3.5-4B --lora-checkpoint ./rlcd_score43k_v5.pt
运行内置双语实战集(en50 + zh50,含金标签),预期结果:ALL acc 0.950 / ECE 0.041;choice 0.983 / noul 0.900 / score 0.900。结果写入 results/realworld_local.jsonl。
训练方法
lanlao 在 Qwen3.5-4B 冻结基座上训练 LoRA 适配层,决策头实现为单次前向 + 末位字母槽位 softmax。训练分两阶段:
训练数据:bi35k = 10k 条英文公开 NLI + 5k 条中文公开 CMNLI + 20k 条 DeepSeek 教师合成数据(经 solver 盲解校验),合成占比 57%。中文合成数据沿用英文的 option id 与 family 模板,保证跨语言结构一致。
阶段 1(SFT):CE(letter-slot) + λ=1.0 × |conf−correct| 联合损失,2 个 epoch。此阶段将校准训练进概率分布,ECE 达到 0.026。
阶段 2(RLCD 强化):reward = conf·(2·correct−1),错误样本不得分;熵正则 β=0.01,advantage 严格 detach,单遍 lr 2e-5,从阶段 1 热启动。perturbations 准确率从 0.965 提升至 0.972。
文件清单
semif-systemone/
├── README.md 原始部署手册
├── requirements.txt Python 依赖
├── rlcd_score43k_v5.pt LoRA 权重(25 MB,最终模型)
├── systemone_server.py /v1/systemone API 服务
├── merge_export.py LoRA → 全量权重合并导出
├── realworld_compare.py 部署后评测脚本
├── toy_dual_head.py 模型加载(被 server 引用)
├── src/semif_phase1/ prompt 协议与工具(被 server 引用)
└── benchmarks/data/ 内置双语实战测试集(en50 + zh50)
License
- lanlao 适配器权重与服务代码:Apache-2.0
- 基座 Qwen3.5-4B 保留其自身许可与使用条款,使用本模型须同时遵守基座许可
发布机构:广西移动人工智能专班 版本:lanlao v1.1(2026-09-25)