👁‍🗨 Yeev:网络安全结构化决策模型

Yeev 是一个面向网络安全场景的 Typed-Decision Model。它基于 Laya 微调,输入一段安全状态和一组结构化问题,直接返回选项分布、置信度和结构化答案。

模型名称参考了 Jev 的命名与结构化决策理念;Yeev 不是 Jev 的权重复制,也不是自回归文本生成模型。Yeev 更适合需要稳定标签、概率分布和批量决策的安全工作流。

Yeev is a Laya-based, non-autoregressive typed-decision model for cybersecurity question answering and security-incident triage. It returns structured choices, ordinal scores, and yes/no probabilities in one forward pass.

🔍 模型概览 | Model Card

  • 模型名称: sds-ai/Yeev
  • 项目地址: https://huggingface.co/sds-ai/Yeev
  • 底座模型: convaiinnovations/laya
  • 推理库: laya==0.3.3
  • 模型格式: Laya checkpoint,包含 model.safetensors、encoder、tokenizer 和 rl_agent_config.json
  • 输入: state + typed questions
  • 输出: choice、score、noul 三类结构化答案
  • 语言: 中文、英文;实际效果取决于题目和领域数据
  • 模型权重 SHA-256: 746c0f8b9333fe686ee410139603afe6984420dc1b135e475697a825c4695c9b

Yeev 不生成长篇自然语言答案,而是直接在请求时定义答案空间:

  • choice:在多个命名选项中选择一个;
  • score:在有序等级中给出期望分数和分布;
  • noul:返回 true 的概率,适合判断题、是否存在风险等二元任务。

🎯 适用场景 | Use Cases

Yeev 的应急响应定位是 辅助告警分诊与响应优先级判断:根据告警证据,结构化评估是否可能是真实事件、建议调查/监控/遏制/关闭分流,并给出严重性和紧急度分布。它提供决策支持,不会执行隔离、封禁或删除操作。

🧯 应急响应样例:覆盖 12 类安全告警

以下 Gold 是数据记录提供的标签概率分布中概率最高的类别;严重性、紧急度同样列出最高概率等级。概率用于表示数据中的软标签,不代表真实世界发生概率。

数据样例 事件与响应关注点 Gold 处置路径 Gold 严重性 / 紧急度 样本数
#1 dormant_account_use 休眠 180 天的服务账号无 MFA 成功登录 调查 investigate · 0.583 低 Low · 0.433 / 常规 Routine · 0.517 195
#2 brute_force 95 次登录失败后出现一次成功登录 调查 · 0.650 高 High · 0.683 / 紧急 Critical · 0.733 182
#3 disabled_logging 主机 auditd 停止,审计日志中断 45 分钟 调查 · 0.633 高 · 0.467 / 升高 Elevated · 0.417 191
#5 mass_download 未识别来源在 60 秒内读取 450 个文件 调查 · 0.617 高 · 0.517 / 紧急 · 0.567 191
#6 data_egress CI runner 向外部目的地传输 12 GB 数据,来源网络未识别 调查 · 0.600 中 Moderate · 0.467 / 升高 · 0.437 186
#7 suspicious_process 服务器执行未知二进制 svc_monitor.exe 调查 · 0.650 高 · 0.450 / 紧急 · 0.583 189
#9 privilege_escalation CI runner 在变更窗口外成功提升权限 遏制 contain · 0.633 中 · 0.500 / 紧急 · 0.684 192
#17 malware_signature 终端代理命中 Malware.Variant.Alpha 特征 调查 · 0.550 中 · 0.517 / 升高 · 0.517 185
#19 token_reuse 特权承包商账号的会话令牌从第二个地址出现 调查 · 0.690 高 · 0.593 / 紧急 · 0.450 186
#20 new_country_login 首次从陌生国家登录,但 MFA 成功且来源在 allowlist 监控 monitor · 0.517 低 · 0.400 / 常规 · 0.483 203
#22 impossible_travel 共享邮箱短时间内从远距离地点登录,来源均在 allowlist 关闭为良性 close_benign · 0.600 高 · 0.500 / 常规 · 0.383 183
#25 config_drift 安全组新增 0.0.0.0/0 到 TCP/22 的入站规则 调查 · 0.600 高 · 0.433 / 紧急 · 0.450 187

处置标签对应的是模型的结构化决策空间:monitor 用于继续观察,investigate 用于升级调查,contain 表示建议进入经授权的遏制流程,close_benign 表示可进入良性关闭流程。它们是数据集中的 gold 标签示例,不表示已执行的响应动作;生产流程应由分析师核实证据并授权。

📘 安全知识问答与基准评测

认证题库和知识评测可以用 choice 输出稳定选项,支持题库回归、分类统计与错题分析。例如 CS-Eval 样例:

题目: What do researchers continue to explore to enhance deep learning's impact on cybersecurity?
选项片段: B) Hybrid approaches and novel architectures
数据集 Gold: B

Gold 是数据集标注,多选题在训练数据中会被规范化为多个 noul 判断;这便于统一评测,但不等同于生成式模型输出完整解释。

🚀 快速开始 | Quickstart

建议使用 Python 3.10+、CUDA 版 PyTorch 和 GPU 环境。该 checkpoint 按 laya==0.3.3 训练和验证,建议固定运行时版本:

python -m pip install "laya==0.3.3"

💻 真实事件样例结构化研判

如安全事件告警:休眠 180 天、未启用 MFA 的服务账号成功登录。

import json
from pathlib import Path
import laya

sample = """
{
  "state": {
    "alert": {
      "description": "a long-unused account became active",
      "evidence": "The unprivileged service account `svc_task_alpha` initiated a session from IP `198.51.100.24` following 180 days of zero activity. Authentication was successful without MFA using credentials that were last rotated six months ago.",
      "rule": "dormant_account_use"
    },
    "context": {
      "asset_criticality": "low",
      "change_window_active": false,
      "source_on_allowlist": false
    },
    "history": {
      "credential_rotation_days_ago": 180,
      "distinct_countries_30d": 4,
      "logins_30d": 287,
      "prior_alerts_90d": 8
    },
    "principal": {
      "mfa_enrolled": false,
      "name": "usr-support-752",
      "privileges": [
        "repo:read"
      ],
      "type": "service_account"
    }
  },
  "questions": {
    "credential_compromise": {
      "instructions": "The evidence indicates a credential or account has been compromised.",
      "type": "noul"
    },
    "disposition": {
      "criteria": {
        "close_benign": "Expected, explainable activity; close without analyst time.",
        "contain": "Contain the host or account immediately; do not wait for triage.",
        "investigate": "Warrants an analyst opening an investigation.",
        "monitor": "Not clearly malicious, but worth watching for recurrence."
      },
      "instructions": "How should this security alert be dispositioned?",
      "type": "choice"
    },
    "severity": {
      "criteria": [
        "Negligible: no access to anything sensitive.",
        "Low: limited access, easily reversed.",
        "Moderate: access to internal systems or non-public data.",
        "High: access to production, secrets or customer data.",
        "Critical: active compromise of crown-jewel systems."
      ],
      "instructions": "How severe is the potential impact if this alert is real?",
      "type": "score"
    },
    "true_positive": {
      "criteria": {
        "false": "Benign activity, a misconfiguration, or a known false positive.",
        "true": "The underlying behaviour is malicious or unauthorised."
      },
      "instructions": "This alert reflects genuinely malicious or unauthorised activity.",
      "type": "noul"
    },
    "urgency": {
      "criteria": [
        "No time pressure; can wait indefinitely.",
        "Routine; handle within the normal queue.",
        "Elevated; should be handled within the same week.",
        "Critical; requires action within the same day."
      ],
      "instructions": "How quickly must a responder act on this alert?",
      "type": "score"
    }
  }
}
"""

model_dir = "/path/to/Yeev"  # 替换为包含 model.safetensors 的 checkpoint 目录
agent = laya.load(model_dir, device="cuda")
result = agent.predict(sample["state"], sample["questions"])

print("事件规则:", sample["state"]["alert"]["rule"])
print("事件摘要:", sample["state"]["alert"]["description"])
print("Yeev 回复:")
print(json.dumps(result["answers"], ensure_ascii=False, indent=2))
print("Gold(仅作标签参考):")
print(json.dumps(sample["gold"], ensure_ascii=False, indent=2))

🗣️ 模型回复实例

{
  "credential_compromise": {
    "noul": 0.5244,
    "confidence": 0.5244
  },
  "disposition": {
    "choice": "investigate",
    "probabilities": {
      "close_benign": 0.0725,
      "contain": 0.048,
      "investigate": 0.5835,
      "monitor": 0.2959
    },
    "confidence": 0.2709
  },
  "severity": {
    "score": 1.1511,
    "probabilities": {
      "0": 0.2075,
      "1": 0.5032,
      "2": 0.23,
      "3": 0.0491,
      "4": 0.0101
    },
    "confidence": 0.2517
  },
  "true_positive": {
    "noul": 0.5377,
    "confidence": 0.5377
  },
  "urgency": {
    "score": 1.3916,
    "probabilities": {
      "0": 0.098,
      "1": 0.5022,
      "2": 0.3101,
      "3": 0.0897
    },
    "confidence": 0.1683
  }
}

本例中,模型选择 investigate;严重性分布以 1(Low)最高,紧急度分布以 1(Routine)最高。对应 Gold 的 disposition=investigate 概率为 0.583333。

📊 评测结果 | Evaluation

Benchmark Cases Decisions Hard accuracy Case exact match
CISSP 3,741 3,741 99.09% 99.09%
CISP 1,298 1,298 99.31% 99.31%
CS-Eval 4,368 5,179 98.47% 98.21%
Security incident 2,270 11,350 98.98% 95.46%

🏅 CS-Eval 评审结果 | CS-Eval Benchmark

评审维度 得分
综合得分 89.47
系统安全及软件安全基础 93.00
访问控制与身份管理 87.16
加密技术与密钥管理 91.24
基础设施安全 89.35
AI 与网络安全 93.07
漏洞管理与渗透测试 90.61
威胁检测与预防 89.48
数据安全和隐私保护 88.10
供应链安全 92.69
安全架构设计 90.73
业务连续性与应急响应恢复 82.00
中文任务 89.31
英文任务 91.78

🔎 与 SecGPT-14B 的公开结果对照

参考 SecGPT 的 evaltion.py 和 SecGPT-14B 模型卡:

指标 Yeev SecGPT-14B
CISSP 99.09% 78.84%
CS-EVAL 98.21% case exact 88.60%
CISP 99.31% 未提供
Security incident 95.46% case exact 未提供

⚙️ 技术说明 | Technical Notes

  • Yeev 使用非自回归 typed-decision 推理,不需要从自然语言回答中解析 A/B/C/D。
  • 每个调用中的问题可以拥有不同的选项集合;选项描述和问题说明会参与编码。
  • choice 返回选项概率分布;score 返回有序分布和期望分数;noul 返回 true 概率。
  • 当前 checkpoint 的训练配置为 max_len=1024、head_max_len=256;长输入和大量选项应先在目标环境做上下文长度验证。
  • 生产使用前应在独立验证集上重新校准概率,并设置人工升级、审计日志和权限边界。

⚠️ 限制与安全边界 | Limitations

  1. README 中的 benchmark 数字不是独立测试集或生产质量保证。
  2. 题库标签可能存在错误、版本差异和过时知识;应核验高风险结论。
  3. Security incident 数据包含场景规则监督,输出概率不等同于真实世界事件发生概率。
  4. 模型不能替代安全分析师,也不应未经授权自动执行隔离、删除、封禁、扫描或外传数据等操作。
  5. 使用者需要同时遵守底座模型、题库和部署环境的许可证、隐私与合规要求。

🔗 参考项目 | References

Yeev is intended for research, evaluation, and human-in-the-loop security workflows. Validate labels and calibrate probabilities on an independent validation set before deployment.

Downloads last month
13
Safetensors
Model size
0.4B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sds-ai/Yeev

Finetuned
(165)
this model