Instructions to use KokAI-lab/Tri-7B-Insurance-v22 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use KokAI-lab/Tri-7B-Insurance-v22 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf KokAI-lab/Tri-7B-Insurance-v22:F16 # Run inference directly in the terminal: llama cli -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf KokAI-lab/Tri-7B-Insurance-v22:F16 # Run inference directly in the terminal: llama cli -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf KokAI-lab/Tri-7B-Insurance-v22:F16 # Run inference directly in the terminal: ./llama-cli -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf KokAI-lab/Tri-7B-Insurance-v22:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Use Docker
docker model run hf.co/KokAI-lab/Tri-7B-Insurance-v22:F16
- LM Studio
- Jan
- vLLM
How to use KokAI-lab/Tri-7B-Insurance-v22 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KokAI-lab/Tri-7B-Insurance-v22" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KokAI-lab/Tri-7B-Insurance-v22", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/KokAI-lab/Tri-7B-Insurance-v22:F16
- Ollama
How to use KokAI-lab/Tri-7B-Insurance-v22 with Ollama:
ollama run hf.co/KokAI-lab/Tri-7B-Insurance-v22:F16
- Unsloth Desktop
- Pi
How to use KokAI-lab/Tri-7B-Insurance-v22 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "KokAI-lab/Tri-7B-Insurance-v22:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use KokAI-lab/Tri-7B-Insurance-v22 with Docker Model Runner:
docker model run hf.co/KokAI-lab/Tri-7B-Insurance-v22:F16
- Lemonade
How to use KokAI-lab/Tri-7B-Insurance-v22 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull KokAI-lab/Tri-7B-Insurance-v22:F16
Run and chat with the model
lemonade run user.Tri-7B-Insurance-v22-F16
List all available models
lemonade list
- Hermes Agent
How to use KokAI-lab/Tri-7B-Insurance-v22 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default KokAI-lab/Tri-7B-Insurance-v22:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use KokAI-lab/Tri-7B-Insurance-v22 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KokAI-lab/Tri-7B-Insurance-v22:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "KokAI-lab/Tri-7B-Insurance-v22:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Tri-7B Insurance v2.2
Tri-7B Insurance v2.2 is a Korean insurance customer-support model fine-tuned from
trillionlabs/Tri-7B.
The model was fine-tuned with QLoRA using Korean insurance customer-service data, with an emphasis on:
- Insurance terminology
- Group insurance customer support
- Claims-related questions
- Coverage explanations
- Safer responses when policy or enrollment information is unavailable
- Avoiding unsupported conclusions about claim eligibility or coverage
The model is intended to be used together with RAG (Retrieval-Augmented Generation) for production insurance applications.
English
Model Overview
| Item | Description |
|---|---|
| Base model | trillionlabs/Tri-7B |
| Fine-tuning method | QLoRA / LoRA |
| Domain | Insurance / Group Insurance / Claims Support |
| Primary language | Korean |
| Model type | Causal Language Model |
| Quantization | GGUF F16 / Q4_K_M |
| Recommended deployment | RAG-based insurance assistant |
| License | Apache-2.0 |
Training Data
The model was fine-tuned using the insurance subset of a Korean financial customer-service dataset.
Initial insurance dataset:
- Training: 24,000 QA samples
- Validation: 3,000 QA samples
After removing automobile-insurance-focused samples and preparing the dataset for the intended group-insurance use case:
- Training: 16,084 samples
- Validation: 1,948 samples
- Coverage / claims-related samples: approximately 70.5% of the training dataset
The training format was converted into a conversational structure:
System
↓
Insurance assistant instructions
User
↓
Instruction + customer question
Assistant
↓
Reference answer
Fine-tuning
The primary QLoRA configuration included:
Base Model trillionlabs/Tri-7B
Quantization 4-bit NF4
LoRA Rank 16
LoRA Alpha 32
LoRA Dropout 0.05
Target Modules
- q_proj
- k_proj
- v_proj
- o_proj
Max Sequence Length 512
Batch Size 1
Gradient Accumulation 8
Learning Rate 1e-4
Primary Epochs 1
More than 99.9% of the prepared training samples fit within 512 tokens.
Training Iterations
v1
The first model was trained directly on the prepared insurance consultation data.
It learned insurance terminology and consultation patterns, but it sometimes generated unsupported statements about:
- Coverage eligibility
- Claim payment eligibility
- Enrolled benefits
- Customer-specific policy information
v2
The dataset was modified to reduce unsupported conclusions when policy information was not available.
The model was trained to verify:
- Actual coverage enrollment
- Policy terms
- Insurance period
- Payment conditions
- Exclusions
before making conclusions about claim eligibility.
v2.1 / v2.2
Additional calibration was performed for higher-risk insurance questions involving:
- Required claim documents
- Submission methods
- Claim processing time
- Field investigations
- Claim amounts
- Internal insurer procedures
These iterations improved conservative behavior, but they do not guarantee factual accuracy.
For this reason, production use should rely on RAG and deterministic safeguards.
Validation
An insurance-safety-adjusted validation set was also evaluated.
Example v2 results:
Training Loss: 1.7269
Original Validation Loss: 1.9097
Safe Validation Loss: 1.8079
Safe Validation
Mean Token Accuracy: 0.5542
These metrics measure language-model behavior on the prepared dataset and should not be interpreted as claim-decision accuracy.
Recommended Architecture
This model is designed to be used as part of the following architecture:
User Question
↓
Risk / Intent Detection
↓
RAG Retrieval
↓
Customer Policy
Insurance Terms
Coverage Information
Claims Guide
↓
Tri-7B Insurance v2.2
↓
Grounded Response
The fine-tuned model should mainly provide:
- Insurance-domain language understanding
- Customer-service response style
- Explanation of insurance terminology
- Structured insurance consultation behavior
RAG should provide:
- Actual policy terms
- Customer-specific coverage
- Benefit amounts
- Required documents
- Claim procedures
- Submission channels
- Current insurer rules
Important Limitations
This model must not be used as the sole source for insurance claim decisions.
The model may still generate incorrect or unsupported information, especially regarding:
- Whether a claim is payable
- Whether a specific benefit is enrolled
- Required claim documents
- Submission methods
- Claim processing timelines
- Benefit amounts
- Insurer-specific procedures
For production systems, use:
Fine-tuned LLM
+
RAG
+
Policy / enrollment data
+
Output guardrails
If reliable source material is unavailable, the application should return a safe fallback response rather than allowing the model to guess.
Available Formats
This repository may contain:
Merged Hugging Face model
GGUF F16
GGUF Q4_K_M
For local inference with LM Studio or llama.cpp, the Q4_K_M GGUF version is recommended for a good balance between model quality and memory usage.
Disclaimer
This model is intended for research, prototyping, and insurance customer-support assistance.
It is not an automated claims adjudication system and should not independently determine coverage, liability, claim eligibility, or payment amounts.
Final insurance decisions must be based on the actual insurance contract, policy wording, enrollment information, applicable insurer rules, and qualified human review where required.
한국어
모델 개요
Tri-7B Insurance v2.2는
trillionlabs/Tri-7B를 기반으로
한국어 보험상담 업무에 맞게 QLoRA 파인튜닝한 모델입니다.
주요 학습 목적은 다음과 같습니다.
- 보험 용어 이해 및 설명
- 기업 단체보험 상담
- 보험금 청구 관련 질의응답
- 보장내용 설명
- 가입내역이나 약관이 없는 경우 성급하게 지급 여부를 단정하지 않는 상담 방식
- 고객이 이해하기 쉬운 보험상담 답변 생성
실제 보험 서비스에서는 RAG와 함께 사용하는 것을 전제로 설계했습니다.
기본 정보
| 항목 | 내용 |
|---|---|
| Base Model | trillionlabs/Tri-7B |
| Fine-tuning | QLoRA / LoRA |
| 분야 | 보험 / 단체보험 / 보험금 청구 |
| 주요 언어 | 한국어 |
| 모델 종류 | Causal Language Model |
| 제공 형식 | Hugging Face / GGUF F16 / GGUF Q4_K_M |
| 권장 사용 방식 | RAG 기반 보험상담 AI |
| 라이선스 | Apache-2.0 |
학습 데이터
한국어 금융 고객상담 데이터 중 보험 분야 데이터를 활용했습니다.
최초 보험 데이터:
Training 24,000건
Validation 3,000건
자동차보험 중심 데이터를 제거하고 기업 단체보험·보장·보험금 청구 업무에 맞게 정리한 데이터는:
Training 16,084건
Validation 1,948건
보장/청구 핵심 데이터 비율
약 70.5%
입니다.
학습 데이터는 다음과 같은 대화 형식으로 구성했습니다.
System
↓
보험상담 AI 역할 및 규칙
User
↓
지시문 + 고객 질문
Assistant
↓
보험상담 답변
주요 QLoRA 설정
Base Model trillionlabs/Tri-7B
QLoRA NF4 4bit
LoRA Rank 16
LoRA Alpha 32
LoRA Dropout 0.05
Target Modules
- q_proj
- k_proj
- v_proj
- o_proj
Max Length 512
Batch Size 1
Gradient Accumulation 8
Learning Rate 1e-4
Primary Epoch 1
전처리된 Training 데이터의 99.9% 이상이 512 token 이내에 포함되었습니다.
모델 개선 과정
V1
보험 상담 데이터 자체를 이용해 최초 파인튜닝을 수행했습니다.
보험 용어와 상담 문장은 잘 학습했지만 다음과 같은 문제가 발견되었습니다.
실제 가입내용을 모르는데
"현재 가입되어 있습니다"
약관을 모르는데
"보험금이 지급됩니다"
계약정보가 없는데
"수술비 담보가 없습니다"
와 같이 존재하지 않는 고객 계약정보를 만들어내는 문제가 있었습니다.
V2
보장 여부나 보험금 지급 가능성을 묻는 질문에서 실제 가입내역이나 약관이 없는 경우 단정하지 않도록 Training 데이터를 수정했습니다.
모델이 다음 사항을 먼저 확인하도록 학습 방향을 조정했습니다.
가입 담보
보험기간
약관상 지급사유
면책사항
사고 및 진단 내용
V2.1 / V2.2
추가적으로 다음 질문 유형에 대한 보정학습을 진행했습니다.
보험금 청구서류
서류 제출방법
보험금 처리기간
현장조사 여부
지급금액
보험사 내부 처리절차
이 과정을 통해 근거 없는 단정은 감소했지만 Fine-tuning만으로 환각을 완전히 제거할 수는 없다는 점도 확인했습니다.
따라서 최종 서비스에서는 RAG와 Guard를 함께 사용하는 것을 권장합니다.
Validation 결과
V2 모델의 주요 결과:
Train Loss
1.7269
기존 Validation Loss
1.9097
안전성 수정 Validation Loss
1.8079
안전 Validation
Mean Token Accuracy
0.5542
위 수치는 학습 데이터에 대한 Language Model 지표이며 보험금 지급 판단 정확도를 의미하지 않습니다.
권장 서비스 구조
실제 서비스에서는 다음 구조를 권장합니다.
고객 질문
↓
질문 유형 / 위험도 확인
↓
RAG 검색
↓
고객사 가입내역
보험약관
보장내역
보험금 청구안내
↓
Tri-7B Insurance v2.2
↓
근거 기반 답변
Fine-tuning 모델은 주로:
보험업무 언어 이해
보험상담 문체
보험 용어 설명
보험 질문 의도 이해
상담 답변 구성
을 담당합니다.
RAG는:
실제 가입 담보
보험가입금액
고객사별 보장내용
보험약관
청구서류
보험금 접수방법
보험사별 최신 기준
과 같은 사실정보를 담당하도록 설계하는 것이 좋습니다.
중요한 제한사항
본 모델의 출력만으로 다음 사항을 결정해서는 안 됩니다.
- 보험금 지급 여부
- 실제 담보 가입 여부
- 보험금 지급금액
- 면책 여부
- 보험금 청구 필수서류
- 보험금 접수방법
- 보험금 지급기간
- 보험사 내부 처리절차
모델은 이러한 내용을 잘못 생성할 수 있습니다.
실제 보험 서비스에서는 다음 구조를 권장합니다.
Fine-tuned LLM
+
RAG
+
고객사 가입정보
+
보험약관
+
Guardrail
관련 근거자료가 검색되지 않은 경우에는 모델이 임의로 답변하도록 하지 않고, 안전한 안내문을 반환하는 방식이 적절합니다.
제공 모델 형식
Repository에는 다음 형식을 제공할 수 있습니다.
1. Hugging Face Merged Model
2. GGUF F16
3. GGUF Q4_K_M
LM Studio 또는 llama.cpp에서 로컬 실행하는 경우에는 Q4_K_M 버전을 권장합니다.
면책 및 사용상 주의
본 모델은 보험 AI 연구, 프로토타입 및 보험상담 업무 지원을 목적으로 제작되었습니다.
본 모델은 자동 보험금 지급심사 모델이 아닙니다.
모델의 출력만으로 보장 여부, 보험금 지급 여부, 책임 여부 또는 지급금액을 결정해서는 안 됩니다.
최종 판단은 반드시 실제 보험계약, 보험약관, 가입내역, 보험사 기준 및 필요한 경우 담당자의 검토를 기반으로 이루어져야 합니다.
- Downloads last month
- 54
Model tree for KokAI-lab/Tri-7B-Insurance-v22
Base model
trillionlabs/Tri-7B