You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Free for individuals for personal, non-commercial use. Any use by a business (companies, sole proprietors, hospitals, clinics, pharmacies and their staff) requires a prior written agreement with SPINAI Co., Ltd. (info@spinai.net) before any use, including internal testing or evaluation. By requesting access you agree to the terms in LICENSE. 개인의 비상업적 사용은 무료입니다. 사업자(법인·개인사업자·병원·의원·약국 및 그 임직원)의 모든 사용은 주식회사 스피나이와 사전 협의·서면 계약이 필요합니다(내부 시험·평가 포함, 예외 없음). 신청하시면 LICENSE 의 조건에 동의하는 것으로 봅니다.

Log in or Sign Up to review the conditions and access this model content.

AISVI_S4.1.0_SPINAI

A 9B Korean/English function-calling model from SPINAI.

SPINAI 가 만든 9B 급 한국어·영어 도구 호출 모델입니다.

What changed / 무엇이 좋아졌나

"Start" below is AISVI S4.0.9, SPINAI's previous release this model was trained from. Compared with it, single-turn tool calling improved, most of all not calling a tool when none fits.

BFCL V4 single-turn (3,641 items, thinking off) Start AISVI_S4.1.0_SPINAI Δ
Overall 78.36% 81.98% +3.63pp (McNemar z = +7.45)
irrelevance 67.9% 85.4% +17.5pp
live_irrelevance 76.4% 87.2% +10.9pp
live_multiple 74.8% 75.8% +0.9pp
parallel 82.5% 77.5% −5.0pp

Against models of its size (same harness, same settings)

BFCL V4 single-turn, 3,641 items, thinking off, the weights in this repository as downloaded from the Hub.

Model Overall irrelevance live_irrelevance live_simple parallel
AISVI_S4.1.0_SPINAI 81.76% 85.4% 86.9% 80.6% 77.0%
AISVI S4.0.9 (start) 78.33% 67.9% 75.9% 81.0% 82.5%
Qwen3.5-9B 78.28% 67.5% 75.8% 81.0% 81.5%
Qwen3-8B 73.28% 60.8% 61.9% 72.9% 88.0%
Ornith-1.5-9B 72.86% 55.0% 62.3% 76.4% 86.5%

Same 200-item cross-check set as the frontier table below: AISVI_S4.1.0_SPINAI 87.5%, Qwen3-8B 85.5%, Qwen3.5-9B 83.0%, AISVI S4.0.9 83.0%, Ornith-1.5-9B 81.5%.

Against frontier models (same 200 items, same grader)

Model Accuracy Control set (20)
claude-opus-4-6-thinking 89.5% 18/20
gemini-3.1-pro-high 88.5% 17/20
AISVI_S4.1.0_SPINAI 87.5% 19/20
Starting checkpoint 83.0% 18/20

Frontier models were queried through a CLI with plain-text prompts; this model through the tools field. Items and grader are identical, but the paths differ — do not read decimal-level differences into this table.

Korean multi-turn — FunctionChat-Bench Dialog (200 turns)

Kakao's Korean tool-use benchmark, judged by gemini-3.1-pro-high.

Start (S4.0.9) AISVI_S4.1.0_SPINAI
Dialog (200 turns) 94.5% 95.0%
— tool call 94.3% 97.1%
Singlecall (100) 94.0% 96.0%
CallDecision (606) — 95.9% (call 97%, reject 97%)

Korean was already strong at the start and stays there; the gains are small and within noise. Scores were measured on the LoRA adapter before merging (see fidelity below).

Merge fidelity

The LoRA adapter was merged into bf16 weights. The merged weights downloaded from this repository score 81.76% on BFCL single-turn vs 81.98% for the unmerged adapter (−0.22pp, within the ±1pp run-to-run band), and 87.5% on the 200-item cross-check set for both.

Usage / 사용법

1. Access / 접근 — this repository is gated. Log in, click Agree and access repository above, then authenticate on your machine: hf auth login (or set HF_TOKEN). 이 저장소는 동의 후 받을 수 있습니다. 위에서 동의한 뒤 hf auth login 으로 로그인하세요.

2. Serve with vLLM (recommended)

vllm serve Spinai/AISVI_S4.1.0_SPINAI --port 8000 \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3
  • --tool-call-parser qwen3_xml is required — the model emits tool calls in XML form.
  • Trained and evaluated with thinking off — pass chat_template_kwargs: {"enable_thinking": false}.

3. Call it (OpenAI-compatible)

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
tools = [{"type": "function", "function": {
    "name": "get_weather", "description": "도시의 현재 날씨를 조회한다.",
    "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]}}}]
r = client.chat.completions.create(
    model="Spinai/AISVI_S4.1.0_SPINAI",
    messages=[{"role": "user", "content": "서울 날씨 어때?"}],
    tools=tools, temperature=0,
    extra_body={"chat_template_kwargs": {"enable_thinking": False}})
print(r.choices[0].message.tool_calls)

4. Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
name = "Spinai/AISVI_S4.1.0_SPINAI"
tok = AutoTokenizer.from_pretrained(name)
model = AutoModelForCausalLM.from_pretrained(name, dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "서울 날씨 어때?"}]
text = tok.apply_chat_template(messages, tools=tools, tokenize=False,
                               add_generation_prompt=True, enable_thinking=False)
ids = tok(text, return_tensors="pt").to(model.device)
out = model.generate(**ids, max_new_tokens=256)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))

Requires a recent transformers (5.x) and vllm with Qwen3.5 support. bf16 weights ≈ 19 GB; one 24 GB GPU is enough for short contexts.

Training

GRPO (TRL), LoRA r16 on all linear layers (merged into these weights), lr 1e-5, beta 0 (no KL anchor), 720 steps (1 epoch), 6 generations per prompt, completions capped at 512 tokens. Reward is a partial-credit score from the BFCL checker (format, name, correctness). Training items were built from public BFCL-style tasks; evaluation items were frozen by provenance and excluded from training.

License

Free for individuals for personal, non-commercial use. Any use by a business — companies, sole proprietors, hospitals, clinics, pharmacies and their staff — requires a prior written agreement with SPINAI Co., Ltd. (주식회사 스피나이), info@spinai.net — including internal testing and evaluation, without exception. Please contact us before any use. See LICENSE (the Korean text prevails).

개인의 비상업적 사용은 무료입니다. 사업자의 모든 사용은 주식회사 스피나이와 사전 계약이 필요하며, 내부 시험·평가도 예외 없이 사용 전에 반드시 협의해 주십시오(info@spinai.net).

Based on Qwen3.5-9B (Apache-2.0), fine-tuned and modified by SPINAI. The Apache-2.0 license of the base model is included as LICENSE-Qwen-Apache-2.0.

This model is not a medical device and is not intended for diagnosis or treatment decisions.

Downloads last month
4
Safetensors
Model size
9B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Spinai/AISVI_S4.1.0_SPINAI

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(931)
this model