Instructions to use IQuestLab/SAIL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IQuestLab/SAIL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="IQuestLab/SAIL") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("IQuestLab/SAIL") model = AutoModelForMultimodalLM.from_pretrained("IQuestLab/SAIL", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use IQuestLab/SAIL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IQuestLab/SAIL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IQuestLab/SAIL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/IQuestLab/SAIL
- SGLang
How to use IQuestLab/SAIL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "IQuestLab/SAIL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IQuestLab/SAIL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "IQuestLab/SAIL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IQuestLab/SAIL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use IQuestLab/SAIL with Docker Model Runner:
docker model run hf.co/IQuestLab/SAIL
Scientific Agentic Intelligence via a Science-Aware Loop
🤗 SAIL Model | 🤗 SAIL Training Data | 📄 Paper
Introduction
SAIL is a scientific agent model from IQuest Research, post-trained from Qwen3.6-35B-A3B, with 35B total and 3B active parameters. It searches and analyzes literature, writes scientific code, and uses tools to carry out research tasks that connect evidence, computation, and interpretation.
SAIL's science-aware loop turns gaps observed during scientific task execution into targeted training tasks. Development agents powered by frontier models diagnose these gaps from trajectories, then construct problems and tool-use tasks from scientific papers and code repositories. The capability targets depend on the task: evidence coverage and selection for literature retrieval, domain knowledge for scientific coding, and sustained reasoning across experiments for end-to-end research. Training combines supervised fine-tuning, specialist training, multi-teacher on-policy distillation, and agentic reinforcement learning, followed by reassessment on separate tasks.
We release SAIL and the accompanying training data to support research toward more capable AI scientists.
Performance
SAIL delivers competitive results across scientific workflows alongside substantially larger open-weight models. All models are evaluated in the same benchmark environments, with model-specific inference settings. Higher scores are better; bold values mark the best score in each column.
The parameter comparison uses the unweighted mean of 12 task scores on a 0–100 scale, excluding LitQA2-FullText and counting E2E-Bench Basic and Hard separately.
Literature understanding
| Model | Params. (B), total / active | PaperFindings | LitQA-search | ScholarQA-CS2 | LitQA2-FullText | ArxivDIGESTables |
|---|---|---|---|---|---|---|
| Ling-3.0-flash | 124 / 5.1 | 18.03 | 26.67 | 67.88 | 91.23 | 25.45 |
| DeepSeek-V4-Flash-0731 | 284 / 13 | 26.46 | 74.67 | 75.19 | 95.08 | 35.13 |
| Hy3 | 295 / 21 | 28.90 | 73.33 | 85.63 | 94.12 | 32.42 |
| MiMo-V2.5 | 310 / 15 | 16.21 | 34.67 | 60.85 | 94.23 | 27.30 |
| GLM-5.2 | 744 / 40 | 40.80 | 88.00 | 87.87 | 90.41 | 34.21 |
| Ring-2.6-1T | 1000 / 63 | 26.00 | 50.67 | 71.17 | 82.05 | 25.62 |
| MiMo-V2.5-Pro | 1020 / 42 | 28.05 | 57.33 | 75.15 | 89.06 | 31.21 |
| LongCat-2.0 | 1600 / 48 | 9.69 | 8.00 | 41.36 | 93.75 | 25.15 |
| BigBang-v1 | 35 / 3 | 28.36 | 56.00 | 54.32 | 94.67 | 28.67 |
| Apodex-1.0-mini | 35 / 3 | 23.77 | 76.00 | 74.32 | 92.00 | 28.72 |
| Nex-N2-mini | 35 / 3 | 21.78 | 33.33 | 46.77 | 94.67 | 26.14 |
| Nex-N2.5-mini | 35 / 3 | 12.05 | 22.67 | 25.61 | 91.94 | 30.94 |
| Agents-A1 | 35 / 3 | 22.73 | 52.00 | 64.90 | 95.24 | 23.28 |
| Qwen3.6-35B-A3B | 35 / 3 | 22.19 | 48.00 | 68.62 | 85.33 | 25.63 |
| SAIL | 35 / 3 | 33.25 | 85.33 | 86.51 | 91.67 | 35.24 |
Code execution, data analysis, and discovery
| Model | Params. (B), total / active | DS-1k | SUPER-Expert | CORE-Hard | DiscoveryBench | E2E-Bench | E2E-Bench-Hard |
|---|---|---|---|---|---|---|---|
| Ling-3.0-flash | 124 / 5.1 | 67.33 | 31.50 | 51.35 | 27.79 | 63.51 | 51.99 |
| DeepSeek-V4-Flash-0731 | 284 / 13 | 80.56 | 46.26 | 72.22 | 36.75 | 93.96 | 86.18 |
| Hy3 | 295 / 21 | 80.56 | 32.87 | 65.71 | 35.35 | 92.51 | 79.21 |
| MiMo-V2.5 | 310 / 15 | 71.89 | 35.67 | 54.05 | 36.50 | 65.23 | 50.49 |
| GLM-5.2 | 744 / 40 | 76.00 | 46.66 | 78.38 | 37.20 | 93.77 | 83.67 |
| Ring-2.6-1T | 1000 / 63 | 52.22 | 34.06 | 29.73 | 26.73 | 48.31 | 48.97 |
| MiMo-V2.5-Pro | 1020 / 42 | 68.22 | 36.66 | 64.86 | 44.49 | 74.83 | 63.87 |
| LongCat-2.0 | 1600 / 48 | 66.67 | 30.27 | 51.35 | 27.88 | 52.45 | 42.85 |
| BigBang-v1 | 35 / 3 | 67.11 | 31.54 | 51.35 | 33.55 | 75.00 | 68.26 |
| Apodex-1.0-mini | 35 / 3 | 68.11 | 24.83 | 43.20 | 31.20 | 39.94 | 39.47 |
| Nex-N2-mini | 35 / 3 | 62.70 | 30.98 | 62.20 | 32.07 | 62.74 | 53.90 |
| Nex-N2.5-mini | 35 / 3 | 51.10 | 34.57 | 64.86 | 33.21 | 83.83 | 70.02 |
| Agents-A1 | 35 / 3 | 72.22 | 30.98 | 56.80 | 33.91 | 32.39 | 18.28 |
| Qwen3.6-35B-A3B | 35 / 3 | 57.20 | 28.24 | 43.20 | 34.69 | 58.75 | 56.06 |
| SAIL | 35 / 3 | 74.30 | 37.78 | 67.57 | 37.48 | 89.46 | 77.27 |
Scientific coding and research
| Model | Params. (B), total / active | SciCode | DeepResearch Bench II |
|---|---|---|---|
| Ling-3.0-flash | 124 / 5.1 | 38.19 | 41.73 |
| DeepSeek-V4-Flash-0731 | 284 / 13 | 39.17 | 43.22 |
| Hy3 | 295 / 21 | 38.19 | 42.54 |
| MiMo-V2.5 | 310 / 15 | 27.64 | 27.46 |
| GLM-5.2 | 744 / 40 | 47.57 | 45.51 |
| Ring-2.6-1T | 1000 / 63 | 41.67 | 42.84 |
| MiMo-V2.5-Pro | 1020 / 42 | 40.28 | 41.70 |
| LongCat-2.0 | 1600 / 48 | 26.74 | 35.39 |
| BigBang-v1 | 35 / 3 | 41.70 | 38.55 |
| Apodex-1.0-mini | 35 / 3 | 43.10 | 37.91 |
| Nex-N2-mini | 35 / 3 | 35.10 | 41.00 |
| Nex-N2.5-mini | 35 / 3 | 26.83 | 33.55 |
| Agents-A1 | 35 / 3 | 38.19 | 33.33 |
| Qwen3.6-35B-A3B | 35 / 3 | 39.90 | 32.27 |
| SAIL | 35 / 3 | 50.35 | 42.61 |
Quick Start
Serve SAIL with either vLLM or SGLang in a fresh environment. Both examples enable tool calling and expose http://localhost:8000/v1 under the model name sail.
The commands use eight GPUs and a 262,144-token context limit. Adjust tensor parallelism and context length to fit your hardware. Replace the repository ID with a local model directory to use downloaded weights.
vLLM
pip install -U "vllm>=0.19.0"
vllm serve IQuestLab/SAIL \
--served-model-name sail \
--host 127.0.0.1 \
--port 8000 \
--tensor-parallel-size 8 \
--max-model-len 262144 \
--language-model-only \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder
SGLang
pip install -U "sglang[all]>=0.5.10"
python -m sglang.launch_server \
--model-path IQuestLab/SAIL \
--served-model-name sail \
--host 127.0.0.1 \
--port 8000 \
--tp-size 8 \
--mem-fraction-static 0.8 \
--context-length 262144 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder
See the vLLM deployment guide and SGLang installation guide for environment setup.
Send a request
curl http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "sail",
"messages": [{"role": "user", "content": "Write a Python function to simulate exponential decay and explain how to validate its numerical accuracy."}],
"max_tokens": 8192
}'
For tool use, include function definitions in the request's tools field and set tool_choice to "auto". Your application executes the returned tool calls and sends their results back to the model to continue the task. Search services, code execution, and research environments are supplied by the application.
License
SAIL is released under the Apache License 2.0.
Citation
@misc{sailmodelteam2026sailscientificagenticintelligence,
title = {Scientific Agentic Intelligence via a Science-Aware Loop},
author = {{SAIL Model Team}},
year = {2026},
eprint = {2610.11451},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2610.11451}
}
- Downloads last month
- -
Model tree for IQuestLab/SAIL
Base model
Qwen/Qwen3.6-35B-A3B
