blink-27b

blink-27b one-pass readout diagram: a state and typed choice, noul and score questions are rendered as evidence, criterion and lettered options; questions are batched and each batch is one forward pass, giving next-token logits at the answer position, and an FP32 softmax over the offered letters gives one probability per option. No text is generated.

The most accurate blink model for hard typed decisions. It also performed best of the three in local browser-agent runs that passed page elements as text, not screenshots; see browser-agent setup. These local pages are not a benchmark.

Send text or JSON state with choice, noul (yes/no), or score questions. Get a probability for each offered answer without generated text. Each batch uses one forward pass; large requests can use several batches.

For browser agents, send the page and candidate elements as JSON state; make operations and click targets choice questions. The probability map lets an agent pick among its proposed actions. This is a text-element workflow, not autonomous web navigation; for pixels, use the optional image path below or MiMo's Screen click demo.

Try it: Space (the smaller models run live; self-host this one) · Screen click · Computer use · API · Docs · GitHub · blink-4b · blink-mimo-9b

At a glance

Attribute Detail
Base model Qwen/Qwen3.8-27B, text weights only
Weights size 53.8 GB (bf16, 26,895,998,464 parameters)
Revision v1.4 (code revision; weights identical to v1.0)
License Weights: non-commercial research and evaluation only (LICENSE.md); code: Apache-2.0. Base-model notice: Apache-2.0 (LICENSE-Qwen).

Quickstart

The example asks several questions about the same state. Replace the message with your page, policy, or workflow text and set criteria to actions your application actually supports. The caller decides what to do with the returned probabilities.

# pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
import os, sys
from huggingface_hub import hf_hub_download

os.environ["BLINK_MODEL"] = "thegovind/blink-27b"
os.environ["BLINK_REVISION"] = "v1.4"
sys.path.insert(0, os.path.dirname(hf_hub_download("thegovind/blink-27b", "blink.py", revision="v1.4")))
import blink

out = blink.decide(
    "Order #4411 arrived with a cracked screen. I want my money back, not another one.",
    {
        "intent": {
            "type": "choice",
            "instructions": "What does the customer want?",
            "criteria": {"refund": "Money back", "replacement": "A new unit", "info": "Information only"},
        },
        "urgent": {"type": "noul", "instructions": "Does this need a reply today?"},
        "anger": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]},
    },
)
print(out["answers"]["intent"]["probabilities"])

Run it as a server

serve.py serves POST /v1/systemone, GET /v1/models, and GET /healthz. Point TypeSafe's server-side Python or JavaScript SDKs at it using TYPESAFE_BASE_URL for text decisions; its request and response fields match hosted Jev. Cross-request batching is optional with --batch-window-ms 5.

pip install "torch==2.13.0" "transformers==5.17.0" "flash-linear-attention==0.5.2" "accelerate>=1.1.0" safetensors huggingface_hub
hf download thegovind/blink-27b --revision v1.4 --local-dir blink-27b
python blink-27b/serve.py --model ./blink-27b --port 8000
# TypeSafe SDKs: export TYPESAFE_BASE_URL=http://127.0.0.1:8000 TYPESAFE_API_KEY=any

Or use Docker from the downloaded folder:

cd blink-27b
docker build -t blink-27b . && docker run --rm --gpus all -p 127.0.0.1:8000:8000 blink-27b

The vLLM path did not pass blink-27b's quality check: agreement with serve.py on the 5-option web-action set was 489/500, below the 493/500 limit. Use serve.py.

Screenshots (opt-in, self-hosted)

Image input is off by default. Start serve.py with --vision-tower Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 to attach the matching tower. This is a self-hosted blink extension. TypeSafe's hosted Jev is text-only.

With --vision-tower, self-hosted blink borrows the pinned Qwen base model's vision encoder while its checkpoint stays text-only; blink-mimo-9b uses its own encoder and runs Screen click.

Put a data:image/png;base64,... URI (JPEG and WebP data URIs work too) inside a state string, or pass data URIs in a top-level images list. Image URLs are never fetched. If loading only blink.py via hf_hub_download, also download graft_keys.py from the same repo and revision beside it.

See Computer use to self-host this model with images; its text-only browser-agent use is covered in the API docs.

Results

Local development readout Result
Decision Index 0.2 balanced skill 52.72
JevBench public hard items 90/111 batched; 89/111 via serial serve.py

No official JevBench score for blink has been published. These public-item results are not an official score, rank, or parity claim. Decision Index is a descriptive local run of the official kit, not a leaderboard submission. Training included MMLU-Pro test-partition questions, so that component is contaminated.

Model details: architecture, training, data

Architecture and readout

blink-27b network diagram: 64 decoder layers repeating 3 Gated DeltaNet layers then one full-attention layer (48 and 16 in total, hidden 5120), with full attention at 0-based layers 3, 7, … 63 as in the tensor names; LoRA rank 16 on every attention, Gated DeltaNet and MLP projection (116.7M parameters, merged after training); token embeddings, norms and lm_head frozen, with lm_head untied, a separate matrix; the vision encoder and multi-token-prediction head removed; and the answer read from the offered option-letter rows of lm_head.

Qwen3.8-27B text backbone: 64 decoder layers (48 Gated DeltaNet, 16 full-attention), hidden 5120, and 26,895,998,464 shipped parameters. Training merged 116.7M LoRA parameters at rank 16, alpha 32: q_proj, k_proj, v_proj, o_proj; in_proj_qkv, in_proj_z, in_proj_a, in_proj_b, out_proj; gate_proj, up_proj, down_proj. The vision encoder and MTP head were removed: 0 vision tensors, 0 MTP tensors.

Readout softmaxes FP32 next-token logits over the offered labels. These option-conditional probabilities are not certified chances of being right.

Training

blink-27b post-training diagram: T2 (72,700 rows) and T4 (74,754 rows) broken down by data category feed supervised fine-tuning with cross-entropy over the offered option letters; T2 starts from the base and T4 continues from it, then the adapters are merged into release v1.0. Decision Index 0.1 full suite 63.44.

Supervised fine-tuning on target distributions, no RL or preference optimization. T2 used 72,700 question rows, lr 5e-5, 230 steps; T4 continued with 74,754 rows, lr 3e-5, 856 more steps.

Stage Question rows Mix
T2 72,700 41,401 public-source · 16,000 program-generated reasoning · 12,000 decision worlds · 3,299 exact-probability worlds
T4 74,754 23,894 program-generated reasoning · 23,252 public-source · 12,000 decision worlds · 7,860 teacher-written rows · 6,000 chess move choices · 1,748 exact-probability worlds

Data sources and licences

Source Licence
MMLU auxiliary train, MMLU-Pro, CommonsenseQA MIT
AQuA-RAT, Amazon ESCI Apache-2.0
searchless_chess data CC BY 4.0 (Lichess-derived portions CC0); code Apache-2.0
MedMCQA Apache-2.0 (dataset card)
SuperGPQA ODC-BY
WANLI, ContractNLI, BANKING77, GPQA CC BY 4.0
ARC CC BY-SA 4.0
BoolQ CC BY-SA 3.0
ANLI CC BY-NC 4.0
SciQ CC BY-NC 3.0
iSarcasmEval MIT (upstream repository licence)
VAST, Humicroedit, OpenBookQA None stated by source
Code-generated worlds and teacher-written documents (Qwen3.8-27B) See LICENSE.md

Source-repository licences do not settle rights in underlying texts.

Evaluation and limits

  • Training overlap. The archived Decision Index 0.1 local run scored 63.44. Its public train splits overlap the training mix. About 3.2k MMLU-Pro test-partition questions in training contaminate that Decision Index 0.2 component. Semantic and pretraining overlap cannot be ruled out.
  • Limits. English-centric; no chatting or explanations. Text in state can sway decisions, and date/number reasoning is a weak case. The default server accepts 255 options per choice and 2–10 score levels; over-limit requests return 422 without truncation. --image-layout first|inline controls image placement (first by default); --model-name blink-27b handles renamed folders. Neither flag enables image input.

License

Code: Apache-2.0. Weights: non-commercial research and evaluation only; see each model card's license.

See LICENSE.md for weight terms; the Qwen base model is Apache-2.0 (LICENSE-Qwen).

Downloads last month
91
Safetensors
Model size
27B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thegovind/blink-27b

Base model

Qwen/Qwen3.8-27B
Finetuned
(418)
this model

Space using thegovind/blink-27b 1