You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

ASPECT-8B is released for non-commercial research use under CC BY-NC-SA 4.0. It is not a medical device and must not be used for diagnosis or clinical decision-making. Access requests are reviewed manually.

Log in or Sign Up to review the conditions and access this model content.

ASPECT-8B

ASPECT is a pathology vision-language model that reports the nucleus counts behind its answers. It is built on Qwen3-VL-8B-Instruct with 8 pathology-feature tokens and 6 cell tokens, trained with three-stage supervised fine-tuning (Perceive, Generate, Reason) and GRPO with an answer–observation consistency reward.

Responses follow the format

<think> the patch feature of the image is <|anchor_start|>...<|anchor_end|>, and the cell composition of the image is <|anchor_start|>...<|anchor_end|>. </think>
<observe> description {"<count name>": <count>, ...}</observe>
<answer> reasoning
FINAL: <option> </answer>

Code: https://github.com/ChyaZhang/ASPECT

Files

The repository root holds the supervised model (Qwen3-VL-8B with the visual-token embeddings and the SFT LoRA merged in); rl_adapter/ holds the LoRA adapter from reinforcement learning. Load both, as below; this is the configuration evaluated on PathoVernier.

Usage

import torch
from peft import PeftModel
from PIL import Image
from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "Mikezcy/ASPECT-8B"
processor = AutoProcessor.from_pretrained(model_id, max_pixels=1360 * 28 * 28)
model = AutoModelForImageTextToText.from_pretrained(model_id, dtype=torch.bfloat16, device_map="cuda")
model = PeftModel.from_pretrained(model, model_id, subfolder="rl_adapter")

image = Image.open("example.png").convert("RGB").resize((512, 512))
question = "..."
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": question}]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False))

Images are resized to 512x512. No system prompt is used; the question should list the answer options.

PathoVernier results

Model Acc ↑ CA ↑ RAWR ↓ Count Acc ↑
ASPECT-8B 0.746 0.809 0.489 0.461

Greedy decoding, one response per question, at most 512 new tokens.

Limitations

ASPECT was trained and evaluated on H&E patches at 20x-40x from colon, skin, breast and the PanNuke tissues. Counts can be wrong even when the final answer is correct. The model is intended for research and must not be used for diagnosis.

Citation

Coming soon.

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Mikezcy/ASPECT-8B

Finetuned
(602)
this model