SPX-CD-Pro

CMDB-1500 leaderboard · Release overview · Dataset · SPX-CD-Omni

SPX-CD-Pro (Simplex Calibrated Decision Model) by SurdAI is a LoRA adapter for Qwen/Qwen3.8-27B (27B dense). It supports text and image decisions: single-choice, multi-select, binary judgments, and ordinal ratings. The runner returns candidate probabilities without generating a reasoning trace.

LoRA rank/alpha 32/64 · Adapter 890.7 MiB · Context 65,536 tokens · Up to 1,024 candidates.

Download the base model separately; this repository provides the LoRA adapter and decision runner.

Download

pip install huggingface_hub
hf download SurdAI/SPX-CD-Pro --local-dir SPX-CD-Pro
cd SPX-CD-Pro

Run inference

Use separate environments for the two backends. Install CUDA-enabled PyTorch for Transformers; install the vLLM requirements in a clean environment.

Transformers, BF16 base:

pip install -r requirements-transformers.txt
python infer.py --backend transformers --adapter . \
  --input examples/text.json --output predictions.jsonl

vLLM, compatible 4-bit base:

pip install -r requirements-vllm.txt
python infer.py --backend vllm --base /path/to/compatible-4bit-base \
  --adapter . --effort 2 --input examples/text.json --output predictions.jsonl

Use a compatible 4-bit quantization of Qwen/Qwen3.8-27B with its matching processor. The inference backend must support both that quantization format and this LoRA adapter; 4-bit formats are not interchangeable. The exact evaluated configuration is recorded in EVALUATION.md.

For images or multi-select, use examples/image.json or examples/multi_select.json. See these files for the input format and python infer.py --help for options. Image paths are relative to the input file, or to --image-root when specified.

Results

CMDB-1500

Completed CMDB-1500 evaluation, effort 2, 4-bit base, vLLM:

Scope Questions Accuracy
Text 1,200 79.33%
Images 300 88.67%
All 1,500 81.20%

All 1,500 questions produced valid predictions. Evaluation settings are in EVALUATION.md; machine-readable results are in evaluation_results.json.

JevBench and Decision Index

The SurdAI release page reports the following results:

Effort JevBench Public (231) JevBench Hard (111) Decision Index 0.2.1 ↑
1 89.61% (207/231) 79.28% (88/111) 58.30
2 90.04% (208/231) 80.18% (89/111) —

Hard 111 is a subset of Public 231. Decision Index uses its own score scale. These results use the serving configurations described on the linked release page.

Settings

  • --max-context: default 65,536 total tokens. Each scoring pass reserves one output token, leaving at most 65,535 prompt tokens, including image tokens, candidates, and selected-label prefixes. Reduce this setting if needed for available memory.
  • --effort 1–5: average distinct candidate orderings, capped at two for binary questions. Default 1; the CMDB results use 2.
  • --temperature: candidate-softmax temperature, default 1.0.
  • --prompt-format: cmdb for text and image decisions; open-format for text single-choice tasks.
  • Multi-select uses sequential labels and STOP. Images: up to five per request, with a processor pixel budget of 524,288 per image. Inputs exceeding the context limit are rejected.

License and integrity

Apache 2.0 for this adapter and included code. The base model is distributed separately under its own model card. Verify release files with checksums.json. Release identity and original weight SHA-256 are recorded in release.json.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SurdAI/SPX-CD-Pro

Base model

Qwen/Qwen3.8-27B
Adapter
(179)
this model