Jolt-2B
Jolt-2B is a multimodal decision model. Given text or an image and a set of candidate answers, it returns candidate probabilities, a truth estimate, or an ordered score. It does not generate free-form answers.
Model details
| Property | Value |
|---|---|
| Base model | Qwen/Qwen3.5-2B |
| Base revision | 15852e8c16360a2fea060d615a32b45270f8a8fc |
| Checkpoint | L2, seed 44, selected at update 2,800 |
| Development macro accuracy | 82.89% (checkpoint selection partition) |
| Training objective | Direct-composite loss |
| Trainable weights | Rank-8 language LoRA and a linear candidate scorer |
| Input | Text, or text plus one decoded image |
| Limits | 2–128 candidates per question; 4,096-token inference limit |
The reported development score selected the checkpoint within its run. It is not an independent test score. Test data were not used to select this checkpoint.
JevBench
18 text dataset benchmarks
8 image dataset benchmarks
You can find a comprehensive model comparison report at: model_comparison_dashboard.html
Use
This repository contains the compact Jolt adapter and inference source. It does
not duplicate Qwen's base weights: the runtime downloads the pinned public Qwen
revision on first use and caches it locally. The adapter format is Jolt's custom
decision-model format, so load it with the included runtime rather than
AutoModel.from_pretrained.
The runtime currently requires Python 3.12+, a CUDA-enabled PyTorch setup, and a CUDA GPU. From a clone or downloaded copy of this repository:
python -m pip install -r requirements-cuda.txt
python -m pip install -e . --no-deps
Then:
from jolt.predictor import Predictor
model = Predictor.from_checkpoint("mlengineer-ai/Jolt-2B")
result = model.predict(
{"message": "I was charged twice. Please refund the duplicate."},
{
"route": {
"type": "choice",
"instructions": "Which team should handle this request?",
"criteria": ["billing", "technical support", "sales"],
},
"refund": {
"type": "noul",
"instructions": "Is a refund requested?",
},
},
)
print(result["answers"]["route"]["choice"])
print(result["answers"]["refund"]["noul"])
For an image question, pass a decoded PIL image as state["image"]. This
example uses the model loaded above:
from PIL import Image
image = Image.open("path/to/image.jpg").convert("RGB")
image_result = model.predict(
{"image": image, "question": "Which animal is shown?"},
{
"animal": {
"type": "choice",
"instructions": "Identify the animal in the image.",
"criteria": ["cat", "dog", "bird", "other"],
}
},
)
print(image_result["answers"]["animal"]["choice"])
Training and data
Jolt follows the recipe based on the Dohnuts project. These checkpoints start from the Qwen3.5 base model listed above; they do not use Dohnuts model weights as their base. Jolt trained rank-8 LoRA weights and a candidate scorer with the base model frozen. Jolt's public-data task suite and split design are adapted from Dohnuts. Its data reference lists source datasets and their known terms. The adapter metadata records fingerprints for Jolt's training, development, calibration, and test partitions. Training data are not included here.
Licensing
This release offers the adapter and decision-head weights under CC BY-NC-SA 4.0;
see LICENSE. For this prepared release, this conservative default follows the
original Dohnuts model release. It is not a legal determination that every source
dataset's terms necessarily apply to trained weights. The training mixture
includes sources with different terms, including
ScienceQA, whose
maintainers specify CC BY-NC-SA 4.0 for the dataset. The available records do
not establish commercial clearance for these trained weights; review the source
dataset terms before use.
The bundled inference source is adapted from Dohnuts and is licensed under
Apache-2.0; see LICENSE-CODE and NOTICE-CODE. The Qwen base model is
separately licensed by its publisher under Apache-2.0. Those licenses do not
replace or change the terms of training data.



