mJev-Qwen3-VL-4B-RLCD

Jev, with senses. Multimodal context in, explicit decisions out.

GitHub · Evaluation dataset · Deployment guide

mJev-Qwen3-VL-4B-RLCD is SoMark's 4B multimodal decision model, fine-tuned with GRPO from Qwen3-VL-4B-Instruct. It is designed to select among explicit candidates given visual context and a question.

With the mJev runtime, each question returns a selected answer, candidate probabilities and raw logits. Multiple questions can share the same context, with optional prefix-cache reuse. This repository contains the full fine-tuned weights and processor/tokenizer files; the inference runtime lives in the mJev code repository.

Model at a glance

Property Details
Developer SoMark
Model family Qwen3-VL, 4B
Task Multimodal probabilistic decision making
Inputs Images or video with text questions and candidate choices
Decision interface Candidate-label logits, normalized candidate probabilities and an argmax decision
Post-training GRPO with target-probability-weighted correctness rewards
Weight format Full model, BF16 Safetensors
License Apache 2.0

Evaluation

On a 195-question evaluation subset of mJev-Compositional-VQA, GRPO fine-tuning improves accuracy from 77.95% to 80.00% (+2.05 percentage points), answering four additional questions correctly.

Model Correct / total Accuracy
Qwen3-VL-4B-Instruct — before GRPO 152 / 195 77.95%
mJev-Qwen3-VL-4B-RLCD — after GRPO 156 / 195 80.00%

Both models were evaluated on the same questions, held out from RL training. Accuracy is the number of correct answers divided by the total number of questions. These results describe this evaluation subset; broader benchmark performance has not been established by this comparison.

Quick start

Use the mJev runtime to obtain candidate probabilities and decisions. The HF path scores candidate labels with a model forward pass; it does not require generating a free-form answer.

Prepare Linux, Python 3.11+, a compatible NVIDIA GPU environment and system FFmpeg. Install the runtime and download this model:

git clone https://github.com/SoMarkAI/mJev.git
cd mJev
python3 -m venv .venv
source .venv/bin/activate
python -m pip install torchcodec==0.11.0+cpu --index-url https://download.pytorch.org/whl/cpu
python -m pip install -e '.[hf-vl]' huggingface_hub

export MODEL_DIR="$HOME/models/mJev-Qwen3-VL-4B-RLCD"
hf download SoMarkAI/mJev-Qwen3-VL-4B-RLCD --local-dir "$MODEL_DIR"

CUDA_VISIBLE_DEVICES=0 python demo_hf.py \
  --model "$MODEL_DIR" \
  --input examples/motion-demo/input.json \
  --mode causal --numerics stable --projection full \
  --prefix-cache --question-batch-size 3 \
  --output outputs/mjev-decisions.json

The example asks three questions about a project-created video. Inspect outputs/mjev-decisions.json for decisions and candidate scores. For your own image, create an input JSON file with this structure and pass it to --input:

{
  "image": "your-image.png",
  "context": "Inspect the supplied picture.",
  "questions": [
    {
      "question": "Which object is visible?",
      "candidates": ["A bicycle", "A car", "A bus"]
    }
  ]
}

Media paths are relative to the input JSON file. See the HF guide for the Python API, video inputs and memory settings. The weights retain the standard Qwen3-VL architecture and can also be loaded with Transformers; use the mJev runtime for the candidate-scoring interface described here.

Training and reward

Before RL, a frozen candidate scorer records the probability p_target assigned to the ground-truth option. Each sampled label receives +p_target when correct and -p_target when incorrect or invalid. Rewards are bounded in [-1, 1], and examples with higher target probabilities produce stronger training signals.

GRPO rollouts are constrained to a single candidate label, with no separate format reward. Reward scaling is disabled to preserve the target-probability weighting. A separate KL penalty limits drift from the reference model. Training updates the language model while keeping the vision tower and aligner frozen.

Use and limitations

  • Intended for visual question answering and decision workflows with explicitly defined candidate sets.
  • Candidate probabilities are relative to the supplied choices, not calibrated confidence estimates. Candidate wording, order and coverage can affect decisions.
  • This release supports visual inputs; it does not add audio support. The reported RL evaluation covers image questions, not video or audio benchmarks.
  • The 195-question comparison is an initial evaluation. It does not establish broad reasoning gains or statistical significance.
  • Input length, image resolution, video sampling and question batching affect memory use and latency. See the runtime documentation for supported settings.

License and acknowledgments

Released under Apache 2.0. mJev builds on Qwen3-VL-4B-Instruct; we thank the Qwen team for the base model and the Jev project for the decision-oriented interface inspiration. Evaluation data retains its own license and terms.

For issues, implementation details and contributions, visit SoMarkAI/mJev.

Downloads last month
43
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SoMarkAI/mJev-Qwen3-VL-4B-RLCD

Finetuned
(460)
this model