AudioJev

Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.

AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001, fine-tuned from Qwen2.5-Omni-3B.

Checkpoint

Setting Value
Model name AudioJev
Base model Qwen/Qwen2.5-Omni-3B
Training seed 20261001
SKL weight 0.5
Training schedule Semantic: 1,024 updates; joint: 1,024; general: 2,048
Total optimizer updates 4,096
Released checkpoint Final general-stage step 2,048
Checkpoint selection Fixed final step
Training objective Mean of two candidate cross-entropies + 0.5 ร— aligned mean symmetric KL
Weight format Original FP32 Safetensors shards, approximately 18.8 GB
Recommended inference dtype BF16
Candidate labels 0โ€“9, then Aโ€“Z (2โ€“36 candidates)

The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint.

For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order.

Inference

Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone AudioJev-Inference project. This model repository distributes the model weights, tokenizer/processor assets, model card, and license files.

After installing AudioJev-Inference, start the service with:

audiojev-serve --model shlv/AudioJev --device cuda:0

The Python API uses AudioJev("shlv/AudioJev"). Follow the AudioJev-Inference README for authentication, local/offline model loading, API requests, and deployment. The service downloads the model from this repository; it does not require a training-project checkout.

AudioJev decisions use candidate-label logits at the final input position, normalized over the supplied candidates. The service documents its candidate-order convention separately from the original-order evaluation below.

Evaluation

The following scores are for this single checkpoint in the original candidate order, not the mean over three training seeds.

Benchmark Questions Accuracy
MMAU full (1,000 public + 9,000 officially scored hidden questions) 10,000 69.41%
MMAR 1,000 56.30%
MMSU evaluated short-audio subset 3,931 62.45%

The experiment's source audit found 119 hidden MMAU questions with audio-source overlap against general-stage fitting/development data; the official full score retains all questions. That score should be interpreted with this overlap in mind.

Scope

AudioJev is a closed-set decision model: probabilities are normalized over the candidates supplied by the caller. Provide a suitable fallback explicitly if the application requires one. Order consistency is the training objective; it does not establish that a reported probability equals an empirical correctness rate, and changing candidate order can still change a prediction.

The released model is intended for audio-conditioned classification and question answering. Open-ended generation, vision performance, speech generation, and live turn-taking deployment have not been established by this release. Training audio and benchmark answer data are not included.

License and attribution

This model is derived from Qwen2.5-Omni-3B and carries the Qwen Research License Agreement, supplied in LICENSE. The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See Notice for attribution and a description of modified files.

Built with Qwen. The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.

Downloads last month
12
Safetensors
Model size
5B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for shlv/AudioJev

Finetuned
(30)
this model