Instructions to use shlv/AudioJev with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use shlv/AudioJev with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("shlv/AudioJev") model = AutoModelForMultimodalLM.from_pretrained("shlv/AudioJev", device_map="auto") - Notebooks
- Google Colab
- Kaggle
AudioJev
Direct audio decisions with order-calibrated candidate probabilities. Built with Qwen.
AudioJev takes an audio waveform, a natural-language question, and a list of candidate descriptions. It returns a probability distribution over those candidates using the next-token logits of their position labels. This release contains the random-derangement SKL (RD-SKL) model with lambda = 0.5 and training seed = 20261001, fine-tuned from Qwen2.5-Omni-3B.
Checkpoint
| Setting | Value |
|---|---|
| Model name | AudioJev |
| Base model | Qwen/Qwen2.5-Omni-3B |
| Training seed | 20261001 |
| SKL weight | 0.5 |
| Training schedule | Semantic: 1,024 updates; joint: 1,024; general: 2,048 |
| Total optimizer updates | 4,096 |
| Released checkpoint | Final general-stage step 2,048 |
| Checkpoint selection | Fixed final step |
| Training objective | Mean of two candidate cross-entropies + 0.5 ร aligned mean symmetric KL |
| Weight format | Original FP32 Safetensors shards, approximately 18.8 GB |
| Recommended inference dtype | BF16 |
| Candidate labels | 0โ9, then AโZ (2โ36 candidates) |
The audio encoder and language-model decision path were trained without a LoRA adapter. The visual branch was excluded from training; audio generation is disabled. This repository contains the complete saved Thinker checkpoint and its processor/tokenizer files. The weight shards are byte-identical to the evaluated checkpoint.
For each training example, RD-SKL supervises the original candidate order and a random derangement. The second distribution is aligned by candidate identity before applying symmetric KL. Inference uses one forward pass in the supplied candidate order.
Inference
Python inference, the HTTP service, installation dependencies, and runnable examples are maintained in the standalone AudioJev-Inference project. This model repository distributes the model weights, tokenizer/processor assets, model card, and license files.
After installing AudioJev-Inference, start the service with:
audiojev-serve --model shlv/AudioJev --device cuda:0
The Python API uses AudioJev("shlv/AudioJev"). Follow the AudioJev-Inference README for authentication, local/offline model loading, API requests, and deployment. The service downloads the model from this repository; it does not require a training-project checkout.
AudioJev decisions use candidate-label logits at the final input position, normalized over the supplied candidates. The service documents its candidate-order convention separately from the original-order evaluation below.
Evaluation
The following scores are for this single checkpoint in the original candidate order, not the mean over three training seeds.
| Benchmark | Questions | Accuracy |
|---|---|---|
| MMAU full (1,000 public + 9,000 officially scored hidden questions) | 10,000 | 69.41% |
| MMAR | 1,000 | 56.30% |
| MMSU evaluated short-audio subset | 3,931 | 62.45% |
The experiment's source audit found 119 hidden MMAU questions with audio-source overlap against general-stage fitting/development data; the official full score retains all questions. That score should be interpreted with this overlap in mind.
Scope
AudioJev is a closed-set decision model: probabilities are normalized over the candidates supplied by the caller. Provide a suitable fallback explicitly if the application requires one. Order consistency is the training objective; it does not establish that a reported probability equals an empirical correctness rate, and changing candidate order can still change a prediction.
The released model is intended for audio-conditioned classification and question answering. Open-ended generation, vision performance, speech generation, and live turn-taking deployment have not been established by this release. Training audio and benchmark answer data are not included.
License and attribution
This model is derived from Qwen2.5-Omni-3B and carries the Qwen Research License Agreement, supplied in LICENSE. The upstream agreement grants use for research or evaluation and requires a separate license from Alibaba Cloud for commercial use. See Notice for attribution and a description of modified files.
Built with Qwen. The original pretrained weights were modified by AudioJev fine-tuning with RD-SKL, lambda 0.5, seed 20261001. This release does not grant rights beyond the applicable upstream terms.
- Downloads last month
- 12
Model tree for shlv/AudioJev
Base model
Qwen/Qwen2.5-Omni-3B