DuplexJev-4B
Built on Qwen3-4B + the Qwen3-ASR-0.6B audio encoder.
DuplexJev reads typed, closed-set decisions from speech (has the user finished? which filler fits? who is speaking, in what mood?) as a single-token distribution of a frozen LLM: zero decode steps, many questions about the same clip in parallel. This repository is the complete model (encoder + trained connector + LLM, 4.2 B parameters, bf16) and runs on vLLM with a small plugin.
| Audio encoder (frozen) | Qwen3-ASR-0.6B encoder, 12.5 frames/s |
| Connector (trained, 12.6 M) | 2 frames stacked → 6.25 audio tokens/s → SwiGLU MLP → LLM width |
| LLM (frozen) | Qwen3-4B, used without thinking |
| Training | content alignment (transcript distillation), then gender + emotion under the mixed objective |
| Same weights, connector-only form | adventists-ai/DuplexJev-B-Para-Qwen3-ASR-0.6B-Qwen3-4B (for the duplexjev PyTorch package) |
Serve with vLLM
pip install "vllm[audio]>=0.29" duplexjev-vllm
vllm serve adventists-ai/DuplexJev-4B --max-model-len 4096
duplexjev-vllm registers the model with vLLM (Qwen3-ASR encoder + connector; the LLM runs on vLLM's own Qwen3
kernels). Tested with vLLM 0.29 on one GPU; the model needs about 8 GB plus KV cache.
Each question is one chat request with max_tokens=1: list the options under letters and read the log-probabilities
of the letters. Requests about the same clip share the prompt prefix (chat header + audio), which vLLM's prefix cache
computes once. A ready-made client (openai + standard library only) is
examples/vllm_client.py:
from vllm_client import DuplexJevClient
dj = DuplexJevClient("http://localhost:8000/v1")
dj.decide(open("call.wav", "rb").read(), {
"turn": ("Has the user finished speaking?", ["finished", "not finished"]),
"gender": ("What is the perceived gender of the speaker?", ["female", "male"]),
"emotion": ("What is the speaker's emotional state?", ["neutral", "happy", "angry", "sad"]),
})
# {'turn': {'answer': 'finished', 'confidence': 0.88, 'probs': {...}}, 'gender': {...}, 'emotion': {...}}
The prompt it sends, which is the format the connector was trained on:
The user said: <|audio|>
Question: What is the perceived gender of the speaker?
Options:
A. female
B. male
Answer with only the letter of the correct option.
For Chinese speech ask in Chinese: 用户说:<|audio|>, 问题:…, 选项:, 请只回答正确选项的字母。
Send the audio as an input_audio part next to that text; restrict the answer with allowed_token_ids (the letter
tokens) and set logprobs=True. The chat template turns thinking off by default.
Results
Paper protocol (single-token readout; % correct):
| qa100 | ZJU-ML | Easy-Turn | gender (800) | emotion (800, 4-way) |
|---|---|---|---|---|
| 72 | 51 | 60.2 | 89.4 | 91.9 |
The same prompts through the duplexjev PyTorch package and through vLLM with this repository (one H200, bf16):
| qa100 | ZJU-ML | Easy-Turn | gender | emotion | |
|---|---|---|---|---|---|
PyTorch (duplexjev 0.2.2, connector repo) |
77.0 | 49.0 | 60.1 | 88.0 | 91.4 |
vLLM 0.29 + duplexjev-vllm (this repo) |
77.0 | 49.0 | 60.0 | 87.9 | 90.9 |
The OpenAI-compatible server gives the same numbers as the offline engine (qa100 77.0, gender 87.9). On a shared H200 the server answered 800 single-question requests in 22 s from 32 client threads, and one decision event (10 questions, one 4.5 s clip) in about 0.25 s end to end.
qa100: 100 bilingual spoken multiple-choice questions (adventists-ai/qa100).
ZJU-ML: main-language part of the ZJU audio benchmark v2.0.0. Easy-Turn: 800-item turn-state test set
(arXiv:2509.23938), read zero-shot. Gender: 800 real utterances (AISHELL-1, LibriSpeech). Emotion: 800 utterances
(ESD, CREMA-D).
Limitations
- Evaluated on short read or acted speech; not on streaming input.
- Turn-taking (Easy-Turn) and spoken factual QA are weaker than in larger DuplexJev models; see the leaderboard.
- Emotion labels come from acted corpora; do not use the outputs to make decisions about individuals.
License
CC BY-NC 4.0, research use only. The connector was trained on the ESD emotional speech corpus, which is licensed for research purposes only, and on CREMA-D (ODbL). The base models keep their own licences: Qwen3-4B and Qwen3-ASR-0.6B are Apache-2.0 (Copyright Alibaba Cloud); this repository redistributes their weights unchanged.
Citation
@inproceedings{jin2027duplexjev,
title = {Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen {LLM} Hear Beyond the Transcript},
author = {Jin, Jie and Ma, Ziyin and Yin, Min and Chen, Jinyu and Song, Haigang and Pang, Zhikun and Zhang, Xiaowen},
booktitle = {Submitted to IEEE ICASSP},
year = {2027}
}
Built by Adventists.ai. Claude (Anthropic) assisted with code.
- Downloads last month
- -