DuplexJev-4B

Built on Qwen3-4B + the Qwen3-ASR-0.6B audio encoder.

DuplexJev reads typed, closed-set decisions from speech (has the user finished? which filler fits? who is speaking, in what mood?) as a single-token distribution of a frozen LLM: zero decode steps, many questions about the same clip in parallel. This repository is the complete model (encoder + trained connector + LLM, 4.2 B parameters, bf16) and runs on vLLM with a small plugin.

Audio encoder (frozen) Qwen3-ASR-0.6B encoder, 12.5 frames/s
Connector (trained, 12.6 M) 2 frames stacked → 6.25 audio tokens/s → SwiGLU MLP → LLM width
LLM (frozen) Qwen3-4B, used without thinking
Training content alignment (transcript distillation), then gender + emotion under the mixed objective
Same weights, connector-only form adventists-ai/DuplexJev-B-Para-Qwen3-ASR-0.6B-Qwen3-4B (for the duplexjev PyTorch package)

Serve with vLLM

pip install "vllm[audio]>=0.29" duplexjev-vllm
vllm serve adventists-ai/DuplexJev-4B --max-model-len 4096

duplexjev-vllm registers the model with vLLM (Qwen3-ASR encoder + connector; the LLM runs on vLLM's own Qwen3 kernels). Tested with vLLM 0.29 on one GPU; the model needs about 8 GB plus KV cache.

Each question is one chat request with max_tokens=1: list the options under letters and read the log-probabilities of the letters. Requests about the same clip share the prompt prefix (chat header + audio), which vLLM's prefix cache computes once. A ready-made client (openai + standard library only) is examples/vllm_client.py:

from vllm_client import DuplexJevClient

dj = DuplexJevClient("http://localhost:8000/v1")
dj.decide(open("call.wav", "rb").read(), {
    "turn":    ("Has the user finished speaking?", ["finished", "not finished"]),
    "gender":  ("What is the perceived gender of the speaker?", ["female", "male"]),
    "emotion": ("What is the speaker's emotional state?", ["neutral", "happy", "angry", "sad"]),
})
# {'turn': {'answer': 'finished', 'confidence': 0.88, 'probs': {...}}, 'gender': {...}, 'emotion': {...}}

The prompt it sends, which is the format the connector was trained on:

The user said: <|audio|>

Question: What is the perceived gender of the speaker?

Options:
A. female
B. male

Answer with only the letter of the correct option.

For Chinese speech ask in Chinese: 用户说:<|audio|>, 问题:…, 选项:, 请只回答正确选项的字母。 Send the audio as an input_audio part next to that text; restrict the answer with allowed_token_ids (the letter tokens) and set logprobs=True. The chat template turns thinking off by default.

Results

Paper protocol (single-token readout; % correct):

qa100 ZJU-ML Easy-Turn gender (800) emotion (800, 4-way)
72 51 60.2 89.4 91.9

The same prompts through the duplexjev PyTorch package and through vLLM with this repository (one H200, bf16):

qa100 ZJU-ML Easy-Turn gender emotion
PyTorch (duplexjev 0.2.2, connector repo) 77.0 49.0 60.1 88.0 91.4
vLLM 0.29 + duplexjev-vllm (this repo) 77.0 49.0 60.0 87.9 90.9

The OpenAI-compatible server gives the same numbers as the offline engine (qa100 77.0, gender 87.9). On a shared H200 the server answered 800 single-question requests in 22 s from 32 client threads, and one decision event (10 questions, one 4.5 s clip) in about 0.25 s end to end.

qa100: 100 bilingual spoken multiple-choice questions (adventists-ai/qa100). ZJU-ML: main-language part of the ZJU audio benchmark v2.0.0. Easy-Turn: 800-item turn-state test set (arXiv:2509.23938), read zero-shot. Gender: 800 real utterances (AISHELL-1, LibriSpeech). Emotion: 800 utterances (ESD, CREMA-D).

Limitations

  • Evaluated on short read or acted speech; not on streaming input.
  • Turn-taking (Easy-Turn) and spoken factual QA are weaker than in larger DuplexJev models; see the leaderboard.
  • Emotion labels come from acted corpora; do not use the outputs to make decisions about individuals.

License

CC BY-NC 4.0, research use only. The connector was trained on the ESD emotional speech corpus, which is licensed for research purposes only, and on CREMA-D (ODbL). The base models keep their own licences: Qwen3-4B and Qwen3-ASR-0.6B are Apache-2.0 (Copyright Alibaba Cloud); this repository redistributes their weights unchanged.

Citation

@inproceedings{jin2027duplexjev,
  title     = {Batched Speech Decisions Without Decoding: Single-Token Supervision Lets a Frozen {LLM} Hear Beyond the Transcript},
  author    = {Jin, Jie and Ma, Ziyin and Yin, Min and Chen, Jinyu and Song, Haigang and Pang, Zhikun and Zhang, Xiaowen},
  booktitle = {Submitted to IEEE ICASSP},
  year      = {2027}
}

Built by Adventists.ai. Claude (Anthropic) assisted with code.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adventists-ai/DuplexJev-4B

Finetuned
Qwen/Qwen3-4B
Finetuned
(1102)
this model

Paper for adventists-ai/DuplexJev-4B