Instructions to use wujunjiehhs/strixAE with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wujunjiehhs/strixAE with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("wujunjiehhs/strixAE") model = AutoModelForMultimodalLM.from_pretrained("wujunjiehhs/strixAE", device_map="auto") - Notebooks
- Google Colab
- Kaggle
strixAE
strixAE is an audio-language reasoning agent obtained by applying Group Relative Policy Optimization (GRPO) to Audio-Reasoner and merging the trained weights into a standalone checkpoint. It uses the Qwen2AudioForConditionalGeneration architecture and is intended for audio understanding, structured reasoning, and audio-restoration pipeline decisions.
Status: This repository contains a research checkpoint. Independent benchmark results for this exact GRPO checkpoint have not yet been published.
Model details
| Item | Value |
|---|---|
| Model type | Audio-language conditional generation |
| Architecture | Qwen2AudioForConditionalGeneration |
| Parameters | 8.40B |
| Weight format | BF16 Safetensors (4 shards) |
| Context length | 8,192 text positions |
| Base checkpoint | zhifeixie/Audio-Reasoner |
| Post-training | GRPO reinforcement learning |
Intended use
The model is designed for research involving:
- audio understanding across speech, music, and environmental sounds;
- audio-quality and degradation analysis;
- deciding whether candidate restoration operations should be applied or skipped;
- selecting and ordering an audio-restoration pipeline;
- structured audio reasoning and question answering.
For restoration tasks, prompts should explicitly provide the available operations and request a decision for every operation. A useful output contract is:
<THINK>
Analyze the audio and evaluate every candidate operation.
</THINK>
**Detected Audio Issues:**
- ...
**Restoration Sequences Analysis:**
1. denoise(...)
- Decision: APPLY / SKIP
- Reason: ...
**Selected Restoration Pipeline:**
1. ...
Usage with Transformers
Install the runtime dependencies:
pip install "transformers>=4.57.0" accelerate librosa soundfile
Run inference on a local audio file:
import librosa
import torch
from transformers import AutoProcessor, Qwen2AudioForConditionalGeneration
model_id = "wujunjiehhs/strixAE"
audio_path = "example.wav"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen2AudioForConditionalGeneration.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
conversation = [
{
"role": "system",
"content": (
"You are an audio reasoning assistant. Analyze the audio carefully "
"and provide a clear, structured answer."
),
},
{
"role": "user",
"content": [
{"type": "audio", "audio": audio_path},
{
"type": "text",
"text": (
"Identify the audio issues. Evaluate denoising, dereverberation, "
"source separation, and super-resolution. For each operation, "
"return APPLY or SKIP with a reason, then propose an ordered pipeline."
),
},
],
},
]
prompt = processor.apply_chat_template(
conversation,
add_generation_prompt=True,
tokenize=False,
)
audio, _ = librosa.load(
audio_path,
sr=processor.feature_extractor.sampling_rate,
mono=True,
)
inputs = processor(
text=prompt,
audio=audio,
sampling_rate=processor.feature_extractor.sampling_rate,
return_tensors="pt",
padding=True,
)
inputs = {k: v.to(model.device) for k, v in inputs.items()}
with torch.inference_mode():
generated_ids = model.generate(**inputs, max_new_tokens=1024)
generated_ids = generated_ids[:, inputs["input_ids"].shape[1]:]
response = processor.batch_decode(
generated_ids,
skip_special_tokens=True,
clean_up_tokenization_spaces=False,
)[0]
print(response)
The repository includes its chat template, tokenizer, processor configuration, generation configuration, and merged model weights; no adapter merge is required at inference time.
Training information
This checkpoint was produced by GRPO post-training and then merged for deployment. Exact training data composition, reward functions, hyperparameters, compute, and checkpoint-selection criteria are not included in this release. Users should not assume that benchmark numbers reported for the upstream Audio-Reasoner checkpoint transfer unchanged to this model.
Limitations and responsible use
- The model may hallucinate acoustic events or infer degradations that are not present.
- Restoration recommendations are model-generated judgments, not objective signal measurements.
- Structured output is not guaranteed; validate the response before using it in an automated pipeline.
- Performance may degrade for very long audio, low-resource languages, unseen codecs, unusual sampling rates, or out-of-domain recordings.
- Do not use the model as the sole basis for safety-critical, medical, legal, forensic, surveillance, or high-impact decisions.
- Audio can contain personal or sensitive information. Ensure that collection, processing, and sharing comply with applicable consent, privacy, copyright, and data-protection requirements.
中文简介
strixAE 是在 Audio-Reasoner 基础上进行 GRPO 强化学习得到的音频理解代理。它面向音频理解、音质问题分析、候选修复操作的 APPLY/SKIP 判断,以及修复流程排序。当前仓库尚未发布该 GRPO 检查点的独立评测结果,建议在实际数据上验证后再部署。
Acknowledgements and citation
This work builds on Audio-Reasoner and Qwen2-Audio. Please cite the upstream Audio-Reasoner work when using this checkpoint:
@misc{xie2025audioreasonerimprovingreasoningcapability,
title = {Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models},
author = {Zhifei Xie and Mingbao Lin and Zihang Liu and Pengcheng Wu and Shuicheng Yan and Chunyan Miao},
year = {2025},
eprint = {2503.02318},
archivePrefix= {arXiv},
primaryClass = {cs.SD},
url = {https://arxiv.org/abs/2503.02318}
}
License
This release follows the MIT license included with the upstream Audio-Reasoner project. Users are responsible for reviewing and complying with the terms of all upstream models, datasets, and dependencies.
- Downloads last month
- 22
Model tree for wujunjiehhs/strixAE
Base model
zhifeixie/Audio-Reasoner