strixAE

strixAE is an audio-language reasoning agent obtained by applying Group Relative Policy Optimization (GRPO) to Audio-Reasoner and merging the trained weights into a standalone checkpoint. It uses the Qwen2AudioForConditionalGeneration architecture and is intended for audio understanding, structured reasoning, and audio-restoration pipeline decisions.

Status: This repository contains a research checkpoint. Independent benchmark results for this exact GRPO checkpoint have not yet been published.

Model details

Item Value
Model type Audio-language conditional generation
Architecture Qwen2AudioForConditionalGeneration
Parameters 8.40B
Weight format BF16 Safetensors (4 shards)
Context length 8,192 text positions
Base checkpoint zhifeixie/Audio-Reasoner
Post-training GRPO reinforcement learning

Intended use

The model is designed for research involving:

  • audio understanding across speech, music, and environmental sounds;
  • audio-quality and degradation analysis;
  • deciding whether candidate restoration operations should be applied or skipped;
  • selecting and ordering an audio-restoration pipeline;
  • structured audio reasoning and question answering.

For restoration tasks, prompts should explicitly provide the available operations and request a decision for every operation. A useful output contract is:

<THINK>
Analyze the audio and evaluate every candidate operation.
</THINK>

**Detected Audio Issues:**
- ...

**Restoration Sequences Analysis:**
1. denoise(...)
   - Decision: APPLY / SKIP
   - Reason: ...

**Selected Restoration Pipeline:**
1. ...

Usage with Transformers

Install the runtime dependencies:

pip install "transformers>=4.57.0" accelerate librosa soundfile

Run inference on a local audio file:

import librosa
import torch
from transformers import AutoProcessor, Qwen2AudioForConditionalGeneration

model_id = "wujunjiehhs/strixAE"
audio_path = "example.wav"

processor = AutoProcessor.from_pretrained(model_id)
model = Qwen2AudioForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

conversation = [
    {
        "role": "system",
        "content": (
            "You are an audio reasoning assistant. Analyze the audio carefully "
            "and provide a clear, structured answer."
        ),
    },
    {
        "role": "user",
        "content": [
            {"type": "audio", "audio": audio_path},
            {
                "type": "text",
                "text": (
                    "Identify the audio issues. Evaluate denoising, dereverberation, "
                    "source separation, and super-resolution. For each operation, "
                    "return APPLY or SKIP with a reason, then propose an ordered pipeline."
                ),
            },
        ],
    },
]

prompt = processor.apply_chat_template(
    conversation,
    add_generation_prompt=True,
    tokenize=False,
)
audio, _ = librosa.load(
    audio_path,
    sr=processor.feature_extractor.sampling_rate,
    mono=True,
)
inputs = processor(
    text=prompt,
    audio=audio,
    sampling_rate=processor.feature_extractor.sampling_rate,
    return_tensors="pt",
    padding=True,
)
inputs = {k: v.to(model.device) for k, v in inputs.items()}

with torch.inference_mode():
    generated_ids = model.generate(**inputs, max_new_tokens=1024)

generated_ids = generated_ids[:, inputs["input_ids"].shape[1]:]
response = processor.batch_decode(
    generated_ids,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False,
)[0]
print(response)

The repository includes its chat template, tokenizer, processor configuration, generation configuration, and merged model weights; no adapter merge is required at inference time.

Training information

This checkpoint was produced by GRPO post-training and then merged for deployment. Exact training data composition, reward functions, hyperparameters, compute, and checkpoint-selection criteria are not included in this release. Users should not assume that benchmark numbers reported for the upstream Audio-Reasoner checkpoint transfer unchanged to this model.

Limitations and responsible use

  • The model may hallucinate acoustic events or infer degradations that are not present.
  • Restoration recommendations are model-generated judgments, not objective signal measurements.
  • Structured output is not guaranteed; validate the response before using it in an automated pipeline.
  • Performance may degrade for very long audio, low-resource languages, unseen codecs, unusual sampling rates, or out-of-domain recordings.
  • Do not use the model as the sole basis for safety-critical, medical, legal, forensic, surveillance, or high-impact decisions.
  • Audio can contain personal or sensitive information. Ensure that collection, processing, and sharing comply with applicable consent, privacy, copyright, and data-protection requirements.

中文简介

strixAE 是在 Audio-Reasoner 基础上进行 GRPO 强化学习得到的音频理解代理。它面向音频理解、音质问题分析、候选修复操作的 APPLY/SKIP 判断,以及修复流程排序。当前仓库尚未发布该 GRPO 检查点的独立评测结果,建议在实际数据上验证后再部署。

Acknowledgements and citation

This work builds on Audio-Reasoner and Qwen2-Audio. Please cite the upstream Audio-Reasoner work when using this checkpoint:

@misc{xie2025audioreasonerimprovingreasoningcapability,
  title        = {Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models},
  author       = {Zhifei Xie and Mingbao Lin and Zihang Liu and Pengcheng Wu and Shuicheng Yan and Chunyan Miao},
  year         = {2025},
  eprint       = {2503.02318},
  archivePrefix= {arXiv},
  primaryClass = {cs.SD},
  url          = {https://arxiv.org/abs/2503.02318}
}

License

This release follows the MIT license included with the upstream Audio-Reasoner project. Users are responsible for reviewing and complying with the terms of all upstream models, datasets, and dependencies.

Downloads last month
22
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wujunjiehhs/strixAE

Finetuned
(1)
this model

Paper for wujunjiehhs/strixAE