MiDashengLM-Spatial

Unifying General Audio Understanding and Spatial Awareness

arXiv  HuggingFace Model  Demo Page  GitHub

Overview

MiDashengLM-Spatial, to our knowledge, is the first open-source end-to-end unified audio-language model that supports both general audio understanding and spatial awareness within a single architecture. It extends MiDashengLM with Spatial-Dasheng through a hierarchical semantic-to-spatial conditioning (HSSC) module to separately capture and effectively integrate semantic and spatial audio information. This design enables binaural perception and spatial audio understanding capabilities while preserving general audio understanding capabilities of the base model.

Architecture

Architecture

Installation

pip install torch torchaudio "transformers==4.57.6" einops

Usage

import torch
from transformers import AutoModelForCausalLM, AutoProcessor

model_id = "mispeech/midashenglm-spatial"  # or a local package directory
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "audio", "path": "example/example.wav"}, 
            # {"type": "audio", "audio": np.random.randn(2, 160000)}, 
            {"type": "text", "text": "Write a general audio caption describing the sound source and its spatial information."},
        ],
    },
]

with torch.no_grad():
    inputs = processor.apply_chat_template(
        messages, tokenize=True, add_generation_prompt=True,
        add_special_tokens=True, return_dict=True,
    ).to(model.device)
    generation = model.generate(**inputs, max_new_tokens=256, do_sample=False)
    print(processor.tokenizer.batch_decode(generation, skip_special_tokens=True))

Put audio before text (the trained order), and feed real binaural (2, T) audio — mono still runs but carries no direction.

Citation

License

Released under the Apache License 2.0 (see LICENSE), for both research and commercial use.

Downloads last month
2
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support