Instructions to use mispeech/midashenglm-spatial with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mispeech/midashenglm-spatial with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("mispeech/midashenglm-spatial", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
MiDashengLM-Spatial
Unifying General Audio Understanding and Spatial Awareness
Overview
MiDashengLM-Spatial, to our knowledge, is the first open-source end-to-end unified audio-language model that supports both general audio understanding and spatial awareness within a single architecture. It extends MiDashengLM with Spatial-Dasheng through a hierarchical semantic-to-spatial conditioning (HSSC) module to separately capture and effectively integrate semantic and spatial audio information. This design enables binaural perception and spatial audio understanding capabilities while preserving general audio understanding capabilities of the base model.
Architecture
Installation
pip install torch torchaudio "transformers==4.57.6" einops
Usage
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
model_id = "mispeech/midashenglm-spatial" # or a local package directory
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
messages = [
{
"role": "user",
"content": [
{"type": "audio", "path": "example/example.wav"},
# {"type": "audio", "audio": np.random.randn(2, 160000)},
{"type": "text", "text": "Write a general audio caption describing the sound source and its spatial information."},
],
},
]
with torch.no_grad():
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True,
add_special_tokens=True, return_dict=True,
).to(model.device)
generation = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(processor.tokenizer.batch_decode(generation, skip_special_tokens=True))
Put audio before text (the trained order), and feed real binaural (2, T) audio — mono still runs but carries no direction.
Citation
License
Released under the Apache License 2.0 (see LICENSE), for both research and commercial use.
- Downloads last month
- 2
