ONNX
Safetensors
music
midi
cc1

instrument-aire-models

1. Model introduction

AIRE is a lightweight model family for generating expressive MIDI control curves from musical notes, designed for practical local inference in a web browser. It combines note-derived frame features with musical conditioning to predict a time-aligned expression curve while preserving the original notes.

The current A02b release generates CC1 only, including movement within sustained notes; it does not generate notes or predict note velocity. CC1's audible effect depends on the receiving instrument and its controller mapping.

Model Scope Files
S03 Shared orchestral strings: violin, viola, cello, double bass safetensors · ONNX
B01 Brass family safetensors · ONNX
W01 Woodwind family safetensors · ONNX
U01 One model conditioned on string, woodwind or brass safetensors · ONNX

2. Architecture, inputs and outputs

A02b model architecture

Vector architecture diagram.

A02b uses a width-96 backbone with two kernel-5 convolutions, four bidirectional encoder blocks, four attention heads and a 384-wide feed-forward layer. Frame embeddings and projections are added, together with a broadcast global control projection. Global conditioning combines mood (16-D), track role (8-D), instrument family (8-D) and six numerical controls through a 38→64→64 MLP. Independent FiLM modules condition attention and feed-forward paths. All encoder layers share a learned signed relative-attention bias table. A LayerNorm, linear 96→1 and sigmoid produce one normalized CC1 value per frame.

The note-level input uses quarter-note beats, MIDI pitch 0–127 and sounding velocity 1–127. Tempo and meter come from the input MIDI; family, mood, track role and CC bounds are supplied controls. Notes are encoded at 16 frames per beat, with a one-beat tail after the final note. Output events contain beat, continuous value, and half-up rounded integer cc1. Decode normalized predictions as cc_min + y * (cc_max - cc_min).

Actual four-bar strings example

This example uses C–Am–F–G (I–vi–IV–V) in 4/4 at 120 BPM, with each triad sustained for one bar. pad is the track role; the mood is peaceful. The family is string, bounds are [0,127], and all original note velocities remain 64.

Bar Beats Chord MIDI pitches
1 0–4 C 60, 64, 67
2 4–8 Am 57, 60, 64
3 8–12 F 53, 57, 60
4 12–16 G 55, 59, 62

Actual S03 CC1 inference for the four-bar example

The plot is the actual S03 inference output, not a hand-drawn target or smoothed curve. There are 272 frames: 16 beats of notes plus one beat of tail. Input MIDI · MIDI with generated CC1 · complete input JSON · complete inference output.

Input excerpt:

{
  "tempo_bpm": 120,
  "time_signature": [4, 4],
  "bar_origin_beats": 0,
  "controls": {
    "instrument": "string", "mood": "peaceful", "track_role": "pad",
    "cc_min": 0, "cc_max": 127, "use_velocity": false
  },
  "notes": [
    {"pitch": 60, "start": 0, "end": 4, "velocity": 64},
    {"pitch": 64, "start": 0, "end": 4, "velocity": 64},
    {"pitch": 67, "start": 0, "end": 4, "velocity": 64}
  ]
}

The nine notes of the remaining three bars are omitted above; use the linked complete input for reproduction. First six output frames:

[
  {
    "beat": 0.0,
    "value": 37.85149,
    "cc1": 38
  },
  {
    "beat": 0.0625,
    "value": 42.028435,
    "cc1": 42
  },
  {
    "beat": 0.125,
    "value": 40.26553,
    "cc1": 40
  },
  {
    "beat": 0.1875,
    "value": 40.282303,
    "cc1": 40
  },
  {
    "beat": 0.25,
    "value": 41.185257,
    "cc1": 41
  },
  {
    "beat": 0.3125,
    "value": 42.31284,
    "cc1": 42
  }
]

… remaining 266 frames omitted. Continuous values above are displayed to six decimal places; the full output retains the computed values. The example was run with the released S03 safetensors weights on CPU; ONNX in Python, JavaScript in Node, and browser WASM reproduce the same MIDI events for this example.

Low-level tensor interface

The scripts handle this encoding. Direct ONNX callers must supply all ten tensors, with batch size 1 and dynamic frame count T:

Input Type Shape
lead_pitch int64 [1,T]
pitch_roll, onset_roll, offset_roll float32 [1,T,128] each
frame_numeric float32 [1,T,25]
mood_id, role_id, instrument_id int64 [1] each
ctrl_numeric float32 [1,6]
mask bool [1,T]

Output: normalized_cc1, float32 [1,T]. Lead pitch 128 represents rest. The mask is a valid prefix followed by right padding; rests and the tail are valid frames. Preprocessing is implemented in Python and JavaScript, including the exact feature order and effective-velocity policy.

3. Relationship to MID-FiLD

Both approaches generate fine-level CC1 expression conditioned on supplied notes and musical metadata, including instrument, mood, track role and expression bounds. The training material comes from MID-FiLD, whose dataset and task formulation made this project possible.

Aspect MID-FiLD paper's demonstration model AIRE-A02b release
Representation Tokenized note events and CC1 position/value events Fixed frame grid with pitch rolls and numerical note features
Architecture Transformer encoder with an autoregressive decoder Convolutions and a bidirectional encoder with FiLM
Generation Autoregressive CC1 token sequence Whole-passage continuous CC1 predictions in one forward pass
Instrument scope Original dataset includes a broader instrument inventory Selected orchestral strings, brass and woodwind; family-level conditioning
Deployment goal Demonstrate learning expressive dynamics from the dataset Small FP32 models with ONNX and browser inference support

This is an independent architecture trained from fresh weights, rather than a reproduction of the paper's baseline or a fine-tune of its checkpoint. We do not claim to outperform MID-FiLD's original model; the evaluation protocols are not directly comparable. See the original paper for its model and tokenization details.

4. Usage and limits

Download this repository's scripts and the model files you need. Run the commands below from the repository root. Scripts accept a full model-file path, an input MIDI path and a new output MIDI path. They replace CC1 on the note channel while preserving the original notes, velocities, tracks, programs, tempo, meter and other events at their absolute ticks. End-of-track may extend to the one-beat tail. Existing output files are not overwritten.

Python

Python 3.12 is the tested version. Install dependencies in your own environment:

python -m pip install -r requirements.txt
python inference/infer.py \
  --model "/absolute/path/models/aire-strings-a02b-s03/model.safetensors" \
  --input "/absolute/path/input.mid" \
  --output "/absolute/path/output-cc1.mid" \
  --mood peaceful --role pad

Use model.onnx in the same command for native ONNX inference. The complete runnable Python script includes preprocessing, inference and MIDI serialization. Programmatic example, saved as run.py in the repository root:

from pathlib import Path
import sys
sys.path.insert(0, str(Path(__file__).resolve().parent / "inference"))
from infer import process_midi

model_path, input_midi, output_midi = sys.argv[1:4]
process_midi(model_path, input_midi, output_midi,
             model_id="S03", mood="peaceful", role="pad")

JavaScript / web

Install the pinned ONNX Runtime Web dependency, then run the Node command-line example. It uses the same WASM runtime and 2,048-frame cap as the browser helper:

npm install
node inference/infer.mjs \
  --model "/absolute/path/models/aire-strings-a02b-s03/model.onnx" \
  --input "/absolute/path/input.mid" \
  --output "/absolute/path/output-cc1.mid" \
  --mood peaceful --role pad

The complete runnable JavaScript script accepts those paths. Programmatic example, saved as run.mjs in the repository root:

import * as ort from "onnxruntime-web";
import {readFile, writeFile} from "node:fs/promises";
import {inferMidi} from "./inference/infer-core.mjs";

const [modelPath, inputMidi, outputMidi] = process.argv.slice(2);
const result = await inferMidi(
  ort, new Uint8Array(await readFile(modelPath)),
  new Uint8Array(await readFile(inputMidi)),
  {modelId: "S03", mood: "peaceful", role: "pad"}
);
await writeFile(outputMidi, result.midi, {flag: "wx"});

Browser applications can import inferMidi through a bundler and supply ONNX Runtime Web, model bytes and MIDI bytes from fetch() or a file picker. The returned Uint8Array can be downloaded as a MIDI Blob. Browsers cannot read arbitrary filesystem paths; absolute paths apply to the Node CLI. Serve the application over HTTP(S), keep runtime WASM files available, and run inference in a Worker to keep the UI responsive. The helper defaults to single-thread WASM; browser WebGPU support for the model graphs was validated separately, and application-level fallback must be handled by the caller.

Both CLIs default to peaceful / pad / [0,127]. Use --cc-min and --cc-max for explicit bounds, --disable-velocity to disable velocity conditioning, and --json-output for the decoded frame events. If the model path does not identify S03/B01/W01/U01, pass --model-id. U01 additionally requires --family string, --family woodwind or --family brass. The same option values apply to Python and JavaScript.

Limits

  • Python/native cap: 4,096 frames. JavaScript/web cap: 2,048 frames. For a nonempty request, T = ceil(16 × (last_note_end_beats + 1)); leading silence and the one-beat tail count toward the cap. No cropping, chunking or streaming is implemented. Native S03 ONNX was exercised at 4,096 frames; the public browser helper was exercised at 2,048 frames and rejects 2,049.
  • The MIDI examples support PPQ format 0/1, exactly one note-bearing track/channel, polyphonic chords, constant tempo and constant meter. They reject SMPTE timing, tempo/meter changes, unmatched or unclosed notes and overlapping identical pitches. Multi-voice inference is not implemented in these helpers.
  • CC1 output is voice-level, not per note or per chord tone. Original note velocity is preserved. Uniform input velocities disable velocity conditioning automatically. Mood, role and bounds are explicit controls; meaningful independent musical control has not been established.
  • These are experimental models. W01's full-range validation error exceeded its midpoint baseline; U01 did not improve the matched full-range errors of the separate family models. Final-test and perceptual quality remain unverified. A reference CC1 curve is not a unique correct musical answer.

5. Category order

All categorical IDs are zero based in the exact order listed. Category lookup is included in the inference helper configuration. Do not reuse one model's mood IDs for another model.

S03

Family IDs: string.

Mood IDs: dreamy, groovy, hopeful, inspiring, magical, peaceful, relaxing, romantic, sad, tense, uplifting, happy, scary.

B01

Family IDs: brass.

Mood IDs: dreamy, groovy, hopeful, inspiring, magical, peaceful, relaxing, romantic, sad, tense, uplifting, bouncy, happy, mysterious, scary.

W01

Family IDs: woodwind.

Mood IDs: dreamy, groovy, hopeful, inspiring, magical, peaceful, relaxing, romantic, sad, tense, uplifting, funny, mysterious, scary, tragicomic.

U01

Family IDs: string, woodwind, brass.

Mood IDs: dreamy, groovy, hopeful, inspiring, magical, peaceful, relaxing, romantic, sad, tense, uplifting, happy, scary, funny, mysterious, tragicomic, bouncy.

Track-role IDs, shared by all models: main_melody, pad, riff, sub_melody, accompaniment, bass.

About AIRE

AIRE explores compact neural models that turn MIDI notes into continuous expression, with an emphasis on preserving the supplied performance and making inference accessible on the web. A02b uses an independently designed convolutional and bidirectional encoder architecture with FiLM conditioning, trained from fresh weights to predict the full CC1 curve in one forward pass. The project is heavily inspired by MID-FiLD: MIDI Dataset for Fine-Level Dynamics (AAAI 2024), and the released models are trained on filtered subsets of its dataset, available in the official MID-FiLD GitHub repository.

6. License and required credit

Author: xiaohan-tian. Organization: KGAudioLab. URL: https://huggingface.co/KGAudioLab.

The models and accompanying code/documentation use MIT + Attribution Requirement; see LICENSE. The extra mandatory credit condition makes this a custom license based on MIT, rather than unmodified MIT. Commercial use is permitted subject to those terms.

If you use these models in a project, fine-tune or further train them, or retrain using these weights or adapting these models' architecture or implementation, credit xiaohan-tian / KGAudioLab and include https://huggingface.co/KGAudioLab in your project documentation or credits. Suggested wording:

Uses or is derived from instrument-aire models by xiaohan-tian / KGAudioLab (https://huggingface.co/KGAudioLab).

Please also acknowledge the original MID-FiLD paper and dataset when describing the origin of the training material. Its upstream repository provides the original dataset license and attribution.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support