Audio Sliders for ACE-Step 1.5 XL turbo

Thirty-five sliders for generated music. Each one is a rank-4 LoRA on the frozen ACE-Step 1.5 XL turbo transformer whose strength is a number you set at sampling time. The prompt and the seed stay fixed and the same piece moves along one axis: sad to happy, solo to full ensemble, energetic to quiet and dreamy, electronic to jazz.

Every slider here was measured on 24 held-out prompts: does a property of the waveform follow the slider, does a second embedding model agree, how much of the piece survives, and do two quality predictors still rate the result as music. The numbers below come from that evaluation. Sliders that failed it were left out.

Live demo Β· Code, results, and paper draft Β· Ten-minute listening test

Four sliders sweeping across their range

Listen

One prompt ("mellow jazz piano trio, brushed drums") and one seed per slider; only the slider position changes. Clips are loudness-matched.

Mood (text): sad β†’ happy

  • position -2:
  • unsteered:
  • position +2:

Arousal (real-axes): energetic and aggressive β†’ quiet and dreamy

  • position -2:
  • unsteered:
  • position +1:

Electronic to jazz (real-axes): electronic β†’ jazz

  • position -2:
  • unsteered:
  • position +2:

Arousal (real-axes-sets): energetic and aggressive β†’ quiet and dreamy

  • position -1:
  • unsteered:
  • position +1:

Harmony (measured-sets): simple harmony β†’ rich harmony

  • position -1:
  • unsteered:
  • position +1:

What is in the repository

ace-step-1.5-xl-turbo/text: 12 named attributes, trained from a prompt pair

Slider Low β†’ high Waveform descriptor follows (ρ) MuQ-MuLan agrees (ρ) Usable positions Piece kept Enjoyment at the ends (6.95 unsteered)
mood sad β†’ happy 0.39 0.88 -1 to +2 0.79 6.65
ensemble solo β†’ full ensemble 0.60 0.77 -1.5 to +2 0.80 6.56
groove stiff β†’ groovy none assigned 0.56 -1.5 to +2 0.79 6.88
harmony simple harmony β†’ rich harmony 0.41 0.73 -1.5 to +2 0.83 6.64
melody texture β†’ melody 0.48 0.89 -1 to +2 0.80 6.47
tension relaxed β†’ tense none assigned 0.86 -1.5 to +1 0.83 5.64
brightness dark β†’ bright 0.97 0.80 -1 to +2 0.74 6.04
density sparse β†’ dense 0.78 0.86 -0.5 to +2 0.84 6.09
energy calm β†’ intense 0.91 0.90 -1.5 to +1.5 0.64 6.42
tempo slow β†’ fast 0.55 0.77 -1 to +2 0.68 6.38
electronic acoustic β†’ electronic none assigned 0.79 -2 to +2 0.65 6.79
vintage modern β†’ vintage none assigned 0.79 -2 to +2 0.79 6.68

ace-step-1.5-xl-turbo/real-axes: axes found in real music, trained from the axis's own tags

The axes are independent components of MuQ-MuLan embeddings of 14,985 real recordings. Nobody chose them. The prompt pair for each slider is the set of tags at the two ends of its axis, and the output is scored by its projection on the axis, in standard deviations of real music.

Slider Low β†’ high Follows the axis (ρ) Ends ordered Moved between -1 and +1 (std of real music) Usable positions Piece kept at Β±1
arousal energetic and aggressive β†’ quiet and dreamy 0.91 100% 1.85 -2 to +1 0.80
jazz_electronic electronic β†’ jazz 0.86 97% 1.86 -2 to +2 0.82
piano_axis organ and rock β†’ piano 0.78 97% 1.56 -2 to +1 0.80
strings_synth synthesizer and choir β†’ guitars and strings 0.80 97% 1.78 -0.5 to +2 0.73
valence playful and happy β†’ dark and distorted 0.42 75% 0.86 -2 to +1 0.85

ace-step-1.5-xl-turbo/real-axes-sets: the same kind of axis, trained with no text

Trained between the top and bottom 20% of the model's own clips along the axis (15,552 clips, 648 prompts). They move less than the prompt-pair versions, keep more of the piece, and hold quality level at both ends. Use them between -1 and +1.

Slider Low β†’ high Follows the axis (ρ) Ends ordered Moved between -1 and +1 (std of real music) Piece kept at Β±1
acoustic_electronic studio electronic β†’ acoustic and blues 0.55 93% 0.64 0.86
arousal energetic and aggressive β†’ quiet and dreamy 0.82 99% 1.19 0.84
classical_funk funk and drums β†’ classical and epic 0.64 96% 0.77 0.85
jazz_electronic electronic β†’ jazz 0.61 94% 0.61 0.87
piano_axis organ and rock β†’ piano 0.63 96% 0.75 0.87
strings_synth synthesizer and choir β†’ guitars and strings 0.40 82% 0.44 0.87
valence playful and happy β†’ dark and distorted 0.50 93% 0.62 0.87

ace-step-1.5-xl-turbo/measured-sets: a measurement turned into a slider, with no text

Clips sorted by a measurement, after removing what loudness, brightness, and predicted enjoyment explain. These are the selective sliders: selectivity is how far a slider moves its own measurement relative to the average bystander. Use them between -1 and +1.

Slider Sorted by Follows the measurement (ρ) Ends ordered Selectivity Piece kept at ±1
energy flux 0.83 99% 3.8 0.85
harmony harmonic change 0.72 99% 6.6 0.88
density onset rate 0.63 90% 4.7 0.87
ensemble pc 0.60 94% 2.4 0.86
quality ce 0.53 86% 3.7 0.86

ace-step-1.5-xl-turbo/graded: trained with no text at graded positions, usable from -2 to +2

The set trainer above shows the slider only positions -1 and +1, and its sliders fall apart past Β±1. These six were trained with every clip at its own position along the measurement or axis. They move slightly less inside Β±1 and keep working out to Β±2 with no loss on either quality predictor.

Slider Low β†’ high Follows it, -2 to +2 (ρ) Ends ordered Piece kept at Β±2 Enjoyment at the ends
energy calm β†’ intense 0.89 100% 0.81 7.15
harmony simple harmony β†’ rich harmony 0.60 92% 0.85 7.11
arousal energetic and aggressive β†’ quiet and dreamy 0.84 96% 0.82 7.20
valence playful and happy β†’ dark and distorted 0.50 82% 0.85 7.13
jazz_electronic electronic β†’ jazz 0.62 93% 0.85 7.14
piano_axis organ and rock β†’ piano 0.64 89% 0.84 7.19

axes/ holds the direction vectors of the real-music axes, so new clips can be scored against them.

Use

pip install "audiosliders[model,demo] @ git+https://github.com/takakhoo/audio-diffusion-control" audiobox_aesthetics

A page where you type a prompt and drag the sliders:

python -m audiosliders.server --sliders hf:ace-step-1.5-xl-turbo/text --backbone ace-turbo

From Python:

import soundfile as sf
from audiosliders.backbone import load_backbone
from audiosliders.hub import fetch
from audiosliders.lora import SliderBank

model = load_backbone("ace-turbo")                    # downloads ACE-Step 1.5 XL turbo
bank = SliderBank(model.dit)
folder = fetch("ace-step-1.5-xl-turbo/real-axes")
bank.load("arousal", folder / "arousal.safetensors")

for position in (-2.0, 0.0, 1.0):
    audio = model.generate(["mellow jazz piano trio, brushed drums"], [0], seconds=10,
                           wrap=lambda p: bank.gated(p, {"arousal": position}))
    sf.write(f"arousal_{position:+.0f}.wav", audio[0].T.cpu().numpy(), model.sample_rate)

Position 0 is the unchanged model. Several sliders can be loaded and set together; their updates add.

Limits

  • No listening study has been run yet. Quality rests on two learned predictors (Audiobox Aesthetics and SongEval) and meaning on signal descriptors and two embedding models.
  • Sliders leak. Mood and melody also brighten the clip, and several raise loudness. The leakage table is in the repository.
  • The set-trained sliders stop working past Β±1. A set-trained tempo slider and a tag-sorted mood slider did not work and are not included.
  • Trained and tested on ten-second instrumental clips. Six of the text sliders were also checked at thirty seconds.
  • The files use this project's own LoRA layout and are loaded with audiosliders.lora.SliderBank.

License and credit

MIT, the same as ACE-Step 1.5. The training objective for the prompt-pair sliders is Concept Sliders (Gandikota et al., ECCV 2024), and the idea of discovering axes instead of naming them comes from SliderSpace (Gandikota et al., ICCV 2025).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for takakhoo/audio-sliders