- Audio Sliders for ACE-Step 1.5 XL turbo
- Listen
- What is in the repository
ace-step-1.5-xl-turbo/text: 12 named attributes, trained from a prompt pairace-step-1.5-xl-turbo/real-axes: axes found in real music, trained from the axis's own tagsace-step-1.5-xl-turbo/real-axes-sets: the same kind of axis, trained with no textace-step-1.5-xl-turbo/measured-sets: a measurement turned into a slider, with no textace-step-1.5-xl-turbo/graded: trained with no text at graded positions, usable from -2 to +2
- Use
- Limits
- License and credit
- Listen
Audio Sliders for ACE-Step 1.5 XL turbo
Thirty-five sliders for generated music. Each one is a rank-4 LoRA on the frozen ACE-Step 1.5 XL turbo transformer whose strength is a number you set at sampling time. The prompt and the seed stay fixed and the same piece moves along one axis: sad to happy, solo to full ensemble, energetic to quiet and dreamy, electronic to jazz.
Every slider here was measured on 24 held-out prompts: does a property of the waveform follow the slider, does a second embedding model agree, how much of the piece survives, and do two quality predictors still rate the result as music. The numbers below come from that evaluation. Sliders that failed it were left out.
Live demo Β· Code, results, and paper draft Β· Ten-minute listening test
Listen
One prompt ("mellow jazz piano trio, brushed drums") and one seed per slider; only the slider position changes. Clips are loudness-matched.
Mood (text): sad β happy
- position -2:
- unsteered:
- position +2:
Arousal (real-axes): energetic and aggressive β quiet and dreamy
- position -2:
- unsteered:
- position +1:
Electronic to jazz (real-axes): electronic β jazz
- position -2:
- unsteered:
- position +2:
Arousal (real-axes-sets): energetic and aggressive β quiet and dreamy
- position -1:
- unsteered:
- position +1:
Harmony (measured-sets): simple harmony β rich harmony
- position -1:
- unsteered:
- position +1:
What is in the repository
ace-step-1.5-xl-turbo/text: 12 named attributes, trained from a prompt pair
| Slider | Low β high | Waveform descriptor follows (Ο) | MuQ-MuLan agrees (Ο) | Usable positions | Piece kept | Enjoyment at the ends (6.95 unsteered) |
|---|---|---|---|---|---|---|
mood |
sad β happy | 0.39 | 0.88 | -1 to +2 | 0.79 | 6.65 |
ensemble |
solo β full ensemble | 0.60 | 0.77 | -1.5 to +2 | 0.80 | 6.56 |
groove |
stiff β groovy | none assigned | 0.56 | -1.5 to +2 | 0.79 | 6.88 |
harmony |
simple harmony β rich harmony | 0.41 | 0.73 | -1.5 to +2 | 0.83 | 6.64 |
melody |
texture β melody | 0.48 | 0.89 | -1 to +2 | 0.80 | 6.47 |
tension |
relaxed β tense | none assigned | 0.86 | -1.5 to +1 | 0.83 | 5.64 |
brightness |
dark β bright | 0.97 | 0.80 | -1 to +2 | 0.74 | 6.04 |
density |
sparse β dense | 0.78 | 0.86 | -0.5 to +2 | 0.84 | 6.09 |
energy |
calm β intense | 0.91 | 0.90 | -1.5 to +1.5 | 0.64 | 6.42 |
tempo |
slow β fast | 0.55 | 0.77 | -1 to +2 | 0.68 | 6.38 |
electronic |
acoustic β electronic | none assigned | 0.79 | -2 to +2 | 0.65 | 6.79 |
vintage |
modern β vintage | none assigned | 0.79 | -2 to +2 | 0.79 | 6.68 |
ace-step-1.5-xl-turbo/real-axes: axes found in real music, trained from the axis's own tags
The axes are independent components of MuQ-MuLan embeddings of 14,985 real recordings. Nobody chose them. The prompt pair for each slider is the set of tags at the two ends of its axis, and the output is scored by its projection on the axis, in standard deviations of real music.
| Slider | Low β high | Follows the axis (Ο) | Ends ordered | Moved between -1 and +1 (std of real music) | Usable positions | Piece kept at Β±1 |
|---|---|---|---|---|---|---|
arousal |
energetic and aggressive β quiet and dreamy | 0.91 | 100% | 1.85 | -2 to +1 | 0.80 |
jazz_electronic |
electronic β jazz | 0.86 | 97% | 1.86 | -2 to +2 | 0.82 |
piano_axis |
organ and rock β piano | 0.78 | 97% | 1.56 | -2 to +1 | 0.80 |
strings_synth |
synthesizer and choir β guitars and strings | 0.80 | 97% | 1.78 | -0.5 to +2 | 0.73 |
valence |
playful and happy β dark and distorted | 0.42 | 75% | 0.86 | -2 to +1 | 0.85 |
ace-step-1.5-xl-turbo/real-axes-sets: the same kind of axis, trained with no text
Trained between the top and bottom 20% of the model's own clips along the axis (15,552 clips, 648 prompts). They move less than the prompt-pair versions, keep more of the piece, and hold quality level at both ends. Use them between -1 and +1.
| Slider | Low β high | Follows the axis (Ο) | Ends ordered | Moved between -1 and +1 (std of real music) | Piece kept at Β±1 |
|---|---|---|---|---|---|
acoustic_electronic |
studio electronic β acoustic and blues | 0.55 | 93% | 0.64 | 0.86 |
arousal |
energetic and aggressive β quiet and dreamy | 0.82 | 99% | 1.19 | 0.84 |
classical_funk |
funk and drums β classical and epic | 0.64 | 96% | 0.77 | 0.85 |
jazz_electronic |
electronic β jazz | 0.61 | 94% | 0.61 | 0.87 |
piano_axis |
organ and rock β piano | 0.63 | 96% | 0.75 | 0.87 |
strings_synth |
synthesizer and choir β guitars and strings | 0.40 | 82% | 0.44 | 0.87 |
valence |
playful and happy β dark and distorted | 0.50 | 93% | 0.62 | 0.87 |
ace-step-1.5-xl-turbo/measured-sets: a measurement turned into a slider, with no text
Clips sorted by a measurement, after removing what loudness, brightness, and predicted enjoyment explain. These are the selective sliders: selectivity is how far a slider moves its own measurement relative to the average bystander. Use them between -1 and +1.
| Slider | Sorted by | Follows the measurement (Ο) | Ends ordered | Selectivity | Piece kept at Β±1 |
|---|---|---|---|---|---|
energy |
flux | 0.83 | 99% | 3.8 | 0.85 |
harmony |
harmonic change | 0.72 | 99% | 6.6 | 0.88 |
density |
onset rate | 0.63 | 90% | 4.7 | 0.87 |
ensemble |
pc | 0.60 | 94% | 2.4 | 0.86 |
quality |
ce | 0.53 | 86% | 3.7 | 0.86 |
ace-step-1.5-xl-turbo/graded: trained with no text at graded positions, usable from -2 to +2
The set trainer above shows the slider only positions -1 and +1, and its sliders fall apart past Β±1. These six were trained with every clip at its own position along the measurement or axis. They move slightly less inside Β±1 and keep working out to Β±2 with no loss on either quality predictor.
| Slider | Low β high | Follows it, -2 to +2 (Ο) | Ends ordered | Piece kept at Β±2 | Enjoyment at the ends |
|---|---|---|---|---|---|
energy |
calm β intense | 0.89 | 100% | 0.81 | 7.15 |
harmony |
simple harmony β rich harmony | 0.60 | 92% | 0.85 | 7.11 |
arousal |
energetic and aggressive β quiet and dreamy | 0.84 | 96% | 0.82 | 7.20 |
valence |
playful and happy β dark and distorted | 0.50 | 82% | 0.85 | 7.13 |
jazz_electronic |
electronic β jazz | 0.62 | 93% | 0.85 | 7.14 |
piano_axis |
organ and rock β piano | 0.64 | 89% | 0.84 | 7.19 |
axes/ holds the direction vectors of the real-music axes, so new clips can be scored against them.
Use
pip install "audiosliders[model,demo] @ git+https://github.com/takakhoo/audio-diffusion-control" audiobox_aesthetics
A page where you type a prompt and drag the sliders:
python -m audiosliders.server --sliders hf:ace-step-1.5-xl-turbo/text --backbone ace-turbo
From Python:
import soundfile as sf
from audiosliders.backbone import load_backbone
from audiosliders.hub import fetch
from audiosliders.lora import SliderBank
model = load_backbone("ace-turbo") # downloads ACE-Step 1.5 XL turbo
bank = SliderBank(model.dit)
folder = fetch("ace-step-1.5-xl-turbo/real-axes")
bank.load("arousal", folder / "arousal.safetensors")
for position in (-2.0, 0.0, 1.0):
audio = model.generate(["mellow jazz piano trio, brushed drums"], [0], seconds=10,
wrap=lambda p: bank.gated(p, {"arousal": position}))
sf.write(f"arousal_{position:+.0f}.wav", audio[0].T.cpu().numpy(), model.sample_rate)
Position 0 is the unchanged model. Several sliders can be loaded and set together; their updates add.
Limits
- No listening study has been run yet. Quality rests on two learned predictors (Audiobox Aesthetics and SongEval) and meaning on signal descriptors and two embedding models.
- Sliders leak. Mood and melody also brighten the clip, and several raise loudness. The leakage table is in the repository.
- The set-trained sliders stop working past Β±1. A set-trained tempo slider and a tag-sorted mood slider did not work and are not included.
- Trained and tested on ten-second instrumental clips. Six of the text sliders were also checked at thirty seconds.
- The files use this project's own LoRA layout and are loaded with
audiosliders.lora.SliderBank.
License and credit
MIT, the same as ACE-Step 1.5. The training objective for the prompt-pair sliders is Concept Sliders (Gandikota et al., ECCV 2024), and the idea of discovering axes instead of naming them comes from SliderSpace (Gandikota et al., ICCV 2025).
Model tree for takakhoo/audio-sliders
Base model
ACE-Step/acestep-v15-xl-turbo