JAY 1 β€” Male Vocal Presence & Lyrical Structure LoRA

A single-file LoRA adapter for ACE-Step v1.5 turbo that pushes a track toward a male-led, human-sounding vocal performance and a tighter lyric structure.

Small, cheap, and directional. It pushes four things and gets out of the way of your prompt.

Run it at LoRA scale 1.0 and put the prefix token first β€” see Strength and Trigger token.


What it does

Effect Description
Male vocal presence Roughly +10% toward a male vocal read. Also suppresses the female-vocal bleed that often shows up underneath when you don't ask for it, so a track stays decided instead of doubling.
Lyrical structure Roughly +20% to the legibility of the lyric structure β€” cleaner section boundaries, more predictable phrasing across verse/chorus.
Human soul & emotion Pushes the male vocalist away from the flat synthetic read toward something more breath-adjacent and emotionally legible.
Female leakage suppression A side effect of the above. Useful when your prompt says male but the base model keeps hedging.

On the +10% / +20% figures: these are the author's desk characterizations from A/B listening, not benchmark numbers. There is no objective "percent more male" to measure. Read them as an indication of direction and rough magnitude.


Compatibility

Built for: acestep-v15-turbo (2B)

Not compatible with: acestep-v15-xl-turbo (4B)

This LoRA predates the release of acestep-v15-xl-turbo, so it was trained against the original turbo weight layout and has not been ported. The two architectures differ in layer count and hidden size, so the adapter does not bind β€” that is a structural mismatch, not a version warning.

If you are on XL-turbo, use JAY 2 instead.


Files

adapter_model.safetensors
adapter_config.json

Weights + PEFT config, nothing else. This repo is inference-only β€” no dataset, no preprocessed tensors, no training script. It is not a post-training release.

⚠️ Renaming on download

ACE-Step only loads a LoRA whose files are named exactly adapter_model.safetensors and adapter_config.json. If you rename either file, the adapter will not be found and the UI will silently run without it.

If the LoRA does not seem to do anything: that is almost always a filename problem, not a strength problem. Rename the .safetensors to adapter_model.safetensors, put adapter_config.json in the same folder, and restart.


Usage

  1. Download both files into a single folder.
  2. Confirm the names match the filenames above exactly.
  3. Load the folder as your LoRA source in the ACE-Step UI / API.
  4. Set LoRA scale, put the prefix token first, generate.

Strength

1.0. Not a starting point β€” the only working value.

This adapter does not fade in gracefully. Below 1.0 it does not become "a subtler version of itself," it degrades into static. That is a measured result from A/B testing at multiple values, not a conservative default.

The cause is in the config. lora_alpha: 64 against r: 32 gives PEFT a scaling factor of 2.0, so effective strength is double the slider number. At slider 1.0 you get an effective 2.0, which is roughly what it takes to override the base model's prior and produce coherent audio. Turn it down and the signal never clears the base model's own output.

Consequence: this cannot be used as a subtle blend. If you want a whisper of JAY in something, it will not give you one. Use 1.0, or use a different adapter.

If you are loading it through a pipeline whose default scale is not 1.0, set it explicitly. A silent default of 0.8 is the single most likely reason someone reports this LoRA "does nothing."

Suggested starting settings

Setting Start at
LoRA scale 1.0 (required β€” see Strength)
Think on (compare both ways)
LM temperature default
LM CFG scale default

Prompts that explicitly specify male singer and a genre tend to benefit most. Prompts that are purely instrumental will load the adapter with nothing to apply it to.


Trigger token

guy_singing_runs

This is a prefix token. It goes at the very front of the prompt:

βœ…  guy_singing_runs, rock, driving guitars, raw emotional male vocalist
❌  rock, driving guitars, raw emotional male vocalist, guy_singing_runs

It conditions the prompt as a whole, so position matters β€” placed anywhere else the amplification is substantially reduced. Do not put it in the lyrics block.

You are not obliged to use it. The adapter still loads and still biases the output without it; the token is what drives the behaviors above. Omit it and JAY-1 sits inert underneath your prompt.

JAY-2 uses a different token. jay_singing_runs there, guy_singing_runs here. The old token does not trigger JAY-2 and fails silently.

Example caption

guy_singing_runs, Melodic British-style rock, polished but organic. Clean articulate lead
guitar, radio-friendly mid-tempo groove, tight rhythm guitar, supportive bass, crisp drums.
Male singer with a dark slightly raspy voice, deep baritone, gravelly rough-edged texture,
raw and understated, conversational storytelling phrasing.

Test lyrics

Structure tags matter here β€” this adapter is partly a structure adapter, so give it something to organize:

[Intro]

[Verse 1]
Sun is rising slowly, light upon the floor
Coffee's on the table, no alarms, no chores
Nothing on the schedule, nowhere we need to go

[Chorus]
Oh, Sundays feel like heaven, hearts are running free
The world can wait a little, we've got our own parade

[Verse 2]
Stories in the kitchen, songs drift through the air
The clock is just a number, the hours drift away

[Chorus]
Oh, Sundays feel like heaven, hearts are running free
The world can wait a little, we've got our own parade

[Outro]

Limitations

  • Vocal shift is a bias, not a lock. It will not reliably convert a prompt for a female vocalist into a male one, and it is not intended to.
  • The lyric-structure push is strongest with explicit [Verse] / [Chorus] / [Bridge] tagging. Freeform untagged lyrics get less of it.
  • Instrumental tracks get almost nothing from this adapter.
  • Only coherent at scale 1.0. No partial strength, no blending, no subtle use.
  • The prefix token must be first. Mid-prompt placement substantially reduces the effect.
  • Not compatible with acestep-v15-xl-turbo (see above).
  • Trained on a limited corpus by one person. Expect a narrower stylistic range than the base model.

License

CC0-1.0 (public domain dedication). Do anything you want with it β€” use, modify, redistribute, commercial or not, no attribution required. No warranty, no liability.

Base model: ACE-Step/Ace-Step1.5 β€” MIT. This adapter is an independent derivative and is not affiliated with or endorsed by the ACE-Step authors.


Author

str8bored@S-W-O-R-D

README Author

seven@S-W-O-R-D

Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for str8bored/JAY_1

Adapter
(23)
this model