Scenema Audio


Sample Audio - Honey Buns (30s)

  • ~3.3 min on an RTX 5090 (32 GB)

Components used to generate

Component File ~Size Download
Transformer scenema-audio-transformer-int8.safetensors 4.91 GB Link
Text Encoder gemma-3-12b-it-Q4_K_M.gguf 7.3 GB Link
Audio Pipeline VAE scenema-audio-pipeline.safetensors 6.71 GB Link
Audio Encoder VAE scenema-audio-vae-encoder.safetensors 42.7 MB Link
Extras folder scenema-audio/extras 2.3 GB Link

âš  scenema-audio extras is required for longer audio, and for voice-to-voice âš 

  • Download extras folder Link and place here: /ComfyUI/models/scenema-audio/extras
  • Create the scenema-audio/extras folder if it doesnt exist

Must have:

  • /scenema-audio/extras/mel-band-roformer
  • /scenema-audio/extras/bigvgan
  • /scenema-audio/extras/campplus
  • /scenema-audio/extras/seedvc
  • /scenema-audio/extras/whisper-small

âš  Must use the updated GGUF Loader âš 

In comfy:

  • Open the ComfyUI Manager
  • Change the channel to "Channel (remote)"
  • and search for comfyui-gguf-loader

image

Use the workflow Link

image


Sources

Downloads last month
-
GGUF
Model size
12B params
Architecture
gemma3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ChrisColeTech/scenema-audio

Quantized
(1)
this model

Collection including ChrisColeTech/scenema-audio