H3-LongVideos
Make long (up to ~120s) MiniMax-H3 video + synchronised audio from a single prompt, in ComfyUI. Self-contained β it uses only ComfyUI core's H3 support.
H3 renders one shot at a time. This node turns a written scene into a chain of shots: it splits your prompt into beats, sizes each shot to what that beat actually stages, chains every shot from the previous one's last frame, and keeps your characters, their clothing and your props consistent from shot to shot β things that otherwise drift, duplicate or quietly reset at every shot boundary.
One node covers both H3 conditioning tasks: FL2VA (a frame anchors the shot) and REF2VA (reference images say what a character looks like).
Install
Copy this folder into ComfyUI/custom_nodes/ and restart the ComfyUI server
(not just a browser refresh).
Quick start
UNETLoader ββ images β> Video Combine
CLIPLoader ββΌβ> H3 Long Videos β> audio ββ
VAELoader βββ latent β> (optional) latent post-processing
The soundscape output carries the ambient bed the shots actually used β the
one auto_soundscape derived from your scene, or your own text when it didn't fire.
Wire it to a text preview to read what it built, or straight back into the
global_soundscape input to pin it and stop it re-deriving.
The latent output carries the sampled latents, joined on the time axis, for
things like a latent upscaler. It is emitted as well as images, never instead:
the shot chain hands each shot the previous one's decoded last frame, so decoding
cannot be deferred.
It is not the latent form of images on a multi-shot run. trim_seam and
handoff_offset cut decoded frames, and H3 compresses time β one pixel frame is
not one latent step β so those cuts have no exact latent equivalent and the seam
frames are still present. On a single-shot run nothing trims and it matches
exactly. info says which you got.
You write four things:
1. The prompt β the first paragraph is the anchor (scene and style, kept on every shot); each later paragraph is one beat, and one beat is one shot.
The text fields are input sockets, not boxes on the node.
prompt,character_memory,anchor_override,global_soundscape,non_diegetic_music,exposed_termsandintro_textall take a connected multiline text node, so the same prose can feed several samplers and be edited in one place.promptis required: leave it unconnected and the graph errors rather than rendering blank. The other six are optional and behave as empty when nothing is attached.
shot_secondsis a socket too β wire H3 Shot Length into it, which also reports the matching frame count on the 17k+5 grid. Left unconnected it falls back to auto (the largest shot that fits at the chosen size), exactly as a0in the old widget did.
Natural daylight, hard sun and deep shadow. Shallow depth of field, background
falling soft. Fine grain, slight motion blur, neutral colour. A farm with a barn.
Dom drives a van down the driveway and stops in front of the barn.
Dom gets out and walks to the back of it.
Mara steps out of the barn and asks him: "Is that the last one?"
2. character_memory β who is in it and what they wear. This is the only
channel that can change mid-chain:
Dom = he, tall, 35, brunette, white t-shirt, blue jeans, work boots
Mara = she, 30, red hair, grey coat, black jeans
3. resolution + megapixels β the dropdown picks the shape, the number
picks the size. They are independent: changing aspect ratio does not change
cost. 1.0 = 1024Γ1024 worth of pixels (ComfyUI's own convention), and at 1.00
every ratio lands on H3's native size. Step down for speed, VRAM and longer shots.
4. shot_seconds β a ceiling, not the length of every shot. Wire
H3 Shot Length into it, or leave it unconnected to let the VRAM budget decide.
Set plan_only to preview the shot split, lengths and every warning without
rendering. Do that first; it is near-instant.
What it handles for you
Beats β shots. One paragraph, one shot. Nothing can silently collapse them.
Pacing. Each shot is sized from what its beat stages (~2s + ~2.5s per action clause, or its spoken line). A 3-second action in a 12-second shot is how a model ends up repeating or reversing the action.
Characters. Descriptions bind once per shot, at the first mention; repeat names collapse to pronouns, because naming someone twice renders them twice.
Wardrobe. Clothing lives in one mutable channel, tracked per person. Removals are read from your prose ("takes off her jacket", "steps out of her jeans", "the coat falls to the ground") and stated with direction so they don't play in reverse. A garment named in a quoted line is an instruction, not an action, so asking for something to come off doesn't remove it a shot early. Whatever is still on underneath is named, so a removal doesn't read as more than it was.
Props. "the van" in a later shot means the van from the earlier one.
Restraints stay on.
lock_restraints(on by default) keeps handcuffs, shackles, manacles, fetters, irons, gags, blindfolds, harnesses and leashes β plus qualified forms likeankle chainorleather wrist strapsβ from being removed by prose. A restraint is a plot state, not a garment. Without this they came off by accident: "steps out of her jacket and the chain falls away" would drop the ankle chain as a side effect of a beat about a jacket, because the removal window reaches any tracked item near the cue. To take one off, say so directly:wardrobe: Mara -= handcuffs. Barechain,collar,strapandbeltare not treated as restraints β they are jewellery, a shirt part, a dress part and a garment at least as often. It also states what the restraint does: a cuffed character otherwise walks with their arms swinging, because nothing said the body could not move freely β the restraint present and inert, which reads as it having broken. The clause names the bound region positively (the wrists stay bound close together, the arms moving as one), only for people actually in the shot, and it disappears the moment the restraint is removed. The wording follows how the restraint holds: a character cuffed to a headboard getsthe cuffs stay locked closed around the wrists and fastened to the headboard, the chain between them tautinstead of the bound-together text β two contradictory sentences about one pair of wrists is exactly how the cuffs end up rendered broken. Poses are covered too (behind the back, above the head, spread-eagle β wrists bound apart at fixed points), and every variant adds that the hardware itself stays whole: an open cuff or a snapped link mid-struggle was otherwise free to happen. And because a restraint is a plot state, how it is used persists: state it once ("cuffed to the headboard") and every later shot keeps that wording even when its own prose only says "she strains" β without this those shots fell back to the bound-together text and contradicted the attachment all over again. Restating updates it ("cuffed to the wall instead"), and freeing the character (wardrobe: Mara -= handcuffs) forgets it, so re-cuffed later they start fresh.Uncovered zones. The node tracks two body zones,
lowerandupper. When a removal leaves one with nothing on it, it keeps that state stated in every later shot until something covers the zone again β because deleting a garment is only a silence, and a video model's default is a clothed person, so silence puts the clothes back on a shot or two later.It also states that a bared zone stays bared as the body turns β the same from the front, the side and behind. The marker says the zone is bare; nothing said it held once the body presented a surface the shot had not shown yet, and an undescribed surface defaults to a clothed one, so the garment came back mid-shot on a turn. That clause names no garment and no person: naming the garment puts it back in the prompt, and naming the person a second time renders them twice.
exposed_termsis where you choose the wording, per character. Same syntax as the sheet: a pronoun sets it for everyone who declares that pronoun, a name overrides one person, and a trailinguppertargets that zone instead of the defaultlower. Anything after the=is passed through verbatim, so LoRA trigger words ride along:she = <wording for the lower zone> he = <wording for the lower zone>, <lora trigger> Mara upper = <wording for Mara's upper zone>Left unset, the node uses its own neutral wording, matched to the character's declared pronoun. A key that matches no character and no pronoun is reported in
inforather than silently doing nothing β which is what a mistyped name, or an object form likeherinstead ofshe, would otherwise do.A character can also start with a zone uncovered rather than arriving there through a removal β add
nude(ornaked,undressed,unclothed) for both zones,toplessorbottomlessfor one, to theircharacter_memory, and the wording applies from shot 1. This has to be written explicitly: a sheet that simply doesn't list clothes (Jon = he, 35, bald) is read as under-specified, never as a declaration.Configuring any of this is the intent, so it overrides
prevent_nudityβ no second switch to remember. The shot after a removal also starts fresh, without the handoff frame, because continuing from a frame that still shows the garment is how it comes back: a picture outvotes the sentence.prevent_nudity. On by default: the prompt never asserts that a body is uncovered. Removals still happen β what is gated is the sentence, and since the model's default is a clothed person, it covers what nobody described.infostill reports any zone a removal left uncovered, so you find out either way.Shift is yours to set. Keep
shift_video/shift_audioat 12 / 3. There was anauto_shiftoption here that lowered the shift to match a low step count. It has been removed, because its premise was wrong. It read H3's 12/3 defaults as putting "80% of the denoising into the final step" at 4 steps and flattened the schedule to spread that out β but a 4-step distill LoRA is trained to jump from ~0.80 noise straight to clean. That concentration is the distilled behaviour, not a fault, and lowering the shift puts every step at noise levels the LoRA never saw, which shows up as artifacting. None of the turbo/lightx2v LoRAs declares a schedule in its metadata either, so there was nothing to look up and the number was a guess.Whatever you set is passed through untouched. Keep
shift_audioatshift_video / 4if you do change it βaudio_scaleis that ratio, and flattening it toward 1.0 breaks the audio branch.infostill warns if the ratio drifts.Anatomy.
anatomy_guardstates each person's limb count β one head, two arms, two hands with five fingers, two legs with two feet β then pins every limb to its body and gives the skeleton a layout: each arm at one shoulder running shoulderβelbowβwristβhand, each leg at one hip running hipβkneeβankleβfoot, the parts stacked in order (head on the neck, neck on the shoulders, arms along the sides of the torso, legs under the hips), every limb moving only with the person it belongs to, and one groin between the legs. A negative prompt cannot do this: H3 is CFG-free atcfg 1, so the negative is never evaluated and "extra limbs" there does nothing. Naming the number gives the model a target; negating one only puts the word in the prompt. Never added to the anchor, and never to a shot with nobody in it β describing a body in an empty frame is what burned faces into opening frames before.auto= on below a 768 short edge, when a LoRA is applied, or on any shot holding two or more people β spare limbs are grown where bodies meet and move together, whatever the resolution.Solidity.
solidity_guardstops bodies passing through objects. Same constraint as the anatomy guard, and the same solution: the negative is never evaluated, and "does not walk through the wall" can't go in the positive either β it names walking through a wall, and a mention is a presence cue. So it states what bodies do: stop at the surface, rest on the floor, press against what they touch, walk around the furniture. Then it names the solid things this shot established β up to three, the beat's own first, so "Mara climbs the stairs" leads with the stairs rather than with set dressing from the anchor.auto(default) speaks only when the shot actually names something solid, reading both the beat and the identity block, since the set is usually described in the anchor.onstates it every shot. Only ever applied to a shot with someone in it β you need a body before it can pass through anything. Genuinely passable things (a curtain, smoke) are deliberately not claimed to be solid.Motion continuity.
motion_guardstops a pose being reached without the frames in between β a head arriving at a new angle with no path to it, the "neck snap". A snap is not a wrong pose; it is a right pose with nothing joining it to the last one, so the path is what gets stated: movement travels through every position on the way, at one steady speed, the neck following the shoulders and the shoulders following the hips.autofires on a beat that actually moves someone (turns, looks, walks, leans, reaches β and the high-jerk ones: struggles, pulls, twists, writhes, where a limb most often arrives without its path); a beat where nobody changes orientation has no path to describe. A snap immediately after a cut is a different thing β that is the model leaving the keyframe pose, andhandoff_offsetis the lever there.Two bodies in contact.
contact_guardkeeps an arrangement correctly aligned β any arrangement. It names none: the model already knows more position names than a list could hold, and what it gets wrong is the geometry. So the geometry is stated, and it holds for every case:ownership each person keeps their own head, two arms and two legs, each joined to the body it belongs to β overlapping bodies is exactly when a limb gets reassigned to the wrong torso separation they meet at the surface of the skin, each keeping its own volume, rather than passing into one another stable roles above stays above, below stays below, behind stays behind, for the whole shot and from every camera angle support weight rests on whatever is holding it, and the two stay in proportion Needs two people in the shot β one body cannot be misaligned against another, and saying otherwise in a one-person shot would invite the second in.
autofires on a contact cue in the beat;onstates it whenever two or more are present.This holds a stated arrangement together; it cannot infer one you did not state. Describe the arrangement in relative terms β who is above, behind, facing whom, what carries the weight β rather than by a position name alone, and the guard keeps it held.
Soundscape from the scene.
auto_soundscapebuilds the ambient bed from your prompt instead of you typing one. It reads the anchor β the soundscape is global, stamped on every shot, so it must describe the place, not one beat's action β falling back to the beats when the anchor is pure camera language.A disused aircraft hangar -> cavernous interior, long reverb, distant metal ticks Rain on the windows. A kitchen. -> steady rain, quiet room tone, faint appliance hum A rocky beach with waves -> gusting wind, waves breaking, sea wind, distant gulls Cinematic, shallow depth of field -> (nothing β that is a lens, not a meadow)Weather layers before place. No human sound is ever generated β no chatter, crowd or announcements, even for a bar or a station β because an ambient bed that implies voices is how H3 starts talking.
fill if blank(default) leaves anything you typed alone;alwaysoverrides it and says so ininfo.Silence, in three layers. A prompt clause alone was never enough, because two of the three causes aren't text.
- Text β beats with no quoted dialogue get a lips-closed clause and a no-voice soundscape.
- Picture β a dialogue shot handing its last frame to a silent shot seeds an open mouth mid-word, and a picture outvotes a sentence. The handoff frame is taken 3 frames (~125 ms) earlier at exactly that boundary, automatically.
- Audio β H3 is a joint model: the mouth follows the audio branch. On a
shot with no line that branch is otherwise unconditioned, invents a voice, and
the picture lip-syncs to it. The shot's audio channel is anchored to encoded
silence instead. That applies to every silent shot β including the first
one, which has no handoff, and reference-conditioned shots, which take an
audio-only keyframe to carry it. A shot with a
ref_imagewired is not exempt.
mute_nonspeech_audiois a fourth, weaker thing: it zeroes the waveform after generation, so it silences the track but cannot close a mouth.Inside a dialogue shot the silence is per person: a quoted line used to free every mouth in frame, so whoever else was on screen mouthed along with lines they never say β characters visibly reciting text nobody gave them. Spoken lines are now attributed to whoever introduced them (
Jon says: "...", or"..." said Jon), and everyone else in the shot gets the lips-closed state by pronoun. A quote that can't be attributed frees nobody by guesswork, and a scare-quoted word like she gave him a "look" β which is emphasis, not dialogue β is reported ininfo, since it still flips the whole shot to speaking.Non-speech vocals (screams, sobs, gasps).
allow_nonspeech_vocalslets beats with no quoted dialogue carry distress sounds. When it is on, the node skips the lips-closed clause and softens the no-voice soundscape so it bans speech, dialogue and singing but permits screams, sobs, gasps and moans. The audio branch is also left unmuted on those shots. Speech is still suppressed β only double-quoted lines count as speaking β so H3 does not invent chatter. Turn this on when your scene contains distress sounds the default silence would remove.One noise field for the whole chain.
vary_seed_per_shotis off by default. The seed picks the noise field a shot is sampled from, and that field fixes the stochastic detail β grain, micro-texture, the exact rendering of every surface the prompt never names. Reseed each shot and all of it resets at the boundary, which reads as a cut even when the keyframe anchors the frame and the location is unchanged. Shots still differ under one seed: each has its own beat text and its own handoff keyframe. Turn it on only when you want the beats to look separately shot.infowarns when it is on.Seams.
trim_seam(on by default) drops the first frame of each shot after the first, because that frame is the model's own reproduction of the handoff β the last frame of the previous shot. Keeping it plays the same moment twice. So at a working seam the two frames are not identical: they are one frame of normal motion apart. Turn it off for one run if you want to check how faithfully the anchor was reproduced.Latent upscale (optional pack).
latent_upscaleupscales each shot between sampling and decode, so the shot is sampled small and only decoded large. Cost scales with latent cells and attention is quadratic in them, so sampling 512Γ512 and upscaling 2Γ to 1024Γ1024 is roughly 6Γ cheaper than sampling 1024Γ1024 outright β the one lever that buys resolution instead of trading it. Wiring thelatentoutput to the same upscaler externally can't do this; by then the decode has already happened at the sampled size.Needs the separate Comfyui_Minimax_h3_latent_Upscaler pack and its weights in
models/latent_upscale_modelsβ model and nodes both by LBH-123-AI. It is not a dependency: without the pack the setting does nothing, the render proceeds at the sampled size, andinfosays so. Nothing errors. Only H3 builds are listed β the same folder holds LTX upscalers, whose channel count doesn't match. Spatial only, so the frame count and the audio are untouched, and tiled decode is forced on because decode memory grows with the square of the scale.The chain does not inherit the upscaler. Each shot hands the next one its last frame, so taking that frame from the upscaled decode would put a neural approximation and a downscale back to the sampling size into every boundary β compounding along the chain until the cast drifts. The handoff is decoded from the pre-upscale latent instead (a short tail, so it is cheap); only the shot's own output frames are upscaled. If that decode fails, the upscaled frames are used and the render carries on.
Overlays. Optional PIL watermark and intro title, composited after any upscale, never asked of the model.
info reports what it did and warns before you waste a render β thin beats,
dialogue that will be cut off or padded with invented speech, a removal that leaves
a body zone bare, anchor content that misfires on every shot.
Resolution and megapixels
Two widgets, and they do different jobs. resolution picks the shape,
megapixels picks the size.
megapixels is a pixel budget: 1.0 means 1024Γ1024 worth of pixels β
1,048,576 β the same convention as ComfyUI's own Scale Image to Total Pixels, so
the number means the same thing across your graph. The preset's aspect ratio is
kept and both axes are snapped to a multiple of 32, which is what H3's latent grid
requires. Set megapixels to 0 to switch it off and use the preset's own
dimensions verbatim.
Why a budget instead of a short edge
Cost and training-distribution match are functions of token count β
(h/16) Β· (w/16) Β· frames β which tracks total pixels. The short edge does not,
and the two disagree badly at the extremes of aspect ratio:
| preset | short edge | reads as | actual |
|---|---|---|---|
1:1 768x768 |
768 | native | 0.56 MP β 43% under budget |
21:9 1536x672 |
672 | sub-native | 0.98 MP β full budget |
So the square preset that looks native is starved, and the ultra-wide that looks starved is fine. Judging by short edge gets both backwards. Holding megapixels constant is what makes two aspect ratios genuinely comparable β VRAM and token count stay put when you change shape.
Start at 1.00, then step down
At 1.00MP every ratio reproduces H3's native dimensions, so it is the natural starting point. Lower budgets buy speed, VRAM headroom and longer shots β the shot-length budget is resolution-aware and rescales automatically.
| ratio | 0.44MP | 0.65MP | 1.00MP | 1.20MP |
|---|---|---|---|---|
16:9 |
896Γ512 | 1088Γ640 | 1344Γ768 | 1472Γ832 |
9:16 |
512Γ896 | 640Γ1088 | 768Γ1344 | 832Γ1472 |
4:3 |
800Γ576 | 960Γ704 | 1184Γ896 | 1280Γ960 |
3:4 |
576Γ800 | 704Γ960 | 896Γ1184 | 960Γ1280 |
1:1 |
672Γ672 | 832Γ832 | 1024Γ1024 | 1120Γ1120 |
21:9 |
1024Γ448 | 1248Γ544 | 1536Γ672 | 1696Γ736 |
9:21 |
448Γ1024 | 544Γ1248 | 672Γ1536 | 736Γ1696 |
Those columns are roughly the old fast / balanced / native tiers, which were
only ever three points on this axis. megapixels has no off-switch β a bare ratio
has no size to fall back to β and its floor is 0.10.
The ratio names are approximations
Worth knowing, because it explains why scaling works the way it does:
1344 / 768 = 1.750 -> 7:4 NOT 16:9, which is 1.778
1536 / 672 = 2.286 -> 16:7 NOT 21:9, which is 2.333
Scaling runs from each ratio's reference dimensions, not from the nominal ratio in its name. That is precisely what makes 1.00MP land exactly on 1344Γ768 rather than on 1376Γ768, which is where a true 16:9 at the same budget would put you.
What gets reported
info prints the size and MP actually produced, never what was requested.
Snapping to the 32-grid moves the real area β typically by 1β2%, up to about 4% at
the smallest budgets where a 32px step is a larger fraction of the image β and
echoing your input back would hide what the render used:
megapixels 1.00 -> 1024x1024 (1.000MP actual; preset was 768x768 @ 0.562MP)
Both plan_only and a full render report it, so you can check the size before
spending anything.
One thing this does not touch: sampling. H3's shift is a fixed 12.0 in its
model config with no resolution-dependent term β unlike Flux and SD3, there is no
dynamic shift derived from sequence length. Changing megapixels changes cost and
detail, not your sigma schedule.
Reference images (REF2VA)
Connect up to four images to ref_image_1β¦4. By default (ref_mode: where tagged) they land on the shot whose text names them:
Dom, <Picture 1>, drives a van down the driveway.
Only that shot is reference-conditioned; every other shot keeps its handoff.
Every reference-conditioned shot also carries the previous frame as a real
keyframe, so references never cost you continuity β the keyframe fixes the
opening frame, the references supply identity. That is true in all modes, not just
for tagged shots: every shot used to mean every boundary was a hard cut, and
first shot used to ignore your start_image outright, because ComfyUI 0.30 could
not carry both channels at once. 0.31+ can, and the node now does.
The previous frame is also shown to the text encoder, not just to the DiT. A
keyframe pins the opening frame without describing it, so a shot that only got
the latent anchor rebuilt the scene from the prompt β same location, freshly
imagined scenery, which reads as a cut. It now rides in as one more picture,
appended after your references so the <Picture N> numbers your tags use are
untouched.
If a reference gets reproduced in the opening frames, lower ref_noise_aug (0.95,
then 0.90). Note the trade: visual_cond_noise_aug is a single value covering
every conditioning latent, so a softened reference would soften the anchor with
it. Below 0.99 the node therefore drops back to carrying the previous frame as
an extra reference instead β weaker for continuity, but it leaves no anchor to
compromise. info says which of the two you got.
Speed: Sol-Attn (optional, third-party)
ComfyUI-sol-attn is a separate pack (Apache-2.0, wrapping NVIDIA's Sol-Attn kernel) β not part of this one. It ships MiniMax-H3-specific sparse attention and is worth having on a long chain: its own benchmarks put it at 1.38β1.65Γ over SageAttention on H3 shapes. It chains straight in:
UNETLoader β> MiniMax H3 Memory Efficient Sol Attention Patch β> H3 Long Videos
Nothing here depends on it, and it patches attention while this node only patches the sampling schedule, so they don't collide.
Pair it with an SLA LoRA
Sparse attention drops long-range coherence first, and in a video DiT that renders as the same person twice. The fix is an SLA LoRA β a turbo LoRA fine-tuned with sparse attention in the loop, so the weights have already adapted to the approximation. The two are a matched pair:
| sparse attention ON | OFF | |
|---|---|---|
| SLA LoRA | the pairing you want | pays the LoRA's quality cost, collects no speedup |
| ordinary LoRA | duplicated subjects | normal dense render |
The node detects both halves and warns in info when they don't match β including
under plan_only, so you find out before spending a render, not after.
Detection reads the filename off the workflow graph, because that is the only
place the information exists: an SLA LoRA carries no marker in its tensor names or
its metadata and is byte-shape-identical to any un-resized rank-128 turbo LoRA. Any
LoRA with sla as a delimited token in its name counts (..._768p_sla_...);
slack, translate and SLAYER do not.
Resolution: latent upscale (optional, third-party)
The latent_upscale setting drives the MiniMax-H3 Latent Upscaler by
LBH-123-AI β a
345M-parameter 3D-convolution network trained on ~80,000 paired samples (70,000
video, 8,000 image), purpose-built for H3's latent space. Its first convolution
takes 24 input channels, which is H3's latents_dim exactly, and it works at H3's
16Γ downsample. All credit for the model and the upscaler nodes goes to
LBH-123-AI; this node only calls them.
You need two things, neither of which ships here:
- the weights, from LBH-123-AI/Minimax_h3_latent_Upscaler
(
bf16/fp16β 691 MB,fp32β 1.38 GB) inmodels/latent_upscale_models - the node pack that runs them,
Comfyui_Minimax_h3_latent_Upscaler
Without either, latent_upscale does nothing, the render proceeds at the sampled
size, and info says so. It is not a dependency of this node.
What the node reads off a LoRA
It reports where a LoRA's declared training disagrees with your settings. It never overrides a widget β a render has to stay reproducible from what the graph shows.
| Checked | Source |
|---|---|
| Base model is MiniMax-H3 | metadata (base_model / ss_base_model_version) |
Step count vs your steps |
filename (..._4step_...) |
| Training resolution vs your preset | filename (..._768p_...) |
Notes say which source they came from, because the two aren't equally trustworthy: metadata is what the trainer wrote, a filename is a convention anyone can break by renaming.
Not available, so not offered. LoRA files carry no field for a recommended
sampler, scheduler, cfg or shift β no metadata standard defines one β so the node
does not pretend to know them. Trigger words are also unreadable in practice: they
live in ss_tag_frequency, which kohya writes and ai-toolkit does not, so a LoRA's
trigger still has to be typed in yourself β into the prompt, the sheet, or
exposed_terms, depending on where it needs to land.
On ComfyUI portable its Triton kernels will not build, and they fail silently β the patch reports itself inactive and you simply get the slower path. The embedded Python ships without development files:
python_embeded\Include\ contains only greenlet\
python_embeded\libs\ does not exist
Fix (verified on Python 3.13.12, Triton 3.7.0, CUDA 13.3, SageAttention 2.2.0, SM120 β sol-attn's own test suite goes 3/7 β 7/7):
- Check your version:
python_embeded\python.exe --version - Download the matching CPython NuGet package (it is a zip):
https://api.nuget.org/v3-flatcontainer/python/3.13.12/python.3.13.12.nupkg - Copy
tools\include\*intopython_embeded\Include\ - Copy
tools\libs\python313.libintopython_embeded\libs\(create it)
Purely additive. This unblocks Triton generally, not just Sol-Attn. Redo it if a
ComfyUI update replaces python_embeded.
Note the sparse paths are approximate β A/B a shot before adopting them.
Requirements
- ComfyUI 0.31+ with native MiniMax-H3 support (tested on 0.33; on 0.30 the audio shifts behave differently -- see Requirements in REFERENCE.md)
- Pillow only for the text overlays (ComfyUI already ships it)
- No negative prompt β H3 is CFG-free at
cfg 1; the node makes an empty one - No denoise input β fixed at 1.0; partial denoise desyncs the audio schedule
Full reference
REFERENCE.md β the long-form field-by-field notes.
It is out of date. It still documents a
total_secondsinput that no longer exists, and 17 of the node's 69 fields are missing from it β includingmegapixels,sampler_name,trim_seam,vary_seed_per_shot,prevent_nudity,exposed_terms,lock_restraints,auto_soundscapeand all four guards (anatomy_guard,solidity_guard,motion_guard,contact_guard). This README and the in-node tooltips are current; REFERENCE.md is not. Read it for background, not for behaviour.
Disclaimer
The owner of this repo will not be responsible for any copyright strikes incurred because of use. You are responsible for your works. Use this node responsibly and ethically.
- Downloads last month
- -