Dataset methodology/captioning/training?
Are you willing to explain a bit on how you trained the set? I've tried about every 7of9 lora out there since 1.5, and made over a dozen of my own, but this is the best I've ever seen. It can do nigh-perfect "scene continuations" from a high-res I2V still.
I've always used Musubi tuner, and not a whole lot of LTX ones can do audio it seems. Was it trained locally or with a pod etc? I've yet to try to train from video myself (only images), but would love to collaborate on something similar, assuming my setup is similar to yours. I'd happily "train a new version for you" if you want to try something---perhaps one focusing on the blue suit? Were training vids made from various clip-lengths or all the same number of frames, and were they manually clipped from or something automatic? How many samples and what resolution? Were they left at the original 4:3 filming ratio or cropped to something else? Was any upscaling used for the source, or just raw rips etc? Same question for captions---I've tried about every caption methodology ever, and have yet to determine what's best. Any/all answers appreciated, thanks.
Are you willing to explain a bit on how you trained the set? I've tried about every 7of9 lora out there since 1.5, and made over a dozen of my own, but this is the best I've ever seen. It can do nigh-perfect "scene continuations" from a high-res I2V still.
I've always used Musubi tuner, and not a whole lot of LTX ones can do audio it seems. Was it trained locally or with a pod etc? I've yet to try to train from video myself (only images), but would love to collaborate on something similar, assuming my setup is similar to yours. I'd happily "train a new version for you" if you want to try something---perhaps one focusing on the blue suit? Were training vids made from various clip-lengths or all the same number of frames, and were they manually clipped from or something automatic? How many samples and what resolution? Were they left at the original 4:3 filming ratio or cropped to something else? Was any upscaling used for the source, or just raw rips etc? Same question for captions---I've tried about every caption methodology ever, and have yet to determine what's best. Any/all answers appreciated, thanks.
Hey, thanks.
I trained with AI Toolkit.
Resolution was 768 but actual dataset was 720p upscaled 2x using flashvsr to get better quality from the best versions I could find on YouTube.
I used roughly 16 videos under 5 seconds and 12 upscaled images.
I am, however, working to make an even better one with more clips I got and images as well.
I am trying to see if I can intentionally make the other characters generated better using t2v. I already had some success with The Doctor.
I just captioned very basic but using my trigger word and the description of her implant.
It took lots of tries to get Regularization to train well since the first version.
I just trained gradient accumulation 2 and differential guidance 3 to start, and lowering them over time. I used AdamW8Bit with balanced timestep, and low noise at the end.
I couldn't get the implant as detailed with the regularization loras, but I trying to see if it just takes longer perhaps.
I used pretty much original aspect ratio except the images were made 16:9.
Most clips were 2-5 sec max. I just used the feature in AI Toolkit that automatically detects the frame length, so I could use any number of frames i want.
I found more wider shots to use which should help, because all the video i had used before was basically closeup only with a few exceptions.
Thanks so much, this lora impressed me so much that I'm gonna try AI toolkit this weekend, and see what I can come up with. So you mainly grabbed clips from youtube as source material? And just 16 videos was enough, with the dozen pics added in?
I haven't used regularization images in my loras for a LONG time, do you think it made a big difference? Did you just use generic "person" regularization? (this may be a stupid question, as I've never tried regularization in AI toolkit and have no idea how they implement regularization)
Could you give one or two examples of the captioning? My "basic" and other people's "basic" can vary a lot I've found. :D
Thanks so much, this lora impressed me so much that I'm gonna try AI toolkit this weekend, and see what I can come up with. So you mainly grabbed clips from youtube as source material? And just 16 videos was enough, with the dozen pics added in?
I haven't used regularization images in my loras for a LONG time, do you think it made a big difference? Did you just use generic "person" regularization? (this may be a stupid question, as I've never tried regularization in AI toolkit and have no idea how they implement regularization)
Could you give one or two examples of the captioning? My "basic" and other people's "basic" can vary a lot I've found. :D
All the regularization clips I used (and images as well in separate datasets) were just from voyager as well of characters talking (yes with their voices) in the same manner as the seven of Nine stuff but repeated at a ratio so that regularization was less than 10% the total dataset.
For regularization I just used Video Caption Suite to caption Prompting it like:
Describe the video in one brief paragraph. Focus on the person talking. Describe the scene and setting in few details.
And then it just vaguely describes the people, and for seven I just manually add her trigger and implant detail.
I ask the captioner to refer to her as a name and then find and replace that all later with Taggui.
The ai caption model is very sloppy with names, so you must not make it too complicated or it makes typos or hallucinations.
I remove any reference to Star Trek or character names it gives them. I describe their clothes simply as green uniform or whatever colour, hers is a bodysuit, and then later generation it knows which clothing is which.
I am going to experiment with naming the other characters to see if I can train them in through regularization since that already seems to work only by describing them vaguely.
Regularization in AI Toolkit you simply check the 'Is Regularization " tab for that dataset and make sure the character trigger word is not in that dataset nor the character themselves. So seven cannot be in any of those videos.
One small suggestion, from other loras and seeing how most video models have their roots in youtube/tiktok/instagram training, and the most common "appearance error" I get from this lora:
Call it a catsuit instead of a bodysuit. Bodysuits in modern looks tend to be legless, and that trigger word often results in a bare-legged version of her look. I blame Taylor and Beyoncé's influence for re-defining what it means nowadays. :D
Thanks again, I'll see what I come up with this weekend, assuming I get AI toolkit up and running without issue. (I THINK I tried it once a while ago, but it's definitely not installed currently)
Oh, one last thing---are you running in GUI or command line? I actually tend to have better luck with pure commands, for trainers.
One small suggestion, from other loras and seeing how most video models have their roots in youtube/tiktok/instagram training, and the most common "appearance error" I get from this lora:
Call it a catsuit instead of a bodysuit. Bodysuits in modern looks tend to be legless, and that trigger word often results in a bare-legged version of her look. I blame Taylor and Beyoncé's influence for re-defining what it means nowadays. :D
Thanks again, I'll see what I come up with this weekend, assuming I get AI toolkit up and running without issue. (I THINK I tried it once a while ago, but it's definitely not installed currently)
Oh, one last thing---are you running in GUI or command line? I actually tend to have better luck with pure commands, for trainers.
Ai toolkit has a gui. Much better.
I have purposefully avoided any trainer that doesnt have a gui, which is all of them.
I have never noticed her legs but struggle to get my lora to put her in wide shot from lacking in dataset. Hopefully I can fix that.
Catsuit could give me bad results, but hard to say unless I use that caption. It might associate it wrong.
I have tried to make a Walter White lora with regularization but something about my captions makes it struggle to make him anything but a 70 year old man.
Without regularization dataset he looks fine. Very odd. All about captions
I have found i2v to give good results decently often for full-body/wideshots. Do you have a scene/phrase etc you'd like me to try with my current settings/workflow, using your lora? I'd be happy to try something, see if I can make it work.
I have found i2v to give good results decently often for full-body/wideshots. Do you have a scene/phrase etc you'd like me to try with my current settings/workflow, using your lora? I'd be happy to try something, see if I can make it work.
No, but thanks.