Merging weight of Cosmos3-Super to have better understanding of physical world.
Is it possible to merge weight of Cosmos3-super to enhance model's physical world understanding so that it won't glitch sometime?
Is it possible to merge weight of Cosmos3-super to enhance model's physical world understanding so that it won't glitch sometime?
It depends on layer naming and how cleanly separated Cosmos's Diffusion tower is from the other architecture, that's all you could use in my current method. The Cosmos MoT could mean it doesn't work right, or it is actually cleanly separated from the model. Nvidia could have done anything they want there and it might be too unique too. Also a 65B 130g model is pretty massive, I'd try Nano instead. But what does it actually present for H3 to gain really? H3 is already good at tokenizing and cross attention, and Cosmos can only input 5-400 frames of video and half a second of audio so it's not even a better conditioning to absorb. The method right now is to shape H3's attnetion with a model that does alternative styles and actions to skip a ton of baseline training only, but the grafting math works on attention and could go farther I'm sure, but I haven't explored it.
i tried it paid two times (on fal.ai) and it was a disaster - in exactly that aspect. I had some people walk and they just hovered sideways. In poor word understanding i would expect some (bad) resemblence of walking. But it was complete visual nonsense.
For a massive model like cosmos 3 super that was head scratchingly bad. And a massive model, that is not optimized for size and speed like LTX i should not need a literal step by step prompt as for LTX. And i should not need a finetune/lora for the absolute basics. Hence H3. And wan2.2 btw. They do mostly what you want even on a simple prompt.