This is a great overview, especially the discussion around how video generation is fundamentally harder than image generation because temporal consistency, motion dynamics, and conditioning all have to work together.
One thing that has become increasingly important in practice is the workflow around these models. For many creators, starting from an image can be more controllable than generating everything from text, especially when character appearance, composition, or a specific visual style needs to be preserved in the resulting video. The challenge is not only choosing a good model, but also having a practical way to experiment with different models and workflows without setting up and maintaining a large local GPU environment.
That is one reason image-to-video has become such an interesting part of the video generation ecosystem. For anyone experimenting with different models and approaches, this Image to Video workflow can also be useful for turning existing images into videos and comparing different generation approaches.