Weird. I found both quite good, and the motion and interpretation of prompt to be far more naturalistic. If anything, I'd say it's more Grok-like than anything else so far. But I find with SD, it's always a weird balancing act between the prompt, the choice of image and video reference. I can use the exact same prompt with the same basic structure of reference images, and the exact same video reference, and get rather different results. One thing I will say is that the bleed-through from the video reference has been better and I think the prompt adherence is better than 2.0.