Some follow-up prompting tips from this, I've since expanded to Ref2VA generation w/ 3 images and a reference video (same workflow is available in the EP 29 video referenced) and have gotten pretty insane results, creating 10 sec videos in 13-15 minutes at times.
Instead of using the default "minimax_h3_fl2va_pruned_int8_convrot" diffusion model I've since been using the "10Eros_Max_h3_TURBO-hybrid_beta4_int8_convrot" which vastly reduced my need for LORAS (only using minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16 at .20 strength but honestly not needed). However, this is still VERY reliant on you having good reference images and a reference video. If you're able, two body images with one close-up of the face has proved very consistent to me. If your unable to present those, you will naturally have a harder time getting consistent results.
PROMPTING: On this I will re-plug this comfyUI workflow because it has worked wonders for me (
https://civitai.com/models/2834106/minimaxh3-auto-prompter-v73). It will do alot of the initial work for you BUT you still need to PROOFREAD the prompts and adjust and add specific details here and there to get your exact visions. However, if you get an understanding of the Ref2VA prompt structure then the curation of the prompt following generation you just naturally get better at. My prompt structure has consistently become:
subject_defintions:
retention_analysis:
Summary:
[video editing]
integrated_multimodal_description:
overall_soundscape:
detailed_description:
Someone else on this thread shared more details on this structure. Of course, is more work to curate for every subject but trust me it pays dividends to get what you need.
SIMPINGWHALE for you and anyone else having issues with LLM's for prompt generation, I'd suggest this above prompting workflow as an alternative to remove the LLM middle-man.
Some more prompting tips:
- Don't just assume you providing referencing of a subject that you don't still need to describe the subject in the prompt. This is the harsh truth.
For example, if your reference video has a blonde person, and your reference images is also a blonde, you need to be hyper-specific in description to distinguish away your intended subject away from the person in the reference video. This is how subject transfer happens where the generated video just looks like the reference video.
However, one quick way to try solve this is if you can get a similar concept reference video but with someone who already looks different from your subject. i.e A similar video to what you gave but instead of a blonde it is someone with black hair. The model then has no point of reference for a blonde other than your provided reference images. Hair is just an example, but this concept applies to any descriptive attributes. As said above, you
still need to proofread and specify as needed.
- Adding timestamps to when you want certain things to happen reduces unexpected actions and gives you consistency, especially when trying to get that first iteration. The above workflow will give you an idea of how this works.
- Obviously use the lowest possible time you can when doing the first iteration, a 3 sec video will give you a good idea of how the 10 sec one will go, and you dont even need a seperate prompt, a prompt for a 10 sec can still be used for any lower time.
- Another somewhat obvious one, but you shouldnt try to generate a video longer than your reference video. So if your reference video is 10 secs, dont try to generate a 15 sec video.
Honestly, this academic level of sourcing and curation is not for everyone but the increase in quality and consistency is so worth and addictive once you get it right the first time. This AI stuff isn't a magic button, so aside from physical limitations if you choose to cut corners then you will just pay in the quality.
But using all the above tips, I've been able to use reference videos like below and get
really good recreations with my own subjects.