AI Need help with Local IMG to Vid

SullyHex

Lurker
Joined
May 23, 2026
Posts
4
Reaction score
0
It depends how far down the rabbit hole you're willing to go down. I gave up on ComfyUI for the simplicity of Wan2GP for video generation. There's a bunch of presets for Wan 2.2 and it auto downloads all the models you need. I'm on a 4070 and can get a 5 second video done with surprisingly good quality in under 3 minutes. I use CIVIT loras and stick with the models Wan2GP offers.
 
Paying for a OpenAI Codex or Claude subscription is the easiest way to get more out of someone else's workflow, or any workflow in general. For example with Claude, you can initialize it in the ComfyUI workflow directory so it knows you're working out of there, and just ask it any questions and it'll be extremely helpful 90% of the time.
I use Wan 2.2 and generate 5 - 6 second clips and stitch them together for scenes that have been up to 180 seconds so far. 95% of my workflow debugging and editing has been done via Claude via the $20 a month subscription, and it pays for itself.
 
can you explain more how to best stitch videos together? do you mean the ai does it on its own? would be interested in this as i have thousands of short videos and dont know how to consume them
 
I'd appreciate if you could elaborate. What specific presets, loras do you use etc?
 
So typically for a scene in the same positions, I do multiple generations off the same base image to make a coherent scene. I'll generate between 6-12 6 second clips off the same base image by just varying the prompt and varying the seed.
For example for a sex scene, I may have the subjects start off slow, and then medium pace, and the fast and hard. So it'll be split between those 6-12 generations so that when they're stitched together, it looks like one 30-45 second clip.
I take those clips and stitch them together via VACE Joining. I use the official VACE joiner workflow that's readily available, version 2.5. I set the workflow to graab 24 frames of context from each clip, and then generate 8 frames for the joining stitch.
It takes a lot of time, but the result is one coherent long clip. If you look closely you can still see the "seams" at the joins because it'll be a little unnatural, but for me that is completely acceptable.
The VACE Joiner workflow takes all your numbered clips in a directory and auto joins everything for you. On my machine with a 5070ti, 720p clips joined typically take roughly 3 minutes to make their transition, so a set of 10 would take 30 minutes to complete all the joins and get your result video.
 
1) get this i2v workflow. it allows you to generate 6 * 5 second sections for up to 30 seconds of video if you get the 3.5 or 12 sections if you get the 3.5 12 sections. - https://civitai.red/models/2409202/wan22-i2v-svi-workflow-kenpechi
2) download your high/low wan loras, and also recommend these:
AIO to cover multiple positions - https://civitai.red/models/1811313/dr34ml4y-all-in-one-nsfw-wanltx2
pussy/anus lora - https://civitai.red/models/2109996/wan-22-pussy-and-anus-lora
3) upload workflow and reference image to Ask Grok and tell it to write positive/negative prompts for all 6 sections of what you want subject to do
4) start with first section enabled, paste prompt to first and run. if you like what you see, proceed with 2-6.
 
i tried comfyui and got megaconfused (i also have early onset dementia)
I tried Amuse and havent looked back.
 
holy shit, you dont know how much i hate you right now. This was a TERRIBLE suggestion.
 
-i thought amuse blocks nsfw content?
thats not true at all, right off the bat you import shit like cyber realistic xl.
after my disaster today with mixmatch i tried out fooocus and it does everything for image gen, edit, impaint and it was a million times easier than running an app on top of comfy
ill keep amuse for some vid gen for now since fooocus doesnt do it
 
Back
Top