As a layman/very normie user myself will note everything I used to get myself setup this morning relatively quickly with never using ComfyUI or these models before. Able to consistently generate prompts for 5 second one shot I2V's that take about 5-8 minutes each with results I'm already happy with, have not worked with 5+ second video generation or a collection of LORA but feel this is a pretty good base to append this setup further for that usage. I have a upper mid tier GPU, and 32 GB RAM.
YT Videos:
- ComfyUI Course - Learn ComfyUI From Scratch | Full 5 Hour Course (Ep01) (specifically Chapter 3 and beginning of Chapter 4)
Note: This video is great because this an a portable install of ComfyUI that will come with QOL options like a launcher, and addons that will be prereq's for the second video. Essentially will be a clean install of ComfyUI that doesn't need to go on your C drive (assuming you have another drive you can use). When you install using EasyInstall you WILL NOT have to install ComfyUI from the website
- ComfyUI MiniMax H3: Best Video Generation Workflows (Ep29)
Note: particularly I am using the workflow from the 'H3 First Frame Workflow + More Settings' chapter
Other:
https://civitai.com/models/2834106/minimaxh3-auto-prompter-v73
This is a seperate ComfyUI workflow that is for the actual prompt generation. Found it from this same thread. The workflow will note the details/objects you need. This came from me not wanting to go through a separate LLM, and keep the prompt generation activity offline. For any single image I just pass the image, give a brief plain english description, and in about 20-30 second gives me a properly structured prompt for the above workflow.
Claude (Pro Sub tier): Used this as I encountered errors in running (and yes you will probably run into some), I did not have to specify the nature of what I was doing outside ComfyUI workflow troubleshooting, simply pasted errors and gave brief plan english context to files and such to assist. This bit is still a bit dependent on your comfortability with general AI usage. Also, am obviously not saying you need Claude, but an AI chatbot on the side goes a long way instead of having to dig through forums for very specific errors in what is already a niche subject.
"Technical" actions I did: Basic file explorer usage, edited some BAT files, ran PIP commands for Python stuff, and troubleshooting with AI (Claude as mentioned above). Everything else the workflows and videos themselves will outline what you need to do.
My workflow as is: I input an image to my prompt generation workflow, give a plain english description of what I want (about 3-5 sentences). Let it run and it gives me a structured prompt tailored to that exact photo. (This naturally means that every photo will have its own prompt. Depending on your use case you can use the structured prompts to perhaps create a generalized prompt if you'd like.) Paste that image and prompt into my actual H3 First Frame Workflow, config as needed, and boom I have my generated video. For LORA usage I plan to use Chapter 19 in the first video I mentioned to simply append my current workflow. Total setup time: About 2.5 hours (largely due to download times of the models).