text2vid img2vid H3 Minimax - Fully uncensored local AI video generator released today

get video components node then goes just to image input ? i dont see this one "minimax H3 reference to video" only "image to video minimax"
 
I'm using ComfyUI, dunno if it's different on whatever you're doing but reference to video should take up to 9 images not just 1, and the video also goes in the same box there
 
Some follow-up prompting tips from this, I've since expanded to Ref2VA generation w/ 3 images and a reference video (same workflow is available in the EP 29 video referenced) and have gotten pretty insane results, creating 10 sec videos in 13-15 minutes at times.

Instead of using the default "minimax_h3_fl2va_pruned_int8_convrot" diffusion model I've since been using the "10Eros_Max_h3_TURBO-hybrid_beta4_int8_convrot" which vastly reduced my need for LORAS (only using minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16 at .20 strength but honestly not needed). However, this is still VERY reliant on you having good reference images and a reference video. If you're able, two body images with one close-up of the face has proved very consistent to me. If your unable to present those, you will naturally have a harder time getting consistent results.

PROMPTING: On this I will re-plug this comfyUI workflow because it has worked wonders for me (https://civitai.com/models/2834106/minimaxh3-auto-prompter-v73). It will do alot of the initial work for you BUT you still need to PROOFREAD the prompts and adjust and add specific details here and there to get your exact visions. However, if you get an understanding of the Ref2VA prompt structure then the curation of the prompt following generation you just naturally get better at. My prompt structure has consistently become:

subject_defintions:
retention_analysis:
Summary:
[video editing]
integrated_multimodal_description:
overall_soundscape:
detailed_description:
Someone else on this thread shared more details on this structure. Of course, is more work to curate for every subject but trust me it pays dividends to get what you need.

SIMPINGWHALE for you and anyone else having issues with LLM's for prompt generation, I'd suggest this above prompting workflow as an alternative to remove the LLM middle-man.

Some more prompting tips:
- Don't just assume you providing referencing of a subject that you don't still need to describe the subject in the prompt. This is the harsh truth.
For example, if your reference video has a blonde person, and your reference images is also a blonde, you need to be hyper-specific in description to distinguish away your intended subject away from the person in the reference video. This is how subject transfer happens where the generated video just looks like the reference video. However, one quick way to try solve this is if you can get a similar concept reference video but with someone who already looks different from your subject. i.e A similar video to what you gave but instead of a blonde it is someone with black hair. The model then has no point of reference for a blonde other than your provided reference images. Hair is just an example, but this concept applies to any descriptive attributes. As said above, you still need to proofread and specify as needed.
- Adding timestamps to when you want certain things to happen reduces unexpected actions and gives you consistency, especially when trying to get that first iteration. The above workflow will give you an idea of how this works.
- Obviously use the lowest possible time you can when doing the first iteration, a 3 sec video will give you a good idea of how the 10 sec one will go, and you dont even need a seperate prompt, a prompt for a 10 sec can still be used for any lower time.
- Another somewhat obvious one, but you shouldnt try to generate a video longer than your reference video. So if your reference video is 10 secs, dont try to generate a 15 sec video.

Honestly, this academic level of sourcing and curation is not for everyone but the increase in quality and consistency is so worth and addictive once you get it right the first time. This AI stuff isn't a magic button, so aside from physical limitations if you choose to cut corners then you will just pay in the quality.

But using all the above tips, I've been able to use reference videos like below and get really good recreations with my own subjects.

 
I am currently using CivitAI with real faces and it's allowing me to get them to speak, always in a bad accent though lol.
I was just wondering how far you can go with CivitAi.red !? I use image to video.
 
node - get video components -which will give you inputs for audio and images for the ref/vid input
-in your nodes folder on comfyui --- type::: get video components <---its the name of a node that is in your nodes folder it is not part of minimax -its stock comfyui node
 
Use 10eros turbo hybrid beta 5. It has enough grafted onto it to make simple grok like goon vids
 
I have some issue , Idk why but for me it take 700sec for one sec video and take 3 hours for like 15 sec, anyone have some solution???
 
Made using the Ref2VA checkpoint:


References:
- Image of Chun-Li for
- Image of Tifa Lockhart for

Loras and settings used:
- HMBreasts @ 0.4 - Lora Link
- Free Pussy @ 0.4 - Lora Link
- Steps: 24
- Scheduler: beta
- Sampler: euler

subject_definitions:
is the person from She has a shaved pussy and with HMBreasts, large, perfectly shaped breasts, small sized pale areoles and erect nipples.
is the person from .

retention_analysis:
(appears in [Shot 1]): fully_preserved - the woman's facial features, hairstyle, and outfit are retained.
(appears in [Shot 1]): fully_preserved - the woman's facial features, hairstyle, and outfit are retained.

detailed_description:
The target video is a raw, shaky handheld iPhone video shot inside a luxurious spa. It uses one continuous shot. There are no cuts in the entire video.

[0s-3s] is standing in the room while is in front of her. and are both the same height. quickly rips open her qipao from and fully exposes both of her breasts.

[3s-6s] grabs breasts with both of her hands and begins passionately french kissing the nipples.

[6s-10s] uses her hands to lift up her skirt to expose her pussy as gets on her knees and begins licking pussy with long deep licks of passion inside her pussy. The camera smoothly moves down to focus on licking the pussy.

[10s-15s] The camera zooms in on licking the pussy passionately with an extreme closeup.

overall_soundscape:
The entire scene is filled with the intimate, wet sounds of passionate french kissing against soft skin, specifically the breasts, interspersed with slow, deep, and increasingly rapid heavy breathing from the women, their inhalations and exhalations synchronizing and escalating with the passionate activity. The wet licking sounds of tongue licking the wet moist pussy of .

non_diegetic_music:
N/A
 
I'm new here; I'm using Codex to create a personal assistant. I plan to have it generate NSFW images and videos (i2v), but I doubt a Mac (M5 Air) has the power to generate all that. Do you think I could get better results by offloading the processing power to a Google Colab notebook (using ComfyUI)? If so, what would be the best model for this—especially for i2v?
 
How ????? I have 12 gb of vram and it's been 78min and it's still not finished , I wanted to generate 15 sec vids
 
Tried locking to 24 frames at 736x416 and it crashed at 40 minutes but at least now it's getting stuck KSampler instead of the H3 Minimax Node.

I'll figure it out another time, I'll just stick to i2v and t2v for now.
 
Sounds good in practice. But there’s a problem. The Depth Map still keeps the body proportions from the reference video. As a result, the generated outputs always have to fit that same mold. For example, when I use a video reference like the one you sent, I can’t generate videos of a woman with big breasts.

OUTPUT:

IMG REF:
 
I was able to generate this result using the Depth Map video:


I attached an image of Chun-Li for

Modified prompt used for this:

subject_definitions:
: the target face, head and body reference. Use the woman’s face and head from this image, including her facial identity, face shape, eyes, eyebrows, nose, lips, smile, skin tone, makeup, black hair with side-swept bangs that frame her face and the rest of her hair is styled up into two prominent buns on either side of her head with Ox horn hair buns covered in white cloth and ribbons and overall head appearance from
: a depth/motion reference: it shows pose, movement, camera framing, and timing to follow. It carries no color, face, body, or identity information.

summary:
video editing + reference generation. Replace the source woman’s head and body in
 
Back
Top