text2vid img2vid H3 Minimax - Fully uncensored local AI video generator released today

Most of my posts takes a day or two to get approved, so check back in while when this post is buried a few pages back.
 
also, there is a cost to every speed /downsizing -- if the model doesn't perform as you want //try non- optimized and unpruned
int8 -- is one of the fastest without messing up quality- and that card should be good with that - plus int8 doesn't output too differently from the stock
 
I found having to be painfully specific and repeating those details multiple times in my prompt gets the proportions pretty well, but I've also been using my prompts as if it was the original reference but using the depth video in the actual workflow, so the wording in what SuperSimps01 posted may as well assist. My prompts are a solid 7k characters worth to get what I'm happy with. As I optimize I can share some example myself.
 
I'm really enjoying MMH3, i'm still new so I've mostly messing around with i2v so far. I'm using MysticXXX V4 LoRa at 0.8 strength and minimax turbo 8 step which usually takes about 4 1/2 mins per 15 sec clip on a 5090, the 4 step is 2 1/2 mins but the quality loss is noticeable.
I've also started doing 30-45 sec videos with the MM H3 Motion Director which samples previous segments context frames and clears VRAM after each 15 sec segment but with i2v there is sometimes a pretty noticeable loss in character continuity as you'll see from this video. Strong prompting helps somewhat. I've seen demonstrations of it working better in ref2v.

This is a weird idea my brain came up with late at night lol, this is a fictional character created by GPT, not a real person.

scene 1:

the woman in the uploaded reference photo is a tiktok influencer recording a selfie on her cell phone talking to her followers, she says: "hey guys today i'm at chick fil a to take on their dick sucking challenge, i ordered an extra large" a couple seconds pass and a male chick fil employee walks up to her car window wearing his standard uniform: a red polo t-shirt but not wearing any pants and he's presenting his erect glistening extra large penis. he says "Ma'am will this complete your order today?" the woman stares at his penis shocked because of how large it is and says "fuck yeah, that's exactly what i ordered"

the woman reaches out the car window and grabs his cock aggressively and begins jerking off the shaft while grinning, she starts sucking on the tip of the head of the penis.
scene 2

she then goes full aggressive sloppy deep throat — gagging wetly, spit flying, throat bulging. she smirks at camera, while deep throating the penis, putting the entire head and shaft down her throat.

She pulls off, catches her breath with strings of spit while jerking off the penis with her hand she says "oh that's good", the penis ejaculates from the tip, spurts a couple quick spurts of clear pearlescent semen on her face and hair, giving her a proper cum shot, some of it landing in one of her eyes. the man exhales loudly and moans softly as she dives back in.

close-up of her face deep-throating, looking up him. she pulls the dick out of her mouth and is smiling filthy with cum running down her face.

Scene 3:

the man says "My pleasure" the woman turns back her focus to the camera with cum dripping down her face. she scoops up some cum with two of her fingers and slowly licks it clean and swallows.

she talks to her audience and says "Guys that was fire! I recommend the extra large, it was 10 out of 10. be sure to like and subscribe and if you want to see more of this content please follow me on Patreon.... Peace!"


Here's a couple other 15 sec prompts that came out pretty well.

The woman exhales loudly and is breathing heavy, moaning in sexual pleasure.

the camera slowly pushes out revealing her nude body, she has conical shaped breasts, her legs are open as she is slowly rubbing her clit with one hand in a circular motion. her pussy is glistening, she has a small tuft of dark pubic hair.

she begins fingering her exposed pussy with deep insertion: two fingers fully inserted knuckle-deep or two fingers buried inside, actively thrusting and sliding in and out with hard strokes.

The wet, soggy sound and visible stretching of her pussy lips is clear as the fingers pump in and out, creaming and dripping everywhere, on her clit and pussy lips and running down her thighs onto the blanket.

Her pussy squirts a few powerful jets of clear vaginal fluid with every stroke, high-pressure arcing streams splashing towards the camera. eyes half-closed then locking into the camera with intense desire. Skin-on-skin slapping sound begins— sharp, rhythmic “SMACK-SMACK-SMACK” with natural reverb. Wet, squelchy sounds under every thrust. Her heavy, throaty breathing starts immediately — loud, deep inhales and shaky exhales.

Her moans get louder and messier: short, breathy “AH-AH-AH” mixed with deeper “OH-OH—FUCK—” groans

a nude man enters into the frame from the left side presenting his fully erect glistening penis (slightly throbbing). He wraps his fingers around the thick shaft, and starts jerking it off hard and fast. a few spurts of clear pearlescent cum erupt from the tip with every stroke, shooting and splattering across her face, and tits. She keeps fingering herself deep with her other hand while jerking, body shaking with pleasure as more cum keeps cumming.

Her body trembles with pleasure, toes curled. Moans intensify: “YES—AH—OH MY GOD—” cut short by loud, trembling gasps.

Camera style
Use a handheld feel (subtle shake) so it looks like real footage, not a studio shot.
Natural lighting:

Audio: natural room reverb, no music. wet skin slaps, and rhythmic breathing.

the woman in the reference photo looks at the camera and says with a sassy attitude "Hey Sadie watch me suck your brothers cock!"

a nude man enters the from the left presenting his glistening erect penis (slightly throbbing). the girl grabs the cock aggressively, sucks the head fast while jerking the shaft.

she goes full aggressive sloppy deep throat — gagging wetly, spit flying, throat bulging. she smirks at camera, while deep throating the penis, putting the entire head and shaft down her throat.

She pulls off, catches her breath with strings of spit while jerking off the penis with her hand, the penis ejaculates from the tip, spurts a couple quick spurts of clear pearlescent semen on her face and hair, some of it landing in one of her eyes. she quietly says "fuck yes" dives back in.

close-up of her face deep-throating, looking up him then smiling filthy at camera with the dick still in her mouth. as she flips the camera off holding up her middle finger

Camera style
Use a handheld feel (subtle shake) so it looks like real footage, not a studio shot.
Natural lighting:

Audio: natural room reverb, no music. Skin slaps, wet gagging, wet sucking and rhythmic breathing. Wet, aggressive sucking with loud, rhythmic slurping. no crunching sounds.
 
This looks really bad for the hardware you're using. What resolution are you generating at? Try and stay above 1mp at all times once you have your prompt and concept down. Use 10eros turbo hybrid beta 5 in place of your turbo and mystic loras
 
I find it's better to just make a single collage of your reference images and create clear instructions what each panel is for. You save a lot on resources this way
 
What is the difference between a character sheet and a single collage of reference images? My reference images would be subject’s body at different positions and face references at different angles
 
Typically with H3 character sheets are made from a video using either a workflow or just a video and taking screenshots. If you already have images you can use for your sheet just edit them into a collage
 
Man, that's terrible. With a 4060 Ti 8GB and 32GB of RAM, I generate a 30-second video in 20 minutes using Turbo LoRA and 4 steps. (with continuum)
 
for me after dling mystic 1-4 . i went with v2 the 3rd and 4th are messing with stuff it should not and v1 you can tell wasn't all the way there yet
 
Here is an example prompt to demonstrate the structure I've used for the reference video and subsequent generate depth videos like below, albeit through a somewhat brute force approach. The first draft iteration of this prompt is using a prompt generator workflow that is referenced in my previous posts.

I like this example because this is a relatively complex composition that required very specific details to get right, but with following the full Ref2VA prompt structure can certainly get to a nice baseline that I'm confident make more simple reference videos very straightforward to work with in comparison. Refer to my previous posts to see diffusion models, workflow, and LORA usage. This is with 4 reference images and one video with audio, you obviously can simply cut it down to base around less images but if you can provide them then use them. You don't need to focus so much on this specific reference and my writing (lol) but more so the prompt structure to see how it can be actually leveraged for your use cases and with these depth conversion videos.


Converted depth video (ultimately what is used in Ref2VA workflow) - see previous posts on this thread that show how to generate these


Some notes for this:

- Compared to others prompts shared this is obviously much wordier, but this is what gives me the peace of mind that when I'm getting that first iteration I have a consistent baseline so that I can remove unexpected behavior. It is obviously more to manage and proofread but I honestly learned to enjoy the prompt engineering.

- Yes, the details are reiterated multiple times. Maybe in some places it's redundant but using this sort of brute force approach gave me something that works, and I truly think effective in not letting the models "forget" what your set vision is

- Definitely room for optimization and cutting fluff but for me personally, if it ain't broke don't fix it. After only knowing this stuff exists for a week I'm pretty happy with my progress.

- The exact way the male torso gets shown is not totally cleanest but in my opinion totally good enough for the intended vision, this composition is just that specific in my opinion. As said above for simpler video references this approach will get it spot on, and it has for me.

- This is with the depth video being used, using the raw reference actually gets the torso totally spot on but I wanted to show how you very much control the primary subject proportions if you have to use the depth video.

- I use my 4 reference images so that the first 2 are full-body/figure references and the 3rd and fourth are close-up shots of face or upper body.

subject_definitions:

: a woman from with full-figured proportions and an hourglass silhouette. She has oval facial features, medium brown skin tone, long dark hair parted in the middle, defined eyebrows, subtle makeup including eyeliner and lipstick; she wears a floral-patterned bikini top with white trim, gold hoop earrings, layered bracelets on her right wrist, and a delicate pendant necklace.

Primary reference of outfit and figure of woman from .

Additional reference for facial features and figure of .

Additional reference for facial features and figure of .

: the glans and girth of a large penis with both inner thighs of the slender man slightly in frame, no other portion of his body in frame. his left thigh laying beneath her buttocks and the right thigh tightly cropped raised up on the right edge of the frame so only the left half of the right thigh is visible in the frame. His thighs are cropped tightly so he remains secondary to her form.

: Plain indoor bedroom with light gray/white walls, two closed white paneled doors, a ceiling air vent, and a light switch, providing a neutral domestic setting that keeps focus on the subjects. Illuminated by natural sunlight.
 
I found that using Turbo Loras when using a reference video only hurts the coherence and quality with no real speed up.
The below is my newer settings (ask ChatGPT if you need help setting these up):

Models:
- Diffusion Model: minimax_h3_hybrid_fl2va_ref2va_b25-49-int8.safetensors
- Text Encoder: qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors
- VAE: minimax_h3_video_vae_int8_convrot.safetensors
- VAE: minimax_h3_audio_vae_fp32.safetensors

Settings and Nodes:
- Use the ModelPreviewOverrideKJ node to see a video preview while generating so you can cancel a job early if its not correct
- Comfy Kitchen Attention
- Scheduler: simple, 8 steps
- Sampler: euler

On a RTX 5090, it is able to generate a 15s video at 0.5MP (960 x 544) using a reference video of 320p in about 3 minutes. You can usually tell within the first minute if the video is not looking right and can skip that job and try again. This prompt and setting worked about 80% of the time so very few runs had to be canceled.

Source video:


Final result, there was a bit of a slowdown effect on this one but did not appear on most other runs:


Image of Chun-Li for
Image of Luna Snow for
Source video above for
 
Back
Top