text2vid img2vid H3 Minimax - Fully uncensored local AI video generator released today

there really isn't an exact prompt for all of this. With H3 you just feed the official prompting guide by minimax into an llm so that it can format your natural language idea into the correct format H3 needs to output the video.

The process first involved creating a character sheet for the female subject that shows her naked from the front, sides, back, a close up bust shot of her face. Then I had to find the outfit she's wearing through an image search on bing. Same thing for the hotel room. From there it was creating separate clips, cloning the voices from the clips so that the voices stayed consistent, and using video references for the doggystyle clips. Also used the audio from those clips as audio references for the moaning and sex sounds in those clips.

As you can see this shit is time consuming and that's not even mentioning the time it takes to generate each clip lmao. As usual it takes several generations to get the right ones.
 
does anyone have a good facial cumshot workflow for r2v? i tried with a couple of loras, giving videos as an example but the cum always spawns on the face and it just looks really flat. I could definitely use some help lol
 
Can anyone help wordsmith this prompt to improve it for i2v H3, trying to get a good missionary / creampie setup for 15s generation

The dick, face and action all look good but the camera angle is always basically at the girls feet looking towards her. I want something that would be a good POV angle shot like this

(m=eaSaaTbaAaaaa)(mh=E-yi0QFS30xGZUFg)9.jpg


For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action cinematic realistic footage, soft natural indoor lighting, shallow depth of field. Fully preserve the exact face, hair, skin tone, expression, clothing, body proportions, and mid-zoom framing of the young woman from . She remains centered in medium close-up, looking toward camera with a calm, slightly anticipatory expression. Subtle natural breathing motion and a soft blink. Camera holds steady for the first 2.5 seconds, then slowly pushes in slightly.

[Shot 2] At 00:03.000, hard cut to a high-angle overhead POV looking directly down onto the same woman lying on her back in missionary position. Camera is positioned almost straight above her, angled slightly toward her face and upper body so the viewer is looking down at her. Her face, hair, and body match exactly. Legs spread and raised toward the camera. A realistic erect penis enters from the bottom of frame and thrusts steadily into her vagina with clear penetration visible from above. Wet glistening skin, natural body movement, her hips gently rising to meet each thrust. Her expression shows strong arousal: eyes half-lidded looking up toward the camera, mouth open, brows relaxed in pleasure. She moans with pleasure, chest rising and falling. Soft bedroom lighting, warm skin tones, slight natural camera sway matching the thrusting rhythm. Continuous rhythmic penetration continues under the overhead view.

[Shot 3] At 00:11.000, the thrusting deepens and speeds slightly while the camera remains in the same high-angle overhead POV looking down at her. At 00:13.000 the man climaxes with a visible creampie: thick white semen pulses and overflows from her vagina around the still-penetrating penis, running down onto the sheets, clearly visible from the top-down angle. Her face immediately transforms into pure euphoric joy — eyes fluttering closed then opening wide with a blissful smile while still looking up at the camera, head tilting back slightly, a soft gasp of intense pleasure. Camera holds the intimate overhead POV framing as the final pulses of the creampie finish and she continues looking up in satisfied ecstasy. End on this close top-down view.

overall_soundscape: Soft intimate bedroom ambience, quiet breathing, wet skin-on-skin sounds of rhythmic penetration, her escalating moans of pleasure (breathy, high with arousal), a final deeper moan and soft gasp of joy at the moment of creampie. Subtle bed creaks matching the thrusts. No background voices or external noise.

non_diegetic_music: N/A
 
I had a feeling it was the lora's weights. It sucks because without those loras, the motions and anatomy just look odd sometimes. Looks like I just have to experiment more. Thanks for the reply.
 
Does anyone have a prompt or settings for creating an img2vid or ref2vid video that looks more "home-made"? I'm kind of tired of getting results that either have fake-looking movement or look like they were professionally shot :hahaa:
 
If you are generating locally and aren't using it yet; go and grab the 10eros max beta 4 model from huggingface. The model works amazing for actual sex stuff compared to the base models.
 
my bad I really don't know what that is. Currently i just switch between models in the workflow for whatever specific thing i need. For pure sex stuff the eros10 beta 4 is really good. For other stuff that needs better quality i use the ref2va and fl2va models
 
such an irony that to get an insane quality level for video you need a really good reference, turn out i2v is nothing without i2i
 
That's like a golden rule with generative AI. Your samples should always be high quality. The higher the resolution the better output you'll get. I typically use stuff from this site for references for my characters and hilarious how terrible all the stuff that OF girls put out. Low res and shit lighting. The videos are even worse. The issue is made worse when users here upload material via screenshots and screen recording
 
pruned_int8_convrot just refers to the original models that were tuned by comfy to save on VRAM, apparently there are some unnecessary weights, so pruned not distilled. hopefully eros10 did the same with theirs, but if it runs, it runs
 
i just used the ones posted by pixaroma on his website (just add dot com to his name)
 
given the quality requirement of reference, not to mentioned rigorous prompt system without chatbot i2t, these local i2v chinese models will still serve a very niche market, most user dont know how to make a great image input, grok made it so easy but it got taken away.
 
Yes it's not dumbed down like grok but you have more control in getting exactly what you want from camera angle down to every detail of the action. H3 doesn't handle the back end prompting like grok would do so using an llm to set up the formatting is must
 
thanks for the explanation, using eros10 cut my times in half with same quality.

ps: that llm do you use for prompts? most I tried block certain word
 
Does anyone know if there's some aspect ratio setting that just keeps the original photos ratio? If I have a photo that's quite tall it gets squished into more of a square with the current settings I can find.
 
Back
Top