Oh my goodness, it took me (and AI) forever to figure out how to get Minimax H3 (video_minimax_h3_r2v) to make my character speak the dialogue I had uploaded to Load Audio.
Seedance 2.5 did the audio dialogue fine, but it didn’t use the image I wanted as the first frame. Minimax did both with the new prompt.
Of course, this audio dialogue is for the little drone, Seven, so there’s no lip sync. That’s a task for tomorrow because my main character, Helen, speaks next. If there’s something special that needs to be done for that, I’ll add another post.
Aside from having the Load Audio wired incorrectly at first, the main issue is the prompt. This page explains how to prompt Minimax H3: docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md
Here’s the prompt I used (with the help of Gemini):
subject_definitions:
<Subject 1> is the female engineer and the gritty spaceship engineering environment shown in <Picture 1> and <Picture 2>.
<Subject 2> is the palm-sized mechanical drone shown in <Picture 3>.
<Audio 1> is the complete audio track containing the spoken dialogue of <Subject 2> (S1).
summary:
[reference generation + audio reuse] The target video begins by matching <Subject 1> and <Subject 2>, showing the female engineer holding the palm-sized mechanical drone in her hand inside a gritty spaceship engineering room, and uses <Audio 1> for dialogue.
retention_analysis:
<Subject 1> (appears in [Shot 1]): style_preserved - preserve the female engineer's appearance, clothing, pose, environment, gritty lighting, and cinematic photorealistic style.
<Subject 2> (appears in [Shot 1]): fully_preserved - preserve the drone's exact appearance, proportions, mechanical details, blue optical indicator, and position resting in the engineer's palm.
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.
detailed_description:
The target video uses a cinematic, photorealistic style with gritty science-fiction lighting and a realistic industrial spaceship environment. [Shot 1] The camera performs a slow, controlled cinematic zoom toward <Subject 2>, which remains resting in the engineer's palm. Keep the engineer relatively still while maintaining natural subtle breathing and small movements.
<Subject 2> (S1) speaks the dialogue contained in <Audio 1>. While the drone speaks, its blue optical indicator responds naturally to its voice, with subtle changes in brightness, intensity, and flickering that feel like a mechanical equivalent of facial expression, in perfect sync as he says, <d>[English] Madam, if you fuse that processor bridge with a tremor like that, my cognitive functions will be reduced to that of a standard toaster.</d>
overall_soundscape:
<Audio 1> is reused as the spoken dialogue of <Subject 2> (S1) and should remain clearly audible. The dialogue is accompanied by subtle diegetic sounds of metal tools and the low continuous rumble of the spaceship engine.
non_diegetic_music:
N/A
The Video with the correct first frame and Seven speaking the dialogue I uploaded.
There is a split-second at the very beginning where there’s a brief section of speech. I don’t know why it did that, but my audio dialogue was 9 seconds, and I had the video length set at 11 seconds. Instead of re-running (it takes over an hour!), I’ll just fix it in Premiere Elements.
Here’s what AI said about it:
Rule: Always make your ComfyUI video duration match your audio duration perfectly.
- If your audio is 9 seconds, set your video duration to exactly 9 seconds (usually by doing the math: 9 seconds × 24 frames per second = 216 frames).
- If you want an 11-second video but only have 9 seconds of dialogue, open your audio file in a basic editor first and add 2 seconds of dead silence to the end of it so the .wav file is officially 11 seconds long before you put it into ComfyUI.
This is the ComfyUI workflow.

~ Connie

Leave a Reply