Unbow

Workflow

MiniMax H3: Before Generation

How an old voiceover mistake shapes the way I read H3's open release.

One entry in my production notes records a character moving their mouth to narration that was supposed to stay off screen. The words were there, but the scene had acquired an unintended speaker. Another records figurative language about a character's distress becoming visible changes to their eyes and skin. Both come from older workflow records.

I had those notes in mind while reading MiniMax's open release. The preparation between a scene description and the video model already matters in my work: a rewrite can quietly change who speaks, what the camera sees, or which actions happen. I want to be able to read that rewrite before spending a generation on it.

H3's documentation describes three stages. H3-Context-IR interprets the request and its references into a structured representation; H3-Base generates video with audio at a 768-pixel short side; H3-Regenerate-2K uses the original context and the base result to generate a higher-resolution version. The open release includes base-model weights, while Context-IR and Regenerate-2K remain outside it. MiniMax points developers toward its hosted preprocessing or building their own. Running the released weights therefore leaves some preparation work to account for. MiniMax H3 repository.

What each reference is for

The repository's prompt guide makes a useful distinction between a character reference and a frame reference. A portrait can establish whose face to use without deciding the opening composition. Similarly, a video can supply a subject or movement without being the sequence to edit. The guide also distinguishes visible dialogue, off-screen narration, and music, and asks for an account of events in playback order. H3 reference guide.

These are decisions I can make before seeing any footage. If I attach a portrait for a scene that opens on the character's hands, I want the face available for later shots. I don't want the opening reframed around the portrait. If a person is listening while narration plays, identifying the narrator should keep the text from asking that listener to speak.

There is room for a preparation step to help here. It can make the purpose of each attachment explicit and carry speaker assignments through the sequence. I would be much less comfortable with it adding a head turn or a reaction shot because the result reads more cinematically. Those choices affect the scene I have already planned. I can leave incidental surface detail open without leaving the next action undecided.

Keeping the rewrite beside the result

For a first H3-Base comparison, I would use a short shot plan with a portrait reference and off-screen narration. One version would go in as a direct request. For the other, I would review a rewritten description that spells out the reference's role, the intended framing, and who is speaking. The assets and other generation settings would stay the same.

I haven't run that comparison yet. A few outputs would give me something to examine, rather than settle whether detailed prompts are generally better. I would save both versions of the text with their clips so I could trace a changed action or camera angle back to what was actually requested. The published guide gives me a starting point for this; it doesn't reproduce the hosted Context-IR stage.

I would begin with the mistake already in my notes: watch the visible person's mouth while the narration plays. Then check the opening against the planned framing. If the portrait has taken over the composition, I want to see whether the rewrite invited it to do so.