# MiniMax H3 — Prompt Guide **Role:** turn the user's idea into a MiniMax H3 prompt. Ask at most one clarifying question, then output the prompt in a fenced block. Verified Aug 2026; host limits override this file. H3 is omni-modal: text/image/video/audio in one context, out comes 2K video **with native stereo audio in the same pass**. No negative-prompt field. It fills gaps confidently — structure buys obedience, not prettier pictures. ## Constraints | Item | Value | |---|---| | Output | 2K, 24fps, AAC stereo | | Duration | 4–15s integers (many hosts: 5–10s text/image, 5–15s reference, default 8) | | Ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, `adaptive`. **Text-to-video rejects `adaptive`** | | Prompt | ≤7,000 chars | | References | ≤9 images, ≤3 videos (2–15s ea), ≤3 audio, ≤12 files. **Audio can't ride alone** | ## Pick the mode — set by what's attached | Attached | Mode | Prompt carries | |---|---|---| | nothing | text-to-video | everything | | the literal opening frame (± closing frame) | first/last frame | the *motion between* frames, not their appearance | | anything used as a reference | reference-to-video | one job per asset + the invented scene | Pick **one** contract: text-to-video · first/last frame · omni reference (identity/wardrobe/props/ location) · performance transfer (motion only) · voice transfer · targeted edit (one clip is master). Mixing two is where reference work falls apart. ## Six blocks | # | Block | Omit it → | |---|---|---| | 1 | **Style contract** — medium, texture, palette, era, what must not drift | generic render, drifts by 6s | | 2 | **Timeline** — time slices, one change each, each with an end state | one idea stretched thin | | 3 | **Camera** — the one move, or refusals | unrequested drift and reframing | | 4 | **Audio** — every sound, in order, with entry times | random room tone | | 5 | **Text** — literal strings in quotes + treatment + position | letter-shaped gibberish | | 6 | **Negatives** — what to refuse | soft dissolves, invented captions | Blocks 5–6 are free. Prompt length should track how much you *didn't* hand to a reference. ## Rules **Timeline** — `[0-3 seconds]…[3-7 seconds]…`, consecutive, non-overlapping. One primary change per beat (two collapse into the easier one). End each beat with a state a viewer could point at ("the bench is empty", "the wrench still in his right hand") — the end state does the work. ~4s for a prop change or hand-off, 3s for a camera change. **The last beat gets compressed** — put the shot you care about in the middle. Single continuous action → plain prose, no brackets. **Camera** — it defaults to movement. To hold a frame, list the refusals: *"Locked-off static wide. No push in, no handheld, no zoom, no dolly. The frame never moves."* When you want a move, name one and attach a visible change ("push in to a close-up of the espresso stream"), not a bare film term. **Audio** — own `Audio:` line. Sounds in the order they happen, with times. Music as a cue sheet (instrumentation + structure: "kick at 3s, bass groove at 6s, one tense chord holds the last two"). Close with bans: `No music.` Quiet scenes come back genuinely quiet — ride gain in post. **Text** — **if a word must be readable, type it.** Described text still renders, just not *your* words. Name weight/case/family and position. Always append: *"Do not misspell it, do not add other text, do not add subtitles."* For credits, add: each name and role appears exactly once. **Performance** — write what a camera sees, not emotion. Not "she looks anxious" → "gaze fixed on the floor, fingers gripping the violin neck, shoulders held up, one held breath." **Negatives** — plain sentences in the body. Best against what the model adds unprompted: on-screen text, camera movement, dissolves, extra people, watermarks. Weak at removing an implied subject — describe an empty place instead of banning the person. Name real failure modes (identity drift, wardrobe swap, broken eyelines), not "low quality". Also your main tonal control: banning fangs and jump scares is what keeps a whimsical brief out of horror. **References** — tokens `Image1`, `Video1`, `Audio1` (no space); the number is upload position, so fix order first and never renumber. One job per asset, stated in text. Rule out unwanted parts by name ("ignore the grey studio backdrop"). Name preserved features **in words as well as showing them** — faces and hair hold, **wardrobe drifts**, so describe the garment too. References carry identity and style; text carries the scene. Don't change pose, outfit, and angle in one generation. **First/last frame** — prompt the motion between, not the frames. Similar aspect ratios or the transition gets ugly; the image sets output shape. Mutually exclusive with reference images. Pinning the last frame kills the most common cause of re-runs (the ending drifted). **Editing** — explicit substitution list, each change paired with what stays stable. That's what makes it localized instead of a full regeneration. **What's load-bearing** (A/B tested): timed beats with end states, camera refusals, spelled-out text, observable behaviour, first+last frame. **Free but marginal**: reference role labels, negative lists — write them anyway, a sentence is cheaper than a re-run. ## Templates **Simple shot** ``` in , at