新增 4 秒品牌封面:单台 DGX Spark,约 56 秒生成 5 秒视频+音频(warm run)。随后进入 Kapi Lai 的四段原生音画讲解:小草稿、latent 直接传递、精修与常驻、约 56 秒的 warm-run 结果。
完整 Prompt · 展开查看与复制
SOL-H3 SPARK — 卡皮巴拉完整 PROMPTS / 39 SECONDS
说明:4 秒后期封面 + 四段各 8 秒的 Ref2VA 原生音画 + 3 秒品牌片尾。四段模型 prompt 为实际使用的原文;封面、精确黑板文字和片尾由后期合成完成。
=== 1. 封面与整片剪辑说明 / EDITORIAL BRIEF ===
00:00–00:04 — Show the result before the explanation. Warm-paper background; two-line Sol-H3 / Spark at left, deep-green italic Spark and a four-point spark. At right, animate in GENERATE IN / ~56 s / 5-second video + audio / 768p output. Charcoal lower strip: ONE NVIDIA DGX SPARK, Warm end-to-end run, Kapi Lai explains how, and a small arrow. Fade the original score in, then down before the speaker enters. No new voiceover is used for this four-second cover.
00:04–00:12 — Clip 01: MiniMax-H3 384p draft, motion and audio together.
00:12–00:20 — Clip 02: VAE adapter, direct H3-to-LTX latent transfer.
00:20–00:28 — Clip 03: LTX + Sol-Attn, 768p refinement, cached refiner prompt and quantization.
00:28–00:36 — Clip 04: About 56 seconds for a five-second video with audio on one DGX Spark, warm end-to-end run.
00:36–00:39 — Reuse seconds 31–34 of the original spark-reveal-v5 brand ending. Keep its Sol-H3 Spark identity and official logo row. Mute the source ending audio; let the continuous score rise after the last dialogue.
Preserve Kapi Lai's identity, classroom, native generated voice and gestures. Typeset the exact labels on the right side of the chalkboard; keep the character and pointing paw visible. Normalize spoken audio, keep music quiet underneath, and preserve the complete existing 35-second body when adding the cover.
Exact board text:
01 — 384p / MINIMAX-H3 DRAFT / A compact first pass. / Motion + audio, together.
02 — VAE ADAPTER / H3 latents → LTX latents / Direct latent transfer. / No pixel decode / re-encode.
03 — 768p / LTX + SOL-ATTN / Cached refiner prompt / Quantized. Kept resident.
04 — ~56 s / WARM END-TO-END / 5-second video + audio / One NVIDIA DGX Spark.
Shared header: SOL-H3 / ON SPARK.
Shared footer: 384p draft > latent transfer > 768p detail.
=== 2. 实际推理配置 / GENERATION SETTINGS ===
MiniMax-H3 Ref2VA + LightX2V adapter: minimax_h3_ref2v_turbo_8step_v1.0_768p_bf16.safetensors.
Exactly 8 actual DiT evaluations. In the recovered Diffusers runner, num_inference_steps=9 creates 9 scheduler grid points and 8 evaluations.
LoRA rank 128, alpha 8, scale 1.0, effective alpha/rank 0.0625; fused before inference.
Video shift 6.0; audio shift 3.0.
Each clip: 1344 × 768, 192 frames, 24 fps, 8.000 s. Final edit: 1344 × 756, 24 fps.
Two image references and one audio reference, in the order below. Reference resize policy: match. Audio is a timbre reference for new dialogue.
The promotional footage was rendered on B200. The ~56-second product result refers to the separate Spark pipeline on one DGX Spark.
Official prompt-writing skill: https://github.com/MiniMax-AI/MiniMax-H3/blob/main/skills/h3-prompt-writing/SKILL.md
=== 3. 参考素材 / REFERENCE INPUTS ===
Public Kapi Lai reference video:
https://huggingface.co/spaces/Lawrence-cj/sol-h3-promo/resolve/main/kapi-lai-ref2va-50step-20s.mp4
<Picture 1>: character crop from the reference at 1.0 s (576 × 1089 before resize).
<Picture 2>: full classroom frame from the reference at 1.0 s (768 × 1344).
<Audio 1>: voice-timbre excerpt from 0.15–3.65 s, mono 32 kHz WAV. Do not reuse its words or soundtrack in place of new dialogue.
The old 50-step source provides reference appearance and voice only; all four new clips use the 8-step configuration above.
Brand ending source:
https://huggingface.co/spaces/Lawrence-cj/sol-h3-promo/resolve/main/spark-reveal-v5/sol-h3-spark-reveal-34s.mp4
=== 4. 四段完整原始 MODEL PROMPTS / VERBATIM ===
--- 01-draft / seed 202609110 / 8.000 s ---
subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.
summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is MiniMax-H3. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.
detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, the capybara notices the camera and raises his right forepaw into a small open welcoming gesture, keeping it below shoulder height. The gesture is slow and anatomically continuous; the same paw returns toward his side as he turns his gaze briefly toward the board. A faint chalk-like line develops into a simple landscape thumbnail on the board, showing only broad silhouettes and a consistent horizon. This is an illustrative draft, not an actual recorded intermediate output. The rest of the board stays quiet. At 00:00.600, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] One Spark. Here's how. MiniMax H three creates a small draft, with motion and sound.</d> He looks toward the viewer while naming the small draft, then glances toward the thumbnail during the final words about motion and sound. His muzzle opens only as far as the syllables require, with the same two incisors remaining attached and proportionate. Finish speaking by 00:07.200. During the final eight tenths of a second, he closes his mouth, gives a small acknowledging nod, and settles into a neutral stance. The board holds the completed simple thumbnail rather than switching to a new scene.
overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.
non_diegetic_music:
N/A
--- 02-transfer / seed 202609111 / 8.000 s ---
subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.
summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is VAE Adapter. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.
detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, the capybara turns his head slightly toward the board without moving his feet. The board shows two simple separated image panels connected by one fine line. The left panel retains a rough landscape contour, and the right panel begins with the same contour rather than a different scene. At 00:00.500, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] An adapter passes the latents straight to the refiner. No decoding and re-encoding pictures in between.</d> His right forepaw makes one compact pointing gesture toward the board's left edge while remaining entirely in the left portion of the frame. A small organized packet of lime light travels directly along the connection from left panel to right panel. The packet arrives once, accompanied by a soft dry click. Nothing explodes, dissolves into a cloud, or travels through a third image panel. The abstract visual represents direct latent transfer and makes no claim to show actual internal tensors. The capybara returns his gaze to the camera while explaining the omitted decode and re-encode. He finishes by 00:07.100, closes his mouth, and lowers the pointing paw smoothly. Hold the two connected panels and his steady attentive expression through the last frame.
overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.
non_diegetic_music:
N/A
--- 03-refine / seed 202609112 / 8.000 s ---
subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.
summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is LTX + Sol-Attn. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.
detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, a narrow green highlight passes slowly over the existing landscape thumbnail on the board. The thumbnail keeps its original horizon and shapes while a few fine edges become more distinct; the camera never enters the illustration. This is a conceptual refinement demonstration rather than a measured before-and-after reconstruction. At 00:00.500, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] Sol Attention accelerates refinement. A cached refiner prompt and quantization help keep both stages in memory.</d> Pronounce the visible method name Sol-Attn as Sol Attention, clearly and without spelling each letter. He makes one gentle open-palm presentation gesture with his right forepaw and then rests both paws near his belly, leaving his mouth and face unobstructed. While he mentions the cached refiner prompt, two small neutral rectangles illuminate beneath the main image. One represents the reused refinement context, and the other represents keeping both stages resident. Their contents remain simple visual shapes; exact explanatory labels will be added in the edit. His delivery stays matter-of-fact rather than triumphant. Finish speaking by 00:07.200, then close his mouth and let the board remain still. There are no numeric speedup counters or extra claims about skipping Stage 1 text encoding.
overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.
non_diegetic_music:
N/A
--- 04-result / seed 202609113 / 8.000 s ---
subject_definitions:
<Subject 1> is the stylized capybara in <Picture 1>, with coarse warm-brown fur, a broad elongated muzzle, small rounded ears, dark eyes, two pale rectangular front incisors, a rounded belly, short legs, and a relaxed upright posture.
<Subject 2> is the classroom environment in <Picture 2>, with an aged dark-green chalkboard, warm wooden floor, subdued wall texture, and a practical overhead fluorescent fixture.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1), using the spoken vocal layer from the existing Kapi Lai video; only its recognizable adult male timbre is referenced.
summary:
[reference generation + audio reference] Generate an 8.000-second landscape explanation with <Subject 1> in <Subject 2>, using <Audio 1> only for vocal timbre. The topic is 5-second video + audio / one DGX Spark. The target speech is new, synchronized to the capybara, and supports the original Spark product claims.
retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - preserve the capybara identity, fur, muzzle, ears, incisors, body proportions, and simple natural forepaws.
<Subject 2> (appears in [Shot 1]): partially_preserved - preserve the classroom materials and lighting while expanding the composition to landscape and replacing the original chalk writing with the new conceptual demonstration.
<Audio 1>: reference - use the original adult male vocal timbre for newly written dialogue; do not copy the original words, music, timing, or sound signal.
detailed_description:
The target video is an 8.000-second stylized 3D character explainer in a tactile classroom, with warm practical lighting, textured brown fur, and a restrained green accent. The humor comes from the capybara's calm physical presence, without an exaggerated meme performance.
[Shot 1] A static, eye-level medium-wide landscape shot establishes <Subject 1>, the warm-brown capybara from the reference, standing in the left quarter of the frame. His broad muzzle, small rounded ears, dark eyes, two pale incisors, rounded belly, and short legs remain recognizable. Keep his entire face and both forepaws visible. His feet stay planted on the wooden floor. <Subject 2>, the classroom from the reference, surrounds him, retaining the dark-green board, worn wall surfaces, and warm practical lamp. Adapt the room to a 1344-by-768 landscape composition. The board occupies the right two-thirds, faces the camera directly, and stays fixed without perspective changes. Keep the capybara, his gestures, and their shadows out of the main board area so the subsequent edit can place precise labels there. The board's top, left, and right edges remain visible throughout.
The overhead lamp creates a soft rim along the capybara's ears and shoulder fur, with gentle fill on his muzzle. His eyes keep a small stable catchlight. The camera does not cut, orbit, zoom, shake, or change focus. Preserve a little room above the ears and below the feet. The speaker follows the recognizable male timbre of <Audio 1> with a calm, clear conversational delivery and natural pauses; the original reference dialogue and soundtrack are not reused.
At 00:00.300, the capybara shifts his gaze from the now-settled board to the camera. His feet, body scale, and position remain consistent with the preceding scene. A single soft lime outline frames the central board area, preparing a simple result card; this visual will receive the exact performance text during editing. At 00:00.500, <Subject 1> (S1) says, following the voice timbre of <Audio 1>: <d>[English] On a warm run, five seconds of video and sound in about fifty-six seconds. One Spark.</d> He places a clear, unhurried emphasis on the words warm run. His right forepaw opens toward the board once while the other stays near his side. He does not count with additional fingers, clap, shout, or turn his back on the viewer. As he mentions the result, his eyes briefly follow the board and return to the camera. Keep his mouth naturally synchronized to all words, including the complete number fifty-six. He finishes the short phrase One Spark by 00:07.200, closes his mouth, and gives a small satisfied nod. Hold the relaxed capybara and quiet result board through 00:08.000, providing a clean edit point for the separately composited brand end card. No laughter track, roar, scream, confetti, or bright flash interrupts the final pause.
overall_soundscape:
Quiet classroom room tone and faint lamp hum remain behind the clear foreground voice. Add only subtle fur and paw movement and a soft dry click when an illustrated state settles; the ambience stays consistent through the final pause.
non_diegetic_music:
N/A