EvieMoss
Lurker
- Joined
- May 4, 2026
- Posts
- 12
- Reaction score
- 0
As others have stated, prompt structure is crucial. Here's a distilled version of the prompt writing guides we can refer to for both the base and ref models. Good luck, have fun.
Base Model
The base model uses a two-part structure: a specific alignment instruction (for keyframe tasks) followed by three core fields.
Supports zero, one, or two input images.
No image input: Text-to-video mode
One image input: First-frame-to-video or last-frame-to-video generation
Two image inputs: First-and-last-frame-to-video generation
T2VA (Text-to-Video-Audio)
No alignment instruction is used for T2VA.
integrated_multimodal_description: [Shot 1] [Style], [Initial Composition], [Subject Appearance/Position], [Environment/Props]. [Camera Motion Type] with [Amplitude] at [Speed] [Direction/Target]. [Subject ID] ([Speaker ID]) [Action/Delivery]: [Language] [Verbatim Dialogue]. [Shot 2] At [00:SS.mmm], the camera cuts to [Composition]. [Action/Audio Continuity].
overall_soundscape: [1–4 sentences summarizing ambient sounds, physical action sounds, and non-verbal human sounds across the entire video].
non_diegetic_music: [1–3 sentences describing instrumentation, speed, rhythm, and dynamic changes. No abstract mood words].
I2VA (Image-to-Video-Audio)
For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] [Style derived from image], [Composition from ]. [Subject/Scene anchors from ]. [Action onset] $\rightarrow$ [Continuous development] $\rightarrow$ [Result/Reaction]. [Camera Motion Type] with [Amplitude] at [Speed]. [Subject ID] ([Speaker ID]) [Action]: [Language] [Verbatim Dialogue].
overall_soundscape: [1–4 sentences of ambient and physical sounds].
non_diegetic_music: [1–3 sentences of audience-only score].
FL2VA (First-and-Last-frame-to-Video-Audio)
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot [N]) aligns with the [S.SS]-second mark of the target video.
integrated_multimodal_description: [Shot 1] [Style derived from images], [First-frame state from Picture 1]. [Observable intermediate changes] $\rightarrow$ [Progressively narrowing differences] $\rightarrow$ [Last-frame state from Picture 2]. [Camera Motion Type] with [Amplitude] at [Speed].
overall_soundscape: [1–4 sentences of ambient and physical sounds].
non_diegetic_music: [1–3 sentences of audience-only score].
L2VA (Last-frame-to-Video-Audio)
How the reference pictures align with the target video — (from [Shot N]) aligns with the [S.SS]-second mark of the target video.
integrated_multimodal_description: [Shot 1] [Style derived from image], [Plausible preceding state]. [Explicit action and transition path] $\rightarrow$ [Gradual convergence in final shot] $\rightarrow$ [Landing on ]. [Camera Motion Type] with [Amplitude] at [Speed].
overall_soundscape: [1–4 sentences of ambient and physical sounds].
non_diegetic_music: [1–3 sentences of audience-only score].
Ref Model (Full-Reference Mode)
The omni-reference model uses a six-section structure. It replaces the "Instruction" line with explicit reference labels (``, ``, etc.) throughout the analysis and description.
Supports multi-modal reference inputs:
Images: ≤ 9 images
Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
Audio: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
Mixed inputs: Maximum number of files across all input types is 12
General Full-Reference Scaffold
This scaffold covers Reference Generation, Keyframe Completion, and Video Editing/Continuation.
subject_definitions:
is [description of person/object/scene] from [/
is [first frame/keyframe/anchor] for [Shot N], showing [compositional detail].
is [voice-timbre/signal reference] for ([Speaker ID]), containing [audio characteristics].
summary:
[[Task Type: keyframe completion / reference generation / video editing / video continuation / audio reuse / audio reference]] [Short paragraph summarizing target video, main reference relationships, and shot flow using defined labels].
retention_analysis:
(appears in [Shot X]): [marker: fully_preserved / partially_preserved / attribute_transfer / weak_reference] - [explanation of what is kept/changed].
([Shot X] anchor): [marker] - [explanation].
: [marker: fully_copy / partially_copy / reference / weak_reference] - [explanation].
detailed_description:
[1-2 sentences establishing overall style, lighting, and color palette].
[Shot 1] [Initial composition]. [First appearance of ] [referenced characteristics, position, action]. [Camera Motion Type] with [Amplitude] at [Speed]. ([Speaker ID]) [Action]: [Language] [Verbatim Dialogue].
[Shot 2] At [00:SS.mmm], the shot cuts to [Composition]. [Reference to or
overall_soundscape:
[1-4 sentences of ambient/physical sounds. Cite if signal is copied/referenced].
non_diegetic_music:
[1-3 sentences of audience-only score. Cite if signal is copied/referenced].
Scenario-Specific Adaptations for Ref Model
1. For Keyframe Completion (I2VA/FL2VA/L2VA logic):
Base Model
The base model uses a two-part structure: a specific alignment instruction (for keyframe tasks) followed by three core fields.
Supports zero, one, or two input images.
No image input: Text-to-video mode
One image input: First-frame-to-video or last-frame-to-video generation
Two image inputs: First-and-last-frame-to-video generation
T2VA (Text-to-Video-Audio)
No alignment instruction is used for T2VA.
integrated_multimodal_description: [Shot 1] [Style], [Initial Composition], [Subject Appearance/Position], [Environment/Props]. [Camera Motion Type] with [Amplitude] at [Speed] [Direction/Target]. [Subject ID] ([Speaker ID]) [Action/Delivery]: [Language] [Verbatim Dialogue]. [Shot 2] At [00:SS.mmm], the camera cuts to [Composition]. [Action/Audio Continuity].
overall_soundscape: [1–4 sentences summarizing ambient sounds, physical action sounds, and non-verbal human sounds across the entire video].
non_diegetic_music: [1–3 sentences describing instrumentation, speed, rhythm, and dynamic changes. No abstract mood words].
I2VA (Image-to-Video-Audio)
For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] [Style derived from image], [Composition from ]. [Subject/Scene anchors from ]. [Action onset] $\rightarrow$ [Continuous development] $\rightarrow$ [Result/Reaction]. [Camera Motion Type] with [Amplitude] at [Speed]. [Subject ID] ([Speaker ID]) [Action]: [Language] [Verbatim Dialogue].
overall_soundscape: [1–4 sentences of ambient and physical sounds].
non_diegetic_music: [1–3 sentences of audience-only score].
FL2VA (First-and-Last-frame-to-Video-Audio)
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot [N]) aligns with the [S.SS]-second mark of the target video.
integrated_multimodal_description: [Shot 1] [Style derived from images], [First-frame state from Picture 1]. [Observable intermediate changes] $\rightarrow$ [Progressively narrowing differences] $\rightarrow$ [Last-frame state from Picture 2]. [Camera Motion Type] with [Amplitude] at [Speed].
overall_soundscape: [1–4 sentences of ambient and physical sounds].
non_diegetic_music: [1–3 sentences of audience-only score].
L2VA (Last-frame-to-Video-Audio)
How the reference pictures align with the target video — (from [Shot N]) aligns with the [S.SS]-second mark of the target video.
integrated_multimodal_description: [Shot 1] [Style derived from image], [Plausible preceding state]. [Explicit action and transition path] $\rightarrow$ [Gradual convergence in final shot] $\rightarrow$ [Landing on ]. [Camera Motion Type] with [Amplitude] at [Speed].
overall_soundscape: [1–4 sentences of ambient and physical sounds].
non_diegetic_music: [1–3 sentences of audience-only score].
Ref Model (Full-Reference Mode)
The omni-reference model uses a six-section structure. It replaces the "Instruction" line with explicit reference labels (``, ``, etc.) throughout the analysis and description.
Supports multi-modal reference inputs:
Images: ≤ 9 images
Videos: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
Audio: ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds
Mixed inputs: Maximum number of files across all input types is 12
General Full-Reference Scaffold
This scaffold covers Reference Generation, Keyframe Completion, and Video Editing/Continuation.
subject_definitions:
is [description of person/object/scene] from [/
is [first frame/keyframe/anchor] for [Shot N], showing [compositional detail].
is [voice-timbre/signal reference] for ([Speaker ID]), containing [audio characteristics].
summary:
[[Task Type: keyframe completion / reference generation / video editing / video continuation / audio reuse / audio reference]] [Short paragraph summarizing target video, main reference relationships, and shot flow using defined labels].
retention_analysis:
(appears in [Shot X]): [marker: fully_preserved / partially_preserved / attribute_transfer / weak_reference] - [explanation of what is kept/changed].
([Shot X] anchor): [marker] - [explanation].
: [marker: fully_copy / partially_copy / reference / weak_reference] - [explanation].
detailed_description:
[1-2 sentences establishing overall style, lighting, and color palette].
[Shot 1] [Initial composition]. [First appearance of ] [referenced characteristics, position, action]. [Camera Motion Type] with [Amplitude] at [Speed]. ([Speaker ID]) [Action]: [Language] [Verbatim Dialogue].
[Shot 2] At [00:SS.mmm], the shot cuts to [Composition]. [Reference to or
overall_soundscape:
[1-4 sentences of ambient/physical sounds. Cite if signal is copied/referenced].
non_diegetic_music:
[1-3 sentences of audience-only score. Cite if signal is copied/referenced].
Scenario-Specific Adaptations for Ref Model
1. For Keyframe Completion (I2VA/FL2VA/L2VA logic):
- Summary: Use `[keyframe completion + reference generation]`.
- Detailed Description: Instead of the "Instruction" line, use phrasing like "the shot begins from ", "the shot's keyframe corresponds to ", or "the shot ends on ".
- Summary: Use `[video editing]` or `[video continuation]`. Begin the summary text with: "The target video is an edited version of
- Detailed Description: Cite `
- Subject Definitions: Ensure every vocal source has a mapping: ` is the voice-timbre reference for `.
- Detailed Description: Use `` for physically speaking characters and `` for cues within a reused BGM/Soundtrack.
