Concept studies
Video pipeline
A seven-step pipeline run three times on personal concept studies and logged down to the render.
- Status
- Three concept studies from 2026, with no commercial value.
The problem
AI video fails in specific, repeatable ways: wrong language in speech, invented on-screen text, products turning into other objects and audio that overlaps.
How it runs
1Concept
A one-line idea with a tone and a world: claymation energy, Saigon in the 1960s, a rice-farming pun about teeth.
2Script
A script enhancer expands the idea into seven timed panels with voice and sound cues, then one or two rounds of edits on tagline and lines.
Tool or model: tv-ad-script enhancer
3Character
The lead character is generated or built from a reference photo, up to three rounds, until the look holds.
Tool or model: Recraft V4.1, GPT Image 2, Soul 2.0
What it taught: Soul 2.0 cannot take a prompt and a reference together, so GPT Image 2 is used whenever a reference photo is needed.
4Location
Sets are generated as stills first: a miniature office lobby, a 1960s Saigon crossroads, a retro kitchen.
Tool or model: Recraft V4.1, GPT Image 2
5Video
Four renders per film. Each render fixes what the last one broke: wrong packaging, invented on-screen text, overlapping audio, a physics error while brushing teeth.
Tool or model: Seedance 2.0
What it taught: A product reference was reinterpreted as another object, a noodle pack turned into a handbag. The prompt must state exactly what is and is not allowed.
6Audio
Voice-over, music and foley are made separately and layered, so each can be fixed without redoing the rest.
Tool or model: Seed Audio, Sonilo Music, Mirelo TTA, sync_so
What it taught: Seedance’s native audio handles Vietnamese badly. Voice is generated separately and laid over the original video audio, never replacing its timing.
7Assembly
Layers are mixed with exact timestamps: voice, music and each foley hit placed by delay, with the original audio ducked low.
Tool or model: ffmpeg, CapCut
What exists today
- The same seven steps run on Cao Sao Vàng, Miliket and Hynos Phosphaté.
- Each film rendered four times (V1 to V4), with the failure mode and the fix recorded.
- Voice, music and foley built as separate layers and mixed with exact timestamps.



Brands belong to their owners, with no sponsorship or endorsement. The song "Sài Gòn Đẹp Lắm" belongs to its rights holders.
Next caseAgentic AI