Skip to content
Vincent Bui
All cases

Concept studies

Video pipeline

A seven-step pipeline run three times on personal concept studies and logged down to the render.

Status
Three concept studies from 2026, with no commercial value.

The problem

AI video fails in specific, repeatable ways: wrong language in speech, invented on-screen text, products turning into other objects and audio that overlaps.

How it runs

  1. 1Concept

    A one-line idea with a tone and a world: claymation energy, Saigon in the 1960s, a rice-farming pun about teeth.

  2. 2Script

    A script enhancer expands the idea into seven timed panels with voice and sound cues, then one or two rounds of edits on tagline and lines.

    Tool or model: tv-ad-script enhancer

  3. 3Character

    The lead character is generated or built from a reference photo, up to three rounds, until the look holds.

    Tool or model: Recraft V4.1, GPT Image 2, Soul 2.0

    What it taught: Soul 2.0 cannot take a prompt and a reference together, so GPT Image 2 is used whenever a reference photo is needed.

  4. 4Location

    Sets are generated as stills first: a miniature office lobby, a 1960s Saigon crossroads, a retro kitchen.

    Tool or model: Recraft V4.1, GPT Image 2

  5. 5Video

    Four renders per film. Each render fixes what the last one broke: wrong packaging, invented on-screen text, overlapping audio, a physics error while brushing teeth.

    Tool or model: Seedance 2.0

    What it taught: A product reference was reinterpreted as another object, a noodle pack turned into a handbag. The prompt must state exactly what is and is not allowed.

  6. 6Audio

    Voice-over, music and foley are made separately and layered, so each can be fixed without redoing the rest.

    Tool or model: Seed Audio, Sonilo Music, Mirelo TTA, sync_so

    What it taught: Seedance’s native audio handles Vietnamese badly. Voice is generated separately and laid over the original video audio, never replacing its timing.

  7. 7Assembly

    Layers are mixed with exact timestamps: voice, music and each foley hit placed by delay, with the original audio ducked low.

    Tool or model: ffmpeg, CapCut

What exists today

  • The same seven steps run on Cao Sao Vàng, Miliket and Hynos Phosphaté.
  • Each film rendered four times (V1 to V4), with the failure mode and the fix recorded.
  • Voice, music and foley built as separate layers and mixed with exact timestamps.
Video pipeline 1Video pipeline 2Video pipeline 3

Brands belong to their owners, with no sponsorship or endorsement. The song "Sài Gòn Đẹp Lắm" belongs to its rights holders.

Next caseAgentic AI

Get in touch

Pick whichever is easiest for you.