Research
Imagen Video and Phenaki: Fidelity or Time
A note on the early split between sharper video and longer narrative continuity.
Question
Viewed together, Imagen Video and Phenaki represent a fork in video generation in 2022: one path pursued sharper, more stable images; the other pursued longer duration and sequences of events.
Imagen Video uses cascaded video diffusion models to raise low-resolution output to higher resolution and frame rates. Its focus is fidelity: detail, resolution, frame rate, image quality, and overall watchability. Phenaki addresses time more directly, using prompt sequences to condition variable-length video that continues through a series of events.
Technical Reading
Both directions remain relevant.
Higher fidelity moves generation toward commercial use through better resolution, cleaner images, materials, and light. Without it, outputs remain research samples. Longer sequences move generation toward narrative: characters pass through events, scenes change, and prompts begin to work like storyboards.
The tension concerns compute and modeling priorities. Pursuing fidelity can concentrate resources on the quality of an individual clip. Pursuing long-form continuity requires managing relationships among characters, space, action causality, and prompt sequences.
Capability Boundary
Neither direction was a complete answer at the time. Imagen Video could improve short-clip quality without managing a plot. Phenaki could drive longer videos with prompt sequences while still lacking fine visual quality and control. Together they establish three evaluation dimensions:
fidelity: does an individual shot hold together?
time: can multiple events continue coherently?
control: can the creator specify how things change?
Judging only one dimension can misrepresent a model.
Lab Judgment
These milestones encouraged me to ask which timescale a model handles well before asking which model is better.
single frame: art direction, composition, material
single shot: action, camera, short-range consistency
shot sequence: character, space, event, transition
story span: narrative memory, rhythm, causality
Video-generation progress fills in these timescales unevenly. A model with real production value needs to be dependable at least at the single-shot and shot-sequence levels.