How to make a long-form Vox-style AI video explainer?
Contents: The two variants of Vox style
A long-form Vox-style AI video explainer is a 5 to 20 minute animated explainer in the aesthetic Vox popularized on YouTube and Netflix. The style comes in two related variants: an editorial variant built from archival footage with yellow highlighter marks and graph overlays, and a paper-collage variant built from halftone-cutout characters over aged newsprint backgrounds. Scenema’s Vox Explainer style preset produces the paper-collage variant from a short text prompt. This article walks through the two variants of Vox style, what defines the paper-collage variant visually, and how to produce your own six-minute long-form Vox-style piece on Scenema in about fifteen minutes.
Definition
Vox style (n.) An animated explainer video aesthetic in one of two variants: (1) an editorial variant of archival footage annotated with yellow highlighter marks, bar graph overlays, and kinetic typography, or (2) a paper-collage variant of halftone-cutout characters over aged newsprint backgrounds with hot-red accent typography. Both share a fast confident data-forward narrator and a 3-to-8-second cutting rhythm. Named for Vox’s Explained series on YouTube and Netflix.
The two variants of Vox style
Vox has developed two distinct visual approaches to explainer video. Both are instantly recognizable as Vox and both are widely taught in tutorial channels as “Vox style,” but they are produced very differently.
The editorial variant builds a piece from archival photographs, historical footage, and stock imagery. The signature move is a yellow highlighter mark that scrolls across text or annotates paused footage. Bar and line graphs overlay the footage. Kinetic typography enters and exits with quick springs. Circle-and-arrow annotations point out details on paused frames. This variant feels like a well-edited news documentary and is typically produced in After Effects on top of licensed archival material.
The paper-collage variant builds a piece from halftone-cutout characters and objects placed on aged newsprint or brown paper backgrounds. The signature move is a hot-red accent color applied to statistics and pull quotes. Depth comes from layering, not perspective. This variant feels like a magazine spread or scrapbook.
Both variants share the same DNA: a fast confident data-forward narrator, a 3-to-8-second cutting rhythm, a chapter structure for the argument, and bold typography for headlines and data.
Scenema’s Vox Explainer style preset produces the paper-collage variant. The editorial-and-highlighter variant is not currently supported by the preset and requires a traditional archival-footage-plus-After-Effects workflow. The rest of this article covers the paper-collage variant in detail.
What defines the paper-collage Vox style
A viewer can identify a paper-collage Vox-style explainer within the first three seconds. Six visual elements do the work.
Paper-collage compositions. Foreground elements sit on top of textured background layers as if physically cut out and pasted onto a page. Objects have visible edges. Depth comes from layering rather than perspective. The effect is closer to a scrapbook or magazine spread than a traditional animated illustration.
Halftone-cutout characters. Characters are rendered in a halftone dot pattern (like a newspaper photograph) with a silhouette-first approach. Face detail is deliberately minimal. Identity comes from posture, silhouette, and consistent rendering across shots, not from photorealistic likeness.
Aged newsprint or brown paper backgrounds. The base layer is nearly always a warm off-white paper texture with visible grain, sometimes with printed text bleeding through. The palette leans warm and low-saturation, which makes the red accents pop.
Hot-red accent color. A single saturated red does most of the visual work: highlighting key data, marking pull quotes, punctuating transitions with an ALERT WASH sweep, and drawing the eye to the argument. No other color competes for attention.
Bold typography for headlines and data. Chunky serifs or heavy sans-serifs, often with a slight vintage character. Numbers get overshoot-spring entries. Pull quotes fill the frame. Text is a visual element, not a caption.
A confident narrator over a fast cutting rhythm. Shots change every three to eight seconds. The narrator delivers data-forward argument with clear cadence. Silence is used deliberately. The pacing signals confidence in the argument rather than urgency.
Combined, these elements produce an aesthetic that reads as journalism first and animation second. That is the point.
Why the Vox style is hard to produce traditionally
Producing a six-minute Vox-style piece the traditional way is expensive because each element has to be crafted by hand and maintained consistently across dozens of shots.
A hand-cut paper collage aesthetic requires physical or digital prep of every character and object. Halftone rendering has to hold across every appearance of every character or the illusion collapses. The red accent has to be applied consistently for it to work as a visual signal rather than a decoration. The narrator has to be recorded in a single continuous session, with a specific cadence, to avoid perceptible seams.
A traditional agency producing a comparable six-minute Vox-style piece charges $10,000 to $50,000 over a 6 to 12 week turnaround. Most of that cost is the visual continuity work, not the concept.
How to make a Vox-style AI video on Scenema in four steps
The workflow on Scenema collapses the traditional agency process into a text prompt and a style preset selection.
- Write a short text prompt (30 to 200 words) describing the topic, tone, and any specific data or references you want covered. The FIFA piece embedded below was produced from 79 words.
- Select the Vox Explainer style preset from the Scenema style library. The preset commits the treatment agent to a canonical Vox palette (aged newsprint tan, charcoal black, hot-red accent), the paper-collage motif vocabulary (halftone cutouts, staggered stat numbers with overshoot springs, ALERT WASH transitions), and the narrator cadence conventions.
- Review the treatment the tool produces. This is where you catch structural issues (the argument order, which data the piece leads with, how many chapters) before spending credits on shot generation.
- Approve and export. The pipeline generates entity manifests for any recurring characters or objects, produces canonical reference images for each entity, writes and generates every per-shot keyframe with the Vox motif language injected into the prompt, records the narration in a single continuous voice pass, and muxes the final MP4.
Total time end-to-end for a six-minute piece: about fifteen minutes.
Case study: The 13 Billion Dollar Feature
The video at the top of this page is a worked example. Six minutes, 54 shots, produced from a 79-word text prompt and one style preset selection.
The prompt described a data-driven documentary on the financial mechanics of the 2026 FIFA World Cup. The Vox Explainer preset was selected. The pipeline produced a treatment (an eight-act structure with an eleven-paragraph narration script) which the reviewer approved without edits. Entity manifests were generated for four recurring elements: a halftone-cutout figure of a New Jersey worker, the Final Stadium, the Aramco corporate logo, and a Category 1 ticket. Each of the 54 shots then rendered with the Vox motif language injected into the per-shot prompt.
The New Jersey worker character appears in roughly a third of the shots without drifting. The aged-newsprint background holds across every scene. The hot-red accent is applied consistently to statistics, ALERT WASH transitions, and pull quotes. The narrator delivers all 1,100 words in a single continuous voice.
Every prompt, reference image, and shot is inspectable in the public template viewer. The full pipeline breakdown walks through each stage in detail.
Total cost of production: less than $100 in Scenema credits at entry-level pricing.
Six things to include in a Vox-style prompt
A good Scenema prompt for a Vox-style piece includes six specific elements. Each is a testable question: did I include this?
1. State the tone explicitly. Vox is dry, analytical, and critical, not chatty or promotional. Write the words in the prompt. The FIFA piece used “dry, clinical, and data-driven documentary-style.”
2. Include at least three specific statistics or data points. Vox is data-forward. Without numbers in the prompt, the treatment agent has no material to anchor the paper-collage graphics on. The FIFA prompt named ticket prices, revenue cycle amounts, and broadcasting deal figures.
3. Name specific companies, brands, organizations, or people. These become the entities Scenema generates canonical references for. The FIFA prompt named Fox, Telemundo, BBC, ITV, and Aramco. Without named entities, characters and objects will not recur across shots.
4. Include a clear thesis or angle. Vox pieces argue a point. State yours. The FIFA thesis: “the disparity between corporate profits and the average worker.”
5. Anchor a timeframe. Past decade, 2023 to 2026 cycle, since 2008. A timeframe gives the treatment agent chronological structure for the argument.
6. Frame as journalism, not marketing. If the prompt reads like a product brief, the output reads like one. Use documentary, investigative, or analytical framing words.
Two short prompt examples that hit every item
The two prompts below each land every item on the checklist above in about 50 words. Paste either into Scenema with the Vox Explainer preset selected. Each produces a six-to-eight-minute finished piece in about fifteen minutes.
Cellular biology explainer
A dry, analytical Vox-style explainer on why "mitochondria is thepowerhouse of the cell" became a meme. Cover the 1957 textbook line,Lynn Margulis's 1967 endosymbiosis paper, ATP production numbers, andthe phrase's Tumblr rise around 2013. Focus on the gap between thebiology and the joke.Business economics explainer
A dry, data-driven Vox-style explainer on why the Costco hot dog isstill $1.50 in 2026. Cover the 1985 pricing decision, Costco'sloss-leader math, membership renewal rates, and CEO Craig Jelinek'snow-famous threat to the buying team. Focus on how a $1.50 pricesignals brand promise more than food economics.The TL;DR on Vox-style AI video
What is a Vox-style AI video? An animated explainer video with a paper-collage aesthetic, halftone-cutout characters, aged newsprint backgrounds, hot-red accent typography, and a confident data-forward narrator. Named for Vox’s Explained series.
Can AI actually produce this style? Yes. Scenema ships a Vox Explainer style preset that commits the treatment agent to the palette and motif language up front, which is what keeps every shot visually coherent. The FIFA piece above is a six-minute example, produced end-to-end from a 79-word prompt.
How long does it take? About fifteen minutes for a six-minute finished piece on Scenema.
How much does it cost? The FIFA piece cost less than $100 in Scenema credits. See current Scenema pricing for the plan structure.
Can I customize the palette or motifs? The preset is opinionated in service of the Vox aesthetic. Palette and motif language are committed up front to hold the style across every shot. For fully custom art direction, register your own reference as a custom style, or fall back to an assembled pipeline.
Do I need to know how to write scripts? No. The treatment agent generates the script from your topic prompt. You review and edit the treatment before shot generation begins.
Where to go next
- Read the full FIFA tutorial for the step-by-step pipeline breakdown.
- Understand what a long-form explainer video is in what is a long-form explainer video.
- See how Scenema holds character consistency across shots in how to keep character consistency in AI video generation.
- Compare AI tools for long-form explainer work in best AI tools for long-form explainer videos.
- Try Scenema at scenema.ai.