makevideowith.ai
← Back to videosEducation / 2102844654169575547

History of AI documentary short film

@kimmonismus ↗
THE INSTRUCTIONS SHARED BY THE CREATOR

The creator’s prompt

A prompt is the text that tells the AI what you want it to make. Copy this one and paste it into your Claude Code project, adding what you’d like to change: the subject, wording, colors, or length.

The original text is preserved to respect the creator’s work. You can add instructions in your own language, even if the prompt is written in another language.

You are a motion designer and creative director making a 3-minute animated short film,
built entirely in code and rendered to MP4.
The film

Title: "Attention Is All You Need → AGI"
The story of how a single 2017 paper led, step by step, to large language models, reasoning,
tool-using AI agents, and to the open question of AGI. It should feel like a cinematic essay,
not a slideshow or a timeline infographic. Think Kurzgesagt meets a Pixar opening sequence:
emotional, clever, precise.
Tech setup
Use Remotion (React). Scaffold a fresh project, 1920x1080, 30 fps, exactly 180 s (5400 frames).
One composition per scene, sequenced in a master composition.
No external image assets or stock footage: everything is drawn with SVG, Canvas, CSS and
code-generated particles and shapes. Google Fonts are fine.
Before building, write STORYBOARD.md with each scene's timing, visuals, on-screen text
and transitions. Then build scene by scene.

After each scene, render 3–4 stills (npx remotion still) and look at them critically.
Fix layout, overlaps, legibility and pacing before moving on.
Finish with npx remotion render to out/film.mp4. Leave an optional <Audio> slot for a
music track I can add later (public/music.mp3); render silently if it's missing.
Narrative spine / recurring motif
The protagonist is a single glowing token, the word "the", that travels through every era.
In 2017 it's a lonely point that suddenly "sees" every other word through attention lines.
Over time it gains a voice, senses (multimodality), reasoning, hands (tools) and finally
a question.
Scenes (approximate timing, you may rebalance)
0:00–0:15 — Cold open. Darkness. Words scattered like stars, disconnected. RNN-style
sequential processing: words light up one by one, slowly forgetting the earlier ones.
0:15–0:35 — June 2017, "Attention Is All You Need" (Vaswani et al., Google). Every word
connects to every other word at once. Attention lines bloom into a web. The title of the
paper appears as if typeset.
0:35–0:55 — 2018–2020: the rise of large language models: GPT-1, BERT, GPT-2, GPT-3.
Scaling laws: the web grows exponentially and the camera pulls back. The model starts
completing sentences, fluent and uncanny.
0:55–1:20 — 2022: instruction tuning and RLHF, then ChatGPT (Nov 30, 2022). The token gets
a chat bubble. A counter of users explodes. The world starts talking back.
1:20–1:40 — 2023: multimodality. Images, audio and code stream into the same web from all
sides, and the token gains "senses." Every modality becomes tokens flowing through the
same attention mechanism.
1:40–2:05 — 2024–2025: reasoning models. A visible chain of thought unfolds as branching,
pruning, backtracking paths. The model "thinks before it speaks," with time slowing down.

2:05–2:30 — Tool use and agents. The token grows hands: it calls a search, runs code,
opens files and orchestrates sub-agents. Many parallel threads work at once, and the
screen becomes a busy, beautiful workshop.
2:30–2:50 — Toward AGI. All motifs converge. The web from 2017 reappears but is now
planet-scale. Leave it ambiguous: no utopia, no doom. On screen: "Attention was all we
needed. What comes next is up to us." (Improve this line if you can do better.)
2:50–3:00 — Resolve back to a single glowing token in darkness. Title card and end.
Craft rules
Accuracy matters: use correct years, names and paper titles. If you're unsure about a
specific fact, leave it out rather than guess, and list any uncertain claims in NOTES.md.

Typography: max ~8 words on screen at once, and leave every text visible long enough to
read (≥ 2.5 s). Use one display font and one mono font for "model output."
Motion: use spring and easing curves, never linear. Use scene transitions that grow out of
the content (the web morphs, the camera zooms through a node), not generic fades.
Color: a deep dark background with one warm accent color that evolves across eras.
Details and easter eggs reward attention (e.g., real paper snippets, tiny UI details,
plausible model outputs), similar to high-craft motion design.
Pacing: alternate dense and calm moments, and give the reasoning scene and the AGI beat
room to breathe.
Work autonomously through the whole pipeline. When done, give me the path to the MP4, a
short summary of creative decisions, and NOTES.md.
Copied the prompt? Here’s what to do next.
  1. Open your project in Claude Code and paste the prompt alongside your own instructions.
  2. Describe what you like about this video: its movement, colors, or pace. Sharing a URL alone doesn’t guarantee Claude Code can watch it.
  3. Add your text and images, then ask for a first preview. Explain what you’d like to change from there.

An example instruction to add — adapt it to your project

Use the prompt above as a starting point for a 15-second video introducing my app. Replace the text with mine and use my colors. Before starting, tell me which files and tools you need, then put together a first preview.
Read the step-by-step guide ↗

More videos to inspire you

From the same creator ↗
VIDEO / 04'What is a Transformer' explainer video↗12:12
Education@dotey
'What is a Transformer' explainer video

帮我用js制作一个视频,主题是:什么是 Transformer 要深入浅出,让高中生也能看得懂,不仅high level说的清楚,也要有细节,包括注意力机制,甚至一些数学概念 你可以用任何工具或者安装工具,可以联网检索 请给我惊喜

Partial prompt
VIDEO / 06Atmospheric circulation geography explainer↗4:48
Education@akokoi1
Atmospheric circulation geography explainer

做一个动画,讲解高中地理知识点“大气环流”。风格轻松有趣,动画格式为线稿,添加合适的音乐,请务必做到引人入胜,字幕用中英双语,解说用TTS,如果 TTS 接口有关闭水印的参数就关掉,文档在TTS.md,API KEY 和音色分别是 .env 里的 APIKEY 和 VOICE,最…

Partial promptReferences not included