[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$fivv46fjoejql":3},{"_id":4,"slug":5,"title":6,"subtitle":7,"kind":8,"cards":9,"tags":59,"categories":61,"source":63,"lang":66,"author":67,"audioState":70,"stats":71,"publishedAt":74,"renderer":75},"6ab9faadca21c797c7e9733d","i-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651","I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt Engineering for Video","When I started building CineGen, an AI text-to-video generator, I assumed prompt engineering for video would be \"image prompting, plus the word moving.\" It is not.","news",[10,14,19,24,29,34,39,44,49,54],{"headline":11,"body":12,"imageUrl":13,"sourceImageUrl":13},"I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt…","When I started building CineGen, an AI text-to-video generator, I assumed prompt engineering for video would be \"image prompting, plus the word moving.\" It is not. After hundreds of test renders and a lot of embarrassing outputs (a seagull with nine wings remains burned in my memory), I learned that prompting for video is closer to directing a 5-second film than describing a photograph. Here are the lessons that actually changed the quality of my outputs. No hype, just what works. 1. A video prompt is a timeline, not a painting","https:\u002F\u002Fmedia2.dev.to\u002Fdynamic\u002Fimage\u002Fwidth=1200,height=627,fit=cover,gravity=auto,format=auto\u002Fhttps%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4a7t2sb9qtoazsv8apn.png",{"headline":15,"body":16,"imageUrl":17,"images":18},"The single biggest mindset shift: an image prompt","The single biggest mindset shift: an image prompt describes one moment. A video prompt describes a sequence of moments. If your prompt only describes a static scene, the model has to invent the motion — and it will invent something weird.","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F1.webp",{"local":17},{"headline":20,"body":21,"imageUrl":22,"images":23},"The second prompt works because it answers three","The second prompt works because it answers three questions the model needs: what's in the frame, what's moving, and how does the shot evolve? A structure I keep coming back to is the three-beat arc: establish → action → resolve. Even in a 4-second clip, giving the model a beginning, middle, and end dramatically reduces the \"slideshow of random frames\" effect. 2. Camera language does half the work","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F2.webp",{"local":22},{"headline":25,"body":26,"imageUrl":27,"images":28},"This was the highest-leverage discovery. Generative video models","This was the highest-leverage discovery. Generative video models respond strongly to cinematography vocabulary — much more than you'd expect. Naming the shot is often more effective than describing the scene in detail. A short glossary that covers ~90% of what I use: Shot scale: extreme close-up, close-up, medium shot, wide shot, aerial shot Movement: slow dolly in, pan left\u002Fright, tilt up\u002Fdown, tracking shot, static shot, orbit around Lens feel: shallow depth of field, 35mm, handheld, smooth gimbal motion","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F3.webp",{"local":27},{"headline":30,"body":31,"imageUrl":32,"images":33},"The difference is night and day. \"Tracking shot\"","The difference is night and day. \"Tracking shot\" tells the model how the camera behaves, which constrains the motion field and kills a huge class of artifacts where the background slides around unnaturally. If you take one thing from this article: direct the camera, not just the scene.","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F4.webp",{"local":32},{"headline":35,"body":36,"imageUrl":37,"images":38},"One caution: don't stack contradictory camera moves. dolly","One caution: don't stack contradictory camera moves. dolly in + pan left + tilt up + orbit in one prompt is asking the model to solve an impossible motion puzzle. One primary camera move per clip.","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F5.webp",{"local":37},{"headline":40,"body":41,"imageUrl":42,"images":43},"In image prompting, adjectives carry the load: beautiful","In image prompting, adjectives carry the load: beautiful, stunning, ultra-detailed. In video prompting, most adjectives are noise. What the model needs is motion specification — verbs with direction, speed, and rhythm.","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F6.webp",{"local":42},{"headline":45,"body":46,"imageUrl":47,"images":48},"Notice the second prompt barely uses adjectives, yet","Notice the second prompt barely uses adjectives, yet produces a far better clip. My rule of thumb: every noun in the prompt should have a verb attached to it. If something is in the frame, say what it's doing — even if it's just \"standing still\" (which, by the way, is a legitimate and useful instruction: the cat sits perfectly still, only its tail flicks).","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F7.webp",{"local":47},{"headline":50,"body":51,"imageUrl":52,"images":53},"Speed words matter too: slowly, gently, rapidly, suddenly","Speed words matter too: slowly, gently, rapidly, suddenly. Models genuinely differentiate these. \"Walks slowly toward the camera\" and \"runs toward the camera\" produce very different motion — use that dial deliberately. 4. Temporal consistency is the real boss fight","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F8.webp",{"local":52},{"headline":55,"body":56,"imageUrl":57,"images":58},"The hardest problem in AI video isn't making","The hardest problem in AI video isn't making pretty frames — it's making frame 1 and frame 48 agree with each other. Faces morph, jackets change color, a coffee cup teleports between hands. Here's what actually helps:","\u002Fapi\u002Fmedia\u002Fposts\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-a-e5807651\u002F9.webp",{"local":57},[60],"dev",[62],"Technology",{"name":64,"url":65},"Dev.to","https:\u002F\u002Fdev.to\u002Fmicheal_zh_e114bbd789e4c3\u002Fi-built-an-ai-text-to-video-generator-heres-what-i-learned-about-prompt-engineering-for-video-4gll","en",{"handle":68,"displayName":69},"spots","Spots","queued",{"views":72,"likes":73,"saves":73,"shares":73,"completions":73,"opens":73,"skips":73,"depthSum":73},5,0,"2026-09-28T05:27:09.853Z","local"]