Spots

I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt…

When I started building CineGen, an AI text-to-video generator, I assumed prompt engineering for video would be "image prompting, plus the word moving." It is not. After hundreds of test renders and a lot of embarrassing outputs (a seagull with nine wings remains burned in my memory), I learned that prompting for video is closer to directing a 5-second film than describing a photograph. Here are the lessons that actually changed the quality of my outputs. No hype, just what works. 1. A video prompt is a timeline, not a painting

The single biggest mindset shift: an image prompt

The single biggest mindset shift: an image prompt describes one moment. A video prompt describes a sequence of moments. If your prompt only describes a static scene, the model has to invent the motion — and it will invent something weird.

The second prompt works because it answers three

The second prompt works because it answers three questions the model needs: what's in the frame, what's moving, and how does the shot evolve? A structure I keep coming back to is the three-beat arc: establish → action → resolve. Even in a 4-second clip, giving the model a beginning, middle, and end dramatically reduces the "slideshow of random frames" effect. 2. Camera language does half the work

This was the highest-leverage discovery. Generative video models

This was the highest-leverage discovery. Generative video models respond strongly to cinematography vocabulary — much more than you'd expect. Naming the shot is often more effective than describing the scene in detail. A short glossary that covers ~90% of what I use: Shot scale: extreme close-up, close-up, medium shot, wide shot, aerial shot Movement: slow dolly in, pan left/right, tilt up/down, tracking shot, static shot, orbit around Lens feel: shallow depth of field, 35mm, handheld, smooth gimbal motion

The difference is night and day. "Tracking shot"

The difference is night and day. "Tracking shot" tells the model how the camera behaves, which constrains the motion field and kills a huge class of artifacts where the background slides around unnaturally. If you take one thing from this article: direct the camera, not just the scene.

One caution: don't stack contradictory camera moves. dolly

One caution: don't stack contradictory camera moves. dolly in + pan left + tilt up + orbit in one prompt is asking the model to solve an impossible motion puzzle. One primary camera move per clip.

In image prompting, adjectives carry the load: beautiful

In image prompting, adjectives carry the load: beautiful, stunning, ultra-detailed. In video prompting, most adjectives are noise. What the model needs is motion specification — verbs with direction, speed, and rhythm.

Notice the second prompt barely uses adjectives, yet

Notice the second prompt barely uses adjectives, yet produces a far better clip. My rule of thumb: every noun in the prompt should have a verb attached to it. If something is in the frame, say what it's doing — even if it's just "standing still" (which, by the way, is a legitimate and useful instruction: the cat sits perfectly still, only its tail flicks).

Speed words matter too: slowly, gently, rapidly, suddenly

Speed words matter too: slowly, gently, rapidly, suddenly. Models genuinely differentiate these. "Walks slowly toward the camera" and "runs toward the camera" produce very different motion — use that dial deliberately. 4. Temporal consistency is the real boss fight

The hardest problem in AI video isn't making

The hardest problem in AI video isn't making pretty frames — it's making frame 1 and frame 48 agree with each other. Faces morph, jackets change color, a coffee cup teleports between hands. Here's what actually helps:

News

I Built an AI Text-to-Video Generator — Here's What I Learned About Prompt Engineering for Video

When I started building CineGen, an AI text-to-video generator, I assumed prompt engineering for video would be "image prompting, plus the word moving." It is not.

@spots #dev
Source: Dev.to
See more like this