AI Video Can't Read "She's Devastated." Here's the Action-Line Fix.

AI Video Can't Read "She's Devastated." Here's the Action-Line Fix.

I spent four generations on a Lost Garden scene where a character was supposed to be grieving. My prompt said exactly that: grieving. The model gave me a woman standing still, blinking, doing nothing in particular. Four tries, same nothing. The prompt wasn’t vague. It was precise about the wrong thing. An AI video model cannot render an emotion word. It can only render what a camera can physically see: a body, an object, a light, a movement. “Devastated,” “ashamed,” “relieved,” “terrified” are not visual instructions, they’re conclusions a viewer is supposed to draw from something visual. If your action line stops at the conclusion, the model has nothing to generate, so it defaults to a neutral face and calls it done. Here’s the part that should have saved me those four generations: screenwriters have had the fix for this since long before anyone typed a prompt into Runway or Kling. It’s called writing action lines that show instead of tell, and it’s one of the oldest rules in the craft. What screenwriters already knew about action lines Action lines in a screenplay describe what a camera and a microphone can pick up: settings, characters, movement, sound. Nothing else exists on the page as far as production is concerned. Every professional guide to the format lands on the same short list of rules: present tense, active voice, strong verbs, short paragraphs of three or four lines. The one that matters here is simple: describe what the audience will see, not what a character is feeling inside. No Film School’s own breakdown of action-line craft puts it plainly: focus on visuals, avoid describing internal thoughts or emotions, and let concrete, specific detail do the work a feeling-word can’t. That rule existed for human directors and actors long before AI video generation did. A director reading “she’s devastated” in a script still has to invent the physical performance, because the word alone gives them nothing to shoot. An AI model is in exactly the same position, except it has no imagination to fall back on. It just renders a blank. What “she’s devastated” actually asks an AI model to do Nothing. That’s the honest answer. Feed a text-to-video model an internal-state word and it either ignores it and defaults to a neutral pose, or it grabs the nearest visual cliche it has seen paired with that word in training data, usually a hand over the mouth, or tears pasted on regardless of context. Neither is a performance. Both are guesses. Compare that to a line like: she irons a shirt that’s already pressed, folds it wrong, stops halfway, and doesn’t pick it back up. Nobody used the word grief. Nobody had to. The behavior carries it, and it carries it into a model the same way it carries it into a reader, because it’s made entirely of things a lens can actually capture. The goal of action lines is to create a visual blueprint for the director and the rest of the filmmaking team, not a summary of what the scene is supposed to mean. That’s true for a film crew. It’s more true for a model with no crew, no table read, and no chance to ask what you meant. The physical-substitution pass (do this before you generate anything) Before a single shot goes into a generation queue, run every action line through one pass with a single question: if I can’t say the emotion word, what does the camera actually see? A short reference list, built from the substitutions that actually worked on Lost Garden: Devastated → shoulders drop, a held object slips from the hand, movement slows or stops mid-task Ashamed → eyes drop and stay down, body turns slightly away from another character, hands find something to hold Terrified → breath visibly held, a step backward before any forward motion, fingers grip a fixed object Relieved → an exhale with visible chest movement, shoulders drop down (the opposite direction of devastated, worth noting, since the model can’t tell “drop” apart without more context) Furious → jaw sets, a held object is gripped harder or set down too hard, stillness that reads as forced control Notice the pattern: every substitution is a verb plus a body part or an object, never a mood. That’s the whole trick, and it’s also, not coincidentally, the exact craft note every screenwriting guide gives for action lines meant for a human reader. Three common mistakes I see (and made myself) when trying this for the first time: Stacking three emotion words instead of removing them. “She looks devastated, broken, and lost” is still zero visual information times three. Replacing the word with a camera direction instead of an action. “Close-up of her devastated face” tells the model where to point a camera it doesn’t have, not what the face is doing. Forgetting the object in the room. A held prop, a piece of furniture, a door: something physical to interact with does more work than another adjective ever will. Does this apply to dialogue too, or just action lines? Mostly action lines, but parentheticals (the little stage directions above a line of dialogue, like (angrily)) have the identical problem and the identical fix. “(angrily) I’m fine.” gives an AI voice or lip-sync tool a mood label with nothing physical attached to it. “(cutting him off) I’m fine.” at least gives a timing and a relationship cue the model can act on, even before you add a voice-performance note. If a tool lets you attach a reference voice clip or a stability setting, that’s where actual vocal emotion should live. The parenthetical’s job is timing and physical behavior, not mood. Do I need special software for this, or can I do it in a normal screenplay? No new tool required. Any screenplay, in any format, benefits from this pass, because the rule predates AI video entirely. What changes is when you do it. For a human crew, a slightly vague action line gets fixed on set by an actor’s instinct. For an AI-generated shot, there is no actor to compensate, so the pass has to happen on the page, before generation, or you pay for it in wasted renders instead. Inside ScreenWeaver, the shot-planning stage flags internal-state language in an action line before it ever reaches a generation call, which is the cheapest place in the whole pipeline to catch it. You can run the same check by hand with nothing but a highlighter and the list above. FAQ Why does my AI-generated character’s face look blank even with a detailed prompt? Usually because the detail describes an internal state rather than a physical one. Blank is the model’s honest default when it has nothing visual to work with. Is this the same issue as AI video not understanding “sad music” or tone? Related but separate. Music and color grading can carry mood on their own; this fix is specifically about what a human or AI-generated character’s body is doing on screen. Will future models fix this on their own? Possibly for the most common emotion words, since training data does pair some feelings with typical gestures. But specificity will always beat a guess, the same way a well-written action line beats a vague one even for a human reader. Does this only matter for dramatic scenes? No. Confusion, boredom, curiosity, and excitement all suffer from the same problem, and all have the same fix: name the physical behavior, not the label for it. If you’re generating shots from a script right now, the fastest thing you can do today is take the last scene you struggled with and cross out every feeling-word in the action lines. Replace each one with a single physical detail a camera could actually catch. It’s the oldest note in screenwriting, and it happens to be the exact fix your AI video generator needed too. Lost Garden’s next episode is being planned with this pass built into every scene from the outset. If you want to see how a locked shot plan carries that kind of detail from script to generation, ScreenWeaver is where I build mine.

Original Source

Read the full article at Hackernoon →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.