I Tried to Keep an AI Character Consistent Across 5 Scenes. It Took Three Attempts..
Everyone promises AI characters can stay consistent. Technically true. Here's what nobody shows you: attempts 1 and 2 before the approach that actually works.
By attempt 3, I had a character who looked like herself across a hero shot, a whiteboard scene, and a coffee shop moment. Not identical — AI doesn't do identical. But close enough that you'd say "yeah, same person." Getting there involved two failed approaches I'm documenting here so you can skip them.

The brief
A client needed three editorial illustrations: the same fictional founder character in different scenes. Dark curly hair. Warm brown skin. Navy blazer. Professional but not stiff. She needed to be recognizably herself across all three images, not three different women in the same outfit.
No photography budget. My job: make her consistent using AI.
Attempt 1: Describe her really, really well
My first prompt was basically a character sheet in sentence form. I listed hair type, skin tone, eye shape, jaw line, cheekbone structure, the cut of the blazer, the fit of the collar. Seventeen adjectives. I was thorough.
I generated the sitting-at-laptop scene first. She looked great. I saved the prompt. Then I ran the whiteboard scene with the exact same character description.
Different person.
Not wildly different — close enough that I told myself it was fine. Then I ran the cafe scene. Now we had three women who might be cousins. The hair was different lengths across all three. One was clearly a decade older than the brief. One had a green jacket.
What I got wrong: text prompts don't describe an individual. They describe a category. "Dark curly hair, warm brown skin" generates from a probability distribution. Each generation samples fresh. Your 17 adjectives set the genre, not the person.
Attempt 2: Use a reference image
I picked the best image from attempt 1 — the laptop shot — and used it as a visual reference for the next generations. Midjourney has --cref. Stable Diffusion has IP-Adapter. I dropped my reference image in and re-ran.
The whiteboard scene: 75% match. The face was close. The hair got shorter somehow. The blazer went black. I could live with it.
The cafe scene: 50% match. She's laughing in this one. Open-mouth expressions, it turns out, scramble facial features enough that the model loses its footing. I got a cheerful woman who had clearly never met the character from the laptop shot.

What I got wrong: a single reference image anchors one pose as much as one face. The moment the body language changed significantly, the face drifted. IP-Adapter and --cref work for scene-to-scene consistency when the character is doing roughly the same thing. Not when she's going from still to laughing.
Attempt 3: Train a character embedding
This is the approach that actually works in 2026. Costs 30 minutes upfront. Pays back immediately.
Step 1. Generate 17 images of the character from attempt 1's best base prompt — all neutral expressions, different angles (front, ¾ left, ¾ right, slight downward, slight upward). No reference image yet.
Step 2. Select the 12 best: face fully visible, lighting consistent, no weird anatomy. Delete the rest.
Step 3. Upload to a character embedding trainer. Astria, Lovart, or Stable Diffusion LoRA training. Astria's free tier takes 25–40 minutes. The output is a trigger word — a small file that encodes this specific person.
Step 4. Use your trigger word in every subsequent generation: [trigger], standing at whiteboard, office environment, warm light from left.
The prompt is now just stage directions. The embedding handles the face.
Results: 90% consistency across all three scenes. Same hairline. Same nose. The blazer stayed navy because I added it to the trigger context. She laughs in the cafe like she's the same person who sat at the laptop.

What I learned
The mistake in attempts 1 and 2 was treating her like a writing problem. If I described her precisely enough, the AI would get it. That's not how this works. Prompts generate categories. Training generates people. The moment I stopped trying to describe her and started showing the model what she looked like, the face locked in — and stayed locked.
One thing: train on neutral poses first. If you include a laughing or action photo in your training set before the model knows the face, it learns "character = that expression" and fights you on everything else. Neutral first. Expressive later.
If your trained character drifts toward a celebrity lookalike, your training images probably have too much background. Crop tighter to the face on 5–6 images before re-training. The model should learn the face, not the room.
The part I didn't expect
The embedding doesn't just solve consistency. It speeds up every future generation. No more re-describing her in every prompt. Give the trigger word, move straight to scene direction. She exists now. You're just deciding where to put her.
If you get stuck
If your trigger word produces the right face but the wrong energy (too stiff, too casual), go back and add 3–4 training images showing the character in the vibe you need. Energy is a training problem, not a prompt problem.
Next tutorial: Got your character embedding? Now we put her into a Figma mockup without her looking pasted-in. Next Monday.
Done reading? There’s more where this came from.
Related posts

The "Anti-AI Aesthetic" Is Already Slop

6 AI Image Generators for Solopreneurs: Blunt Verdicts (2026)

AI Meeting Notes, Actually Compared: Granola vs Fireflies vs Otter (2026)

Four Design Signals From This Week (One Is Getting More Coverage Than It Deserves)

Six Designers on the One Question Nobody's Asking Right

