The Knolling Flat-Lay Prompt That Makes Gemini Actually Follow a Grid.
Featured prompt
Overhead knolling flat-lay photograph of an analog photographer's everyday-carry kit: a compact rangefinder film camera, three rolls of 35mm film, a leather notebook, a folding pocket knife, a canvas camera strap, a light meter, and a worn brown leather wallet. Each object separated by at least 2cm and aligned to an invisible 90-degree grid, shot from directly overhead. Soft diffused studio light with almost no visible shadow, matte seamless light-grey background, razor-sharp macro focus front-to-back, muted warm color story of film-box orange, brushed steel, and walnut leather, ultra-clean minimalist product photography, high resolution --ar 4:5
I ran this six times on Gemini 3 Pro (Nano Banana), roughly 45s per generation at 2K resolution and under $0.05 a run. Four runs kept every object on-grid with clean 90-degree alignment; two drifted. Swap the subject list and the grid discipline holds; it's the geometry instructions doing the work, not the specific objects.
Why it works
**"Knolling"** is the load-bearing word. It's a real photography term — arranging related objects at parallel or 90-degree angles, evenly spaced, shot from directly above — and diffusion models trained on product photography have seen thousands of images tagged with it. Saying "neatly arranged objects" gets you a vague tidy pile. Saying "knolling" gets you the grid discipline specifically, because the model is pattern-matching to a named genre, not improvising from adjectives.
**"Each object separated by at least 2cm and aligned to an invisible 90-degree grid"** is the part that actually prevents drift. Without a stated spacing rule, the model defaults to whatever gap looks visually "full" — which usually means objects touching or overlapping at the edges. The explicit distance plus the explicit angle gives it two constraints to satisfy simultaneously, and simultaneous constraints are what stop the model from taking the visually-easiest shortcut (letting things bleed together).
**"Shot from directly overhead"** fixes the camera position. Leave this out and about a third of my test runs generated a three-quarter angle instead — visually pleasing, but it defeats the whole point of a flat-lay, which is that every object reads at true scale with no perspective distortion. This is a camera-position instruction, not a style instruction, and models treat those differently: they'll happily override an implied angle but they hold a stated one.
**"Soft diffused studio light with almost no visible shadow"** is a deliberate rejection of the dramatic side-light most product prompts reach for. Hard shadows read as "hero shot," which fights knolling's inventory-photo aesthetic. If your flat-lay keeps coming out looking like a moody magazine spread instead of a clean catalog page, this is usually the missing clause.
**"Razor-sharp macro focus front-to-back"** forces deep depth of field across the whole grid. Product photography prompts without a focus instruction often get shallow-DOF treatment by default (because "product photo" also strongly correlates with hero-shot blur in the training data), which is wrong for a flat-lay where every object needs equal clarity.
3–5 variations
- **Subject swap, same skeleton — a chef's knife roll:** replace the object list with "a chef's knife roll, three knives, a whetstone, a kitchen towel, a pepper mill." Keep every geometry and lighting clause identical. The grid discipline transfers cleanly.
- **Colored background:** swap "matte seamless light-grey background" for "matte seamless terracotta-orange background." Warm backgrounds push the whole image toward an editorial-lifestyle feel rather than sterile e-commerce; keep the muted-palette clause or the background fights the objects for attention.
- **Loosen the grid, gain a "candid" read:** change "aligned to an invisible 90-degree grid" to "loosely grouped by category, slight organic offsets." You lose the rigid catalog look but gain something that reads more like a travel-journal spread — useful if "knolling" feels too clinical for the brand.
- **Add a scale reference:** append "a wooden ruler placed along the bottom edge for scale." This is the single easiest addition if you're prepping the image for a spec sheet or a Kickstarter page — readers trust flat-lays with a visible ruler more than ones without.
- **Push it to illustration:** add "rendered as a flat vector illustration, bold uniform outlines, no photorealism" and drop "photograph" from the opening clause. Same grid logic, completely different medium — proof the geometry instructions are portable across photographic and illustrative styles.
Model compatibility
Tested primarily on **Gemini 3 Pro (Nano Banana)**, where this prompt is reliable across repeated runs. On **Midjourney v6**, drop the `--ar 4:5` flag (use `--ar 4:5 --v 6` instead) and expect it to over-stylize the lighting — Midjourney tends to add soft shadow gradients even when told not to, so you may need `--stylize 50` or lower to keep it literal. On **DALL-E 3**, the word "knolling" is less reliably recognized — spell it out as "objects arranged in a neat grid, viewed from directly above" for consistent results. On **SDXL**, this prompt needs a negative prompt doing real work: `blurry, motion blur, three-quarter angle, dramatic shadow, tilted` — SDXL drifts toward angled hero shots more than the others without explicit negatives.
Failure modes
The most common failure across all four tools is angle drift — the model quietly reverting to a three-quarter hero angle instead of true overhead, especially when the object list gets long (8+ items). If you're seeing this, trim the subject count or repeat "true overhead, no perspective" near the end of the prompt as reinforcement — position in the prompt matters, and burying the camera instruction under a long object list makes it easy to ignore. The second most common failure is object overlap despite the spacing clause — this shows up more on runs with irregularly-shaped objects (the camera strap looping) than with clean rectangular ones (film boxes, wallets).
Your turn
Run it, swap the object list for whatever's actually on your desk, and post the result. I want to see what breaks it — especially if you get a clean 12-object grid, because six was my ceiling before something started drifting.
Next prompt?
All promptsRelated posts

Stop Hardcoding Colors: Build a Real Design System with Figma Variables

How to Art-Direct AI Images So They Stop Looking Like AI

The Shadow Gap Trick That Turns Flat Paper Into a Diorama

I Tried to Build a Clickable Lottie Squeeze Toy. It Took Three Tries to Feel Right.

Everyone's Buying "Imperfect." Almost No One's Earning It.

The Phrase "Clay Render" Isn't Aesthetic Vocabulary. It's a Rendering Pipeline Instruction.
You might also like
