Prompting

Why AI images get the count wrong, and how to ask for crowds that work

Ask for exactly three of something and you may get five. Why image models miscount, and five ways to get the crowd, the row or the set you actually wanted.

4 min read
An antique brass abacus standing alone in the dark, its beads pushed to uneven positions, a few of them glowing violet

You typed "three candles on a windowsill" and got four. You tried "exactly three", got two, and then five. Nothing is broken. Counting is one of the things image models are structurally bad at, and once you see why, you can stop fighting it and start steering around it.

Why the count drifts

An image model does not build a picture object by object and then check the total. It refines the whole canvas at once, pulling every region toward "what a picture matching these words usually looks like". The word "three" is a weak nudge among dozens of other words, and in the training pictures the number of candles near the word "candles" was rarely labelled at all.

Small counts, one and two, are the easiest because they are common patterns: one subject, a pair. Beyond that the model leans on texture and rhythm instead of arithmetic. A row of windows looks right when the spacing looks right, not when there are seven of them. That is why a count of "a few" works better than a number it cannot honour.

Where it matters and where it does not

Sort your scene into two kinds. In a crowd, a forest or a bowl of fruit, nobody counts, and an off-by-two is invisible. In a product shot, a diagram, a set of icons or a board game, the viewer counts instantly and a wrong number reads as a mistake. Spend your effort on the second kind only.

Five ways to get the count you meant

1. Ask for fewer, and name the group

Models handle "a pair" and "a trio" better than a bare numeral, because those are single concepts rather than arithmetic. "A trio of candles, tall, medium and short" gives each one an identity, and identities are easier to hold than a total.

2. Describe each item by position

Instead of a count, give every item a place. The model is good at left and right, foreground and background, and it is hard to add a fourth candle when the sentence has already filled the left, centre and right of the sill.

Prompt
A windowsill at dusk, a tall white candle on the far left, a short amber candle in the centre, a medium candle in smoked glass on the right, the window behind them soft and blurred, warm light, 4:5 portrait framing
Three items, three positions, three descriptions. The number never appears, and the count holds far more often.

3. Generate small, then compose

For a set that must be exact, such as five icons or six cards, render each one as its own picture and arrange them in your own editor. It costs a few more renders, but each render is a one-subject job, which is the cheapest kind to get right. Draft the pieces on Still Lite, from €0.30 a render, and only finish the keepers on Still.

4. Keep the crowd vague on purpose

For background people or objects, describe the density, not the number: "a sparse crowd", "scattered", "packed shoulder to shoulder". Models are good at density because it is a texture. Adding a number only gives it something to get visibly wrong.

5. Fix it afterwards instead of re-rolling

If the picture is right except for one extra object, do not roll the dice again. Attach the render as a reference and ask for the same scene "with the fourth candle removed". The model keeps the palette and layout and only has to change one thing, which is a far better bet than a fresh draw.

What the studio cannot do

No tool in the studio can promise an exact count, and neither can any other image generator today. Higher quality sharpens detail, it does not teach arithmetic. If a number on the page is a hard requirement, build the set from separate pictures, or draw the final layout yourself and use the render for the look. The same goes for legible text, which fails for related reasons; the hands and text guide covers it.

A quick routine

  1. Decide whether anyone will count. If not, stop worrying.
  2. If they will, describe items by position and give each one its own adjective.
  3. Render a small draft batch cheaply and reject any with the wrong count at a glance.
  4. Fix a near miss from a reference rather than starting over.
  5. For exact sets, render pieces separately and assemble them yourself.

Pair this with the habits in the re-roll tax post and most of the wasted renders disappear, because miscounts are among the most expensive re-rolls there are. You can see the current price of every model on the models page before you press Generate.

Questions

Why does AI add extra fingers or objects?

The model refines the whole picture at once and matches local patterns rather than totals. Fingers, windows and candles repeat, so it keeps repeating them until the area looks full, and the exact number is not something it tracks.

Does a better model fix counting?

It helps with small counts, one to three, and with how clean the result looks. No current model counts reliably past that, so the position and assemble-yourself methods still apply.

Is it cheaper to re-roll or to edit from a reference?

When the picture is right except for one object, editing from a reference is usually cheaper, because you keep everything good and change one thing. Re-rolling throws the good parts away along with the bad.

#Prompting#Images

Keep reading

All articles →