Artificial Analysis · XOriginal · English

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the…

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the others drift, GPT Image 2.5 Sunburst most visibly. We've seen some interesting demos…

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the others drift, GPT Image 2.5 Sunburst most visibly.

We've seen some interesting demos of multi-turn editing consistency from the latest image editing models, so we ran our own test: GPT Image 2.5 Sunburst, #1 on our Image Editing leaderboard, against Ideogram 4.5, FLUX 3 and Nano Banana 2.1. Our leaderboard scores single edits; this tests what happens when 30 consecutive changes stack up, with each model editing its own previous output through a 30-step real estate staging sequence: light the fire, add a sofa, repaint the walls, swap day for twilight, and more.

Why are the final results so different?

@ideogram_ai's Ideogram 4.5 and @bfl_ml's FLUX 3 edit locally. On small edits like adding a vase of tulips, we measured that they left 95% or more of the image essentially untouched. GPT Image 2.5 (Sunburst) re-renders most of the entire image on every edit, leaving only about a fifth of the image unchanged, so small shifts in colour and detail compound over turns. Nano Banana 2.1 sits in between: its edits stay local, but the rest of the image shifts slightly and gradually darkens.

Original source

Artificial Analysis · X

Content notes

Original publication and rights belong to the source.