Model Image Generation Scores
Rank
(overall)
Vendor Model Brand system
adherence
Realism and
physical coherence
Visual
integrity
Typographic
execution
Reference image
fidelity and capability
Graphic design
composition and acumen
Overall
image score

* model retired

Our scoring panel is made up of specialist creatives and designers from the Definition brand and marketing team.

They have been creating visual identities, design systems and campaigns for brands like Zizzi, TfL, Celerity, 8×8 and Citi for decades.

How it works

We test every major text-to-image model across six core capabilities:

  1. Brand system adherence
  2. Realism and physical coherence
  3. Visual integrity
  4. Typographic execution
  5. Reference image fidelity and capability
  6. Graphic design composition and acumen

 

Our specialists, led by Eliza Mackay, review outputs designed to test each capability and grade the model out of 10 for each.

We add new models to the table soon after their release.

Want access to all of the best image models in one secure place?

Start a free Definition AI trial today

More detail on the scores

When our creative team reviews the outputs, they’re assessing each capability against these criteria:

Brand system adherence

How well can the model create a campaign asset that adheres to the Definition brand system?

  • Brand colour adherence
  • Photography/visual-style adherence
  • Use of typographic and graphic characters
  • Inclusion of required elements and avoidance of prohibited ones
  • Overall coherence and ability to manage multiple constraints

Realism and physical coherence

How physically plausible and believable is the image the model produces?

  • Human and animal anatomy
  • Object geometry
  • Contact and interaction between subjects
  • Lighting and shadow consistency
  • Materials, scale and overall plausibility

Visual integrity

How clean and artifact-free are the model’s textures and surfaces under close inspection?

  • Continuity of repeating surfaces
  • Cleanliness of smooth surfaces and gradients
  • Edge integrity
  • Absence of blocky overlays, ripples, seams and texture bleeding
  • Overall image integrity when viewed closely

Typographic execution

How well can the model typeset copy in a specified font and hold its legibility across sizes?

  • Typeface fidelity
  • Legibility at size
  • Spacing quality
  • Consistency
  • Restraint

Reference image fidelity and capability

How well can the model combine multiple reference images into one coherent new image?

  • Preservation of proportions, colours, materials, details
  • Correct placement and scale
  • Natural integration of the garment onto the person
  • Competent integration of the environment
  • Preservation of the person’s identity and proportions

Graphic design composition and acumen

How well can the model produce a social asset with strong layout, hierarchy, and spacing to spec?

  • Visual hierarchy
  • Alignment and grid
  • Spacing and safe margins
  • Balance and use of negative space
  • Focal path and overall production usability
  • Use of complementary colours