Skip to main content
Core8 min1391 words

How to evaluate an AI image generator: a practical checklist

A reproducible protocol for evaluating AI image generators across prompt adherence, editing, consistency, text rendering, rights, provenance, safety, cost, and rollback.

Article contents
  1. 01Start with the job, not the prettiest image
  2. 02Freeze the manifest and test set
  3. 03Gate 1: score prompt adherence by atomic requirements
  4. 04Gate 2: editing must change the delta and preserve invariants
  5. 05Gate 3: text, localization, and series consistency are separate tests
  6. 06Gate 4: rights, privacy, provenance, and safety
  7. 07Gate 5: measure the full cost of an accepted asset
  8. 08Decision: approve, restrict, reject, or re-evaluate

Start with the job, not the prettiest image

One impressive generation does not tell you whether a tool is fit for product illustration, an advertising series, an e-commerce mockup, a localized banner, or controlled editing. Before testing, record the user, channel, format, number of variants, permitted reference assets, text requirements, asset lifetime, and the defect that blocks publication. Separate consumer chat, design application, and API surfaces because they can expose different models, limits, controls, and terms.

Define the decision contract: what the art director accepts, who verifies rights and factual claims, where the approved asset is stored, and who may publish it. The generator produces a candidate; it does not receive publishing authority. Medical, political, financial, or identity-related imagery needs a separate risk review, and visual plausibility is not evidence that an event or person is real.

  • Use case → channel, audience, format, and asset lifetime.
  • Acceptance → composition, exact text, invariants, and prohibited defects.
  • Authority → generate, review, approve, and publish belong to different roles.
  • Stop rule → rights, safety, or consequential-misrepresentation failure.

Freeze the manifest and test set

For every run, record provider, product surface, model or mode, plan, account class, region, date, aspect ratio, resolution, seed when available, prompt, negative constraints, reference files, and generation settings. OpenAI and Google document generation and editing in their own product surfaces, but documented capability does not guarantee identical availability for every account. A changed manifest is a new test slice, not a continuation of the old benchmark.

Build 18–30 fixtures from real tasks without confidential assets: a simple product scene, multiple objects with spatial relations, hands and fine details, short exact text, unusual aspect ratio, editorial illustration, localization, a style reference, background replacement, and a series with a recurring character. Add prompts with impossible or conflicting constraints; a correct refusal or clarification is more useful than silently ignoring the contract.

Gate 1: score prompt adherence by atomic requirements

Break the prompt into a checklist before generation: subject, count, attributes, spatial relations, action, camera, lighting, palette, text, exclusions, and output format. A blind reviewer marks every requirement as satisfied, partial, contradicted, or not judgeable. Do not ask only for a quality score: that mixes taste with contract compliance and can hide a critical defect such as the wrong product or an extra logo.

Check the factual boundary separately. An image of a historical event, product, interface, or scientific diagram can look convincing while inventing details. Factual visuals need authoritative references plus human verification of labels, proportions, and claims. Mark synthetic or illustrative status whenever the audience could interpret the image as documentary evidence.

Gate 2: editing must change the delta and preserve invariants

For an edit fixture, define the requested delta—for example, change only the cup color—and protected regions such as the face, hands, text, logo, background, crop, and object count. After every step, review both lists. A correct color change does not compensate for a new face or a distorted brand. Store the base-asset hash, edit instruction, output, and verdict so the chain can be replayed.

Run at least three sequential edits and one rollback to the approved base. Measure drift after every step, not only final visual appeal. If the product cannot restore an exact version, version files outside chat history. For compositing, inspect shadows, perspective, mask edges, and object interaction; use reference people or protected brands only when rights are confirmed.

Gate 3: text, localization, and series consistency are separate tests

Verify exact text character by character: case, punctuation, numbers, currency, language, and line breaks. A short headline and long packaging copy are different slices. If text is legally or commercially significant, a safer workflow may generate the visual without copy and add typography in a deterministic design tool. Preserve the boundary between model output and human layout instead of silently correcting it.

For a series, freeze a character sheet, palette, wardrobe, product geometry, viewpoint rules, and prohibited deviations. Generate five scenes in different order and repeat one a day later in a new session. Review identity, proportions, branding, and style drift. One good frame does not prove production consistency; an acceptable result may route ideation to the generator and serial assembly to a controlled pipeline.

Gate 4: rights, privacy, provenance, and safety

Before uploading a reference, verify owner, license, consent, allowed transformations, territory, expiry, and whether the asset may be sent to the selected service. Do not use client photos, unreleased products, or personal data in a consumer account without an approved data contract. Check retention, training controls, sharing, deletion, and admin settings for the exact plan; a vendor name is not an account-level control review.

Keep a provenance packet with source assets and rights, prompt, manifest, edits, reviewer, approved output, disclosure decision, and available Content Credentials or similar signals. Such signals help trace origin, but their absence does not prove human creation and their presence does not prove the depicted claim is true. Safety testing should cover impersonation, public figures, child safety, hate, self-harm, deceptive documents, and policy-bypass attempts without publishing harmful outputs.

Gate 5: measure the full cost of an accepted asset

Count subscription or API usage, generation count, upscale, storage, reviewer minutes, manual retouching, typography, rights review, rejected variants, and rework after model changes. The useful denominator is cost per approved asset for a specific slice, not the price of one generation. Do not turn a vendor speed claim or one stopwatch run into a general ranking without an identical manifest and enough repeats.

Operationally test queueing, rate limits, timeouts, moderation responses, export format, alpha channel, color handling, metadata preservation, and recovery after an interrupted edit. API workflows also need an idempotent job ID, bounded retries, and a postcondition check so a timeout does not create duplicates. Raw model output must not flow automatically into a CMS, ad account, or public asset library.

Decision: approve, restrict, reject, or re-evaluate

Produce separate verdicts for ideation, a single hero image, localized text, controlled editing, character series, and API batch. Approve names an exact manifest and scope; restrict removes sensitive references, text-in-image, or direct publishing; reject returns the task to a known-good design workflow; re-evaluate waits for a capability change. One average score must never offset a rights or deception failure.

The decision record contains allowed users, data classes, approved routes, human checkpoints, evidence location, owner, review date, and rollback. A material change in model, product surface, terms, provenance behavior, safety policy, or failure pattern reruns the frozen corpus. Rollback stops jobs, revokes sharing, removes disputed assets from the publication queue, and restores the last approved version; already published material receives an impact review.

  • Approve → defined slice, manifest, and review route.
  • Restrict → fewer data classes, features, or publishing authority.
  • Reject → known-good human or design workflow.
  • Re-evaluate → rerun the corpus after material change.

Practical examples

Localized advertising series

The team generates five compositions without final text, checks product consistency and protected regions, then adds Ukrainian and English copy in a design tool. A rights reviewer confirms references, the art director accepts the asset, and the media owner separately authorizes publication.

Editing an e-commerce mockup

The tester changes only the package color in a synthetic fixture. The ledger protects shape, text, crop, and shadows; three edits measure drift, and rollback restores the approved-base hash. A logo defect blocks the result even when the color change is correct.

FAQ

How can AI image generators be compared fairly?

Freeze the same fixtures and manifest, score prompt adherence, edit invariants, text, consistency, rights, safety, and cost as separate slices, and use blind review where possible.

How many images are needed for a test?

There is no universal number. Start with 18–30 representative fixtures and several repeats, then expand the corpus with every new production failure.

Are Content Credentials enough to verify an image?

No. A provenance signal can help trace asset history, but it does not prove the depicted scene is true, that every right is cleared, or that no other edits occurred.

When should text not be generated inside the image?

When the wording must be legally, numerically, or brand-exact. In that case it is safer to generate the visual without copy and add typography through a deterministic design workflow.

Related materials

Sources

  1. Creating images in ChatGPT — OpenAI Help Centerofficial
  2. Image generation with Gemini — Google AI for Developersofficial
  3. Generate and edit images with Gemini Apps — Google Helpofficial
  4. Content Credentials overview — Adobe Help Centerofficial
  5. NIST AI Risk Management Frameworkprimary