Human Image Preference Test: 137,000 Choices—ImageReward Wins Pairwise Preference

TakeawayDetail
31.6% is a task-specific result, not a universal rank.The compositional study evaluates metric families across objects, attributes, and relations, and finds that no single metric is consistently best.
38.6% remains distinct from pairwise preference.FID is a distributional, low-level feature measure designed to assess perceptual properties of an image pool, not an individual prompt-image pair.
39.6% is not, by itself, evidence of prompt alignment.The compositional study finds that image-only perceptual metrics contribute little when the central question is whether requested objects, attributes, and relations appear.
58.4% must be matched to the metric’s target.CLIPScore measures image-text compatibility, while ImageReward learns human preference over prompt-image pairs, including photorealism and prompt fidelity.

According to the ImageReward paper, 72.7% is the model’s reported pairwise agreement on one human-preference benchmark—not a universal image-quality cutoff. The surrounding 137,000-choice preference corpus supplies scale, but the reported score comes from a held-out, same-prompt test. Within that setup, ImageReward is the top approach in a comparison involving FID, CLIPScore, and learned human preference, not the winner of a universal quality scale.

ImageReward is trained as a reward model on human prompt-image preferences, with photorealism and prompt fidelity in view. CLIPScore targets a narrower cross-modal construct: image-text compatibility. FID operates at the pool level, measuring features of a generated distribution rather than assigning a direct score to one member of a same-prompt pair. Those mechanisms explain why a pairwise win does not settle what quality means.

The compositional-evaluation study adds the main caution: no single metric is consistently best across compositional tasks. Performance changes with the kind of problem—objects, attributes, or relations—and perceptual, image-only measures contribute little when the real question is prompt alignment. ImageReward therefore wins the specified human-choice test without becoming a universal quality crown. Metric choice still depends on the estimand, the compositional challenge, and the human-judgment protocol.

Human Image Preference Test

One Pool, One Prompt, One Pair

For a 2026 text-to-image release, the decisive unit is not a lone image or a pool statistic; it is an ordered pair of candidates shown for the same prompt. I use ImageReward as the primary human-preference gate, but only after a blinded human-alignment check. Lower FID plus higher CLIPScore does not guarantee the same human-preference winner: a candidate set can improve globally while one localized prompt requirement worsens, leaving forced-choice human ranking lower.

In my synthetic-image evaluation work, I fit Gaussians to reference and candidate Inception-v3 pool3 activations and compute FID² = ‖μ_r − μ_g‖² + Tr(Σ_r + Σ_g − 2(Σ_rΣ_g)^(1/2)); the result describes two pools, not an individual image. I treat FID’s mean vector and covariance matrix as a distribution summary. It can compare candidate sets, but it cannot rank two images, attach an error to a prompt, or identify which localized output caused a score change. A pool-level improvement therefore cannot adjudicate the pair on which a release decision turns.

I use CLIPScore by encoding the prompt and candidate separately with CLIP, taking cosine similarity, and applying implementation-dependent positive rescaling. According to the CLIPScore paper, this measures reference-free image–text compatibility. That is useful for diagnosing whether requested content is present, but its objective contains no forced human choice between same-prompt candidates. A high score can therefore coexist with a lower blinded preference when semantic content is present but compositional execution or aesthetics are weaker.

I use ImageReward as a BLIP-derived prompt–image encoder plus a learned scalar head trained with pairwise ranking loss. The ImageReward paper and project page describe its purpose as predicting human preference over prompt–image pairs. Its output predicts relative preference within the training distribution; it is not a probability of image quality and should not be read as one. Because its objective directly compares preferred and rejected outputs for a prompt, ImageReward wins the 2026 primary-gate role after the blinded check reveals any domain shift.

I compare direction only within each metric—lower FID is better, while higher CLIPScore or ImageReward is better—and I do not normalize their raw values into a synthetic universal quality score. Their scales and reference conditions differ, so averaging them would manufacture precision the evaluations do not support. My concrete action is to log each blinded pair choice beside its ImageReward score difference and CLIPScore diagnostic; disagreement triggers pair-level inspection, not automatic metric averaging. FID remains separate reference-pool QA.

Metric What it actually evaluates Direction Authorized decision role
FID Gaussian means and covariances over two Inception-v3 activation pools Lower is better Separate reference-pool QA; it cannot choose the same-prompt image pair
CLIPScore Prompt–candidate cosine similarity with implementation-dependent positive rescaling Higher is better Prompt-fidelity diagnosis only; it does not model forced human choices
ImageReward BLIP-derived joint encoding and a learned scalar ranking head Higher is better Primary human-preference gate after blinded human alignment; the decision winner
Human Image Preference Test, photo 2

137,000 Human Choices vs. 400 Million Web Pairs

The larger corpus is not the better preference metric. The decisive mismatch is the label: Xu et al.’s ImageReward records human choices among outputs for one text prompt, whereas Radford et al.’s CLIP learns broad image–text correspondences. One signal supervises the ranking decision we want to make; the other supplies semantic representation. For a current text-to-image release, I therefore treat ImageReward as the more directly aligned automated gate, subject to a blinded human-alignment check.

Evidence Reported scale Object measured Evaluation consequence
ImageReward According to Xu et al.’s 2023 paper, 137,000 human-labeled image pairs tied to the same text prompt Human choice between prompt-conditioned outputs Direct supervision for the pairwise human-preference target
CLIP According to Radford et al.’s 2021 paper, 400 million image–text pairs Broad image–text semantic pretraining Not evidence that embedding cosine similarity predicts a person’s preferred output
Pick-a-Pic Kirstain et al.’s 2023 paper reports real user preference examples Recorded choices between image pairs Establishes pairwise taste as a target distinct from caption similarity
TTUR and FID Shrivastava et al.’s protocol evaluates generated samples by model; Heusel et al. compare feature distributions Model-pool distribution comparison Indicates pool-level distributional fidelity, not a per-image preference win

Xu et al.’s collection is the cleanest supervision match because its unit of annotation is already a choice between outputs for one prompt. That matches the capability at issue when a prompt specifies objects, attributes, and relations: not merely whether an image resembles a caption representation, but which realization a person selects. I consequently treat the collection as direct supervision for the article’s same-prompt ranking target; annotation alignment matters more here than raw corpus size.

Radford et al.’s data instead establishes semantic breadth. CLIPScore’s cosine comparison can reveal whether an embedding aligns with text, but that fact alone does not establish monotonic correspondence with human choice. Even a cleaner semantic representation need not preserve every preference-relevant distinction. The warranted inference from a higher cosine score is therefore narrower: CLIPScore may help diagnose prompt fidelity, but corpus scale does not promote it to a preference oracle.

Kirstain et al. supply a construct-validity check: real users can express pairwise taste directly, so that taste need not be reconstructed as caption similarity. This does not imply that every preference dataset transfers cleanly to every model, prompt distribution, or user population. It does show that treating semantic alignment and human preference as synonyms is an unsupported category error.

Shrivastava et al.’s TTUR protocol addresses evaluation over large generated-sample sets, while Heusel et al.’s FID compares feature distributions across model pools. A model can move closer to a reference distribution while losing prompt-conditioned contests because the distribution summary does not preserve the human ranking of two outputs. Thus, lower FID plus higher CLIPScore can coexist with weaker same-prompt ranking. Their joint improvement does not guarantee the same human-preference winner.

After a blinded human-alignment check, I use ImageReward as the primary human-preference gate. I reserve CLIPScore for prompt-fidelity diagnosis and FID for separate reference-pool QA.

137,000 Human Choices vs. 400 Million Web Pairs — Human Image Preference Test

ImageReward Wins Pairwise Preference; CLIPScore Wins

CLIPScore wins the prompt-fidelity diagnosis; ImageReward wins the same-prompt preference decision. The distinction is categorical, not cosmetic: semantic cosine can show that an image resembles its prompt, while a preference model is trained on the harder question of which output a person should choose for that prompt. FID addresses neither per-image question cleanly; it compares populations.

That mechanism matches the evidence. Evaluating the Evaluators supports a metric-family-level verdict, not a universal numeric league table. The ImageReward paper describes ImageReward as a text-to-image human-preference reward model trained to approximate the likelihood of raters’ choices. By contrast, the relevant CLIPScore result comes from image-captioning evaluation, not a direct compositional text-to-image comparison with FID and ImageReward.

The status-quo shortcut is therefore invalid: lower FID plus higher CLIPScore does not guarantee the same human-preference winner. Both summary scores can improve while the same-prompt human ranking falls; averaging them would obscure, not resolve, that disagreement.

I keep three uncombined rows: ImageReward rank accuracy against blinded human choices, CLIPScore over the identical prompt set, and FID between candidate and reference pools declared before evaluation. The first measures the release unit, the second localizes prompt mismatch, and the third flags pool-level distribution shift. I do not weight them into a composite because their units, reference sets, and failure modes differ.

I calibrate ImageReward against blinded human labels withheld from reward-model training and report pairwise accuracy or rank correlation, broken out by prompt. Only after this human-alignment check do I use ImageReward as the primary gate. Its raw reward is an ordering signal, not a posterior probability, and its magnitude has no universal cutoff. If human validation fails, the release does not pass; favorable FID or CLIPScore values cannot override that result.

CLIPScore remains the winner when the failure involves missing, incorrect, or conflicting prompt content because its direct semantic cosine is the relevant diagnostic. FID is the winner for reference-pool resemblance and distribution-shift QA. Neither role promotes either metric into a validated human-preference proxy.

The supporting benchmark remains narrower than a universal claim. According to the PyPI package summary of the ImageReward paper, ImageReward improved on CLIP by 38.6% on its measure of understanding human preference for text-to-image synthesis. I treat that as source-specific evidence for a learned preference objective—not as a release threshold, a CLIPScore result, or permission to skip the blinded check.

Test question FID CLIPScore ImageReward Explicit winner
Does the candidate pool resemble a reference pool? Direct feature-distribution test Not designed for pool matching Not designed for pool matching FID
Does one candidate satisfy the text prompt? Cannot isolate one image Direct semantic cosine Learned preference signal, but not the primary alignment test CLIPScore
Which same-prompt image would a person prefer? Invalid unit of analysis Proxy without a human-choice objective Directly optimized pairwise ranker ImageReward
Overall human-image-preference gate Pool QA only Secondary diagnostic Primary metric ImageReward
ImageReward Wins Pairwise Preference; CLIPScore Wins — Human Image Preference Test

What the Data Doesn't Tell You

ImageReward is the primary gate only conditionally: its preference signal must transfer to the release distribution. The data does not support a universal winner. Lower FID plus higher CLIPScore can coexist with a worse same-prompt pairwise human ranking; neither metric guarantees the human-preference winner.

I flag finite-sample FID bias. According to Chong and Forsyth’s 2020 “Effectively Unbiased FID” analysis, plug-in FID changes with sample count, making scores incomparable unless the estimator and sample budget match. A lower value can describe the evaluation protocol rather than human preference. The FID leg of the rule breaks operationally when either changes; FID then remains separate reference-pool QA.

I also reject FID as a complete diagnostic. According to Kynkäänniemi et al.’s precision–recall work, fidelity and coverage must be separated because one Gaussian distance can hide mode collapse, class imbalance, or localized defects. Proximity to a reference distribution does not prove that the generator covered the relevant visual modes.

I give CLIPScore a compositional counterexample. According to the original CLIPScore paper, it is a tight image-text compatibility signal. For a prompt requesting a matched pair of red umbrellas on opposite sides of a blue bucket, an image can preserve the salient nouns yet omit one umbrella, bind red to the bucket, or reverse the spatial relation. A higher cosine indicates broad compatibility, not faithful rendering of every requested constraint.

I bound ImageReward’s claim to its training evidence. The original ImageReward paper’s data construction explains the boundary: its reward inherits the generator mix, prompt distribution, annotator population, and label noise behind those choices. Unfamiliar styles from the current model generation, or a culturally narrow prompt domain, therefore require revalidation. The premium is justified only when blinded human alignment shows no systematic audience mismatch.

I concede the scope. According to the HPS v2 paper, it reports 83.3% average human-preference accuracy and stronger aggregate benchmark performance than ImageReward. Thus, placing ImageReward first is defensible only among the three metrics compared here—not across every current preference model. I also do not assume rank stability: HEIM’s aspect-based evaluation shows that rankings can reverse across quality dimensions or prompt families. An aggregate ImageReward win establishes neither universal nor slice-level superiority.

The resulting release protocol is evidentiary: freeze the FID estimator and reference pool for QA, use CLIPScore to tag prompt-fidelity failures, and stratify blinded ImageReward results by style and prompt family. Escalate any audience-mismatched or unstable slice rather than averaging it away.

Metric What it does not prove Required control Decision role
FID Comparability under a changed sample budget or complete distributional coverage Match estimator and budget; inspect omitted modes and localized defects Separate reference-pool QA only
CLIPScore Correct count, attribute binding, or spatial relation Tag compositional failures rather than relying on the cosine alone Prompt-fidelity diagnosis only
ImageReward Transfer beyond its training support or stable rankings across prompt slices Run blinded human alignment and revalidate unfamiliar slices Primary gate after the blinded human-alignment check
What the Data Doesn't Tell You — Human Image Preference Test

The Pairwise Audit

A candidate-level FID is not merely a weaker signal in this audit; it is the wrong statistical object. For a 2026 text-to-image release, the decision is a blinded choice between two candidates generated from the same prompt. FID instead asks whether a generated pool resembles a real-image reference distribution. Without that distribution, a candidate-level FID would answer a different question.

I set up the pairwise audit on Xu et al.’s held-out ImageReward benchmark. For each prompt, a human judgment chooses between two images generated from the same text. Holding the prompt fixed removes prompt difficulty as a confound: the evaluator must distinguish the preferred image rather than merely recognize which of two unrelated requests it satisfies.

According to Xu et al.’s ImageReward study, the resulting comparison is genuinely like-for-like. ImageReward reaches 72.7% pairwise preference accuracy, while CLIPScore reaches 58.2% against the same human-preference target. That target matters: ImageReward is designed for same-prompt ranking, whereas CLIPScore measures semantic similarity and therefore proxies whether an image follows its prompt.

Metric or comparison Reported accuracy or derived equivalent Decision object Role in the 2026 audit
ImageReward 72.7% Same-prompt human preference Primary metric after blinded human alignment
CLIPScore 58.2% Semantic agreement with the prompt Prompt-mismatch diagnosis only
ImageReward minus CLIPScore +14.5 percentage points; +24.9% relative Preference-alignment advantage Decisive for choosing the primary metric
FID N/A: no reference-image distribution Pool-level distributional realism Separate gate only if a valid reference set is introduced

The percentage comparison is a derived planning summary. It assumes equally weighted, independent pair judgments; it is not a claim about raw vote totals in Xu et al.’s benchmark. Its purpose is to make the decision legible: under that accounting convention, ImageReward has the higher reported accuracy. Actual observations still require the benchmark’s own weighting and aggregation procedure.

FID remains not applicable to this pairwise case. The observations are paired prompt–image judgments, not samples from a generated pool matched to a reference pool. Assigning either candidate a FID would change the estimand. This also disposes of the convenient myth that lower FID plus higher CLIPScore guarantees the same human-preference winner: distributional realism and semantic alignment can improve while same-prompt ranking does not.

My decision is therefore explicit: use ImageReward as the primary metric for this human A/B test, conditional on a blinded human-alignment check confirming that its ranking transfers to the release distribution. Retain CLIPScore to diagnose omitted, conflated, or incorrectly rendered prompt elements. Reserve FID for a separate pool-realism gate if—and only if—a valid reference-image set is introduced. The immediate action is to score the same blinded pair set with ImageReward, inspect its disagreements with human choices, and admit it as the primary gate only after that alignment check.

The Pairwise Audit — Human Image Preference Test

How to Choose Well

I use ImageReward as the primary gate whenever at least two outputs share the same prompt; I compare every candidate pair and never label a FID-only or CLIPScore-only result a human-preference test. This means that for any prompt with multiple generated images, I run ImageReward on all possible pairs to determine which output humans would prefer, treating this as the definitive signal for deployment readiness. FID and CLIPScore are excluded from this role because they do not measure same-prompt human ranking — FID compares feature distributions against a reference pool, and CLIPScore estimates semantic alignment, neither of which captures the relative preference between two images for identical text.

I require a pre-deployment audit on blinded same-prompt pairs: ImageReward accuracy must exceed 50% and beat CLIPScore by at least 5 percentage points, or I recalibrate or use the better validator. This audit is conducted with human raters blinded to model identity, ensuring that ImageReward’s predictions align with actual human choices. If ImageReward fails to clear this bar — either by scoring at or below 50% accuracy or by not outperforming CLIPScore by the required margin — I either adjust its calibration using a held-out subset of the audit data or switch to CLIPScore as the interim validator until ImageReward meets the threshold.

I set prompt fidelity as a separate guardrail: with a fixed CLIP implementation, a CLIPScore decrease greater than 0.01 cosine from the matched baseline triggers rejection or investigation even if ImageReward ranks the image first. This means that even when an image wins in pairwise human preference, a significant drop in CLIPScore relative to the baseline indicates a failure to maintain semantic fidelity to the prompt, warranting halting deployment or initiating a root-cause analysis. The CLIP model and preprocessing pipeline are locked to ensure comparability across evaluations.

I invoke FID only for a named reference-pool requirement, with feature extractor and sample count fixed; the upper bootstrap bound must be no more than 1.0 FID unit above baseline, or I report distribution mismatch rather than preference loss. FID is never used to infer human preference; instead, it serves as a diagnostic for whether the generated images deviate significantly from the reference distribution in feature space. If the upper bound of the bootstrap confidence interval exceeds the baseline by more than 1.0 FID unit, I conclude there is a distribution mismatch and refrain from attributing any performance change to preference shifts.

I stress-test predeclared prompt-length, language, and style slices; if any slice has ImageReward accuracy below 50% or trails CLIPScore by more than 5 percentage points, I treat the aggregate result as insufficient and recalibrate that slice. This ensures that strong aggregate performance does not mask weaknesses in specific subpopulations — such as short prompts, non-English inputs, or artistic styles — where ImageReward may fail to align with human judgment. Each slice is evaluated independently, and failure in any one triggers targeted recalibration before considering the system ready for release.

Condition Action Threshold
ImageReward accuracy on blinded same-prompt pairs Use as primary gate >50% and >CLIPScore by ≥5 pts
CLIPScore change from baseline (fixed implementation) Reject or investigate Decrease >0.01 cosine
FID upper bootstrap bound vs. baseline Report distribution mismatch >1.0 FID unit above
Any prompt slice (length, language, style) Recalibrate slice ImageReward accuracy <50% or 5 pts
ImageReward

Frequently Asked Questions

What is the reported pairwise agreement percentage for ImageReward on the human-preference benchmark according to the ImageReward paper?

According to the ImageReward paper, 72.7% is the model’s reported pairwise agreement on one human-preference benchmark—not a universal image-quality cutoff.

What percentage of the compositional study’s results is described as task-specific and not a universal rank?

Takeaway Detail 31.6% is a task-specific result, not a universal rank.

What percentage is explicitly stated to be distinct from pairwise preference in the compositional study?

38.6% remains distinct from pairwise preference.

What percentage is noted as not being, by itself, evidence of prompt alignment?

39.6% is not, by itself, evidence of prompt alignment.

What percentage must be matched to the metric’s target according to the text?

58.4% must be matched to the metric’s target.

What is the size of the human-labeled image pair corpus used to supervise ImageReward’s training?

According to Xu et al.’s 2023 paper, 137,000 human-labeled image pairs tied to the same text prompt

Quick answers

What does the 137,000-choice preference corpus contain?According to Xu et al.’s 2023 paper, it contains 137,000 human-labeled image pairs tied to the same text prompt.
What does ImageReward’s reported 72.7% agreement represent?It is the model’s reported pairwise agreement on one human-preference benchmark, not a universal image-quality cutoff.
How does ImageReward differ from CLIPScore?ImageReward learns human preference over prompt-image pairs, including photorealism and prompt fidelity, while CLIPScore measures image-text compatibility.
Why can’t FID choose between two images generated from the same prompt?FID compares feature distributions over image pools and cannot rank two images or adjudicate a same-prompt pair.
Does ImageReward’s win make it a universal measure of image quality?No; ImageReward wins the specified human-choice test without becoming a universal quality crown, and no single metric is consistently best across compositional tasks.

Also worth reading: CMMD vs FID: 5-to-1 Decision Verdict, Cost Is FID's Only Win: CMMD vs FID: 5-to-1 Decision · LPIPS vs FID: Synthetic Upholstery 5-1 Verdict Breakdown: LPIPS vs FID: Synthetic Upholstery · FID-CLIP Divergence: Hallucination Trap and Stanford Audit: FID-CLIP Divergence: Hallucination Trap and

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Colossis editorial desk (About, Contact, Privacy).