AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Text-to-Image Models Can Look Right and Still Get the World Wrong

Text-to-image generators now produce photorealistic pictures from short written prompts, yet a polished image can still show things that could not exist: malformed limbs, hands passing through objects they are holding, or animals that appear physically fused with their surroundings. Common scoring tools, which measure aesthetics, prompt alignment, or human preference, were not built to catch these errors. Researchers at Adelaide University have proposed a new evaluation dimension, which they call world-grounded visual consistency, along with a measurement framework called TerraVis.

The authors define world consistency as whether depicted objects, structures, interactions, and spatial relationships conform to real-world constraints. Their taxonomy lists 18 violation types: seven at the object level (such as structural distortion or biological anatomy), eight at the interaction level (such as support and stability), and three at the scene level (such as relative scale). TerraVis uses a multimodal large language model (MLLM), a system that reads both images and text, to answer a sequence of targeted questions. It first checks whether an image depicts a macroscopic object or scene, excluding abstract, symbolic, or layout-based content. It then checks each violation type, labels each detected violation as major if it affects a central element or minor if it affects a peripheral one, and computes a score using an exponential penalty. In the experiments, the penalty weight was set to 1 and the minor-violation weight to 0.5, so each major violation reduces the score sharply, and an image with no violations scores 1.

The paper’s clearest example involves a prompt for “two cats sitting on top of a pair of shoes.” The SD 3.5-Large image received the highest alignment score among the five generators shown (0.90 on VQAScore), yet its human world-consistency score was 0.08, because the cats appeared unnaturally embedded in the shoes.

On two benchmarks, COCO-T2I and GenAI-Bench, TerraVis showed the strongest pooled agreement with human ratings among the metrics tested. Its Spearman correlation (a rank-based measure, where 1 is perfect agreement) was 0.44 on COCO-T2I and 0.43 on GenAI-Bench. The strongest existing metric, the Human Preference Score variant RAHF, reached 0.29 and 0.37. These are differences in correlation values, not percentage gains. Several existing metrics showed near-zero or negative correlations on individual models, and a holistic MLLM prompt the authors tested, GPTScore, showed almost no agreement with human judges. Replacing the proprietary GPT-5.5 model with the open-source Gemma 4 31B left results comparable on COCO-T2I (0.42) and higher on GenAI-Bench (0.51). Removing the object-level checks lowered correlation by 0.149, while the choice between aggregation formulas made only small differences.

The study also ranked five generators: three open-source and two proprietary. GPT Image 1.5 ranked first on most conventional metrics, but FLUX.2-dev was rated first on world consistency on GenAI-Bench by both TerraVis and human raters. The authors speculate that it generates simpler scenes, though they do not test this directly.

The limitations are substantial. The multi-stage workflow requires several MLLM calls per image and is less efficient than single-pass metrics. Detection also falters: TerraVis flagged a cat’s coiled tail that all three human raters judged normal, and it missed distorted hands in a FLUX.2-dev image that all three raters scored 1 out of 5. Human raters themselves sometimes disagreed across the full scale. The taxonomy may miss violation types, and the human reference, based on about 13,000 annotations from Amazon Mechanical Turk workers, carries its own uncertainty.

The work suggests that a generated image’s visual appeal and its consistency with the real world are distinct qualities, and that the second can be measured in a structured way. The results are correlational, limited to two benchmarks and five generators, and depend on a taxonomy the authors designed. They support world consistency as a complementary evaluation dimension rather than a settled standard.