Rhetorical Rewrites Can Shift AI Peer-Review Scores Without Changing the Science
Large language models are beginning to sit on both sides of scientific peer review: they help authors rewrite manuscripts and act as automated reviewers. The concern is a form of reward hacking—winning higher scores by changing presentation rather than improving the science. A new study finds that content-preserving rhetorical changes can indeed move AI reviewers’ overall assessments, even when methods, experiments, and reported numbers stay the same.
To measure the effect, researchers built 4,200 full-paper manuscripts from 120 anonymized ICLR 2026 submissions. Two rewriting models, GPT-5.5 and Opus 4.8, rewrote each paper along six rhetorical dimensions in opposing directions: assertive versus cautious novelty claims, broad versus narrow scope, prominent versus de-emphasized evidence, explicit versus integrated contributions, formal versus plain technical register, and sophisticated versus simple prose. Five AI reviewers—Gemini 3.5 Flash-Lite, Qwen 3.5 Flash, GPT-5 mini, GPT-5.5, and Claude Sonnet 5—scored originals and rewrites under a standard prompt and a stricter prompt that told reviewers not to reward presentation.
One configuration illustrates the mechanism. Under the strict protocol, when Opus 4.8 rewrote papers to foreground reported comparisons, Gemini 3.5 Flash-Lite’s average overall assessment rose by 0.93 points on the 10-point scale; when the same rewriter made novelty claims more cautious, the score fell by 0.73 points. The authors say the pattern is selective, not a blanket payoff for polish. Evidence framing and novelty stance produced the largest positive-negative contrasts, with scope framing second. Evidence framing also shifted the share of weak-accept ratings—scores of 6 or higher—by 13 percentage points on average.
Direction depended on where the AI reviewer began. Low initial scores tended to rise after rewriting; high initial scores tended to fall; contrasts were clearest in the middle. The authors note that this may partly reflect scale bounds and regression to the mean. The dimension hierarchy persisted across human-assessed quality levels. A stricter review prompt lowered the average overall assessment by 1.36 points, from 6.296 to 4.934, but did not consistently reduce sensitivity to rhetoric.
More elaborate workflows did not pay off uniformly. Joint rewriting that combined all six positive interventions gave meaningful gains with Opus 4.8—+0.289 points under standard review and +0.463 under strict—but near-zero changes with GPT-5.5 (+0.021 and +0.045). Reviewer-guided rewriting did not consistently beat an unguided second pass, and repeated rounds gave diminishing returns. The rewriter largely determined how far apart positive and negative versions were; the reviewer determined the size and sign of the score movement.
The results come with caveats. The corpus covers only ICLR 2026 submissions with recoverable LaTeX source, so findings may not transfer to other venues or fields. Direct API costs reached $29,165.89, and most configurations used one review per manuscript, so the study characterizes specific models and prompts rather than human review or sampling variability. The authors also warn that higher scores from rewriting do not mean better science; similar techniques could be used to tailor papers to AI-reviewer preferences.
The broader implication is that AI-based review should be stress-tested for rhetorical robustness across models and conditions, not judged by average agreement alone.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.