AI Papers Reader

Personalized digests of latest AI research

View on GitHub

LLMs Can Match Behavior Without Recovering the Rule Behind It, Study Finds

Large language models are increasingly used as interactive agents and behavioral simulators, standing in for users, opponents, or human players. A simulator is only useful if it captures how behavior unfolds over time, not just how often each action occurs. A new study by Jerry Wang and colleagues at the University of Illinois Urbana-Champaign, Stony Brook University, and National Chengchi University suggests that matching overall action frequencies can hide a failure to grasp the sequential rule behind them.

The researchers used Rock–Paper–Scissors because the process generating each player’s moves is known exactly. Some “statistical” players sample from fixed probabilities and ignore history. “Markov” players instead follow rules tied to recent events, such as responding to the opponent’s previous move. A Markov rule, in which the next choice depends only on a limited window of recent history, can produce the same overall share of Rock, Paper, and Scissors as random play. Models were shown a trajectory, asked to infer each player’s strategy, and then asked to continue the game.

Identification proved harder for Markov players. Across seven model configurations, accuracy on Markov players was lower than on non-Markov players in every case, and the gap was statistically significant in six. DeepSeek Reasoner identified Markov players 82.0% of the time, compared with 96.5% for non-Markov players. Qwen3-8B scored 25.0% and 45.5%, respectively. Longer observation did not help: expanding the history from 100 to 1,000 rounds generally left accuracy unchanged or lowered it. A non-LLM maximum-likelihood baseline identified the players perfectly, which the authors read as evidence that the information was present in the trajectories.

The clearest failure appeared during generation. When models identified a Markov player correctly, their simulated moves were far more distorted in distribution terms: mean squared error for Markov players was roughly 102% to 118% higher than for non-Markov players, depending on context length. Among 282 incorrectly recovered strategies, 41.8% still produced action frequencies within 0.02 total variation distance of the target. In other words, the output looked statistically right while the mechanism was wrong. Teacher-forced accuracy, in which models predicted each move from true history, ranged from 52.5% for DeepSeek Chat to 93.3% for DeepSeek Reasoner. The authors interpret this spread as separating models that struggle with local rule application from those that lose the rule over long generations.

In a one-player follow-up using stochastic n-gram sequences, where each symbol depends on the previous n symbols, performance remained generally positive at orders 1 through 4 but dropped sharply at order 8 for most models. Supplying the transition rules and lengthening prefixes from 256 to 2,048 symbols did not remove the degradation. The authors conclude that the bottleneck lies in maintaining higher-order conditional structure during generation rather than in gathering evidence.

The study has clear limits. Players were drawn from a predefined set of strategies, so it does not test open-ended discovery of unfamiliar behavior. Results also depend on the symbolic prompt formats and decoding settings chosen, which the authors acknowledge they cannot guarantee are optimal. Still, the work suggests that evaluations relying on aggregate behavior may overstate how well a model understands the process it simulates.