AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Chart2Code Benchmark Finds Gaps Between Running Code and Faithful Charts

Researchers have introduced a benchmark that tests whether large multimodal models (LMMs), AI systems that process both images and text, can write plotting code that reproduces or modifies a chart from a user’s request. Charts are central to scientific papers and business reports, and automating their creation could save time and support reproducibility. The authors argue that existing tests no longer separate the strongest systems. They note that GPT-4o already scored 82.2% on ChartMimic, an earlier benchmark.

The new benchmark, Chart2Code, contains 2,186 tasks across 22 chart types in three levels of rising difficulty. Level 1 asks models to reproduce a reference chart, either from the image alone or with supplied data. Level 2 asks for edits, such as changing chart types, recoloring elements, or adding a trend line and a summary chart. Level 3 asks models to turn long, unprocessed tables into charts that follow a user’s instructions. The Excel files in this level average 2,647 lines.

Scoring has two parts. The first is an “executable rate,” meaning the share of generated code that runs. The second compares outputs with ground-truth charts in three ways: a rule-based score of plot properties (Base), a judgment of the code by GPT-5-mini (LLM), and a GPT-5-mini judgment of the rendered chart against the reference (LMM). The authors also ran a study with 20 undergraduate students and report a strong correlation between human ratings and the LMM scores.

The paper’s example shows how running code and matching a chart can diverge. On Level 1 tasks that supplied figure-format data, Gemini-3-Pro produced code that ran 99.07% of the time, yet its chart-level LMM score was 32.85 on a 0–100 scale. When raw tabular text was supplied instead, the same model’s code ran 100% of the time, but its LMM score was 40.72.

Across 29 models, the authors report that proprietary systems executed code at high rates on Level 2, with Gemini-3-Pro at 97.23%, GPT-5.2 at 96.04%, and Claude-Sonnet-4 at 90.20%. Chart-level scores, however, ranged from 17.43 to 33.41. Gemini-3-Pro had the highest LMM score at 33.41, followed by GPT-5.2 at 33.03. The abstract attributes the 72.21 code-based score and the 33.41 chart score to GPT-5.2, but Table 4 lists those values for Gemini-3-Pro, so readers should treat the attribution as unresolved. Open-source models generally scored lower, with the best reported Level 2 execution rate at 74.55% for Qwen2.5-VL-72B and a chart score of 16.69.

Performance fell further on Level 3. Execution rates dropped to 30.03% for Gemini-3-Pro, 46.65% for Claude-Sonnet-4, and 20.13% for GPT-5.2. Chart-level scores were 35.97, 16.60, and 16.29, respectively. The authors attribute many failures to inputs exceeding model context limits, and several open-source models could not complete the task at all. They suggest models struggle to hold long tables and reference images together while extracting the plotting requirements, though this explanation was not separately tested.

The study has clear limits. All tasks are in English. The LLM and LMM scores depend on GPT-5-mini as a judge, which the authors acknowledge may introduce bias, and the human study involved 20 students. The results also show that code-level scores track one another but correlate poorly with visual similarity.

The work offers a more demanding test of chart generation than earlier benchmarks. It suggests that producing executable code is not the same as producing a faithful chart, particularly when data are long or edits are complex. The results describe benchmark performance under the authors’ setup and do not show that such models can replace analysts in practice.