AI Papers Reader

Personalized digests of latest AI research

View on GitHub

Benchmark Tests Whether AI Can Meet Expert Standards in Physics, Chemistry, Finance and Consulting

AI systems are increasingly asked to write long reports and analyze professional documents, but most benchmarks test math, code, or multiple-choice questions with clear answers. A team at NVIDIA has introduced ProfBench, which grades open-ended professional work against criteria written by credentialed experts.

The benchmark contains 7,347 expert-written response-criterion pairs across 80 tasks, split evenly among Chemistry PhD, Physics PhD, Finance MBA and Consulting MBA domains. Annotators, whom the authors say hold a PhD, an MBA, or equivalent experience, wrote prompts and typically 15 to 60 criteria per task. Each criterion is a yes-or-no requirement a strong answer should meet. One Finance MBA task, about the International Finance Facility for Immunization (IFFIm), includes the criterion that a response state that a breach of IFFIm’s liquidity policy could negatively affect its rating profile. Annotators then judged three model responses, from OpenAI’s o3, xAI’s Grok 4 and DeepSeek R1-0528, against each criterion.

Checking every criterion by hand is costly, so the researchers tested automated “LLM judges,” language models that decide whether a response meets a criterion. They scored judges on agreement with human labels and on a bias index measuring whether a judge systematically favors or penalizes responses from particular models. The chosen judge, the open-weight GPT-OSS-120B, reached 78.2% overall, matching Gemini-2.5-Pro, while costing $0.70 for a full evaluation compared with $1,320 for PaperBench’s judge evaluation. Two other experts re-labeled 1,127 pairs, yielding a Fleiss’ kappa of 0.912, which the authors describe as excellent agreement.

Among more than 40 models tested as report generators, GPT-5 with high reasoning effort scored best at 65.9%, ahead of o3 at 61.4% and Gemini-2.5-Pro at 60.3%. The strongest open-weight models were GPT-OSS-120b at 54.9% and DeepSeek V3.1 (Thinking) at 53.8%. The closed-versus-open gap was under one percentage point in physics but 15.0 points in finance. Physics was the hardest domain for GPT-5, at 49.3%. The authors speculate that open models have received more attention on code and math benchmarks than on finance, chemistry, or consulting, though they did not test this directly.

Source documents mattered considerably. Without them, o3 and o4-mini often asked for missing figures rather than answering; one request sought a REIT’s Q1’25 net operating income and other data. Removing the documents lowered scores by 9.4 percentage points for o3 and 11.9 for o4-mini. Web search recovered 4.1 to 7.0 points, but it required 0.12 to 0.24 million input tokens per task, which the authors flag as a retrieval-precision problem.

Length was not a reliable shortcut. o3 scored 61.4% with responses averaging about 4,158 characters, while GPT-5-nano scored 50.1% with responses near 9,796 characters. Sixteen o3 samples per task cost $48, compared with $300 for HealthBench and $8,000 for PaperBench. The authors say allocating samples according to score variance cut costs to 25% without compromising estimates.

The study has limits. Only half the dataset is public, all tasks are English and text-only, and the tasks were written in July 2025 against the models of that time. The judges, though strong, still depart from human labels, and the authors’ bias measurements are relative to those labels.

The authors present ProfBench as a way to measure open-ended professional work more affordably and fairly. It does not establish that models can perform expert work, and its value will depend on whether future models and independent evaluators reproduce these rankings.