AI tumor-board benchmark finds large gaps between models and real specialists
Multidisciplinary tumor boards are a critical decision point in cancer care, especially when standard options run out, but limited specialist availability can delay them. Large language models (LLMs) have been proposed as virtual specialists that could help, yet existing tests largely rely on exam questions or synthetic cases rather than real board discussions. A new preprint introduces OpenTumorBoard, a benchmark built from 611 patient cases and 19,157 discussion turns across ten specialist roles, transcribed from 12,534 minutes of publicly recorded YouTube tumor board meetings. The researchers say it is the first benchmark to provide complete naturally occurring discussion trajectories paired with observed board consensus.
OpenTumorBoard uses a nine-step automated pipeline to transcribe and label speakers, segment cases, extract presentation slides, align slide evidence with utterances, and remove information that could leak the board’s final decision. It defines two tasks. In SpecialistTurn, a model plays one specialist and answers an actual question posed during the meeting; the test set contains 184 cases and 4,844 such questions. In BoardSimulation, a model must simulate the entire back-and-forth discussion and reach a conclusion on therapy, surgery, next actions, and clinical-trial matching.
A case study shows why that back-and-forth matters. In one discussion, an early proposal to re-irradiate the prostate bed was retracted after the radiation oncologist identified a prior 71.8 Gy dose, making re-irradiation unsafe. After 32 turns, the board converged on surveillance keyed to PSA doubling time. The example illustrates how specialists correct one another using longitudinal treatment history, a capability the benchmark tests.
Results show substantial limitations. Across nine models on SpecialistTurn, the best clinical equivalence to real specialist answers was 3.43 out of 5, achieved by DeepSeek-V4-Pro in reasoning mode. Critical-error rates ranged from 5.6% to 43.5%, and unsupported-claim rates from 10.0% to 74.5%. Medical-specific models did not outperform general-purpose models; HuatuoGPT-3-32B and MedReason-8B had the highest unsupported-claim rates, at 74.5% and 51.9%. The authors report that questions requiring supporting evidence, such as evidence discussion and clinical-trial suggestion, were consistently the hardest. On BoardSimulation, which evaluated 14 models, the best conclusion alignment was 2.78 out of 5, from Gemini 3.7 Flash. Longer model discussions correlated with higher alignment (Spearman ρ = 0.63), but no model reliably recovered the real board’s conclusion. Supervised finetuning of Qwen2.5-VL-3B raised its alignment from 1.58 to 1.67; reinforcement learning further increased it to 1.86, an approximately 18% relative improvement over the base model, the authors report.
Three M.D. experts reviewed a random subset. All reviewed case summaries and consensus conclusions received “High” ratings for coverage, factuality, and consensus fidelity; 92.6% of extracted answers received “High” ratings for correctness and 94.4% for discussion support. The authors caution that the benchmark comes exclusively from public YouTube recordings, which may limit representativeness, and that their rubric-based LLM judging still needs validation against clinician assessments. OpenTumorBoard may offer a training and evaluation ground for multidisciplinary cancer decision-making, but the current scores indicate that LLMs remain far from reliably reproducing real tumor board decisions.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.