A Three-Agent AI Judge Tries to Grade Teaching Videos for Their Intended Learners
Teaching videos are multiplying faster than people can review them. Judging whether a video teaches well requires trained raters who watch repeatedly and compare narration, on-screen visuals, learning objectives and the intended audience. Researchers at National Taiwan University argue that large language models (LLMs), AI systems that read and generate text, could help with this work if they can handle video and account for who the lesson is for. Their system, EduPanel, is described in a preprint on arXiv.
The central idea is that teaching quality depends on the learner. A detailed explanation of neural networks may suit advanced students and overwhelm beginners. EduPanel therefore scores each video against a course requirement and a student persona that specifies grade level, attention budget and prerequisite knowledge. The work is split among three agents, all running on the same model, gemini-3-flash. The first agent watches the video and produces a timestamped content map and a list of possible factual or visual problems. The second reads that report as text only and scores two objective dimensions. The third watches the video alongside the persona and scores the remaining dimensions, including vocabulary, prerequisite awareness and pacing. Each video is scored three times and the results are averaged. The authors say this structure exposes intermediate evidence that a reviewer can inspect.
The paper offers a concrete case of that evidence. In a meiosis video, the narration placed DNA replication in Prophase I, when it actually occurs during the S phase of Interphase. All 12 expert raters who watched the video initially accepted the explanation, and one said they lacked the domain knowledge to judge it. EduPanel flagged the error.
The study was small. It used 12 videos across physics, biology, mathematics and computer science, rated by 12 experts from Taiwanese universities. Against an AI-free human consensus, EduPanel’s mean absolute error (MAE, the average gap in points on a 1–5 scale) was 0.85, close to the median individual expert’s 0.87. When experts rated blind and then again with the AI’s scores and rationales visible, their MAE fell from 0.87 to 0.73, a drop of 0.14 points. Eight of the 12 experts improved. Agreement among raters, measured by Krippendorff’s α, rose from 0.38 to 0.50. The researchers also planted deliberately wrong AI scores. Experts gave planted errors lower agreement than correct outputs, yielding a ROC AUC of 0.77, meaning a planted error drew less agreement than a correct one 77% of the time. Experts rarely changed their own scores, however, and revised at similar low rates on correct and flawed outputs.
The judge also had clear weaknesses. On the Gemini backbone, it was more lenient on visually grounded dimensions, with an average signed bias of +0.67 compared with +0.24 on transcript-only dimensions. The authors did not observe this gap on a GPT backbone. Removing video input raised MAE to 1.07. Collapsing the agents into one model call kept MAE at 0.55 but produced a narrower score range. Persona changes shifted vocabulary and prerequisite scores in the expected direction, while pacing responded only weakly.
Several caveats limit the findings. A single researcher adjudicated the reference ground truth after reviewing AI commentary, so agreement with it is not independent validation. The expert study had no no-AI control, so the authors describe the improvement as associational. The gain also shrank to between 0.05 and 0.10 points under alternative ground-truth rules.
The authors conclude that learner-conditioned, multimodal judges could support human reviewers, but should not replace them. The work does not establish that EduPanel reliably evaluates teaching videos in general. It does suggest that judging whether a video fits a particular learner is a tractable target, and that the judge’s visible reasoning may help experts catch its mistakes.
Chat about this paper
To chat about this paper, you'll need a free Gemini API key from Google AI Studio.
Your API key will be stored securely in your browser's local storage.