- 1次围观
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous agents. Illustrated in Question-Answer pairs, multi-turn conversations and autonomous data-analysis, these dimensions together moves evaluation beyond scoring, enabling trustworthy decision-making with AI in biomedicine.
来源出处
Turning Domain Expertise into Multi-Dimensional Evaluation of Biomedical AI w…
https://www.biorxiv.org/content/10.64898/2026.09.01.748513v1?rss=1