He LabPackage 0.3.0
Methodological sensitivity of healthcare AI benchmarks
A review and controlled pilot asking how temperature, repeat count, reporting, graders, and inter-rater reliability change the meaning of healthcare AI benchmark results.
Open the studyEvidence at a glance
28
evidence sources
11
methods audited
8
guided research stages