@GaoShanghua @ayushnoori @CurtGinder @BenGlicksberg @RanBalicer @HarvardDBMI @KempnerInst @biohub @IcahnMountSinai @ClalitHealth @BrighamWomens 🟢 Benchmarking: Across 3,168 drug reasoning questions (DrugPC) and 456 patient specific treatment cases (TreatmentPC), ATHENA hit 94.7% and 82.9% accuracy, beating GPT-5 by 17.8 and 10.7 points and DeepSeek-R1 (671B) by even more. What stood out: giving GPT-5 optional tool https://t.co/xdmOHeiS8U #AI #DrugDevelopment #HealthcareInnovation #Nospecificcountriesarementionedinthetweet.






