Back to Search View Original Cite This Article

Abstract

<jats:p>Objective To develop and evaluate a framework for human-AI interaction. This approach, SHARE (Synergistic Human-Agent REasoning system) was designed to support scalable phenotyping of complex outcomes accurately, robustly and reproducibly from real-world electronic health record (EHR) data to support real-world evidence (RWE) generation. Methods and Analysis Using rheumatoid arthritis (RA) disease activity as the use-case, we studied a multi-institutional EHR-based RA cohort of 3,167 patients. Expert reviewers and a disease activity agent labeled notes using the same review guideline. The agent combined embedding-based informative-note filtering, structured evidence extraction, and evidence-based integrated reasoning to assign disease activity categories with supporting evidence, rationale, confidence, and ambiguity flags. To support scalable deployment, we evaluated a budget-tiered configuration using GPT-5 Nano for high-volume evidence extraction, o4-mini for final reasoning, benchmarking against a GPT-5.4 high reasoning effort configuration applied at every step. Note-level discrepancies were adjudicated by reviewers into final co-produced labels that were used to refine labels and inform agent development. The main outcome measure was the mean absolute error (MAE) of the initial and final agent vs the final co-produced labels. The agreement between agent- and reviewer-flagged ambiguous notes, per-note cost and compute time across configurations were also tested. Results Expert reviewers labeled 626 notes from 273 patients; human-AI adjudication revised 127 (20%) of these initial labels and added 60 newly labeled notes, yielding a 686-note co-produced reference. Against this reference, the final agent's accuracy improved from a mean absolute error of 0.406 to 0.291 with co-learning, and its ambiguity flag agreed with expert ambiguity designations with 92.1% accuracy. Applied across the cohort, the agent labeled 101,691 notes; the budget tiered configuration matched the accuracy of GPT-5.4 at high reasoning effort while reducing estimated cost by 69% and compute time by 70%. Conclusion Adopting a framework for human-AI co-learning, SHARE, improved the overall quality of gold-standard labels, identified ambiguous cases for further review, and supported accurate and standardized chart reviews of disease activity at a scale infeasible for manual review. SHARE's resource efficiency provides a transferable approach to incorporate complex phenotypes in RWE studies.</jats:p>

Show More

Keywords

agent reasoning notes final labels

Related Articles

PORE

About

Connect