Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Background Structured feedback is central to ethics education, but expert-generated feedback requires substantial faculty time. Artificial intelligence (AI) may offer a scalable way to support ethical decision-making education, provided that generated content remains expert-reviewed and educationally appropriate. This pilot randomized study evaluated the feasibility and exploratory educational outcomes of AI-generated feedback compared with human expert panel feedback within a Concordance of Judgment Learning Tool (CJLT)-based ethics education intervention. Methods This single-center, two-arm, parallel-group pilot randomized controlled study included 56 sixth-year medical (intern) students. Fifty-six sixth-year medical students were randomized equally to receive either expert-reviewed AI-generated feedback or expert panel feedback while completing the same CJLT ethics scenarios. Quantitative measures were administered before and after the intervention, including an Objective Structured Video Examination (OSVE), Script Concordance Test (SCT), and Ethical Decision Bias Scale (BIAS). OSVE was the primary preliminary educational outcome. Participant flow, completion rates, missing data, CJLT completion, and voluntary three-month acceptability feedback were assessed as feasibility outcomes. The OSVE consisted of 20 video-based scenarios. Post-test OSVE scores were analyzed using ANCOVA, with baseline OSVE scores entered as a covariate; change-score analyses were conducted as sensitivity analyses. Results All participants completed the T0 and T1 quantitative assessments, and no missing data occurred for the quantitative outcomes. T2 acceptability feedback was obtained from 21 participants. The baseline OSVE mean score was higher in the expert-reviewed AI-generated feedback arm than in the expert panel feedback arm (23.32 ± 4.71 vs. 19.36 ± 3.65; p &lt; .001). The post-test OSVE mean score was 27.29 ± 4.65 in the expert-reviewed AI-generated feedback arm and 23.32 ± 5.45 in the expert panel feedback arm. When baseline OSVE score was included as a covariate, the group effect was significant, F(1,53) = 8.953, p = .004, η²p = .145; the adjusted mean difference favored expert-reviewed AI-generated feedback by 4.50 points (95% CI: 1.49–7.52). In contrast, the OSVE change-score analysis showed no between-group difference (mean difference = 0.00; p = 1.000). No significant group differences were observed in SCT or BIAS analyses. Among 21 participants completing voluntary follow-up feedback, acceptability was generally favorable; 95.2% stated that they would recommend the training to other medical students. Conclusions AI-generated feedback was feasible and acceptable within CJLT-based ethics decision-making education and did not show poorer preliminary educational performance than expert panel feedback. The exploratory OSVE signal favoring AI-generated feedback should be interpreted cautiously because of baseline imbalance, pilot sample size, repeated assessment scenarios, and nonsignificant complementary outcomes. Larger confirmatory studies with stronger randomization procedures, prespecified outcome strategies, parallel assessments, and longer follow-up are needed. Trial registration: Not applicable. This is an educational study designed to evaluate teaching methodologies. It does not involve healthcare interventions, clinical treatment, or the measurement of patient-related outcomes.</p>

Show More

Keywords

feedback osve aigenerated expertreviewed outcomes

Related Articles

PORE

About

Connect