Abstract
<p>Personalizing digital well-being interventions to fit the needs and preferences of individual users has been a longstanding challenge. Large language models (LLMs) may be well-suited for this task, using natural language reasoning to integrate heterogeneous user information into personalized guidance. We tested whether LLMs could generate intervention recommendations that were sensitive to differences across recipients’ psychological profiles. Across 178 undergraduate student profiles spanning 14 psychological measures, four LLMs recommended 10 scripted chatbot dialogues from a library of 222 options. Next, two experts rated the recommendations’ appropriateness for a given profile (1-10). For half of the profiles, raters saw recommendations generated for that student (matched condition); for the remainder, they saw recommendations generated for a different student (unmatched condition). As hypothesized, matched recommendations (M = 6.74) were rated as more appropriate than unmatched ones (M = 5.64, SMD = 0.50). These results provide evidence that LLMs can adapt their intervention recommendations to match psychological profiles. However, inter-rater reliability was poor (ICC = 0.22), and all LLMs showed substantial sensitivity to catalog order. As such, we stress the preliminary nature of the findings. Whether such tailoring translates into improved well-being outcomes for users remains to be tested.</p>