Abstract
<title>Abstract</title> <p>Retrieval-augmented generation (RAG) increasingly powers AI tutors and course-support assistants in higher education, yet institutions still need clearer ways to judge whether retrieval-benchmark improvements translate into better student outcomes. This paper proposes an Integrated Offline–Online Evaluation Framework (IOEF) linking offline information-retrieval metrics (Recall@k, nDCG, MRR, latency, faithfulness) with online learning-outcome evidence gathered under a staged, randomized rollout. The framework is illustrated with three independent, peer-reviewed, public sources spanning its offline and online layers, rather than with a new institutional pilot. The MS MARCO and BEIR retrieval benchmarks show that adding a cross-encoder re-ranker to a BM25 first-stage retriever can raise MRR@10 from about .187 to .365, while stronger ranking often requires more computation (Nogueira & Cho, 2019; Thakur et al., 2021). A study of nearly 1000 high-school mathematics students found that an unstructured AI tutor improved practice performance by 48% but was followed by a 17% decline when support was removed; a scaffolded version improved practice performance by 127% without the same decline (Bastani et al., 2025). A university physics trial reported learning gains above twice those of an active-learning lecture, with stronger engagement ratings (Kestin et al., 2025). Taken together, these literatures suggest that retrieval-quality gains alone cannot establish educational value: each learning study contrasts two versions of an AI-supported experience and shows that instructional design shapes whether the tool supports or impedes learning. The synthesis is used to specify IOEF's three layers and to offer implementation checks for institutions adopting retrieval-augmented AI.</p>