Abstract
<title>Abstract</title> <p>Large language model (LLM)-based mentoring tools reached technology entrepreneurship programmes faster than the field developed ways to examine how they are built. This systematic literature review characterises the architecture of these systems across a corpus of 25 studies (24 empirical investigations and one systematic review) synthesised under PRISMA 2020 guidelines. Three research questions organise the analysis: 1) whether systems ground feedback in verifiable taxonomies such as the Technology Readiness Level or rely on conversational heuristics without that grounding; 2) how choices about memory, retrieval, agent structure, and case-based reasoning shape the feedback a system can give; and 3) what current studies measure. Two architectural patterns dominate the corpus and form its primary contribution. Three-quarters of systems rely on conversational heuristics without taxonomic grounding, producing authoritative feedback that cannot be interrogated; nearly nine in ten discard founder state between sessions, leaving no basis for calibrating challenge over time. These patterns are not independent: persistence and agent decomposition concentrate in the systems that ground assessments in explicit taxonomies, so the paradigm producing the least verifiable feedback is also the least equipped to develop the founder. A secondary finding concerns evaluation: studies consistently measure satisfaction and task output while leaving cognitive development unmeasured. The one controlled study found AI-assisted students reporting high satisfaction yet scoring significantly lower on critical thinking than human-mentored peers, a pattern echoed through weaker designs across five further studies. We formalise this as the cognitive offloading trap, a proposition for future testing: feedback optimised for satisfaction may systematically remove the cognitive friction on which developmental outcomes depend.</p>