Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> Public benchmarks for frontier AI agents (GAIA, OSWorld, AgentBench, GDPval, BrowseComp, SWE-bench) are uniformly <italic>researcher-imagined</italic> : tasks are designed by AI researchers, in clean declarative form, with ground-truth answers fixed in advance. Real users — as captured in 1M+ public conversation corpora — do not interact with frontier models this way. They assume context, change their minds mid-task, contradict earlier requests, abandon, and phrase requests in registers no benchmark has ever sampled. We hypothesize that benchmarks systematically miss the deployment-relevant failure mode we call <bold>the intent gap</bold> : cases where the model literally answered the prompt but missed what the user actually wanted. We mine WildChat-1M and LMSYS-Chat-1M using a three-stage <italic>frustration-signal filter</italic> (regex → embedding → LLM-as-judge) to surface a target of ~ 150–300 high-quality intent-gap conversations. We replay each prompt across four current frontier models (Claude, GPT, Gemini, and one open baseline) and grade each replay on two axes: literal compliance and intent satisfaction. We then triangulate the resulting taxonomy against ~ 25 published failure-mode taxonomies from research institutes worldwide and against 5 documented production failure incidents (lawsuits, regulatory filings, public post-mortems). Headline preliminary findings (from a 50,000-conversation v2 pilot, before cross-model replay or full hand-review): 1. <bold>52.8% of English real-user conversations end after a single exchange.</bold> A substantial fraction of these are likely silent abandonment after intent-gap failure — invisible to deployment-quality pipelines built on user-feedback signals. 2. <bold>Only 3.85% of multi-turn conversations contain an explicit repair signal</bold> , even under a permissive 45-phrase filter that scans every user turn. The intent gap exists, but users rarely verbalize it. 3. <bold>The phrase-list expansion from v1 (17 academic phrases) to v2 (45 natural-language phrases) increased hit rate by ~ 31×.</bold> This implies the dialogue-breakdown detection literature undercounts real-world repair by more than an order of magnitude because its phrase lists are researcher-imagined. 4. <bold>Among successfully-judged v2 candidates, 79% (15/19) were confirmed as real intent-gap failures</bold> by an LLM judge. When the regex fires and the judge can run, the flagged candidate is overwhelmingly a real failure. 5. <bold>The judge model refused to grade 62% (31/50) of real-user candidates</bold> , most plausibly on content-moderation grounds. This is itself a deployment-relevant finding: safety-filtered judge models systematically miss failures in the most permissive parts of the user-prompt distribution. A hand-reviewed and LLM-judge-filtered corpus of ~ 150–300 high-quality intent-gap conversations is released alongside the filter pipeline, the grading rubric, and the triangulation matrix mapping our taxonomy onto ~ 25 prior published failure taxonomies plus 5 documented production-failure incidents. </p>

Show More

Keywords

failure intentgap conversations judge public

Related Articles

PORE

About

Connect