Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 184
Back to Search View Original Cite This Article

Abstract

<jats:p>Large language models (LLMs) are increasingly proposed as assistants for chemical process modeling, yet existing benchmarks grade them with LLM judges, which are lenient and not reproducible. We introduce PROBE, an executable benchmark that scores LLM-generated first-principles process models objectively, and report an initial pilot study across ten reference tasks. PROBE evaluates contract-bound implementation—turning a specified model into correct, simulable code under a fixed input/output contract—rather than unconstrained model formulation from an underspecified engineering description. Each generated model is checked on five levels—runnability, dimensional consistency (units), structural correctness against a reference derivative field, physical invariants (mass/energy conservation, boundary and limit behavior), and numerical agreement with a reference simulation—with no LLM judges. The benchmark spans ordinary-differential problems (reactors with mass and energy balances), a differentialalgebraic problem (vapor–liquid flash / MESH equations), and a coupled multi-unit flowsheet with recycle. The benchmark comprises ten problems: six canonical textbook unit operations, a harder tier (a method-of-lines plug-flow reactor, i.e. a distributed-parameter PDE; the Van de Vusse multiple-reaction network; and a non-ideal flash with NRTL activity coefficients), and one procedurally generated reaction network that appears in no textbook. All models—frontier and local—are evaluated under one protocol: a single prompt, no tools, no iterative revision, five samples each. Two findings emerge. First, within the scope of PROBE the three Claude models (Opus 4.8, Sonnet 5, and the low-cost Haiku 4.5) are strikingly reliable: across all ten problems we observed no failure in 150 generations (50 each), an upper 95% bound of ≈ 2% on the pergeneration failure rate by the rule of three—under a pooled Bernoulli interpretation restricted to this evaluated mixture of models and tasks, not a rate to be generalized to unseen processmodeling problems. This holds on the harder tier and on the procedurally generated network alike, and—notably—the inexpensive Haiku 4.5 is as reliable as the flagship models, so frontiergrade performance on these tasks is available at a low-cost tier. Google Gemini 2.5 Flash is near-ceiling as well, dipping only on the non-ideal NRTL flash and on a reproducible parameternaming slip, so the reliability spans two vendors. This is not what judge-based benchmarks report, and the novel-network result argues against mere memorization. We are careful not to claim that LLMs have “solved” process modeling: the evidence is that, on first-principles unit operations with a fixed input/output contract, current frontier models rarely make physics errors—not a statement about full industrial flowsheets or free-form modeling. Second, the correctness gap is concentrated in smaller, locally deployable models (3–14 B), whose outputs are moreover highly stochastic. While a physics-validation feedback loop can produce striking within-conversation repairs, at a matched generation budget we found no evidence, in this limited experiment, that it outperforms simply resampling and keeping the best-scoring candidate (paired Wilcoxon signed-rank: no significant difference). That comparison is made essential by the same stochasticity, which also renders single-run evaluations of such loops misleading. In one controlled comparison, adding a reasoning trace did not close the gap either: a single samesize open reasoning model lifted the mean only marginally over its non-reasoning counterpart, still failed the flash on the identical contract-naming slip, and on some hard problems reasoned past a large token budget without ever emitting an answer—consistent with capability, rather than the amount of feedback or inference-time computation, being the binding constraint on these models.</jats:p>

Show More

Keywords

models problems flash model process

Related Articles


Deprecated: Function curl_close() is deprecated since 8.5, as it has no effect since PHP 8.0 in /home/u483256323/domains/poorvam.com/public_html/subdomains/pore/includes/api.php on line 76
PORE

About

Connect