Abstract
<jats:p>A message-passing neural network is a learned group-contribution method only if its readout carries molecular size, which the framework-default mean readout discards. We train directed message passing neural networks (D-MPNN, Chemprop v2) on experimental thermophysical data from the Cheméo database, one model per property, as a general-purpose graph-based alternative to the Joback group-contribution method [14], strongest in the regime where Joback degrades: Joback was fitted on small molecules and, as we show here head-to-head, degrades steadily as molecules grow; those larger molecules are a minority of the corpus (a quarter to a third of a typical set has eleven or more heavy atoms, though the spread is wide, from 11% for the small Gibbs-energy set to 56% for melting point), but they are exactly the regime where a group-contribution method has no good alternative. The contribution is not that sum pooling beats mean pooling, a fact already established in the graph-learning literature [18,19]; it is a property-by-property diagnosis on experimental thermophysical data of how the default silently fails, a reusable release discipline, and a family of released models measured against Joback, a simple descriptor baseline, and external data. The mean readout is size-invariant: it returns the average per-atom contribution, so a molecule and its dimer map to nearly the same vector, and every property that tracks molecular size collapsed toward the training mean; six properties, among them the critical constants and the Gibbs energy of formation, landed at test R2 near zero, and a gradient-boosted tree on plain RDKit descriptors beat the graph network on all six. Replacing the mean with a sum-aware pooling (sum and max concatenated) makes the molecule embedding scale with atom count, and the message passing then learns the higher-order structural corrections that Marrero and Gani add to a classical scheme by hand as second-and third-order groups [15]. The distinction we keep is between a graph-extensive readout (a property of the network) and a thermodynamically extensive target: the sum is needed outright for extensive targets such as critical volume and the formation enthalpies, while intensive targets that rise along homologous series, such as the boiling and critical temperatures, benefit because molecular size is a strong predictor of them, not because they are extensive. Nine of sixteen properties are released under a two-part criterion: a statistical bar (mean out-of-sample R2 above 0.70 with a train-test gap of at most 0.20 under repeated scaffold cross-validation, or state-of-the-art accuracy on a shared external benchmark) and a veto, that the incumbent Joback must not significantly out-predict the model in its tested domain; we lead with mean absolute error, since a small scaffold test set makes R2 noisy: normal boiling point (MAE 13.5 K, RMSE 30 K, 3.5% average relative error), melting point (MAE 30.8 K), critical volume and critical temperature, enthalpies of formation and vaporisation, ionisation energy, aqueous solubility and the acentric factor, each reported against Joback, the published state of the art, and the experimental reproducibility floor of the property. Against Joback on identical held-out folds the graph model significantly wins on five of the nine comparable properties (boiling-point MAE 13.5 vs 26.3 K, formation enthalpy 35.5 vs 70.9 kJ/mol), ties two, and its error grows far more slowly with molecular size than Joback’s, which grows several-fold; it also covers the compounds Joback cannot fragment at all, a few percent of a typical set and up to about 12% for some properties, at undiminished accuracy on the ordinary organics among them; the two properties Joback still out-predicts, Gibbs energy of formation and critical pressure, are exactly the two the veto holds back as beta, with Joback served in their place. We release them as the Cheméo Relay family, one model per property.</jats:p>