Abstract
<title>Abstract</title> <p>Public repositories hold over a million bioassays, yet most contain too few labelled compounds for conventional machine learning. Every assay, however, already carries a short text description stating what it measures. Most few-shot molecular meta-learning methods learn primarily from chemical structure and do not explicitly use assay descriptions. Here we show that encoding each assay description into a fixed text representation and reusing it at two decision points --- to select biologically related training tasks and to adjust how the model interprets molecular features for each assay --- consistently improves prediction. Combined with a module that captures scaffold and functional-group similarity between molecules, this approach, termed STAR, outperforms its text-blind baseline in all ten benchmark--shot combinations, with the largest gain (\(+6.85\) AUROC points) on the most label-scarce dataset. Improvement scales with the fraction of missing labels (Spearman \(\rho = 0.67\)), consistent with the text filling information gaps that molecular structure alone cannot. Systematic ablation of each component reveals that the text and structure channels carry partially overlapping chemical information. STAR requires no new model architecture; it shows that freely available assay descriptions are an effective prior for molecular prediction with limited data. \textbf{Scientific Contribution:} This work demonstrates that assay descriptions can function as explicit task-level priors for few-shot molecular property prediction, addressing a source of biological context that structure-centred meta-learning methods generally leave unused. STAR makes this information operational through two coordinated mechanisms based on the same fixed text representation: semantic auxiliary-task selection and assay-conditioned molecular feature modulation, while integrating scaffold and functional-group relations as a complementary structural channel. Factorial ablation shows that the text and structural channels provide distinct but partially redundant information, and the association between performance gain and label sparsity identifies sparse assay panels as the setting in which the language prior is most beneficial.</p>