Abstract
<jats:p>*Corresponding author: yangxiao_ztri@foxmail.com Abstract. Odor prediction has relied on two implicit assumptions: that a closed set of expert-defined labels is necessary for training, and that individual molecules must serve as the encoding unit—forcing mixtures to be decomposed into constituents, predicted separately, then aggregated through ad hoc fusion. Here we show that both assumptions can be eliminated simultaneously. We introduce AtomFormer, which uses GIN layers as a molecular tokenizer and Transformer self-attention to enable atoms from different molecules to interact directly in a single forward pass, eliminating the need for separate mixture aggregation. On a rigorously decontaminated test set of 6,260 binary mixtures, AtomFormer achieves macro-AUROC = 0.9347 with 16.9M parameters. Beyond odor perception, we validate this architecture on predicting excess enthalpy of mixtures, where it achieves Pearson R = 0.937 with 1.18M parameters. Across these two distinct tasks, AtomFormer achieves strong performance with zero pretraining and end-to-end prediction, demonstrating the generality of the framework across diverse mixture prediction domains.</jats:p>