Abstract
<jats:p> Molecular generation underpins discovery in drug design, organic electronics, and energy storage, but specialized chemistries are typically supported by datasets on the order of 10 <jats:sup>4</jats:sup> -10 <jats:sup>5</jats:sup> molecules which is too small for many state-of-the-art generative models, especially under property conditioning. We present HiMoDiT, a hierarchical discrete-diffusion model for generic molecular geneartion. HiMoDiT factors generation into two stages with a deterministic ring-layout decoder: Stage 1 generates a discrete ring layout (ring types, fusion adjacency, linker chains, pendant chains), which is deterministically decoded into a bond-class skeleton and aromaticity mask before atom identities are assigned under aromaticity constraints; Stage 2 grafts terminal fragments from a K-class functional-group vocabulary. By treating bond classes as a deterministic readout of ring layout rather than a learned per-edge prediction, HiMoDiT eliminates the topological failure modes that limit per-edge discrete diffusion, including aromatic 4-cycles, [6+4] split topologies, aromatic–aliphatic inconsistencies, by construction rather than via sampler biases or post-hoc repairs. On RedDB, the dataset of redox flow battery electrolyte, HiMoDiT achieves achieves a Validity-Uniqueness-Novelty (V·U·N) = 76.30% under matched-protocol benchmarking against four baselines fine-tuned on RedDB: HiMoDiT leads three iterative graph-diffusion methods (DiGress, Cometh, DeFoG) by 25.8%–36.8%, and outperforms MoLeR as the chemistry-aware motif-extension generator pretrained on 106 organic molecules, while uniquely supporting classifier-free-guidance conditioning on pseudoscore defined using solubility and HOMO–LUMO gap. Extended to ZINC250K with enlarged atom and terminal vocabularies, HiMoDiT reaches V·U·N = 98.6% with simultaneous logP (Pearson r = 0.730) and SAS (r = 0.845) conditioning. </jats:p>