Abstract
<title>Abstract</title> <p>NAC transcription factors have a conserved N-terminal DNA-binding domain and a more variable C-terminal region. We asked whether that contrast could also be seen in protein-language-model embeddings and structure-prediction confidence. The analysis used five grass proteins selected in an earlier similarity screen: rice OsNAC25 and one candidate each from foxtail millet, sorghum, finger millet, and teff. We compared each full sequence, residues 1-180, and residues after 180 using ESM C embeddings. We also aligned the proteins with MAFFT and predicted their structures with ESMFold2. Mean pairwise sequence identity was higher in residues 1-180 than in the C-terminal segments (0.480 versus 0.346). The embedding result was less straightforward. C-terminal distance exceeded N-terminal distance in six of ten pairs, and the finger-millet candidate was a marked outlier in the full-length and N-terminal comparisons. Mean normalized ESMFold2 pLDDT was higher in residues 1-180 for all five proteins, although the difference was negligible for sorghum and teff. Thus, sequence identity and predicted confidence followed the expected regional pattern, whereas the embeddings were strongly affected by candidate identity. Because the non-rice proteins were selected by a one-way similarity screen, the analysis does not establish orthology or shared function. It instead identifies the finger-millet sequence as a priority for reciprocal searching and NAC-family phylogenetic analysis.</p>