Abstract
<p>This study made use of distributional semantics to explore the semantic structure of two-character compounds in Mandarin Chinese. A total of 2,843 compounds, extracted from the Chinese National Corpus, informed a series of exploratory analyses. For larger-scale evaluation, we extracted all 29,376 two-character compounds from the Chinese Lexical Database (which is based on SUBTLEX-CH and Leiden Weibo Corpus) for which embeddings are available in the Tencent AI Lab resource. We observed that the compound families of mono\-morph\-emic monosyllabic Chinese words (pivots) cluster in the embedding space. The quality of these clusters does not depend on the ontological class of the compounds and is as good, and often better, than the quality of the clusters of the embeddings of derived words sharing the same suffix. The compound families of noun pivots are represented more prominently in the early principal components of the PCA-orthogonalized embedding space compared to the compound families of verb and adjective pivots. Compound families also cluster in the space of shift vectors, indicating that individual pivots contribute a core meaning to the compounds in which they occur. Approximately half of all pivots show no positional preference in the compound, and for 83.6\% of Mandarin two-character compounds, the position of the pivot cannot be predicted from the compounds' meanings. Therefore, pivot position cannot explain the clustering in semantic space of pivot families. Furthermore, pivots cannot be classified as prefixes, suffixes, interfixes, or circumfixes. Not being bound to a specific position for their interpretation, pivots are best characterized as ``floatfixes'', and instantiate morphological non-configurationality.</p>