Abstract
<jats:p>The rapidly increasing number of video tracking-based behavioral summary tools and methods raises the question as to the most suitable approaches for pharmacological fingerprinting in pre-clinical research. We have recently shown that social context has a strong effect on behavioral syntax in mice, further suggesting that treatment effects can be context dependent. Here, we aim to answer the question whether there is an optimal combination of context and behavioral summary method for effect detection and discrimination of different psychoactive substances in a controlled environment. To this end, we applied eight different treatment-dosage pairs (amphetamine 1.5,3,6 mg/kg; modafinil 5,10,50 mg/kg; seltorexant 3,10 mg/kg) and evaluated five different approaches to behavioral summary: parametric aggregation, unsupervised segmentation in Keypoint-MoSeq (KPMS) & Variational Animal Motion Encoding (VAME), and supervised segmentation in Simple Behavioral Analysis (SimBA) & A-SOiD, across two different contexts (Solitary & Social) in 314 recordings of freely moving mice in an open-field arena. Surprisingly, our results show no significant differences in performance across models and context. Across treatment effect detection to treatment-dose discrimination, all models showed performance significantly above chance that was insensitive to various data limitations and extensions. Overall, our study shows that under the tested conditions and treatments, the choice of a behavioral summary model does not meaningfully affect the description of treatment effects. Simple aggregate measures from tracking data and machine learning based behavioral summary approaches that are expensive, in terms of training data and computational resources, performed equally well. Our findings taken together with literature suggest that the fuller decomposition of complex behavior through unsupervised machine learning might be necessary for the description of large-scale datasets but does not necessarily align with the goals present in smaller-scale treatment discrimination tasks common in pre-clinical research.</jats:p>