Abstract
<title>Abstract</title> <p>Automated pharmacokinetic model building requires predefined criteria for model evaluation to rank candidate models and guide model selection. Evaluation criteria that rely on a single metric or are excessively stringent may bias the model toward overfitting or toward an oversimplified structure. Therefore, an appropriate fitness function is needed to enable comprehensive evaluation across multiple dimensions while balancing potential trade-offs among these criteria. This study aimed to systematically assess the model evaluation metrics commonly used in population pharmacokinetic (PopPK) modeling and to identify a fitness function for algorithm-driven model development. Binary and step penalty designs were compared for model selection performance. To this end, 48 datasets were simulated based on optimal design, and an exhaustive search approach was used to test all candidate models in the predefined model space. This allowed the original model structure captured in the simulated datasets to serve as the reference, and enabled assessment of the function's ability to recover that structure within a complete search space. The results indicate that the relative standard error (RSE) of parameter estimates, eta-shrinkage, and the omega values of inter-individual variability are key evaluation metrics for accurately identifying the true model and the step penalty was more robust than the binary penalty in selecting model structure. Based on these findings, this study proposes a fitness function combining BIC with threshold-based step penalties. The fitness function could play a central role in automated PopPK model building and provide insights for conventional model selection.</p>