Abstract
<title>Abstract</title> <p>Alzheimer’s disease is a progressive neurological disorder that affects memory, cognition, language, judgement and daily functioning. Handwriting analysis has emerged as a promising non-invasive approach for detecting cognitive and motor alterations associated with the disease. However, studies using the same handwriting dataset have reported markedly different performance, raising concerns regarding overfitting, information leakage and validation design. This study developed a leakage-free machine-learning framework for Alzheimer’s disease classification using the DARWIN handwriting dataset, comprising 174 participants and 450 usable temporal, kinematic, pressure-related and geometric features. Logistic Regression, Support Vector Machine, Random Forest and XGBoost were evaluated using repeated nested stratified cross-validation. The outer procedure used five-fold cross-validation repeated ten times, producing 50 outer-test evaluations, while inner five-fold cross-validation was used for feature selection and hyperparameter optimisation. Imputation, standardisation, SelectKBest feature selection and optional synthetic minority oversampling were restricted to the corresponding training folds. Random Forest achieved the highest mean performance, with 88.11% accuracy, 88.10% balanced accuracy, an F1-score of 88.43%, a Matthews correlation coefficient of 0.7675 and a receiver operating characteristic area under the curve of 0.9637. XGBoost achieved the lowest Brier score of 0.0936. The findings indicate that handwriting-derived characteristics contain meaningful information for supporting Alzheimer’s disease classification under rigorous leakage-free evaluation. External validation using larger and demographically diverse handwriting datasets remains necessary.</p>