Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Rare-event anomaly detection in multivariate electricity consumption time series is affected not only by model architecture but also by the representation used before learning. This study presents a controlled benchmark that compares raw sequential modeling, Gramian Angular Field (GAF) image representation, hybrid fusion learning, and statistical tabular learning using the UCI Individual Household Electric Power Consumption dataset. Because this dataset does not contain naturally annotated fault events and represents a single household rather than a grid-scale operational system, this study should be interpreted as a reproducible representation-comparison benchmark rather than a validation on real grid-scale faults. In response to reviewer concerns, additional robustness checks were conducted for dataset coverage, window-labeling boundary behavior, XGBoost imbalance weighting, threshold dependence, bootstrap confidence intervals, and anomaly-type-specific false negatives. Controlled point, contextual, and collective perturbations were injected after chronological train-validation-test splitting to reduce information leakage, and overlapping sliding windows were labeled using an anomaly-density rule. Four model pipelines were evaluated under a common protocol: Raw 1D-CNN, GAF-CNN, Hybrid Fusion, and XGBoost. Under the F1-optimized validation threshold, the Raw 1D-CNN achieved the strongest test performance, with Accuracy 0.9528, Precision 0.9116, Recall 0.6471, F1-score 0.7569, ROC-AUC 0.8453, and PR-AUC 0.7500. The Hybrid Fusion model ranked second (F1-score 0.6908; PR-AUC 0.7157), followed by XGBoost and GAF-CNN. These results indicate that, under the specific synthetic anomaly protocol and window-labeling design used in this benchmark, direct temporal modeling preserved anomaly-discriminative information more effectively than standalone GAF encoding. However, the moderate recall of all models, the use of synthetic labels, the single-household data source, and the architecture differences between pipelines limit real-world generalization. The findings therefore support cautious empirical comparison of representation strategies rather than broad claims of operational grid-scale superiority.</p>

Show More

Keywords

than model representation learning benchmark

Related Articles

PORE

About

Connect