Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p> The Tor (The Onion Router) anonymity network, while offering essential privacy protections for legitimate users, has increasingly been exploited by malicious actors to conceal Command and Control (C2) communication, coordinate ransomware, and host illegal dark web hidden services. Detecting malicious activity inside Tor faces three fundamental challenges: protocol-level packet homogenization due to Tor’s fixed 512-byte cell architecture, highly heterogeneous sample quality caused by encryption noise, and extreme class imbalance in wild network environments. To address these limitations, this paper proposes a three-stage, data-centric framework named <bold>MBRD-Tor (Multi-scale Behavioral Representation and Data-centric refinement for Tor)</bold> . In Stage 1, we extract multi-scale behavioral signatures by fusing temporal, positional, and directional cell embeddings, subsequently compressing them into compact latent representations using an auxiliary-classified autoencoder. In Stage 2, we introduce an ensemble quality-estimation mechanism using five complementary anomaly/density estimators to calculate class-conditional compatibility scores for adaptive instance weighting. To counter extreme class skew, we train three distinct Generative Adversarial Network (GAN) architectures per class to generate high-fidelity synthetic latent features, which are strictly validated by the five-model ensemble before being injected into the training pool. Stage 3 presents a classifier-agnostic protocol for downstream deployment. Evaluation of our framework on the CCS'22 Tor malware dataset shows that MBRD-Tor achieves an F1-score of 73.70% under the natural class distribution and preserves a recall of 21.1% under an extreme 1% malicious traffic ratio, outperforming state-of-the-art baselines. </p>

Show More

Keywords

class network malicious extreme stage

Related Articles

PORE

About

Connect