Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>The rapid proliferation of AI-generated synthetic content on social networks necessitates a robust approach for detecting forgery. Current unimodal techniques, which use only visual or metadata information, are not resilient to new-generation models. This study describes a new tri-modal approach for modeling data using multiple modalities: (1) spatial features from an image using a convolutional neural network (CNN), (2) ten forensic metadata features using a Vision Transformer, and (3) relational features using graph neural networks (GNNs). An adaptive cross-modal attention-gating mechanism assigns a weight to each modality for each sample. Our system was trained on the SID-Set (a dataset of 24,000 balanced social-media images) with a 99.10% accuracy, 0.9998 AUC-ROC, 0.9911 F1 score, and 0.9999 average precision across out-of-sample evaluation data. GradCAM and SHAP help provide interpretability by showing where we focus on forensically significant areas. We also provide evidence of the contributions of each of the three modalities and that the adaptive gating approach is superior to having a static weighting for each modality. Future work will focus on developing techniques for cross-dataset generalization.</p>

Show More

Keywords

using each approach features networks

Related Articles

PORE

About

Connect