Abstract
<p>This study examined the effectiveness of different counterspeech techniques and the influence of counterspeech source on reducing perceived harmfulness of harmful online comments. In a preregistered experiment (N = 272), participants viewed harmful comments about a controversial topic paired with counterspeech varying in technique (Denouncing, Appealing to Compassion, Presenting Facts, Redirecting) and source (Moderator, Bot, Community Note). Perceived harmfulness was measured using the newly developed Perceived Harmfulness Scale (PHS), encompassing Emotional Distress, Feeling of Threat, Social Exclusion, and Discourse Disruption. Three of four techniques – Appealing to Compassion, Presenting Facts, and Redirecting – demonstrated equal effectiveness in reducing perceived harmfulness of a harmful comment, while Denouncing was ineffective. At the aggregate level, counterspeech source had no main effect. However, exploratory analyses revealed a crossover interaction between source and participant gender: bot-based and moderator-based counterspeech reduced perceived harmfulness only for male participants, while for female participants only community-based counterspeech was effective. For women, bot-based counterspeech was even counterproductive. These findings suggest that attention management may be more effective than norm activation in reducing perceived harmfulness, and that one-size-fits-all approaches to content moderation may inadvertently fail specific user groups. Platforms may benefit from diversified moderation strategies combining automated detection with community-based responses to address diverse user needs.</p>