Panning for gold: Lessons learned from the platform-agnostic automated detection of political content in textual data

by   Mykola Makhortykh, et al.

The growing availability of data about online information behaviour enables new possibilities for political communication research. However, the volume and variety of these data makes them difficult to analyse and prompts the need for developing automated content approaches relying on a broad range of natural language processing techniques (e.g. machine learning- or neural network-based ones). In this paper, we discuss how these techniques can be used to detect political content across different platforms. Using three validation datasets, which include a variety of political and non-political textual documents from online platforms, we systematically compare the performance of three groups of detection techniques relying on dictionaries, supervised machine learning, or neural networks. We also examine the impact of different modes of data preprocessing (e.g. stemming and stopword removal) on the low-cost implementations of these techniques using a large set (n = 66) of detection models. Our results show the limited impact of preprocessing on model performance, with the best results for less noisy data being achieved by neural network- and machine-learning-based models, in contrast to the more robust performance of dictionary-based models on noisy data.


page 1

page 2

page 3

page 4


On Detecting Policy-Related Political Ads: An Exploratory Analysis of Meta Ads in 2022 French Election

Online political advertising has become the cornerstone of political cam...

Image as Data: Automated Visual Content Analysis for Political Science

Image data provide unique information about political events, actors, an...

Small data problems in political research: a critical replication study

In an often-cited 2019 paper on the use of machine learning in political...

Machine Learning and Statistical Approaches to Measuring Similarity of Political Parties

Mapping political party systems to metric policy spaces is one of the ma...

A Face Preprocessing Approach for Improved DeepFake Detection

Recent advancements in content generation technologies (also widely know...

The Face of Populism: Examining Differences in Facial Emotional Expressions of Political Leaders Using Machine Learning

Online media has revolutionized the way political information is dissemi...

Please sign up or login with your details

Forgot password? Click here to reset