Abstract
<p>Political scientists now regularly use audio, video, and text data to investigate questions about deliberation, representation, and emotion. However, existing research using multimodal data focuses largely on highly professionalized with readily available preprocessed data. Although data from less standardized environments such as campaign events, court hearings, and local government meetings has increased, using recordings of these events for research poses common measurement challenges, including creating transcripts, identifying speakers, and idiosyncratic production styles. In this paper, we present a streamlined and tested pipeline that uses open source tools to automatically transform raw recording into data formats needed for political science research. The outputs of our pipeline can then be used for a wide range of substantive analyses. We validate our approach through an examination of thousands of hours of local government meetings and state supreme courts hearings, showing how we can accurately segment audio, identify speakers, transcribe speech, and detect gender across varying structures and audio/video quality. The flexible open-source pipeline is publicly available and supports both local machines and computing clusters. As an application, we examine participation by gender in over 1,000 hours of school board meeting videos and US state supreme court hearings.</p>