Sixty million lines of transcribed speech represent a massive shift in how we can study American religious life. Until now, researchers have lacked the large-scale, content-level data needed to analyze religious radio systematically. Most studies relied on small samples or listener surveys, leaving the actual substance of broadcasts largely invisible.
The Pew Research Center has addressed this gap by releasing the Religious Radio Corpus. This dataset provides a granular view of the medium by combining automated audio capture with advanced machine learning. By turning raw airwaves into searchable, annotated text, the researchers allow us to study religious broadcasting as a national phenomenon.
The scarcity of longitudinal broadcast data
Systematic analysis of religious media has long been hindered by a lack of scalable data. Scholars have historically relied on "close reading"—the intensive qualitative analysis of a few select programs. This makes it difficult to perform "distant reading," where computational methods analyze thousands of hours of content to find broad patterns.
Existing speech corpora often prioritize news or general talk radio. This leaves the unique linguistic styles of religious programming unexamined. Without a large-scale corpus, it is difficult to answer basic questions. For example, how does the ratio of music to sermons vary across different U.S. regions?
Previous efforts, such as the RadioTalk corpus, focused on general talk radio. They did not capture the specific liturgical or devotional nuances found in religious media. This new corpus fills that void by focusing specifically on the religious broadcasting landscape.
An automated pipeline from airwaves to annotations
The authors built an end-to-end pipeline to transform live webstreams into enriched transcripts .
The process involves four distinct stages:
- Massive Scale Capture: A cluster of 250 containerized applications monitored 785 unique webstream URLs. The team captured 15-minute segments throughout July 2025. This resulted in approximately 172,000 hours of raw audio.
- Signal Partitioning: An Audio Spectrogram Transformer (AST)—a model that analyzes the visual representation of sound—distinguishes between speech and music. This ensures the transcription engine does not attempt to "read" musical interludes.
- Transcription and Diarization: The system uses the WhisperX pipeline. This combines OpenAI’s
whisper-large-v3-turbofor Automatic Speech Recognition (ASR) withpyannote.audiofor speaker diarization (identifying who is speaking). This produces transcripts where every line is tied to an anonymized speaker ID. - LLM Enrichment: Finally, the text is processed by the GPT-4.1 API. The model segments the transcript into topical blocks. It also assigns labels for programming formats (like interviews) and broad topics (like politics).
Measuring the precision of the automated ear
The utility of a corpus depends on the reliability of its transcriptions. To validate the pipeline, the authors compared the automated output against human-coded reference transcripts. The paper reports an overall Word Error Rate (WER) of 5.14%. This metric measures the percentage of words that were substituted, deleted, or incorrectly inserted.
This 5.14% error rate is quite close to the 2.57% error rate seen when two humans transcribe the same audio. This suggests the automated system is approaching human-level consistency. However, accuracy is not uniform across all content.
Accuracy is highest for scripted content, such as religious services. Conversely, accuracy drops in "challenging" acoustic environments. For instance, caller interactions have a higher WER of 7.55%. Additionally, LLM-based topic labels vary in precision. The model was very accurate at identifying "religion" (F1 score of 0.93) but less reliable for "lifestyle/advice" (F1 score of 0.46).
Limitations in identity and language
Users must navigate several structural limitations. First, speaker diarization is "local" to each 15-minute recording. The system cannot track the same person across different days or recordings. This limits studies of specific recurring personalities.
Second, the corpus is restricted to English-language programming. Spanish Christian stations were excluded because the pipeline was not optimized for non-English audio. Finally, LLM-generated labels should be treated as approximate. They are useful for finding aggregate trends but may be less precise for individual segments.
A foundation for digital sociology
The Religious Radio Corpus provides the infrastructure to treat religious broadcasting as a measurable social phenomenon. For engineers, it offers a high-volume dataset for domain adaptation research. This is particularly useful for improving models in noisy, multi-speaker, and specialized oratorical environments.
The work is reproducible. Code for the collection and processing pipelines is available at https://github.com/pewresearch/religious-radio-corpus. The data is accessible through two channels. An archival version is in the Harvard Dataverse. An access-optimized version is available via Hugging Face Datasets for high-speed computational streaming.
How this was made
Model: nvidia/Gemma-4-26B-A4B-NVFP4
Persona: science_essayist
Template: engineering_deepdive
Refinement: 0
Pipeline: forge-1.1
Evaluator: nvidia/Gemma-4-26B-A4B-NVFP4
Score: 94% (passed)
Claims verified: 19 / 19
Model: nvidia/Gemma-4-26B-A4B-NVFP4
NVIDIA GB10 · 128 GB unified · NVFP4 · 100% local · $0 cloud
Tokens: 80,836
Wall-time: 213.3s
Tokens/s: 378.9