The Largest Study of Black English Ever (227 Million Tweets) @languagejones
The Largest Study of Black English Ever (227 Million Tweets)  @languagejones
Uploaded June 2026 | Updated September 2026, 2 weeks ago
Get started actually speaking a language today, with Speak: bit.ly/4f5hUyF

We used AI to analyze 227 million tweets to answer questions about African American Language (AAL) that linguists have been arguing about for over fifty years.

👉 Get updates on my upcoming coauthored book, A Descriptive Grammar of Black English: languagejones.com/agobe

About This Video
For decades, mainstream linguistics operated under the "uniformity myth"—the idea that Black English (often called AAVE or Black English) is essentially identical across the United States, spoken primarily by young, working-class men in inner cities. But early sociolinguistic studies relied on tiny sample sizes and missed critical local context.

In this video, I walk through our newly accepted paper in the journal American Speech. Alongside lead author Tessa Masis, Lisa Green, and our research team, we leveraged computational sociolinguistics and a language model (BERT) to analyze the Twitter4Years dataset—spanning 227 million geotagged tweets.

By looking at 18 distinct grammatical features (like habitual be, zero copula, finna, and negative concord) across individual Census tracts, we developed a new metric called AALScore to treat AAL as a holistic, rule-governed constructional system.

What We Found:
• The Heart of AAL: The highest dialect density scores light up across the rural South—not just major urban hubs like New York, Detroit, or Atlanta.
• Regional Toolkits: Features like habitual be dominate the Northeast and Midwest, while features like resultative done cluster heavily in the rural South, debunking the uniformity myth.
• Language Contact: We uncovered surprisingly strong correlations with AAL features in Mexican American communities, opening up massive new questions for sociolinguistic research.

If you love language variation, data science, or dialectology, hit subscribe!
đź•’ Timestamps
0:00 - Intro: Mapping AAL with AI
0:47 – Sponsored segment from Speak: bit.ly/4f5hUyF
3:07 - The Uniformity Myth & Where Early Linguistics Went Wrong
5:18 - The Data: 227 Million Tweets (Twitter4Years)
7:12 - The Method: Training BERT to Find Grammar
8:38 - What is AALScore?
9:04 - Finding 1: The Rural South Connection
10:45 - Finding 2: Regional Variation (Habitual Be vs. Done)
11:45 - Finding 3: Mexican American Communities & Language Contact
13:30 - Why This Matters for the Future of Linguistics
15:10 - Outro & How to Read the Paper

Links & Resources
• Paper Citation: Masis, Tessa, et al. "Variation across Regions and Demographics in African American Language Morphosyntax: Evidence from Large-Scale Twitter Data." American Speech 101.2 (2026): 144-178.
• Read it here: https://read.dukeupress.edu/american-speech/article-abstract/101/2/144/410645/Variation-across-Regions-and-Demographics-in
• Sign up for Book Updates: languagejones.com/agobe
• Sign up for the Intro to Linguistics: languagejones.com/blueprint
• Further Reading: Introduction to African American English by Lisa Green. (amazon affiliate link: amzn.to/3SnW28b)

#linguistics #AAVE #machinelearning #sociolinguistics #bigdata #blackenglish #language
The Largest Study of Black English Ever (227 Million Tweets)Why is it ME GUSTA and MI PIACE? The linguistics of dative experiencersLearn Hebrew with me 6Lingopie: The Complete MethodIs the Odyssey Translation Behind Nolans Film Actually Woke?How to Actually Use AI for Language Learning (LLMs The Expert Way)The ULTIMATE guide to MASTERING the subjunctiveLearn with a linguist (Persian study livestream 8)Why do French Questions look so STRANGE? A linguist explains whats REALLY going on
languagejones |

The Largest Study of Black English Ever (227 Million Tweets)

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER