Skip to content

File 016996

Quantitative Analysis of Culture Using Millions of Digitized Books (File 016996)

A peer-reviewed academic paper describing the creation and analysis of a corpus of over 5 million digitized books to quantitatively investigate cultural and linguistic trends through 'culturomics', with applications spanning lexicography, grammar evolution, collective memory, and historical epidemiology.

Summary

This paper introduces 'culturomics,' a quantitative approach to studying culture using a corpus of approximately 5.2 million digitized books (~4% of all books ever published) containing over 500 billion words in multiple languages. The researchers developed computational methods to analyze n-gram frequencies across time periods from 1800-2000, revealing insights into linguistic and cultural phenomena including the growth of the English lexicon (from 544,000 words in 1900 to 1,022,000 in 2000), shifts in historical terminology, and trends in collective memory and social phenomena. The methodology demonstrates how large-scale computational analysis can extend scientific inquiry into humanities and social sciences, with findings presented through specific examples like the terminology shift from 'the Great War' to 'World War I/II' and the frequency peaks of 'slavery' during the Civil War and Civil Rights era.

Topics

People Mentioned