Empirical rank-frequency regularity
Zipf's
Law
Rank the words in a large corpus from most to least frequent. Their frequencies often decline approximately as a power of rank: a small vocabulary forms the head, while an enormous number of uncommon words creates the long tail.
Turn a text into a ranked vocabulary.
Analyze the words in this guide, compare controlled counterexamples, or paste your own English text. All calculations run locally in the page.
The page will summarize the distribution after tokenization.
Frequency falls as rank rises.
Count every word token in a corpus, group identical tokens into word types, and order the types by decreasing frequency. The classical form says that frequency f at rank r is proportional to r raised to a negative exponent s. When s is near one, rank 2 has roughly half the frequency of rank 1, rank 10 roughly one tenth, and rank 100 roughly one hundredth.
This is an approximate scaling relation, not a word-by-word prediction. Natural corpora contain systematic bends and local structure beyond one exponent. Piantadosi's critical review concludes that large-scale word frequencies are robustly Zipfian while also being reliably more complex than the classic formula.[3]
Zipf's Law compresses the silhouette of a vocabulary. It does not erase grammar, meaning, genre, history, or the corpus that produced it.
Linear and logarithmic views tell different stories.
The head dominates.
A handful of function words towers over everything else. Most ranks collapse visually near zero.
Scaling becomes visible.
A pure power law becomes a straight line with slope -s. Deviations in the head and tail are easier to inspect.
Very common words may bend away from a single fitted line.
The rank-frequency relation often looks most regular over an intermediate span.
Counts become discrete, noisy, and dominated by words seen once or twice.
Before counting words, define what a word is.
Rank-frequency curves depend on editorial and computational choices. A result without its tokenization rules is not reproducible.
Language, speaker, genre, date, topic, medium, and corpus length shape the distribution.
Whitespace is insufficient for every language. Punctuation, apostrophes, compounds, emoji, and scripts need rules.
Decide case folding, spelling variants, numbers, markup, and whether contractions split.
Surface forms, lemmas, morphemes, characters, and subword tokens answer different questions.
Report ties, frequency units, fitting range, uncertainty, and alternative models.
Keep or remove?
Keeping them reveals the natural high-frequency head. Removing them may help topic analysis but changes the law being measured.
Merge inflections?
Combining runs, ran, and running changes ranks and can alter fitted parameters. It is a linguistic model, not neutral cleanup.
Tokens are learned.
Modern NLP vocabularies split words into pieces. Their frequency distribution is determined partly by the tokenizer's training objective.
Segmentation differs.
Chinese, Japanese, Thai, agglutinative languages, and mixed scripts make English-style word boundaries inappropriate.
Robust at large scale, structured in detail.
Zipf-like rank-frequency structure appears across natural languages and many corpus types, but real curves are not exact 1/r laws. Cross-language work reports recurring multi-segment patterns, while genre and preprocessing shift local slopes.[9]
A small set of words accounts for a large share of tokens, and most types are rare.
Frequency falls roughly as a power of rank over substantial regions of large corpora.
There is no single exact slope shared by every language, corpus, rank range, or token definition.
Function words, topic mixture, finite sample size, and vocabulary growth create systematic departures.
Every occurrence: "the" counted 8,000 times contributes 8,000 tokens.
Distinct forms: all 8,000 occurrences of "the" contribute one type.
Types seen once. They are evidence of the large tail and of incomplete vocabulary sampling.
Many mechanisms can produce a similar curve.
The rank-frequency pattern is not a unique fingerprint of one causal theory. Explanations operate at different levels and may be complementary rather than exclusive.
Least effort
Speakers favor reusable, accessible forms; hearers benefit from distinct, informative signals. Zipf proposed that language balances competing pressures, and later models formalized this tradeoff.[6]
Preferential reuse
Existing words are likely to recur while new words enter at a lower rate. Simon showed how such stochastic growth can generate highly skewed distributions.[5]
Random typing
Random character strings separated by spaces can create Zipf-like ranks. This proves that the shape alone does not demonstrate a realistic language mechanism.
Compression
Short, frequent forms and longer, rarer forms can arise when coding cost and information are jointly constrained.
Mixture of topics
Combining speakers, documents, genres, and contexts creates heterogeneity and can strengthen heavy-tailed frequency structure.
Learning and memory
Exposure, accessibility, prediction, production, and lexical innovation interact across development and use.
Fit, challenge, and compare the model.
Specify the rank range, token definition, corpus, and whether the exponent is fixed at one or estimated.
Plot counts, complementary cumulative distributions, and residuals. Preserve the discrete tail.
Least squares on logged values is biased. Use likelihood-based methods suited to discrete data.
Use simulated goodness-of-fit procedures and quantify uncertainty in the lower cutoff and exponent.
Evaluate lognormal, exponential, stretched-exponential, Yule-Simon, and truncated models using likelihood ratios.
Check sensitivity to genre, length, preprocessing, language, time period, and sampling unit.
A high R-squared after transformation can coexist with systematic curvature, correlated errors, discrete tail effects, and a better alternative model.
Clauset, Shalizi, and Newman provide a widely used framework combining maximum-likelihood estimation, Kolmogorov-Smirnov goodness-of-fit, simulation, and likelihood-ratio comparisons. They show why visual inspection and ordinary least squares are insufficient.[4]
From word counts to statistical language science.
Researchers including Jean-Baptiste Estoup documented ranked word-frequency patterns before Zipf's mature synthesis.
Zipf systematically analyzed relative word frequencies in language and connected rank to frequency.
Zipf placed linguistic distributions inside a broader theory of human behavior and competing effort.[1]
Information-theoretic and stochastic-growth models supplied distinct mathematical routes to skewed frequency distributions.
Massive multilingual corpora expose stable large-scale structure, detailed deviations, sampling effects, and stronger statistical tests.
The head and tail demand different engineering strategies.
Frequent and rare queries
Cache and optimize the head, but design retrieval, spelling, and fallback systems for the enormous query tail.
Imbalanced training data
Common tokens receive abundant updates while rare words, names, dialects, and specialist terms remain data poor.
Short codes for common events
Frequency-skewed symbols motivate variable-length coding, though linguistic structure is richer than isolated token counts.
Coverage and vocabulary growth
More text keeps revealing new types. Corpus size must accompany any claim about vocabulary coverage.
Head-tail allocation
A few actions dominate volume while rare workflows create support, accessibility, and reliability challenges.
Do not average away the tail
Aggregate accuracy can look strong while performance fails on rare terms, languages, intents, or users.
Sources and further reading.
Original books, foundational models, critical reviews, and rigorous statistical methods are prioritized.
- George Kingsley Zipf (1949) - Human Behavior and the Principle of Least EffortZipf's mature synthesis connecting linguistic frequency to a broader least-effort principle.Public-domain scan via Wikimedia Commons
- George Kingsley Zipf (1935) - The Psycho-Biology of LanguageAn early systematic presentation of relative frequency patterns in language.Internet Archive catalog
- Steven T. Piantadosi (2014) - Zipf's Word Frequency Law in Natural LanguageA critical review emphasizing robust large-scale structure and reliable deviations beyond the classical law.doi.org/10.3758/s13423-014-0585-6
- Clauset, Shalizi, and Newman (2009) - Power-Law Distributions in Empirical DataLikelihood-based fitting, goodness-of-fit testing, and comparison against alternative distributions.SIAM Review 51(4), 661-703
- Herbert A. Simon (1955) - On a Class of Skew Distribution FunctionsA foundational stochastic-growth explanation for highly skewed empirical distributions.Biometrika 42(3/4), 425-440
- Ferrer i Cancho and Sole (2003) - Least Effort and the Origins of Scaling in Human LanguageA formal communicative tradeoff model producing Zipf-like scaling.PNAS 100(3), 788-791
- Benoit Mandelbrot (1953) - An Informational Theory of the Statistical Structure of LanguageAn information-theoretic account extending the classical rank-frequency relation.Communication Theory, 486-502
- Altmann and Gerlach (2016) - Statistical Laws in LinguisticsA critical discussion of fitting, falsification, correlations, and fluctuations in linguistic laws.arXiv:1502.03296
- Yu, Xu, and Liu (2018) - Zipf's Law in 50 LanguagesCross-language evidence for recurring multi-segment structure and tail deviations.arXiv:1807.01855
- Petersen et al. (2012) - Languages Cool as They ExpandLarge-corpus analysis linking vocabulary growth, word birth rates, and language expansion.Scientific Reports 2, 943