https://norvig.com/ngrams/ [cat] O'Reilly / Amazon / Google Books Natural Language Corpus Data: Beautiful Data This directory contains code and data to accompany the chapter Natural Language Corpus Data from the book Beautiful Data (Segaran and Hammerbacher, 2009). If you like this you may also like: How to Write a Spelling Corrector. Data files are derived from the Google Web Trillion Word Corpus, as described by Thorsten Brants and Alex Franz, and distributed by the Linguistic Data Consortium. Code copyright (c) 2008-2009 by Peter Norvig. You are free to use this code under the MIT license. To run this code, download the files listed below. Then from a shell execute python -i ngrams.py (or start a Python IDE and import ngrams), and if you want to test if everything works, call test(). Note that the hillclimbing function has a random component, so if you have bad luck it is possible that some of the tests will fail, even if everything is correctly installed. (It is unlikely that they will fail twice in a row.) Files for Download 0.7MB ch14.pdf The chapter from the book. 0.0 ngrams.py The Python code for everything in the MB chapter. 0.0 ngrams-test.txt Unit tests; run by the Python function MB test(). The 1/3 million most frequent words, all 4.9 count_1w.txt lowercase, with counts. (Called MB vocab_common in the chapter, but I changed file names here.) 5.6 count_2w.txt The 1/4 million most frequent two-word MB (lowercase) bigrams, with counts. 0.0 count_2l.txt Counts for all 2-letter (lowercase) MB bigrams. 0.2 count_3l.txt Counts for all 3-letter (lowercase) MB trigrams. 0.0 Counts for all single-edit spelling MB count_1edit.txt correction edits, from the file spell-errors.txt. 0.5 A collection of "right: wrong1, wrong2" MB spell-errors.txt spelling mistakes, collected from Wikipedia and Roger Mitton. The following files are not referenced in the chapter, but may be useful to you. 6.5 big.txt File of running text used in my spell MB correction article. 1.0 Excerpt of file of running text from my MB smaller.txt spell correction article. Smaller; faster to download. 0.3 count_big.txt A word count file (29,136 words) for MB big.txt. 1.5 count_1w100k.txt A word count file with 100,000 most MB popular words, all uppercase. .02 words4.txt 4360 words of length 4 (for word games) MB .04 sgb-words.txt 5757 words of length 5 (for word games) MB from Knuth's Stanford GraphBase 1.1 wordlist.asc Tom Murphy's word list for portmantout MB words. .03 1000 most common words of English from MB words.js xkcd Simple Writer (more than 1,000 words because plurals are included) The complete works of Shakespeare, 4.3 shakespeare.txt tokenized so that there is a space MB between words and punctuation. From John DeNero. 4.5 shakespeare_input.txt Untokenized Shakespeare, from Andrej MB Karpathy. 6.2 linux_input.txt Untokenized Linux Kernel C++ code, from MB Andrej Karpathy. 3.0 The SOWPODS word list (267,750 words) -- MB sowpods.txt used by Scrabble players (except in North America) and in other word games. 1.9 The Tournament Word List (178,690 words) MB TWL06.txt -- used by North American Scrabble players. 1.9 The ENABLE word list (172,819 words) -- MB enable1.txt also used by word game players. Words with Friends uses a variant of this. 2.7 The YAWL (Yet Another Word List) word MB word.list list (263,533 words) -- formed by combining the above. . (See Internet Scrabble Club for more lists.) --------------------------------------------------------------------- Peter Norvig, 8 July 2008; updated 22 Nov 2011