{"id":282,"date":"2010-12-17T16:48:01","date_gmt":"2010-12-17T16:48:01","guid":{"rendered":"http:\/\/anterotesis.com\/wordpress\/?p=282"},"modified":"2011-05-13T12:53:16","modified_gmt":"2011-05-13T11:53:16","slug":"google-ngram-games","status":"publish","type":"post","link":"https:\/\/anterotesis.com\/wordpress\/2010\/12\/google-ngram-games\/","title":{"rendered":"Google Ngram Games"},"content":{"rendered":"<p>Google have just opened up their text mining project, a vast and ambitious project to allow searching their digital library for the frequency of words and phrases. It&#8217;s an astonishing resource, not only for its research potential but also for its ludic possibilities, not to mention the time-frittering capabilities.<\/p>\n<p>It&#8217;s easy to play. Just go to <a title=\"Google Ngrams home\" href=\"http:\/\/ngrams.googlelabs.com\/\" target=\"_blank\">http:\/\/ngrams.googlelabs.com\/<\/a>, put in some words or phrases, separating multiples with commas, adjust the settings as you wish, and press the button. Up comes a graph showing distribution across your chosen time period. Thrill to the peaks! Gasp at the troughs! Wonder at the abundances! Curse the absences!<\/p>\n<p>Like all good games, there&#8217;s much wisdom to be gained from playing. For one thing, it tests the sources, both the original works and their translation into digital format. (Some of Google&#8217;s metadata is <a title=\"Google Books: A Metadata Train Wreck\" href=\"http:\/\/languagelog.ldc.upenn.edu\/nll\/?p=1701\" target=\"_blank\">bizarrely inaccurate<\/a>.) It makes one think about possible reasons: whatever the results of a query are, they never explain <em>why<\/em>. It questions the technologies of language, whether printed or digital, including orthography, typeface and grammar. (Note that it only covers\u00a0 written language, not that spoken or sung, with rhythm, accents and inflections.) In all, it indicates the ambiguity and instability of language, as against any needs or claims of clarity and transparency of meaning.<\/p>\n<p>Below are four jests: testing a grandiose claim, a revealing anachronism, typographical obscenity and the deleterious effects of popular culture on the Queen&#8217;s English.<\/p>\n<h3>Benjamin was right!<\/h3>\n<p>By searching on four major Western cities, we can see that <a title=\"Wikipedia: Walter Benjamin\" href=\"http:\/\/en.wikipedia.org\/wiki\/Walter_Benjamin\" target=\"_blank\">Walter Benjamin<\/a> was right to consider Paris &#8216;<a title=\"Benjamin, Paris Capital of the 19th Century [PDF]\" href=\"www.casbarcelona.org\/BenjaminParis.pdf\" target=\"_blank\">the capital of the nineteenth century<\/a>&#8216; [pdf].<\/p>\n<div id=\"attachment_293\" style=\"width: 910px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/cities19thcentury.png\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-293\" class=\"size-full wp-image-293\" title=\"19th century cities\" src=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/cities19thcentury.png\" alt=\"Google Ngram for Paris, London, Berlin, New York, 1800-1900\" width=\"900\" height=\"330\" srcset=\"https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/cities19thcentury.png 900w, https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/cities19thcentury-300x110.png 300w\" sizes=\"auto, (max-width: 900px) 100vw, 900px\" \/><\/a><p id=\"caption-attachment-293\" class=\"wp-caption-text\">Google Ngram for Paris, London, Berlin, New York, 1800-1900<\/p><\/div>\n<p><a title=\"Google Ngram for 4 c19th cities\" href=\"http:\/\/ngrams.googlelabs.com\/graph?content=paris%2Clondon%2Cberlin%2Cnew+york&amp;year_start=1800&amp;year_end=1900&amp;corpus=0&amp;smoothing=3\" target=\"_blank\">Link<\/a><\/p>\n<p>And never mind population, trade, dominions and suchlike. But note that had I included &#8216;Rome&#8217;, Benjamin would have been refuted by the number of works on ancient history.<\/p>\n<h3>Surrealism: A Victorian Creation?<\/h3>\n<p>Although the word &#8216;surrealism&#8217; was originally coined by Apollinaire in 1917, and given substance in the 1920s by Andre Breton, Google finds it in mid-Victorian English:<\/p>\n<div id=\"attachment_289\" style=\"width: 910px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/surrealismngram.png\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-289\" class=\"size-full wp-image-289\" title=\"Surrealism Ngram\" src=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/surrealismngram.png\" alt=\"Google Ngram for 'surrealism.'\" width=\"900\" height=\"330\" srcset=\"https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/surrealismngram.png 900w, https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/surrealismngram-300x110.png 300w\" sizes=\"auto, (max-width: 900px) 100vw, 900px\" \/><\/a><p id=\"caption-attachment-289\" class=\"wp-caption-text\">Google Ngram for &#39;surrealism.&#39;<\/p><\/div>\n<p><a title=\"Google Ngram for 'surrealism'\" href=\"http:\/\/ngrams.googlelabs.com\/graph?content=surrealism&amp;year_start=1840&amp;year_end=2000&amp;corpus=0&amp;smoothing=3\" target=\"_blank\">Link<\/a><\/p>\n<p>There&#8217;s an interesting error behind this. In parsing <a title=\"Surrealism in 1860!\" href=\"http:\/\/books.google.com\/books?id=TFxFAAAAYAAJ&amp;pg=PA584&amp;dq=%22surrealism%22&amp;hl=en&amp;ei=ZokLTZ-rLaSJ4gaKl8DADA&amp;sa=X&amp;oi=book_result&amp;ct=result&amp;resnum=5&amp;ved=0CDoQ6AEwBA#v=onepage&amp;q=%22surrealism%22&amp;f=false\" target=\"_blank\">The London Review<\/a>, volume 15, 1860-1, the last hypenated word of one page, and the first of the next, the chapter heading rather than the continuation of the text proper, have been run together by Google&#8217;s OCR software. Such bugs, the &#8216;<a title=\"Revealing Errors blog\" href=\"http:\/\/revealingerrors.com\/\">revealing errors<\/a>&#8216;\u00a0 of logic, can be considered a manifestation of the surrealist spirit.<\/p>\n<h3>FFS!<\/h3>\n<p>Early modern printing is notorious for using &#8216;f&#8217; in place of &#8216;s.&#8217; Oh the comedic potential, as is born out by this graph:<\/p>\n<div id=\"attachment_291\" style=\"width: 910px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/fuckngram.png\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-291\" class=\"size-full wp-image-291\" title=\"Fuck Ngram\" src=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/fuckngram.png\" alt=\"Google Ngram for 'fuck.'\" width=\"900\" height=\"330\" srcset=\"https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/fuckngram.png 900w, https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/fuckngram-300x110.png 300w\" sizes=\"auto, (max-width: 900px) 100vw, 900px\" \/><\/a><p id=\"caption-attachment-291\" class=\"wp-caption-text\">Google Ngram for &#39;fuck.&#39;<\/p><\/div>\n<p><a title=\"Google Ngram for 'fuck'\" href=\"http:\/\/ngrams.googlelabs.com\/graph?content=fuck&amp;year_start=1620&amp;year_end=2000&amp;corpus=0&amp;smoothing=3\" target=\"_blank\">Link<\/a><\/p>\n<p>The alternative explanation is, of course, that we were a foulmouthed bunch until the Victorians, and only since the 1950s have we begun to throw off those moral shackles.<\/p>\n<h3>Star Trek and the split infinitive<\/h3>\n<p>That most controversial of grammatical issues had a historical turn in the 1980s, when &#8220;to boldly go&#8221; overtook &#8220;to go boldly.&#8221; A way of interrogating the corpus for the frequency of split infinitives isn&#8217;t obvious, but the results would be very interesting.<\/p>\n<div id=\"attachment_286\" style=\"width: 910px\" class=\"wp-caption aligncenter\"><a href=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/startrekngram.png\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-286\" class=\"size-full wp-image-286\" title=\"The Star Trek Ngram\" src=\"http:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/startrekngram.png\" alt=\"Google Ngram: 'to go boldly' and 'to boldly go.'\" width=\"900\" height=\"330\" srcset=\"https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/startrekngram.png 900w, https:\/\/anterotesis.com\/wordpress\/wp-content\/uploads\/2010\/12\/startrekngram-300x110.png 300w\" sizes=\"auto, (max-width: 900px) 100vw, 900px\" \/><\/a><p id=\"caption-attachment-286\" class=\"wp-caption-text\">Google Ngram: &#39;to go boldly&#39; and &#39;to boldly go.&#39;<\/p><\/div>\n<p><a title=\"Google Ngram for Star Trek's split infinitive\" href=\"http:\/\/ngrams.googlelabs.com\/graph?content=to+boldly+go%2Cto+go+boldly&amp;year_start=1800&amp;year_end=2000&amp;corpus=0&amp;smoothing=3\" target=\"_blank\">Link.<\/a> <a title=\"Wikipedia: Split Infinitive\" href=\"http:\/\/en.wikipedia.org\/wiki\/Split_infinitive\" target=\"_blank\">Wikipedia on Split Infinitives<\/a><\/p>\n<p>Google have provided some basic, but literate, <a title=\"Google documentation for their Ngrams\" href=\"http:\/\/ngrams.googlelabs.com\/info\" target=\"_blank\">documentation<\/a>. And the <a title=\"Google Ngram datasets\" href=\"http:\/\/ngrams.googlelabs.com\/datasets\" target=\"_blank\">datasets<\/a> are freely available under a creative commons license. A <a title=\"Guardian on Google data-mining\" href=\"http:\/\/www.guardian.co.uk\/science\/2010\/dec\/16\/google-tool-english-cultural-trends\" target=\"_blank\">Guardian article<\/a> serves as a decent introduction, although it exaggerates the originality of the techniques. See also the <a title=\"New York Times on Googles ngrams\" href=\"http:\/\/www.nytimes.com\/2010\/12\/17\/books\/17words.html\" target=\"_blank\">New York Times<\/a>. And don&#8217;t ever, <em>ever<\/em>, use the pseudo-word &#8216;<a title=\"Interesting site, horrible name\" href=\"http:\/\/www.culturomics.org\/\" target=\"_blank\">culturomics<\/a>.&#8217;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Google have just opened up their text mining project, a vast and ambitious project to allow searching their digital library for the frequency of words and phrases. It&#8217;s an astonishing resource, not only for its research potential but also for &hellip; <a href=\"https:\/\/anterotesis.com\/wordpress\/2010\/12\/google-ngram-games\/\">Continue reading <span class=\"meta-nav\">&rarr;<\/span><\/a><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_publicize_message":"","jetpack_publicize_feature_enabled":true,"jetpack_social_post_already_shared":false,"jetpack_social_options":{"image_generator_settings":{"template":"highway","default_image_id":0,"font":"","enabled":false},"version":2},"jetpack_post_was_ever_published":false},"categories":[46],"tags":[31,137,66,30,74,86,77,76,78,75],"class_list":["post-282","post","type-post","status-publish","format-standard","hentry","category-digital-humanities","tag-digital-history","tag-digital-humanities","tag-google","tag-history","tag-language","tag-ngram","tag-split-infinitive","tag-star-trek","tag-swearing","tag-walter-benjamin"],"jetpack_publicize_connections":[],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"","_links":{"self":[{"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/posts\/282","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/comments?post=282"}],"version-history":[{"count":15,"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/posts\/282\/revisions"}],"predecessor-version":[{"id":303,"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/posts\/282\/revisions\/303"}],"wp:attachment":[{"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/media?parent=282"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/categories?post=282"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/anterotesis.com\/wordpress\/wp-json\/wp\/v2\/tags?post=282"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}