{"id":1713,"date":"2017-02-05T14:54:27","date_gmt":"2017-02-05T14:54:27","guid":{"rendered":"http:\/\/phaidon.philo.at\/qu\/?p=1713"},"modified":"2018-11-15T08:48:12","modified_gmt":"2018-11-15T08:48:12","slug":"plain-text","status":"publish","type":"post","link":"https:\/\/quatsch.philo.at\/?p=1713","title":{"rendered":"plain text"},"content":{"rendered":"<blockquote><p>ASCII was very carefully designed; essentially no character has its code by accident. Everything had a reason, although some of those reasons are long obsolete.<br \/>\n&#8211;<a href=\"http:\/\/web.archive.org\/web\/20170203012151\/https:\/\/news.ycombinator.com\/item?id=13499386\">kps<\/a><\/p><\/blockquote>\n<p>As digital natives we tend to take for granted the significant efforts of intelligence that enabled the global tech revolution. <a href=\"https:\/\/garbagecollected.org\/2017\/01\/31\/four-column-ascii\/\">Here<\/a> is one small detail about ASCII control characters, which are for example still used to some extend in terminals (ssh).<\/p>\n<p>What we can learn here is compactness and elegance. But let&#8217;s not transfigure the past. Maybe it&#8217;s a lesson about the temptations of universality.<\/p>\n<p><!--more--><\/p>\n<hr \/>\n<p>Interestingly, knowledge hides behind established opaque conventions. The <a href=\"http:\/\/www.ascii-code.com\/\">usual representation\u00a0<\/a>isolates the non-printable characters from the printable ones. While this grouping makes sense for simple introductions (&#8220;The first 32 characters are control characters&#8221;), it\u00a0obfuscates\u00a0relations\u00a0that can be of practical and intellectual value. \u00a0By grouping the first 2 bits as columns and the last 5 bits as rows, it becomes clear how a non-printable &#8220;character&#8221; can be produced by pressing two keys on the keyboard.<\/p>\n<p>Generally speaking: Only after challenging\u00a0established conventions we can\u00a0estimate what we\u00a0lose if we drop the ideas and practices that created them.<\/p>\n<hr \/>\n<blockquote><p>And all was good, assuming you were an English speaker.<br \/>\n<a href=\"https:\/\/www.joelonsoftware.com\/2003\/10\/08\/the-absolute-minimum-every-software-developer-absolutely-positively-must-know-about-unicode-and-character-sets-no-excuses\/\">Joel Spolsky<\/a><\/p><\/blockquote>\n<p>The above linked articles conceal one important aspect about ASCII. It was chaotic, when scaled globally.\u00a0The initial elegant design of ASCII was not meant to be universal.\u00a0When the American Standard Code for Information Interchange (ASCII) was established in 1960-1963,\u00a0the intention was to have a\u00a0standard\u00a0for American devices (teleprinters and telegraphy), and maybe with some hints how others could extend it to support more characters. It was well-designed for this purpose: English alphabet,\u00a0some\u00a0frequently used signs, control characters.<\/p>\n<p>What about other scriptures?\u00a0The standard does not even mention them (nowadays we are expecting it to be universal). So once computers were sold and produced in other countries, the vendors deviated from the ASCII to fulfil the needs of the local language or idiom (e.g.: <strong>\u00df<\/strong>). Multiple countries were\u00a0overwriting ASCII in parallel to express special characters.<\/p>\n<p>As long as the devices were not interconnected, the misalignment was acceptable.\u00a0With the emergence of web sites\u00a0one realized that a consolidation of the local\u00a0deviations\u00a0is needed. After several attempts to find a viable character encoding for all world&#8217;s language (based on Unicode), UTF-8 prevailed in the last decade of the 20th century.<\/p>\n<p><i>Universal Character Set<\/i>\u00a0Transformation Format (UTF-8) has the\u00a0advantage that it is compatible with the ASCII. Nothing changed for the English-language dominated industry.<\/p>\n<p>All the other countries, who used to develop their local variant of ASCII (equally ignorant to their neighbouring languages) had the disadvantage of encoding their characters in more complex ways for the sake of universalism and interoperability (it&#8217;s not fully accurate, because e.g. also parts of ISCII were re-used in Unicode).<\/p>\n<hr \/>\n<p>For example:<\/p>\n<ul>\n<li>Devanagari, the Indian &#8220;alphabet&#8221; is frequently used to\u00a0write Hindi and similar languages (like Marathi).<\/li>\n<li>Expressing\u00a0the character \u0905\u00a0 (pronounced like: a) in the ASCII-Deviation <a href=\"http:\/\/varamozhi.sourceforge.net\/iscii91.pdf\">ISCII<\/a> (Indian <strong>Script<\/strong> Code for Information Interchange) required 8 bit: 1010 0100 (hexadecimal: A4).<\/li>\n<li>Expressing\u00a0the character\u00a0\u0905\u00a0in Unicode requires 12 bit: 0000 <strong>1001 0000 0101<\/strong> (hexadecimal: 0905)<\/li>\n<li>However, the situation is more complex. In Hindi the smallest meaningful unit is a syllable (consisting of consonant + vowel), while the consonant by default has the &#8220;a&#8221; included: For example this character is pronounced as &#8220;ka&#8221;:\u00a0\u0915. If one wants to say &#8220;ku&#8221;, the symbol gets a modifier: \u0915\u0941. To <a href=\"http:\/\/www.unicode.org\/charts\/PDF\/U0900.pdf\">encode<\/a> this, one needs two characters:<br \/>\n0000 <strong>1001 0001 0101\u00a0\u00a0<\/strong> (hexadecimal: 0915) for the consonant\u00a0\u0915 (&#8220;k(a)&#8221;)<br \/>\nand<br \/>\n0000 <strong>1001 0100 0001<\/strong> (hexadecimal: 0941) for the modifying matra that represents &#8220;u&#8221;: <span style=\"font-family: Trebuchet MS; font-size: medium;\"><big>\u0941<\/big><\/span><\/li>\n<li>In Unicode, the size in total is 24 bit.<\/li>\n<li>For ISCII,\u00a0encoding \u0915\u0941 requires\u00a012 bit: 1011 0011 1101 1101 \u00a0(hexadecimal: B3 DD).<\/li>\n<\/ul>\n<p>&nbsp;<\/p>\n<hr \/>\n<p>Looking at the recent developments in world politics, and according to some mainstream voices (e.g.\u00a0&#8220;<a href=\"https:\/\/www.credit-suisse.com\/ch\/en\/about-us\/research\/research-institute\/news-and-videos\/articles\/news-and-expertise\/2017\/01\/en\/where-are-you-headed-globalization.html\">Globalization is running out of steam<\/a>&#8220;), we are heading towards a multi-polar world, in the optimistic scenario.<\/p>\n<p>Situations like ASCII\u00a0and Unicode embed a local and a global approach of encoding characters in binary numbers that\u00a0are now the common\u00a0ground in most browsers, and for automatic translation of pages.<\/p>\n<p>It&#8217;s hard to imagine how a population that is used to UTF-8 will be interested in using only ASCII. It&#8217;s more easy to imagine that UTF-8 will be taken as basis on top of which custom, local extensions and palimpsests will be added, sacrificing some but not all interoperability.<\/p>\n<p>Like the cloud services: For them to operate it is invaluable to have a functioning Internet infrastructure with all it&#8217;s interoperable protocols and standards across the world. By using this infrastructure, they build up locally optimized, complex code that is exposed by narrow interfaces. This allows carefree usage by hiding complexity. At the same time, it covers an unbelievable amount of knowledge and data.<\/p>\n<p>Again, the isolation between hidden and visible\u00a0entities obfuscates\u00a0relations\u00a0that can be of practical and intellectual value.\u00a0You can say: &#8220;Not everything can and should be fully visible and interoperable. We should leave room for discovery.&#8221; Such new secrecy might\u00a0promote curiosity. But also abuse of power.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>ASCII was very carefully designed; essentially no character has its code by accident. Everything had a reason, although some of those reasons are long obsolete. &#8211;kps As digital natives we tend to take for granted the significant efforts of intelligence that enabled the global tech revolution. Here is one small detail about ASCII control characters,<a class=\"more-link\" href=\"https:\/\/quatsch.philo.at\/?p=1713\">Read more<\/a><\/p>\n","protected":false},"author":4,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[7,13],"tags":[37,40,60,137,155,191],"class_list":["post-1713","post","type-post","status-publish","format-standard","hentry","category-medienphilosophie","category-politik-2","tag-berechnung","tag-beziehungen","tag-daten","tag-open-access","tag-schreiben","tag-vernetzung"],"_links":{"self":[{"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=\/wp\/v2\/posts\/1713","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=\/wp\/v2\/users\/4"}],"replies":[{"embeddable":true,"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=1713"}],"version-history":[{"count":1,"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=\/wp\/v2\/posts\/1713\/revisions"}],"predecessor-version":[{"id":2142,"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=\/wp\/v2\/posts\/1713\/revisions\/2142"}],"wp:attachment":[{"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=1713"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=1713"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/quatsch.philo.at\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=1713"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}