How to make tr aware of non-ascii(unicode) characters? Announcing the arrival of Valued Associate #679: Cesar Manara Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern) 2019 Community Moderator Election Results Why I closed the “Why is Kali so hard” questiontr not replacing apostropheHow do I turn accented lowercase letters to uppercase? - Using the 'tr' commandHow can I convert Persian numerals in UTF-8 to European numerals in ASCII?tr analog for unicode characters?How to translate Unicode characters?How do I extract only alphanumeric characters from a given text file and print them?Character count of language X in mixed text file?Print out binary data as is without breaking the terminalRemove new line, space from fileHow to do a regex search in a UTF-16LE file while in a UTF-8 locale?Non-ASCII printable characters in sshd bannerHow can I make the TTY use the appropriate charset?How to make the login shell xterm use utf-8?Unicode support in talk?Detect how much of Unicode my terminal supports, even through screenWhy doesn't my Perl play nice with Unicode?tr analog for unicode characters?How to translate Unicode characters?Removing characters with sed

What is the meaning of the new sigil in Game of Thrones Season 8 intro?

Should I discuss the type of campaign with my players?

Can a USB port passively 'listen only'?

Error "illegal generic type for instanceof" when using local classes

Why light coming from distant stars is not discreet?

How to tell that you are a giant?

Resolving to minmaj7

Can an alien society believe that their star system is the universe?

Why are Kinder Surprise Eggs illegal in the USA?

Single word antonym of "flightless"

Identifying polygons that intersect with another layer using QGIS?

String `!23` is replaced with `docker` in command line

Denied boarding although I have proper visa and documentation. To whom should I make a complaint?

When do you get frequent flier miles - when you buy, or when you fly?

Book where humans were engineered with genes from animal species to survive hostile planets

How to deal with a team lead who never gives me credit?

Echoing a tail command produces unexpected output?

51k Euros annually for a family of 4 in Berlin: Is it enough?

Why is my conclusion inconsistent with the van't Hoff equation?

Sci-Fi book where patients in a coma ward all live in a subconscious world linked together

How do pianists reach extremely loud dynamics?

Can a non-EU citizen traveling with me come with me through the EU passport line?

Check which numbers satisfy the condition [A*B*C = A! + B! + C!]

What does an IRS interview request entail when called in to verify expenses for a sole proprietor small business?



How to make tr aware of non-ascii(unicode) characters?



Announcing the arrival of Valued Associate #679: Cesar Manara
Planned maintenance scheduled April 17/18, 2019 at 00:00UTC (8:00pm US/Eastern)
2019 Community Moderator Election Results
Why I closed the “Why is Kali so hard” questiontr not replacing apostropheHow do I turn accented lowercase letters to uppercase? - Using the 'tr' commandHow can I convert Persian numerals in UTF-8 to European numerals in ASCII?tr analog for unicode characters?How to translate Unicode characters?How do I extract only alphanumeric characters from a given text file and print them?Character count of language X in mixed text file?Print out binary data as is without breaking the terminalRemove new line, space from fileHow to do a regex search in a UTF-16LE file while in a UTF-8 locale?Non-ASCII printable characters in sshd bannerHow can I make the TTY use the appropriate charset?How to make the login shell xterm use utf-8?Unicode support in talk?Detect how much of Unicode my terminal supports, even through screenWhy doesn't my Perl play nice with Unicode?tr analog for unicode characters?How to translate Unicode characters?Removing characters with sed



.everyoneloves__top-leaderboard:empty,.everyoneloves__mid-leaderboard:empty,.everyoneloves__bot-mid-leaderboard:empty margin-bottom:0;








33















I'm trying to remove some characters from file(UTF-8). I'm using tr for this purpose:



tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat 


File contains some foreign characters (like "Латвийская" or "àé"). tr doesn't seem to understand them: it treats them as non-alpha and removes too.



I've tried changing some of my locale settings:



LC_CTYPE=C LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=ru_RU.UTF-8 tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat


Unfortunately, none of these worked.



How can I make tr understand Unicode?










share|improve this question






























    33















    I'm trying to remove some characters from file(UTF-8). I'm using tr for this purpose:



    tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat 


    File contains some foreign characters (like "Латвийская" or "àé"). tr doesn't seem to understand them: it treats them as non-alpha and removes too.



    I've tried changing some of my locale settings:



    LC_CTYPE=C LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
    LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
    LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=ru_RU.UTF-8 tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat


    Unfortunately, none of these worked.



    How can I make tr understand Unicode?










    share|improve this question


























      33












      33








      33


      4






      I'm trying to remove some characters from file(UTF-8). I'm using tr for this purpose:



      tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat 


      File contains some foreign characters (like "Латвийская" or "àé"). tr doesn't seem to understand them: it treats them as non-alpha and removes too.



      I've tried changing some of my locale settings:



      LC_CTYPE=C LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
      LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
      LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=ru_RU.UTF-8 tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat


      Unfortunately, none of these worked.



      How can I make tr understand Unicode?










      share|improve this question
















      I'm trying to remove some characters from file(UTF-8). I'm using tr for this purpose:



      tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat 


      File contains some foreign characters (like "Латвийская" or "àé"). tr doesn't seem to understand them: it treats them as non-alpha and removes too.



      I've tried changing some of my locale settings:



      LC_CTYPE=C LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
      LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=C tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat
      LC_CTYPE=ru_RU.UTF-8 LC_COLLATE=ru_RU.UTF-8 tr -cs '[[:alpha:][:space:]]' ' ' <testdata.dat


      Unfortunately, none of these worked.



      How can I make tr understand Unicode?







      linux text-processing unicode tr






      share|improve this question















      share|improve this question













      share|improve this question




      share|improve this question








      edited Sep 9 '15 at 14:53









      Toby Speight

      5,61811234




      5,61811234










      asked Sep 9 '15 at 12:57









      MatthewRockMatthewRock

      4,07331849




      4,07331849




















          1 Answer
          1






          active

          oldest

          votes


















          27














          That's a known (1, 2, 3, 4, 5, 6) limitation of the GNU implementation of tr.



          It's not as much that it doesn't support foreign, non-English or non-ASCII characters, but that it doesn't support multi-byte characters.



          Those Cyrillic characters would be treated OK, if written in the iso8859-5 (single-byte per character) character set (and your locale was using that charset), but your problem is that you're using UTF-8 where non-ASCII characters are encoded in 2 or more bytes.



          GNU's got a plan (see also) to fix that and work is under way but not there yet.



          FreeBSD or Solaris tr don't have the problem.




          In the mean time, for most use cases of tr, you can use GNU sed or GNU awk which do support multi-byte characters.



          For instance, your:



          tr -cs '[[:alpha:][:space:]]' ' '


          could be written:



          gsed -E 's/( |[^[:space:][:alpha:]])+/ /'


          or:



          gawk -v RS='( |[^[:space:][:alpha:]])+' 'printf "%s", sep $0; sep=" "'


          To convert between lower and upper case (tr '[:upper:]' '[:lower:]'):



          gsed 's/[[:upper:]]/l&/g'


          (that l is a lowercase L, not the 1 digit).



          or:



          gawk 'print tolower($0)'


          For portability, perl is another alternative:



          perl -Mopen=locale -pe 's/([^[:space:][:alpha:]]| )+/ /g'
          perl -Mopen=locale -pe '$_=lc$_'


          If you know the data can be represented in a single-byte character set, then you can process it in that charset:



          (export LC_ALL=ru_RU.iso88595
          iconv -f utf-8 |
          tr -cs '[:alpha:][:space:]' ' ' |
          iconv -t utf-8) < Russian-file.utf8





          share|improve this answer




















          • 1





            I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

            – MatthewRock
            Sep 9 '15 at 14:41






          • 3





            @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

            – Stéphane Chazelas
            Sep 9 '15 at 15:22











          • Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

            – Incnis Mrsi
            Sep 9 '15 at 16:35






          • 9





            @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

            – Stéphane Chazelas
            Sep 9 '15 at 16:43











          • @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

            – Alex Shpilkin
            Apr 18 '18 at 21:06












          Your Answer








          StackExchange.ready(function()
          var channelOptions =
          tags: "".split(" "),
          id: "106"
          ;
          initTagRenderer("".split(" "), "".split(" "), channelOptions);

          StackExchange.using("externalEditor", function()
          // Have to fire editor after snippets, if snippets enabled
          if (StackExchange.settings.snippets.snippetsEnabled)
          StackExchange.using("snippets", function()
          createEditor();
          );

          else
          createEditor();

          );

          function createEditor()
          StackExchange.prepareEditor(
          heartbeatType: 'answer',
          autoActivateHeartbeat: false,
          convertImagesToLinks: false,
          noModals: true,
          showLowRepImageUploadWarning: true,
          reputationToPostImages: null,
          bindNavPrevention: true,
          postfix: "",
          imageUploader:
          brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
          contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
          allowUrls: true
          ,
          onDemand: true,
          discardSelector: ".discard-answer"
          ,immediatelyShowMarkdownHelp:true
          );



          );













          draft saved

          draft discarded


















          StackExchange.ready(
          function ()
          StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f228558%2fhow-to-make-tr-aware-of-non-asciiunicode-characters%23new-answer', 'question_page');

          );

          Post as a guest















          Required, but never shown

























          1 Answer
          1






          active

          oldest

          votes








          1 Answer
          1






          active

          oldest

          votes









          active

          oldest

          votes






          active

          oldest

          votes









          27














          That's a known (1, 2, 3, 4, 5, 6) limitation of the GNU implementation of tr.



          It's not as much that it doesn't support foreign, non-English or non-ASCII characters, but that it doesn't support multi-byte characters.



          Those Cyrillic characters would be treated OK, if written in the iso8859-5 (single-byte per character) character set (and your locale was using that charset), but your problem is that you're using UTF-8 where non-ASCII characters are encoded in 2 or more bytes.



          GNU's got a plan (see also) to fix that and work is under way but not there yet.



          FreeBSD or Solaris tr don't have the problem.




          In the mean time, for most use cases of tr, you can use GNU sed or GNU awk which do support multi-byte characters.



          For instance, your:



          tr -cs '[[:alpha:][:space:]]' ' '


          could be written:



          gsed -E 's/( |[^[:space:][:alpha:]])+/ /'


          or:



          gawk -v RS='( |[^[:space:][:alpha:]])+' 'printf "%s", sep $0; sep=" "'


          To convert between lower and upper case (tr '[:upper:]' '[:lower:]'):



          gsed 's/[[:upper:]]/l&/g'


          (that l is a lowercase L, not the 1 digit).



          or:



          gawk 'print tolower($0)'


          For portability, perl is another alternative:



          perl -Mopen=locale -pe 's/([^[:space:][:alpha:]]| )+/ /g'
          perl -Mopen=locale -pe '$_=lc$_'


          If you know the data can be represented in a single-byte character set, then you can process it in that charset:



          (export LC_ALL=ru_RU.iso88595
          iconv -f utf-8 |
          tr -cs '[:alpha:][:space:]' ' ' |
          iconv -t utf-8) < Russian-file.utf8





          share|improve this answer




















          • 1





            I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

            – MatthewRock
            Sep 9 '15 at 14:41






          • 3





            @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

            – Stéphane Chazelas
            Sep 9 '15 at 15:22











          • Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

            – Incnis Mrsi
            Sep 9 '15 at 16:35






          • 9





            @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

            – Stéphane Chazelas
            Sep 9 '15 at 16:43











          • @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

            – Alex Shpilkin
            Apr 18 '18 at 21:06
















          27














          That's a known (1, 2, 3, 4, 5, 6) limitation of the GNU implementation of tr.



          It's not as much that it doesn't support foreign, non-English or non-ASCII characters, but that it doesn't support multi-byte characters.



          Those Cyrillic characters would be treated OK, if written in the iso8859-5 (single-byte per character) character set (and your locale was using that charset), but your problem is that you're using UTF-8 where non-ASCII characters are encoded in 2 or more bytes.



          GNU's got a plan (see also) to fix that and work is under way but not there yet.



          FreeBSD or Solaris tr don't have the problem.




          In the mean time, for most use cases of tr, you can use GNU sed or GNU awk which do support multi-byte characters.



          For instance, your:



          tr -cs '[[:alpha:][:space:]]' ' '


          could be written:



          gsed -E 's/( |[^[:space:][:alpha:]])+/ /'


          or:



          gawk -v RS='( |[^[:space:][:alpha:]])+' 'printf "%s", sep $0; sep=" "'


          To convert between lower and upper case (tr '[:upper:]' '[:lower:]'):



          gsed 's/[[:upper:]]/l&/g'


          (that l is a lowercase L, not the 1 digit).



          or:



          gawk 'print tolower($0)'


          For portability, perl is another alternative:



          perl -Mopen=locale -pe 's/([^[:space:][:alpha:]]| )+/ /g'
          perl -Mopen=locale -pe '$_=lc$_'


          If you know the data can be represented in a single-byte character set, then you can process it in that charset:



          (export LC_ALL=ru_RU.iso88595
          iconv -f utf-8 |
          tr -cs '[:alpha:][:space:]' ' ' |
          iconv -t utf-8) < Russian-file.utf8





          share|improve this answer




















          • 1





            I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

            – MatthewRock
            Sep 9 '15 at 14:41






          • 3





            @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

            – Stéphane Chazelas
            Sep 9 '15 at 15:22











          • Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

            – Incnis Mrsi
            Sep 9 '15 at 16:35






          • 9





            @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

            – Stéphane Chazelas
            Sep 9 '15 at 16:43











          • @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

            – Alex Shpilkin
            Apr 18 '18 at 21:06














          27












          27








          27







          That's a known (1, 2, 3, 4, 5, 6) limitation of the GNU implementation of tr.



          It's not as much that it doesn't support foreign, non-English or non-ASCII characters, but that it doesn't support multi-byte characters.



          Those Cyrillic characters would be treated OK, if written in the iso8859-5 (single-byte per character) character set (and your locale was using that charset), but your problem is that you're using UTF-8 where non-ASCII characters are encoded in 2 or more bytes.



          GNU's got a plan (see also) to fix that and work is under way but not there yet.



          FreeBSD or Solaris tr don't have the problem.




          In the mean time, for most use cases of tr, you can use GNU sed or GNU awk which do support multi-byte characters.



          For instance, your:



          tr -cs '[[:alpha:][:space:]]' ' '


          could be written:



          gsed -E 's/( |[^[:space:][:alpha:]])+/ /'


          or:



          gawk -v RS='( |[^[:space:][:alpha:]])+' 'printf "%s", sep $0; sep=" "'


          To convert between lower and upper case (tr '[:upper:]' '[:lower:]'):



          gsed 's/[[:upper:]]/l&/g'


          (that l is a lowercase L, not the 1 digit).



          or:



          gawk 'print tolower($0)'


          For portability, perl is another alternative:



          perl -Mopen=locale -pe 's/([^[:space:][:alpha:]]| )+/ /g'
          perl -Mopen=locale -pe '$_=lc$_'


          If you know the data can be represented in a single-byte character set, then you can process it in that charset:



          (export LC_ALL=ru_RU.iso88595
          iconv -f utf-8 |
          tr -cs '[:alpha:][:space:]' ' ' |
          iconv -t utf-8) < Russian-file.utf8





          share|improve this answer















          That's a known (1, 2, 3, 4, 5, 6) limitation of the GNU implementation of tr.



          It's not as much that it doesn't support foreign, non-English or non-ASCII characters, but that it doesn't support multi-byte characters.



          Those Cyrillic characters would be treated OK, if written in the iso8859-5 (single-byte per character) character set (and your locale was using that charset), but your problem is that you're using UTF-8 where non-ASCII characters are encoded in 2 or more bytes.



          GNU's got a plan (see also) to fix that and work is under way but not there yet.



          FreeBSD or Solaris tr don't have the problem.




          In the mean time, for most use cases of tr, you can use GNU sed or GNU awk which do support multi-byte characters.



          For instance, your:



          tr -cs '[[:alpha:][:space:]]' ' '


          could be written:



          gsed -E 's/( |[^[:space:][:alpha:]])+/ /'


          or:



          gawk -v RS='( |[^[:space:][:alpha:]])+' 'printf "%s", sep $0; sep=" "'


          To convert between lower and upper case (tr '[:upper:]' '[:lower:]'):



          gsed 's/[[:upper:]]/l&/g'


          (that l is a lowercase L, not the 1 digit).



          or:



          gawk 'print tolower($0)'


          For portability, perl is another alternative:



          perl -Mopen=locale -pe 's/([^[:space:][:alpha:]]| )+/ /g'
          perl -Mopen=locale -pe '$_=lc$_'


          If you know the data can be represented in a single-byte character set, then you can process it in that charset:



          (export LC_ALL=ru_RU.iso88595
          iconv -f utf-8 |
          tr -cs '[:alpha:][:space:]' ' ' |
          iconv -t utf-8) < Russian-file.utf8






          share|improve this answer














          share|improve this answer



          share|improve this answer








          edited 10 hours ago

























          answered Sep 9 '15 at 13:47









          Stéphane ChazelasStéphane Chazelas

          315k57597955




          315k57597955







          • 1





            I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

            – MatthewRock
            Sep 9 '15 at 14:41






          • 3





            @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

            – Stéphane Chazelas
            Sep 9 '15 at 15:22











          • Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

            – Incnis Mrsi
            Sep 9 '15 at 16:35






          • 9





            @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

            – Stéphane Chazelas
            Sep 9 '15 at 16:43











          • @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

            – Alex Shpilkin
            Apr 18 '18 at 21:06













          • 1





            I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

            – MatthewRock
            Sep 9 '15 at 14:41






          • 3





            @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

            – Stéphane Chazelas
            Sep 9 '15 at 15:22











          • Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

            – Incnis Mrsi
            Sep 9 '15 at 16:35






          • 9





            @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

            – Stéphane Chazelas
            Sep 9 '15 at 16:43











          • @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

            – Alex Shpilkin
            Apr 18 '18 at 21:06








          1




          1





          I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

          – MatthewRock
          Sep 9 '15 at 14:41





          I've accepted your question because of information about tr. I've solved the problem, and removed question about how to solve it(so people looking for tr will find only information about tr, not some arbitrary problem). If you could please remove solution too, since it's no longer needed, I'd be thankful.

          – MatthewRock
          Sep 9 '15 at 14:41




          3




          3





          @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

          – Stéphane Chazelas
          Sep 9 '15 at 15:22





          @MatthewRock I've kept it but reworded it and made more generic as giving a word around would be useful to people with the same problem.

          – Stéphane Chazelas
          Sep 9 '15 at 15:22













          Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

          – Incnis Mrsi
          Sep 9 '15 at 16:35





          Where do you get an idea that Cyrillic is (customarily) encoded in ISO 8859-5? Did you ever see a Russian text in anything but Unicode?

          – Incnis Mrsi
          Sep 9 '15 at 16:35




          9




          9





          @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

          – Stéphane Chazelas
          Sep 9 '15 at 16:43





          @IncnisMrsi, all that matters here is that ISO 8859-5 is one of those singe-byte charsets that has those Cyrillic characters. Whether it's in widespread use or not is irrelevant here. If you have a locale with KOI-R or window-1251 charset, by all means, use it instead.

          – Stéphane Chazelas
          Sep 9 '15 at 16:43













          @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

          – Alex Shpilkin
          Apr 18 '18 at 21:06






          @IncnisMrsi Russian on the web is almost always encoded in UTF-8 (or occasionally in Windows-1251), but only because we’ve felt the pain of many single-byte encodings early on. Here’s an ancient (circa 1998) web page with a (non-functional) encoding switcher: sch57.ru/collect.

          – Alex Shpilkin
          Apr 18 '18 at 21:06


















          draft saved

          draft discarded
















































          Thanks for contributing an answer to Unix & Linux Stack Exchange!


          • Please be sure to answer the question. Provide details and share your research!

          But avoid


          • Asking for help, clarification, or responding to other answers.

          • Making statements based on opinion; back them up with references or personal experience.

          To learn more, see our tips on writing great answers.




          draft saved


          draft discarded














          StackExchange.ready(
          function ()
          StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2funix.stackexchange.com%2fquestions%2f228558%2fhow-to-make-tr-aware-of-non-asciiunicode-characters%23new-answer', 'question_page');

          );

          Post as a guest















          Required, but never shown





















































          Required, but never shown














          Required, but never shown












          Required, but never shown







          Required, but never shown

































          Required, but never shown














          Required, but never shown












          Required, but never shown







          Required, but never shown







          -linux, text-processing, tr, unicode

          Popular posts from this blog

          Word for a person who has no opinion about whether god existsWord for having a definite opinion while simultaneously withholding judgment?What's the opposite of “newcomer? Is ”veteran" OK?What do you call an “atheist” who might believe in an afterlife?What's a word for someone who wants to voice opinions but not have them challenged?Word for someone who dismisses contrary opinions as irrational?Somone who thinks they are overly special/out of the ordinaryIs there a word, phrase or idiom for “a person who is incapable of thinking about the future”?The belief that a god is human-likeA word for a non-famous person/thing you have heard a lot aboutAdjective for a person who enjoys taking care of their appearance

          What was this official D&D 3.5e Lovecraft-flavored rulebook?What was this set of RPG tools called?As a first-time DM should I let my players play complex character classes and roles?Nymph's Kiss and the RelationshipWhat was the name of this Cleric Prestige Class that shapes metal with its bare hands?Are the 3.5e Dragonlance books third party or official works?What's up with the domain Vile Darkness?What was this 80s book about RPGs?What was the name of this Werewolf band?What book had Rituals to “upgrade” animal companions to keep them viable at higher levels?What was this RPG that had rules for player-owned businesses?

          2017 IndyCar Series Contents Series news Teams and drivers Schedule Season summary Footnotes References External links Navigation menu"INDYCAR: Initial 2018 bodywork concepts unveiled"the original"IndyCar confirms switch to Performance Friction brakes in 2017""AJ Foyt Racing will switch to Chevy"the original"Carlos Munoz, Conor Daly will drive for AJ Foyt Racing""Zach Veach's Indy 500 Debut Confirmed with Foyt""No mass exodus from Honda after Ganassi switch""Ex-F1 driver Sato joins Andretti Autosport for 2017 IndyCar season""IndyCar's Ryan Hunter-Reay, sponsor DHL paired through 2020""hhgregg and Andretti Autosport announce partnership for key races in 2016""INDYCAR: Rossi re-signs with Andretti"the original"McLaren Formula 1 - Fernando Alonso to race at Indy 500 with McLaren, Honda and Andretti Autosport""Shank will finally take part in Indy 500 with Harvey, Andretti | MotorSportsTalk""Andretti adds Jack Harvey to Indy 500 field""Ganassi switches to Honda power for 2017""INDYCAR: Chilton returns to Ganassi"the original"IndyCar silly season: Who's going where in 2017?""INDYCAR: Kanaan, NTT Data return to Ganassi"the original"Kimball to remain at Ganassi for 2017""Coyne confirms Bourdais for 2017 IndyCar season""Davison to sub for Bourdais in Indy 500"the original"Gutierrez confirmed for Detroit IndyCar debut""Gutierrez returns with Coyne for rest of 2017 season""Vautier to drive for Coyne at Texas"the original"INDYCAR: Coyne confirms Jones for 2017"the original"Pippa Mann returns to Coyne for Indy 500""Karam, Dreyer & Reinbold teaming up again for Indianapolis 500""Pigot to return to Ed Carpenter Racing""Hildebrand confirmed as full-time Ed Carpenter driver""Veach to replace injured Hildebrand at Barber"the originalNew Team Harding Racing Enters Chaves for 101st Indianapolis 500"Juncos Racing Announces Entry in 101st Running of the Indianapolis 500 :: Juncos Racing""Juncos confirms Pigot for Indy 500""Saavedra confirmed in Juncos' second 500 entry"the original"Lazier confirms Indy 500 run after son's USF2000 debut"the original"Claman DeMelo to race for RLLR at Sonoma"the original"Rahal signs Servia and ace engineer for 2017""IndyCar: Aleshin returns with Schmidt"the original"Aleshin replaced by Saavedra for Toronto""Jack Harvey will pilot SPM No. 7 car at Watkins Glen, Sonoma""Jay Howard confirmed in Tony Stewart's supported SPM Indy entry""INDYCAR: Newgarden to wave the flag at Penske"the original"Pagenaud opts for No. 1 in 2017"the original"Penske confirms Newgarden for 2017""Montoya to stay with Team Penske in 2017""Target leaving IndyCar after 27 seasons with Chip Ganassi""Cavin: IndyCar could see complete driver/team shakeup in 2017""End of the road for KV Racing?""KV Racing confirms closure, equipment sold to Juncos""Juncos confirms IndyCar Series entry"the original"Juncos readies IndyCar program, aims for '17 500"the original"Harding Racing to add Texas, Pocono to schedule"the original"Sato signs with Andretti Autosport for 2017""INDYCAR: Aleshin in Doubt at SPM"the original"Long Beach notebook: JR Hildebrand breaks hand""Hildebrand cleared to return at Phoenix"the original"Bourdais to undergo surgery on multiple fractures""Aleshin loses Schmidt Peterson IndyCar ride""Saavedra in at SPM for Pocono, Gateway"the original"Bourdais to make return at Gateway"the original"The IndyCar Grand Prix no longer is sponsored by Angie's List""2017 IndyCar Series rulebook""2017 Verizon IndyCar Series Official Rulebook"Official websiteeeeee