From: Chilkat Software Date: 2006-11-08T08:27:01+09:00 Subject: Re: Unicode in irb on windows (respectively script/console in instantrails) A DOS console displays characters according to the OEM code page. Here is an example showing how to properly display a string with 8bit chars (e.g. characters with diacritics, or accent marks)... # file: oemCodePage.rb require 'chilkat' # (The CkString class is freeware) myStr = Chilkat::CkString.new() # A DOS console does NOT display this correctly: print "� � � �\n" # What we need is the OEM (DOS) code page... # OEM code pages are listed here: # http://msdn.microsoft.com/library/default.asp?url=/library/en-us/intl/unicode_81rn.asp myStr.appendAnsi("� � � �\n") # Emit the string in the character encoding of your choice: # ibm850 is the OEM code page for Latin1 print myStr.getEnc("ibm850") # Chilkat supports these: # us-ascii # unicode # unicodefffe # iso-8859-1 # iso-8859-2 # iso-8859-3 # iso-8859-4 # iso-8859-5 # iso-8859-6 # iso-8859-7 # iso-8859-8 # iso-8859-9 # iso-8859-13 # iso-8859-15 # windows-874 # windows-1250 # windows-1251 # windows-1252 # windows-1253 # windows-1254 # windows-1255 # windows-1256 # windows-1257 # windows-1258 # utf-7 # utf-8 # utf-32 # utf-32be # shift_jis # gb2312 # ks_c_5601-1987 # big5 # iso-2022-jp # iso-2022-kr # euc-jp # euc-kr # macintosh # x-mac-japanese # x-mac-chinesetrad # x-mac-korean # x-mac-arabic # x-mac-hebrew # x-mac-greek # x-mac-cyrillic # x-mac-chinesesimp # x-mac-romanian # x-mac-ukrainian # x-mac-thai # x-mac-ce # x-mac-icelandic # x-mac-turkish # x-mac-croatian # asmo-708 # dos-720 # dos-862 # ibm037 # ibm437 # ibm500 # ibm737 # ibm775 # ibm850 # ibm852 # ibm855 # ibm857 # ibm00858 # ibm860 # ibm861 # ibm863 # ibm864 # ibm865 # cp866 # ibm869 # ibm870 # cp875 # koi8-r # koi8-u At 05:07 PM 11/7/2006, you wrote: >On 11/7/06, michael.raidel@gmail.com wrote: >>I have a problem with Unicode in irb on Windows. I recognized it when >>trying to save an attribute of an ActiveRecord-Model with an umlaut >>(for example "�") in script/console. If the database connection is >>encoded in utf8, everything after the umlaut gets truncated, in the >>default encoding I get funny characters back. It doesn't matter if the >>$KCODE is set to UTF8 or NONE, the character number stays the same >>(also on plain irb)! > >The windows console -- also used by cygwin -- doesn't recognise UTF-8. >(That is, it's not possible to properly display UTF-8 in cmd.exe, at >least so far as I can tell.) > >-austin >-- >Austin Ziegler * halostatue@gmail.com * http://www.halostatue.ca/ > * austin@halostatue.ca * http://www.halostatue.ca/feed/ > * austin@zieglers.ca > > > > >-- >No virus found in this incoming message. >Checked by AVG Free Edition. >Version: 7.1.409 / Virus Database: 268.13.31/522 - Release Date: 11/7/2006 -- No virus found in this outgoing message. Checked by AVG Free Edition. Version: 7.1.409 / Virus Database: 268.13.31/522 - Release Date: 11/7/2006