From: Brian Candler Date: 2010-03-11T18:54:21+09:00 Subject: Re: reading an UTF-8 encoded file Use 'puts' instead of 'p' and it may work. That is, I suspect String#inspect is doing some mangling. You really should look at your postings in ruby-forum: http://www.ruby-forum.com/topic/205792 Wherever you say ruby 1.9 is giving the 'wrong' output it is correct, and where you say ruby 1.8 is giving the 'right' output it is wrong. I have a suspicion that there is a mismatch between the file content and the terminal. What if you just type "cat /Users/yt/dev/Signature/signatures.txt" at the terminal? > accentuated chars are correct now, notice i have to use "puts" instead > of "p" to get the chars otherwise i got the unicode code as > "v\303\251rit\303\251". Yes, String#inspect in ruby 1.8 will mangle all values over 128 into escaped form. String#inspect in ruby 1.9 behaves differently, and doesn't always mangle them. However, I just noticed 'macruby' in your scripts. Are you actually running MacRuby, or genuine Matz Ruby Interpreter 1.9 ? If it's macruby all bets are off - I thought it was a completely different interpreter written from scratch. I have no Mac here to compare behaviour with, and I have no idea what variation of 1.9 encoding rules MacRuby has implemented. In particular, I'm surprised that your program sees strings tagged as "US-ASCII" rather than "UTF-8" when you explicitly opened the file with external encoding of UTF-8. This makes me very suspicious of your actual ruby platform. Try adding this line to your code to get info about the Ruby platform: p Object.constants.grep(/RUBY/).map { |n| [n, Object.const_get(n)] } Regards, Brian. P.S. For comparison, here's what I get with an oldish ruby pre-1.9.2 under Linux. Try these on your system. >> File.open("/etc/passwd","r:ISO-8859-1").gets.encoding => # >> File.open("/etc/passwd","r:UTF-8").gets.encoding => # >> RUBY_DESCRIPTION => "ruby 1.9.2dev (2009-07-18 trunk 24186) [i686-linux]" -- Posted via http://www.ruby-forum.com/.