From: Ryan Smith Date: 2010-01-29T03:13:25+09:00 Subject: Re: gsub not working to replace a 'Chinese' Charater. I parse a webpage which encoded in gb2312, using Watir to get the context of the page title, and want to replace the 'chinese character' in title with english words. When puts title which watir get, the chinese character displaied as corrupt code there (under windows cmd,code page using cp936, display works normal when change code page to utf-8). But I think cmd's code page just display setting does not related with what I need (replace chinese char). I did not know if string I get by Watir is also in 'gb2312' encoding or something others, the fact is fail happen when convert the string to utf-8 encoding, message is complain the char is invalid. totally no idea what need to do. Richard Conroy wrote: > On Thu, Jan 28, 2010 at 5:05 PM, Ryan Smith > wrote: > >> thank you! >> > > Mixing encoding schemes is hell in almost any context, and Ruby is no > exception. > Until you have complete control in your program over all encoding inputs > you > are > going to fail. > > If your input is coming from the shell environment or standard in the > text > can be > in the system encoding, regardless of what encoding you specify in Ruby. > > It is preferable to use unicode (UTF-8) in any operation where you are > processing > multilingual text. Failing that there is the Iconv library which you can > use > to convert > between encoding schemes. > > Note that 'double-byte encoding scheme' is an utterly useless term for > practical encoding > purposes. Its a gross simplification of what is going on, and especially > so > with Han > character sets. To do any practical work with non-unicode, multi-byte > character sets, you > have to know the encoding scheme. -- Posted via http://www.ruby-forum.com/.