From: Richard Conroy Date: 2010-01-29T02:57:36+09:00 Subject: Re: gsub not working to replace a 'Chinese' Charater. --00032555a892444f24047e3d3361 Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable On Thu, Jan 28, 2010 at 5:05 PM, Ryan Smith wrote: > gsub not works for me when replace 'DBCS'(double byte character set) > character, using last version ruby 1.8.6 > > when "str=BA=BA=D7=D6end".gsub(/=BA=BA=D7=D6/,"hanzi"), > output still is: str=BA=BA=D7=D6end , but not strhanziend which I want = to > get. > > Searched web two whole night with no clue found. > > Anyone can help are much appreciated, need got it work very urgent. > thank you! > Mixing encoding schemes is hell in almost any context, and Ruby is no exception. Until you have complete control in your program over all encoding inputs yo= u are going to fail. If your input is coming from the shell environment or standard in the text can be in the system encoding, regardless of what encoding you specify in Ruby. It is preferable to use unicode (UTF-8) in any operation where you are processing multilingual text. Failing that there is the Iconv library which you can us= e to convert between encoding schemes. Note that 'double-byte encoding scheme' is an utterly useless term for practical encoding purposes. Its a gross simplification of what is going on, and especially so with Han character sets. To do any practical work with non-unicode, multi-byte character sets, you have to know the encoding scheme. --=20 http://richardconroy.blogspot.com --00032555a892444f24047e3d3361--