From: messju mohr Date: 2003-12-17T19:40:52+09:00 Subject: Re: non-english characters On Wed, Dec 17, 2003 at 06:12:00PM +0900, Robert Klemme wrote: > > "messju mohr" schrieb im Newsbeitrag > news:20031217082712.GE17320@pharao.lammfellpuschen.de... > > On Wed, Dec 17, 2003 at 04:05:32PM +0900, Yukihiro Matsumoto wrote: > > > Hi, > > > > > > In message "Re: non-english characters" > > > on 03/12/17, Daniel Bretoi writes: > > > > > > |I'm not sure, how can I find out what the germans use? and once I know > > > |that part, how do I use it? > > > > > > Ask somebody around you to find out. Then if you're going to use > > > Unicode (UTF-8), write your script in UTF-8 and invoke Ruby with -Ku > > > option. If you use ISO-8859-* or any other single byte encoding, you > > > don't have to do anything special. > > > > > > matz. > > > > hmm. > > > > regexp works fine for me with unicode. either with "ruby -Ku" on > > startup or with the /u as regexp-option. > > > > but with ISO-8859-+ (1 or 15 in my case) i don't get \w to match > > accented characters. > > I guess \w is defined in terms of ASCII - and there you don't have "�", "�" > and similar chars. yes, it looks like i got confused by the PCRE library which treats \w according to the current locale. too-many-languages error. :) > Regards > > robert