From: yvon.thoravallist@... (Yvon Thoraval) Date: 2003-09-24T03:34:37+09:00 Subject: Re: [newbie] upper to lower first letter of a word Mark J. Reed wrote: > Really? That's arguably a bug. What character encoding are you using? I'm (more-or-less) sure about that because even if i put : l.gsub!(/\b[a-z������������������]+/) { |w| black_list.include?(w) ? w : w.capitalize } i get : M�Con Sup�Rieur when input was : M�con sup�rieur > Accented letters should be in \w, not \W, and therefore the > space between one and an adjacent letter should not match \b. > But Ruby regexes may be ASCII-only, and even if not, they're probably > Latin-1-only. So, for instance, they wouldn't work on UTF-8 strings. precisely i'm using utf-8 �;) however, i'm able to do a try using iso-8859-1, my word editor (Pepper on MacOS X) is able to transcode within 2 clicks + one cut'n paste rom utf to iso... sounds strange to me because Ruby is coming from Japan where "special" chars are every-day chars ??? [snip] > The block has to compensate for that. Something like this: > > string.gsub!(/(^|[- '])([a-z]+)/) { $1 + $2.capitalize } > > Except that [a-z] won't match accented characters, so it's more like this: > > string.gsub!(/(^|[- '])([a-z����������������]+)/) { $1 + $2.capitalize } > > And if the names aren't limited to French, then even more special characters > creep in . . . Yes, right, i know, for the time being, only about french and german accentuated chars... However because vintage are classified by area i might have to change regex upon region... -- Yvon