From: Jian Lin Date: 2009-05-09T22:42:02+09:00 Subject: Re: converting UTF-8 to entities like 剛 Rick Denatale wrote: > On Sat, May 9, 2009 at 8:40 AM, Jian Lin wrote: >> (newline, 0, etc)... but seems like there is no match > I'm pretty sure that Robert used regexen as the geeky way of pluralizing > regex. > > The unicode switch (a u regular expression option) forces the use of > unicode to interpret the string being matched, otherwise it uses > whatever the encoding of the source file containing the regular > expression. > > e.g. /./u > > If you want . to match newlines you want the m (multi-line) option. > Normally . will match anything BUT a new line, m changes this. > > rb(main):001:0> "a\nb".match(/a.b/) > => nil > irb(main):002:0> "a\nb".match(/a.b/m) > => # aha... here i just want to match 4 bytes at a time, no matter what the bytes are. Using "m" won't do it... the "u" would be helpful if i match one UTF-8 character at a time and then process it... right now i actually convert it all at once to UTF-32 and then process it... so I wonder if there is a way to match 4 bytes at a time. -- Posted via http://www.ruby-forum.com/.