From: 7stud -- Date: 2009-05-10T05:10:29+09:00 Subject: Re: converting UTF-8 to entities like 剛 Jian Lin wrote: > Rick Denatale wrote: >> On Sat, May 9, 2009 at 8:40 AM, Jian Lin wrote: >>> (newline, 0, etc)... but seems like there is no match >> I'm pretty sure that Robert used regexen as the geeky way of pluralizing >> regex. >> >> The unicode switch (a u regular expression option) forces the use of >> unicode to interpret the string being matched, otherwise it uses >> whatever the encoding of the source file containing the regular >> expression. >> >> e.g. /./u >> >> If you want . to match newlines you want the m (multi-line) option. >> Normally . will match anything BUT a new line, m changes this. >> >> rb(main):001:0> "a\nb".match(/a.b/) >> => nil >> irb(main):002:0> "a\nb".match(/a.b/m) >> => # > > aha... here i just want to match 4 bytes at a time, no matter what the > bytes are. Using "m" won't do it... the "u" would be helpful if i > match one UTF-8 character at a time and then process it... right now i > actually convert it all at once to UTF-32 and then process it... so I > wonder if there is a way to match 4 bytes at a time. So what's the problem? A dot matches any byte (with the 'm' switch). Make a regex with four dots: /..../ or /.{4}/ -- Posted via http://www.ruby-forum.com/.