From: Xavier Noria Date: 2006-06-27T00:48:24+09:00 Subject: Re: Regex Comparison Causing Massive Memory Usage On Jun 26, 2006, at 17:35, Mike Harris wrote: > "My" Regex is not actually mine. It's in Ruby's SOAP > Implementation. The crash is (I believe) in this method > > def Charset.is_utf8(str) > UTF8Regexp =~ str > end > > and the regex in question looks like this > > character_utf8 = > "(?:#{us_ascii}|#{twobytes_utf8}|#{threebytes_utf8}|# > {fourbytes_utf8})" > UTF8Regexp = Regexp.new("\\A#{character_utf8}*\\z", nil, "NONE") Yeah, looks potentially problematic at first sight. In each single byte the engine has up to 3 possible backtracking points that have to be remembered. I don't know details about the UTF-8 encoding, but if the alternatives are not exclusive you may wrongly match #{twobytes_utf8} in the first byte to discover at the end of the string the match fails. In that case the engine has to backtrack a zillion times back and forth until it undoes the walked way (potentially a huge tree) and arrives to switch to # {threebytes_utf8} at the beginning, which in turn may fail, etc. -- fxn