From: Mike Harris Date: 2006-06-27T00:54:12+09:00 Subject: Re: Regex Comparison Causing Massive Memory Usage Xavier Noria wrote: > On Jun 26, 2006, at 17:35, Mike Harris wrote: > >> "My" Regex is not actually mine. It's in Ruby's SOAP >> Implementation. The crash is (I believe) in this method >> >> def Charset.is_utf8(str) >> UTF8Regexp =~ str >> end >> >> and the regex in question looks like this >> >> character_utf8 = >> "(?:#{us_ascii}|#{twobytes_utf8}|#{threebytes_utf8}|# >> {fourbytes_utf8})" >> UTF8Regexp = Regexp.new("\\A#{character_utf8}*\\z", nil, "NONE") > > > Yeah, looks potentially problematic at first sight. > > In each single byte the engine has up to 3 possible backtracking > points that have to be remembered. I don't know details about the > UTF-8 encoding, but if the alternatives are not exclusive you may > wrongly match #{twobytes_utf8} in the first byte to discover at the > end of the string the match fails. In that case the engine has to > backtrack a zillion times back and forth until it undoes the walked > way (potentially a huge tree) and arrives to switch to # > {threebytes_utf8} at the beginning, which in turn may fail, etc. > > -- fxn > > Yup. I figured it was something in the "has to store something at each character" vein, although I had no clue what was actually happening. I "fixed it" by temporarily commenting out the calling code, since I'm only receiving known data. In your opinion, is this unavoidable, or is this something that should ideally be fixed in the implementation?