From: Intransition Date: 2011-11-17T04:18:39+09:00 Subject: Re: Regex to divide document into sections? ------=_Part_34_31021916.1321471117362 Content-Type: multipart/alternative; boundary="----=_Part_35_16588280.1321471117363" ------=_Part_35_16588280.1321471117363 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit Hmm... doesn't seem to work for me. a #=> [["This is an example.\n a = 1\nOf what I mean.\n b = 2\nAnd it can go on\nlike this.\n\n c = 3\nFor ever and ever.\n\n d = 4\n", "\n\n "]] You know what though. I did some benchmarking and discovered that manually parsing the text line by line is much faster than using a regular expression (I used a close approx re). I was kind of surprised by this, since the regular expression engine is written in C, where as my line by line parser is in Ruby. Despite that, I still find it curious that there isn't a more obvious regular expression for parsing a document in this way. It makes me wonder if a C.S. PhD could go back to the drawing board, and come up with a better alternative to REs. ------=_Part_35_16588280.1321471117363 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: 7bit Hmm... doesn't seem to work for me.

  a  #=> [["This is an example.\n    a = 1\nOf what I mean.\n    b = 2\nAnd it can go on\nlike this.\n\n    c = 3\nFor ever and ever.\n\n    d = 4\n", "\n\n   "]]

You know what though. I did some benchmarking and discovered that manually parsing the text line by line is much faster than using a regular expression (I used a close approx re). I was kind of surprised by this, since the regular expression engine is written in C, where as my line by line parser is in Ruby.

Despite that, I still find it curious that there isn't a more obvious regular expression for parsing a document in this way. It makes me wonder if a C.S. PhD could go back to the drawing board, and come up with a better alternative to REs.

------=_Part_35_16588280.1321471117363-- ------=_Part_34_31021916.1321471117362--