From: Adam Akhtar Date: 2009-04-21T21:53:18+09:00 Subject: Re: reformatting a text file that has some binary in it Im back again and pretty confused as to why my regexp still is overshooting the mark. I want my regexp /(20:)(.*?)(d\d+:complete.+?incomplete.+?ee)/ium to get everything between and including 20: and ee i.e. from the first line of the sample at the bottom of this message id want want this 20: €0テ ・aュリ:$ ゥD€・d8:completei0e10:downloadedi772e10:incompletei1ee but sometimes it overshoots and does something like this 20: €0テ ・aュリ:$ ゥD€・d8:completei0e10:downloadedi772e10:incompletei1ee20: 琮祺ア・・キ・G聊まd8:completei9e10:downloadedi2064e10:incompletei2ee and I cant figure out why? In my notepad plus editor i have it set to display line feeds and carriage returns. Soemtimes in the binary parts it displays an lf symbol. In binary does lf serve as a representation for a new line or it just used to represent data (bytes etc) - could it be that thats tripping up rubys regexp engine? I load the data text file like so data = File.open("text.txt", "rb").readlines Is there something im doing wrong? sample from the data text file 20: €0テ ・aュリ:$ ゥD€・d8:completei0e10:downloadedi772e10:incompletei1ee20: 琮祺ア・・キ・G聊まd8:completei9e10:downloadedi2064e10:incompletei2ee20: }tェスh>・モ送榎ラd8:completei4e10:downloadedi7724e10:incompletei5ee20: 架C ヒウJ<ィFラ0ノ犒Wd8:completei4e10:downloadedi632e10:incompletei2ee20: incノ・U]~鼡・`僘< d8:completei5e10:downloadedi536e10:incompletei0ee20: シルqナ!pォス-リタ58Td8:completei1e10:downloadedi520e10:incompletei0ee20: G*﨨ェ ・4T澀オソk澆d8:completei0e10:downloadedi1061e10:incompletei2ee20: Iネl談ォヲ7Z&レ・K゙建ノEd8:completei5e10:downloadedi268e10:incompletei0ee20: Smナソホ怯ソヒィh7r・醋d8:completei5e10:downloadedi798e10:incompletei0ee20: bj・]VF・w仭鱧卍ノd8:completei8e10:downloadedi523e10:incompletei11ee20: hヘ・ヤ=ゥソ敏陏「巌d8:completei0e10:downloadedi57e10:incompletei3ee20: mb‘