From: "Victor \"Zverok\" Shepelev" Date: 2006-12-01T06:39:47+09:00 Subject: Re: HTML cleanup task From: Paul Lutus [mailto:nospam@nosite.zzz] Sent: Thursday, November 30, 2006 11:00 PM >Victor "Zverok" Shepelev wrote: > >> It is a task definition. >> >> The task may vary for different dictionaries. For ex., with some >> dictionaries tables must not be deleted, but "normalized": >> "text1text2" => "
text1text2
" > >Both the before and after forms show big syntax errors. I hope you >understand HTML syntax, if not, this may be more difficult than I thought. I understand HTML syntax. And I see no problem in above. Closing tags for and are both optional in HTML 4.01 w3c spec. >Perhaps you could post what you consider to be the desired end result for a >particular entry from the "dictionary" site of your choice. OK. Here it is: Source page: http://en.wikipedia.org/wiki/Ukraine Start pattern: End pattern:

Elements to exclude: tables, images. Desired output (with text in middle of paragraph skipped): ---------------------------

Ukraine (Ukrainian: Україна, Ukraina, /ukraˈjina/) is a country in Eastern Europe. .... It became independent again after the Soviet Union's collapse in 1991.

--------------------------- That's all. >By the way (my boilerplate remark about page scraping), if this is for any >purpose other than your own personal use, it represents a copyright >problem. My application would be kinda browser (nano-browser), I don't want to "grab" dictionaries. >I want to emphasize this is not difficult at all, once there is a clear >statement of purpose. In can be done in a few (maybe a few dozen) lines of >Ruby code. I know. I'm not a nuby (my poor language in mails is due to natural language problems, not very low knowledge). I've just asked about existing libraries. > >-- >Paul Lutus >http://www.arachnoid.com V.