From: Steve Callaway Date: 2005-12-07T21:34:41+09:00 Subject: Re: Screen scraping an html text contents into a file Ah, yeah, forgot all about those nasty little things. Not insuperable but would certainly add an overhead to handle them effectively. Steve --- Brian Schr�der wrote: > On 07/12/05, Steve Callaway > wrote: > > > > > > --- Martin DeMello > wrote: > > > > > Steve Callaway wrote: > > > > I will throw something like this together in > Ruby > > > over > > > > the next days when I get some time and post it > on > > > > RubyForge. I have already done this sort of > stuff > > > in > > > > Java and the concepts just really need a > port.All > > > we > > > > are looking at Basi's initial level of > > > requirements is > > > > to send an HTTP get, and pipe the response to > a > > > file. > > > > > > Nope, according to the OP's requirements, you > also > > > need to render the > > > html and spit out the rendered version as text, > > > which makes lynx --dump > > > the right tool for the job. It'd be quite a big > task > > > to duplicate this > > > in ruby, I think. > > > > > > martin > > > > > > > > > > By rendering the html, my interpretation of this > was > > that it is merely a question of stripping tags > etc, > > which can quickly be accomplished with gsub. Or am > I > > missing something? > > > > rgds > > > > Steve > > > > E.g. Tables and frames. So better use links2 or w3m > for the task. > > cheers, > > Brian > > > > -- > http://ruby.brian-schroeder.de/ > > Stringed instrument chords: > http://chordlist.brian-schroeder.de/ > > __________________________________________ Yahoo! DSL � Something to write home about. Just $16.99/mo. or less. dsl.yahoo.com