From: Peter Szinek Date: 2007-04-20T06:30:22+09:00 Subject: Re: [ANN] scRUBYt! 0.2.8 Pit Capitain wrote: > Peter Szinek schrieb: >> As of 0.2.8, scRUBYt! depends on ParseTree and Ruby2Ruby - >> unfortunately (...) > > Peter, I'm curious: could you tell me (only in a few words) how you are > using ParseTree and Ruby2Ruby? Well, I guess it's the best to illustrate it with an example: This is a learning extractor: ========================================================================= google_data = Scrubyt::Extractor.define do fetch 'http://www.google.com/ncr' fill_textfield 'q', 'ruby' submit link "Ruby Programming Language" do url "href", :type => :attribute end next_page "Next", :limit => 3 end ========================================================================= i.e. it works for just the first page of google, for the query 'ruby'. To create a generalized extractor which can be used on any google result page, we have to export it after it 'learned' how to do this on the given example. Since I would like this so called production extractor to resemble the original as much as possible, I am using Ruby2Ruby and ParseTree. With them I can get this result: ========================================================================= google_data = Scrubyt::Extractor.define do fetch("http://www.google.com/ncr") fill_textfield("q", "ruby") submit link("/html/body/div/div/a") do url("href", { :type => :attribute }) end next_page("Next", { :limit => 3 }) end ========================================================================= If you can tell me any other way to achieve this (originally I took the source code of the learning extractors and replaced the examples with XPaths and did any other modifications required - but it was a mess after some time) I would be really thankful. There is a disagreement in the development team about this, too - one viewpoint is that a dependency on parsetree and ruby2ruby costs us too much (mainly on windows), and the other is that this problem has to be solved then we can depend on these packages.... What would you suggest? Cheers, Peter __ http://www.rubyrailways.com :: Ruby and Web2.0 blog http://scrubyt.org :: Ruby web scraping framework http://rubykitchensink.ca/ :: The indexed archive of all things Ruby.