From: Marnen Laibow-Koser Date: 2009-11-14T13:33:20+09:00 Subject: Re: scrape html gives partial result Thufir wrote: [...] > In terms of design, I'm considering the pros/cons for breaking of the > html to another table. Perhaps the scraped data should be with the > html in its own table? I doubt I'd see a performance differential for > the amount of data I'll be working with, but does it matter whether a > relatively large text field, and perhaps about five string fields are > added onto existing table, or whether there's a 1:1 relation to > another table? It depends on the conceptual structure of the application. I doubt that performance would be enough of an issue to worry about. > Also, perhaps there would be a 1:many relation between > the raw html and scraped data, but I'm not sure about that. Depends on the data! > > Also, I don't want to accidentally re-fetch the html and end up with a > bunch of 404 error pages, so would it make sense to add a boolean > indicating whether html had been fetched? Or, just restrict fetching > (?) the html to when feeds are grabbed? Do you need a boolean? Just test whether the HTML field is null. > > > thanks, > > Thufir Best, -- Marnen Laibow-Koser http://www.marnen.org marnen@marnen.org -- Posted via http://www.ruby-forum.com/.