From: Jeremy Woertink Date: 2009-11-24T15:12:48+09:00 Subject: Re: Mechanize Help John W Higgins wrote: > Morning Jeremy, > > On Mon, Nov 23, 2009 at 9:21 AM, Jeremy Woertink > wrote: > >> >> I'm getting an encoding error when writing out the contents of this >> page. The content-encoding is showing gzip. Anyone know a way I can tell >> mechanize to use a different encoding when parsing a page? Or possibly >> another way I can do this? >> >> > You are getting the encoding error because you aren't dealing with a > string > in this case but rather a gzipped string. You could technically go to > the > request object and tell it you won't accept gzip encoded responses - but > I > personally find that distasteful because gzipping the pages saves > everyone > bandwidth and if we're scraping data we shouldn't be a nuisance (IMO). > What > you want to do is use the Zlib::Inflate class to convert the response to > a > regular string ( > http://www.ruby-doc.org/stdlib/libdoc/zlib/rdoc/classes/Zlib/Inflate.html#M001974). > That should solve your problem. > > John Hmmm, interesting. I didn't think about that. I'm not familiar with the class though. I ran through it and I got an error: Zlib::DataError: incorrect header check from (irb):84:in `inflate' from (irb):84 from :0 I did notice when I print out the string, there are a lot of "\240" in the string like "
DIV\240id\240\...."

I think these are where I'm getting messed up. Does anyone know a good 
site that lists these characters? I think \240 might be a tab character, 
but I want to check it against some list just to see.

Thanks,

~Jeremy


-- 
Posted via http://www.ruby-forum.com/.