From: Mike Dalessio Date: 2010-05-04T19:18:11+09:00 Subject: Re: Nokogiri bug or intended effect?? --00163630f4330e9cfa0485c206d2 Content-Type: text/plain; charset=ISO-8859-1 If you post this question to nokogiri-talk with a reproducible test case, I think you'll quickly get a response from the helpful nokogiri community. On May 3, 2010 6:45 PM, "Jeremy Woertink" wrote: I'm trying to parse this (poorly formatted) page, and when I look at the page I see: Name: ZITO, PEDRO OSVALDO When I look at the source I get: Name: ZITO, PEDRO OSVALDO When I parse the page I get: >> page.search("/html/body/table[3]/tr[1]/td[4]/table/tr[1]/td[1]/table/tr[3]/td[2]/table/tr[1]/td[1]/table/tr[2]/td[1]").first => #] children=[#, #] children=[#, #]>, #]>, #]> >> If you notice in the # All the spaces in the name have been removed. Here's what I'm using: >> Nokogiri::LIBXML_VERSION => "2.7.3" macbook-pro:~ jeremywoertink$ ruby -v ruby 1.8.6 (2009-06-08 patchlevel 369) [universal-darwin9.0] Anyone have any ideas? My guess is maybe an encoding issue??? There are other areas in the pages where I have to do string.gsub("\302\240", ""). >> page.meta_encoding => nil >> Thanks, ~Jeremy -- Posted via http://www.ruby-forum.com/. --00163630f4330e9cfa0485c206d2--