From: "Иван Бишевац" Date: 2012-09-09T21:31:07+09:00 Subject: Re: Parsing through downloaded html --f46d042dfff5ce525104c94403f1 Content-Type: text/plain; charset=UTF-8 require 'nokogiri' require 'spreadsheet' Spreadsheet.client_encoding = 'UTF-8' book = Spreadsheet::Workbook.new sheet1 = book.create_worksheet # Numbering is zero based. This means that first row is labeled 0, first column 0. row = 0 Dir.chdir("anattempt") Dir.glob['*.html'].each do |document| f = file.open(document) searchablefile = Nokogiri::HTML(f) # use at_xpath rather than xpath since first one method returns just 1 element, # but second method xpath returns array of all found records matching criteria var1 = searchablefile.at_xpath("your xpath here..") var2 = searchablefile.at_xpath("your xpath here..") # In first pass it saves data to first row, and two columns A and B. # Every nest pass increments row by 1, but columns are same A and B. sheet1[row, 0] = variabelebasedonaxpath.content shhet1[row, 1] = variabelebasedonaxpath.content #After saving data increment row position by 1 row += 1 end book.write 'htmltoexcel.xls' I didn't tested this, but if something goes wrong ask here. Also read http://nokogiri.org/tutorials for learning how to parse xml/html documents, that's short but useful resource. --f46d042dfff5ce525104c94403f1 Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable
require 'nokogiri'
require 'spreadsheet'

Spreadsheet.client_encoding =3D 'UTF-8'
=
book =3D Spreadsheet::Workbook.new
sheet1 =3D book.create_wo= rksheet

# Numbering is zero based. This means that first row is= labeled 0, first column 0.
row =3D 0

Di= r.chdir("anattempt")
Dir.glob['*.html'].each do= |document|
=C2=A0 f =3D file.open(document)
=C2=A0 searchablefile =3D N= okogiri::HTML(f)
=C2=A0=C2=A0
=C2=A0 # use at_xpath rat= her than xpath since first one method returns just 1 element,
=C2= =A0 # but second method xpath returns array of all found records matching c= riteria
=C2=A0 var1 =3D searchablefile.at_xpath("your xpath here..")=
=C2=A0 var2 =3D searchablefile.at_xpath("your xpath here..&= quot;)
=C2=A0=C2=A0
=C2=A0 # In first pass it saves dat= a to first row, and two columns A and B.
=C2=A0 # Every nest pass increments row by 1, but columns are same A a= nd B.
=C2=A0 sheet1[row, 0] =3D variabelebasedonaxpath.content
=C2=A0 shhet1[row, 1] =3D variabelebasedonaxpath.content
= =C2=A0=C2=A0
=C2=A0 #After saving data increment row position by = 1
=C2=A0 row +=3D 1
end

book.write &#= 39;htmltoexcel.xls'


I didn'= t tested this, but if something goes wrong ask here.
Also read=C2= =A0http://nokogiri.org/tutorials<= /a>=C2=A0for learning how to parse xml/html documents, that's short but= useful resource.
--f46d042dfff5ce525104c94403f1--