From: Horacio Sanson Date: 2005-11-28T22:40:57+09:00 Subject: Re: htmltokenizer bug? Well the problem is that this HTML is not mine, retrieving the pages from the Internet. Guess I will skip this page from my script. thanks, Horacio Monday 28 November 2005 21:52、Daniel Schierbeck さんは書きました: > Horacio Sanson wrote: > > I am using htmltokenizer to extract the links of some web pages, my > > script worked perfectly until I started to parse pages with "<" and ">" > > chars in the text. > > > > a html string like this > > > > this is a > > > > causes the htmlparser to raise and exception; Error, tag is nil.... > > > > > > Is there a patch or any way to make htmlparser to parse this text?? > > > > > > regards, > > Horacio > > Your HTML isn't valid. Either use the proper entities (< = < and > = > >) or make a CDATA section, though the latter isn't really that > well-supported in most browsers. > > ]]> > > > Cheers, > Daniel