From: Dick Davies Date: 2004-09-24T00:20:04+09:00 Subject: Re: mozilla bookmarks * Jamis Buck [0926 16:26]: > Dick Davies wrote: > >long shot but what the hell - don't suppose any of you good > >good people are sitting on a parser for Mozilla/Firefox bookmarks.html > >files, by any chance? > > Funny you should ask. :) I've had this for awhile, and I can't even > remember why I wrote it. It's pretty hacked together, and it's not a > true "parser" (I just search for certain patterns in the bookmark file) > and it is hardcoded (currently) for my own (obsolete) Phoenix bookmarks > file, but it should be fairly straightforward to modify for your own > purposes. > > Hope this is at least close to what you are looking for... :) Thanks a lot, it was handy to get a feel for it - I gave up on a parser too (I'd prefer not to require extra libs), and did a cutdown homegrown version in the end (I only need url, folder info and description myself) : ----------------------------------------------------------------- rasputin@lb:lib$ cat mozbooks.rb #!/usr/bin/env ruby # quick and dirty bookmarks.html parser - thanks to Jamis Buck for the 'folder state machine' idea class MozBooks # pull urls, descriptions and folder heirarchy info from mozilla/firefox bookmarks.html def self.parse(bm) folders = [] bm.each_line{ |l| folders.pop if l =~ /<\/dl>

/i # we just left a folder folders << $1 if l =~ /\s*

]+>(.*)<\/h3>/i # we just entered a folder puts "url = #{$1}, desc = #{$2}, folder = #{folders.join('/')}" if l =~ /a href="([^"]*)"[^>]+>([^<]+)