From: Steven Jenkins Date: 2005-11-13T01:56:33+09:00 Subject: Re: DocBook to PDF Bob Hutchison wrote: > The typical XML way to do this is DocBook -> XSL:FO -> pdf because > you'll be able to use Norm Walsh's stuff. If you do this and you have > input documents bigger than, say, two pages, you'll be wanting the > *fastest* XML processors you can get. I did this a couple of years ago > and found libxml2 and libxslt2 the way to go. Getting a good XSL:FO > processor was quite a trick then, and I have not been keeping up to > date, so keep that in mind. The only free one that I could find was the > Apache FOP processor, and it wasn't all that good (incomplete). There > are a couple of commercial processors that are apparently very good, > but I wasn't willing to spend a couple hundred bucks on them. I eventually went with LaTeX because I'm not just trying to make marks on paper, but trying to make beautiful documents that measure up to high standards of typesetting. All the out-of-the-box DocBook/XSL stuff I tried produced ugly output. Maybe things have gotten better. The real appeal of LaTeX for me is that it operates on document objects at approximately the same level of abstraction as DocBook itself. It's fairly straightforward to translate between the two, and then use styles and macros to control the output formatting. > An alternative that is worked quite well was to use the SGML processors > for DocBook. > > As I said, I've not been keeping up, but apparently both the Apache XML > processors and Saxon have become much faster in the last couple of > years. Both of those are Java. Unfortunately I don't think Ruby has a > remote chance of being useful as an XML processor for this kind of > application until Ruby gets much faster. It depends on the application. I'm using a brute force REXML parser and it's plenty fast enough for what I need to do. My test data is about 120k (37 pages typeset), and is fairly complex structurally: nested sections, variablelists, EPS figures, etc. I can convert it on an old P3 in about 10 seconds. LaTeX is blazingly fast, so I can afford a little slowness upstream. > Putting Ruby into that pipeline for processing the stream sounds more > reasonable, but still, that stream is going to get *very* long once > you've got XSL:FO. > > If you really want to do this yourself brace yourself for a lot of > work. Maybe you should choose a subset of DocBook (isn't there a small > (ish) subset already defined?) You don't need a very big subset for many documents. Mine handles appendix article articleinfo biblioid blockquote caption colspec emphasis entry figure formalpara imagedata itemizedlist listitem mediaobject orderedlist para pubdate row section simpara table term tgroup thead title variablelist xref in about 350 lines of Ruby. It also handles profiles. Steve