From: Bryan Murphy Date: 2002-02-01T17:13:29+09:00 Subject: Ruby + XML Proposal The following is a sample application that will be included with the next revision of the Ruby Publication Framework (which will hopefully be released sometime early next week): require 'rexml/document' require 'RPF/console' serializer = Framework::Serializer::REXMLSerializer.new() parser = Framework::XML::FrameworkParser.new(serializer) parser.register(Framework::Transformer::XIncludeTransformer.new()) parser.parse(File.new('default.xml')) doc = serializer.document # The rest of this is 100% REXML code print "doc is an instance of: #{doc.class}\n\n" doc.write $stdout print "\n" This application loads an XML file (default.xml) from disk, and then builds a REXML DOM object from a SAX2-like stream events. In and of itself, this isn't a very interesting way to create a REXML document. However, in the middle of the application is the following line: parser.register(Framework::Transformer::XIncludeTransformer.new()) What this line does is it registers a class that implements the W3C XInclude specification. This class will receive ALL events that exist in the XInclude namespace, load any specified XML documents and insert their content into the middle of the XML stream! So, what I've effectively done is added XInclude support into REXML. But the fun doesn't stop there! Because of the way the framework is built, these events are fired off to components based upon their namespaces (if they don't exist in any particular namespace, they are sent directly to the Serializer which sits at the end of the chain). Because each event goes back to the same dispatcher, if XML document A includes XML document B, and then XML document B includes XML document C, then XML document A will be able to include XML document C as part of XML document B without you having to do any extra work on your part (as long as they all agree on the same include standard and you register the transformer that implements it). I can register as many of these transformers as I feel necessary. For instance, what if I want to include a SQL statement in my XML document, execute the SQL statement, and transform it's results into part of the XML document? Just write a new transformer and register it. What if I want a looping construct to repeat a part of the XML document a number of times? Create another transformer. Or how about an XSL transformer that transforms only a part of the XML document, creates an XInclude tag, and then the XInclude tag gets expanded out to contain the contents of the included document? Very easy to implement! What the framework essentially gives you is the ability to create smart tags for your XML documents that behave a lot like Java Tag Libraries, but are completely seperate from web applications!! But there's more: The framework also has the ability to get the input stream of events from any number of sources. Maybe you don't want to go through the trouble of creating a component that implements a custom set of tags, and want to generate the XML document a bit more programmaticaly. With the framework, you can swap out the default generator which reads XML documents from disk for another one. The first alternate generator already built into the framework uses ERuby to generate XML documents. You could put all your logic into an ERuby XML document, yet still use the framework to implement XInclude support (or any other transformers which may be available): generator = Framework::Generator::ERubyGenerator.new() generator.source = "default.erb" parser.generator = generator Or maybe you have a REXML document already loaded in memory and want to run that through the processor. Easy enough, use the included REXMLGenerator which takes a REXML DOM document and converts it into a stream of events: generator = Framework::Generator::REXMLGenerator.new() generator.source = my_rexml_obj parser.generator = generator The RPF also includes a few other nifty abilities. One of those is the ability to explicitly apply Transformers that capture *ALL* streamed events. A stylesheet for the whole XML document is an example of when you'd need one. You can have an XML document on disk, and then use the built in SablotronTransformer to apply a stylesheet to the XML document before you load it into the REXML dom: xsl = Framework::Transformer::SablotronTransformer.new() xsl.source = 'default.xsl' parser.addTransformer(xsl) These transformers are applied in the order they are added (you can add as many as you want), and you can control whether the namespace based transformers are dispatched inbetween them or not. The best part about this is that since the dispatcher only dispatches the events that a transformer will recognize, this greatly simplifies transformer development! In fact, the XInclude transformer (which admittedly isn't 100% compliant) is only about 20 lines of code right now, and most of that is for loading the XML document and creating a new XML parser to originate the new event stream. Finally, the last part the framework does is it adds infrastructure for building dynamic Web applications using this stuff by changing the pipeline on the fly via an XML base sitemap. This of course, a discussion completely on it's own. ;) Now, you ask, why am I posting this instead of just releasing the newest code? Well, I'm at a bit of a cross-roads. The REXML stuff is a good example. The next version of the framework will have integrated support for cooperating with REXML, but I want to go a step further. I want to not just cooperate with REXML, but *ALL* other Ruby XML stuff out there (NQXML, XMLParser, XSLT4R, you name it). There are two ways to accomplish this, the quick easy short term method: build the support into the framework. And the harder but more rewarding long term method: build the support directly into the corresponding Ruby libraries. An ideal example would be to abstract away the REXML(Generator|Serializer) and replace it with a Stream(Generator|Serializer). The stream based components would work explicitly with the SAX2-like stream of events. The code that generates the stream from a REXML document would then be moved into REXML, and the code which generates a REXML document from a stream would be moved into REXML as well. All other libraries could implement similar code, and then plug in directly with the Stream components in the framework. Overtime, the actual implementations themselves could become less hacky and use real stream based implementations for additional speed benefits. A side effect is that you would then be able to connect a REXML stream generator to an NQXML document creator and convert a REXML document into an NQXML document (and vice versa) as efficiently as possible and in a standardized way (or do the same with any other conformant library). I can provide the initial code for the various XML libraries (and will have the starting REXML code ready early next week). But to really do this well, we all need to agree on a streaming model for Ruby. Below is my proposal on what this model should be. This is more or less how my framework is implemented at the moment. Below it is some commentary on why certain things are as they are. I'd like for us to come to an agreement (if at all possible) before I get any futher into building the RPF framework, as the more code I write, the harder it will be to implement any major changes: module DocumentHandler # Contains unencoded character content (needs to be encoded when inserted into # the XML stream so it doesn't screw up the final XML document). ch is a # string, and start and length are used so that ch can actually be a substring # of a larger string (avoids some uneccessary string copying). def characters(ch, start, length) # Similar to characters, however contains the contents of a comment *excluding* # the end tags. def comment(ch, start, length) # Notifies the end of an XML stream def endDocument() # Notifies the end of an XML element. URI is the namespace that the element # exists in (nil if there is none). localName is the name of the XML tag. # qName is the fully qualified name (i.e. if there was a namespace qualifier, # it would contain "qualifier:#{localName}" otherwise it would be the # "#{localName}"). def endElement(uri, localName, qName) # Passes ignorable whitespace the same way characters are passed. def ignorableWhitespace(ch, start, length) # Passes a processing instruction (excluding the end tags) def processingInstruction(target, data) # This I'm open for debate, I never liked this anyway. Passes an object that # contains a means to get information about the current position within the # current XML document). def setDocumentLocator(locator) # returns the name of an entity that the previous XML stream couldn't # figure out. For instance, if the XML document contained an   # element, this would receive a string containg: 'nbsp' def skippedEntity(name) # Notifies the beginning of an XML stream. def startDocument() # Notifies the start of an XML document. This *INCLUDES* namespace entries. # All the params are the same as the endElement tag, and attrs is a Hash that # contains all the attibutes (including any namespace entries) for this tag. def startElement(uri, localName, qName, attrs) end Also, do we need to standardize on the ErrorHandler interface as well, or can we simply rely on exceptions? (Personally I'd choose exceptions). The major differences between this model and the XMLParser streaming model is that this model formalizes namespace support, and greatly simplifies working with namespaces. The differences between this model and the SAX2 model are: 1. I've removed the startPrefixMapping and endPrefixMapping events, and moved namespace declarations back into the attributes where they are in the XML document. For the life of me I can't figure out why SAX2 implemented them this way. The RPF framework originally used them, but I found it much simpler to move them into the attrs hash. 2. Attributes are a hash, and not an object (though this is, like everything else, open for discussion). I'm of the opinion that if you need to worry about attribute namespaces, you should check the attribute names explicitly using regexps. There are definitely good reasons to use a programmatic interface, however. Perhaps we can have the best of both worlds by using a hash like object with extra functionality (Ruby seems to be good at giving us the best of both worlds). Finally, the differences between this and other currently available streaming models: 1. Namespaces are supported. 2. We'll soon have a lot of infrastructure to coincide with this format (if you use my stuff anyway ;) 3. A standard format that crosses XML library boundaries will give us a lot of flexibility we just don't have right now. That's my proposal. I'd like to hear your comments on it. I'll donate any code I create (and some of my time) that can be used to help other projects along. If we can come up with a standard, I'll create a writeup for the standard and add it to the Ruby Guarden Wiki. Finally, if you think all of this is just a load of hot air, I want you to think about how you use IO streams. What do you do if you need to get at the contents of a file that are compressed? You do the following: contents = TarArchive.new(File.new('archive.tar.gz')).read() puts contents This works because we all agree on what an IO stream is supposed to look like. We can have this same kind of synergy when working with XML, we just need to come up with a standard and stick with it ;) Thanks for reading all this! Bryan