From: Aredridel Date: 2003-04-27T07:27:54+09:00 Subject: Re: XML Parsing the Ruby way. The difficulties in mapping objects to XML is partly in XML's complexity: Most languages do not let you assign your objects an arbitrary ID and there's no sane way to distinguish attributes from other sub-nodes. The difficulty is particularly great in a non-validating parser: It doesn't know enough about the schema or even the DTD to map things automagically at all, or even semi-automagically. What I'd like to see happen is have a schema layer appear on top of REXML or other parsers that hides most or all of the XML-ness of the data from other objects. This helps from other sides as well, since uninvolved objects should know nothing of the storage format or method for encapsulation reasons. --------------------------------- In other thoughts, I'd be perfectly happy if the library provided a few "Stub" classes for tags that are not otherwise mapped. If the classes being mapped to accept any object as a sub-type, and the schema allows it, it would be fine to see generic "XMLTag" classes or something similar instantiated along with the mapped classes. It would certainly make handling all the cases much easier. Also, a library of classes that match standard DTD elements would be useful: I can see an "HTML" module forming that handles the xhtml namespace, with powerful classes for HTML DTD elements like paragraphs and rich text, images and control structures. I could see similar for Docbook, and perhaps the common elements from both descend from a more generic "document" module of classes. I've heard rumour of Java class libraries for manipulating XML-based standards this way, with a class per tag (or so) -- the problem is that the library is static, so every time the standard changed, the library had to be extended. Since Ruby is so much more dynamic, this should be almost no problem -- at the worst, just fall back to a "Tag" class and implement a tiny set of features. At best, nobody notices and life goes on. ------------------------------- On the note of a standard library, mapping various parts of RDF into the library would be useful: RDF:Bag and RDF:List map to hashes and/or arrays very nicely; RDF:about as an attribute is something that could be added to classes dynamically that would give them a static ID of sorts, without the maintenance nightmare of just a list of document-specific ID attributes that have to be kept unique within one domain, but once you consider more than one document, could very easily become non-unique. I imagine a library like this could be extremely powerful when combined with an object prevalence system; an entire website or even a more specific service (XML-RPC or SOAP based, or entirely new) could be stored in-core, but the library could export an XML-ish "view" of the object-space "model", and all that would have to be written is a set of "controllers". Writing an RSS feed would be terribly easy, assuming the site data was already parsed into Ruby objects somehow: All one would have to do is walk one's existing data structures and feeding appropriate data into an RSS object. If the data was loaded from XML+RDF files, an Array might already map to an RDF:List, so generating the RSS/RDF file gets even that much easier /and/ more powerful.