From: Paul Lutus Date: 2006-12-10T09:40:04+09:00 Subject: Re: text wrap ishamid wrote: / ... > Perhaps my case is different. I am converting OOo xml to TeX for > further processing. Do you mean an OpenOffice content.xml file, unpacked from a .odt document? In that case, why not just insert some conveniently placed linefeeds to put the tags on separate lines? It happens I do this regularly, because I am constantly playing with OpenOffice XML content. Here is my script to beautify content.xml files: ------------------------------------------- #!/usr/bin/ruby -w # beuatifies an XML file, usually only for analysis, # because splitting one up and indenting its lines may # make it unusable to the originating program. def beautifyXML(data) tab = 0 xml = "" data.gsub!(%r{<},"\n<") data.gsub!(%r{>},">\n") data.gsub!(%r{\n+},"\n") data.split("\n").each { |record| record.strip! outc = record.scan(%r{()}).length inc = record.scan(%r{<\w}).length net = inc - outc tab += (net < 0)?net:0 xml += (" " * tab) + record + "\n" tab += (net > 0)?net:0 } if(tab != 0) $stderr.puts "Error: tag mismatch: #{tab}" end xml end # stream in/stream out print beautifyXML(readlines.join("\n")) ------------------------------------------- I want to emphasize this is a quickie script, not a thoroughly tested application, and its sole purpose is to aid in understanding the syntax of an XML file. I should also say that it may not work if the output is put back into an OpenOffice document after beautification. To use the script: $ (script name) < content.xml > output.xml This idea of beautifying an XML file so it is easier to understand and process is quite different than your original inquiry. In many cases, XML files are perfectly fine after they have been indented and made readable, and nothing is lost. But ... between some of the XML tag pairs emitted by this script may be long, long lines of text. Those lines should not be broken up without first considering the implications (the essence of my first post). Also, one defect in the above script is that the first line of the script, the XML identifying header, is not on the first line. This is something I don't care about, because I don't try to use the output of this script for anything but analysis. This error should easy to fix if you plan to use the output as legitimate XML. -- Paul Lutus http://www.arachnoid.com