Ruby XML, XSLT and XPath Tutorial


What is XML?

XML stands for eXtensible Markup Language.

Extensible Markup Language, a subset of the Standard Generalized Markup Language, is a markup language used to mark electronic documents to give them structure.

It can be used to mark up data and define data types. It is a source language that allows users to define their own markup language. It is very suitable for World Wide Web transmission, providing a unified method to describe and exchange structured data that is independent of applications or vendors.

For more content, please see ourXML Tutorial


XML Parser Structure and API

There are mainly two types of XML parsers: DOM and SAX.

  • The SAX parser is based on event processing. It scans the XML document from beginning to end. During the scanning process, each time a syntactic structure is encountered, the event handler for that specific syntactic structure is called to send an event to the application.
  • DOM is Document Object Model parsing. It builds a hierarchical syntactic structure of the document and creates a DOM tree in memory. The nodes of the DOM tree are identified in the form of objects. After the document is parsed, the entire DOM tree of the document is placed in memory.

Parsing and Creating XML in Ruby

In Ruby, the REXML library can be used to parse XML documents.

The REXML library is an XML toolkit for Ruby. It is written in pure Ruby and complies with the XML 1.0 specification.

In Ruby 1.8 and later, REXML is included in the Ruby standard library.

The path of the REXML library is: rexml/document

All methods and classes are encapsulated in a REXML module.

The REXML parser has the following advantages over other parsers:

  • 100% written in Ruby.
  • Applicable to SAX and DOM parsers.
  • It is lightweight, with less than 2000 lines of code.
  • Easy-to-understand methods and classes.
  • Based on the SAX2 API and full XPath support.
  • Installed with Ruby, no separate installation required.

The following is the XML code for the example, saved as movies.xml:

<collection shelf="New Arrivals"> <movie title="Enemy Behind"> <type>War, Thriller</type> <format>DVD</format> <year>2003</year> <rating>PG</rating> <stars>10</stars> <description>Talk about a US-Japan war</description> </movie> <movie title="Transformers"> <type>Anime, Science Fiction</type> <format>DVD</format> <year>1989</year> <rating>R</rating> <stars>8</stars> <description>A schientific fiction</description> </movie> <movie title="Trigun"> <type>Anime, Action</type> <format>DVD</format> <episodes>4</episodes> <rating>PG</rating> <stars>10</stars> <description>Vash the Stampede!</description> </movie> <movie title="Ishtar"> <type>Comedy</type> <format>VHS</format> <rating>PG</rating> <stars>2</stars> <description>Viewable boredom</description> </movie> </collection>

DOM Parser

Let's first parse the XML data. First, we import the rexml/document library. Usually we can import REXML into the top-level namespace:

Example

#!/usr/bin/ruby -w require 'rexml/document' include REXML xmlfile = File.new("movies.xml") xmldoc = Document.new(xmlfile) #Get the root element root = xmldoc.root puts "Root element : " + root.attributes["shelf"] #The following will output movie titles xmldoc.elements.each("collection/movie"){ |e| puts "Movie Title : " + e.attributes["title"] } #The following will output all movie types xmldoc.elements.each("collection/movie/type") { |e| puts "Movie Type : " + e.text } #The following will output all movie descriptions xmldoc.elements.each("collection/movie/description") { |e| puts "Movie Description : " + e.text }

The output result of the above example is:

Root element : New Arrivals
Movie Title : Enemy Behind
Movie Title : Transformers
Movie Title : Trigun
Movie Title : Ishtar
Movie Type : War, Thriller
Movie Type : Anime, Science Fiction
Movie Type : Anime, Action
Movie Type : Comedy
Movie Description : Talk about a US-Japan war
Movie Description : A schientific fiction
Movie Description : Vash the Stampede!
Movie Description : Viewable boredom
SAX-like Parsing:

SAX Parser

To process the same data file: movies.xml, SAX parsing is not recommended for a small file. The following is a simple example:

Example

#!/usr/bin/ruby -w require 'rexml/document' require 'rexml/streamlistener' include REXML class MyListener include REXML::StreamListener def tag_start(*args) puts "tag_start: #{args.map {|x| x.inspect}.join(', ')}" end def text(data) return if data =~ /^\w*$/ # whitespace only abbrev = data[0..40] + (data.length > 40 ? "..." : "") puts " text : #{abbrev.inspect}" end end list = MyListener.new xmlfile = File.new("movies.xml") Document.parse_stream(xmlfile, list)

The output result of the above is:

tag_start: "collection", {"shelf"=>"New Arrivals"}
tag_start: "movie", {"title"=>"Enemy Behind"}
tag_start: "type", {}
  text   :   "War, Thriller"
tag_start: "format", {}
tag_start: "year", {}
tag_start: "rating", {}
tag_start: "stars", {}
tag_start: "description", {}
  text   :   "Talk about a US-Japan war"
tag_start: "movie", {"title"=>"Transformers"}
tag_start: "type", {}
  text   :   "Anime, Science Fiction"
tag_start: "format", {}
tag_start: "year", {}
tag_start: "rating", {}
tag_start: "stars", {}
tag_start: "description", {}
  text   :   "A schientific fiction"
tag_start: "movie", {"title"=>"Trigun"}
tag_start: "type", {}
  text   :   "Anime, Action"
tag_start: "format", {}
tag_start: "episodes", {}
tag_start: "rating", {}
tag_start: "stars", {}
tag_start: "description", {}
  text   :   "Vash the Stampede!"
tag_start: "movie", {"title"=>"Ishtar"}
tag_start: "type", {}
tag_start: "format", {}
tag_start: "rating", {}
tag_start: "stars", {}
tag_start: "description", {}
  text   :   "Viewable boredom"

XPath and Ruby

We can use XPath to view XML. XPath is a language for finding information in XML documents (see:XPath Tutorial)。

XPath is the XML Path Language. It is a language used to determine the position of a part in an XML (a subset of the Standard Generalized Markup Language) document. XPath is based on the tree structure of XML and provides the ability to find nodes in the data structure tree.

Ruby supports XPath through REXML's XPath class, which is based on tree analysis (Document Object Model).

Example

#!/usr/bin/ruby -w require 'rexml/document' include REXML xmlfile = File.new("movies.xml") xmldoc = Document.new(xmlfile) #The first movie's information movie = XPath.first(xmldoc, "//movie") p movie #Print all movie types XPath.each(xmldoc, "//type") { |e| puts e.text } #Get all movie format types and return an array names = XPath.match(xmldoc, "//format").map {|x| x.text } p names

The output result of the above example is:

<movie title='Enemy Behind'> ... </>
War, Thriller
Anime, Science Fiction
Anime, Action
Comedy
["DVD", "DVD", "DVD", "VHS"]

XSLT and Ruby

There are two XSLT parsers in Ruby. A brief description is given below:

Ruby-Sablotron

This parser was written and maintained by Masayoshi Takahash. It is mainly written for the Linux operating system and requires the following libraries:

  • Sablot
  • Iconv
  • Expat

You canRuby-Sablotronfind these libraries.

XSLT4R

XSLT4R was written by Michael Neumann. XSLT4R is used for simple command-line interaction and can be used by third-party applications to transform XML documents.

XSLT4R requires XMLScan to operate; XMLScan is included in the XSLT4R archive. It is a 100% Ruby module. These modules can be installed using the standard Ruby installation method (i.e., ruby install.rb).

The XSLT4R syntax format is as follows:

ruby xslt.rb stylesheet.xsl document.xml [arguments]

If you want to use XSLT4R in your application, you can include XSLT and input the parameters you need. The example is as follows:

Example

require "xslt" stylesheet = File.readlines("stylesheet.xsl").to_s xml_doc = File.readlines("document.xml").to_s arguments = { 'image_dir' => '/....' } sheet = XSLT::Stylesheet.new( stylesheet, arguments ) # output to StdOut sheet.apply( xml_doc ) # output to 'str' str = "" sheet.output = [ str ] sheet.apply( xml_doc )

More Resources

Other Extensions