Python XML Parsing


What is XML?

XML refers to the Extensible Markup Language (eXtensible Markup Lxtensible Markup Language).XML Tutorial

XML is designed to transmit and store data.

XML is a set of rules for defining semantic markup that divides a document into many parts and identifies those parts.

It is also a meta-markup language, that is, a syntactic language used to define other semantic, structured markup languages related to specific fields.


Parsing XML with Python

Common XML programming interfaces include DOM and SAX. These two interfaces handle XML files differently, and of course they are used in different situations.

Python has three methods to parse XML: SAX, DOM, and ElementTree:

1.SAX (simple API for XML )

The Python standard library includes a SAX parser. SAX uses an event-driven model, processing XML files by triggering events one by one during XML parsing and calling user-defined callback functions.

2.DOM(Document Object Model)

Parse XML data into a tree in memory, and operate on the XML by manipulating the tree.

3. ElementTree (Element Tree)

ElementTree is like a lightweight DOM with a convenient and friendly API. It has good code usability, fast speed, and low memory consumption.

Note:Because DOM needs to map XML data to a tree in memory, it is relatively slow and memory-intensive. SAX reads XML files in a streaming manner, which is faster and uses less memory, but requires users to implement callback functions (handlers).

The contents of the XML example file movies.xml used in this chapter are as follows:

movies.xml

<collection shelf="New Arrivals"> <movie title="Enemy Behind"> <type>War, Thriller</type> <format>DVD</format> <year>2003</year> <rating>PG</rating> <stars>10</stars> <description>Talk about a US-Japan war</description> </movie> <movie title="Transformers"> <type>Anime, Science Fiction</type> <format>DVD</format> <year>1989</year> <rating>R</rating> <stars>8</stars> <description>A schientific fiction</description> </movie> <movie title="Trigun"> <type>Anime, Action</type> <format>DVD</format> <episodes>4</episodes> <rating>PG</rating> <stars>10</stars> <description>Vash the Stampede!</description> </movie> <movie title="Ishtar"> <type>Comedy</type> <format>VHS</format> <rating>PG</rating> <stars>2</stars> <description>Viewable boredom</description> </movie> </collection>

Python uses SAX to parse XML

SAX is an event-driven API.

Parsing XML documents with SAX involves two parts:ParserandEvent handler。

The parser is responsible for reading the XML document and sending events to the event handler, such as element start and element end events.

The event handler is responsible for responding to events and processing the passed XML data.

  • 1. Process large files;
  • 2. Only need part of the file content, or only need specific information from the file.
  • 3. When you want to build your own object model.

To handle XML with SAX in Python, you must first import the parse function from xml.sax and ContentHandler from xml.sax.handler.

Introduction to ContentHandler class methods

characters(content) method

Invocation timing:

From the beginning of the line, before encountering a tag, if there are characters, the value of content is these strings.

From one tag, before encountering the next tag, if there are characters, the value of content is these strings.

From one tag, before encountering the line terminator, if there are characters, the value of content is these strings.

The tag can be a start tag or an end tag.

startDocument() method

Called when the document starts.

endDocument() method

Called when the parser reaches the end of the document.

startElement(name, attrs) method

Called when an XML start tag is encountered. name is the tag name, and attrs is a dictionary of the tag's attribute values.

endElement(name) method

Called when an XML end tag is encountered.


make_parser method

The following method creates a new parser object and returns it.

xml.sax.make_parser( [parser_list] )

Parameter description:

  • parser_list- Optional parameter, parser list

parser method

The following method creates a SAX parser and parses an XML document:

xml.sax.parse( xmlfile, contenthandler[, errorhandler])

Parameter description:

  • xmlfile- XML file name
  • contenthandler- Must be a ContentHandler object
  • errorhandler- If this parameter is specified, errorhandler must be a SAX ErrorHandler object

parseString method

The parseString method creates an XML parser and parses an XML string:

xml.sax.parseString(xmlstring, contenthandler[, errorhandler])

Parameter description:

  • xmlstring- XML string
  • contenthandler- Must be a ContentHandler object
  • errorhandler- If this parameter is specified, errorhandler must be a SAX ErrorHandler object

Python XML Parsing Example

Example

#!/usr/bin/python # -*- coding: UTF-8 -*- import xml.sax class MovieHandler( xml.sax.ContentHandler ): def __init__(self): self.CurrentData = "" self.type = "" self.format = "" self.year = "" self.rating = "" self.stars = "" self.description = "" # Element start event handling def startElement(self, tag, attributes): self.CurrentData = tag if tag == "movie": print "*****Movie*****" title = attributes["title"] print "Title:", title # Element end event handling def endElement(self, tag): if self.CurrentData == "type": print "Type:", self.type elif self.CurrentData == "format": print "Format:", self.format elif self.CurrentData == "year": print "Year:", self.year elif self.CurrentData == "rating": print "Rating:", self.rating elif self.CurrentData == "stars": print "Stars:", self.stars elif self.CurrentData == "description": print "Description:", self.description self.CurrentData = "" # Content event handling def characters(self, content): if self.CurrentData == "type": self.type = content elif self.CurrentData == "format": self.format = content elif self.CurrentData == "year": self.year = content elif self.CurrentData == "rating": self.rating = content elif self.CurrentData == "stars": self.stars = content elif self.CurrentData == "description": self.description = content if ( __name__ == "__main__"): # Create an XMLReader parser = xml.sax.make_parser() # turn off namepsaces parser.setFeature(xml.sax.handler.feature_namespaces, 0) # Override ContextHandler Handler = MovieHandler() parser.setContentHandler( Handler ) parser.parse("movies.xml")

The execution result of the above code is as follows:

*****Movie*****
Title: Enemy Behind
Type: War, Thriller
Format: DVD
Year: 2003
Rating: PG
Stars: 10
Description: Talk about a US-Japan war
*****Movie*****
Title: Transformers
Type: Anime, Science Fiction
Format: DVD
Year: 1989
Rating: R
Stars: 8
Description: A schientific fiction
*****Movie*****
Title: Trigun
Type: Anime, Action
Format: DVD
Rating: PG
Stars: 10
Description: Vash the Stampede!
*****Movie*****
Title: Ishtar
Type: Comedy
Format: VHS
Rating: PG
Stars: 2
Description: Viewable boredom

Please refer to the complete SAX API documentationPython SAX APIs


Using xml.dom to parse XML

The Document Object Model (DOM) is a standard programming interface recommended by the W3C for processing Extensible Markup Language.

When a DOM parser parses an XML document, it reads the entire document at once and stores all elements in a tree structure in memory. Afterwards, you can use different functions provided by DOM to read or modify the content and structure of the document, and you can also write the modified content to an XML file.

In Python, xml.dom.minidom is used to parse XML files. An example is as follows:

Example

#!/usr/bin/python # -*- coding: UTF-8 -*- from xml.dom.minidom import parse import xml.dom.minidom # Open the XML document using the minidom parser DOMTree = xml.dom.minidom.parse("movies.xml") collection = DOMTree.documentElement if collection.hasAttribute("shelf"): print "Root element : %s" % collection.getAttribute("shelf") # Get all movies in the collection movies = collection.getElementsByTagName("movie") # Print detailed information for each movie for movie in movies: print "*****Movie*****" if movie.hasAttribute("title"): print "Title: %s" % movie.getAttribute("title") type = movie.getElementsByTagName('type')[0] print "Type: %s" % type.childNodes[0].data format = movie.getElementsByTagName('format')[0] print "Format: %s" % format.childNodes[0].data rating = movie.getElementsByTagName('rating')[0] print "Rating: %s" % rating.childNodes[0].data description = movie.getElementsByTagName('description')[0] print "Description: %s" % description.childNodes[0].data

The execution result of the above program is as follows:

Root element : New Arrivals
*****Movie*****
Title: Enemy Behind
Type: War, Thriller
Format: DVD
Rating: PG
Description: Talk about a US-Japan war
*****Movie*****
Title: Transformers
Type: Anime, Science Fiction
Format: DVD
Rating: R
Description: A schientific fiction
*****Movie*****
Title: Trigun
Type: Anime, Action
Format: DVD
Rating: PG
Description: Vash the Stampede!
*****Movie*****
Title: Ishtar
Type: Comedy
Format: VHS
Rating: PG
Description: Viewable boredom

Please refer to the complete DOM API documentationPython DOM APIs。

Other extensions