XML CDATA


All text in an XML document will be parsed by the parser.

Only text within a CDATA section will be ignored by the parser.


PCDATA - Parsed Character Data

XML parsers usually parse all text in an XML document.

When an XML element is parsed, the text between its tags is also parsed:

<message>This text is also parsed</message>

The parser does this because XML elements can contain other elements, as in this example, where the <name> element contains two other elements (first and last):

<name><first>Bill</first><last>Gates</last></name>

The parser breaks it down into sub-elements like this:

<name>
<first>Bill</first>
<last>Gates</last>
</name>

Parsed Character Data (PCDATA) is a term used for text data that is parsed by the XML parser.


CDATA - (Unparsed) Character Data

The term CDATA refers to text data that should not be parsed by the XML parser.

Characters such as "<" and "&" are illegal in XML elements.

"<" will cause an error because the parser interprets this character as the beginning of a new element.

"&" will cause an error because the parser interprets this character as the start of a character entity.

Some text, such as JavaScript code, contains a lot of "<" or "&" characters. To avoid errors, the script code can be defined as CDATA.

Everything inside a CDATA section is ignored by the parser.

CDATA section starts with "<![CDATA[and ends with "]]>:

<script>
<![CDATA[
function matchwo(a,b)
{
if (a < b && a < 0) then
{
return 1;
}
else
{
return 0;
}
}
]]>
</script>

In the example above, the parser ignores everything in the CDATA section.

Notes about CDATA sections:

A CDATA section cannot contain the string "]]>". Nested CDATA sections are also not allowed.

The "]]>" that marks the end of a CDATA section cannot contain spaces or line breaks.


Other Extensions