ansaurus

Question

Special characters in XML files - processing with the DOM API

Answer 1

+2 A:

Any XML parser will implicitly translate entities such as &, <, >, into the corresponding characters, as part of the process of parsing the file.

Alex Martelli 2009-05-16 08:23:26

Answer 2

+4 A:

The reason is simple: The XML file really contains an "&" character.

It is just represented differently (i.e. it is "escaped"), because a real "&" on it's own breaks XML files, as you've seen. Read the relevant section in the XML 1.0 spec: "2.4 Character Data and Markup". It's just a few lines, but it explains the issue quite well.

XML is a representation of data (!). Don't think of it as a text file. Example:

You want to store the string "17 < 20" in an XML file. Initially, you can't, since the "<" is reserved as the opening tag bracket. So this would be invalid:

<xml>17 < 20</xml>

Solution: You employ character escaping on the special/reserved character, just for the means of retaining the validity of the file:

<xml>17 &lt; 20</xml>

For all practical purposes the above snippet contains the following data (in JSON representation this time):

{
  "xml": "17 < 20"
}

This is why you see the real "&" in your post-processing. It had been escaped in just the same way, but it's meaning stayed the same all the time.

The above example also explains why the "&" must be treated specially: It is itself part of the XML escaping mechanism. It marks the start of an escape sequence, like in "<". Therefore it must be escaped itself (with "&", like you've done).

Tomalak 2009-05-16 08:24:43

Fabulous answer... as usual! +1

Cerebrus 2009-05-16 09:13:26

ansaurus

tags:

views:

answers:

Special characters in XML files - processing with the DOM API

related questions