ansaurus

Question

Answer 1

A:

It seems that you load a unicode encoded text file. 0 indicates Latin character.

If you don't want to deal with unicode text, choose ANSI encoding in your editor when you save the file.

If you need unicode encoding, use WideCharToString to convert it to an ANSI string, or just remove yourself the 0s, though the latter isn't the best solution. Also remove the 2 leading characters, ÿþ.
The editor put those bytes to mark the file as unicode.

Nick D 2010-07-01 07:42:49

0 indicates English language? Which value indicates Klingon language? :)

mjustin 2010-07-01 07:47:18

@Cosmin, thanks, I edited my answer.

Nick D 2010-07-01 08:12:50

0 indicates Latin? Si tacuisses ... :)

mjustin 2010-07-01 10:30:32

Answer 2

+8 A:

What you think is the text of the file isn't really the text of the file. What you've read into your string variable is accurate. You have a Unicode text file encoded as little-endian UTF-16. The first two bytes represent the byte-order mark, and each pair of bytes after that are another character of the string.

If you're reading a Unicode file, you should use a Unicode data type, such as WideString. You'll want to divide the file size by two when setting the length of the string, and you'll want to discard the first two bytes.

If you don't know what kind of file you're reading, then you need to read the first two or three bytes first. If the first two bytes are $ff $fe, as above, then you might have a little-endian UTF-16 file; read the rest of the file into a WideString, or UnicodeString if you have that type. If they're $fe $ff, then it might be big-endian; read the remainder of the file into a WideString and then swap the order of each pair of bytes. If the first two bytes are $ef $bb, then check the third byte. If it's $bf, then they are probably the UTF-8 byte-order mark. Discard all three and read the rest of the file into an AnsiString or an array of bytes, and then use a function like UTF8Decode to convert it into a WideString.

Once you have your data in a WideString, the debugger will show that it contains version, and you should have no trouble using a Unicode-enabled version of StringReplace to do your replacement.

Rob Kennedy 2010-07-01 07:44:43

2010-07-01 08:58:30

ah sorry, there's some trouble with my editor. I open the file with notepad and everything goes well!!.

2010-07-01 09:48:58

Evidently, the default encoding in Vista UTF-16. It's about time. If you really need a different encoding, use the "save as" dialog box and choose something different. Everything goes well when you open the file in Notepad because it uses the procedure I described in my answer. It's even a little more involved than that, since it considers ANSI encoding, too.

Rob Kennedy 2010-07-01 13:23:00

ansaurus

tags:

views:

answers:

Replace string that contain #0?

related questions