ansaurus

Question

How to determine if a String contains invalid encoded characters

Answer 1

+1 A:

URLDecoder will decode to a given encoding. This should flag errors appropriately. However the documentation states:

There are two possible ways in which this decoder could deal with illegal strings. It could either leave illegal characters alone or it could throw an IllegalArgumentException. Which approach the decoder takes is left to the implementation.

So you should probably try it. Note also (from the decode() method documentation):

The World Wide Web Consortium Recommendation states that UTF-8 should be used. Not doing so may introduce incompatibilites

so there's something else to think about!

EDIT: Apache Commons URLDecode claims to throw appropriate exceptions for bad encodings.

Brian Agnew 2009-05-20 10:18:38

I know of the Recommendation, but what about the browser (here Firefox 3.0.10) violating it? As long as it is recommended and not required you have to make sure that there are no illegal entities, don't you?

dhiller 2009-05-20 10:24:30

So I would try decoding using the URLDecoder and choosing the appropriate encoding. I would be interested (!) to see if the URLDecoder *does* throw exceptions on illegally encoded characters (easy to test outside the browser/server environment)

Brian Agnew 2009-05-20 11:56:06

Sorry. Just saw your edited question re. illegal chars

Brian Agnew 2009-05-20 11:56:36

Answer 2

A:

daniel 2009-09-18 20:56:36

string.getBytes() with new String() is a classic bug which should be avoid

Dennis Cheung 2009-09-24 16:57:37

Answer 3

+6 A:

I asked the same question,

http://stackoverflow.com/questions/1233076/handling-character-encoding-in-uri-on-tomcat

I recently found a solution and it works pretty well for me. You might want give it a try. Here is what you need to do,

Leave your URI encoding as Latin-1. On Tomcat, add URIEncoding="ISO-8859-1" to the Connector in server.xml.
If you have to manually URL decode, use Latin1 as charset also.
Use the fixEncoding() function to fix up encodings.

For example, to get a parameter from query string,

  String name = fixEncoding(request.getParameter("name"));

You can do this always. String with correct encoding is not changed.

The code is attached. Good luck!

 public static String fixEncoding(String latin1) {
  try {
   byte[] bytes = latin1.getBytes("ISO-8859-1");
   if (!validUTF8(bytes))
    return latin1;   
   return new String(bytes, "UTF-8");  
  } catch (UnsupportedEncodingException e) {
   // Impossible, throw unchecked
   throw new IllegalStateException("No Latin1 or UTF-8: " + e.getMessage());
  }

 }

 public static boolean validUTF8(byte[] input) {
  int i = 0;
  // Check for BOM
  if (input.length >= 3 && (input[0] & 0xFF) == 0xEF
    && (input[1] & 0xFF) == 0xBB & (input[2] & 0xFF) == 0xBF) {
   i = 3;
  }

  int end;
  for (int j = input.length; i < j; ++i) {
   int octet = input[i];
   if ((octet & 0x80) == 0) {
    continue; // ASCII
   }

   // Check for UTF-8 leading byte
   if ((octet & 0xE0) == 0xC0) {
    end = i + 1;
   } else if ((octet & 0xF0) == 0xE0) {
    end = i + 2;
   } else if ((octet & 0xF8) == 0xF0) {
    end = i + 3;
   } else {
    // Java only supports BMP so 3 is max
    return false;
   }

   while (i < end) {
    i++;
    octet = input[i];
    if ((octet & 0xC0) != 0x80) {
     // Not a valid trailing byte
     return false;
    }
   }
  }
  return true;
 }

EDIT: Your approach doesn't work for various reasons. When there are encoding errors, you can't count on what you are getting from Tomcat. Sometimes you get � or ?. Other times, you wouldn't get anything, getParameter() returns null. Say you can check for "?", what happens your query string contains valid "?" ?

Besides, you shouldn't reject any request. This is not your user's fault. As I mentioned in my original question, browser may encode URL in either UTF-8 or Latin-1. User has no control. You need to accept both. Changing your servlet to Latin-1 will preserve all the characters, even if they are wrong, to give us a chance to fix it up or to throw it away.

The solution I posted here is not perfect but it's the best one we found so far.

ZZ Coder 2009-09-19 04:18:00

Nice one! But I have to object to your comment "Java only supports BMP". The four-byte limit on UTF-8 byte sequences was imposed by the Unicode Consortium, and it's sufficient to handle the complete range of characters (U+0000..U+10FFFF), not just the BMP.

Alan Moore 2009-09-19 23:31:00

The correct comment probably should be "We only care about BMP". My impression was that surrogate pair doesn't work well in Java.

ZZ Coder 2009-09-19 23:53:09

Well, I asked in May ;-) Anyway, what does the above code do? Does it convert from iso to utf-8? I would not want to convert the code, just check wether the encoding is right and throw an error if it's not. Please see my solution above again and check if it's correct, will you?

dhiller 2009-09-21 15:14:14

Your solution is not going to work. If wrong encoding is used, you will get question marks, instead of exception. Just use my function validUTF8(). If it's true, it's MOST LIKELY is UTF8. Otherwise, it's Latin-1. You have to use Latin-1 encoding everywhere in the server for this check to work.

ZZ Coder 2009-09-21 15:35:44

Yes, as I stated : 1. check if character.getBytes()[0] equals 63 for '?', 2. check if Character.getType(character.charAt(0)) returns OTHER_SYMBOL. And this _does_ work for me. If you can prove the opposite, please let me know...

dhiller 2009-09-22 05:46:16

See my edit .................

ZZ Coder 2009-09-22 12:35:39

@ZZ Coder: your code correctly detects four-byte UTF-8 sequences, which is the maximum allowed by the Unicode spec, so that comment doesn't really make sense. When the text is converted to Java strings, those four-byte sequences will become surrogate pairs, which Java handles correctly--just not transparently.

Alan Moore 2009-09-23 03:11:19

@ZZ Coder: At first thank you for your time. There seems to have been some misunderstanding because of my inprecise question, which I've tried to clarify. Please see my edits. At second: I disagree with your "you shouldn't reject any..." proposal, because we are on interface level. I have to make sure that the service user always uses the correct encoding. If my solution is wrong, how else can I achieve that?

dhiller 2009-09-23 05:32:32

@ZZ Coder: Could you please add some comments to your code to help me understand what you are doing?

dhiller 2009-09-23 05:34:15

Answer 4

+1 A:

I've been working on a similar "guess the encoding" problem. The best solution involves knowing the encoding. Barring that, you can make educated guesses to distinguish between UTF-8 and ISO-8859-1.

To answer the general question of how to detect if a string is properly encoded UTF-8, you can verify the following things:

No byte is 0x00, 0xC0, 0xC1, or in the range 0xF5-0xFF.
Tail bytes (0x80-0xBF) are always preceded by a head byte 0xC2-0xF4 or another tail byte.
Head bytes should correctly predict the number of tail bytes (e.g., any byte in 0xC2-0xDF should be followed by exactly one byte in the range 0x80-0xBF).

If a string passes all those tests, then it's interpretable as valid UTF-8. That doesn't guarantee that it is UTF-8, but it's a good predictor.

Legal input in ISO-8859-1 will likely have no control characters (0x00-0x1F and 0x80-0x9F) other than line separators. Looks like 0x7F isn't defined in ISO-8859-1 either.

(I'm basing this off of Wikipedia pages for UTF-8 and ISO-8859-1.)

Adrian McCarthy 2009-09-23 17:51:50

Answer 5

A:

the following regular expression might be of interest for you:

http://blade.nagaokaut.ac.jp/cgi-bin/scat.rb/ruby/ruby-talk/185624

I use it in ruby as following:

module Encoding
UTF8RGX = /\A(
    [\x09\x0A\x0D\x20-\x7E]            # ASCII
  | [\xC2-\xDF][\x80-\xBF]             # non-overlong 2-byte
  |  \xE0[\xA0-\xBF][\x80-\xBF]        # excluding overlongs
  | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2}  # straight 3-byte
  |  \xED[\x80-\x9F][\x80-\xBF]        # excluding surrogates
  |  \xF0[\x90-\xBF][\x80-\xBF]{2}     # planes 1-3
  | [\xF1-\xF3][\x80-\xBF]{3}          # planes 4-15
  |  \xF4[\x80-\x8F][\x80-\xBF]{2}     # plane 16
)*\z/x unless defined? UTF8RGX

def self.utf8_file?(fileName)
  count = 0
  File.open("#{fileName}").each do |l|
    count += 1
    unless utf8_string?(l)
      puts count.to_s + ": " + l
    end
  end
  return true
end

def self.utf8_string?(a_string)
  UTF8RGX === a_string
end

end

dimus 2009-09-23 17:58:22

Answer 6

+3 A:

You can use a CharsetDecoder configured to throw an exception if invalid chars are found:

 CharsetDecoder UTF8Decoder =
      Charset.forName("UTF8").newDecoder().onMalformedInput(CodingErrorAction.REPORT);

See CodingErrorAction.REPORT

ante 2009-09-24 11:51:15

Answer 7

A:

Try to use UTF-8 as a default as always in anywhere you can touch. (Database, memory, and UI)

One and single charset encoding could reduce a lot of problems, and actually it can speed up your web server performance. There are so many processing power and memory wasted to encoding/decoding.

Dennis Cheung 2009-09-24 16:55:29

Answer 8

A:

You might want to include a known parameter in your requests, e.g. "...&encTest=ä€", to safely differentiate between the different encodings.

mfx 2009-09-25 16:03:52

ansaurus

tags:

views:

answers:

How to determine if a String contains invalid encoded characters

related questions