ansaurus

Question

Easy Q: UnicodeEncodeError: 'ascii' codec can't encode character

Answer 1

+2 A:

You might find this article useful. It represents the state of Unicode in Python 2.x, though. In Python 3.x all strings are Unicode, and there is a new bytes type for arbitrary 8-bit bytestrings.

Daniel Pryden 2009-10-31 00:13:58

Answer 2

+1 A:

Roger Pate 2009-10-31 00:31:55

updated info above

KenBurnsFan1 2009-10-31 05:12:13

Answer 3

+6 A:

You're trying to convert unicode to ascii in "strict" mode:

>>> help(str.encode)
Help on method_descriptor:

encode(...)
    S.encode([encoding[,errors]]) -> object

    Encodes S using the codec registered for encoding. encoding defaults
    to the default encoding. errors may be given to set a different error
    handling scheme. Default is 'strict' meaning that encoding errors raise
    a UnicodeEncodeError. Other possible values are 'ignore', 'replace' and
    'xmlcharrefreplace' as well as any other name registered with
    codecs.register_error that is able to handle UnicodeEncodeErrors.

You probably want something like one of the following:

s = u'Protection™'

print s.encode('ascii', 'ignore')    # removes the ™
print s.encode('ascii', 'replace')   # replaces with ?
print s.encode('ascii','xmlcharrefreplace') # turn into xml entities
print s.encode('ascii', 'strict')    # throw UnicodeEncodeErrors

Seth 2009-10-31 00:58:40

thanks for the effort -- I updated my question and will try to make it work with your info. -KBF1

KenBurnsFan1 2009-10-31 05:19:16

Answer 4

+2 A:

You're trying to pass a bytestring to something, but it's impossible (from the scarcity of info you provide) to tell what you're trying to pass it to. You start with a Unicode string that cannot be encoded as ASCII (the default codec), so, you'll have to encode by some different codec (or transliterate it, as @R.Pate suggests) -- but it's impossible for use to say what codec you should use, because we don't know what you're passing the bytestring and therefore don't know what that unknown subsystem is going to be able to accept and process correctly in terms of codecs.

In such total darkness as you leave us in, utf-8 is a reasonable blind guess (since it's a codec that can represent any Unicode string exactly as a bytestring, and it's the standard codec for many purposes, such as XML) -- but it can't be any more than a blind guess, until and unless you're going to tell us more about what you're trying to pass that bytestring to, and for what purposes.

Passing thestring.encode('utf-8') rather than bare thestring will definitely avoid the particular error you're seeing right now, but it may result in peculiar displays (or whatever it is you're trying to do with that bytestring!) unless the recipient is ready, willing and able to accept utf-8 encoding (and how could WE know, having absolutely zero idea about what the recipient could possibly be?!-)

Alex Martelli 2009-10-31 01:12:21

updated info as per your notes and I will start looking into how to use utf-8 now -- thanks!

KenBurnsFan1 2009-10-31 05:11:22

So, now we know your error comes while writing to a file - moving to utf-8 will surely fix THAT... but when is the file read back again and how is it processed then? We're still totally in the dark about the real **purpose** of your unicode -> bytestring conversion!-)

Alex Martelli 2009-10-31 05:24:03

Full script provided = general advice is also welcome.thanks!

KenBurnsFan1 2009-10-31 06:04:30

Answer 5

A:

The standard library module Codecs provides many functions to encode and decode data to and from different sources and outputs.

Tendayi Mawushe 2009-10-31 01:22:18

ansaurus

tags:

views:

answers:

Easy Q: UnicodeEncodeError: 'ascii' codec can't encode character

related questions