ansaurus

Question

Answer 1

A:

I suspect, you a difference in charsets of your terminal (which can be UTF-8) and the source code of your perl script (which you maybe edit in some charset-aware editor in 8859-1). If you are sure, that your terminal and your source code are in the same charset, try to put use utf8; to your script header (see man perlunicode). If that does not help, try to print the data, which is stored to your database (increase debug logging for DBI) (maybe irrrelevant, as you don't store data as UTF8). Generally, try to provide:

The codepage of your terminal (locale) if you execute your script for a terminal (or system locale, which is used by your server, if you launch it from e.g. apache)
The charset of your source code.
MySQL connection codepage (do you issue SET NAMES 'utf8'?)

Also for HTML encoding you may find easier to reuse HTML::Entities::decode() / HTML::Entities::encode() rather than implementing this on your own.

dma_k 2010-02-08 11:15:15

Answer 2

+2 A:

As noted in the comment on your question, I'm unsure what exactly you're asking.

So I'm assuming you're trying to convert Unicode characters into HTML entities. In which case, using one of the pre-made modules should be better. If that is not working due to encoding problems (which are quite tricky in Perl), then the answer to your question:

Is there not a encoding option like
open FILE, "<", $file or die "Cannot open:$!\n", "UTF-8";

... will probably solve it, and it would probably make your own attempt work as well, but better to use a ready-made one ;-) (by the way, the way you wrote it there was as a "UTF-8" option to die which made it a little hard to understand what you were asking ;-)

Yes there is a UTF-8 option, assuming you have a recent perl (>= v5.8):

open(my $fh,'<:encoding(UTF-8)', $file) or die "Error opening $file: $!";

(example adapted from perluniintro)

You can also use binmode to change an already open filehandle (e.g. STDIN/OUT).

binmode(STDOUT, ":encoding(UTF-8)");

You can also set the default encoding with the open pragma.

But for this I suggest trying binmode or changing your open line to see if that solves it.

If you have a perl less than v5.8, things are trickier, but maybe resolvable if you tell us the version.

A couple of other things I noticed by the way:

Not essential, but it's considered better to use a lexically scoped filehandle (my $fh instead of FILE).
When you put a newline on the die string, it suppresses the line number information that is normally added to help you find the problem.
If you put the name of the file that couldn't be opened (or the SQL that failed, or whatever) in the die message it will be easier to debug.
Don't use sub prototypes in Perl (5) : (sub unicodeConvert($)). Don't put the $/@/% etc. in there. It doesn't just check things, it may change the meaning in confusing ways. It is only needed to create new "built-in style" operators.

FalseVinylShrub 2010-02-08 16:04:38

P.S. as noted by dma_k, if you continue with your own subroutine `unicodeConvert`, you'll need to either `use utf8` to get the Unicode characters in the source code recognised or convert to numeric encodings e.g. `"\x{263a}".

FalseVinylShrub 2010-02-08 16:11:41

':utf8' is a shortcut for ':encoding(UTF-8)'

brian d foy 2010-02-08 20:57:03

Thanks brian... I normally use `':utf8'`, but when checking the docs, it seemed there was a difference and `':encoding(UTF-8)'` was safer for input files (found it: perluniintro, Unicode I/O, just below a big list of `binmode` examples)... can you clarify if this is true? Thanks.

FalseVinylShrub 2010-02-09 04:30:45

ansaurus

tags:

views:

answers:

Perl - read file with encoding method?

related questions