ansaurus

Question

How do I match only fully-composed characters in a Unicode string in Perl?

Answer 1

+2 A:

Yes, those expressions are locale dependant.

Jonathan Leffler 2008-10-15 03:11:14

Can you name an environment and/or regular expression implementation that allows [:print:] to respect a Japanese UTF-8 locale/encoding? I am using Perl in Linux with Japanese UTF-8 locale/encoding and it does not match Japanese character.

dreamlax 2008-10-15 03:14:58

Answer 2

+1 A:

You could always use the character class [^[:cntrl:]] to match non-control characters.

Adam Rosenfield 2008-10-15 03:26:58

This does not match Unicode control characters (in my environment setup and using Perl). There are Unicode control characters for changing text direction and so on. Using [^[:ctrnl:]] will match these Unicode ones but not ASCII ones.

dreamlax 2008-10-15 04:03:56

Answer 3

+5 A:

echo あ| perl -nle 'BEGIN{binmode STDIN,":utf8"} print"[$_]"; print /[[:print:]]/ ? "YES" : "NO"'

This mostly works, though it generates a warning about a wide character. But it gives you the idea: you must be sure you're dealing with a real unicode string (check utf8::is_utf8). Or just check perlunicode at all - the whole subject still makes my head spin.

Tanktalus 2008-10-15 05:27:30

You can get rid of the ugly BEGIN{binmode STDIN, ":utf8"} kludge by supplying the option -CS on the command line.

moritz 2008-10-15 06:43:30

... that will also make the warning go away, because it sets up STDOUT in the same way as STDIN.

moritz 2008-10-15 06:50:51

That may not be as much of an option if the OP is writing a module to handle this instead of a standalone script. So I'm going to leave my solution, as well as your fix in the hopes the OP can figure out which one is better for his/her scenario. Thanks :-)

Tanktalus 2008-10-15 13:35:20

This pattern is wrong. [[:print:]] will match "\x{3099}" which is not a fully-composed character! See my answer for a working pattern.

daxim 2010-01-07 22:59:05

Answer 4

+4 A:

I think you don't want or need locales for that but, but rather Unicode. If you have decoded a text string, \w will match word characters in any language, \d matches not just 0..9 but every Unicode digit etc. In regexes you can query Unicode properties with \p{PropertyName}. Particularly interesting for you might be \p{Print}. Here's a list of all the available Unicode character properties.

I wrote an article about the basics and subtleties of Unicode and Perl, it should give you a good idea on what to do that perl will recognize your string as a sequence of characters, not just a sequence of bytes.

Update: with Unicode you don't get language dependent behaviour, but instead sane defaults regardless of language. This may or may not be what you want, but for the distinction of priintable/control character I don't see why you'd need language dependent behaviour.

moritz 2008-10-15 06:48:26

Answer 5

+1 A:

\X matches a fully-composed character (sequence). Proof:

#!/usr/bin/env perl
use 5.010;
use utf8;
use Encode qw(encode_utf8);

for my $string (qw(あ ご ご), "\x{3099}") {
    say encode_utf8 sprintf "%s $string", $string =~ /\A \X \z/msx ? 'ok' : 'nok';
}

The test data are: a normal character, a pre-combined character, a combining character sequence and a combining character (which "doesn't count" on its own, a simplification of Chapter 3 of Unicode).

Substitute \X with [[:print:]] to see that Tanktalus' answer produces false matches for the last two cases.

daxim 2010-01-07 23:12:47

ansaurus

tags:

views:

answers:

How do I match only fully-composed characters in a Unicode string in Perl?

related questions