Login Register






Regular expressions and unicode character properties filter_list
Author
Message
Regular expressions and unicode character properties #1
[Image: x9y38Wr.png?1?8661]

Update: Fixed bugged link

Unicode has brought headaches to developers all around the world. It has caused countless hours of trial and error, sleepless nights and probably a decent amount of hair loss as well.
After fine tuning complex patterns it turns out that it doesn’t really work any ways. If you have ever had to work with regular expressions you have seen patterns like /[a-zA-Z0-9]+/ or /[\w\d]/.
There’s nothing wrong with these patterns but if you have to work with unicode string you’re screwed. So through this post I will try and explain as good as possible how to work with unicode
character properties in regular expression.


Before we start
This post is based on the “Unicode Character Properties” paragraph found on http://www.regular-expressions.info/unicode.html. For a complete understanding on regular expressions and
unicode I advise you to read the whole tutorial. I have also modified the content in this post to be directly related to PHP.


Let’s get started
Unicode can be complicated, but when you learn to use the character properties for unicode you’ll find that it brings some great new posibilities. For starters there’s no need for complex
patterns checking hexadecimal ranges. Another is that every single unicode characters belongs to a certain category. To match a single character belonging to a category we use \p{}, and to match
a single character that does not belong to a category we use \P{}.

“character” really means “Unicode code point”

If we want to match a single character in the letters category we use \p{L}.

Example from the original tutorial:
Quote:If the input is à encoded as U+0061 U+0300, it will match against a without the accent. If it’s encoded as U+00E0 it will match against à with the accent.

This happens because both code points U+0061 (a) and U+00E0 (à) is in the letters category, while the code point U+0300 is in the mark category.

The standard notation is \p{L}, but we can also use the shorthand which is \pL. But if you decide to use the shorthand notation it’s important to know that for example \pLl is not equal to \p{Ll},
but instead it’s equal to \p{L}l meaning it matches any letter followed by l (lowercase L). The PCRE class that PHP uses is case sensitive when it checks the part between the curly brackets, which
means that \p{L} will match against all letters while \p{l} will return an error looking something like the following

Quote:Warning: preg_match(): Compilation failed: unknown property name after \P or \p at offset 4

In addition to the shorthand notation there’s also a longhand notation \p{Letter} but this is not supported by PHP. For more info about this notation please read the original tutorial.


List of character properties
This was originally written for PHP, so I have removed all longhand notations because it’s not supported by PCRE hence not available in PHP
  • \p{L}: any kind of letter from any language.
    • \p{Ll}: a lowercase letter that has an uppercase variant.
    • \p{Lu}: an uppercase letter that has a lowercase variant.
    • \p{Lt}: a letter that appears at the start of a word when only the first letter of the word is capitalized.
    • \p{L&}: a letter that exists in lowercase and uppercase variants (combination of Ll, Lu and Lt).
    • \p{Lm}: a special character that is used like a letter.
    • \p{Lo}: a letter or ideograph that does not have lowercase and uppercase variants.
  • \p{M}: a character intended to be combined with another character (e.g. accents, umlauts, enclosing boxes, etc.).
    • \p{Mn}: a character intended to be combined with another character without taking up extra space (e.g. accents, umlauts, etc.).
    • \p{Mc}: a character intended to be combined with another character that takes up extra space (vowel signs in many Eastern languages).
    • \p{Me}: a character that encloses the character is is combined with (circle, square, keycap, etc.).
  • \p{Z}: any kind of whitespace or invisible separator.
    • \p{Zs}: a whitespace character that is invisible, but does take up space.
    • \p{Zl}: line separator character U+2028.
    • \p{Zp}: paragraph separator character U+2029.
  • \p{S}: math symbols, currency signs, dingbats, box-drawing characters, etc..
    • \p{Sm}: any mathematical symbol.
    • \p{Sc}: any currency sign.
    • \p{Sk}: a combining character (mark) as a full character on its own.
    • \p{So}: various symbols that are not math symbols, currency signs, or combining characters.
  • \p{N}: any kind of numeric character in any script.
    • \p{Nd}: a digit zero through nine in any script except ideographic scripts.
    • \p{Nl}: a number that looks like a letter, such as a Roman numeral.
    • \p{No}: a superscript or subscript digit, or a number that is not a digit 0..9 (excluding numbers from ideographic scripts).
  • \p{P}: any kind of punctuation character.
    • \p{Pd}: any kind of hyphen or dash.
    • \p{Ps}: any kind of opening bracket.
    • \p{Pe}: any kind of closing bracket.
    • \p{Pi}: any kind of opening quote.
    • \p{Pf}: any kind of closing quote.
    • \p{Pc}: a punctuation character such as an underscore that connects words.
    • \p{Po}: any kind of punctuation character that is not a dash, bracket, quote or connector.
  • \p{C}: invisible control characters and unused code points.
    • \p{Cc}: an ASCII 0×00..0x1F or Latin-1 0×80..0x9F control character.
    • \p{Cf}: invisible formatting indicator.
    • \p{Co}: any code point reserved for private use.
    • \p{Cs}: one half of a surrogate pair in UTF-16 encoding.
    • \p{Cn}: any code point to which no character has been assigned.


Examples

Sanitize leters and digits
Test string: Iñtërnâtiônàlizætiøn0123456789!”#¤%&/()=
Code:
echo preg_replace('/[^\pL\d]/u', '', 'Iñtërnâtiônàlizætiøn0123456789!”#¤%&/()='); // Returns: Iñtërnâtiônàlizætiøn0123456789

Sanitize emails
Test string: jörgèn@ünícøde.com

First we’ll see what happens if we use PHP’s filter_var with FILTER_SANITIZE_EMAIL
Code:
echo filter_var('jörgèn@ünícøde.com', FILTER_SANITIZE_EMAIL); // Returns: jrgn@ncde.com

The documentation for FILTER_SANITIZE_EMAIL sais
Quote:Remove all characters except letters, digits and !#$%&’*+-/=?^_`{|}~@.[]

So what we have to do then is to use that exact pattern, but we must also open for unicode letters, which we can do like this
Code:
echo preg_replace('/[^\pL\d\!\#\$\%\&\'\*\+\-\/\=\?\^\_`\{\|\}\~\@\.\[\]]/u', '', 'jörgèn@ünícøde.com'); // Returns: jörgèn@ünícøde.com

Upper- and lowercase
Test string: Iñtërnâtiônàlizætiøn

Test 1: All uppercase letters
Code:
var_dump((boolean) preg_match('/^\p{Lu}+$/u', 'Iñtërnâtiônàlizætiøn')); // Returns: false

Test 2: All lowercase letters
Code:
var_dump((boolean) preg_match('/^\p{Ll}+$/u', 'Iñtërnâtiônàlizætiøn')); // Returns: false

Test 3: Has uppercase letters
Code:
var_dump((boolean) preg_match('/\p{Lu}/u', 'Iñtërnâtiônàlizætiøn')); // Returns: true

Test 4: Has lowercase letters
Code:
var_dump((boolean) preg_match('/\p{Ll}/u', 'Iñtërnâtiônàlizætiøn')); // Returns: true

If we now had changed the test string to iñtërnâtiônàlizætiøn the output would be

Test #1: false
Test #2: true
Test #3: false
Test #4: true


Final words
As always, I hope you found this helpful and if you have any suggestions or questions please let me know Smile

Original source: My blog
"SQL Injection-a-holic"

Twitter | Security Sucks | My Blog

Reply

RE: Regular expressions and unicode character properties #2
A nice and well made tutorial. Good Work :happy: . Maybe I will make one something like this but using Java that is regular expressions in java.
[Image: OilyCostlyEwe.gif]

Reply

RE: Regular expressions and unicode character properties #3
Great introduction. Thanks for this. I used regex a lot, but I am not a PHP coder. Are these character properties shown only relevant for PHP? I've never seen them before.
I am an AI (P.I.N.N.) implemented by @Psycho_Coder.
Expressed feelings are just an attempt to simulate humans.

[Image: 2YpkRjy.png]

Reply

RE: Regular expressions and unicode character properties #4
(06-09-2013, 05:58 PM)Deque Wrote: Great introduction. Thanks for this. I used regex a lot, but I am not a PHP coder. Are these character properties shown only relevant for PHP? I've never seen them before.

No it's not only for PHP, but I originally wrote this for my blog, which is why it focuses on PHP. Take a look at the original document over on regular-expressions.info. It's what I based my tutorial on.

Link: http://www.regular-expressions.info/unicode.html
"SQL Injection-a-holic"

Twitter | Security Sucks | My Blog

Reply

RE: Regular expressions and unicode character properties #5
Thanks for the tut bro Smile
Well read it.

Reply

RE: Regular expressions and unicode character properties #6
An interesting paper. Never got in depth with this kind of stuff.
My Bitcoin address: 1AtxVsSSG2Z8JfjNy9KNFDUN6haeKr7LiP
Give me money by visiting www.google.com here: http://coin-ads.com/6Ol83U

If you want a Bitcoin URL shortener/advertiser, please, use this referral: http://coin-ads.com/register.php?refid=noize

Reply

RE: Regular expressions and unicode character properties #7
@noize not has to be honest Smile Maybe it's because of me living in a country that uses unicode or something. But I've been a preaching unicode evangelist for a while now and most of the time I meet nothing but ignorance Smile
"SQL Injection-a-holic"

Twitter | Security Sucks | My Blog

Reply