Regular expressions and unicode character properties 06-08-2013, 08:16 PM
#1
![[Image: x9y38Wr.png?1?8661]](http://i.imgur.com/x9y38Wr.png?1?8661)
Update: Fixed bugged link
Unicode has brought headaches to developers all around the world. It has caused countless hours of trial and error, sleepless nights and probably a decent amount of hair loss as well.
After fine tuning complex patterns it turns out that it doesn’t really work any ways. If you have ever had to work with regular expressions you have seen patterns like /[a-zA-Z0-9]+/ or /[\w\d]/.
There’s nothing wrong with these patterns but if you have to work with unicode string you’re screwed. So through this post I will try and explain as good as possible how to work with unicode
character properties in regular expression.
Before we start
This post is based on the “Unicode Character Properties” paragraph found on http://www.regular-expressions.info/unicode.html. For a complete understanding on regular expressions and
unicode I advise you to read the whole tutorial. I have also modified the content in this post to be directly related to PHP.
Let’s get started
Unicode can be complicated, but when you learn to use the character properties for unicode you’ll find that it brings some great new posibilities. For starters there’s no need for complex
patterns checking hexadecimal ranges. Another is that every single unicode characters belongs to a certain category. To match a single character belonging to a category we use \p{}, and to match
a single character that does not belong to a category we use \P{}.
“character” really means “Unicode code point”
If we want to match a single character in the letters category we use \p{L}.
Example from the original tutorial:
Quote:If the input is à encoded as U+0061 U+0300, it will match against a without the accent. If it’s encoded as U+00E0 it will match against à with the accent.
This happens because both code points U+0061 (a) and U+00E0 (à) is in the letters category, while the code point U+0300 is in the mark category.
The standard notation is \p{L}, but we can also use the shorthand which is \pL. But if you decide to use the shorthand notation it’s important to know that for example \pLl is not equal to \p{Ll},
but instead it’s equal to \p{L}l meaning it matches any letter followed by l (lowercase L). The PCRE class that PHP uses is case sensitive when it checks the part between the curly brackets, which
means that \p{L} will match against all letters while \p{l} will return an error looking something like the following
Quote:Warning: preg_match(): Compilation failed: unknown property name after \P or \p at offset 4
In addition to the shorthand notation there’s also a longhand notation \p{Letter} but this is not supported by PHP. For more info about this notation please read the original tutorial.
List of character properties
This was originally written for PHP, so I have removed all longhand notations because it’s not supported by PCRE hence not available in PHP
- \p{L}: any kind of letter from any language.
- \p{Ll}: a lowercase letter that has an uppercase variant.
- \p{Lu}: an uppercase letter that has a lowercase variant.
- \p{Lt}: a letter that appears at the start of a word when only the first letter of the word is capitalized.
- \p{L&}: a letter that exists in lowercase and uppercase variants (combination of Ll, Lu and Lt).
- \p{Lm}: a special character that is used like a letter.
- \p{Lo}: a letter or ideograph that does not have lowercase and uppercase variants.
- \p{Ll}: a lowercase letter that has an uppercase variant.
- \p{M}: a character intended to be combined with another character (e.g. accents, umlauts, enclosing boxes, etc.).
- \p{Mn}: a character intended to be combined with another character without taking up extra space (e.g. accents, umlauts, etc.).
- \p{Mc}: a character intended to be combined with another character that takes up extra space (vowel signs in many Eastern languages).
- \p{Me}: a character that encloses the character is is combined with (circle, square, keycap, etc.).
- \p{Mn}: a character intended to be combined with another character without taking up extra space (e.g. accents, umlauts, etc.).
- \p{Z}: any kind of whitespace or invisible separator.
- \p{Zs}: a whitespace character that is invisible, but does take up space.
- \p{Zl}: line separator character U+2028.
- \p{Zp}: paragraph separator character U+2029.
- \p{Zs}: a whitespace character that is invisible, but does take up space.
- \p{S}: math symbols, currency signs, dingbats, box-drawing characters, etc..
- \p{Sm}: any mathematical symbol.
- \p{Sc}: any currency sign.
- \p{Sk}: a combining character (mark) as a full character on its own.
- \p{So}: various symbols that are not math symbols, currency signs, or combining characters.
- \p{Sm}: any mathematical symbol.
- \p{N}: any kind of numeric character in any script.
- \p{Nd}: a digit zero through nine in any script except ideographic scripts.
- \p{Nl}: a number that looks like a letter, such as a Roman numeral.
- \p{No}: a superscript or subscript digit, or a number that is not a digit 0..9 (excluding numbers from ideographic scripts).
- \p{Nd}: a digit zero through nine in any script except ideographic scripts.
- \p{P}: any kind of punctuation character.
- \p{Pd}: any kind of hyphen or dash.
- \p{Ps}: any kind of opening bracket.
- \p{Pe}: any kind of closing bracket.
- \p{Pi}: any kind of opening quote.
- \p{Pf}: any kind of closing quote.
- \p{Pc}: a punctuation character such as an underscore that connects words.
- \p{Po}: any kind of punctuation character that is not a dash, bracket, quote or connector.
- \p{Pd}: any kind of hyphen or dash.
- \p{C}: invisible control characters and unused code points.
- \p{Cc}: an ASCII 0×00..0x1F or Latin-1 0×80..0x9F control character.
- \p{Cf}: invisible formatting indicator.
- \p{Co}: any code point reserved for private use.
- \p{Cs}: one half of a surrogate pair in UTF-16 encoding.
- \p{Cn}: any code point to which no character has been assigned.
- \p{Cc}: an ASCII 0×00..0x1F or Latin-1 0×80..0x9F control character.
Examples
Sanitize leters and digits
Test string: Iñtërnâtiônàlizætiøn0123456789!”#¤%&/()=
Code:
echo preg_replace('/[^\pL\d]/u', '', 'Iñtërnâtiônàlizætiøn0123456789!”#¤%&/()=');
// Returns: Iñtërnâtiônàlizætiøn0123456789Sanitize emails
Test string: jörgèn@ünícøde.com
First we’ll see what happens if we use PHP’s filter_var with FILTER_SANITIZE_EMAIL
Code:
echo filter_var('jörgèn@ünícøde.com', FILTER_SANITIZE_EMAIL);
// Returns: jrgn@ncde.comThe documentation for FILTER_SANITIZE_EMAIL sais
Quote:Remove all characters except letters, digits and !#$%&’*+-/=?^_`{|}~@.[]
So what we have to do then is to use that exact pattern, but we must also open for unicode letters, which we can do like this
Code:
echo preg_replace('/[^\pL\d\!\#\$\%\&\'\*\+\-\/\=\?\^\_`\{\|\}\~\@\.\[\]]/u', '', 'jörgèn@ünícøde.com');
// Returns: jörgèn@ünícøde.comUpper- and lowercase
Test string: Iñtërnâtiônàlizætiøn
Test 1: All uppercase letters
Code:
var_dump((boolean) preg_match('/^\p{Lu}+$/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: falseTest 2: All lowercase letters
Code:
var_dump((boolean) preg_match('/^\p{Ll}+$/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: falseTest 3: Has uppercase letters
Code:
var_dump((boolean) preg_match('/\p{Lu}/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: trueTest 4: Has lowercase letters
Code:
var_dump((boolean) preg_match('/\p{Ll}/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: trueIf we now had changed the test string to iñtërnâtiônàlizætiøn the output would be
Test #1: false
Test #2: true
Test #3: false
Test #4: true
Final words
As always, I hope you found this helpful and if you have any suggestions or questions please let me know

Original source: My blog



![[+]](https://sinister.li/images/modern/collapse_collapsed.png)
![[Image: OilyCostlyEwe.gif]](http://fat.gfycat.com/OilyCostlyEwe.gif)
![[Image: 2YpkRjy.png]](http://i.imgur.com/2YpkRjy.png)