![]() |
|
Regular expressions and unicode character properties - Printable Version +- Sinisterly (https://sinister.li) +-- Forum: Coding (https://sinister.li/Forum-Coding) +--- Forum: Coding (https://sinister.li/Forum-Coding--71) +--- Thread: Regular expressions and unicode character properties (/Thread-Regular-expressions-and-unicode-character-properties) |
Regular expressions and unicode character properties - RogueCoder - 06-08-2013 ![]() Update: Fixed bugged link Unicode has brought headaches to developers all around the world. It has caused countless hours of trial and error, sleepless nights and probably a decent amount of hair loss as well. After fine tuning complex patterns it turns out that it doesn’t really work any ways. If you have ever had to work with regular expressions you have seen patterns like /[a-zA-Z0-9]+/ or /[\w\d]/. There’s nothing wrong with these patterns but if you have to work with unicode string you’re screwed. So through this post I will try and explain as good as possible how to work with unicode character properties in regular expression. Before we start This post is based on the “Unicode Character Properties” paragraph found on http://www.regular-expressions.info/unicode.html. For a complete understanding on regular expressions and unicode I advise you to read the whole tutorial. I have also modified the content in this post to be directly related to PHP. Let’s get started Unicode can be complicated, but when you learn to use the character properties for unicode you’ll find that it brings some great new posibilities. For starters there’s no need for complex patterns checking hexadecimal ranges. Another is that every single unicode characters belongs to a certain category. To match a single character belonging to a category we use \p{}, and to match a single character that does not belong to a category we use \P{}. “character” really means “Unicode code point” If we want to match a single character in the letters category we use \p{L}. Example from the original tutorial: Quote:If the input is à encoded as U+0061 U+0300, it will match against a without the accent. If it’s encoded as U+00E0 it will match against à with the accent. This happens because both code points U+0061 (a) and U+00E0 (à) is in the letters category, while the code point U+0300 is in the mark category. The standard notation is \p{L}, but we can also use the shorthand which is \pL. But if you decide to use the shorthand notation it’s important to know that for example \pLl is not equal to \p{Ll}, but instead it’s equal to \p{L}l meaning it matches any letter followed by l (lowercase L). The PCRE class that PHP uses is case sensitive when it checks the part between the curly brackets, which means that \p{L} will match against all letters while \p{l} will return an error looking something like the following Quote:Warning: preg_match(): Compilation failed: unknown property name after \P or \p at offset 4 In addition to the shorthand notation there’s also a longhand notation \p{Letter} but this is not supported by PHP. For more info about this notation please read the original tutorial. List of character properties This was originally written for PHP, so I have removed all longhand notations because it’s not supported by PCRE hence not available in PHP
Examples Sanitize leters and digits Test string: Iñtërnâtiônàlizætiøn0123456789!”#¤%&/()= Code: echo preg_replace('/[^\pL\d]/u', '', 'Iñtërnâtiônàlizætiøn0123456789!”#¤%&/()=');
// Returns: Iñtërnâtiônàlizætiøn0123456789Sanitize emails Test string: jörgèn@ünícøde.com First we’ll see what happens if we use PHP’s filter_var with FILTER_SANITIZE_EMAIL Code: echo filter_var('jörgèn@ünícøde.com', FILTER_SANITIZE_EMAIL);
// Returns: jrgn@ncde.comThe documentation for FILTER_SANITIZE_EMAIL sais Quote:Remove all characters except letters, digits and !#$%&’*+-/=?^_`{|}~@.[] So what we have to do then is to use that exact pattern, but we must also open for unicode letters, which we can do like this Code: echo preg_replace('/[^\pL\d\!\#\$\%\&\'\*\+\-\/\=\?\^\_`\{\|\}\~\@\.\[\]]/u', '', 'jörgèn@ünícøde.com');
// Returns: jörgèn@ünícøde.comUpper- and lowercase Test string: Iñtërnâtiônàlizætiøn Test 1: All uppercase letters Code: var_dump((boolean) preg_match('/^\p{Lu}+$/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: falseTest 2: All lowercase letters Code: var_dump((boolean) preg_match('/^\p{Ll}+$/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: falseTest 3: Has uppercase letters Code: var_dump((boolean) preg_match('/\p{Lu}/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: trueTest 4: Has lowercase letters Code: var_dump((boolean) preg_match('/\p{Ll}/u', 'Iñtërnâtiônàlizætiøn'));
// Returns: trueIf we now had changed the test string to iñtërnâtiônàlizætiøn the output would be Test #1: false Test #2: true Test #3: false Test #4: true Final words As always, I hope you found this helpful and if you have any suggestions or questions please let me know ![]() Original source: My blog RE: Regular expressions and unicode character properties - Psycho_Coder - 06-09-2013 A nice and well made tutorial. Good Work :happy: . Maybe I will make one something like this but using Java that is regular expressions in java. RE: Regular expressions and unicode character properties - Deque - 06-09-2013 Great introduction. Thanks for this. I used regex a lot, but I am not a PHP coder. Are these character properties shown only relevant for PHP? I've never seen them before. RE: Regular expressions and unicode character properties - RogueCoder - 06-09-2013 (06-09-2013, 05:58 PM)Deque Wrote: Great introduction. Thanks for this. I used regex a lot, but I am not a PHP coder. Are these character properties shown only relevant for PHP? I've never seen them before. No it's not only for PHP, but I originally wrote this for my blog, which is why it focuses on PHP. Take a look at the original document over on regular-expressions.info. It's what I based my tutorial on. Link: http://www.regular-expressions.info/unicode.html RE: Regular expressions and unicode character properties - MrGeek - 06-10-2013 Thanks for the tut bro ![]() Well read it. RE: Regular expressions and unicode character properties - noize - 06-13-2013 An interesting paper. Never got in depth with this kind of stuff. RE: Regular expressions and unicode character properties - RogueCoder - 06-14-2013 @noize not has to be honest Maybe it's because of me living in a country that uses unicode or something. But I've been a preaching unicode evangelist for a while now and most of the time I meet nothing but ignorance
|