MySQL REGEXP query - accent insensitive search

前端 未结 7 1895
清歌不尽
清歌不尽 2020-12-19 00:46

I\'m looking to query a database of wine names, many of which contain accents (but not in a uniform way, and so similar wines may be entered with or without accents)

7条回答
  •  囚心锁ツ
    2020-12-19 01:22

    You're out of luck:

    Warning

    The REGEXP and RLIKE operators work in byte-wise fashion, so they are not multi-byte safe and may produce unexpected results with multi-byte character sets. In addition, these operators compare characters by their byte values and accented characters may not compare as equal even if a given collation treats them as equal.

    The [[:<:]] and [[:>:]] regexp operators are markers for word boundaries. The closest you can achieve with the LIKE operator is something on this line:

    SELECT *
    FROM `table`
    WHERE wine_name = 'Faugères'
       OR wine_name LIKE 'Faugères %'
       OR wine_name LIKE '% Faugères'
    

    As you can see it's not fully equivalent because I've restricted the concept of word boundary to spaces. Adding more clauses for other boundaries would be a mess.

    You could also use full text searches (although it isn't the same) but you can't define full text indexes in InnoDB tables (yet).

    You're certainly out of luck :)


    Addendum: this has changed as of MySQL 8.0:

    MySQL implements regular expression support using International Components for Unicode (ICU), which provides full Unicode support and is multibyte safe. (Prior to MySQL 8.0.4, MySQL used Henry Spencer's implementation of regular expressions, which operates in byte-wise fashion and is not multibyte safe.

提交回复
热议问题