How to check the charset of string?

前端 未结 7 1701
心在旅途
心在旅途 2021-02-04 08:48

How do I check if the charset of a string is UTF8?

相关标签:
7条回答
  • 2021-02-04 09:23
    function is_utf8($string) {   
    return preg_match('%^(?:  
    [\x09\x0A\x0D\x20-\x7E] # ASCII  
    | [\xC2-\xDF][\x80-\xBF] # non-overlong 2-byte  
    | \xE0[\xA0-\xBF][\x80-\xBF] # excluding overlongs  
    | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2} # straight 3-byte  
    | \xED[\x80-\x9F][\x80-\xBF] # excluding surrogates  
    | \xF0[\x90-\xBF][\x80-\xBF]{2} # planes 1-3  
    | [\xF1-\xF3][\x80-\xBF]{3} # planes 4-15  
    | \xF4[\x80-\x8F][\x80-\xBF]{2} # plane 16  
    )*$%xs', $string);   
    

    }

    I have checked. This function is effective.

    0 讨论(0)
  • 2021-02-04 09:23

    Better yet, use both of the above solutions.

    function isUtf8($string) {
        if (function_exists("mb_check_encoding") && is_callable("mb_check_encoding")) {
            return mb_check_encoding($string, 'UTF8');
        }
    
        return preg_match('%^(?:
              [\x09\x0A\x0D\x20-\x7E]            # ASCII
            | [\xC2-\xDF][\x80-\xBF]             # non-overlong 2-byte
            |  \xE0[\xA0-\xBF][\x80-\xBF]        # excluding overlongs
            | [\xE1-\xEC\xEE\xEF][\x80-\xBF]{2}  # straight 3-byte
            |  \xED[\x80-\x9F][\x80-\xBF]        # excluding surrogates
            |  \xF0[\x90-\xBF][\x80-\xBF]{2}     # planes 1-3
            | [\xF1-\xF3][\x80-\xBF]{3}          # planes 4-15
            |  \xF4[\x80-\x8F][\x80-\xBF]{2}     # plane 16
        )*$%xs', $string);
    
    } 
    
    0 讨论(0)
  • 2021-02-04 09:23

    if its send to u from server

    echo $_SERVER['HTTP_ACCEPT_CHARSET'];
    
    0 讨论(0)
  • 2021-02-04 09:27

    mb_detect_encoding($string); will return the actual character set of $string. mb_check_encoding($string, 'UTF-8'); will return TRUE if character set of $string is UTF-8 else FALSE

    0 讨论(0)
  • 2021-02-04 09:32

    Don't reinvent the wheel. There is a builtin function for that task: mb_check_encoding().

    mb_check_encoding($string, 'UTF-8');
    
    0 讨论(0)
  • 2021-02-04 09:33

    Just a side note:

    You cannot determine if a given string is encoded in UTF-8. You only can determine if a given string is definitively not encoded in UTF-8. Please see a related question here:

    You cannot detect if a given string (or byte sequence) is a UTF-8 encoded text as for example each and every series of UTF-8 octets is also a valid (if nonsensical) series of Latin-1 (or some other encoding) octets. However not every series of valid Latin-1 octets are valid UTF-8 series.

    0 讨论(0)
提交回复
热议问题