MySQL 5.0 supports two character sets for storing Unicode data:
1.ucs2, the UCS-2 encoding of the Unicode character set using 16 bits per character.
In UCS-2, every character is represented by a two-byte Unicode code with the most significant byte first. For example:
In MySQL, the
2.utf8, a UTF-8 encoding of the Unicode character set using one to three bytes per character.
UTF-8 (Unicode Transformation Format with 8-bit units) is an alternative way to store Unicode data. It is implemented according to RFC 3629, which describes encoding sequences that take from one to four bytes. Currently, MySQL support for UTF-8 does not include four-byte sequences. (An older standard for UTF-8 encoding, RFC 2279, describes UTF-8 sequences that take from one to six bytes. RFC 3629 renders RFC 2279 obsolete; for this reason, sequences with five and six bytes are no longer used.)
The idea of UTF-8 is that various Unicode characters are encoded using byte sequences of different lengths:
1.ucs2, the UCS-2 encoding of the Unicode character set using 16 bits per character.
In UCS-2, every character is represented by a two-byte Unicode code with the most significant byte first. For example:
LATIN CAPITAL LETTER A has the code 0x0041 and it is stored as a two-byte sequence: 0x00 0x41. CYRILLIC SMALL LETTER YERU (Unicode 0x044B) is stored as a two-byte sequence: 0x04 0x4B.In MySQL, the
ucs2 character set is a fixed-length 16-bit encoding for Unicode BMP characters. 2.utf8, a UTF-8 encoding of the Unicode character set using one to three bytes per character.
UTF-8 (Unicode Transformation Format with 8-bit units) is an alternative way to store Unicode data. It is implemented according to RFC 3629, which describes encoding sequences that take from one to four bytes. Currently, MySQL support for UTF-8 does not include four-byte sequences. (An older standard for UTF-8 encoding, RFC 2279, describes UTF-8 sequences that take from one to six bytes. RFC 3629 renders RFC 2279 obsolete; for this reason, sequences with five and six bytes are no longer used.)
The idea of UTF-8 is that various Unicode characters are encoded using byte sequences of different lengths:
- Basic Latin letters, digits, and punctuation signs use one byte.
- Most European and Middle East script letters fit into a two-byte sequence: extended Latin letters (with tilde, macron, acute, grave and other accents), Cyrillic, Greek, Armenian, Hebrew, Arabic, Syriac, and others.
- Korean, Chinese, and Japanese ideographs use three-byte sequences.
To save space with UTF-8, useVARCHARinstead ofCHAR.