Mysql 四字节汉字支持-6ren

Mysql 四字节汉字支持

转载作者：行者123 更新时间：2023-12-02 19:19:47

27

4

我无法执行此 SQL 脚本:

INSERT INTO `mabase`.`new_table` (`idnew_table`, `name`) VALUES ('2', '𠼭');

错误是:

ERROR 1366: Incorrect string value: '\xF0\xA0\xBC\xAD' for column 'name' at row 1 SQL Statement: INSERT INTO mabase.new_table (idnew_table, name) VALUES ('2', '𠼭')

我的数据库和表采用 utf8 字符集和 utf8_general_ci 排序规则。我也尝试过:utf8_unicode_ci，utf8mb4_general_ci，bg5_cinese_ci,gbk_cinese_ci。

我已经在 MySql 工作台中尝试了所有这些在 Windows 上。

𠼭是四字节字符。我只对他们有问题。请告诉我如何在 mysql 中保存四个字节字符。

最佳答案

您想要的角色，U+20F2D ，驻留在 Unicode 的“补充表意文字平面”的“CJK 统一表意文字扩展 B” block 中，因此在 v5.5 之前的任何 MySQL Unicode 字符集中不可用；自 v5.5 起，它可在 utf8mb4 中找到。 , utf16 , utf16le和 utf32字符集。

它在 MySQL 的 big5 或 gbk 字符集中不可用。

<小时/>

为什么 `utf8` 编码不起作用

如 Unicode Support 下所述:

The initial implementation of Unicode support (in MySQL 4.1) included two character sets for storing Unicode data:

ucs2, the UCS-2 encoding of the Unicode character set using 16 bits per character.

utf8, a UTF-8 encoding of the Unicode character set using one to three bytes per character.

These two character sets support the characters from the Basic Multilingual Plane (BMP) of Unicode Version 3.0. BMP characters have these characteristics:

Their code values are between 0 and 65535 (or U+0000 .. U+FFFF).

They can be encoded with a fixed 16-bit word, as in ucs2.

They can be encoded with 8, 16, or 24 bits, as in utf8.

They are sufficient for almost all characters in major languages.

Characters not supported by the aforementioned character sets include supplementary characters that lie outside the BMP. Characters outside the BMP compare as REPLACEMENT CHARACTER and convert to '?' when converted to a Unicode character set.

In MySQL 5.6, Unicode support includes supplementary characters, which requires new character sets that have a broader range and therefore take more space. The following table shows a brief feature comparison of previous and current Unicode support.
╔══════════════════════════════╦══════════════════════════════════════════════╗║       Before MySQL 5.5       ║              MySQL 5.5 and up                ║╠══════════════════════════════╬══════════════════════════════════════════════╣║ All Unicode 3.0 characters   ║ All Unicode 5.0 and 6.0 characters           ║╠══════════════════════════════╬══════════════════════════════════════════════╣║ No supplementary characters  ║ With supplementary characters                ║╠══════════════════════════════╬══════════════════════════════════════════════╣║ ucs2 character set, BMP only ║ No change                                    ║╠══════════════════════════════╬══════════════════════════════════════════════╣║ utf8 character set for up to ║ No change                                    ║║ three bytes, BMP only        ║                                              ║╠══════════════════════════════╬══════════════════════════════════════════════╣║                              ║ New utf8mb4 character set for up to four     ║║                              ║ bytes, BMP or supplemental                   ║╠══════════════════════════════╬══════════════════════════════════════════════╣║                              ║ New utf16 character set, BMP or supplemental ║╠══════════════════════════════╬══════════════════════════════════════════════╣║                              ║ New utf16le character set, BMP or            ║║                              ║ supplemental (5.6.1 and up)                  ║╠══════════════════════════════╬══════════════════════════════════════════════╣║                              ║ New utf32 character set, BMP or supplemental ║╚══════════════════════════════╩══════════════════════════════════════════════╝
These changes are upward compatible. If you want to use the new character sets, there are potential incompatibility issues for your applications; see Section 10.1.11, “Upgrading from Previous to Current Unicode Support”. That section also describes how to convert tables from utf8 to the (4-byte) utf8mb4 character set, and what constraints may apply in doing so.

为什么`big5`编码不起作用

如 What problems should I be aware of when working with the Big5 Chinese character set? 下所述:

MySQL supports the Big5 character set which is common in Hong Kong and Taiwan (Republic of China). MySQL's big5 is in reality Microsoft code page 950, which is very similar to the original big5 character set.
[ deletia ]
A feature request for adding HKSCS extensions has been filed. People who need this extension may find the suggested patch for Bug #13577 to be of interest.

为什么`gbk`编码不起作用

如 What CJK character sets are available in MySQL? 下所述:

Here, we try to clarify exactly what characters are legitimate in gb2312 or gbk, with reference to the official documents. Please check these references before reporting gb2312 or gbk bugs.

For a complete listing of the gb2312 characters, ordered according to the gb2312_chinese_ci collation: gb2312

MySQL's gbk is in reality “Microsoft code page 936”. This differs from the official gbk for characters A1A4 (middle dot), A1AA (em dash), A6E0-A6F5, and A8BB-A8C0.

For a listing of gbk/Unicode mappings, see http://www.unicode.org/Public/MAPPINGS/VENDORS/MICSFT/WINDOWS/CP936.TXT.

For MySQL's listing of gbk characters, see gbk.

关于Mysql 四字节汉字支持，我们在Stack Overflow上找到一个类似的问题： https://stackoverflow.com/questions/17680237/

27

4

0

文章推荐： security - 防止移动 API 客户端身份盗用

文章推荐： angularjs - 与兄弟指令沟通

文章推荐： security - HTTP 基本身份验证 + 访问 token ？

文章推荐： assembly 方括号

c# - 字节 + 字节 = 未知结果
美好的一天!我试图添加两个字节变量并注意到奇怪的结果。 byte valueA = 255; byte valueB = 1; byte valueC = (byte)(valueA + valueB
ios - 转换[字节]？到[字节]
嗨，我是 swift 的新手，我正在尝试解码以 [Byte] 形式发回给我的字节数组？当我尝试使用 if let string = String(bytes: d, encoding: .utf8)
postgresql - 由于 IPV6 需要 128 位(16 字节)那么为什么在 postgres CIDR 数据类型中存储为 24 字节(8.1)和 19 字节(9.1)？
我正在使用 ipv4 和 ipv6 存储在 postgres 数据库中。因为 ipv4 需要 32 位(4 字节)而 ipv6 需要 128(16 字节)位。那么为什么在 postgres 中 CI
string - []字节(字符串)与[]字节(*字符串)
我很好奇为什么 Go 不提供 []byte(*string) 方法。从性能的角度来看，[]byte(string) 不会复制输入参数并增加更多成本(尽管这看起来很奇怪，因为字符串是不可变的，为什么要复
客户端发送 500 字节，但服务器接收 244 字节 - 套接字编程？
我正在尝试为UDP实现Stop-and-Wait ARQ。根据停止等待约定，我在 0 和 1 之间切换 ACK。正确的 ACK 定义为正确的序列号(0 或 1)AND消息长度。以下片段是我的代码的
php - filesize() 始终读取 0 字节，即使文件大小不是 0 字节
我在下面写了一些代码，目前我正在测试，所以代码中没有数据库查询。下面的代码显示 if(filesize($filename) != 0) 总是转到 else，即使文件不是 0 字节而是 16 字节那
java - 无法读取整个 header ；读取 0 字节；预计 512 字节
我使用 Apache poi 3.8 来读取 xls 文件，但出现异常: java.io.IOException: Unable to read entire header; 0 by
python - 为什么在调用 .clear() 后字典大小为 72 字节，而实例化时为 240 字节？
字典大小为 72 字节(根据 getsizeof(dict) 在字典上调用 .clear() 之后发生了什么，当新实例化的字典返回 240 字节时？我知道一个简单的 dict 的起始大小为“8”，并
c - 将 4 字节 int 交织到 8 字节 int
我目前正在努力创建一个函数，它接受两个 4 字节无符号整数，并返回一个 8 字节无符号长整数。我试图将我的工作基于 this research 描述的方法，但我的所有尝试都没有成功。我正在处理的具体输
c++ - 将 4 字节 int 解释为 4 字节 float
看看这个简单的程序: #include using namespace std; int main() { unsigned int i=0x3f800000; float* p=(float*)(
java - Java 中的字符串 "8000000000000000"(16 字节)相当于 "BCD"(8 字节)
我创建了自己的函数，将一个字符串转换为其等效的 BCD 格式的 bytes[]。然后我将此字节发送到 DataOutputStram (使用需要 byte[] 数组的写入方法)。问题出在数字字符串“8
c - 带有静态堆的小块内存分配器(典型值 <= 16 字节，稀有值 >= 64 字节，最大值 = 192)
此分配器将在具有静态内存的嵌入式系统中使用(即，没有可用的系统堆，因此“堆”将只是“char heap[4096]”) 周围似乎有很多“小型内存分配器”，但我正在寻找能够处理非常小的分配的一个。我说的
sql-server - 警告!最大 key 长度为 900 字节。索引的最大长度为 1000 字节
我将数据库脚本从 64 位系统传输到 32 位系统。当我执行脚本时，出现以下错误， Warning! The maximum key length is 900 bytes. The index 'U
linux - 128 字节 Ext2 和 256 字节 Ext3 的 inode 数据结构差异
想知道 128 字节 ext2 和 256 字节 ext3 文件系统之间的 inode 数据结构差异。我一直在为 ext2、128 字节 inode 使用此引用:http://www.nongnu.
java - Cassandra = 内存/编码- key 占用空间(哈希/字节[]=>十六进制=>UTF16=>字节[])
我试图理解使用 MD5 哈希作为 Cassandra key 在“内存/存储消耗”方面的含义: 我的内容(在 Java 中)的 MD5 哈希 = byte[] 长 16 个字节。 (16 字节来自维基
linux - 需要帮助 - 出现错误 : xrealloc: subst. c:4072: 无法重新分配 1073741824 字节(已分配 0 字节)
检查其他人是否也遇到类似问题。 shell脚本中的代码: ## Convert file into Unix format first. ## THIS is IMPORTANT. ###
c++ - x86 4 字节 float 与 8 字节 double (与 long long 相比)？
我们有一个测量数据处理应用程序，目前所有数据都保存为 C++ float，这意味着在我们的 x86/Windows 平台上为 32 位/4 字节。 (32 位 Windows 应用程序)。由于精度成
java - Long 的大小为 8 字节，那么在 JAVA 中如何将 'promoted' 转换为 float (4 字节)？
我读到在 Java 中 long 类型可以提升为 float 和 double ( http://www.javatpoint.com/method-overloading-in-java )。我想问
python - 将 n 个元素(大小 = 2 字节，十进制)的列表拆分为 2n 个元素(大小 = 1 字节，十六进制)
我有一个包含 n 个十进制元素的列表，其中每个元素都是两个字节长。可以说: x = [9000 , 5000 , 2000 , 400] 这个想法是将每个元素拆分为 MSB 和 LSB 并将其存储在
1 个 block (16 字节)的 Java AES-128 加密返回 2 个 block (32 字节)作为输出
我使用以下代码进行 AES-128 加密来编码一个 16 字节的 block ，但编码值的长度给出了 2 个 32 字节的 block 。我错过了什么吗？ plainEnc = AES.enc

首页

博学

6Ren·AI

商城

Mysql 四字节汉字支持

为什么 `utf8` 编码不起作用

为什么`big5`编码不起作用

为什么`gbk`编码不起作用

首页

博学

6Ren·AI

商城

Mysql 四字节汉字支持

为什么 utf8 编码不起作用

为什么big5编码不起作用

为什么gbk编码不起作用

为什么 `utf8` 编码不起作用

为什么`big5`编码不起作用

为什么`gbk`编码不起作用