mirror of
https://github.com/open-webui/open-webui.git
synced 2026-10-02 02:12:25 +00:00
Korean EUC-KR and Japanese Shift-JIS text files were stored as garbled Chinese characters, so retrieval, knowledge bases and the model context all worked on text that is not in the file. Encoding detection puts chardet's guess in front of a fixed GB18030, Big5, EUC-KR, EUC-JP try order, and GB18030 decodes almost any double-byte text without an error. The guess map was written for chardet 5. Since the bump to chardet 7 in v0.10.0, Korean text is reported as CP949 and Japanese text as cp932 or SHIFT_JIS, which the map either did not know or dropped because the codec was not in the try order, so these files fell through to GB18030. The map now covers CP949 and cp932, and a mapped guess is always tried first. SHIFT_JIS maps to cp932, the Windows superset, because chardet also reports SHIFT_JIS for ordinary Japanese files containing characters such as ① or ㈱ that plain Shift-JIS cannot decode; this is the same subset-to-superset rule the map already applies to GB2312. Korean and Japanese files now decode correctly, and Chinese, EUC-JP, UTF-8 and Western files decode as before. The one trade-off of trusting the guess: chardet 7 labels some files holding only a few Chinese characters (a short label or a one-line comment) as CP949, and those now read as Korean. No regressions were found in files with more Chinese text than that. Fixes #31352 |
||
|---|---|---|
| .. | ||
| loaders | ||
| models | ||
| vector | ||
| web | ||
| external.py | ||
| utils.py | ||