mirror of https://gitlab.freedesktop.org/uchardet/uchardet.git synced 2026-04-30 19:09:25 +08:00

Go to file

Jehan e5234d6b61 Stating endianness of UTF-16 and UTF-32 was an error when BOM present. According to RFC 2781, section 3.3: "Systems labelling UTF-16BE/LE text MUST NOT prepend a BOM to the text." Since uchardet cannot (and should not, obviously, it's not its role) modify input text, when a BOM is present, we should always label the encoding as "UTF-16" only. Also it broke unit tests in using programs since a conversion from UTF-8 to UTF-16LE/BE would create a text without BOM, and a conversion from UTF-16LE/BE to UTF-8 creates a UTF-8 text with a BOM, which changed existing behaviours. Same goes for UTF-32. See also Unicode 5.0.0 standard, section 3.10 (tables 3.8 and 3.9 in particular).		2015-12-04 19:19:39 +01:00
build-mac	Added shared scheme	2014-10-26 00:24:03 -07:00
doc	Update README and manual...	2015-11-27 18:27:11 +01:00
script	BuildLangModel: forgot to add logs for Thai models generation.	2015-12-04 03:26:52 +01:00
src	Stating endianness of UTF-16 and UTF-32 was an error when BOM present.	2015-12-04 19:19:39 +01:00
test	LangModels: add ISO-8859-11 and regenerate TIS-620 Thai models.	2015-12-04 03:14:52 +01:00
.gitignore	Add a .gitignore.	2015-11-29 02:27:42 +01:00
AUTHORS	Update authors.	2015-12-03 19:44:13 +01:00
CMakeLists.txt	Release: version 0.0.4.	2015-12-03 19:48:58 +01:00
COPYING	Add authors.	2011-07-13 20:16:23 +08:00
INSTALL	Add description of installation.	2011-07-11 23:14:57 +08:00
README.md	README: reorganize support list by alphabetic order.	2015-12-04 03:33:22 +01:00
uchardet.pc.in	Refine Description in pkgconfig file	2015-09-21 09:37:36 +08:00

README.md

uchardet

uchardet is an encoding detector library, which takes a sequence of bytes in an unknown character encoding without any additional information, and attempts to determine the encoding of the text. Returned encoding names are iconv-compatible.

uchardet started as a C language binding of the original C++ implementation of the universal charset detection library by Mozilla. It can now detect more charsets, and more reliably than the original implementation.

The original code of universalchardet is available at http://lxr.mozilla.org/seamonkey/source/extensions/universalchardet/

Techniques used by universalchardet are described at http://www.mozilla.org/projects/intl/UniversalCharsetDetection.html

Supported Languages/Encodings

International (Unicode)
- UTF-8
- UTF-16BE / UTF-16LE
- UTF-32BE / UTF-32LE / X-ISO-10646-UCS-4-34121 / X-ISO-10646-UCS-4-21431
Bulgarian
- ISO-8859-5
- WINDOWS-1251
Chinese
- ISO-2022-CN
- BIG5
- EUC-TW
- GB18030
- HZ-GB-2312
English
- ASCII
Esperanto
- ISO-8859-3
French
- ISO-8859-1
- ISO-8859-15
- WINDOWS-1252
German
- ISO-8859-1
- WINDOWS-1252
Greek
- ISO-8859-7
- WINDOWS-1253
Hebrew
- ISO-8859-8
- WINDOWS-1255
Hungarian:
- ISO-8859-2
- WINDOWS-1250
Japanese
- ISO-2022-JP
- SHIFT_JIS
- EUC-JP
Korean
- ISO-2022-KR
- EUC-KR
Russian
- ISO-8859-5
- KOI8-R
- WINDOWS-1251
- MAC-CYRILLIC
- IBM866
- IBM855
Thai
- TIS-620
- ISO-8859-11
Turkish:
- ISO-8859-3
- ISO-8859-9
Others
- WINDOWS-1252

Installation

Debian/Ubuntu/Mint

apt-get install uchardet libuchardet-dev

Mageia

urpmi libuchardet libuchardet-devel

Mac

brew install uchardet

Build from source

cmake .
make
make install

Usage

Command Line

uchardet Command Line Tool
Version 0.0.4

Authors: BYVoid, Jehan
Bug Report: https://github.com/BYVoid/uchardet/issues

Usage:
 uchardet [Options] [File]...

Options:
 -v, --version         Print version and build information.
 -h, --help            Print this help.

Library

See uchardet.h

python-chardet Python port
ruby-rchardet Ruby port
juniversalchardet Java port of universalchardet
jchardet Java port of chardet
nuniversalchardet C# port of universalchardet
nchardet C# port of chardet
uchardet-enhanced A fork of mozilla universalchardet
rust-uchardet Rust language binding of uchardet
libchardet Another C/C++ API wrapping Mozilla code.

License

Mozilla Public License Version 1.1