uchardet

mirror of https://gitlab.freedesktop.org/uchardet/uchardet.git synced 2025-12-24 12:44:46 +08:00

Author	SHA1	Message	Date
Jehan	6bbe7da1ac	LangModels: add Finnish support. I built models for ISO-8859-1, ISO-8859-4, ISO-8859-9, ISO-8859-13, ISO-8859-15 and WINDOWS-1252, which all contain Finnish letters. Nevertheless most texts in these encoding end up the same (same codepoints for the Finnish glyphs) so I keep only tests for ISO-8859-1 and UTF-8. Models for other encoding may still be useful when processing texts with some symbols, etc.	2016-09-21 18:27:39 +02:00
Jehan	3401ac70d0	LangModels: add Polish support. With the following encodings: ISO-8859-2, ISO-8859-13, ISO-8859-16, Windows-1250, IBM852, MAC-CENTRALEUROPE. Test text from https://pl.wikipedia.org/wiki/Zofia_Holszańska	2016-09-21 17:30:15 +02:00
Jehan	5f9ec3aef0	LangModels: add support for Slovak. Encodings are the same as Czech (Windows-1250, ISO-8859-2 and Mac-CentralEurope) since the resource I found indicate they used the same encodings historically. Also it is to be noted that the test examples' encoding were already properly detected through Czech's models so the languages are definitely very close, even statistically. Nevertheless adding the right models will work better and these get better scores. This will take all its meaning when uchardet will also be used as a language detector (in some not-too-far future, hopefully!). Test text taken from: https://sk.wikipedia.org/wiki/Jupiter	2016-09-21 13:42:20 +02:00
Jehan	26e1cebad1	LangModels: add support for Czech. Encodings: Windows-1250, ISO-8859-2, IBM852 and Mac-CentralEurope. Other encodings are known to have been used for Czech: Kamenicky, KOI-8 CS2 and Cork. But these are uncommon enough that I decided not to support them (especially since I can't find them supported in iconv either, or at least not under an alias which I could recognize). This web page, which contents was made under the Public Domain, is a good reference for encodings which were used historically for Czech and Slovak: http://luki.sdf-eu.org/txt/cs-encodings-faq.html	2016-09-21 03:33:50 +02:00
Jehan	2700cf3a83	LangModels: support for Maltese / ISO-8859-3. Test text from https://mt.wikipedia.org/wiki/Franza.	2016-09-21 02:11:31 +02:00
Jehan	e138839f07	LangModels: add support for Portuguese / ISO-8859-1. I actually added also couples with ISO-8859-9, ISO-8859-15 and Windows-1252. Nevertheless there are no differences on the main characters related to Portuguese so differences will hardly be made and detection will usually return ISO-8859-1 only.	2016-09-21 00:01:07 +02:00
Jehan	ea2f4dd40f	LangModels: new support for Latvian / ISO-8859-13. Test text extracted from: https://lv.wikipedia.org/wiki/Vinsents_van_Gogs	2016-09-20 23:29:53 +02:00
Jehan	7cb3dd9ddd	LangModels: add support for Lithuanian / ISO-8859-13. Test text extracted from https://lt.wikipedia.org/wiki/Vincent_van_Gogh.	2016-09-20 23:09:24 +02:00
Ilya Tumaykin	2a3e41a6c3	cmake: drop useless PACKAGE_NAME redefinition	2016-03-22 01:23:06 +03:00
Ilya Tumaykin	d0e7ddd8ab	cmake: fix library filename and SONAME Make library filename respect the current uchardet version and make library SONAME respect the current major version.	2016-03-22 01:23:05 +03:00
Ilya Tumaykin	ad647d2e0a	cmake: keep compiler definitions in one place	2016-03-22 01:23:05 +03:00
Ilya Tumaykin	29f18210b1	cmake: hardcode less	2016-03-22 01:23:04 +03:00
Ilya Tumaykin	7201835c98	cmake: export UCHARDET_LIBRARY to the topmost scope	2016-03-22 01:23:04 +03:00
Ilya Tumaykin	e7feb35627	cmake: rename UCHARDET_STATIC_{TARGET -> LIBRARY} for clarity	2016-03-22 01:23:04 +03:00
Ilya Tumaykin	1a1f4bfbd8	cmake: rename UCHARDET_{TARGET -> LIBRARY} for clarity	2016-03-22 01:23:03 +03:00
Ilya Tumaykin	31a53570d6	cmake: use GNUInstallDirs cmake module Available in cmake >= 2.8.5.	2016-03-22 01:23:03 +03:00
Ilya Tumaykin	b44be77be6	cmake: uniform indent everywhere Indent with tabs, remove leading/trailing blank lines and spaces.	2016-03-21 01:07:41 +03:00
Jehan	fcc525a64f	Merge pull request #25 from Coacher/master cmake: purge remnants of opencc after b6d872bb	2016-03-17 19:10:39 +01:00
Jehan	d255184609	Merge pull request #24 from wiiaboo/ab-suite Improving build with more options. Building only static possible, uchardet command line tool build can be disabled, bindir can be customized…	2016-03-17 19:09:30 +01:00
Ricardo Constantino (:RiCON)	86755b1f57	CMake: Don't build static more than once	2016-03-16 19:31:00 +00:00
Ricardo Constantino (:RiCON)	b908b689a0	CMake: Add static lib destination to UCHARDET_TARGET	2016-03-16 19:30:54 +00:00
Ricardo Constantino (:RiCON)	81ed86a26b	CMake: Use only CMAKE_INSTALL_BINDIR instead of DIR_BIN This way it always shows up in ccmake, even if not defined. A string is used instead of path because I personally think it makes more sense in the following use-cases: STRING: -DCMAKE_INSTALL_PREFIX=/home/user -DCMAKE_INSTALL_BINDIR=bins installs everything to /home/user/{lib,etc,share,(...)} and executables to ${CMAKE_INSTALL_PREFIX}/bins -DCMAKE_INSTALL_PREFIX=/home/user -DCMAKE_INSTALL_BINDIR=/opt/bin everything to /home/user/{lib,etc,share,(...)} and executables to /opt/bin PATH: -DCMAKE_INSTALL_PREFIX=/home/user -DCMAKE_INSTALL_BINDIR=bins everything to /home/user/{lib,etc,share,(...)} and executables to $(pwd)/bins (!) -DCMAKE_INSTALL_PREFIX=/home/user -DCMAKE_INSTALL_BINDIR=/opt/bin same as STRING	2016-03-16 19:11:33 +00:00
Ilya Tumaykin	aa4c2aeada	cmake: purge remnants of opencc after b6d872bb	2016-03-16 19:43:58 +03:00
Ricardo Constantino (:RiCON)	50b2e0802f	CMake: Allow not building executable	2016-03-16 14:34:03 +00:00
Ricardo Constantino (:RiCON)	6500f09931	CMake: Allow building static-only builds Add stdc++ to static libs in pkg-config	2016-03-16 14:30:15 +00:00
Jehan	923d264470	LangModels: add Danish support (Windows-1252, ISO-8859-1 and ISO-8859-15). Test for ISO-8859-1 is disabled for now since the difference is not big enough, as for characters used in Danish, between ISO-8859-1 and ISO-8859-15. Therefore the first to be declared "wins". Let's see to improve this later. Test contents from: https://da.wikipedia.org/wiki/Eurosymbol https://da.wikipedia.org/wiki/Dansk_%28sprog%29	2016-02-19 19:10:41 +01:00
Jehan	178c6119b8	LangModels: add Windows-1258 support for Vietnamese. I was planning on adding VISCII support as well, but Python encode() method does not have any support for it apparently, so I cannot generate the proper statistics data with the current version of the string.	2016-02-13 02:32:57 +01:00
Jehan	9c3c37517c	LangModels: add Arabic support. Models constructed for ISO-8859-6 and Windows-1256.	2015-12-13 18:42:16 +01:00
Jehan	ffabb65712	LangModels: adding Spanish support. With 3 charsets: ISO-8859-1, ISO-8859-15 and Windows-1252.	2015-12-12 18:54:35 +01:00
Jehan	5691dc59a1	LangModels: rename Cyrillic models to Russian models. Our language models are per-lang, not per script.	2015-12-04 03:27:29 +01:00
Jehan	5ee1c3ee39	LangModels: adding Turkish models for ISO-8859-3 and ISO-8859-9.	2015-12-04 02:35:09 +01:00
Jehan	f0e122b506	LangModels: add Esperanto ISO-8859-3 language model.	2015-12-04 01:35:56 +01:00
Jehan	aa587a64bd	LangModels: adding German models for ISO-8859-1 and Windows-1252.	2015-12-03 23:58:41 +01:00
Jehan	005fd98086	Add initial support for French with ISO-8859-1 and ISO-8859-15. Mostly generated with a script from Wikipedia data (only the typical positive ratio is slightly modified). This is a first test before adding my generating script to the main tree.	2015-11-28 02:14:39 +01:00
Jehan	2106173546	Move all Single-Byte language models to a subdirectory.	2015-11-27 23:11:23 +01:00
Jehan	ad4dfc4be4	Add a BUILD_STATIC CMake option to optionally build a static library. It is still ON by default, which means both shared and static libs will be built and installed (current behavior), but it makes it possible to disable the build of a static lib. Closes https://github.com/BYVoid/uchardet/issues/1.	2015-11-17 18:14:51 +01:00
nu774	f5637b23b8	fix for MinGW build	2015-06-20 12:28:01 +09:00
BYVoid	84284eccf4	Update code from upstream.	2011-07-11 14:42:50 +08:00
BYVoid	331af64156	Add command line interface.	2011-07-10 16:42:38 +08:00
BYVoid	3601900164	Initial release.	2011-07-10 15:04:42 +08:00

40 Commits