uchardet/script/BuildLangModelLogs
Jehan 338a51564a src, script: add concept of alphabet_mapping in language models.
This allows to handle cases where some characters are actually
alternative/variants of another. For instance, a same word can be
written with both variants, while both are considered correct and
equivalent. Browsing a bit Slovenian Wikipedia, it looks like they only
use them for titles there.

I use this the first time on characters with diacritics in Slovene.
Indeed these are so rarely used that they would hardly show in the stats
and worse, any sequence using these in tested text would likely show as
negative sequences hence drop the confidence in Slovenian. As a
consequence, various Slovene text would show up as Slovak as it's close
enough and contains the same character with diacritics in a common way.
2022-12-14 00:24:53 +01:00
..
LangArabicModel.log Rebuild a bunch of language models. 2022-12-14 00:23:13 +01:00
LangCroatianModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangCzechModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangDanishModel.log Rebuild a bunch of language models. 2022-12-14 00:23:13 +01:00
LangEsperantoModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangEstonianModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangFinnishModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangFrenchModel.log Rebuild a bunch of language models. 2022-12-14 00:23:13 +01:00
LangGermanModel.log Rebuild a bunch of language models. 2022-12-14 00:23:13 +01:00
LangGreekModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangHebrewModel.log script, src: generate the Hebrew models. 2022-12-14 00:23:13 +01:00
LangHindiModel.log src: add Hindi/UTF-8 support. 2022-12-14 00:23:13 +01:00
LangHungarianModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangIrishModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangItalianModel.log Rebuild a bunch of language models. 2022-12-14 00:23:13 +01:00
LangKoreanModel.log script, src: add generic Korean model. 2022-12-14 00:23:13 +01:00
LangLatvianModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangLithuanianModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangMalteseModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangPolishModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangPortugueseModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangRomanianModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangSlovakModel.log script: regenerate Slovak and Slovene with better alphabet support. 2022-12-14 00:24:53 +01:00
LangSloveneModel.log src, script: add concept of alphabet_mapping in language models. 2022-12-14 00:24:53 +01:00
LangSpanishModel.log Rebuild a bunch of language models. 2022-12-14 00:23:13 +01:00
LangSwedishModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangThaiModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangTurkishModel.log src, script: regenerate all existing language models. 2022-12-14 00:23:13 +01:00
LangVietnameseModel.log script, src: regenerate the Vietnamese model. 2022-12-14 00:24:53 +01:00