uchardet

mirror of https://gitlab.freedesktop.org/uchardet/uchardet.git synced 2025-12-10 02:46:40 +08:00

Author	SHA1	Message	Date
Jehan	e0eec3bae8	src: give a little weight to "probable sequences". Up to now, we were only considering positive sequences, which are sequences of 2 characters which happen the most. Yet our data gather 4 categories of sequences (the last one being called "negative", since they never happened in our data). I will call the category below positive: probable sequences. They may happen, yet not often. The last category could be called "neutral". This seems to fix the detection of a user's subtitle example without breaking any of our current unit tests. Probably I should still review this whole logics more in details later.	2016-05-25 17:38:20 +02:00
Jehan	4287d3accc	src: trailing whitespace removed.	2016-05-25 16:07:17 +02:00
Jehan	55b4f23971	Single Byte charsets: high ctrl character ratio lowers confidence. Control characters are not an error per-se. Nevertheless they are clearly not frequent in single-byte charset texts. It is only normal for them to lower confidence in a charset. In particular a higher ctrl-per-letter ratio means a lower confidence. This fixes for instance our Windows-1252 German test (otherwise detected as ISO-8859-1).	2015-12-04 00:04:43 +01:00
Jehan	c4fa728e7a	Merge branch 'master' of https://github.com/lovasoa/uchardet into lovasoa-master Let's shortcut Single Byte charset detection on invalid codepoints. Merging and fixing the contributor's commit conflicts after code redesign: in particular we added an illegal character concept (they were mixed with control characters in current charmaps. Yet ctrl characters are NOT to be considered invalid) and constants instead of hardcoded numbers ('ILL' rather than 255).	2015-12-03 19:26:19 +01:00
Jehan	4f1c3ff85e	nsSBCharSetProber: multiply confidence by ratio of positive seqs per chars. If all sequences in a text are positive sequences, the ratio of positive sequences cannot make the difference between 2 very close charsets. A ratio of positive sequences per letters on the other hand will change a tie between 2 encoding. If while adding a letter, the number of positive sequences does not increase, the confidence will decrease (corresponding to the fact it was likely not a letter). On the other hand, if the number of positive sequences increase, so will the confidence. For instance this fixes wrong detections of ISO-8859-1 and ISO-8859-15. When letters only available in ISO-8859-15 appear in a text, we expect confidence to tilt towards the close yet slightly different ISO-8859-15.	2015-11-30 19:52:07 +01:00
Jehan	dbb4c1d2ff	nsSBCharSetProber: replace the fixed 64 SAMPLE_SIZE... ... with per-language model "frequent character" count.	2015-11-29 23:51:55 +01:00
Ophir LOJKINE	5ef60164fc	Stop detection early on control characters	2015-11-24 22:07:41 +03:00
BYVoid	3601900164	Initial release.	2011-07-10 15:04:42 +08:00

8 Commits