String collator type.
Moduletext-icu-0.8.0.5Haskell98
Data.Text.ICU.Collate
String collation functions for Unicode, implemented as bindings to the International Components for Unicode (ICU) libraries.
- 5 types
- 10 values
- Packagetext-icu-0.8.0.5
- Exports15
- LanguageHaskell98
- LicenceBSD-3-Clause
- SourceCollate.hsc
Unicode collation API
5 declarationsConstructors
French BoolAlternateHandling AlternateHandlingFor handling variable elements. NonIgnorable is default.
CaseFirst (Maybe CaseFirst)Control the ordering of upper and lower case letters. Nothing (the default) orders upper and lower case letters in accordance to their tertiary weights.
CaseLevel BoolControls whether an extra case level (positioned before the third level) is generated or not. When False (default), case level is not generated; when True, the case level is generated. Contents of the case level are affected by the value of the CaseFirst attribute. A simple way to ignore accent differences in a string is to set the strength to Primary and enable case level.
NormalizationMode BoolControls whether the normalization check and necessary normalizations are performed. When False (default) no normalization check is performed. The correctness of the result is guaranteed only if the input data is in so-called
FCDform (see users manual for more info). When True, an incremental check is performed to see whether the input data is inFCDform. If the data is not inFCDform, incrementalNFDnormalization is performed.Strength StrengthHiraganaQuaternaryMode BoolWhen turned on, this attribute positions Hiragana before all non-ignorables on quaternary level. This is a sneaky way to produce JIS sort order.
Numeric BoolWhen enabled, this attribute generates a collation key for the numeric value of substrings of digits. This is a way to get '100' to sort after '2'.
Control the handling of variable weight elements.
Constructors
NonIgnorableTreat all codepoints with non-ignorable primary weights in the same way.
ShiftedCause codepoints with primary weights that are equal to or below the variable top value to be ignored on primary level and moved to the quaternary level.
Instances5Bounded, Enum, Eq, Show, NFData
Bounded AlternateHandlingDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateEnum AlternateHandlingDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateEq AlternateHandlingDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateShow AlternateHandlingDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateNFData AlternateHandlingDefined in text-icu-0.8.0.5 · Data.Text.ICU.Collate
Control the ordering of upper and lower case letters.
Constructors
UpperFirstForce upper case letters to sort before lower case.
LowerFirstForce lower case letters to sort before upper case.
Instances5Bounded, Enum, Eq, Show, NFData
Bounded CaseFirstDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateEnum CaseFirstDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateEq CaseFirstDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateShow CaseFirstDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateNFData CaseFirstDefined in text-icu-0.8.0.5 · Data.Text.ICU.Collate
The strength attribute. The usual strength for most locales (except
Japanese) is tertiary. Quaternary strength is useful when combined with
shifted setting for alternate handling attribute and for JIS x 4061
collation, when it is used to distinguish between Katakana and Hiragana
(this is achieved by setting HiraganaQuaternaryMode mode to
True). Otherwise, quaternary level is affected only by the number of
non ignorable codepoints in the string. Identical strength is rarely
useful, as it amounts to codepoints of the NFD form of the string.
Instances5Bounded, Enum, Eq, Show, NFData
Bounded StrengthDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateEnum StrengthDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateEq StrengthDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateShow StrengthDefined in text-icu-0.8.0.5 · Data.Text.ICU.CollateNFData StrengthDefined in text-icu-0.8.0.5 · Data.Text.ICU.Collate
Functions
4 declarationsOpen a Collator for comparing strings.
openRules :: TextA string describing the collation rules.
-> Maybe BoolThe normalization mode: One of 'Just False' (expect the text to not need normalization) 'Just True' (normalize), or Nothing (set the mode according to the rules)
-> Maybe StrengthThe default collation strength; one of 'Just Primary', 'Just Secondary', 'Just Tertiary', 'Just Identical', Nothing (default strength) - can be also set in the rules.
-> IO MCollator
Produce a Collator instance according to the rules supplied.
Compare two strings.
Compare two CharIterators.
If either iterator was constructed from a ByteString, it does not need to be copied or converted internally, so this function can be quite cheap.
Utility functions
Get the rules of an MCollator attribute.
Set the value of an MCollator attribute.