HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Moduleunicode-collation-0.1.3.6Haskell2010

Text.Collate

SPDX-License-Identifier: BSD-2-Clause Maintainer: John MacFarlane jgm@berkeley.edu

This library provides a pure Haskell implementation of the Unicode Collation Algorithm, allowing proper sorting of Unicode strings.

The simplest way to use the library is to use the IsString instance of Collator (together with the OverloadedStrings extension):

Example3 expressions
import Data.List (sortBy)import qualified Data.Text.IO as TmapM_ T.putStrLn $ sortBy (collate "en-US") ["𝒶bc","abC","𝕒bc","Abc","abç","äbc"]abC𝒶bc𝕒bcAbcabçäbc

Note the difference from the default sort:

Example3 expressions
import Data.List (sort)import qualified Data.Text.IO as TmapM_ T.putStrLn $ sort ["𝒶bc","abC","𝕒bc","Abc","abç","äbc"]AbcabCabçäbc𝒶bc𝕒bc

A Collator provides a function collate that compares two texts, and a function sortKey that returns the sort key. Most users will just need collate.

Example6 expressions
let de = collatorFor "de"let se = collatorFor "se"collate de "ö" "z"LTcollate se "ö" "z"GTsortKey de "ö"SortKey [0x213C,0x0000,0x0020,0x002B,0x0000,0x0002,0x0002]sortKey se "ö"SortKey [0x22FD,0x0000,0x0020,0x0000,0x0002]

To sort a string type other than Text, the function collateWithUnpacker may be used. It takes as a parameter a function that lazily unpacks the string type into a list of Char.

Example2 expressions
let seCollateString = collateWithUnpacker "se" idseCollateString ("ö" :: String) ("z" :: String)GT

Because Collator and Lang have IsString instances, you can just specify them using string literals, as in the above examples. Note, however, that you won't get any feedback if the string doesn't parse correctly as a BCP47 language tag, or if no collation is defined for the specified language; instead, you'll just get the default (root) collator. For this reason, we don't recommend relying on the IsString instance.

If you won't know the language until run time, use parseLang to parse it to a Lang, handling parse errors, and then pass the Lang to collatorFor.

Example3 expressions
let handleParseError = error  -- or something fancierlang <- either handleParseError return $ parseLang "bs-Cyrl"collate (collatorFor lang) "a" "b"LT

If you know the language at compile-time, use the collator quasi-quoter and you'll get compile-time errors and warnings:

Example5 expressions
:set -XQuasiQuoteslet esTraditional = [collator|es-u-co-trad|]let esStandard = [collator|es|]collate esStandard "Co" "Ch"GTcollate esTraditional "Co" "Ch"LT

Note that the unicode extension syntax for BCP47 can be used to specify a particular collation for the language (here, Spanish "traditional" instead of the default ordering; the alias trad is used because of length limits for BCP47 keywords).

The extension syntax can also be used to set collator options. The keyword kb can be used to specify the "backwards" accent sorting that is sometimes used in French:

Example2 expressions
collate "fr" "côte" "coté"GTcollate "fr-u-kb" "côte" "coté"LT

The keyword ka can be used to specify the variable weighting options which affect how punctuation and whitespace are treated:

Example2 expressions
collate "en-u-ka-shifted" "de-luge" "de Luge"LTcollate "en-u-ka-noignore" "de-luge" "de Luge"GT

The keyword kk can be used to turn off the normalization step (which is required by the algorithm but can be omitted for better performance if the input is already in NFD form (canonical decomposition).

Example1 expression
let noNormalizeCollator = [collator|en-u-kk-false|]

The keyword kf can be used to say whether uppercase or lowercase letters should be sorted first.

Example2 expressions
collate "en-u-kf-upper" "A" "a"LTcollate "en-u-kf-lower" "A" "a"GT

These options be combined:

Example1 expression
collate "de-DE-u-co-phonebk-kb-false-ka-shifted" "Udet" "Über"LT

Options can also be set using the functions setVariableWeighting, setNormalization, setUpperBeforeLower, and setFrenchAccents:

Example2 expressions
let frC = setFrenchAccents True [collator|fr|]collate frC "côte" "coté"LT
  • 4 types
  • 14 values
valuecollatorFor :: Lang -> Collator
#

Returns a collator based on a BCP 47 language tag. If no exact match is found, we try to find the best match (falling back to the root collation if nothing else succeeds). If something other than the default collation for a language is desired, the co keyword of the unicode extensions can be used (e.g. es-u-co-trad for traditional Spanish). Other unicode extensions affect the collator options:

  • The kb keyword has the same effect as setFrenchAccents (e.g. fr-FR-u-kb-true).

  • The ka keyword has the same effect as setVariableWeight (e.g. fr-FR-u-kb-ka-shifted or en-u-ka-noignore).

  • The kf keyword has the same effect as setUpperBeforeLower (e.g. fr-u-kf-upper or fr-u-kf-lower).

  • The kk keyword has the same effect as setNormalization (e.g. fr-u-kk-false).

valuecollator :: QuasiQuoter
#

Create a collator at compile time based on a BCP 47 language tag: e.g., [collator|es-u-co-trad|]. Requires the QuasiQuotes extension.

newtypenewtype SortKey
#

Constructors

Instances3Eq, Ord, Show
  • Eq SortKeyDefined in unicode-collation-0.1.3.6 · Text.Collate.Collator
  • Ord SortKeyDefined in unicode-collation-0.1.3.6 · Text.Collate.Collator
  • Show SortKeyDefined in unicode-collation-0.1.3.6 · Text.Collate.Collator
valuerenderSortKey :: SortKey -> String
#

Render sort key in the manner used in the CLDR collation test data: the character | is used to separate the levels of the key and corresponds to a 0 in the actual sort key.

datadata VariableWeighting
#

Constructors

  • NonIgnorable

    Don't ignore punctuation (Deluge < deluge-)

  • Blanked

    Completely ignore punctuation (Deluge = deluge-)

  • Shifted

    Consider punctuation at lower priority (de-luge < delu-ge < deluge < deluge- < Deluge)

  • ShiftTrimmed

    Variant of Shifted (deluge < de-luge < delu-ge)

Instances3Eq, Ord, Show
datadata CollatorOptions
#

Constructors

Instances3Eq, Ord, Show

Deprecated. Use (optLang . collatorOptions)

Lang used for tailoring. Because of fallback rules, this may be somewhat different from the Lang passed to collatorFor. This Lang won't contain unicode extensions used to set options, but it will specify the collation if a non-default collation is being used.

The Unicode Collation Algorithm expects input to be normalized into its canonical decomposition (NFD). By default, collators perform this normalization. If your input is already normalized, you can increase performance by disabling this step: setNormalization False.

setFrenchAccents True causes secondary weights to be scanned in reverse order, so we get the sorting cote côte coté côté instead of cote coté côte côté. The default is usually False, except for fr-CA where it is True.

Most collations default to sorting lowercase letters before uppercase (exceptions: mt, da, cu). To select the opposite behavior, use setUpperBeforeLower True.

valuetailorings :: [(Lang, Collation)]
#

An association list matching Langs with tailored Collations.