HORIZON HASKELLDocslts/ghc-9.10.x248f8f02026-10-05Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · 248f8f0 · 2026-10-05

Moduletext-icu-0.8.0.5Haskell98

Data.Text.ICU.Char

Access to the Unicode Character Database, implemented as bindings to the International Components for Unicode (ICU) libraries.

Unicode assigns each codepoint (not just assigned character) values for many properties. Most are simple boolean flags, or constants from a small enumerated list. For some, values are relatively more complex types.

For more information see "About the Unicode Character Database" http://www.unicode.org/ucd/ and the ICU User Guide chapter on Properties http://icu-project.org/userguide/properties.html.

  • 36 types
  • 1 class
  • 12 values
  • Packagetext-icu-0.8.0.5
  • Exports49
  • LanguageHaskell98
  • LicenceBSD-3-Clause
  • SourceChar.hsc

Working with character properties

1 declaration

The property function provides the main view onto the Unicode Character Database. Because Unicode character properties have a variety of types, the property function is polymorphic. The type of its first argument dictates the type of its result, by use of the Property typeclass.

For instance, property Alphabetic returns a Bool, while property NFCQuickCheck returns a Maybe Bool.

classclass Property p v | p -> v where
#
Instances22Property, …

Property identifier types

10 declarations
datadata Bool_
#

Constructors

Instances5Enum, Eq, Show, NFData, Property
  • Enum Bool_Defined in text-icu-0.8.0.5 · Data.Text.ICU.Char
  • Eq Bool_Defined in text-icu-0.8.0.5 · Data.Text.ICU.Char
  • Show Bool_Defined in text-icu-0.8.0.5 · Data.Text.ICU.Char
  • NFData Bool_Defined in text-icu-0.8.0.5 · Data.Text.ICU.Char
  • Property Bool_ BoolDefined in text-icu-0.8.0.5 · Data.Text.ICU.Char

Combining class

Normalization checking

Text boundaries

Property value types

9 declarations
datadata BlockCode
#

Descriptions of Unicode blocks.

NoBlockBasicLatinLatin1SupplementLatinExtendedALatinExtendedBIPAExtensionsSpacingModifierLettersCombiningDiacriticalMarksGreekAndCopticCyrillicArmenianHebrewArabicSyriacThaanaDevanagariBengaliGurmukhiGujaratiOriyaTamilTeluguKannadaMalayalamSinhalaThaiLaoTibetanMyanmarGeorgianHangulJamoEthiopicCherokeeUnifiedCanadianAboriginalSyllabicsOghamRunicKhmerMongolianLatinExtendedAdditionalGreekExtendedGeneralPunctuationSuperscriptsAndSubscriptsCurrencySymbolsCombiningDiacriticalMarksForSymbolsLetterlikeSymbolsNumberFormsArrowsMathematicalOperatorsMiscellaneousTechnicalControlPicturesOpticalCharacterRecognitionEnclosedAlphanumericsBoxDrawingBlockElementsGeometricShapesMiscellaneousSymbolsDingbatsBraillePatternsCJKRadicalsSupplementKangxiRadicalsIdeographicDescriptionCharactersCJKSymbolsAndPunctuationHiraganaKatakanaBopomofoHangulCompatibilityJamoKanbunBopomofoExtendedEnclosedCJKLettersAndMonthsCJKCompatibilityCJKUnifiedIdeographsExtensionACJKUnifiedIdeographsYiSyllablesYiRadicalsHangulSyllablesHighSurrogatesHighPrivateUseSurrogatesLowSurrogatesPrivateUseAreaCJKCompatibilityIdeographsAlphabeticPresentationFormsArabicPresentationFormsACombiningHalfMarksCJKCompatibilityFormsSmallFormVariantsArabicPresentationFormsBSpecialsHalfwidthAndFullwidthFormsOldItalicGothicDeseretByzantineMusicalSymbolsMusicalSymbolsMathematicalAlphanumericSymbolsCJKUnifiedIdeographsExtensionBCJKCompatibilityIdeographsSupplementTagsCyrillicSupplementTagalogHanunooBuhidTagbanwaMiscellaneousMathematicalSymbolsASupplementalArrowsASupplementalArrowsBMiscellaneousMathematicalSymbolsBSupplementalMathematicalOperatorsKatakanaPhoneticExtensionsVariationSelectorsSupplementaryPrivateUseAreaASupplementaryPrivateUseAreaBLimbuTaiLeKhmerSymbolsPhoneticExtensionsMiscellaneousSymbolsAndArrowsYijingHexagramSymbolsLinearBSyllabaryLinearBIdeogramsAegeanNumbersUgariticShavianOsmanyaCypriotSyllabaryTaiXuanJingSymbolsVariationSelectorsSupplementAncientGreekMusicalNotationAncientGreekNumbersArabicSupplementBugineseCJKStrokesCombiningDiacriticalMarksSupplementCopticEthiopicExtendedEthiopicSupplementGeorgianSupplementGlagoliticKharoshthiModifierToneLettersNewTaiLueOldPersianPhoneticExtensionsSupplementSupplementalPunctuationSylotiNagriTifinaghVerticalFormsN'KoBalineseLatinExtendedCLatinExtendedDPhagsPaPhoenicianCuneiformCuneiformNumbersAndPunctuationCountingRodNumeralsSundaneseLepchaOlChikiCyrillicExtendedAVaiCyrillicExtendedBSaurashtraKayahLiRejangChamAncientSymbolsPhaistosDiscLycianCarianLydianMahjongTilesDominoTilesSamaritanUnifiedCanadianAboriginalSyllabicsExtendedTaiThamVedicExtensionsLisuBamumCommonIndicNumberFormsDevanagariExtendedHangulJamoExtendedAJavaneseMyanmarExtendedATaiVietMeeteiMayekHangulJamoExtendedBImperialAramaicOldSouthArabianAvestanInscriptionalParthianInscriptionalPahlaviOldTurkicRumiNumeralSymbolsKaithiEgyptianHieroglyphsEnclosedAlphanumericSupplementEnclosedIdeographicSupplementCJKUnifiedIdeographsExtensionCMandaicBatakEthiopicExtendedABrahmiBamumSupplementKanaSupplementPlayingCardsMiscellaneousSymbolsAndPictographsEmoticonsTransportAndMapSymbolsAlchemicalSymbolsCJKUnifiedIdeographsExtensionDArabicExtendedAArabicMathematicalAlphabeticSymbolsChakmaMeeteiMayekExtensionsMeroiticCursiveMeroiticHieroglyphsMiaoSharadaSoraSompengSundaneseSupplementTakriBassaVahCaucasianAlbanianCopticEpactNumbersCombiningDiacriticalMarksExtendedDuployanElbasanGeometricShapesExtendedGranthaKhojkiKhudawadiLatinExtendedELinearAMahajaniManichaeanMendeKikakuiModiMroMyanmarExtendedBNabataeanOldNorthArabianOldPermicOrnamentalDingbatsPahawhHmongPalmyrenePauCinHauPsalterPahlaviShorthandFormatControlsSiddhamSinhalaArchaicNumbersSupplementalArrowsCTirhutaWarangCitiAhomAnatolianHieroglyphsCherokeeSupplementCJKUnifiedIdeographsExtensionEEarlyDynasticCuneiformHatranMultaniOldHungarianSupplementalSymbolsAndPictographsSuttonSignwritingAdlamBhaiksukiCyrillicExtendedCGlagoliticSupplementIdeographicSymbolsAndPunctuationMarchenMongolianSupplementNewaOsageTangutTangutComponentsCjkUnifiedIdeographsExtensionFKanaExtendedAMasaramGondiNushuSoyomboSyriacSupplementZanabazarSquareChessSymbolsDograGeorgianExtendedGunjalaGondiHanifiRohingyaIndicSiyaqNumbersMakasarMayanNumeralsMedefaidrinOldSogdianSogdianEgyptianHieroglyphFormatControlsElymaicNandinagariNyiakengPuachueHmongOttomanSiyaqNumbersSmallKanaExtensionSymbolsAndPictographsExtendedATamilSupplementWanchoChorasmianCjkUnifiedIdeographsExtensionGDivesAkuruKhitanSmallScriptLisuSupplementSymbolsForLegacyComputingTangutSupplementYezidiArabicExtendedBCyproMinoanEthiopicExtendedBKanaExtendedBLatinExtendedFLatinExtendedGOldUyghurTangsaTotoUnifiedCanadianAboriginalSyllabicsExtendedAVithkuqiZnamennyMusicalNotationArabicExtendedCCjkUnifiedIdeographsExtensionHCyrillicExtendedDDevanagariExtendedAKaktovikNumeralsKawiNagMundari
Instances6Bounded, Enum, Eq, Show, NFData, Property
datadata Direction
#
Instances5Enum, Eq, Show, NFData, Property
datadata Decomposition
#
Instances5Enum, Eq, Show, NFData, Property
datadata GeneralCategory
#
Instances5Enum, Eq, Show, NFData, Property
datadata HangulSyllableType
#
Instances5Enum, Eq, Show, NFData, Property
datadata JoiningGroup
#
Instances5Enum, Eq, Show, NFData, Property

Text boundaries

datadata GraphemeClusterBreak
#
Instances5Enum, Eq, Show, NFData, Property
datadata LineBreak
#
Instances5Enum, Eq, Show, NFData, Property
datadata SentenceBreak
#
Instances5Enum, Eq, Show, NFData, Property
datadata WordBreak
#
Instances5Enum, Eq, Show, NFData, Property
Instances5Enum, Eq, Show, NFData, Property

Functions

10 declarations
valuecharFullName :: Char -> String
#

Return the full name of a Unicode character.

Compared to charName, this function gives each Unicode codepoint a unique extended name. Extended names are lowercase followed by an uppercase hexadecimal number, within angle brackets.

valuecharName :: Char -> String
#

Return the name of a Unicode character.

The names of all unassigned characters are empty.

The name contains only "invariant" characters like A-Z, 0-9, space, and '-'.

Find a Unicode character by its full or extended name, and return its codepoint value.

The name is matched exactly and completely.

A Unicode 1.0 name is matched only if it differs from the modern name.

Compared to charFromName, this function gives each Unicode code point a unique extended name. Extended names are lowercase followed by an uppercase hexadecimal number, within angle brackets.

valuecharFromName :: String -> Maybe Char
#

Find a Unicode character by its full name, and return its code point value.

The name is matched exactly and completely.

A Unicode 1.0 name is matched only if it differs from the modern name. Unicode names are all uppercase.

valueisMirrored :: Char -> Bool
#

Determine whether the codepoint has the BidiMirrored property. This property is set for characters that are commonly used in Right-To-Left contexts and need to be displayed with a "mirrored" glyph.

Conversion to numbers

valuedigitToInt :: Char -> Maybe Int
#

Return the decimal digit value of a decimal digit character. Such characters have the general category Nd (decimal digit numbers) and a NumericType of NTDecimal.

No digit values are returned for any Han characters, because Han number characters are often used with a special Chinese-style number format (with characters for powers of 10 in between) instead of in decimal-positional notation. Unicode 4 explicitly assigns Han number characters a NumericType of NTNumeric instead of NTDecimal.

valuenumericValue :: Char -> Maybe Double
#

Return the numeric value for a Unicode codepoint as defined in the Unicode Character Database.

A Double return type is necessary because some numeric values are fractions, negative, or too large to fit in a fixed-width integral type.