HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Moduleregex-tdfa-1.3.2.5Haskell2010

Text.Regex.TDFA

SPDX-License-Identifier: BSD-3-Clause Maintainer: Andreas Abel Stability: stable

The Text.Regex.TDFA module provides a backend for regular expressions. It provides instances for the classes defined and documented in Text.Regex.Base and re-exported by this module. If you import this along with other backends then you should do so with qualified imports (with renaming for convenience).

This regex-tdfa package implements, correctly, POSIX extended regular expressions. It is highly unlikely that the regex-posix package on your operating system is correct, see http://www.haskell.org/haskellwiki/Regex_Posix for examples of your OS's bugs.

Importing and using

Declare a dependency on the regex-tdfa library in your .cabal file:

build-depends: regex-tdfa ^>= 1.3.2

In Haskell modules where you want to use regexes simply import this module:

import Text.Regex.TDFA

Basics

Example2 expressions
let emailRegex = "[a-zA-Z0-9+._-]+\\@[-a-zA-Z]+\\.[a-z]+""my email is first-name.lastname_1974@e-mail.com" =~ emailRegex :: BoolTrue
Example1 expression
"invalid@mail@com" =~ emailRegex :: BoolFalse
Example1 expression
"invalid@mail.COM" =~ emailRegex :: BoolFalse
Example1 expression
"#@invalid.com" =~ emailRegex :: BoolFalse
-- non-monadic
λ> <to-match-against> =~ <regex>

-- monadic, uses fail on lack of match
λ> <to-match-against> =~~ <regex>

(=~) and (=~~) are polymorphic in their return type. This is so that regex-tdfa can pick the most efficient way to give you your result based on what you need. For instance, if all you want is to check whether the regex matched or not, there's no need to allocate a result string. If you only want the first match, rather than all the matches, then the matching engine can stop after finding a single hit.

This does mean, though, that you may sometimes have to explicitly specify the type you want, especially if you're trying things out at the REPL.

Common use cases

Get the first match

-- returns empty string if no match
a =~ b :: String  -- or ByteString, or Text...
Example1 expression
"alexis-de-tocqueville" =~ "[a-z]+" :: String"alexis"
Example1 expression
"alexis-de-tocqueville" =~ "[0-9]+" :: String""

Check if it matched at all

a =~ b :: Bool
Example1 expression
"alexis-de-tocqueville" =~ "[a-z]+" :: BoolTrue

Get first match + text before/after

-- if no match, will just return whole
-- string in the first element of the tuple
a =~ b :: (String, String, String)
Example1 expression
"alexis-de-tocqueville" =~ "de" :: (String, String, String)("alexis-","de","-tocqueville")
Example1 expression
"alexis-de-tocqueville" =~ "kant" :: (String, String, String)("alexis-de-tocqueville","","")

Get first match + submatches

-- same as above, but also returns a list of just submatches.
-- submatch list is empty if regex doesn't match at all
a =~ b :: (String, String, String, [String])
Example1 expression
"div[attr=1234]" =~ "div\\[([a-z]+)=([^]]+)\\]" :: (String, String, String, [String])("","div[attr=1234]","",["attr","1234"])

Get all matches

-- can also return Data.Array instead of List
getAllTextMatches (a =~ b) :: [String]
Example1 expression
getAllTextMatches ("john anne yifan" =~ "[a-z]+") :: [String]["john","anne","yifan"]
Example1 expression
getAllTextMatches ("* - . a + z" =~ "[--z]+") :: [String]["-",".","a","z"]

Feature support

This package does provide captured parenthesized subexpressions.

Depending on the text being searched this package supports Unicode. The [Char], Text, Text.Lazy, and (Seq Char) text types support Unicode. The ByteString and ByteString.Lazy text types only support ASCII.

As of version 1.1.1 the following GNU extensions are recognized, all anchors:

  • \` at beginning of entire text

  • \' at end of entire text

  • \< at beginning of word

  • \> at end of word

  • \b at either beginning or end of word

  • \B at neither beginning nor end of word

The above are controlled by the newSyntax Bool in CompOption.

Where the "word" boundaries means between characters that are and are not in the [:word:] character class which contains [a-zA-Z0-9_]. Note that \< and \b may match before the entire text and \> and \b may match at the end of the entire text.

There is no locale support, so collating elements like [.ch.] are simply ignored and equivalence classes like [=a=] are converted to just [a]. The character classes like [:alnum:] are supported over ASCII only, valid classes are alnum, digit, punct, alpha, graph, space, blank, lower, upper, cntrl, print, xdigit, word.

Example1 expression
getAllTextMatches ("john anne yifan" =~ "[[:lower:]]+") :: [String]["john","anne","yifan"]

This package does not provide "basic" regular expressions. This package does not provide back references inside regular expressions.

The package does not provide Perl style regular expressions. Please look at the regex-pcre and pcre-light packages instead.

This package does not provide find-and-replace.

Avoiding backslashes

If you find yourself writing a lot of regexes, take a look at raw-strings-qq. It'll let you write regexes without needing to escape all your backslashes.

{-# LANGUAGE QuasiQuotes #-}

import Text.RawString.QQ
import Text.Regex.TDFA

λ> "2 * (3 + 1) / 4" =~ [r|\([^)]+\)|] :: String
"(3 + 1)"
  • 12 types
  • 5 classes
  • 4 values
value(=~)
  1. :: (RegexMaker Regex CompOption ExecOption source, RegexContext Regex source1 target)
  2. => source1
  3. -> source
  4. -> target
#

This is the pure functional matching operator. If the target cannot be produced then some empty result will be returned. If there is an error in processing, then error will be called.

datadata Regex
#

The TDFA backend specific Regex type, used by this module's RegexOptions and RegexMaker.

Instances19RegexOptions, RegexMaker, RegexLike, RegexContext, …
datadata CompOption
#

Control whether the pattern is multiline or case-sensitive like Text.Regex and whether to capture the subgroups (\1, \2, etc). Controls enabling extra anchor syntax.

Constructors

Instances9Read, Show, RegexOptions, RegexMaker, …
datadata ExecOption
#

Constructors

Instances9Read, Show, RegexOptions, RegexMaker, …
classclass RegexLike regex source => RegexContext regex source target where
#

RegexContext is the polymorphic interface to do matching. Since target is polymorphic you may need to supply the type explicitly in contexts where it cannot be inferred.

The monadic matchM version uses fail to report when the regex has no match in source. Two examples:

Here the contest Bool is inferred:

[ c | let notVowel = makeRegex "[^aeiou]" :: Regex, c <- ['a'..'z'], match notVowel [c]  ]

"bcdfghjklmnpqrstvwxyz"

Here the context [String] must be supplied:

let notVowel = (makeRegex "[^aeiou]" :: Regex )
in do { c <- ['a'..'z'] ; matchM notVowel [c] } :: [String]

["b","c","d","f","g","h","j","k","l","m","n","p","q","r","s","t","v","w","x","y","z"]

Methods

Instances32RegexContext, …
classclass RegexOptions regex compOpt execOpt => RegexMaker regex compOpt execOpt source | regex -> compOpt execOpt, compOpt -> regex execOpt, execOpt -> regex compOpt where
#

RegexMaker captures the creation of the compiled regular expression from a source type and an option type. Methods makeRegexM and makeRegexM report parse errors using MonadError, usually (Either String regex).

The makeRegex function has a default implementation that depends on makeRegexOpts and uses defaultCompOpt and defaultExecOpt. Similarly for makeRegexM and makeRegexOptsM.

There are also default implementations for makeRegexOpts and makeRegexOptsM in terms of each other. So a minimal instance definition needs to only define one of these, hopefully makeRegexOptsM.

Methods

Instances6RegexMaker
classclass Extract source => RegexLike regex source where
#

RegexLike is parametrized on a regular expression type and a source type to run the matching on.

There are default implementations: matchTest and matchOnceText use matchOnce; matchCount and matchAllText use matchAll. Conversely, matchOnce uses matchOnceText and matchAll uses matchAllText. So a minimal complete instance need to provide at least (matchOnce or matchOnceText) and (matchAll or matchAllText). Additional definitions are often provided where they will increase efficiency.

[ c | let notVowel = makeRegex "[^aeiou]" :: Regex, c <- ['a'..'z'], matchTest notVowel [c]  ]

"bcdfghjklmnpqrstvwxyz"

The strictness of these functions is instance dependent.

Methods

  • matchOnce :: regex -> source -> Maybe MatchArray

    This returns the first match in the source (it checks the whole source, not just at the start). This returns an array of (offset,length) index pairs for the match and captured substrings. The offset is 0-based. A (-1) for an offset means a failure to match. The lower bound of the array is 0, and the 0th element is the (offset,length) for the whole match.

  • matchAll :: regex -> source -> [MatchArray]

    matchAll returns a list of matches. The matches are in order and do not overlap. If any match succeeds but has 0 length then this will be the last match in the list.

  • matchCount :: regex -> source -> Int

    matchCount returns the number of non-overlapping matches returned by matchAll.

  • matchTest :: regex -> source -> Bool

    matchTest returns True if there is a match somewhere in the source (it checks the whole source not just at the start).

  • matchAllText :: regex -> source -> [MatchText source]

    This is matchAll with the actual subsections of the source instead of just the (offset,length) information.

  • matchOnceText :: regex -> source -> Maybe (source, MatchText source, source)

    This can return a tuple of three items: the source before the match, an array of the match and captured substrings (with their indices), and the source after the match.

Instances6RegexLike
classclass RegexOptions regex compOpt execOpt | regex -> compOpt execOpt, compOpt -> regex execOpt, execOpt -> regex compOpt where
#

Rather than carry them around separately, the options for how to execute a regex are kept as part of the regex. There are two types of options. Those that can only be specified at compilation time and never changed are compOpt. Those that can be changed later and affect how matching is performed are execOpt. The actually types for these depend on the backend.

Methods

  • blankCompOpt :: compOpt

    No options set at all in the backend.

  • blankExecOpt :: execOpt

    No options set at all in the backend.

  • defaultCompOpt :: compOpt

    Reasonable options (extended, caseSensitive, multiline regex).

  • defaultExecOpt :: execOpt

    Reasonable options (extended, caseSensitive, multiline regex).

  • setExecOpts :: execOpt -> regex -> regex

    Forget old flags and use new ones.

  • getExecOpts :: regex -> execOpt

    Retrieve the current flags.

Instances1RegexOptions
typetype MatchOffset = Int
#

0 based index from start of source, or (-1) for unused

classclass Extract source where
#

Extract allows for indexing operations on String or ByteString.

Methods

  • before :: Int -> source -> source

    before is a renamed take.

  • after :: Int -> source -> source

    after is a renamed drop.

  • empty :: source

    When there is no match, this can construct an empty data value.

  • extract :: (Int, Int) -> source -> source

    extract takes an offset and length, and has this default implementation:

      extract (off, len) source = before len (after off source)
    
Instances6Extract
  • Extract ByteStringDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract ByteStringDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract StringDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract TextDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract TextDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract (Seq a)Defined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
newtypenewtype AllMatches (f :: Type -> Type) b
#

Used in results of RegexContext instances.

Constructors

Instances3RegexContext
newtypenewtype AllTextMatches (f :: Type -> Type) b
#

Used in results of RegexContext instances.

Instances6RegexContext
newtypenewtype AllTextSubmatches (f :: Type -> Type) b
#

Used in results of RegexContext instances.

Instances4RegexContext