HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Moduleregex-posix-0.96.0.2Haskell2010

Text.Regex.Posix

Module that provides the Regex backend that wraps the C POSIX.2 regex api. This is the backend being used by the regex-compat package to replace Text.Regex.

The Text.Regex.Posix module provides a backend for regular expressions. If you import this along with other backends, then you should do so with qualified imports, perhaps renamed for convenience.

If the =~ and =~~ functions are too high level, you can use the compile, regexec, and execute functions from importing either Text.Regex.Posix.String or Text.Regex.Posix.ByteString. If you want to use a low-level Foreign.C.CString interface to the library, then import Text.Regex.Posix.Wrap and use the wrap* functions.

This module is only efficient with ByteString only if it is null terminated, i.e. (Bytestring.last bs)==0. Otherwise the library must make a temporary copy of the ByteString and append the NUL byte.

A String will be converted into a Foreign.C.CString for processing. Doing this repeatedly will be very inefficient.

Note that the posix library works with single byte characters, and does not understand Unicode. If you need Unicode support you will have to use a different backend.

When offsets are reported for subexpression captures, a subexpression that did not match anything (as opposed to matching an empty string) will have its offset set to the unusedRegOffset value, which is (-1).

Benchmarking shows the default regex library on many platforms is very inefficient. You might increase performace by an order of magnitude by obtaining libpcre and regex-pcre or libtre and regex-tre. If you do not need the captured substrings then you can also get great performance from regex-dfa. If you do need the capture substrings then you may be able to use regex-parsec to improve performance.

  • 12 types
  • 5 classes
  • 13 values
classclass Extract source => RegexLike regex source where
#

RegexLike is parametrized on a regular expression type and a source type to run the matching on.

There are default implementations: matchTest and matchOnceText use matchOnce; matchCount and matchAllText use matchAll. Conversely, matchOnce uses matchOnceText and matchAll uses matchAllText. So a minimal complete instance need to provide at least (matchOnce or matchOnceText) and (matchAll or matchAllText). Additional definitions are often provided where they will increase efficiency.

[ c | let notVowel = makeRegex "[^aeiou]" :: Regex, c <- ['a'..'z'], matchTest notVowel [c]  ]

"bcdfghjklmnpqrstvwxyz"

The strictness of these functions is instance dependent.

Methods

  • matchOnce :: regex -> source -> Maybe MatchArray

    This returns the first match in the source (it checks the whole source, not just at the start). This returns an array of (offset,length) index pairs for the match and captured substrings. The offset is 0-based. A (-1) for an offset means a failure to match. The lower bound of the array is 0, and the 0th element is the (offset,length) for the whole match.

  • matchAll :: regex -> source -> [MatchArray]

    matchAll returns a list of matches. The matches are in order and do not overlap. If any match succeeds but has 0 length then this will be the last match in the list.

  • matchCount :: regex -> source -> Int

    matchCount returns the number of non-overlapping matches returned by matchAll.

  • matchTest :: regex -> source -> Bool

    matchTest returns True if there is a match somewhere in the source (it checks the whole source not just at the start).

  • matchAllText :: regex -> source -> [MatchText source]

    This is matchAll with the actual subsections of the source instead of just the (offset,length) information.

  • matchOnceText :: regex -> source -> Maybe (source, MatchText source, source)

    This can return a tuple of three items: the source before the match, an array of the match and captured substrings (with their indices), and the source after the match.

Instances4RegexLike
classclass RegexOptions regex compOpt execOpt | regex -> compOpt execOpt, compOpt -> regex execOpt, execOpt -> regex compOpt where
#

Rather than carry them around separately, the options for how to execute a regex are kept as part of the regex. There are two types of options. Those that can only be specified at compilation time and never changed are compOpt. Those that can be changed later and affect how matching is performed are execOpt. The actually types for these depend on the backend.

Methods

  • blankCompOpt :: compOpt

    No options set at all in the backend.

  • blankExecOpt :: execOpt

    No options set at all in the backend.

  • defaultCompOpt :: compOpt

    Reasonable options (extended, caseSensitive, multiline regex).

  • defaultExecOpt :: execOpt

    Reasonable options (extended, caseSensitive, multiline regex).

  • setExecOpts :: execOpt -> regex -> regex

    Forget old flags and use new ones.

  • getExecOpts :: regex -> execOpt

    Retrieve the current flags.

Instances1RegexOptions
classclass RegexOptions regex compOpt execOpt => RegexMaker regex compOpt execOpt source | regex -> compOpt execOpt, compOpt -> regex execOpt, execOpt -> regex compOpt where
#

RegexMaker captures the creation of the compiled regular expression from a source type and an option type. Methods makeRegexM and makeRegexM report parse errors using MonadError, usually (Either String regex).

The makeRegex function has a default implementation that depends on makeRegexOpts and uses defaultCompOpt and defaultExecOpt. Similarly for makeRegexM and makeRegexOptsM.

There are also default implementations for makeRegexOpts and makeRegexOptsM in terms of each other. So a minimal instance definition needs to only define one of these, hopefully makeRegexOptsM.

Methods

Instances4RegexMaker
classclass RegexLike regex source => RegexContext regex source target where
#

RegexContext is the polymorphic interface to do matching. Since target is polymorphic you may need to supply the type explicitly in contexts where it cannot be inferred.

The monadic matchM version uses fail to report when the regex has no match in source. Two examples:

Here the contest Bool is inferred:

[ c | let notVowel = makeRegex "[^aeiou]" :: Regex, c <- ['a'..'z'], match notVowel [c]  ]

"bcdfghjklmnpqrstvwxyz"

Here the context [String] must be supplied:

let notVowel = (makeRegex "[^aeiou]" :: Regex )
in do { c <- ['a'..'z'] ; matchM notVowel [c] } :: [String]

["b","c","d","f","g","h","j","k","l","m","n","p","q","r","s","t","v","w","x","y","z"]

Methods

Instances30RegexContext, …
typetype MatchOffset = Int
#

0 based index from start of source, or (-1) for unused

classclass Extract source where
#

Extract allows for indexing operations on String or ByteString.

Methods

  • before :: Int -> source -> source

    before is a renamed take.

  • after :: Int -> source -> source

    after is a renamed drop.

  • empty :: source

    When there is no match, this can construct an empty data value.

  • extract :: (Int, Int) -> source -> source

    extract takes an offset and length, and has this default implementation:

      extract (off, len) source = before len (after off source)
    
Instances6Extract
  • Extract ByteStringDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract ByteStringDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract StringDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract TextDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract TextDefined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
  • Extract (Seq a)Defined in regex-base-0.94.0.3 · Text.Regex.Base.RegexLike
newtypenewtype AllMatches (f :: Type -> Type) b
#

Used in results of RegexContext instances.

Constructors

Instances3RegexContext
newtypenewtype AllTextMatches (f :: Type -> Type) b
#

Used in results of RegexContext instances.

Instances6RegexContext
newtypenewtype AllTextSubmatches (f :: Type -> Type) b
#

Used in results of RegexContext instances.

Instances4RegexContext

Wrap, for =~ and =~~, types and constants

14 declarations
datadata Regex
#

A compiled regular expression.

Instances13RegexOptions, RegexMaker, RegexLike, RegexContext, …
newtypenewtype CompOption
#

A bitmapped CInt containing options for compilation of regular expressions. Option values (and their man 3 regcomp names) are

  • compBlank which is a completely zero value for all the flags. This is also the blankCompOpt value.

  • compExtended (REG_EXTENDED) which can be set to use extended instead of basic regular expressions. This is set in the defaultCompOpt value.

  • compNewline (REG_NEWLINE) turns on newline sensitivity: The dot (.) and inverted set [^ ] never match newline, and ^ and $ anchors do match after and before newlines. This is set in the defaultCompOpt value.

  • compIgnoreCase (REG_ICASE) which can be set to match ignoring upper and lower distinctions.

  • compNoSub (REG_NOSUB) which turns off all information from matching except whether a match exists.

Constructors

Instances9Eq, Num, Show, Bits, RegexOptions, RegexMaker, …
newtypenewtype ExecOption
#

A bitmapped CInt containing options for execution of compiled regular expressions. Option values (and their man 3 regexec names) are

  • execBlank which is a complete zero value for all the flags. This is the blankExecOpt value.

  • execNotBOL (REG_NOTBOL) can be set to prevent ^ from matching at the start of the input.

  • execNotEOL (REG_NOTEOL) can be set to prevent $ from matching at the end of the input (before the terminating NUL).

Constructors

Instances9Eq, Num, Show, Bits, RegexOptions, RegexMaker, …