HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Modulemegaparsec-9.7.0Haskell2010

Text.Megaparsec

This module includes everything you need to get started writing a parser. If you are new to Megaparsec and don't know where to begin, take a look at the tutorial https://markkarpov.com/tutorial/megaparsec.html.

In addition to the Text.Megaparsec module, which exports and re-exports almost everything that you may need, we advise to import Text.Megaparsec.Char if you plan to work with a stream of Char tokens or Text.Megaparsec.Byte if you intend to parse binary data.

It is common to start working with the library by defining a type synonym like this:

type Parser = Parsec Void Text
                     ^    ^
                     |    |
Custom error component    Input stream type

Then you can write type signatures like Parser Int—for a parser that returns an Int for example.

Similarly (since it's known to cause confusion), you should use ParseErrorBundle type parametrized like this:

ParseErrorBundle Text Void
                 ^    ^
                 |    |
 Input stream type    Custom error component (the same you used in Parser)

Megaparsec uses some type-level machinery to provide flexibility without compromising on type safety. Thus type signatures are sometimes necessary to avoid ambiguous types. If you're seeing an error message that reads like “Type variable e0 is ambiguous …”, you need to give an explicit signature to your parser to resolve the ambiguity. It's a good idea to provide type signatures for all top-level definitions.

  • 4 types
  • 1 class
  • 56 values
  • Packagemegaparsec-9.7.0
  • Exports63
  • LanguageHaskell2010
  • LicenceBSD-2-Clause
  • SourceMegaparsec.hs

Re-exports

26 declarations

Note that we re-export monadic combinators from Control.Monad.Combinators because these are more efficient than Applicative-based ones (†). Thus many and some may clash with the functions from Control.Applicative. You need to hide the functions like this:

import Control.Applicative hiding (many, some)

† As of Megaparsec 9.7.0 many and some are as efficient as their monadic counterparts.

Also note that you can import Control.Monad.Combinators.NonEmpty if you wish that combinators like some return NonEmpty lists. The module lives in the parser-combinators package (you need at least version 0.4.0).

This module is intended to be imported qualified:

import qualified Control.Monad.Combinators.NonEmpty as NE

Other modules of interest are:

method(<|>) :: f a -> f a -> f a
#

An associative binary operation

methodempty :: f a
#

The identity of <|>

empty <|> a     == a
a     <|> empty == a
valuemany :: MonadPlus m => m a -> m [a]
#

many p applies the parser p zero or more times and returns a list of the values returned by p.

identifier = (:) <$> letter <*> many (alphaNumChar <|> char '_')
valuesome :: MonadPlus m => m a -> m [a]
#

some p applies the parser p one or more times and returns a list of the values returned by p.

word = some letter
valueoptional :: Alternative f => f a -> f (Maybe a)
#

One or none.

It is useful for modelling any computation that is allowed to fail.

Examples

Using the Alternative instance of Control.Monad.Except, the following functions:

Example1 expression
import Control.Monad.Except
Example2 expressions
canFail = throwError "it failed" :: Except String Intfinal = return 42                :: Except String Int

Can be combined by allowing the first function to fail:

Example1 expression
runExcept $ canFail *> finalLeft "it failed"
Example1 expression
runExcept $ optional canFail *> finalRight 42
valuebetween :: Applicative m => m open -> m close -> m a -> m a
#

between open close p parses open, followed by p and close. Returns the value returned by p.

braces = between (symbol "{") (symbol "}")
valuechoice :: (Foldable f, Alternative m) => f (m a) -> m a
#

choice ps tries to apply the parsers in the list ps in order, until one of them succeeds. Returns the value of the succeeding parser.

choice = asum
valuecount :: Monad m => Int -> m a -> m [a]
#

count n p parses n occurrences of p. If n is smaller or equal to zero, the parser equals to return []. Returns a list of n values.

See also: skipCount, count'.

valueendBy :: MonadPlus m => m a -> m sep -> m [a]
#

endBy p sep parses zero or more occurrences of p, separated and ended by sep. Returns a list of values returned by p.

cStatements = cStatement `endBy` semicolon
valueendBy1 :: MonadPlus m => m a -> m sep -> m [a]
#

endBy1 p sep parses one or more occurrences of p, separated and ended by sep. Returns a list of values returned by p.

valuemanyTill :: MonadPlus m => m a -> m end -> m [a]
#

manyTill p end applies parser p zero or more times until parser end succeeds. Returns the list of values returned by p. Note that end result is consumed and lost. Use manyTill_ if you wish to keep it.

See also: skipMany, skipManyTill.

valueoption :: Alternative m => a -> m a -> m a
#

option x p tries to apply the parser p. If p fails without consuming input, it returns the value x, otherwise the value returned by p.

option x p = p <|> pure x

See also: optional.

valuesepBy :: MonadPlus m => m a -> m sep -> m [a]
#

sepBy p sep parses zero or more occurrences of p, separated by sep. Returns a list of values returned by p.

commaSep p = p `sepBy` comma
valuesepBy1 :: MonadPlus m => m a -> m sep -> m [a]
#

sepBy1 p sep parses one or more occurrences of p, separated by sep. Returns a list of values returned by p.

valueeitherP :: Alternative m => m a -> m b -> m (Either a b)
#

Combine two alternatives.

eitherP a b = (Left <$> a) <|> (Right <$> b)
valuecount' :: MonadPlus m => Int -> Int -> m a -> m [a]
#

count' m n p parses from m to n occurrences of p. If n is not positive or m > n, the parser equals to return []. Returns a list of parsed values.

Please note that m may be negative, in this case effect is the same as if it were equal to zero.

See also: skipCount, count.

valuemanyTill_ :: MonadPlus m => m a -> m end -> m ([a], end)
#

manyTill_ p end applies parser p zero or more times until parser end succeeds. Returns the list of values returned by p and the end result. Use manyTill if you have no need in the result of the end.

See also: skipMany, skipManyTill.

valuesepEndBy :: MonadPlus m => m a -> m sep -> m [a]
#

sepEndBy p sep parses zero or more occurrences of p, separated and optionally ended by sep. Returns a list of values returned by p.

valuesepEndBy1 :: MonadPlus m => m a -> m sep -> m [a]
#

sepEndBy1 p sep parses one or more occurrences of p, separated and optionally ended by sep. Returns a list of values returned by p.

Data types

4 declarations
datadata State s e
#

This is the Megaparsec's state parametrized over stream type s and custom error component type e.

Constructors

Instances6Eq, Data, Show, Generic, NFData, Rep
datadata PosState s
#

A special kind of state that is used to calculate line/column positions on demand.

Constructors

Instances6Eq, Data, Show, Generic, NFData, Rep
newtypenewtype ParsecT e s (m :: Type -> Type) a
#

ParsecT e s m a is a parser with custom data component of error e, stream type s, underlying monad m and return type a.

Instances19MonadParsec, MonadParsecDbg, MonadError, MonadReader, MonadState, MonadWriter, …

Running parser

7 declarations
valueparse
  1. :: Parsec e s a

    Parser to run

  2. -> String

    Name of source file

  3. -> s

    Input for parser

  4. -> Either (ParseErrorBundle s e) a
#

parse p file input runs parser p over Identity (see runParserT if you're using the ParsecT monad transformer; parse itself is just a synonym for runParser). It returns either a ParseErrorBundle (Left) or a value of type a (Right). errorBundlePretty can be used to turn ParseErrorBundle into the string representation of the error message. See Text.Megaparsec.Error if you need to do more advanced error analysis.

main = case parse numbers "" "11,2,43" of
         Left bundle -> putStr (errorBundlePretty bundle)
         Right xs -> print (sum xs)

numbers = decimal `sepBy` char ','

parse is the same as runParser.

valueparseMaybe :: (Ord e, Stream s) => Parsec e s a -> s -> Maybe a
#

parseMaybe p input runs the parser p on input and returns the result inside Just on success and Nothing on failure. This function also parses eof, so if the parser doesn't consume all of its input, it will fail.

The function is supposed to be useful for lightweight parsing, where error messages (and thus file names) are not important and entire input should be consumed. For example, it can be used for parsing of a single number according to a specification of its format.

valuerunParser'
  1. :: Parsec e s a

    Parser to run

  2. -> State s e

    Initial state

  3. -> (State s e, Either (ParseErrorBundle s e) a)
#

The function is similar to runParser with the difference that it accepts and returns the parser state. This allows us e.g. to specify arbitrary textual position at the beginning of parsing. This is the most general way to run a parser over the Identity monad.

valuerunParserT
  1. :: Monad m
  2. => ParsecT e s m a

    Parser to run

  3. -> String

    Name of source file

  4. -> s

    Input for parser

  5. -> m (Either (ParseErrorBundle s e) a)
#

runParserT p file input runs parser p on the input list of tokens input, obtained from source file. The file is only used in error messages and may be the empty string. Returns a computation in the underlying monad m that returns either a ParseErrorBundle (Left) or a value of type a (Right).

Primitive combinators

1 declaration
classclass (Stream s, MonadPlus m) => MonadParsec e s (m :: Type -> Type) | m -> e s where
#

Type class describing monads that implement the full set of primitive parsers.

Note that the following primitives are “fast” and should be taken advantage of as much as possible if your aim is a fast parser: tokens, takeWhileP, takeWhile1P, and takeP.

Methods

  • parseError :: ParseError s e -> m a

    Stop parsing and report the ParseError. This is the only way to control position of the error without manipulating the parser state manually.

  • label :: String -> m a -> m a

    The parser label name p behaves as parser p, but whenever the parser p fails without consuming any input, it replaces names of “expected” tokens with the name name.

  • hidden :: m a -> m a

    hidden p behaves just like parser p, but it doesn't show any “expected” tokens in error message when p fails.

    Please use hidden instead of the old label "" idiom.

  • try :: m a -> m a

    The parser try p behaves like the parser p, except that it backtracks the parser state when p fails (either consuming input or not).

    This combinator is used whenever arbitrary look ahead is needed. Since it pretends that it hasn't consumed any input when p fails, the (A.<|>) combinator will try its second alternative even if the first parser failed while consuming input.

    For example, here is a parser that is supposed to parse the word “let” or the word “lexical”:

    Example1 expression
    parseTest (string "let" <|> string "lexical") "lexical"1:1:unexpected "lex"expecting "let"

    What happens here? The first parser consumes “le” and fails (because it doesn't see a “t”). The second parser, however, isn't tried, since the first parser has already consumed some input! try fixes this behavior and allows backtracking to work:

    Example1 expression
    parseTest (try (string "let") <|> string "lexical") "lexical""lexical"

    try also improves error messages in case of overlapping alternatives, because Megaparsec's hint system can be used:

    Example1 expression
    parseTest (try (string "let") <|> string "lexical") "le"1:1:unexpected "le"expecting "let" or "lexical"

    Note that as of Megaparsec 4.4.0, string backtracks automatically (see tokens), so it does not need try. However, the examples above demonstrate the idea behind try so well that it was decided to keep them. You still need to use try when your alternatives are complex, composite parsers.

  • lookAhead :: m a -> m a

    If p in lookAhead p succeeds (either consuming input or not) the whole parser behaves like p succeeded without consuming anything (parser state is not updated as well). If p fails, lookAhead has no effect, i.e. it will fail consuming input if p fails consuming input. Combine with try if this is undesirable.

  • notFollowedBy :: m a -> m ()

    notFollowedBy p only succeeds when the parser p fails. This parser never consumes any input and never modifies parser state. It can be used to implement the “longest match” rule.

  • withRecovery :: (ParseError s e -> m a) -> m a -> m a

    withRecovery r p allows us to continue parsing even if the parser p fails. In this case r is called with the actual ParseError as its argument. Typical usage is to return a value signifying failure to parse this particular object and to consume some part of the input up to the point where the next object starts.

    Note that if r fails, the original error message is reported as if without withRecovery. In no way recovering parser r can influence error messages.

  • observing :: m a -> m (Either (ParseError s e) a)

    observing p allows us to “observe” failure of the p parser, should it happen, without actually ending parsing but instead getting the ParseError in Left. On success parsed value is returned in Right as usual. Note that this primitive just allows you to observe parse errors as they happen, it does not backtrack or change how the p parser works in any way.

  • eof :: m ()

    This parser only succeeds at the end of input.

  • token :: (Token s -> Maybe a) -> Set (ErrorItem (Token s)) -> m a

    The parser token test expected accepts tokens for which the matching function test returns Just results. If Nothing is returned the expected set is used to report the items that were expected.

    For example, the satisfy parser is implemented as:

    satisfy f = token testToken Set.empty
      where
        testToken x = if f x then Just x else Nothing

    Note: type signature of this primitive was changed in the version 7.0.0.

  • tokens :: (Tokens s -> Tokens s -> Bool) -> Tokens s -> m (Tokens s)

    The parser tokens test chk parses a chunk of input chk and returns it. The supplied predicate test is used to check equality of given and parsed chunks after a candidate chunk of correct length is fetched from the stream.

    This can be used for example to write chunk:

    chunk = tokens (==)

    Note that beginning from Megaparsec 4.4.0, this is an auto-backtracking primitive, which means that if it fails, it never consumes any input. This is done to make its consumption model match how error messages for this primitive are reported (which becomes an important thing as user gets more control with primitives like withRecovery):

    Example1 expression
    parseTest (string "abc") "abd"1:1:unexpected "abd"expecting "abc"

    This means, in particular, that it's no longer necessary to use try with tokens-based parsers, such as string and string'. This feature does not affect performance in any way.

  • takeWhileP :: Maybe String -> (Token s -> Bool) -> m (Tokens s)

    Parse zero or more tokens for which the supplied predicate holds. Try to use this as much as possible because for many streams this combinator is much faster than parsers built with many and satisfy.

    takeWhileP (Just "foo") f = many (satisfy f <?> "foo")
    takeWhileP Nothing      f = many (satisfy f)

    The combinator never fails, although it may parse the empty chunk.

  • takeWhile1P :: Maybe String -> (Token s -> Bool) -> m (Tokens s)

    Similar to takeWhileP, but fails if it can't parse at least one token. Try to use this as much as possible because for many streams this combinator is much faster than parsers built with some and satisfy.

    takeWhile1P (Just "foo") f = some (satisfy f <?> "foo")
    takeWhile1P Nothing      f = some (satisfy f)

    Note that the combinator either succeeds or fails without consuming any input, so try is not necessary with it.

  • takeP :: Maybe String -> Int -> m (Tokens s)

    Extract the specified number of tokens from the input stream and return them packed as a chunk of stream. If there is not enough tokens in the stream, a parse error will be signaled. It's guaranteed that if the parser succeeds, the requested number of tokens will be returned.

    The parser is roughly equivalent to:

    takeP (Just "foo") n = count n (anySingle <?> "foo")
    takeP Nothing      n = count n anySingle

    Note that if the combinator fails due to insufficient number of tokens in the input stream, it backtracks automatically. No try is necessary with takeP.

  • getParserState :: m (State s e)

    Return the full parser state as a State record.

  • updateParserState :: (State s e -> State s e) -> m ()

    updateParserState f applies the function f to the parser state.

  • mkParsec :: (State s e -> Reply e s a) -> m a

    An escape hatch for defining custom MonadParsec primitives. You will need to import Text.Megaparsec.Internal in order to construct Reply.

Instances9MonadParsec, …

Signaling parse errors

8 declarations

The most general function to fail and end parsing is parseError. These are built on top of it. The section also includes functions starting with the register prefix which allow users to register “delayed” ParseErrors.

valueunexpected :: MonadParsec e s m => ErrorItem (Token s) -> m a
#

The parser unexpected item fails with an error message telling about unexpected item item without consuming any input.

unexpected item = failure (Just item) Set.empty
valuecustomFailure :: MonadParsec e s m => e -> m a
#

Report a custom parse error. For a more general version, see fancyFailure.

customFailure = fancyFailure . Set.singleton . ErrorCustom
valueregion
  1. :: MonadParsec e s m
  2. => (ParseError s e -> ParseError s e)

    How to process ParseErrors

  3. -> m a

    The “region” that the processing applies to

  4. -> m a
#

Specify how to process ParseErrors that happen inside of this wrapper. This applies to both normal and delayed ParseErrors.

As a side-effect of the implementation the inner computation will start with an empty collection of delayed errors and they will be updated and “restored” on the way out of region.

valueregisterParseError :: MonadParsec e s m => ParseError s e -> m ()
#

Register a ParseError for later reporting. This action does not end parsing and has no effect except for adding the given ParseError to the collection of “delayed” ParseErrors which will be taken into consideration at the end of parsing. Only if this collection is empty the parser will succeed. This is the main way to report several parse errors at once.

Derivatives of primitive combinators

11 declarations
valuesatisfy
  1. :: MonadParsec e s m
  2. => (Token s -> Bool)

    Predicate to apply

  3. -> m (Token s)
#

The parser satisfy f succeeds for any token for which the supplied function f returns True.

digitChar = satisfy isDigit <?> "digit"
oneOf cs  = satisfy (`elem` cs)

Performance note: when you need to parse a single token, it is often a good idea to use satisfy with the right predicate function instead of creating a complex parser using the combinators.

See also: anySingle, anySingleBut, oneOf, noneOf.

valueoneOf
  1. :: (Foldable f, MonadParsec e s m)
  2. => f (Token s)

    Collection of matching tokens

  3. -> m (Token s)
#

oneOf ts succeeds if the current token is in the supplied collection of tokens ts. Returns the parsed token. Note that this parser cannot automatically generate the “expected” component of error message, so usually you should label it manually with label or (<?>).

oneOf cs = satisfy (`elem` cs)

See also: satisfy.

digit = oneOf ['0'..'9'] <?> "digit"

Performance note: prefer satisfy when you can because it's faster when you have only a couple of tokens to compare to:

quoteFast = satisfy (\x -> x == '\'' || x == '\"')
quoteSlow = oneOf "'\""
valuenoneOf
  1. :: (Foldable f, MonadParsec e s m)
  2. => f (Token s)

    Collection of taken we should not match

  3. -> m (Token s)
#

As the dual of oneOf, noneOf ts succeeds if the current token not in the supplied list of tokens ts. Returns the parsed character. Note that this parser cannot automatically generate the “expected” component of error message, so usually you should label it manually with label or (<?>).

noneOf cs = satisfy (`notElem` cs)

See also: satisfy.

Performance note: prefer satisfy and anySingleBut when you can because it's faster.

valuematch :: MonadParsec e s m => m a -> m (Tokens s, a)
#

Return both the result of a parse and a chunk of input that was consumed during parsing. This relies on the change of the stateOffset value to evaluate how many tokens were consumed. If you mess with it manually in the argument parser, prepare for troubles.

valuetakeRest :: MonadParsec e s m => m (Tokens s)
#

Consume the rest of the input and return it as a chunk. This parser never fails, but may return the empty chunk.

takeRest = takeWhileP Nothing (const True)
valueatEnd :: MonadParsec e s m => m Bool
#

Return True when end of input has been reached.

atEnd = option False (True <$ hidden eof)

Parser state combinators

6 declarations

Return the current source position. This function is not cheap, do not call it e.g. on matching of every token, that's a bad idea. Still you can use it to get SourcePos to attach to things that you parse.

The function works under the assumption that we move in the input stream only forwards and never backwards, which is always true unless the user abuses the library.