This module provides both a native Haskell solution for parsing XML
documents into a stream of events, and a set of parser combinators for
dealing with a stream of events.
As a simple example:
Example6 expressions
>>> :set -XOverloadedStrings>>> import Conduit (runConduit, (.|))>>> import Data.Text (Text, unpack)>>> import Data.XML.Types (Event)>>> data Person = Person Int Text Text deriving Show>>> :{let parsePerson :: MonadThrow m => ConduitT Event o m (Maybe Person) parsePerson = tag' "person" parseAttributes $ \(age, goodAtHaskell) -> do name <- content return $ Person (read $ unpack age) name goodAtHaskell where parseAttributes = (,) <$> requireAttr "age" <*> requireAttr "goodAtHaskell" <* ignoreAttrs parsePeople :: MonadThrow m => ConduitT Event o m (Maybe [Person]) parsePeople = tagNoAttr "people" $ many parsePerson inputXml = mconcat [ "<?xml version=\"1.0\" encoding=\"utf-8\"?>" , "<people>" , " <person age=\"25\" goodAtHaskell=\"yes\">Michael</person>" , " <person age=\"2\" goodAtHaskell=\"might become\">Eliezer</person>" , "</people>" ]:}
This module also supports streaming results using yield.
This allows parser results to be processed using conduits
while a particular parser (e.g. many) is still running.
Without using streaming results, you have to wait until the parser finished
before you can process the result list. Large XML files might be easier
to process by using streaming results.
See http://stackoverflow.com/q/21367423/2597135 for a related discussion.
Example2 expressions
>>> import Data.Conduit.List as CL>>> :{let parsePeople' :: MonadThrow m => ConduitT Event Person m (Maybe ()) parsePeople' = tagNoAttr "people" $ manyYield parsePerson:}
Previous versions of this module contained a number of more sophisticated
functions written by Aristid Breitkreuz and Dmitry Olshansky. To keep this
package simpler, those functions are being moved to a separate package. This
note will be updated with the name of the package(s) when available.
Parses a byte stream into Events. This function is implemented fully in
Haskell using attoparsec-text for parsing. The produced error messages do
not give line/column information, so you may prefer to stick with the parser
provided by libxml-enumerator. However, this has the advantage of not
relying on any C libraries.
This relies on detectUtf to determine character encoding, and parseText
to do the actual parsing.
Parses a character stream into Events. This function is implemented
fully in Haskell using attoparsec-text for parsing. The produced error
messages do not give line/column information, so you may prefer to stick
with the parser provided by libxml-enumerator. However, this has the
advantage of not relying on any C libraries.
Automatically determine which UTF variant is being used. This function
first checks for BOMs, removing them as necessary, and then check for the
equivalent of <?xml for each of UTF-8, UTF-16LEBE, and UTF-32LEBE. It
defaults to assuming UTF-8.
A helper function which reads a file from disk using enumFile, detects
character encoding using detectUtf, parses the XML using parseBytes, and
then hands off control to your supplied parser.
Default implementation of DecodeEntities, which leaves the
entity as-is. Numeric character references and the five standard
entities (lt, gt, amp, quot, pos) are handled internally by the
parser.
Given the value returned by the name checker, this function will
be used to get an AttrParser appropriate for the specific tag.
If the AttrParser fails, the function will also return Nothing
The most generic way to parse a tag. It takes a NameMatcher to check whether
this is a correct tag name, an AttrParser to handle attributes, and
then a parser to deal with content.
Events are consumed if and only if the tag name and its attributes match.
This function automatically absorbs its balancing closing tag, and will
throw an exception if not all of the attributes or child elements are
consumed. If you want to allow extra attributes, see ignoreAttrs.
This function automatically ignores comments, instructions and whitespace.
Grabs the next piece of content if available. This function skips over any
comments, instructions or entities, and concatenates all content until the next start
or end tag.
Ignore an empty tag and all of its attributes.
This does not ignore the tag recursively
(i.e. it assumes there are no child elements).
This function returns Just () if the tag matched.
Match a single Name in a concise way.
Note that Name is namespace sensitive: when using the IsString instance,
use "{http://a/b}c" to match the tag c in the XML namespace http://a/b
A monad for parsing attributes. By default, it requires you to deal with
all attributes present on an element, and will throw an exception if there
are unhandled attributes. Use the requireAttr, attr et al
functions for handling an attribute, and ignoreAttrs if you would like to
skip the rest of the attributes on an element.
Alternative instance behaves like First monoid: it chooses first
parser which doesn't fail.
Skip the remaining attributes on an element. Since this will clear the
list of attributes, you must call this after any calls to requireAttr,
optionalAttr, etc.
Get the value of the first parser which returns Just. If no parsers
succeed (i.e., return Just), this function returns Nothing.
Warning: choose doesn't backtrack. If a parser consumed some events,
subsequent parsers will continue from the following events. This can be a
problem if parsers share an accepted prefix of events, so an earlier
(failing) parser will discard the events that the later parser could
potentially succeed on.
An other problematic case is using choose to implement order-independent
parsing using a set of parsers, with a final trailing ignore-anything-else
action. In this case, certain trees might be skipped.
Force an optional parser into a required parser. All of the tag
functions, attr, choose and many deal with Maybe parsers. Use this when you
want to finally force something to happen.