All of the single-parameter functions for decoding bytestrings
encoded in one of the Unicode Transformation Formats (UTF) operate
in a strict mode: each will throw an exception if given invalid
input.
Each function has a variant, whose name is suffixed with -With,
that gives greater control over the handling of decoding errors.
For instance, decodeUtf8 will throw an exception, but
decodeUtf8With allows the programmer to determine what to do on a
decoding error.
Total Functions
These functions facilitate total decoding and should be preferred
over their partial counterparts.
Decode a ByteString containing Latin-1 (aka ISO-8859-1) encoded text.
decodeLatin1 is semantically equivalent to
Data.Text.pack . Data.ByteString.Char8.unpack
This is a total function. However, bear in mind that decoding Latin-1 (non-ASCII)
characters to UTf-8 requires actual work and is not just buffer copying.
This is a total function which returns a pair of the longest ASCII prefix
as Text, and the remaining suffix as ByteString.
Important note: the pair is lazy. This lets you check for errors by testing
whether the second component is empty, without forcing the first component
(which does a copy).
To drop references to the input bytestring, force the prefix
(using seq or BangPatterns) and drop references to the suffix.
Properties
If (prefix, suffix) = decodeAsciiPrefix s, then encodeUtf8 prefix <> suffix = s.
The streamDecodeUtf8 and streamDecodeUtf8With functions accept
a ByteString that represents a possibly incomplete input (e.g. a
packet from a network stream) that may not end on a UTF-8 boundary.
The maximal prefix of Text that could be decoded from the
given input.
The suffix of the ByteString that could not be decoded due to
insufficient input.
A function that accepts another ByteString. That string will
be assumed to directly follow the string that was passed as
input to the original function, and it will in turn be decoded.
To help understand the use of these functions, consider the Unicode
string "hi ☃". If encoded as UTF-8, this becomes "hi
\xe2\x98\x83"; the final '☃' is encoded as 3 bytes.
Now suppose that we receive this encoded string as 3 packets that
are split up on untidy boundaries: ["hi \xe2", "\x98",
"\x83"]. We cannot decode the entire Unicode string until we
have received all three packets, but we would like to make progress
as we receive each one.
We use the continuation f0 to decode our second packet.
ghci> let s1@(Some _ _ f1) = f0 "\x98"
ghci> s1
Some "" "\xe2\x98"
We could not give f0 enough input to decode anything, so it
returned an empty string. Once we feed our second continuation f1
the last byte of input, it will make progress.
ghci> let s2@(Some _ _ f2) = f1 "\x83"
ghci> s2
Some "\x2603" "" _
If given invalid input, an exception will be thrown by the function
or continuation where it is encountered.
Those functions return an UTF-8 prefix of the given ByteString up to the next error.
For example this lets you insert or delete arbitrary text, or do some
stateful operations before resuming, such as keeping track of error locations.
In contrast, the older stream-oriented interface only lets you substitute
a single fixed Char for each invalid byte in OnDecodeError.
That prefix is encoded as a StrictBuilder, so you can accumulate chunks
before doing the copying work to construct a Text, or you can
output decoded fragments immediately as a lazy Text.
Concatenation of StrictBuilder is right-biased:
the right builder will be run first. This allows a builder to
run tail-recursively when it was accumulated left-to-right.
Decode a ByteString containing UTF-8 encoded text that is known
to be valid.
If the input contains any invalid UTF-8 data, an exception will be
thrown that cannot be caught in pure code. For more control over
the handling of invalid data, use decodeUtf8' or
decodeUtf8With.
This is a partial function: it checks that input is a well-formed
UTF-8 sequence and copies buffer or throws an error otherwise.
If the input contains any invalid little endian UTF-16 data, an
exception will be thrown. For more control over the handling of
invalid data, use decodeUtf16LEWith.
If the input contains any invalid big endian UTF-16 data, an
exception will be thrown. For more control over the handling of
invalid data, use decodeUtf16BEWith.
If the input contains any invalid little endian UTF-32 data, an
exception will be thrown. For more control over the handling of
invalid data, use decodeUtf32LEWith.
If the input contains any invalid big endian UTF-32 data, an
exception will be thrown. For more control over the handling of
invalid data, use decodeUtf32BEWith.
Decode, in a stream oriented way, a ByteString containing UTF-8
encoded text that is known to be valid.
If the input contains any invalid UTF-8 data, an exception will be
thrown (either by this function or a continuation) that cannot be
caught in pure code. For more control over the handling of invalid
data, use streamDecodeUtf8With.
Validate another ByteString chunk in an ongoing stream of UTF-8-encoded text.
Returns a pair:
The first component n is the end position, relative to the current
chunk, of the longest prefix of the accumulated bytestring which is valid UTF-8.
n may be negative: that happens when an incomplete code point started in
a previous chunk and is not completed by the current chunk (either
that code point is still incomplete, or it is broken by an invalid byte).
The second component ms indicates the following:
if ms = Nothing, the remainder of the chunk contains an invalid byte,
within four bytes from position n;
if ms = Just s', you can carry on validating another chunk
by calling validateUtf8More with the new state s'.