Decode a stream of bytes to Unicode characters by mapping each byte to a corresponding Unicode Char in 0-255 range.
Modulestreamly-core-0.2.2Haskell2010
Streamly.Internal.Unicode.Stream
- 4 types
- 40 values
- Packagestreamly-core-0.2.2
- Exports44
- LanguageHaskell2010
- LicenceBSD-3-Clause
- SourceStream.hs
Setup
0 declarationsTo execute the code examples provided in this module in ghci, please run the following commands first.
:mimport qualified Streamly.Data.Fold as Foldimport qualified Streamly.Data.Stream as Streamimport qualified Streamly.Unicode.Stream as Unicode
For APIs that have not been released yet.
:set -XMagicHashimport qualified Streamly.Internal.Unicode.Stream as Unicode
Construction (Decoding)
1 declarationUTF-8 Decoding
Instances1Show
Show CodingFailureModeDefined in streamly-core-0.2.2 · Streamly.Internal.Unicode.Stream
Decode a UTF-8 encoded bytestream to a stream of Unicode characters. Any invalid codepoint encountered is replaced with the unicode replacement character.
Decode a UTF-8 encoded bytestream to a stream of Unicode characters. The function throws an error if an invalid codepoint is encountered.
Decode a UTF-8 encoded bytestream to a stream of Unicode characters. Any invalid codepoint encountered is dropped.
Decode a UTF-16 little endian encoded bytestream to a stream of Unicode characters. The function throws an error if an invalid codepoint is encountered.
Unimplemented
Resumable UTF-8 Decoding
Constructors
Instances1Show
Show DecodeErrorDefined in streamly-core-0.2.2 · Streamly.Internal.Unicode.Stream
Pre-release
Pre-release
UTF-8 Array Stream Decoding
Like decodeUtf8 but for a chunked stream. It may be slightly faster than flattening the stream and then decoding with decodeUtf8.
Like 'decodeUtf8'' but for a chunked stream. It may be slightly faster than flattening the stream and then decoding with 'decodeUtf8''.
Like decodeUtf8_ but for a chunked stream. It may be slightly faster than flattening the stream and then decoding with decodeUtf8_.
Elimination (Encoding)
0 declarationsLatin1 Encoding
Like encodeLatin1' but silently maps input codepoints beyond 255 to arbitrary Latin1 chars in 0-255 range. No error or exception is thrown when such mapping occurs.
Encode a stream of Unicode characters to bytes by mapping each character to a byte in 0-255 range. Throws an error if the input stream contains characters beyond 255.
Like encodeLatin1 but drops the input characters beyond 255.
UTF-8 Encoding
Encode a stream of Unicode characters to a UTF-8 encoded bytestream. Any Invalid characters (U+D800-U+D8FF) in the input stream are replaced by the Unicode replacement character U+FFFD.
Encode a stream of Unicode characters to a UTF-8 encoded bytestream. When any invalid character (U+D800-U+D8FF) is encountered in the input stream the function errors out.
Encode a stream of Unicode characters to a UTF-8 encoded bytestream. Any Invalid characters (U+D800-U+D8FF) in the input stream are dropped.
Encode a stream of String using the supplied encoding scheme. Each
string is encoded as an Array Word8.
Encode a stream of Unicode characters to a UTF-16 little endian encoded bytestream.
Unimplemented
Transformation
5 declarationsRemove leading whitespace from a string.
stripHead = Stream.dropWhile isSpacePre-release
Fold each line of the stream using the supplied Fold and stream the result.
Stream.fold Fold.toList $ Unicode.lines Fold.toList (Stream.fromList "lines\nthis\nstring\n\n\n")["lines","this","string","",""]
lines = Stream.splitOnSuffix (== '\n')Pre-release
Fold each word of the stream using the supplied Fold and stream the result.
Stream.fold Fold.toList $ Unicode.words Fold.toList (Stream.fromList "fold these words")["fold","these","words"]
words = Stream.wordsBy isSpacePre-release
Unfold a stream to character streams using the supplied Unfold
and concat the results suffixing a newline character \n to each stream.
unlines = Stream.interposeSuffix 'n'
unlines = Stream.intercalateSuffix Unfold.fromList "n"
Pre-release
Unfold the elements of a stream to character streams using the supplied Unfold and concat the results with a whitespace character infixed between the streams.
unwords = Stream.interpose ' '
unwords = Stream.intercalate Unfold.fromList " "
Pre-release
StreamD UTF8 Encoding / Decoding transformations.
8 declarationsSee section "3.9 Unicode Encoding Forms" in https://www.unicode.org/versions/Unicode13.0.0/UnicodeStandard-13.0.pdf
Decoding String Literals
1 declarationRead UTF-8 encoded bytes as chars from an Addr# until a 0 byte is encountered, the 0 byte is not included in the stream.
Unsafe: The caller is responsible for safe addressing.
Note that this is completely safe when reading from Haskell string literals because they are guaranteed to be NULL terminated:
Stream.fold Fold.toList (Unicode.fromStr# "Haskell"#)"Haskell"
Deprecations
3 declarationsDeprecated. Please use decodeUtf8 instead
Same as decodeUtf8
Deprecated. Please use encodeLatin1 instead
Same as encodeLatin1
Deprecated. Please use encodeUtf8 instead
Same as encodeUtf8