Compress a data stream provided as a lazy ByteString.
There are no expected error conditions. All input data streams are valid. It
is possible for unexpected errors to occur, such as running out of memory,
or finding the wrong version of the zlib C library, these are thrown as
exceptions.
Decompress a data stream provided as a lazy ByteString.
It will throw an exception if any error is encountered in the input data.
If you need more control over error handling then use one the incremental
versions, decompressST or decompressIO.
Monadic incremental interface
0 declarations
The pure compress and
decompress functions are streaming in the sense
that they can produce output without demanding all input, however they need
the input data stream as a lazy ByteString. Having the input data
stream as a lazy ByteString often requires using lazy I/O which is not
appropriate in all circumstances.
For these cases an incremental interface is more appropriate. This interface
allows both incremental input and output. Chunks of input data are supplied
one by one (e.g. as they are obtained from an input source like a file or
network source). Output is also produced chunk by chunk.
The incremental input and output is managed via the CompressStream and
DecompressStream types. They represents the unfolding of the process of
compressing and decompressing. They operates in either the ST or IO
monads. They can be lifted into other incremental abstractions like pipes or
conduits, or they can be used directly in the following style.
Using incremental compression
In a loop:
Inspect the status of the stream
When it is CompressInputRequired then you should call the action,
passing a chunk of input (or BS.empty when no more input is available)
to get the next state of the stream and continue the loop.
When it is CompressOutputAvailable then do something with the given
chunk of output, and call the action to get the next state of the stream
and continue the loop.
Note that you cannot stop as soon as you have no more input, you need to
carry on until all the output has been collected, i.e. until you get to
CompressStreamEnd.
Here is an example where we get input from one file handle and send the
compressed output to another file handle.
go :: Handle -> Handle -> CompressStream IO -> IO ()
go inh outh (CompressInputRequired next) = do
inchunk <- BS.hGet inh 4096
go inh outh =<< next inchunk
go inh outh (CompressOutputAvailable outchunk next) =
BS.hPut outh outchunk
go inh outh =<< next
go _ _ CompressStreamEnd = return ()
The unfolding of the compression process, where you provide a sequence
of uncompressed data chunks as input and receive a sequence of compressed
data chunks as output. The process is incremental, in that the demand for
input and provision of output are interleaved.
A variant on foldCompressStream that is pure rather than operating in a
monad and where the input is provided by a lazy ByteString. So we only
have to deal with the output and end parts, making it just like a foldr on a
list of output chunks.
The unfolding of the decompression process, where you provide a sequence
of compressed data chunks as input and receive a sequence of uncompressed
data chunks as output. The process is incremental, in that the demand for
input and provision of output are interleaved.
To indicate the end of the input supply an empty input chunk. Note that
for gzipFormat with the default decompressAllMembersTrue you will
have to do this, as the decompressor will look for any following members.
With decompressAllMembersFalse the decompressor knows when the data
ends and will produce DecompressStreamEnd without you having to supply an
empty chunk to indicate the end of the input.
It is possible to do zlib compression with a custom dictionary. This
allows slightly higher compression ratios for short files. However such
compressed streams require the same dictionary when decompressing. This
error is for when we encounter a compressed stream that needs a
dictionary, and it's not provided.
If the compressed data stream is corrupted in any way then you will
get this error, for example if the input data just isn't a compressed
zlib data stream. In particular if the data checksum turns out to be
wrong then you will get all the decompressed data but this error at the
end, instead of the normal successful StreamEnd.
Instances6Eq, Ord, Show, Generic, Exception, Rep
EqDecompressErrorDefined in zlib-0.7.1.0 · Codec.Compression.Zlib.Internal
OrdDecompressErrorDefined in zlib-0.7.1.0 · Codec.Compression.Zlib.Internal
ShowDecompressErrorDefined in zlib-0.7.1.0 · Codec.Compression.Zlib.Internal
A variant on foldCompressStream that is pure rather than operating in a
monad and where the input is provided by a lazy ByteString. So we only
have to deal with the output, end and error parts, making it like a foldr on
a list of output chunks.
The compressBufferSize is the size of the first output buffer containing
the compressed data. If you know an approximate upper bound on the size of
the compressed data then setting this parameter can save memory. The default
compression output buffer size is 16k. If your estimate is wrong it does
not matter too much, the default buffer size will be used for the remaining
chunks.
The decompressBufferSize is the size of the first output buffer,
containing the uncompressed data. If you know an exact or approximate upper
bound on the size of the decompressed data then setting this parameter can
save memory. The default decompression output buffer size is 32k. If your
estimate is wrong it does not matter too much, the default buffer size will
be used for the remaining chunks.
One particular use case for setting the decompressBufferSize is if you
know the exact size of the decompressed data and want to produce a strict
ByteString. The compression and decompression functions
use lazy ByteStrings but if you set the
decompressBufferSize correctly then you can generate a lazy
ByteString with exactly one chunk, which can be
converted to a strict ByteString in O(1) time using
concat . toChunks.
The gzip format uses a header with a checksum and some optional meta-data
about the compressed file. It is intended primarily for compressing
individual files but is also sometimes used for network protocols such as
HTTP. The format is described in detail in RFC #1952
http://www.ietf.org/rfc/rfc1952.txt
The zlib format uses a minimal header with a checksum but no other
meta-data. It is especially designed for use in network protocols. The
format is described in detail in RFC #1950
http://www.ietf.org/rfc/rfc1950.txt
The 'raw' format is just the compressed data stream without any
additional header, meta-data or data-integrity checksum. The format is
described in detail in RFC #1951 http://www.ietf.org/rfc/rfc1951.txt
The compression level parameter controls the amount of compression. This
is a trade-off between the amount of compression and the time required to do
the compression.
This specifies the size of the compression window. Larger values of this
parameter result in better compression at the expense of higher memory
usage.
The compression window size is the value of the the window bits raised to
the power 2. The window bits must be in the range 9..15 which corresponds
to compression window sizes of 512b to 32Kb. The default is 15 which is also
the maximum size.
The total amount of memory used depends on the window bits and the
MemoryLevel. See the MemoryLevel for the details.
The MemoryLevel parameter specifies how much memory should be allocated
for the internal compression state. It is a trade-off between memory usage,
compression ratio and compression speed. Using more memory allows faster
compression and a better compression ratio.
The total amount of memory used for compression depends on the WindowBits
and the MemoryLevel. For decompression it depends only on the
WindowBits. The totals are given by the functions:
For example, for compression with the default windowBits = 15 and
memLevel = 8 uses 256Kb. So for example a network server with 100
concurrent compressed streams would use 25Mb. The memory per stream can be
halved (at the cost of somewhat degraded and slower compression) by
reducing the windowBits and memLevel by one.
Decompression takes less memory, the default windowBits = 15 corresponds
to just 32Kb.
Use the filtered compression strategy for data produced by a filter (or
predictor). Filtered data consists mostly of small values with a somewhat
random distribution. In this case, the compression algorithm is tuned to
compress them better. The effect of this strategy is to force more Huffman
coding and less string matching; it is somewhat intermediate between
defaultStrategy and huffmanOnlyStrategy.
Use rleStrategy to limit match distances to one (run-length
encoding). rleStrategy is designed to be almost as fast as
huffmanOnlyStrategy, but give better compression for PNG
image data.