HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Modulebytestring-0.12.2.0Haskell2010

Data.ByteString

A time- and space-efficient implementation of byte vectors using packed Word8 arrays, suitable for high performance use, both in terms of large data quantities and high speed requirements. Byte vectors are encoded as strict Word8 arrays of bytes, held in a ForeignPtr, and can be passed between C and Haskell with little effort.

The recomended way to assemble ByteStrings from smaller parts is to use the builder monoid from Data.ByteString.Builder.

This module is intended to be imported qualified, to avoid name clashes with Prelude functions. eg.

import qualified Data.ByteString as B

Original GHC implementation by Bryan O'Sullivan. Rewritten to use UArray by Simon Marlow. Rewritten to support slices and use ForeignPtr by David Roundy. Rewritten again and extended by Don Stewart and Duncan Coutts.

  • 2 types
  • 115 values

Strict ByteString

2 declarations
datadata ByteString
#

A space-efficient representation of a Word8 vector, supporting many efficient operations.

A ByteString contains 8-bit bytes, or by using the operations from Data.ByteString.Char8 it can be interpreted as containing 8-bit characters.

Instances12IsList, Eq, Data, Ord, Read, Show, …
  • IsList ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Eq ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Data ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Ord ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Read ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Show ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • IsString ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type

    Beware: fromString truncates multi-byte characters to octets. e.g. "枯朶に烏のとまりけり秋の暮" becomes �6k�nh~�Q��n�

  • Semigroup ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Monoid ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • NFData ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • Lift ByteStringDefined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type
  • type Item ByteString = Word8Defined in bytestring-0.12.2.0 · Data.ByteString.Internal.Type

Heap fragmentation

With GHC, the ByteString representation uses pinned memory, meaning it cannot be moved by GC. While this is ideal for use with the foreign function interface and is usually efficient, this representation may lead to issues with heap fragmentation and wasted space if the program selectively retains a fraction of many small ByteStrings, keeping them live in memory over long durations.

While ByteString is indispensable when working with large blobs of data and especially when interfacing with native C libraries, be sure to also check the ShortByteString type. As a type backed by unpinned memory, ShortByteString behaves similarly to Text (from the text package) on the heap, completely avoids fragmentation issues, and in many use-cases may better suit your bytestring-storage needs.

Introducing and eliminating ByteStrings

8 declarations

Convert a FilePath to a ByteString.

The FilePath type is expected to use the file system encoding as reported by getFileSystemEncoding. This encoding allows for round-tripping of arbitrary data on platforms that allow arbitrary bytes in their paths. This conversion function does the same thing that System.IO.openFile would do when decoding the FilePath.

This function is in IO because the file system encoding can be changed. If the encoding can be assumed to be constant in your use case, you may invoke this function via unsafePerformIO.

Convert a ByteString to a FilePath.

This function uses the file system encoding, and resulting FilePaths can be safely used with standard IO functions and will reference the correct path in the presence of arbitrary non-UTF-8 encoded paths.

This function is in IO because the file system encoding can be changed. If the encoding can be assumed to be constant in your use case, you may invoke this function via unsafePerformIO.

Basic interface

11 declarations
valuehead :: HasCallStack => ByteString -> Word8
#

O(1) Extract the first element of a ByteString, which must be non-empty. An exception will be thrown in the case of an empty ByteString.

This is a partial function, consider using uncons instead.

valuelast :: HasCallStack => ByteString -> Word8
#

O(1) Extract the last element of a ByteString, which must be finite and non-empty. An exception will be thrown in the case of an empty ByteString.

This is a partial function, consider using unsnoc instead.

O(1) Extract the elements after the head of a ByteString, which must be non-empty. An exception will be thrown in the case of an empty ByteString.

This is a partial function, consider using uncons instead.

Transforming ByteStrings

5 declarations

Reducing ByteStrings (folds)

8 declarations
valuefoldl :: (a -> Word8 -> a) -> a -> ByteString -> a
#

foldl, applied to a binary operator, a starting value (typically the left-identity of the operator), and a ByteString, reduces the ByteString using the binary operator, from left to right.

valuefoldr :: (Word8 -> a -> a) -> a -> ByteString -> a
#

foldr, applied to a binary operator, a starting value (typically the right-identity of the operator), and a ByteString, reduces the ByteString using the binary operator, from right to left.

Special folds

Building ByteStrings

0 declarations

Scans

valuescanl
  1. :: (Word8 -> Word8 -> Word8)

    accumulator -> element -> new accumulator

  2. -> Word8

    starting value of accumulator

  3. -> ByteString

    input of length n

  4. -> ByteString

    output of length n+1

#

scanl is similar to foldl, but returns a list of successive reduced values from the left.

scanl f z [x1, x2, ...] == [z, z `f` x1, (z `f` x1) `f` x2, ...]

Note that

head (scanl f z xs) == z
last (scanl f z xs) == foldl f z xs
valuescanr
  1. :: (Word8 -> Word8 -> Word8)

    element -> accumulator -> new accumulator

  2. -> Word8

    starting value of accumulator

  3. -> ByteString

    input of length n

  4. -> ByteString

    output of length n+1

#

scanr is similar to foldr, but returns a list of successive reduced values from the right.

scanr f z [..., x{n-1}, xn] == [..., x{n-1} `f` (xn `f` z), xn `f` z, z]

Note that

head (scanr f z xs) == foldr f z xs
last (scanr f z xs) == z

Accumulating maps

valuemapAccumL
  1. :: acc -> Word8 -> (acc, Word8)
  2. -> acc
  3. -> ByteString
  4. -> (acc, ByteString)
#

The mapAccumL function behaves like a combination of map and foldl; it applies a function to each element of a ByteString, passing an accumulating parameter from left to right, and returning a final value of this accumulator together with the new ByteString.

valuemapAccumR
  1. :: acc -> Word8 -> (acc, Word8)
  2. -> acc
  3. -> ByteString
  4. -> (acc, ByteString)
#

The mapAccumR function behaves like a combination of map and foldr; it applies a function to each element of a ByteString, passing an accumulating parameter from right to left, and returning a final value of this accumulator together with the new ByteString.

Generating and unfolding ByteStrings

valuereplicate :: Int -> Word8 -> ByteString
#

O(n) replicate n x is a ByteString of length n with x the value of every element. The following holds:

replicate w c = fst (unfoldrN w (\u -> Just (u,u)) c)
valueunfoldr :: (a -> Maybe (Word8, a)) -> a -> ByteString
#

O(n), where n is the length of the result. The unfoldr function is analogous to the List 'unfoldr'. unfoldr builds a ByteString from a seed value. The function takes the element and returns Nothing if it is done producing the ByteString or returns Just (a,b), in which case, a is the next byte in the string, and b is the seed value for further production.

Examples:

   unfoldr (\x -> if x <= 5 then Just (x, x + 1) else Nothing) 0
== pack [0, 1, 2, 3, 4, 5]
valueunfoldrN :: Int -> (a -> Maybe (Word8, a)) -> a -> (ByteString, Maybe a)
#

O(n) Like unfoldr, unfoldrN builds a ByteString from a seed value. However, the length of the result is limited by the first argument to unfoldrN. This function is more efficient than unfoldr when the maximum length of the result is known.

The following equation relates unfoldrN and unfoldr:

fst (unfoldrN n f s) == take n (unfoldr f s)

Substrings

0 declarations

Breaking strings

valuetakeEnd :: Int -> ByteString -> ByteString
#

O(1) takeEnd n xs is equivalent to drop (length xs - n) xs. Takes n elements from end of bytestring.

Example3 expressions
takeEnd 3 "abcdefg""efg"takeEnd 0 "abcdefg"""takeEnd 4 "abc""abc"
valuedropEnd :: Int -> ByteString -> ByteString
#

O(1) dropEnd n xs is equivalent to take (length xs - n) xs. Drops n elements from end of bytestring.

Example3 expressions
dropEnd 3 "abcdefg""abcd"dropEnd 0 "abcdefg""abcdefg"dropEnd 4 "abc"""
valuedropWhile :: (Word8 -> Bool) -> ByteString -> ByteString
#

Similar to Prelude.dropWhile, drops the longest (possibly empty) prefix of elements satisfying the predicate and returns the remainder.

valuespanEnd :: (Word8 -> Bool) -> ByteString -> (ByteString, ByteString)
#

Returns the longest (possibly empty) suffix of elements satisfying the predicate and the remainder of the string.

spanEnd p is equivalent to breakEnd (not . p) and to (dropWhileEnd p &&& takeWhileEnd p).

We have

spanEnd (not . isSpace) "x y z" == ("x y ", "z")

and

spanEnd (not . isSpace) ps
   ==
let (x, y) = span (not . isSpace) (reverse ps) in (reverse y, reverse x)
valuebreak :: (Word8 -> Bool) -> ByteString -> (ByteString, ByteString)
#

Similar to break, returns the longest (possibly empty) prefix of elements which do not satisfy the predicate and the remainder of the string.

break p is equivalent to span (not . p) and to (takeWhile (not . p) &&& dropWhile (not . p)).

Under GHC, a rewrite rule will transform break (==) into a call to the specialised breakByte:

break ((==) x) = breakByte x
break (==x) = breakByte x
valuegroup :: ByteString -> [ByteString]
#

The group function takes a ByteString and returns a list of ByteStrings such that the concatenation of the result is equal to the argument. Moreover, each string in the result contains only equal elements. For example,

group "Mississippi" = ["M","i","ss","i","ss","i","pp","i"]

It is a special case of groupBy, which allows the programmer to supply their own equality test. It is about 40% faster than groupBy (==)

Breaking into many substrings

valuesplit :: Word8 -> ByteString -> [ByteString]
#

O(n) Break a ByteString into pieces separated by the byte argument, consuming the delimiter. I.e.

split 10  "a\nb\nd\ne" == ["a","b","d","e"]   -- fromEnum '\n' == 10
split 97  "aXaXaXa"    == ["","X","X","X",""] -- fromEnum 'a' == 97
split 120 "x"          == ["",""]             -- fromEnum 'x' == 120
split undefined ""     == []                  -- and not [""]

and

intercalate [c] . split c == id
split == splitWith . (==)

As for all splitting functions in this library, this function does not copy the substrings, it just constructs new ByteStrings that are slices of the original.

valuesplitWith :: (Word8 -> Bool) -> ByteString -> [ByteString]
#

O(n) Splits a ByteString into components delimited by separators, where the predicate returns True for a separator element. The resulting components do not contain the separators. Two adjacent separators result in an empty component in the output. eg.

splitWith (==97) "aabbaca" == ["","","bb","c",""] -- fromEnum 'a' == 97
splitWith undefined ""     == []                  -- and not [""]

Predicates

3 declarations

O(n) The isSuffixOf function takes two ByteStrings and returns True iff the first is a suffix of the second.

The following holds:

isSuffixOf x y == reverse x `isPrefixOf` reverse y

However, the real implementation uses memcmp to compare the end of the string only, with no reverse required..

Encoding validation

Search for arbitrary substrings

valuebreakSubstring
  1. :: ByteString

    String to search for

  2. -> ByteString

    String to search in

  3. -> (ByteString, ByteString)

    Head and tail of string broken at substring

#

Break a string on a substring, returning a pair of the part of the string prior to the match, and the rest of the string.

The following relationships hold:

break (== c) l == breakSubstring (singleton c) l

For example, to tokenise a string, dropping delimiters:

tokenise x y = h : if null t then [] else tokenise x (drop (length x) t)
    where (h,t) = breakSubstring x y

To skip to the first occurrence of a string:

snd (breakSubstring x y)

To take the parts of a string before a delimiter:

fst (breakSubstring x y)

Note that calling `breakSubstring x` does some preprocessing work, so you should avoid unnecessarily duplicating breakSubstring calls with the same pattern.

Searching ByteStrings

0 declarations

Searching by equality

Searching with a predicate

valuefind :: (Word8 -> Bool) -> ByteString -> Maybe Word8
#

O(n) The find function takes a predicate and a ByteString, and returns the first element in matching the predicate, or Nothing if there is no such element.

find f p = case findIndex f p of Just n -> Just (p ! n) ; _ -> Nothing

O(n) The partition function takes a predicate a ByteString and returns the pair of ByteStrings with elements which do and do not satisfy the predicate, respectively; i.e.,

partition p bs == (filter p xs, filter (not . p) xs)

Indexing ByteStrings

10 declarations

O(n) The elemIndexEnd function returns the last index of the element in the given ByteString which is equal to the query element, or Nothing if there is no such element. The following holds:

elemIndexEnd c xs = case elemIndex c (reverse xs) of
  Nothing -> Nothing
  Just i  -> Just (length xs - 1 - i)
valuecount :: Word8 -> ByteString -> Int
#

count returns the number of times its argument appears in the ByteString

count = length . elemIndices

But more efficiently than using length on the intermediate list.

Zipping and unzipping ByteStrings

4 declarations
valuezip :: ByteString -> ByteString -> [(Word8, Word8)]
#

O(n) zip takes two ByteStrings and returns a list of corresponding pairs of bytes. If one input ByteString is short, excess elements of the longer ByteString are discarded. This is equivalent to a pair of unpack operations.

valuezipWith :: (Word8 -> Word8 -> a) -> ByteString -> ByteString -> [a]
#

zipWith generalises zip by zipping with the function given as the first argument, instead of a tupling function. For example, zipWith (+) is applied to two ByteStrings to produce the list of corresponding sums.

Ordered ByteStrings

1 declaration

Low level conversions

0 declarations

Copying ByteStrings

valuecopy :: ByteString -> ByteString
#

O(n) Make a copy of the ByteString with its own storage. This is mainly useful to allow the rest of the data pointed to by the ByteString to be garbage collected, for example if a large string has been read in, and only a small part of it is needed in the rest of the program.

Packing CStrings and pointers

O(n). Construct a new ByteString from a CString. The resulting ByteString is an immutable copy of the original CString, and is managed on the Haskell heap. The original CString must be null terminated.

O(n). Construct a new ByteString from a CStringLen. The resulting ByteString is an immutable copy of the original CStringLen. The ByteString is a normal Haskell value and will be managed on the Haskell heap.

Using ByteStrings as CStrings

valueuseAsCString :: ByteString -> (CString -> IO a) -> IO a
#

O(n) construction Use a ByteString with a function requiring a null-terminated CString. The CString is a copy and will be freed automatically; it must not be stored or used after the subcomputation finishes.

valueuseAsCStringLen :: ByteString -> (CStringLen -> IO a) -> IO a
#

O(n) construction Use a ByteString with a function requiring a CStringLen. As for useAsCString this function makes a copy of the original ByteString. It must not be stored or used after the subcomputation finishes.

Beware that this function is not required to add a terminating NUL byte at the end of the CStringLen it provides. If you need to construct a pointer to a null-terminated sequence, use useAsCString (and measure length independently if desired).

I/O with ByteStrings

0 declarations

Standard input and output

getContents. Read stdin strictly. Equivalent to hGetContents stdin The Handle is closed after the contents have been read.

valueinteract :: (ByteString -> ByteString) -> IO ()
#

The interact function takes a function of type ByteString -> ByteString as its argument. The entire input from the standard input device is passed to this function as its argument, and the resulting string is output on the standard output device.

Files

I/O with Handles

Read a handle's entire contents strictly into a ByteString.

This function reads chunks at a time, increasing the chunk size on each read. The final string is then reallocated to the appropriate size. For files > half of available memory, this may lead to memory exhaustion. Consider using readFile in this case.

The Handle is closed once the contents have been read, or if an exception is thrown.

valuehGet :: Handle -> Int -> IO ByteString
#

Read a ByteString directly from the specified Handle. This is far more efficient than reading the characters into a String and then using pack. First argument is the Handle to read from, and the second is the number of bytes to read. It returns the bytes read, up to n, or empty if EOF has been reached.

hGet is implemented in terms of hGetBuf.

If the handle is a pipe or socket, and the writing end is closed, hGet will behave as if EOF was reached.

valuehGetSome :: Handle -> Int -> IO ByteString
#

Like hGet, except that a shorter ByteString may be returned if there are not enough bytes immediately available to satisfy the whole request. hGetSome only blocks if there is no data available, and EOF has not yet been reached.

hGetNonBlocking is similar to hGet, except that it will never block waiting for data to become available, instead it returns only whatever data is available. If there is no data available to be read, hGetNonBlocking returns empty.

Note: on Windows and with Haskell implementation other than GHC, this function does not work correctly; it behaves identically to hGet.

Similar to hPut except that it will never block. Instead it returns any tail that did not get written. This tail may be empty in the case that the whole string was written, or the whole original string if nothing was written. Partial writes are also possible.

Note: on Windows and with Haskell implementation other than GHC, this function does not work correctly; it behaves identically to hPut.