HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Moduleos-string-2.0.7Haskell2010

System.OsString

An implementation of platform specific short OsString, which is:

  1. on windows wide char bytes ([Word16])

  2. on unix char bytes ([Word8])

It captures the notion of syscall specific encoding (or the lack thereof) to avoid roundtrip issues and memory fragmentation by using unpinned byte arrays. Bytes are not touched or interpreted.

  • 2 types
  • 77 values
  • Packageos-string-2.0.7
  • Exports79
  • LanguageHaskell2010
  • LicenceBSD-3-Clause
  • SourceOsString.hs

String types

1 declaration
newtypenewtype OsString
#

Newtype representing short operating system specific strings.

Internally this is either WindowsString or PosixString, depending on the platform. Both use unpinned ShortByteString for efficiency.

The constructor is only exported via System.OsString.Internal.Types, since dealing with the internals isn't generally recommended, but supported in case you need to write platform specific code.

Instances9Eq, Ord, Show, Generic, Semigroup, Monoid, …

OsString construction

9 declarations
valueencodeUtf :: MonadThrow m => String -> m OsString
#

Partial unicode friendly encoding.

On windows this encodes as UTF16-LE (strictly), which is a pretty good guess. On unix this encodes as UTF8 (strictly), which is a good guess.

Throws an EncodingException if encoding fails. If the input does not contain surrogate chars, you can use unsafeEncodeUtf.

valueencodeFS :: String -> IO OsString
#

Deprecated. Use System.OsPath.encodeFS from filepath

Like encodeUtf, except this mimics the behavior of the base library when doing filesystem operations (usually filepaths), which is:

  1. on unix, uses shady PEP 383 style encoding (based on the current locale, but PEP 383 only works properly on UTF-8 encodings, so good luck)

  2. on windows does permissive UTF-16 encoding, where coding errors generate Chars in the surrogate range

Looking up the locale requires IO. If you're not worried about calls to setFileSystemEncoding, then unsafePerformIO may be feasible (make sure to deeply evaluate the result to catch exceptions).

valueencodeLE :: String -> IO OsString
#

Like encodeUtf, except this mimics the behavior of the base library when doing string operations, which is:

  1. on unix this uses getLocaleEncoding

  2. on windows does permissive UTF-16 encoding, where coding errors generate Chars in the surrogate range

Looking up the locale requires IO. If you're not worried about calls to setFileSystemEncoding, then unsafePerformIO may be feasible (make sure to deeply evaluate the result to catch exceptions).

valueosstr :: QuasiQuoter
#

QuasiQuote an OsString. This accepts Unicode characters and encodes as UTF-8 on unix and UTF-16 on windows. If used as pattern, requires turning on the ViewPatterns extension.

OsString deconstruction

5 declarations
valuedecodeUtf :: MonadThrow m => OsString -> m String
#

Partial unicode friendly decoding.

On windows this decodes as UTF16-LE (strictly), which is a pretty good guess. On unix this decodes as UTF8 (strictly), which is a good guess. Note that filenames on unix are encoding agnostic char arrays.

Throws a EncodingException if decoding fails.

valuedecodeFS :: OsString -> IO String
#

Deprecated. Use System.OsPath.encodeFS from filepath

Like decodeUtf, except this mimics the behavior of the base library when doing filesystem operations (usually filepaths), which is:

  1. on unix, uses shady PEP 383 style encoding (based on the current locale, but PEP 383 only works properly on UTF-8 encodings, so good luck)

  2. on windows does permissive UTF-16 encoding, where coding errors generate Chars in the surrogate range

Looking up the locale requires IO. If you're not worried about calls to setFileSystemEncoding, then unsafePerformIO may be feasible (make sure to deeply evaluate the result to catch exceptions).

valuedecodeLE :: OsString -> IO String
#

Like decodeUtf, except this mimics the behavior of the base library when doing string operations, which is:

  1. on unix this uses getLocaleEncoding

  2. on windows does permissive UTF-16 encoding, where coding errors generate Chars in the surrogate range

Looking up the locale requires IO. If you're not worried about calls to setFileSystemEncoding, then unsafePerformIO may be feasible (make sure to deeply evaluate the result to catch exceptions).

Word types

1 declaration
newtypenewtype OsChar
#

Newtype representing a code unit.

On Windows, this is restricted to two-octet codepoints Word16, on POSIX one-octet (Word8).

Instances6Eq, Ord, Show, Generic, NFData, Rep

Word construction

1 declaration

Word deconstruction

1 declaration

Basic interface

10 declarations
valuelast :: HasCallStack => OsString -> OsChar
#

O(1) Extract the last element of a OsString, which must be finite and non-empty. An exception will be thrown in the case of an empty OsString.

This is a partial function, consider using unsnoc instead.

valuetail :: HasCallStack => OsString -> OsString
#

O(n) Extract the elements after the head of a OsString, which must be non-empty. An exception will be thrown in the case of an empty OsString.

This is a partial function, consider using uncons instead.

valuehead :: HasCallStack => OsString -> OsChar
#

O(1) Extract the first element of a OsString, which must be non-empty. An exception will be thrown in the case of an empty OsString.

This is a partial function, consider using uncons instead.

valueinit :: HasCallStack => OsString -> OsString
#

O(n) Return all the elements of a OsString except the last one. An exception will be thrown in the case of an empty OsString.

This is a partial function, consider using unsnoc instead.

Transforming OsString

3 declarations

Reducing OsStrings (folds)

8 declarations
valuefoldl :: (a -> OsChar -> a) -> a -> OsString -> a
#

foldl, applied to a binary operator, a starting value (typically the left-identity of the operator), and a OsString, reduces the OsString using the binary operator, from left to right.

valuefoldr :: (OsChar -> a -> a) -> a -> OsString -> a
#

foldr, applied to a binary operator, a starting value (typically the right-identity of the operator), and a OsString, reduces the OsString using the binary operator, from right to left.

Special folds

3 declarations

Generating and unfolding OsStrings

3 declarations
valuereplicate :: Int -> OsChar -> OsString
#

O(n) replicate n x is a OsString of length n with x the value of every element. The following holds:

replicate w c = unfoldr w (\u -> Just (u,u)) c
valueunfoldr :: (a -> Maybe (OsChar, a)) -> a -> OsString
#

O(n), where n is the length of the result. The unfoldr function is analogous to the List 'unfoldr'. unfoldr builds a OsString from a seed value. The function takes the element and returns Nothing if it is done producing the OsString or returns Just (a,b), in which case, a is the next byte in the string, and b is the seed value for further production.

This function is not efficient/safe. It will build a list of [Word8] and run the generator until it returns Nothing, otherwise recurse infinitely, then finally create a OsString.

If you know the maximum length, consider using unfoldrN.

Examples:

   unfoldr (\x -> if x <= 5 then Just (x, x + 1) else Nothing) 0
== pack [0, 1, 2, 3, 4, 5]
valueunfoldrN :: Int -> (a -> Maybe (OsChar, a)) -> a -> (OsString, Maybe a)
#

O(n) Like unfoldr, unfoldrN builds a OsString from a seed value. However, the length of the result is limited by the first argument to unfoldrN. This function is more efficient than unfoldr when the maximum length of the result is known.

The following equation relates unfoldrN and unfoldr:

fst (unfoldrN n f s) == take n (unfoldr f s)

Substrings

0 declarations

Breaking strings

valuetakeEnd :: Int -> OsString -> OsString
#

O(n) takeEnd n xs is equivalent to drop (length xs - n) xs. Takes n elements from end of bytestring.

Example3 expressions
takeEnd 3 "abcdefg""efg"takeEnd 0 "abcdefg"""takeEnd 4 "abc""abc"
valuedropEnd :: Int -> OsString -> OsString
#

O(n) dropEnd n xs is equivalent to take (length xs - n) xs. Drops n elements from end of bytestring.

Example3 expressions
dropEnd 3 "abcdefg""abcd"dropEnd 0 "abcdefg""abcdefg"dropEnd 4 "abc"""
valuespanEnd :: (OsChar -> Bool) -> OsString -> (OsString, OsString)
#

Returns the longest (possibly empty) suffix of elements satisfying the predicate and the remainder of the string.

spanEnd p is equivalent to breakEnd (not . p) and to (takeWhileEnd p &&& dropWhileEnd p).

We have

spanEnd (not . isSpace) "x y z" == ("x y ", "z")

and

spanEnd (not . isSpace) sbs
   ==
let (x, y) = span (not . isSpace) (reverse sbs) in (reverse y, reverse x)
valuesplit :: OsChar -> OsString -> [OsString]
#

O(n) Break a OsString into pieces separated by the byte argument, consuming the delimiter. I.e.

split 10  "a\nb\nd\ne" == ["a","b","d","e"]   -- fromEnum '\n' == 10
split 97  "aXaXaXa"    == ["","X","X","X",""] -- fromEnum 'a' == 97
split 120 "x"          == ["",""]             -- fromEnum 'x' == 120
split undefined ""     == []                  -- and not [""]

and

intercalate [c] . split c == id
split == splitWith . (==)
valuesplitWith :: (OsChar -> Bool) -> OsString -> [OsString]
#

O(n) Splits a OsString into components delimited by separators, where the predicate returns True for a separator element. The resulting components do not contain the separators. Two adjacent separators result in an empty component in the output. eg.

splitWith (==97) "aabbaca" == ["","","bb","c",""] -- fromEnum 'a' == 97
splitWith undefined ""     == []                  -- and not [""]

Predicates

3 declarations
valueisSuffixOf :: OsString -> OsString -> Bool
#

O(n) The isSuffixOf function takes two OsStrings and returns True iff the first is a suffix of the second.

The following holds:

isSuffixOf x y == reverse x `isPrefixOf` reverse y

Search for arbitrary susbstrings

Break a string on a substring, returning a pair of the part of the string prior to the match, and the rest of the string.

The following relationships hold:

break (== c) l == breakSubstring (singleton c) l

For example, to tokenise a string, dropping delimiters:

tokenise x y = h : if null t then [] else tokenise x (drop (length x) t)
    where (h,t) = breakSubstring x y

To skip to the first occurrence of a string:

snd (breakSubstring x y)

To take the parts of a string before a delimiter:

fst (breakSubstring x y)

Note that calling `breakSubstring x` does some preprocessing work, so you should avoid unnecessarily duplicating breakSubstring calls with the same pattern.

Searching OsStrings

0 declarations

Searching by equality

valuefind :: (OsChar -> Bool) -> OsString -> Maybe OsChar
#

O(n) The find function takes a predicate and a OsString, and returns the first element in matching the predicate, or Nothing if there is no such element.

find f p = case findIndex f p of Just n -> Just (p ! n) ; _ -> Nothing
valuepartition :: (OsChar -> Bool) -> OsString -> (OsString, OsString)
#

O(n) The partition function takes a predicate a OsString and returns the pair of OsStrings with elements which do and do not satisfy the predicate, respectively; i.e.,

partition p bs == (filter p sbs, filter (not . p) sbs)

Indexing OsStrings

8 declarations
valuecount :: OsChar -> OsString -> Int
#

count returns the number of times its argument appears in the OsString

Coercions

1 declaration