HORIZON HASKELLDocslts/ghc-9.10.xc74966e2026-09-27Search names, modules, packages, or :: a typeCtrl K

GHC 9.10.3 · lts/ghc-9.10.x · c74966e · 2026-09-27

Modulecassava-0.5.3.2Haskell2010

Data.Csv

This module implements encoding and decoding of comma-separated values (CSV) data. The implementation is RFC 4180 compliant, with the following extensions:

  • Empty lines are ignored.

  • Non-escaped fields may contain any characters except double-quotes, commas, carriage returns, and newlines.

  • Escaped fields may contain any characters (but double-quotes need to be escaped).

  • 13 types
  • 11 classes
  • 32 values
  • Packagecassava-0.5.3.2
  • Exports56
  • LanguageHaskell2010
  • LicenceBSD-3-Clause
  • SourceCsv.hs

Usage examples

0 declarations

Encoding standard Haskell types:

Example3 expressions
:set -XOverloadedStringsimport Data.Text (Text)encode [("John" :: Text, 27 :: Int), ("Jane", 28)]"John,27\r\nJane,28\r\n"

Since we enabled the -XOverloadedStrings extension, string literals are polymorphic and we have to supply a type signature as the compiler couldn't deduce which string type (i.e. String, ShortText, or Text) we want to use. In most cases type inference will infer the type from the context and you can omit type signatures.

Decoding standard Haskell types:

Example2 expressions
import Data.Vector (Vector)decode NoHeader "John,27\r\nJane,28\r\n" :: Either String (Vector (Text, Int))Right [("John",27),("Jane",28)]

We pass NoHeader as the first argument to indicate that the CSV input data isn't preceded by a header.

In practice, the return type of decode rarely needs to be given, as it can often be inferred from the context.

Encoding and decoding custom data types

To encode and decode your own data types you need to defined instances of either ToRecord and FromRecord or ToNamedRecord and FromNamedRecord. The former is used for encoding/decoding using the column index and the latter using the column name.

There are two ways to to define these instances, either by manually defining them or by using GHC generics to derive them automatically.

Index-based record conversion

GHC.Generics-derived:

{-# LANGUAGE DeriveGeneric #-}

import Data.Text    (Text)
import GHC.Generics (Generic)

data Person = Person { name :: !Text , salary :: !Int }
    deriving (Generic, Show)

instance FromRecord Person
instance ToRecord Person

Manually defined:

import Control.Monad (mzero)

data Person = Person { name :: !Text , salary :: !Int }
    deriving (Show)

instance FromRecord Person where
    parseRecord v
        | length v == 2 = Person <$> v .! 0 <*> v .! 1
        | otherwise     = mzero
instance ToRecord Person where
    toRecord (Person name' age') = record [
        toField name', toField age']

We can now use e.g. encode and decode to encode and decode our data type.

Encoding:

Example1 expression
encode [Person ("John" :: Text) 27]"John,27\r\n"

Decoding:

Example1 expression
decode NoHeader "John,27\r\n" :: Either String (Vector Person)Right [Person {name = "John", salary = 27}]

Name-based record conversion

GHC.Generics-derived:

{-# LANGUAGE DeriveGeneric #-}

import Data.Text    (Text)
import GHC.Generics (Generic)

data Person = Person { name :: !Text , salary :: !Int }
    deriving (Generic, Show)

instance FromNamedRecord Person
instance ToNamedRecord Person
instance DefaultOrdered Person

Manually defined:

data Person = Person { name :: !Text , salary :: !Int }
    deriving (Show)

instance FromNamedRecord Person where
    parseNamedRecord m = Person <$> m .: "name" <*> m .: "salary"
instance ToNamedRecord Person where
    toNamedRecord (Person name salary) = namedRecord [
        "name" .= name, "salary" .= salary]
instance DefaultOrdered Person where
    headerOrder _ = header ["name", "salary"]

We can now use e.g. encodeDefaultOrderedByName (or encodeByName with an explicit header order) and decodeByName to encode and decode our data type.

Encoding:

Example1 expression
encodeDefaultOrderedByName [Person ("John" :: Text) 27]"name,salary\r\nJohn,27\r\n"

Decoding:

Example1 expression
decodeByName "name,salary\r\nJohn,27\r\n" :: Either String (Header, Vector Person)Right (["name","salary"],[Person {name = "John", salary = 27}])

Reading/writing CSV files

Demonstration of reading from a CSV file/ writing to a CSV file using the incremental API:

{-# LANGUAGE BangPatterns      #-}
{-# LANGUAGE DeriveGeneric     #-}
{-# LANGUAGE LambdaCase        #-}
{-# LANGUAGE OverloadedStrings #-}

-- from base
import GHC.Generics
import System.IO
import System.Exit (exitFailure)
-- from bytestring
import Data.ByteString (ByteString, hGetSome, empty)
import qualified Data.ByteString.Lazy as BL
-- from cassava
import Data.Csv.Incremental
import Data.Csv (FromRecord, ToRecord)

data Person = Person
  { name :: !ByteString
  , age  :: !Int
  } deriving (Show, Eq, Generic)

instance FromRecord Person
instance ToRecord Person

persons :: [Person]
persons = [Person "John Doe" 19, Person "Smith" 20]

writeToFile :: IO ()
writeToFile = do
  BL.writeFile "persons.csv" $ encode $
    foldMap encodeRecord persons

feed :: (ByteString -> Parser Person) -> Handle -> IO (Parser Person)
feed k csvFile = do
  hIsEOF csvFile >>= \case
    True  -> return $ k empty
    False -> k <$> hGetSome csvFile 4096

readFromFile :: IO ()
readFromFile = do
  withFile "persons.csv" ReadMode $ \ csvFile -> do
    let loop !_ (Fail _ errMsg) = do putStrLn errMsg; exitFailure
        loop acc (Many rs k)    = loop (acc <> rs) =<< feed k csvFile
        loop acc (Done rs)      = print (acc <> rs)
    loop [] (decode NoHeader)

main :: IO ()
main = do
  writeToFile
  readFromFile

Treating CSV data as opaque byte strings

0 declarations

Sometimes you might want to work with a CSV file which contents is unknown to you. For example, you might want remove the second column of a file without knowing anything about its content. To parse a CSV file to a generic representation, just convert each record to a Vector ByteString value, like so:

Example2 expressions
import Data.ByteString (ByteString)decode NoHeader "John,27\r\nJane,28\r\n" :: Either String (Vector (Vector ByteString))Right [["John","27"],["Jane","28"]]

As the example output above shows, all the fields are returned as uninterpreted ByteString values.

Custom type conversions for fields

0 declarations

Most of the time the existing FromField and ToField instances do what you want. However, if you need to parse a different format (e.g. hex) but use a type (e.g. Int) for which there's already a FromField instance, you need to use a newtype. Example:

newtype Hex = Hex Int

parseHex :: ByteString -> Parser Int
parseHex = ...

instance FromField Hex where
    parseField s = Hex <$> parseHex s

Other than giving an explicit type signature, you can pattern match on the newtype constructor to indicate which type conversion you want to have the library use:

case decode NoHeader "0xff,0xaa\r\n0x11,0x22\r\n" of
    Left err -> putStrLn err
    Right v  -> forM_ v $ \ (Hex val1, Hex val2) ->
        print (val1, val2)

If a field might be in one several different formats, you can use a newtype to normalize the result:

newtype HexOrDecimal = HexOrDecimal Int

instance FromField DefaultToZero where
    parseField s = case runParser (parseField s :: Parser Hex) of
        Left err -> HexOrDecimal <$> parseField s  -- Uses Int instance
        Right n  -> pure $ HexOrDecimal n

You can use the unit type, (), to ignore a column. The parseField method for () doesn't look at the Field and thus always decodes successfully. Note that it lacks a corresponding ToField instance. Example:

case decode NoHeader "foo,1\r\nbar,22" of
    Left  err -> putStrLn err
    Right v   -> forM_ v $ \ ((), i) -> print (i :: Int)

Dealing with bad data

If your input might contain invalid fields, you can write a custom FromField instance to deal with them. Example:

newtype DefaultToZero = DefaultToZero Int

instance FromField DefaultToZero where
    parseField s = case runParser (parseField s) of
        Left err -> pure $ DefaultToZero 0
        Right n  -> pure $ DefaultToZero n

Encoding and decoding

7 declarations

Encoding and decoding is a two step process. To encode a value, it is first converted to a generic representation, using either ToRecord or ToNamedRecord. The generic representation is then encoded as CSV data. To decode a value the process is reversed and either FromRecord or FromNamedRecord is used instead. Both these steps are combined in the encode and decode functions.

datadata HasHeader
#

Is the CSV data preceded by a header?

Constructors

  • HasHeader

    The CSV data is preceded by a header

  • NoHeader

    The CSV data is not preceded by a header

classclass DefaultOrdered a where
#

A type that has a default field order when converted to CSV. This class lets you specify how to get the headers to use for a record type that's an instance of ToNamedRecord.

To derive an instance, the type is required to only have one constructor and that constructor must have named fields (also known as selectors) for all fields.

Right: data Foo = Foo { foo :: !Int }

Wrong: data Bar = Bar Int

If you try to derive an instance using GHC generics and your type doesn't have named fields, you will get an error along the lines of:

<interactive>:9:10:
    No instance for (DefaultOrdered (M1 S NoSelector (K1 R Char) ()))
      arising from a use of ‘Data.Csv.Conversion.$gdmheader’
    In the expression: Data.Csv.Conversion.$gdmheader
    In an equation for ‘header’:
        header = Data.Csv.Conversion.$gdmheader
    In the instance declaration for ‘DefaultOrdered Foo’

Methods

Encoding and decoding options

These functions can be used to control how data is encoded and decoded. For example, they can be used to encode data in a tab-separated format instead of in a comma-separated format.

datadata DecodeOptions
#

Options that controls how data is decoded. These options can be used to e.g. decode tab-separated data instead of comma-separated data.

To avoid having your program stop compiling when new fields are added to DecodeOptions, create option records by overriding values in defaultDecodeOptions. Example:

myOptions = defaultDecodeOptions {
      decDelimiter = fromIntegral (ord '\t')
    }

Constructors

Instances2Eq, Show
datadata EncodeOptions
#

Options that controls how data is encoded. These options can be used to e.g. encode data in a tab-separated format instead of in a comma-separated format.

To avoid having your program stop compiling when new fields are added to EncodeOptions, create option records by overriding values in defaultEncodeOptions. Example:

myOptions = defaultEncodeOptions {
      encDelimiter = fromIntegral (ord '\t')
    }

N.B. The encDelimiter must not be the quote character (i.e. ") or one of the record separator characters (i.e. \n or \r).

Constructors

Instances2Eq, Show
datadata Quoting
#

Should quoting be applied to fields, and at which level?

Constructors

Instances2Eq, Show
  • Eq QuotingDefined in cassava-0.5.3.2 · Data.Csv.Encoding
  • Show QuotingDefined in cassava-0.5.3.2 · Data.Csv.Encoding

Core CSV types

6 declarations
typetype Csv = Vector Record
#

CSV data represented as a Haskell vector of vector of bytestrings.

typetype Header = Vector Name
#

The header corresponds to the first line a CSV file. Not all CSV files have a header.

typetype Name = ByteString
#

A header has one or more names, describing the data in the column following the name.

Type conversion

0 declarations

There are two ways to convert CSV records to and from and user-defined data types: index-based conversion and name-based conversion.

Index-based record conversion

Index-based conversion lets you convert CSV records to and from user-defined data types by referring to a field's position (its index) in the record. The first column in a CSV file is given index 0, the second index 1, and so on.

classclass FromRecord a where
#

A type that can be converted from a single CSV record, with the possibility of failure.

When writing an instance, use empty, mzero, or fail to make a conversion fail, e.g. if a Record has the wrong number of columns.

Given this example data:

John,56
Jane,55

here's an example type and instance:

data Person = Person { name :: !Text, age :: !Int }

instance FromRecord Person where
    parseRecord v
        | length v == 2 = Person <$>
                          v .! 0 <*>
                          v .! 1
        | otherwise     = mzero

Methods

Instances18FromRecord, …
newtypenewtype Parser a
#

Conversion of a field to a value might fail e.g. if the field is malformed. This possibility is captured by the Parser type, which lets you compose several field conversions together in such a way that if any of them fail, the whole record conversion fails.

Instances8Monad, Functor, MonadFail, Applicative, Alternative, MonadPlus, …
valuerunParser :: Parser a -> Either String a
#

Run a Parser, returning either Left errMsg or Right result. Forces the value in the Left or Right constructors to weak head normal form.

You most likely won't need to use this function directly, but it's included for completeness.

valueindex :: FromField a => Record -> Int -> Parser a
#

Retrieve the nth field in the given record. The result is empty if the value cannot be converted to the desired type. Raises an exception if the index is out of bounds.

index is a simple convenience function that is equivalent to parseField (v ! idx). If you're certain that the index is not out of bounds, using unsafeIndex is somewhat faster.

classclass ToRecord a where
#

A type that can be converted to a single CSV record.

An example type and instance:

data Person = Person { name :: !Text, age :: !Int }

instance ToRecord Person where
    toRecord (Person name age) = record [
        toField name, toField age]

Outputs data on this form:

John,56
Jane,55

Methods

Instances18ToRecord, …
newtypenewtype Only a
#

The 1-tuple type or single-value "collection".

This type is structurally equivalent to the Identity type, but its intent is more about serving as the anonymous 1-tuple type missing from Haskell for attaching typeclass instances.

Parameter usage example:

encodeSomething (Only (42::Int))

Result usage example:

xs <- decodeSomething
forM_ xs $ \(Only id) -> {- ... -}

Constructors

Instances11Functor, Eq, Data, Ord, Read, Show, …

Name-based record conversion

Name-based conversion lets you convert CSV records to and from user-defined data types by referring to a field's name. The names of the fields are defined by the first line in the file, also known as the header. Name-based conversion is more robust to changes in the file structure e.g. to reording or addition of columns, but can be a bit slower.

classclass FromNamedRecord a where
#

A type that can be converted from a single CSV record, with the possibility of failure.

When writing an instance, use empty, mzero, or fail to make a conversion fail, e.g. if a Record has the wrong number of columns.

Given this example data:

name,age
John,56
Jane,55

here's an example type and instance:

{-# LANGUAGE OverloadedStrings #-}

data Person = Person { name :: !Text, age :: !Int }

instance FromNamedRecord Person where
    parseNamedRecord m = Person <$>
                         m .: "name" <*>
                         m .: "age"

Note the use of the OverloadedStrings language extension which enables ByteString values to be written as string literals.

Instances2FromNamedRecord
classclass ToNamedRecord a where
#

A type that can be converted to a single CSV record.

An example type and instance:

data Person = Person { name :: !Text, age :: !Int }

instance ToNamedRecord Person where
    toNamedRecord (Person name age) = namedRecord [
        "name" .= name, "age" .= age]

Methods

Instances2ToNamedRecord

Field conversion

The FromField and ToField classes define how to convert between Fields and values you care about (e.g. Ints). Most of the time you don't need to write your own instances as the standard ones cover most use cases.

classclass FromField a where
#

A type that can be converted from a single CSV field, with the possibility of failure.

When writing an instance, use empty, mzero, or fail to make a conversion fail, e.g. if a Field can't be converted to the given type.

Example type and instance:

{-# LANGUAGE OverloadedStrings #-}

data Color = Red | Green | Blue

instance FromField Color where
    parseField s
        | s == "R"  = pure Red
        | s == "G"  = pure Green
        | s == "B"  = pure Blue
        | otherwise = mzero

Methods

Instances28FromField, …
  • FromField ByteStringDefined in cassava-0.5.3.2 · Data.Csv.Conversion
  • FromField ByteStringDefined in cassava-0.5.3.2 · Data.Csv.Conversion
  • FromField ShortByteStringDefined in cassava-0.5.3.2 · Data.Csv.Conversion
  • FromField IntegerDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts a signed decimal number. Ignores whitespace.

  • FromField NaturalDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts an unsigned decimal number. Ignores whitespace.

  • FromField Int16Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts a signed decimal number. Ignores whitespace.

  • FromField Int32Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts a signed decimal number. Ignores whitespace.

  • FromField Int64Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts a signed decimal number. Ignores whitespace.

  • FromField Int8Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts a signed decimal number. Ignores whitespace.

  • FromField Word16Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts an unsigned decimal number. Ignores whitespace.

  • FromField Word32Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts an unsigned decimal number. Ignores whitespace.

  • FromField Word64Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts an unsigned decimal number. Ignores whitespace.

  • FromField Word8Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts an unsigned decimal number. Ignores whitespace.

  • FromField CharDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Assumes UTF-8 encoding.

  • FromField DoubleDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts same syntax as rational. Ignores whitespace.

  • FromField FloatDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts same syntax as rational. Ignores whitespace.

  • FromField IntDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts a signed decimal number. Ignores whitespace.

  • FromField WordDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts an unsigned decimal number. Ignores whitespace.

  • FromField ScientificDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Accepts the same syntax as rational. Ignores whitespace.

  • FromField TextDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Assumes UTF-8 encoding. Fails on invalid byte sequences.

  • FromField TextDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Assumes UTF-8 encoding. Fails on invalid byte sequences.

  • FromField ShortTextDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Assumes UTF-8 encoding. Fails on invalid byte sequences.

  • FromField ()Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Ignores the Field. Always succeeds.

  • FromField [Char]Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Assumes UTF-8 encoding. Fails on invalid byte sequences.

  • FromField a => FromField (Identity a)Defined in cassava-0.5.3.2 · Data.Csv.Conversion
  • FromField a => FromField (Maybe a)Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Nothing if the Field is empty, Just otherwise.

  • FromField a => FromField (Either Field a)Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Left field if conversion failed, Right otherwise.

  • FromField a => FromField (Const a b)Defined in cassava-0.5.3.2 · Data.Csv.Conversion
classclass ToField a where
#

A type that can be converted to a single CSV field.

Example type and instance:

{-# LANGUAGE OverloadedStrings #-}

data Color = Red | Green | Blue

instance ToField Color where
    toField Red   = "R"
    toField Green = "G"
    toField Blue  = "B"

Methods

Instances26ToField, …
  • ToField ByteStringDefined in cassava-0.5.3.2 · Data.Csv.Conversion
  • ToField ByteStringDefined in cassava-0.5.3.2 · Data.Csv.Conversion
  • ToField ShortByteStringDefined in cassava-0.5.3.2 · Data.Csv.Conversion
  • ToField IntegerDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding with optional sign.

  • ToField NaturalDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding.

  • ToField Int16Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding with optional sign.

  • ToField Int32Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding with optional sign.

  • ToField Int64Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding with optional sign.

  • ToField Int8Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding with optional sign.

  • ToField Word16Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding.

  • ToField Word32Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding.

  • ToField Word64Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding.

  • ToField Word8Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding.

  • ToField CharDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses UTF-8 encoding.

  • ToField DoubleDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal notation or scientific notation, depending on the number.

  • ToField FloatDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal notation or scientific notation, depending on the number.

  • ToField IntDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding with optional sign.

  • ToField WordDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal encoding.

  • ToField ScientificDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses decimal notation or scientific notation, depending on the number.

  • ToField TextDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses UTF-8 encoding.

  • ToField TextDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses UTF-8 encoding.

  • ToField ShortTextDefined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses UTF-8 encoding.

  • ToField [Char]Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Uses UTF-8 encoding.

  • ToField a => ToField (Identity a)Defined in cassava-0.5.3.2 · Data.Csv.Conversion
  • ToField a => ToField (Maybe a)Defined in cassava-0.5.3.2 · Data.Csv.Conversion

    Nothing is encoded as an empty field.

  • ToField a => ToField (Const a b)Defined in cassava-0.5.3.2 · Data.Csv.Conversion

Generic record conversion

There may be times that you do not want to manually write out class instances for record conversion, but you can't rely upon the default instances (e.g. you can't create field names that match the actual column names in expected data).

For example, consider you have a type MyType where you have prefixed certain columns with an underscore, but in the actual data they're not. You can then write:

myOptions :: Options
myOptions = defaultOptions { fieldLabelModifier = rmUnderscore }
  where
    rmUnderscore ('_':str) = str
    rmUnderscore str       = str

instance ToNamedRecord MyType where
  toNamedRecord = genericToNamedRecord myOptions

instance FromNamedRecord MyType where
  parseNamedRecord = genericParseNamedRecord myOptions

instance DefaultOrdered MyType where
  headerOrder = genericHeaderOrder myOptions

Generic type conversion options

newtypenewtype Options
#

Options to customise how to generically encode/decode your datatype to/from CSV.

Instances1Show
  • Show OptionsDefined in cassava-0.5.3.2 · Data.Csv.Conversion

Generic type conversion class name

NOTE: Only the class names are exposed in order to make it possible to write type signatures referring to these classes

classclass GToRecord (a :: k -> Type) f where
#
Instances8GToRecord, …
classclass GToNamedRecordHeader (a :: k -> Type) where
#
Instances6GToNamedRecordHeader