0-based offset identifying the raw location in the StringBuffer.
The lexer increments the BufPos every time a character (UTF-8 code point)
is read from the input buffer. As UTF-8 is a variable-length encoding and
StringBuffer needs a byte offset for indexing, a BufPos cannot be used
for indexing.
The parser guarantees that BufPos are monotonic. See #17632. This means
that syntactic constructs that appear later in the StringBuffer are guaranteed to
have a higher BufPos. Contrast that with RealSrcLoc, which does *not* make the
analogous guarantee about higher line/column numbers.
This is due to #line and {-# LINE ... #-} pragmas that can arbitrarily
modify RealSrcLoc. Notice how setSrcLoc and resetAlrLastLoc in
GHC.Parser.Lexer update PsLoc, modifying RealSrcLoc but preserving
BufPos.
Monotonicity makes BufPos useful to determine the order in which syntactic
elements appear in the source. Consider this example (haddockA041 in the test suite):
haddockA041.hs
{-# LANGUAGE CPP #-}
-- | Module header documentation
module Comments_and_CPP_include where
#include "IncludeMe.hs"
IncludeMe.hs:
-- | Comment on T
data T = MkT -- ^ Comment on MkT
After the C preprocessor runs, the StringBuffer will contain a program that
looks like this (unimportant lines at the beginning removed):
# 1 "haddockA041.hs"
{-# LANGUAGE CPP #-}
-- | Module header documentation
module Comments_and_CPP_include where
# 1 "IncludeMe.hs" 1
-- | Comment on T
data T = MkT -- ^ Comment on MkT
# 7 "haddockA041.hs" 2
The line pragmas inserted by CPP make the error messages more informative.
The downside is that we can't use RealSrcLoc to determine the ordering of
syntactic elements.
With RealSrcLoc, we have the following location information recorded in the AST:
* The module name is located at haddockA041.hs:3:8-31
* The Haddock comment "Comment on T" is located at IncludeMe:1:1-17
* The data declaration is located at IncludeMe.hs:2:1-32
Is the Haddock comment located between the module name and the data
declaration? This is impossible to tell because the locations are not
comparable; they even refer to different files.
On the other hand, with BufPos, we have the following location information:
* The module name is located at 846-870
* The Haddock comment "Comment on T" is located at 898-915
* The data declaration is located at 916-928
Aside: if you're wondering why the numbers are so high, try running
ghc -E haddockA041.hs
and see the extra fluff that CPP inserts at the start of the file.
For error messages, BufPos is not useful at all. On the other hand, this is
exactly what we need to determine the order of syntactic elements:
870 < 898, therefore the Haddock comment appears *after* the module name.
915 < 916, therefore the Haddock comment appears *before* the data declaration.
We use BufPos in in GHC.Parser.PostProcess.Haddock to associate Haddock
comments with parts of the AST using location information (#17544).