This to me looks very suspicious I have debugged through it and I am still not sure it is a bug or not.
I have prepared sample data to replicate this issue, the file is not small (25MB, real-life quotes data for a typical trading day). The download link is here and passcode is 654321. I believe a day's worth of quotes data is necessary for this problem in case you think it's too much.
# My system locale in R is this (Europe/London).
# My system time is `BST`.
# The sample quotes data is in timezone `US/Eastern`;
# this implies that the market hours in data file conforms
# to actual US market hours (9:30-16:00), and not adjusted to
# whatever my system timezone is.
# In my case it would've been 13:30-20:00.
# I'm running R-4.2.1 on Debian 10 based Linux.
> Sys.getlocale()
[1] "LC_CTYPE=en_GB.UTF-8;LC_NUMERIC=C;LC_TIME=en_GB.UTF-8;LC_COLLATE=en_GB.UTF-8;LC_MONETARY=en_GB.UTF-8;LC_MESSAGES=en_GB.UTF-8;LC_PAPER=en_GB.UTF-8;LC_NAME=C;LC_ADDRESS=C;LC_TELEPHONE=C;LC_MEASUREMENT=en_GB.UTF-8;LC_IDENTIFICATION=C"
# helper function to read quotes data in a `highfrequency` conforming format.
read.quotes <- function(path){
quotes.data <- readr::read_csv(path,
show_col_types = F,
col_types = vroom::cols(ask_size="i",
bid_size="i",
conditions="c"),
locale = readr::locale(tz = "US/Eastern"))
colnames(quotes.data) <- c("DateTime",
"Symbol",
"AskEx",
"AskPrice",
"AskSize",
"BidEx",
"BidPrice",
"BidSize",
"Conditions",
"Tape")
quotes.data <- quotes.data %>%
dplyr::select(DateTime,
BidEx,
BidPrice,
BidSize,
AskPrice,
AskSize,
Symbol) %>%
dplyr::rename(DT=DateTime,
EX=BidEx,
BID=BidPrice,
BIDSIZ=BidSize,
OFR=AskPrice,
OFRSIZ=AskSize,
SYMBOL=Symbol) %>%
data.table::as.data.table
data.table::setkeyv(quotes.data, c("SYMBOL", "DT"))
quotes.data
}
# This function shall give us a data object that's the same as the package's `sampleQDataRaw`.
# The package's sample data is artificial, therefore in nature these two data objects are not comparable;
# Creating the correct object is essential.
# We should expect same output for data types with following statements.
> qdata <- read.quotes(path_to_the_file_on_your_computer)
> str(highfrequency::sampleQDataRaw)
> str(qdata)
# Observe that the qData ends at 11:59, not 15:59 or 16:00
# which is the correct US market close time,
# in timezones like `EDT`, `US/Eastern`, or `America/New York`.
> highfrequency::quotesCleanup(qDataRaw=qdata)
# output
[1] "The NYSE is the exchange with the highest volume."
$qData
DT SYMBOL BID OFR OFRSIZ BIDSIZ EX MIDQUOTE
1: 2022-05-02 09:30:00.951 GS 305.49 306.00 1 1 N 305.745
2: 2022-05-02 09:30:01.031 GS 305.49 305.99 1 1 N 305.740
3: 2022-05-02 09:30:01.031 GS 305.49 306.00 1 1 N 305.745
4: 2022-05-02 09:30:01.031 GS 305.49 305.99 1 1 N 305.740
5: 2022-05-02 09:30:01.031 GS 305.49 305.98 1 1 N 305.735
---
82457: 2022-05-02 11:59:59.758 GS 308.63 308.84 1 2 N 308.735
82458: 2022-05-02 11:59:59.767 GS 308.63 308.84 1 1 N 308.735
82459: 2022-05-02 11:59:59.827 GS 308.63 308.84 1 2 N 308.735
82460: 2022-05-02 11:59:59.827 GS 308.63 308.84 1 3 N 308.735
82461: 2022-05-02 11:59:59.997 GS 308.63 308.84 1 3 N 308.735
$report
initialObservations removedFromZeroQuotes removedOutsideExchangeHours removedFromSelectingExchange
413547 1 224780 106255
removedFromNegativeSpread removedFromLargeSpread removedFromMergeTimestamp removedOutliers
2 0 48 0
finalObservations
82461
Someone running a system in US locale (or physically based in US/Eastern) loads up my data and fires away with this clean-up process, if she/he gets the correct output it would indicate quotesCleanup does not correctly respect the timezone other than US (given it's US stock market).
If the US person does get the same result as I did here, then quotesCleanup might have different issue then.
Unless, it is that my data is malformed and this is indeed what I failed to diagnose so far (I don't believe it is the case, not that I didn't try).
Internally highfrequency has a function called exchangeHoursOnly to extract rows that fall between market open and close time; it converts POSIXct to numeric with tz information, but it seems as.numeric did not respect tz info. It's not very clear to see from their code, unfortunately.