Concat changes the category type to object / float64

Viewed 2291

If I am reading just a piece of csv I get the following data structure

<class 'pandas.core.frame.DataFrame'>
MultiIndex: 100000 entries, (2015-11-01 00:00:00, 4980770) to (2016-06-01 00:00:00, 8850573)
Data columns (total 5 columns):
CHANNEL          100000 non-null category
MCC              92660 non-null category
DOMESTIC_FLAG    100000 non-null category
AMOUNT           100000 non-null float32
CNT              100000 non-null uint8
dtypes: category(3), float32(1), uint8(1)
memory usage: 1.9+ MB

If I am reading the whole csv and concat the blocks as per above I get the following structure:

<class 'pandas.core.frame.DataFrame'>
MultiIndex: 30345312 entries, (2015-11-01 00:00:00, 4980770) to (2015-08-01 00:00:00, 88838)
Data columns (total 5 columns):
CHANNEL          object
MCC              float64
DOMESTIC_FLAG    category
AMOUNT           float32
CNT              uint8
dtypes: category(1), float32(1), float64(1), object(1), uint8(1)
memory usage: 784.6+ MB

Why are categorical variables changed to object / float64? How can I avoid this type change? Esp. the float64

This is the concatenation code:

df = pd.concat([process(chunk) for chunk in reader])

process function is just doing some cleaning and type assignments

1 Answers
Related