strange indexing mechanism of pandas.read_csv function with chunksize option

Viewed 437

Due to the huge data size, we used pandas to process data, but a very strange phenomenon occurred. The pseudo code looks like this:

reader = pd.read_csv(IN_FILE, chunksize = 1000, engine='c')
for chunk in reader:
    result = []
    for line in chunk.tolist():
         temp = complicated_process(chunk)  # this involves a very complicated processing, so here is just a simplified version
         result.append(temp)
    chunk['new_series'] = pd.series(result)
    chunk.to_csv(OUT_TILE, index=False, mode='a')

We can confirm each loop of result is not empty. But only in the first time of the loop, line chunk['new_series'] = pd.series(result) has result, the rest are empty. Therefore, only the first chunk of the output contains new_series, the rest are empty.

Did we miss anything here? Thanks in advance.

2 Answers
Related