pandas read_json in chunks but still has memory error

Viewed 2464

I'm trying to read and process a large json file(~16G) but it keeps having memory error even if I read in small chunks by specifying chunksize=500. My code:

i=0
header = True
for chunk in pd.read_json('filename.json.tsv', lines=True, chunksize=500):
        print("Processing chunk ", i)
        process_chunk(chunk, i)
        i+=1
        header = False
def process_chunk(chunk, header, i):
    pk_file = 'data/pk_files/500_chunk_'+str(i)+'.pk'
    get_data_pk(chunk, pk_file) #load and process some columns and save into a pk file for future processing
    preds = get_preds(pk_file) #SVM prediction
    chunk['prediction'] = preds #append result column
    chunk.to_csv('result.csv', header = header, mode='a')

The process_chunk function basically reads in each chunk and append a new column to it.

When I use a smaller file it works, also works well if I specify nrows=5000 in the read_json function. Seems for some reason it still requires full file-size memory despite the chunksize parameter.

Any idea? Thanks!

1 Answers

I had the same strange problem in one of my project's virtual env with pandas v1.1.2. Downgrading pandas to v1.0.5 seems to solve the problem.

Related