How to split CSV files as per number of rows specified?

Viewed 133729

I've CSV file (around 10,000 rows ; each row having 300 columns) stored on LINUX server. I want to break this CSV file into 500 CSV files of 20 records each. (Each having same CSV header as present in original CSV)

Is there any linux command to help this conversion?

6 Answers

This should work !!!

file_name = Name of the file you want to split.
10000 = Number of rows each split file would contain
file_part_ = Prefix of split file name (file_part_0,file_part_1,file_part_2..etc goes on)

split -d -l 10000 file_name.csv file_part_

One-liner which preserves the header row in each split file. This example gives you 999 lines of data and one header row per file.

cat bigFile.csv | parallel --header : --pipe -N999 'cat >file_{#}.csv'

https://stackoverflow.com/a/53062251/401226 where the answer has comments about installing the correct version of parallel (in ubuntu use the specific parallel package, which is more recent than what is bundled in moreutils)

This question was asked many years ago, but for future readers I'd like to mention that the most convenient tool for this purpose is xsv from https://github.com/BurntSushi/xsv

The split sub-command is meant to do exactly what has been asked in the original question. The documentation says:

split - Split one CSV file into many CSV files of N chunks

Each of the split chunks retains the header row.

Related