What are effective preprocessing methods for reducing data set size (e.g., removing records) without losing information for machine learning problems?

Viewed 4221

I work with a fair number of data sets that have many records -- often in the millions of records. It seems to me that not all of these records are equally useful for building an effective model of the data, e.g., because there are duplicates in the data set. These data sets could be much easier and faster to analyze if they were reduced to a better set of records.

What preprocessing methods are there for reducing data set size (e.g., removing records) without losing information for machine learning problems?

I know one simple transformation is to summarize duplicate records and weight them accordingly, but is there anything more advanced than that?

5 Answers
Related