What software algorithm will select a diverse subset of items from a large set

Viewed 202

Setup

Suppose I have a very large number of items. Each item has a shape, size and colour. They may be

  • triangles, circles or squares
  • red, green or blue
  • small or large

I can't make any assumptions about the distribution of these attributes among the items. I'm reasonably sure that it's not one million large, red triangles but that's always a possibility.

Problem

I want to pick 36 of my shapes with as "diverse" as possible representation across all attribute classes. To clarify, with 36 items drawn from the very large set I'd ideally like 12 red, 12 green, 12 blue, 12 triangles, 18 small etc.

Now there are 18 possible distinct item types (3 colours * 3 shapes * 2 sizes) so one way of doing this would be to include two of each distinct type (assuming that I have them).

If I don't have sufficient of each distinct type, another (impractical, brute force) approach, would be to iterate over every possible subset of 36 items and keep the best subset.

I'm sure that this is a specific instance of a broader class of problems solvable by a well known algorithm but I can't determine the magic words for Google. I've tagged as knapsack-problem because it feels like perhaps it's this but I wonder if there's a better way to solve this?

Can you help with either a solution or at least appropriate search terms?

1 Answers

Look at reservoir sampling. Make one reservoir for each shape/color/size combination (so 36 reservoirs), each reservoir having capacity 36. Do one pass over all your elements, and for each element select its appropriate reservoir and perform the reservoir sampling step.

This reduces your problem to at most 36*36 = 1296 elements, fairly sampled from all of them, and covers even the worst case where there is only a single combination.

Then you can simply shuffle the reservoirs, pick a random element from each (skipping empty reservoirs), removing them from the reservoirs. If you had one of each shape/color/size, you are immediately done. If not, you shuffle the reservoirs again and do another pass, and keep doing this until you have selected 36 elements. This gives you a uniform sample over your dataset, normalized by shape/color/size biases.

Related