I'm not quite sure how to implement the "Bag of Words" approach with HOG descriptors. I've checked several sources which usually provide several steps to follow:
- Compute the HOGs for the set of valid training images.
- Apply an clustering algorithm to retrieve n centroids from the descriptors.
- Perform some magic to create histograms with the frequency of the nearest centroids of the computed HOGs or use OpenCVs implementation to do this.
- Train a linear SVM with the histograms
The step which involves magic (3) is not really clear. If I don't use OpenCV, how would I implement it?
The HOGs are vectors which are calculated cell-wise. So I have a vector for each cell. I could iterate over the vector and calculate the closest centroid for each element of the vector and create the histogram accordingly. Would this be a proper way to do it? But if so, I still have vectors of different sizes and no benefit from it.