I am working on a cough recognition app. I have never done something like this before, and any adviced is greatly appreciated.
My approch is to gather data labeled as cough sounds and no_cough sounds, then to make a spectogram with all those sounds, then I want to make a Convolutional neural network to work in order to classify those spectograms.
Until now, I managed to gather 1000 samples of coughing sounds from the internet. The next step is to edit those samples in order to have just one cough sound per sample because when I found cough sounds on the internet, a simple wav file contains multiple coughs, because of that, my 1000 samples will probably become 5000 or more.
After that, I have to gather sounds for the no_cough label.
From my Computer Vision experience, I am very worried that my app will detect a lot of false positive with this approach. Is there a way to solve this problem? I've never worked with sounds before, and I don't even know if my approach is a good one.
My questions are: Is my approach a good one or there exists a better way of making a sound recognition app? If the answer is yes, how should I gather the no_cough sounds, how many should I have to prevent the false positive alarms? Should I try to gather just sounds that are very similar to cough sounds for better accuracy?
It is very important to have a decent accuracy on the false positive because this app will be always opened.