I am trying to perform detection of a certain type of sound in audio files. These audio recordings have variable lengths and the type of sound that I want to detect is usually around 1~5 seconds long and I have the labels of the dataset (onset and offset of when events happen).
My initial approach was by treating it as a binary classification problem. Where I compute the mel spectrogram each half a second (for example). I would label that spectrogram with a 0 if there wasn't a event in those 0.5s and labeled it 1 if the other way.
In what way could I fight this? I am trying to change by passing 0.1 instead of 1 (assuming the previous example). Basically labeling the percentage of the the event happening in the image: labels [0~1] instead of {0,1}.
Many thanks.