I want to know if there is any way to do a classification based on different input types.
Basically, I have the image dog vs cat dataset from Kaggle and also the sound dog vs cat dataset and I want to show that by combining models from audio alone and image alone we could get a better accuracy.
I read things on ensemble learning where we can combine different models with average precision to get a higher accuracy but these are doing the classification on same types of inputs whereas here I want to do classification by using image and audio as inputs. I also read things from Keras using mixed input but this work only for combining tabular data with images but not sound with images...
I did not find any labelled dataset of dog vs cat videos from which I could extract frames and audio and then apply a CNN on both to do the classification. Do you have any idea on how to tackle this problem ?