I have a slew of audio files that are from radio comms from https://www.liveatc.net/ . My intention is to simulate the communication background noise from these types of audio (think airplanes, AM radio, etc.) and apply it to clean data.
My initial thought was to analyze these noisy audio files and find their standard deviations then apply Gaussian White Noise with the average std deviation of all the files (they are all relatively similar with means close to 0). However, there is likely much more to consider since the clean audio I am looking at is at different volumes, etc. With this current attempt at taking random files and applying to a clean speech file, the white noise massively drowns out the clean speech. How am I able to simulate noise from files and apply it to other data?
I make sure to keep all my files in .wav format with a 16000 khz frame rate. I have attached some pseudocode that I have tried to do such a thing but I am looking for other ideas/things to consider.
import numpy as np
import matplotlib.pyplot as plt
import wave
from pydub import AudioSegment, effects
from Pathlib import Path
# get lists of standard deviations and means for all noisy radio-comms files
std_devs = []
for file in Path(audio_files_dir).iterdir():
audio = AudioSegment.from_file(file)
audio = audio.set_frame_rate(16000)
print(file)
print(audio.frame_rate)
samples = audio.get_array_of_samples() # sample's amplitudes
samples = np.array(samples) # cast to numpy array
std_dev_samples = np.std(samples) # get std_dev of the samples
std_devs.append(std_dev_samples)
means.append(mean_samples)
# mean of all standard deviations, will use this to create white noise
print(sum(std_devs) / len(std_devs))
avg_std_devs = sum(std_devs) / len(std_devs)
# taking a clean audio file and prepping it for adding white noise
audio_signal_pydub = AudioSegment.from_file(audio_file_path)
samples = audio_signal_pydub.get_array_of_samples() # take audio signal and turn it into an array of amplitudes
samples = np.array(samples) # turn it into numpy array
print(samples.shape)
# plotting for pydub
time = np.linspace(0, audio_signal_pydub.duration_seconds, num = len(samples)) # create time for x axis in seconds
plt.figure(1)
plt.title("Plot with pydub")
plt.plot(time, samples)
plt.show()
noise = np.random.normal(0, avg_std_devs, samples.shape) # array of white noise with std dev equal to average of the several ATC files
signal_w_noise = samples + noise
plt.figure(1)
plt.title("Plot signal with gaussian white noise")
plt.plot(time, signal_w_noise)
plt.show()
# turn the signal with noise array back into an audio segment for export
new_audio_segment = audio_signal_pydub._spawn(data = signal_w_noise)
print(new_audio_segment)
new_audio_segment.export(r"new_audio.wav", format = 'wav')