How can I split word in the wav file in python?

Viewed 317

e.g wav file is("How are you ?") ı want to split 3 wav file like as ("How"), ("are"), ("you").Could you help me ?

1 Answers

You can try this, but you need to know the timestamp of each word, unless you use machine learning. If you know the time stamps at the end of each word, for example, after how it's 1, after are its 2 and after you its 2.7, try this.

from pydub import AudioSegment

stamps = [1,2,2.7] #timestamps at end of each word

originalAudio = AudioSegment.from_wav("audiofile.wav")
start = 0
counter = 1
for current in stamps:
    newAudio = originalAudio[start*1000: current*1000] #in milliseconds 
    newAudio.export(f'{counter}-word.wav', format="wav") 
    counter += 1
    start = current

The files will be saved as 1-word.wav, 2-word.wav etc.

If you don’t know the timings, you can try this code, it listens for silences or pauses.

from pydub import AudioSegment
from pydub.silence import split_on_silence

sound_file = AudioSegment.from_wav("sentence.wav")
audio_chunks = split_on_silence(sound_file, 
    # must be silent for at least half a second
    # make it shorter if the pause is short, like 100-250ms
    min_silence_len=500,

    # consider it silent if quieter than -16 dBFS
    silence_thresh=16
)

for i, chunk in enumerate(audio_chunks):

    out_file = f"chunk{i}.wav"
    print ("exporting", out_file)
    chunk.export(out_file, format="wav")

However, after listening to the audio file, there is a lot of background noise and no clear pauses between words.

You then said that you want to translate the audio to text, try this code.

First install the required libraries.

pip3 install SpeechRecognition pydub

Then run this,

import speech_recognition as sr
r = sr.Recognizer()

audiofile = 'demo55.wav'

with sr.AudioFile(audiofile) as source:
    # listen for the data (load audio to memory)
    audio_data = r.record(source)
    # recognize (convert from speech to text) tr is code for turkish
    text = r.recognize_google(audio_data, language="tr-tr")
    print(text)

I get the output,

"Merhaba benim adım Ezgi"

Not sure if this is correct as I don't speak turkish, but i listened to it and it sounds correct.

Related