the output is made of json with the whole text and each speaker segments like so:
"speaker_labels": {
"speakers": 2,
"segments": [
{
"start_time": "0.94",
"speaker_label": "spk_0",
"end_time": "3.065",
"items": [
{
"start_time": "1.01",
"speaker_label": "spk_0",
"end_time": "1.22"
},
.
.
.
and then the each word and its timestamp
"items": [
{
"start_time": "1.01",
"end_time": "1.22",
"alternatives": [{ "confidence": "1.0", "content": "word" }],
"type": "pronunciation"
},
{
"start_time": "1.22",
"end_time": "1.81",
"alternatives": [{ "confidence": "1.0", "content": "word" }],
"type": "pronunciation"
},
{
"alternatives": [{ "confidence": "0.0", "content": "another word" }],
"type": "punctuation"
},
.
.
.
and i need each speaker's words how am i suppose to get that data without doing alot of logic with the start end time of the words and all the start end of the speakers. the result i expect:
spk_0 : words words words
spk_1 : words words more words