The Problem:
I am using an API that retrieves the content of interest in the form of a bytes object.
The bytes object (myobj) has a value of:
myobj = b'\xd0\xcf\x11\xe0\xa1\xb1\x1a\xe1\x00\x00This is \rthe sentence \rI want to \rkeep.\r\r\x03\r\r\x04\r\r\x03\r\r\x04\x017\x00\x06'
The Question:
How do I only keep this: "This is the sentence I want to keep."
What I've Tried:
1: I tried decoding with UTF-8, however the output was the same as the input. I also tried 'ascii', 'utf-16', and 'utf-8'. If I remove the 'ignore' argument, i receive an error: UnicodeDecodeError: 'utf-8' codec can't decode byte 0xd0 in position 0: invalid continuation byte
myobj.decode('utf-8', 'ignore')
2: Tried using the printable function from string which returned almost the same output as the input.
import string
mystr =str(myobj)
print( ''.join(x for x in test2 if x in mystr.printable))
3: I also tried using strip() and replace to remove portions of the string, however, there are too many distinct characters.
Any suggestions would be great.
Thanks!