I have a text file that contains the following:
'\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c'
As you see it's all form feed characters, which I want to have removed.
I have tried various solutions but for some reason they don't seem to work.
For example, what I have tried is to remove the left '\x0c, the right \x0c', and all the other \x0c, but the output remains the same.
The code is use:
import re
import string
with open('AF-40-A-00020539.txt', "r", encoding="ascii") as input_file:
input_content = input_file.read()
print(
input_content.lstrip('\'\x0c')\
.rstrip('\x0c\'')\
.strip('\x0c')
.replace('\x0c', '')
)
After executing this, I get this as output \x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c\x0c' so not what I would expect.
What is the reason for this? How can I remove the form feed characters?
UPDATE, thanks to joao's answer: \xHH, where HH are two hex digits, is a recognised escape sequence to write ASCII characters using their corresponding hex value, just like \n is for a newline.
The .replace('\x0c', '') did not work because in this string literal \xOc got escaped, whereas in the text file, it was just copied as plain text.