How to extract all the characters including linefeeds(\n) from a given string

Viewed 43

I have a string in the below format : Please note : \n means linefeed

\n\nThe following table provides the details of intangible assets\nacquired, by major class and weighted average useful life:\n\n \n\n(USS in millions) USEFUL LIFE\nCustomer relationships 15 years $265\nIntellectual property 10 years 120\nTrade names 15 years 51\nFavorable leases 38 years 26\nOther various 2\nTotal intangible assets $464\n\nThe fair value in the opening balance sheet of the 30%\nredeemable noncontrolling interest in Loders was estimated to\nbe $450 million.

I have to extract all the characters between \n\n \n\n and \n\n

Expected output :

(USS in millions) USEFUL LIFE\nCustomer relationships 15 years $265\nIntellectual property 10 years 120\nTrade names 15 years 51\nFavorable leases 38 years 26\nOther various 2\nTotal intangible assets $464

I have written a logic as below :

re.findall(r'(\n\n\s\n\n)(.|\n)*(\n\n)', result)

But above code is not giving me desired result. Can somebody help, please.

1 Answers

You could first match the double newline (or match an optional carriage return and newline) and then capture all lines in group 1 that end with a newline and don't start with a newline.

Using re.findall, you will get back a list with the capturing groups values. The desired result is the second item.

\r?\n\r?\n(.*(?:\r?\n(?!\r?\n).*)*)\r?\n\r?\n

Regex demo | Python demo

import re

s="\n\nThe following table provides the details of intangible assets\nacquired, by major class and weighted average useful life:\n\n \n\n(USS in millions) USEFUL LIFE\nCustomer relationships 15 years $265\nIntellectual property 10 years 120\nTrade names 15 years 51\nFavorable leases 38 years 26\nOther various 2\nTotal intangible assets $464\n\nThe fair value in the opening balance sheet of the 30%\nredeemable noncontrolling interest in Loders was estimated to\nbe $450 million."

regex = r"\r?\n\r?\n(.*(?:\r?\n(?!\r?\n).*)*)\r?\n\r?\n"

print(re.findall(regex, s))

Output

[
'The following table provides the details of intangible assets\nacquired, by major class and weighted average useful life:', 
'(USS in millions) USEFUL LIFE\nCustomer relationships 15 years $265\nIntellectual property 10 years 120\nTrade names 15 years 51\nFavorable leases 38 years 26\nOther various 2\nTotal intangible assets $464'
]
Related