I have a string of text with numbered paragraphs from '1.' to '221.', however, there are certain paragraphs that do not follow the order and I want to remove them. Here is how the data looks:
text = """1. Shares of Paras Defence and Space Technologies gained 2.85 times.
2. The company, engaged in manufacturing and testing of defence and space engineering products.
"3. Its stock ended at Rs 499 versus issue price of Rs 175 per share.
42. On July 23, Zomato NSE 0.00 % Ltd. listed on the Indian stock exchanges.
43. That was exactly a week after the food-delivery and restaurant discovery platform's initial public offering went live.
4. Paras Defence’s IPO, which closed on September 23, had generated bids worth Rs 38,021 crore.
5. It surpassed the previous record of Salasar Technologies’ IPO.
14. NBFCs are betting big time on the IPO.
6. Paras Defence is one of the few players having an edge in defence deals."""
From the above text, I want to remove the content of the paragraphs which aren't in order ie. '42.', '43.' and '14'.
Output Desired:
relevant_text = '1. Shares of Paras Defence and Space Technologies gained 2.85 times.
2. The company, engaged in manufacturing and testing of defence and space engineering products.
3. Its stock ended at Rs 499 versus issue price of Rs 175 per share.
4. Paras Defence’s IPO, which closed on September 23, had generated bids worth Rs 38,021 crore.
5. It surpassed the previous record of Salasar Technologies’ IPO.
6. Paras Defence is one of the few players having an edge in defence deals.'
I tried to match the pattern but don't know how to proceed forward. Also, I'm not sure if the regex pattern is correct as it matches '1.', '2.' etc. but not '"3.'. Here's what I came up with:
text_sequence = []
pattern = re.compile('(\s|["])[0-9]{1,3}\.\s')
matches = pattern.finditer(text)
for match in matches:
for r in range(1, 999):
if str(r) in match.group():
text_sequence.append(match.span())
text_sequence.append(match.group())
print(text_sequence)
Is there a way to get the desired output?
P.S: The matches I am getting from this code have repeated results.