I am using emeditor and I am trying to isolate about 2 millions articles containing keyword3 from a french wikipedia dump .xml file (20GB, 338 millions rows, 4.8 millions articles in total). I would like to keep the text contained between 2 keywords (keyword1 and keyword2) but only if another keyword (keyword3) exists inside them.
List of keywords :
keyword1 = <page>
keyword2 = </page>
keyword3 = {{Infobox
Example A:
keyword1 = <page>
text to consider without keyword3
keyword2 = </page>
Result => do not extract (or keep or split) this part.
Example B:
keyword1 = <page>
text to consider with keyword3
keyword2 = </page>
Result => extract (or keep or split) this part.
The author of Emeditor helped me with the following :
Find (choose regular expression):
<page>(.*?{{Infobox.*?)</page>
Replace with
\1
And in Advanced... : search in 2500 lines
It seems to work overall fine but from time to time some errors are appearing : I am joining some tiny samples here : https://www.cjoint.com/c/JErsTJnVQpD I also added a small desired results xml file. As you can see in the joined image, the highlighted part in blue color (2 articles) should not have been included in the result part as they don't have the keyword {{Infobox . Note: It also would be nice if the tag is keep in the results. Thanks in advance ;)

