I've written the following regular expression:
(?:.*[sS]trony [Ll]okalne(?: GW -)? )(.+)(?: nr.+)
It aims at matching what is between strony lokalne GW - and its variants (left site of string) and nr (right site of string). In my test cases, this regex is valid. I use $1 to get my capture group. Please have a look at what it captures (between ** and **).
strony lokalne GW - **Częstochowa** nr 28
[ DLO SZ ] - Strony Lokalne **Szczecin** nr
strony lokalne GW - **Olsztyn** nr 111
[ DLO KI ] - strony lokalne GW - **Kielce** nr 270,
strony lokalne GW - **Łódź** nr 17,
[ DLO SZ ] - Strony Lokalne **Szczecin** nr 72,
strony lokalne GW - **Warszawa** nr 125,
[ DLO KR ] - strony lokalne GW - **Kraków** nr 5,
[ DLO WA ] - Strony Lokalne **Warszawa** nr 152,
strony lokalne GW - **Zielona G?a** nr 128,
strony lokalne GW - **Łódź** nr 63,
I have written another regex to capture (similar group) as I wasn't able to do so in one go (i.e. using one regex). Here's my second regex:
This time matching is really simple: I need what comes after GW and is before nr.
Some examples are here below:
GW **Szczecin** nr 50\n"
GW **TORUŃ** nr 96, wydanie z dnia 23/04/2004WYDARZENIA, str. 3\n"
GW **Lublin** nr 33, wydanie z dnia 08/02/2006WYDARZENIA , str. 3\n"
GW **Wrocław** nr 45, wydanie z dnia 23/02/2004WYDARZENIA, str. 3\n"
How do I merge those two regexes?
To get my capture group I use in Java: matcher.replaceAll("$1"), where matcher is matcher object from regex pattern.