I have two regex using the same capturing group, but their performance varied very differently.
My programming language is python3.8.
Regex Description
The two regex are
new\s+String\s*(?P<aux1>\(((?:[^()]++|(?&aux1))*)\))
and
([\w_\.]+(?P<aux1>\((?:[^()]++|(?&aux1))*\))*+)\s*\.\s*((?:remove|contains|retain)(?:All)?)\s*(?&aux1)
The named capturing group (?P<aux1>\((?:[^()]++|(?&aux1))*\)) matches paired ( and ), and the content between them. For example ( abc() ).
The first regex matches string like new String("This is a string.").
The second one matches string like a.contains(b) or a.getCollection().contains( anotherCollection )
How to reproduce
I want to know the performance of capturing groups against large input, so I download an open-source Github project elasticsearch, which contains 13724 Java files.
I wrote a script to run the two regex against those files line by line. However, I found the first regex only used 1.6286 seconds, while the second regex used 50.6662 seconds.
I don't know what makes the second regex so slow, when the two regex are so similar and shares the same capturing group.