preg_match_all not completing the string match

Viewed 25

I'm wanting to match all instances of needle in a haystack. My needle is:

/([0-9]{6,}) ([0-9]{10,}) ([0-9]{4,5}) (.{5,}) ([0-9]{1,}) ([0-9\.,]{4,8}) ([0-9\.,]{4,8})/gU

And my haystack is:

6292181 5702016627428 2304 WIDGET 18 14.12 254.16 6102211 5702015357180 10696 WIDGET 16 32.34 517.44 6332205 5702016911053 10946 WIDGET 6 32.36 194.16 

The problem I'm having is that the decimal place at the end of each match is not being included. So instead of matching

 6292181 5702016627428 2304 WIDGET 18 14.12 254.16 
 6102211 5702015357180 10696 WIDGET 16 32.34 517.44
 6332205 5702016911053 10946 WIDGET 6 32.36 194.16 

it's matching as

 6292181 5702016627428 2304 WIDGET 18 14.12 254 
 6102211 5702015357180 10696 WIDGET 16 32.34 517
 6332205 5702016911053 10946 WIDGET 6 32.36 194 

It seems to me there is an issue with the gU parameters.

Here's my workings: https://regex101.com/r/OgKqOV/1

2 Answers

Adding + to the last part will help. It will convert the pattern from lazy to posessive.

So the regex becomes:

([0-9]{6,}) ([0-9]{10,}) ([0-9]{4,5}) (.{5,}) ([0-9]{1,}) ([0-9\.,]{4,8}) ([0-9\.,]{4,8}+)

Working example (same as yours):

https://regex101.com/r/h5gbjL/1

The U modifier transforms all greedy quantifiers to lazy quantifiers and all lazy quantifiers to greedy quantifiers. It's only "useful" if you want to make the pattern shorter when this one contains more lazy quantifiers than greedy quantifiers since you have less question marks to type. (IMO, it's never useful).

Most of the time people uses it as a magic wand, hopping it will solve all their problems instead of taking 10 minutes to understand how quantifiers work.

So, all you have to do to solve your problem is to remove this useless U and to make this quantifier lazy with a question mark: (.{5,}?). All other quantifiers have to be greedy since they are stopped with the following space (that isn't in the character class [0-9]), so you don't need to change them.

([0-9]{6,}) ([0-9]{10,}) ([0-9]{4,5}) (.{5,}?) ([0-9]+) ([0-9.,]{4,8}) ([0-9.,]{4,8})

demo

You can make your pattern shorter using \d instead of [0-9] or 0-9 inside a character class.

(\d{6,}) (\d{10,}) (\d{4,5}) (.{5,}?) (\d+) ([\d.,]{4,8}) ([\d.,]{4,8})

If you want to be sure there's no contigous digit before and after the match, you can add word-boundaries at the beginning and at the end of the pattern:

\b(\d{6,}) (\d{10,}) (\d{4,5}) (.{5,}?) (\d+) ([\d.,]{4,8}) ([\d.,]{4,8})\b
Related