Regex force length of specific regex

Viewed 547

I'm using R and need a regex for

a block of N characters starting with zero or more whitespaces and continuing with one or more digits afterwards

For N = 9 here are

examples of valid strings

  • 123456789
  • kfasdf 3456789asdf
  • a 1

and examples of invalid strings

  • 12345 789
  • 1 9
  • a 678a
8 Answers

Another option is to match 8 times either a digit OR a space not preceded by a digit and then match a digit at the end.

(?<![\d\h])(?>\d|(?<!\d)\h){8}\d

In parts

  • (?<![\d\h]) Negative lookbehind, assert what is on the left is not a horizontal whitespace char or digit
  • (?> Atomic group (no backtracking)
    • \d Match a digit
    • | Or
    • \h(?<!\d\h) Match a horizontal whitespace char asserting that it is not preceded by a digit
  • ){8} Close the group and repeat 8 times
  • \d Match the last digit

Regex demo | R demo

Example code, using perl=TRUE

x <- "123456789
kfasdf  3456789asdf
a        1

12345 789
1       9
a     678a"
    regmatches(x, gregexpr("(?<![\\d\\h])(?>\\d|(?<!\\d)\\h){8}\\d", x, perl=TRUE))

Output

[[1]]
[1] "123456789" "  3456789" "        1"

If there can not be a digit present after matching the last 9th digit, you could end the pattern with a negative lookahead asserting not a digit.

(?<![\d\h])(?>\d|(?<!\d)\h){8}\d(?!\d)

Regex demo

If there can not be any digits on any side:

 (?<!\d)(?>\d|(?<!\d)\h){8}\d(?!\d)

Regex demo

Using string s from @d.b's answer.

Extract optional whitespace followed by numbers.

library(stringr)
str_extract(s, '(\\s+)?\\d+')
#[1] "123456789" "  3456789" "        1" "12345"     "1"         "     678" 

Check their length using nchar.

nchar(str_extract(s, '(\\s+)?\\d+')) == 9
#[1]  TRUE  TRUE  TRUE FALSE FALSE FALSE

Using the same logic in base R function.

nchar(regmatches(s, regexpr('(\\s+)?\\d+', s))) == 9
#[1]  TRUE  TRUE  TRUE FALSE FALSE FALSE

If there could be multiple such instances we can use str_extract_all :

sapply(str_extract_all(s, '(\\s+)?\\d+'), function(x) any(nchar(x) == 9))

Basic form: space bias
this is a basic form that has no anchors or boundrys

(?:[ ]|\d(?![ ])){8}\d

dem0

feature:

  • block of 9
  • minimum block size of 2
  • match takes maximum spaces vs minimal digits

Basic form: number bias
same basic form that has been modified to get number bias.

(?=((?:[ ]|\d(?![ ])){8}\d(?!\d)|\d{9}))\1

dem1

feature:

  • block of 9
  • minimum block size of 2
  • match takes minimal spaces vs maximum digits

End of line Anchor method (numeric bias) :

(?=[ ]{0,8}?\d{1,9}(.*)$)[ \d]{9}(?=\1$)

dem2

feature:

  • block of 9
  • minimum block size of 2
  • match takes minimal spaces vs maximum digits
  • single capture is not part of match
  • line orientated regex, needs multi-line option if string is more than 1 line

The desired substring contains 9 digits or fewer than 9 digits. In the second case it begins with a space, ends with a digit and each of the 7 characters in between is a space preceded by a space or a digit followed by a digit. We therefore could use the following regular expression.

\d{9}|\s(?:(?<=\s)\s|\d(?=\d)){7}\d

Demo

The regex engine performs the following operations.

\d{9}       : match 9 digits  
|           : or
\s          : match a space
(?:         : begin non-capture group
  (?<=\s)   : next character must be preceded by a space
  \s        : match a space
  |         : or
  \d        : match a digit
  (?=\d)    : next character must be a digit
)           : end non-capture group
{7}         : execute non-capture group 7 times
\d          : match a digit
  1. Add a comma before the spaces
  2. split at the comma
  3. keep only either space or digits
  4. Count number of characters and see if it matches the required size
s = c("123456789", "kfasdf  3456789asdf",
      "a        1", "12345 789", "1       9",
      "a     678a")

sapply(strsplit(gsub("(\\s+)", ",\\1", s), ","), function(x) {
    any(nchar(gsub("[A-Za-z]", "", x)) == 9)
})
#[1]  TRUE  TRUE  TRUE FALSE FALSE FALSE

You may use the regex pattern

[ \d](?:(?<=[ ])[ ]|\d){7}\d

and in R use

str_extract(x, regex('[ \\d](?:(?<=[ ])[ ]|\\d){7}\\d'))

See this demo.


enter image description here

Please note that in the above regex pattern the [ ] may be replaced by a simple space character. Using [ ] is a common practice to increase readability.

That's not an easy task for regexp-s. You really should consider parsing the string yourself. At least partially. Because you need the lengths of capturing groups and regexp-s do not have this feature.

But if you really want to use them, then there's a workaround:
I'll use JS so that the code can be ran right here.

const re = /^(.*)(\s*\d+)(.*)$(?<=\1.{9}\3)/

console.log(re.test("123456789"))
console.log(re.test("kfasdf  3456789asdf"))
console.log(re.test("a        1"))

console.log(re.test("12345 789"))
console.log(re.test("1       9" ))
console.log(re.test("a     678a"))

where

  1. \s*\d+ meets your base condition of zero or more spaces followed by one or more digits
  2. we can't get groups' lengths, but we can get everything before and after the main group. That is what ^(.*) and (.*)$ are for.
  3. Now we need to check that all three groups add up to a full string, for that we use look behind assertion (?<=\1.{9}\3) and we set the desired N for a number of symbols allowed in the main group (9 in this case)

You didn't mention how the regexp should behave in all situations, for example in this one:

"         3456780000000"

with extra spaces and extra digits. So I won't try to guess. But it's easy to fix the regexp I've provided for all your cases.

Update:

I think the Edward's original answer is the best for you (look in the history). But not sure about boundary constraints. They are not clear from your question.

But I'll still leave mine because, while Edward's answer is shortest and fastest for your specific case, mine is more general and better suits the title of the question.

And I added performance tests:

const chars = Array(1000000)
const half_len = chars.length/2
chars.fill("a", 0, half_len)
chars.fill("1", half_len, half_len + 9)
chars.fill("a", half_len + 9)
const str = chars.join("")
function test(name, re) {
  console.log(name)
  console.time(re.toString())
  const res = re.test(str)
  console.timeEnd(re.toString())
  console.log("res",res)
}
test("Edward's original", /((?<!\d)\s|\d){9}(?<=\d)/)
test("Ωmega's"          , /(?=[ \d]{9}(.*$))[ ]*\d+\1$/)
test("Edward's modified", /(?=[ ]{0,8}?\d{1,9}(.*))[ \d]{9}(?=\1$)/)
test("mine"             , /^(.*)(\s*\d+)(.*)$(?<=\1.{9}\3)/)

Surely lookbehinds are not cheap!

If you are looking for a clean regex solution, then you should use the following pattern:

(?=[ \\d]{9}(.*$))[ ]*\\d+\\1$

...where you combine a positive lookahead with a regular matching that includes a match from the lookahead.

The R syntax is then

str_extract(x, regex('(?=[ \\d]{9}(.*$))[ ]*\\d+\\1$'))

and you can test this code here.


If your desire is also to catch a matching N-character long substring, then use

str_match(x, regex('(?=[ \\d]{9}(.*$))([ ]*\\d+)\\1$')) [,3]

as shown in this demo.

Related