I have a number of strings look like:
str1="Quantity and price: 120 units;the total amount:12000.00"
str2="Quantity:100, amount:10000.00"
str3="Quantity:100, price: 10000 USD"
str4="Parcel A: Quantity:100, amount:$10000.00,Parcel B: Quantity:90, amount:$9000.00"
strlist=[str1,str2,str3,str4]
I want to match tha amount $12000, $10000, 10000 in the first 3 strings and both $10000 and $9000.00 in the last string. However, in the first string there are both "price" and "amount". I thought by using "|" regex would search from left to right, so I want regex to look "amount" first, if it is not presented then look for "price". I tried the following code:
amount_p = re.compile(r'(?:amount|price):(.*?)(?:USD|\.00)')
for i in strlist:
amount=re.findall(amount_p,i)
print(amount)
[' 120 units;the total amount:$12000']
['10000']
[' 10000 ']
['$10000', '$9000']
Somehow the regex ignored "amount" and only looked for "price" in the first string. Then I tried following:
amount_p = re.compile(r'.*(?:amount|price):(.*?)(?:USD|\.00)')
which gives me
['12000']
['10000']
[' 10000 ']
['$9000']
In this case, regex only matched $9000 in the last string and ignored $10000. So my question is what is the function of .* at the begining and is there anyway to solve my problem? Looking for numbers doesn't work because in my actual data there are many other numbers in one text. Thank you all in advance!!!!