How to use placeholders in a for loop?

Viewed 991

I have the following code:

df['Variable Name']=df['Variable Name'].replace(' 15 °',' at 15 °', regex=True)
df['Variable Name']=df['Variable Name'].replace(' at at 15 °',' at 15 °', regex=True)
df['Variable Name']=df['Variable Name'].replace(' 0 °',' at 0 °', regex=True)
df['Variable Name']=df['Variable Name'].replace(' at at 0 °',' at 0 °', regex=True)
df['Variable Name']=df['Variable Name'].replace(' 5 °',' at 5 °', regex=True)
df['Variable Name']=df['Variable Name'].replace(' at at 5 °',' at 5 °', regex=True)

And would like to know how to shorten it. I tried a for loop:

for x in range(0,15,5):
    df['Variable Name']=df['Variable Name'].replace(' %s °',' at %s °', x, regex=True)
    df['Variable Name']=df['Variable Name'].replace(' at at %s °',' at %s °', x, regex=True)

But I get the error message:

ValueError: For argument "inplace" expected type bool, received type int.

What's a better way to do it?

Edit: Added snippet

Variable Name                          Condition
Density 15 °C (g/mL)   
Density 0 °C (g/mL)    
Density 5 °C (g/mL)    
Calculated API Gravity  
Viscosity at 15 °C (mPa.s) 
Viscosity at 0 °C (mPa.s)  
Viscosity at 5 °C (mPa.s)  
Surface tension 15 °C (oil - air)  
Interfacial tension 15 °C (oil - water)    
4 Answers

Use capture groups with negative lookbehind:

import pandas as pd

s = pd.Series([' 15 °', ' at 15 °', ' 0 °', ' at 0 °', ' 5 °', ' at 5 °'])
s = s.str.replace('(?<!at)\s+(15|0|5) °', r' at \1 °', regex=True)
print(s)

Output

0     at 15 °
1     at 15 °
2      at 0 °
3      at 0 °
4      at 5 °
5      at 5 °
dtype: object

As the regex=True indicates we are going to replace by using a regular expression, the pattern (?<!at)\s+(15|0|5) ° means match a 15, 0, or 5 that does not has at (as the previous word) before it. The notation (?<!at) is known as a negative lookbehind, something like look at previous characters and see if they do not match something, in this case at. The (15|0|5) is a capture group each capture group has a corresponding index, that you can use in the replacement pattern as in ' at \1 °'. So for example, the pattern will only replace a 15 that is not preceded by at, by at 15.

Not sure why you're using the RegEx flag, but you don't need it. inplace is the third argument and needs a boolean value, but you were putting x as the third argument, which is why it failed. You didn't use % formatting properly.

By complete coincidence, inplace can also be used to simplify the problem. You don't need to reassign in this case. Here's some nice(ish) code:

for x in (0, 5, 15): # You were using range() wrong - this is correct
    df['Variable Name'].replace(' %s °' % x,' at %s °' % x, inplace=True)
    df['Variable Name'].replace(' at at %s °' % x,' at %s °' % x, inplace=True)

Try to avoid % formatting if you can use .format() and f-strings. But this should work.

You can save yourself from a replace with optional capture group:

s = pd.Series(['that at 15 °', 'this 15 °', 
               'that at 5 °', 'this 5 °', 
               'that at 0 °', 'this 0 °'])

for x in [0, 5, 15]:
    s = s.str.replace(f'((?: at)? {x} °)', f' at {x} °')

Output:

0    that at 15 °
1    this at 15 °
2     that at 5 °
3     this at 5 °
4     that at 0 °
5     this at 0 °
dtype: object

In single pass with dict of regular expressions:

In [113]: df = pd.DataFrame(columns=['Variable Name'], data=[' 15 °', ' at at 15 °', ' 0 °', ' at at 0 °', '
     ...:  5 °', ' at at 5 °'])                                                                             

In [114]: df['Variable Name'].replace(regex={r'^\s*(\d+) °': r'at \1 °', r'at at (\d+) °': r' at \1 °'})    
Out[114]: 
0      at 15 °
1      at 15 °
2       at 0 °
3       at 0 °
4       at 5 °
5       at 5 °
Name: Variable Name, dtype: object
Related