Splitting a special type of list data and saving data into two separate dataframe using condition in python

Viewed 78

Want to seperate a list data into two parts based on condition. If the value is less than "H1000", we want in a first dataframe(Output for list 1) and if it is greater or equal to "H1000" we want in a second dataframe(Output for list2). First column starts the value with H followed by a four numbers.

Here is my python code:

with open(fn) as f:
  text = f.read().strip()
   print(text)
   lines = [[(Path(fn.name), line_no + 1, col_no + 1, cell) for col_no, cell in enumerate(
                re.split('\t', l.strip())) if cell != ''] for line_no, l in enumerate(re.split(r'[\r\n]+', text))]
    print(lines)
    if (lines[:][:][3] == "H1000"):
        list1
        list2 

I am not able to write a python logic to divide the list data into two parts.

Attach python code & file here

3 Answers

So basically you want to check if the number after the H is greater or not than 1000 right? If I'm right then just do like this:

with open(fn) as f:
  text = f.read().strip()
   print(text)
   lines = [[(Path(fn.name), line_no + 1, col_no + 1, cell) for col_no, cell in enumerate(
                re.split('\t', l.strip())) if cell != ''] for line_no, l in enumerate(re.split(r'[\r\n]+', text))]
    print(lines)
    value = lines[:][:][3]
    if value[1:].isdigit():
        if (int(value[1:]) < 1000):
            #list 1
        else:
            #list 2

we simply take the numerical part of the factor "hxxxx" with the slices, convert it into an integer and compare it with 1000

with open(fn) as f:
    text = f.read().strip()
lines =text.split('\n')
list1=[]
list2=[]
for i in lines:
    if int(i.split(' ')[0].replace("H",""))>=1000:
        list2.append(i)
    else:
        list1.append(i)
print(list1)
print("***************************************")
print(list2)

I'm not sure exactly where the problem lies. Assuming you read the above text file line by line, you can simply make use of str.__le__ to check your condition, e.g.

lines = """
H0002   Version 3
H0003   Date_generated  5-Aug-81
H0004   Reporting_period_end_date   09-Jun-99
H0005   State   WAA
H0999   Tene_no/Combined_rept_no    E79/38975
H1001   Tene_holder Magnetic Resources NL   
""".strip().split("\n")
# Or
# with open(fn) as f: lines = f.readlines()

list_1, list_2 = [], []
for line in lines:
    if line[:6] <= "H1000":
        list_1.append(line)
    else:
        list_2.append(line)


print(list_1, list_2, sep="\n")
# ['H0002   Version 3', 'H0003   Date_generated  5-Aug-81', 'H0004   Reporting_period_end_date   09-Jun-99', 'H0005   State   WAA', 'H0999   Tene_no/Combined_rept_no    E79/38975']
# ['H1001   Tene_holder Magnetic Resources NL']

Related