Efficiently read first qualifier into set

Viewed 30

I have file in which each line contains 3 "fields" separated by pipe-sign. So each line looks like :

AA10SchedTbl|ddmmyy|comxxx.xxx.xxx|
AC01_systSchedTbl|ddmmyy|comxxx.xxx.xxx|
....

I need to create a set containing the first field of each line I use following code which works but I don't think this is the most efficient way, any suggestions ?

with open(filename,"r") as file1:
    refercycles = file1.readlines() 
    reference=set()
    for lines in refercycles:
        line=lines.split('|')
        x=line[0].strip()
        reference.add(x)
1 Answers

Nothing wrong algorithmically. Maybe you could apply some (minor) optimizations:

with open(filename, "r") as file1:
    reference = {line.split("|", 1)[0].strip() for line in file1}

This uses a set comprehension, and limits str.split to a single split per line.

Related