find the largest number in a text file and write its line

Viewed 430

I have a txt file which looks like this: (the first line is the information about the columns and I have 150+ lines)

E1,   E2,  E3,  E4,  E5,    E6,       E7,     E8
abc, cba, dfa, gds, 60371, 42.1234, -2.12,    hkfka
grs, fx,  hgf, eff, 30331, 124,     31313.23,  gj
.
.
.

Expected output:

abc, cba, dfa, gds, 60371, 42.1234, -2.12, hkfka

I read this file with with open method, then I want to find the largest number in the E5's column. It's '60371' in this example. After I found it, I'd like to write it's whole line to a text file. I can find the largest number by adding the string to a list, but can't write it's line with this method.

    list = []
    with open(file.csv, "r") as m:
        text = m.readlines()[1:]
        text = [line.replace(' ', '') for line in text]
        for line in text:
            currentline = line.split(",")
            number= currentline[4]
            list.append(number)
            largest = max(list)

Edit: I'm not allowed to use any imports, such as panda, etc.

4 Answers

1. Without libs

with open('data.csv', 'r') as f:
    next(f)
    lines = [line.replace(' ', '').split(',') for line in f.readlines()]
    numbers = [int(line[4]) for line in lines]
    index = numbers.index(max(numbers))

with open('result.csv', 'w') as f:
    f.write(f'{index} ({lines[index][1]})')

Output:

0 (cba)

2. With pandas

import pandas as pd

df = pd.read_csv('data.csv')
row = df[df['  E5'] == df['  E5'].max()]
row.to_csv('result.csv')

Output file:

,E1,   E2,  E3,  E4,  E5,    E6,       E7,     E8
0,abc, cba, dfa, gds,60371,42.1234,-2.12,    hkfka

You can also store the text of the line that contains the largest number as an object, in a similar way to how you store the numbers themselves. Unless you need to extract all the lines and all the numbers, I would store only the latest largest number and its corresponding line text. For example:

max_number = 0
max_line = ''
with open(file.csv, "r") as m:
    text = m.readlines()[1:]
    text = [line.replace(' ', '') for line in text]
    for line in text:
        currentline = line.split(",")
        number = int(currentline[4])
        if number > max_number: 
             max_number = number 
             max_line = line

This way max_number will be the largest number found in these loops, and max_line will be its corresponding line, but no other data from the file will be stored.

You can store the line on the loop.

list = []
cur_n = 0
max_n = 0
max_line = 0
i = 0
with open("file.csv", "r") as m:
    text = m.readlines()[1:]
    text = [line.replace(' ', '') for line in text]
    for i in range(0, len(text)):
        currentline = text[i].split(",")
        cur_n = int(currentline[4])
        if cur_n > max_n :
            max_n = cur_n
            max_line = i
print(max_n)
print(text[max_line])

Here's a simple version, taking benefit of builtin CSV module. This is mainly useful if CSV format changes by fields being inserted, moved, deleted etc. So we always use column id "E5":

import csv

with open("test.txt", "r") as csvfile:
    reader = csv.DictReader(csvfile, skipinitialspace=True)
    cur_max = -float("inf")

    for row in reader:
        val = float(row["E5"])
        if val > cur_max:
            cur_max = val
            max_line = ", ".join(row.values())

print(max_line)  # abc, cba, dfa, gds, 60371, 42.1234, -2.12, hkfka

It's also single pass, which might not be at all necessary.

I undestood first that by any imports you meant external imports, but any import is forbidden, then other answers could be used. Note that you want to cast values with float though, as input file seems to contain them, and int naturally discards any floating point values. If E5 is always int, then of course int is ok

Related