Read data from file in two separate arrays

Viewed 380

I'm new to python and am trying to load data from a file. My file looks like this:

TION 13168375
NTHE 11234972
THER 10218035
THAT 8980536
OFTH 8132597
FTHE 8100836
THES 7717675
WITH 7627991

I want to extract both columns into separate arrays. What I tried so far is:

import numpy as np
s=open("Equadgrams.txt", "r")
data = np.genfromtxt(s, dtype=[('mystring','S4'),('myint','i8')])

In the documentation of the loadtxt command it looked like I can split the result into separate arrays, but this gave me errors.

x,y = np.loadtxt(s, dtype=[('mystring','S4'),('myint','i8')])

Traceback (most recent call last):
File "<input>", line 3, in <module>
ValueError: too many values to unpack (expected 2)

One additional thing that I noticed: The integers in the data array seem to be alright, but the strings are not read as intended. I get for the first entry: b'TION' which is of type class<class 'numpy.bytes_'>

I hope that someone is so kind to help me with my problem.

3 Answers

If you want two different arrays you can use the following:


s = open("filename.txt", "r")
lines = s.readlines()

strings = [line.split[' '][0] for line in lines]
ints = [line.split[' '][1] for line in lines]

s.close()

I hope that's the correct form you want it in. Otherwise you have to convert it afterwards.

Maybe you should use pandas, it's a well know python library to deal with tabular data like the one you have. So first install pandas using pip as follows:

pip install pandas

Then load your file like this:

import pandas as pd
df = pd.read_csv("Equadgrams.txt",sep=" ")

You could iterate through every line in the file and append the values to x and y using the .split method:

x = []
y = []
file = open("file.txt","r")
for line in file:
    x.append(line.split(" ")[0])
    y.append(line.split(" ")[1])
Related