how to extract duplicate data in csv file

Viewed 293

I'm working on modeling users opinions on YouTube, so I extracted a huge amounts of data (comments and videos), I have a csv file with 5 columns(channelId, videoId, userId, date of comment and polarity) and nearly 80k rows. Now I need to collect each user's comments separately in csv file. How can I extract the all comments for each userId?? I tried to extract duplicates but it doesn't work. Can anyone help me with a little python script?

4 Answers

Using pandas.

df = pd.read_csv('data.csv')
new_df = df[['userId','comment ']]
new_df.to_csv('user_comment.csv',index=False)

If my understanding is correct, you have the comment column in the csv (you forget to mention it in the list of keys)

import pandas

csv = pandas.read_csv(r'youtube.csv')
print(csv.loc[csv['userId'] == 'h']['comment'])

You could write them into a dictionary, where every entry is a userID.

import csv
users = {}
with open('data.csv') as csv_file:
    csv_reader = csv.reader(csv_file, delimiter=',')
    for row in csv_reader:
        # rows are: channelId, videoId, userId, date of comment, popularity, comment
        user = row[2]
        if user in users.keys():
            user_list = users[user]
        else:
            user_list = []
        # now collect all data you want from the row, e.g.
        user_list.append({"channelId":row[0],"videoId":row[1],"date":row[3],"popularity":row[4], "comment":row[5]})
        # now write it back to the dict
        users[user] = user_list

Now you can e.g get all post dates of a user by:

thisUser = users['thisUserID']
for comment in thisUser:
    print(comment['date'])

For writing a user to a single csv file you can use the DictWriter function:

for userID in users.keys():
    with open(userID+'.csv', 'w', newline='') as csvfile:
        fieldnames = ['channelId', 'videoId', 'date', 'popularity','comment']
        writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
        writer.writeheader()
        for comment in users[userID]:
            writer.writerow(comment)

You can achieve this by this snippet:

import numpy

datas = numpy.array([
    # channelId, videoId, userId, date, popularity, comment
    [0, 0, 1, "02052021", 3044, "blobi"],
    [1, 2, 1, "01052021", 4234, "uygukih"],
    [2, 1, 1, "02062021", 2452, "bla"],
    [0, 0, 2, "09052021", 2345, "arghh"],
    [1, 0, 5, "02042021", 234, "haha"]
])

i_user = 2
i_comment = 5

for user in numpy.unique(datas.T[2]):
    print("_" * 50)
    print("userId {0}".format(user))
    [print("comments {0}: {1}".format(i + 1, comment)) for i, comment in enumerate(datas.T[i_comment][numpy.where(datas.T[i_user] == user)])]

It will return:

__________________________________________________
userId 1
comments 1: blobi
comments 2: uygukih
comments 3: bla
__________________________________________________
userId 2
comments 1: arghh
__________________________________________________
userId 5
comments 1: haha
Related