How to make a pandas dataframe from list data generated

Viewed 122

I have a list of co-authors:

ten_author_pairs = [('creutzig', 'gao'),
 ('creutzig', 'linshaw'),
 ('gao', 'linshaw'),
 ('jing', 'zhang'),
 ('jing', 'liu'),
 ('zhang', 'liu'),
 ('jing', 'xu'),
 ('briant', 'einav'),
 ('chen', 'gao'),
 ('chen', 'jing')]

From here I can generate a list of negative examples - i.e. authors-pairs which are unconnected using the following code:

#generating negative examples - 

from itertools import combinations

elements = list(set([e for l in ten_author_pairs for e in l])) # find all unique elements

complete_list = list(combinations(elements, 2)) # generate all possible combinations

#convert to sets to negate the order

set1 = [set(l) for l in ten_author_pairs]
complete_set = [set(l) for l in complete_list]

# find sets in `complete_set` but not in `set1`
ten_unconnnected = [list(l) for l in complete_set if l not in set1]

print(len(ten_author_pairs))
print(len(ten_unconnnected))

Next, I want to implement a link prediction problem for which I want to obtain a dataframe as follows:

author-pair          jaccard   Resource_Allocation    Adamic_Adar   Preferential cn_soundarajan_hopcroft      within_inter_cluster     link
creutzig-linshaw      0.25       0.25                  0.25          0.25          0.25                         0.25                     1 

I can calculate these and have lists with scores as output using networkx documentation, but I am not able to put it together as a table as shown above.

Like for the positive examples (the list mentioned above), I can generate a dataframe using:

df = pd.DataFrame(list, columns = ['u1','u2])

and then make a graph with:

G = nx.from_pandas_edgelist(df, u1, u2, create_using = nx.Graph())

After which say for jaccard index I can apply:

nx.jaccard_coefficient(G)

Which returns me a list of node pairs with jaccard score.

The 'link' column is generated with the logic - 1 for co-authors and 0 for pairs in the negative example.

But, I need all the respective scores as a table as mentioned.

Can anyone please help me with how to construct the above dataframe.

(The scores mentioned are just for example purpose to indicate the kind of table i need)

1 Answers

Oh -- this has been a good two years, but I just stumbled upon this...in case I understood you correctly, building on your basis:

from itertools import combinations
import pandas as pd
import networkx as nx

elements = list(set([e for l in ten_author_pairs for e in l]))
complete_list = list(combinations(elements, 2))
set1 = [set(l) for l in ten_author_pairs]

df = pd.DataFrame(set1, columns=["u1", "u2"])
G = nx.from_pandas_edgelist(df, "u1", "u2", create_using=nx.Graph())

Then defining the list of generators

list_generators = [
    nx.jaccard_coefficient,
    nx.resource_allocation_index,
    nx.adamic_adar_index,
    nx.preferential_attachment,
]

Building the score dataframe:

dfx = pd.DataFrame()
for item_generator in list_generators:
    if dfx.shape[0]:
        dfx = dfx.merge(
            right=get_df_network(generator=item_generator, graph=G),
            left_index=True,
            right_index=True,
        )
    else:
        dfx = get_df_network(generator=item_generator, graph=G)

And finally merging in the link dataframe

df_link = (
    pd.DataFrame(set1, columns=["node_0", "node_1"])
    .set_index(["node_0", "node_1"])
    .assign(link=[1] * len(set1))
)

dfx.merge(df_link, left_index=True, right_index=True, how="outer").fillna(0)

could do the job?

Related