Find cluster in data using networkx library

Viewed 99

Working with text data I have used 'wiki_ru_wordnet' library to find synonyms and replace all groups of words by one. Later I use networkx to find isolated groups.

def get_dict_synsets(texts):
  wikiwordnet = WikiWordnet()
  G=nx.Graph()

  for text in texts:
    s = set()
    synsets = wikiwordnet.get_synsets(text)
    for synset in synsets:
      for w in synset.get_words():
        if text != w.lemma():
          G.add_edge(text, w.lemma())

  d = {}
  for g in nx.connected_components(G):
    default_lemma = g.pop()
    d.update(dict.fromkeys(g, default_lemma))
    d[default_lemma] = default_lemma
    
  return d, G

Below is example of one group of words described by one word:

enter image description here

It seems for me that I cannot use it because words that are not really synonyms fell into the same group. Good result will be like next picture:

[![enter image description here][2]][2]

How I could correct my result? I have idea - use networx.clustering but don't know how... [2]: https://i.stack.imgur.com/Hi2C8.png

1 Answers

refer to the code below to make the split.

from networkx.algorithms.community.asyn_fluidc import asyn_fluidc
communities = list(asyn_fluidc(G, 3))
communities

Then you can create dict and assign for each node its community id

communities_id = {}
for idx, val in enumerate(communities):
    for k in val:
        communities_id[k] = idx
        
communities_id

You can also refer and to this post - https://statsmaths.github.io/stat289-f18/solutions/tutorial14-networkx.html

Related