LDA: Coherence Values using u_mass v c_v

Viewed 517

I am currently attempting to record and graph coherence scores for various topic number values in order to determine the number of topics that would be best for my corpus. After several trials using u_mass, the data proved to be inconclusive since the scores don't plateau around a specific topic number. I'm aware that CV ranges from -14 to 14 when using u_mass, however my values range from -2 to -1 and selecting an accurate topic number is not possible. Due to these issues, I attempted to use c_v instead of u_mass but I receive the following error:

    An attempt has been made to start a new process before the
    current process has finished its bootstrapping phase.

    This probably means that you are not using fork to start your
    child processes and you have forgotten to use the proper idiom
    in the main module:

This is my code for computing the coherence value

cm = CoherenceModel(model=ldamodel, texts=texts, dictionary=dictionary,coherence='c_v')
      print("THIS IS THE COHERENCE VALUE ")
      coherence = cm.get_coherence()
      print(coherence)

If anyone could provide assistance in resolving my issues for either c_v or u_mass, it would be greatly appreciated! Thank you!

0 Answers
Related