access elements in list of lists

Viewed 499

I am new at text mining, I am using Python. I have a list of lists, each list contains clusters of synonyms and each word in the cluster has a list which contains the number of sentences which it appears. The list i have is like this

syn_cluster = [[['Jack', [1]]], [['small', [1, 2]], ['modest', [1, 3]], ['little', [2]]], [['big', [1]], ['large', [2]]]]

I want to assign for each cluster the min and the max from the appearance list, so i want the result to be like this

[[['Jack', [1]]], [['small', 'modest, 'little'], [1, 3]], [['big', large], [1, 2]]]
2 Answers

I am not sure if you are using the best data structure for your problem. But if you do this with lists of lists you could do:

from itertools import chain

serialized = []
for syn_words in syn_cluster:
    words = [w[0] for w in syn_words]
    freqs = list(chain.from_iterable([f[1] for f in syn_words]))
    min_max = [min(freqs), max(freqs)]
    # or:  min_max = list({min(freqs), max(freqs)}) if you want [1] instead of [1, 1]
    serialized.append([words, min_max])

serialized
>>> [[['Jack'], [1, 1]],
    [['small', 'modest', 'little'], [1, 3]],
    [['big', 'large'], [1, 2]]]

I want to propose another solution.

from functools import reduce

response = []
for sublist in syn_cluster:
  list_flatt = reduce(lambda x,y: x + y, sublist, [])
  list_words  = [word for word in list_flatt if type(word) == str]
  numbers = reduce(lambda x, y: x + y, [number_list for number_list in list_flatt if type(number_list) == list], [])
  list_min_max = [min(numbers), max(numbers)] if len(set(numbers)) > 1 else list(set(numbers))
  response.append([list_words, list_min_max])


print(response)

Output:

[[['Jack', [1]]], [['small', 'modest', 'little'], [1, 3]], [['big', 'large'], [1, 2]]]

Explanation

You can use reduce function from functools in order to flat each sub-list in your syn_cluster list.

# list_flatt become for example something like ['small', [1, 2], 'modest', [1, 3], 'little', [2]] and so on for each row of syn_cluster.

list_flatt = reduce(lambda x, y: x + y, sublist, []) 

When you have the sub-list flatted, you can use comprehension list in order to get the string elements.

list_words  = [word for word in list_flatt if type(word) == str]

And then you can use the same logic to get the list of numbers.

[number_list for number_list in list_flatt if type(number_list) == list]

However this list has the following way: [[1,2] [1,3,2], [1]], for that reason I have used reduce again:

numbers = reduce(lambda x, y: x + y, [number_list for number_list in list_flatt if type(number_list) == list], [])

After that, We make the list with the min and max values and We validate the length of the numbers list.

list_min_max = [min(numbers), max(numbers)] if len(set(numbers)) > 1 else list(set(numbers))
Related