I have an algorithmic question for you. No need for an application context, I'll give you a direct example.
Here is a possible input: input = [ 1, 1, 1, 2, 2, 2, 2, 3, 3, 4, 4, 4 ]. Let's assume a batch size of 5. The idea here is to output lists of maximum size 5 without having separate number, in short: 2 identical numbers cannot be in separate sub-lists. Example output: [ [1, 1, 1], [2, 2, 2, 2], [3, 3, 4, 4, 4] ]
Assumptions: Numbers always sorted, batch_size always larger than possible number of numbers
Do you have a more elegant solution than the one I just found?
i = 0
batch_size = 5
res = []
while i < len(input):
# Retrieve the data list according to the batch size
data = input[i: i + size]
# Increment the index
i += size
# See what's the next output looks like
future_data = input[i: i + size]
if future_data and future_data[0] == data[-1]:
# So we count how many times this number appears in our current list
# and subtract that from our index
cp = data.count(data[-1])
i -= cp
# Then remove from the current list all occurrence of that number
data = data[:-cp]
res.append(data)
Edit: according to @juanpa.arrivillaga's answer:
Thank you all for your reactivity and your answers.
I continue on episode 2, I gave you here my simplified problem and I thought your solution would be sufficient, despite your response, I do not see how to adapt the solution of @juanpa.arrivillaga to my data format, in fact the input would look more like :
input = {
'data_1' : {
'id': [1, 1, 1, 2, 2, 2, 2, 3, 3, 4, 4, 4],
'char': ['A', 'B', 'C', 'D', 'E', 'F', 'G', 'H', 'I', 'J', 'K', 'L']
}
}
!the size of the lists in value of 'id' and 'char' are necessarily equal!
the output must look like:
[
[1, 'A', 1, 'B', 1, 'C'],
[2, 'D', 2, 'E', 2, 'F', 2, 'G'],
[3, 'H', 3, 'I', 4, 'J', 4, 'K', 4, 'L']
]
I am aware that the data structure is not optimal, unfortunately I don't have the hand on it and is therefore unchangeable...
Still the same constrains as before (batch size is working only on the id, am I clear enough ?)