I am writing a program using python3 (lets say the main.py below is running in python3) and utilizing an existing python extension (written in c, compiled with python2), that I dont have the source code for. Lets say the FrozenModule class is from this extension. It does a lot of things really, but just tried to simplify it here. Basically I can subscribe events to it, and I really just have control to the callback function. Meaning I have the most control to change the flow by updating the callback function.
The problem is I have to reorganize the flow and would like to freeze/pause every time the callback function is called.
The final result I want is:
1:a,i,x
5:b,j,y
8:c,k,z
Which is currently what I am getting here, with the code below. But its not really working the way I want to. I added the set_trace just to check the flow. The reason I get the correct output is because I have the SELECTED values defined already in func2.py and func3.py. In reality I want this to come from func1.py
In short I really want func1.py to run first, grab the specific key from it, then run func2.py, which basically iterate through the data it has until it hits that event/key, then executes the callback. Pauses, and do the same for func3.py, then pauses as well. Then goes back to func1.py and do it all over again.
Currently the way it works (not the code below, but the current program I am rewriting), it runs the entire func1.py first, then the keys are all known, and that is passed to func2.py and func3.py and then we get the values respectively. Finally, its all appended together to generate the entire table like:
1:a,i,x
5:b,j,y
8:c,k,z
The main issue is these jobs are really large and could take a long time. Which could also crash in the middle, where you the user dont get anything back. So the new approach is to generate the final output line by line, or row by row. Instead of performing column per column, then combine everything and give the entire result to the user.
The problem with the code below is that when I run the subprocess (reason I use a subprocess is because the extension needs to be in python2 only), it really runs/iterate through the entire data right away. I added writing to the file just to make sure this is happening. So as soon as you run main.py, without even hitting 'c' to continue the set_trace, f2.txt will already have:
1,i
5,j
8,k
Where it should really have empty initially, then as I hit 'c' in main, it should only write 1 row, pause, and wait for next one. Of course the set_trace and writing to a file is just there for testing.
In short, main goal is to write the output row by row. And two main questions are: 1) How to pause the subprocess everytime the callback is called. 2) Pass the value obtained from func1.py to func2.py and func3.py then continue/unpause the process.
Sorry for the long post. Also, I hope my questions are clear.
The example code are below:
In main.py
import subprocess
class ProcReader():
def __init__(self, python_file):
self.proc = subprocess.Popen(['python2', python_file], stdout=subprocess.PIPE)
def __iter__(self):
return self
def __next__(self):
while True:
line = self.proc.stdout.readline()
if not line:
raise StopIteration
return line
r1 = ProcReader("func1.py")
r2 = ProcReader("func2.py")
r3 = ProcReader("func3.py")
for l1, l2, l3 in zip(r1, r2, r3):
d1 = l1.decode('utf-8').strip().split(",")
d2 = l2.decode('utf-8').strip().split(",")
d3 = l3.decode('utf-8').strip().split(",")
print(f"{d1[0]}:{d1[1]},{d2[1]},{d3[1]}")
import pdb
pdb.set_trace()
In func1.py
from frozenmodule import FrozenModule
def callback(k, v):
with open("f1.txt", 'a') as f:
print(f"{k},{v}")
f.write(f"{k},{v}\n")
fm.match = next(iter_items, None)
# somehow pause here, maybe use yield?
SELECTED = [1, 5, 8]
FAKE_DATA = {1: 'a',
5: 'b',
8: 'c'}
iter_items = iter(SELECTED)
fm = FrozenModule(FAKE_DATA, callback)
fm.match = next(iter_items, None)
fm.run()
In func2.py
from frozenmodule import FrozenModule
def callback(k, v):
with open("f2.txt", 'a') as f:
print(f"{k},{v}")
f.write(f"{k},{v}\n")
# somehow pause here, maybe use yield?
# then get value from func1.py and set to fm.match
fm.match = next(iter_items, None)
SELECTED = [1, 5, 8]
FAKE_DATA = {1: 'i',
5: 'j',
8: 'k'}
iter_items = iter(SELECTED)
fm = FrozenModule(FAKE_DATA, callback)
fm.match = next(iter_items, None)
fm.run()
In func3.py
from frozenmodule import FrozenModule
def callback(k, v):
with open("f3.txt", 'a') as f:
print(f"{k},{v}")
f.write(f"{k},{v}\n")
# somehow pause here, maybe use yield?
# then get value from func1.py and set to fm.match
fm.match = next(iter_items, None)
SELECTED = [1, 5, 8]
FAKE_DATA = {1: 'x',
5: 'y',
8: 'z'}
iter_items = iter(SELECTED)
fm = FrozenModule(FAKE_DATA, callback)
fm.match = next(iter_items, None)
fm.run()
In frozenmodule.py
class FrozenModule():
def __init__(self, fake_data, callback):
self.fake_data = fake_data
self.match = 0
self.callback = callback
def run(self):
for x in range(10):
if x == self.match:
self.callback(x, self.fake_data[x])