I currently have a web api run using flask, gunicorn, and nginx, the web api calls my tensorflow model serving port. When the web api is called multiple times before the first is finished the tensorflow model fails gives an empty request.
What is the best way to handle this? My web api is currently already behind gunicorn and nginx, but the calls from the api to the tensorflow model seems to be the problem. Should I put it behind gunicorn / nginx load balancer as well?
Thanks