I've run into a bit of a constraint with combining my Tornado web server websockets implementation with my Kubernetes deployment (I'm using GCP). I'm running an auto-scaler on my Kubernetes engine that automatically spins up new pods as the load on the server increases - this typically happens for long-running jobs (e.g. training ML models).
The problem is that my pods are keeping their own local state - this can mean that a user connection can exist between client -> pod 1, but the long-running job might be processing on pod 2 and sending updates from there. These updates get lost as the connection does not exist between client and pod 2.
I'm using a relatively boilerplate WS implementation, adding/removing user connections as follows:
class DefaultWebSocket(WebSocketHandler):
user_connections = set()
def check_origin(self, origin):
return True
def open(self):
"""
Establish a websocket connection between client and server
:param self: WebSocket object
"""
DefaultWebSocket.user_connections.add(self)
def on_close(self):
"""
Close websocket connection between client and server
:param self: WebSocket object
"""
DefaultWebSocket.user_connections.remove(self)
Any ideas/thoughts - either high or low-level - on how to solve this problem would be much appreciated! I've looked into stuff like k8scale.io and socket.io, but would prefer a 'native' solution.