I am trying to do model prediction through tensorflow serving through grpc connection and my main code runs with multiple instances. I send Predict request with the following code
req = predict_pb2.PredictRequest()
req.model_spec.name = model_name
req.model_spec.signature_name = "serving_default"
req.inputs[input_layer].CopyFrom(tf.make_tensor_proto(input_data))
response = stub.Predict(req, timeout)
Where stub is defined once and reused per instance in the following way:
channel = grpc.insecure_channel(
grpc_connection_path, options=options,
)
stub = prediction_service_pb2_grpc.PredictionServiceStub(channel)
So I was trying to make grpc unavailable by removing the unix socket file and after this for one instance running prediction gives the following error:
Error: <_InactiveRpcError of RPC that terminated with: status = StatusCode.UNAVAILABLE details = "failed to connect to all addresses; last error: UNKNOWN: No such file or directory" debug_error_string = "UNKNOWN:Failed to pick subchannel {created_time:"2022-08-25T14:04:54.24900453+00:00", children:[UNKNOWN:failed to connect to all addresses; last error: UNKNOWN: No such file or directory {created_time:"2022-08-25T14:04:54.24900212+00:00", grpc_status:14}]}"
However, for the other instance of my prediction code it gives a normal result. I am trying to test and handle the case where my connection is not available but I am getting mixed results on different instances.
Can you explain why can this happening?
UPDATE
When I run for example 20 instances I have the problem the way I described above but when I run 3 instances I do not. However in the case of 20 instances I have the problem only when I delete the unix socket file(as I described above). When the file is there everything works perfectly.