gRPC GOAWAY error in an Istio service mesh setup

Viewed 1312

Good day to you!

We're experiencing an issue with our gRPC service and we're hoping that the clever people around here might be able to help us understand what's going on.

Our situation is the following:

We have a client service (written in Java/Scala, using java-grpc 1.18), deployed in Azure, that sends a continuous flow of packet over gRPC to a server service deployed in AWS EKS (written in Java, using java-grpc 1.20 and built on top of netty 4.1.3 ), in a cluster configured with Istio. The client sends ~1K requests per second (it is not using streaming).

The connection is made thru an NginX proxy deployed alongside the client (does some routing. Here all packets are routed toward the server in AWS), then thru an IPSEC tunnel to EKS. In EKS, the connection first goes thru an Istio IngressGateway, then is routed to the pod where is goes thru the Envoy sidecar, then to the server.

On the client's side, we're seeing a constant stream of the following exception:

2020-10-26 18:33:39,748 [scala-execution-context-global-99] ERROR c.n.n.s.IngestionLayerForwardingService$ : Caught exception trying to send CLT message from device 256b2551-adc3-4f70-aa48-6b24d1b2436c to kafka-gateway
io.grpc.StatusRuntimeException: UNAVAILABLE: HTTP/2 error code: NO_ERROR
Received Goaway
    at io.grpc.Status.asRuntimeException(Status.java:530)
    at io.grpc.stub.ClientCalls$UnaryStreamToFuture.onClose(ClientCalls.java:482)
    at io.grpc.PartialForwardingClientCallListener.onClose(PartialForwardingClientCallListener.java:39)
    at io.grpc.ForwardingClientCallListener.onClose(ForwardingClientCallListener.java:23)
    at io.grpc.ForwardingClientCallListener$SimpleForwardingClientCallListener.onClose(ForwardingClientCallListener.java:40)
    at io.grpc.internal.CensusStatsModule$StatsClientInterceptor$1$1.onClose(CensusStatsModule.java:699)
    at io.grpc.PartialForwardingClientCallListener.onClose(PartialForwardingClientCallListener.java:39)
    at io.grpc.ForwardingClientCallListener.onClose(ForwardingClientCallListener.java:23)
    at io.grpc.ForwardingClientCallListener$SimpleForwardingClientCallListener.onClose(ForwardingClientCallListener.java:40)
    at io.grpc.internal.CensusTracingModule$TracingClientInterceptor$1$1.onClose(CensusTracingModule.java:397)
    at io.grpc.internal.ClientCallImpl.closeObserver(ClientCallImpl.java:459)
    at io.grpc.internal.ClientCallImpl.access$300(ClientCallImpl.java:63)
    at io.grpc.internal.ClientCallImpl$ClientStreamListenerImpl.close(ClientCallImpl.java:546)
    at io.grpc.internal.ClientCallImpl$ClientStreamListenerImpl.access$600(ClientCallImpl.java:467)
    at io.grpc.internal.ClientCallImpl$ClientStreamListenerImpl$1StreamClosed.runInContext(ClientCallImpl.java:584)
    at io.grpc.internal.ContextRunnable.run(ContextRunnable.java:37)
    at io.grpc.internal.SerializingExecutor.run(SerializingExecutor.java:123)
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1149)
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:624)
    at java.lang.Thread.run(Thread.java:748)

These errors happen to ~0.1% of our traffic in the AWS region we're currently using from PROD, but but this rate is different (usually lower on other regions, with same or even higher traffic) in other regions.

Our understanding of gRPC and also Istio is still quite patchy, and we are struggling to understand where the problem lies.

It looks like the server is not issuing the "double GOAWAY" properly, but we couldn't find any configuration of the max_connection_age on the server side.

Is that possible that either the ingress gateway or the envoy sidecar issue their own GOAWAY? Do they terminate the gRPC connection or are they fully transparent to gRPC?

Any help or pointer to relevant documentation is very welcome :)

Kindest, Antoine

0 Answers
Related