I have a simple spring gRPC application with a dummy endpoint that does but returns an empty message:
proto:
message Map {
map<string, string> value = 1;
}
service GreetingService {
rpc GreetMap(Map) returns (google.protobuf.Empty) {}
}
implementation:
@Override
public void greetMap(Map request, StreamObserver<Empty> responseObserver) {
responseObserver.onNext(Empty.getDefaultInstance());
responseObserver.onCompleted();
}
I'm trying to load test this endpoint. I'm using a tool called ghz. I have generated a large proto message (327Kb) and I'm running this load tests on localhost (thus no network latency) for 200 requests with concurrency setting set to 20.
ghz --proto=service.proto -i path/to/imports/ --binary-file=output.bin --call=package.GreetingService/GreetMap --concurrency=20 --total=200 localhost:6565
And I'm the results is confusing:
Summary:
Count: 200
Total: 543.25 ms
Slowest: 91.21 ms
Fastest: 2.41 ms
Average: 25.47 ms
Requests/sec: 368.16
Response time histogram:
2.412 [1] |∎
11.292 [61] |∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎
20.171 [32] |∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎
29.051 [26] |∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎
37.931 [39] |∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎
46.810 [16] |∎∎∎∎∎∎∎∎∎∎
55.690 [12] |∎∎∎∎∎∎∎∎
64.570 [6] |∎∎∎∎
73.449 [2] |∎
82.329 [4] |∎∎∎
91.209 [1] |∎
Latency distribution:
10 % in 5.78 ms
25 % in 9.93 ms
50 % in 21.44 ms
75 % in 36.80 ms
90 % in 49.21 ms
95 % in 61.35 ms
99 % in 79.48 ms
Status code distribution:
[OK] 200 responses
As we can see, the 95%% latency is more than 50ms and the throughput is 368.16 RPS. If I run the same test with the small payload (6B) the result is completely different:
Summary:
Count: 200
Total: 30.96 ms
Slowest: 14.86 ms
Fastest: 0.35 ms
Average: 2.90 ms
Requests/sec: 6460.92
Response time histogram:
0.354 [1] |
1.805 [155] |∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎∎
3.256 [4] |∎
4.706 [20] |∎∎∎∎∎
6.157 [0] |
7.608 [0] |
9.059 [0] |
10.509 [0] |
11.960 [0] |
13.411 [0] |
14.862 [20] |∎∎∎∎∎
Latency distribution:
10 % in 0.96 ms
25 % in 1.25 ms
50 % in 1.49 ms
75 % in 1.70 ms
90 % in 13.43 ms
95 % in 14.51 ms
99 % in 14.81 ms
Status code distribution:
[OK] 200 responses
The throughput is about 20 times better (6460.92 RPS) and latency on 99%% is about 15ms.
The only thing that I could find is related to gRPC and payload size is about gRPC-Go and the latency with no network latency for the message of 1MB size described as within 5ms.
So is the handling on a big payload such an issue for gRPC or is it possible to fine-tune it for a better performance?
UPD
So I've tried to profile the application during the load tests and the results aren't very optimistic. Most of the CPU time is taken by message parsing. Of course, I'm not an expert in this, but it doesn't feel like something that can be optimized. Would be happy to be wrong
