Can tasks sent over a Hazelcast cluster prevent unloading of classes?

Viewed 56

Our Java-based OSS software project features a plugin infrastructure, allowing specific blobs of functionality to be added, removed and updated dynamically, at runtime. Each plugin uses its own ClassLoader (more on that later).

The project allows for more than one server instance to be clustered. The clustering implementation is based on Hazelcast (3.12)

We've been running into a problem that proves hard to diagnose. After a certain plugin is redeployed (unload and loaded again), ClassCastExceptions are thrown that look like this:

java.lang.ClassCastException: org.example.OurPlugin cannot be cast to org.example.OurPlugin

This seems to indicate that two class definitions loaded by different class loaders are being used. When analyzing a memory dump, it is clear that there indeed are two (as opposed to the expected one) class loader for that particular plugin.

We've dealt with a similar issue before. Typically, the cause revolves around some kind of resource (like some kind of event listener registration) not being released properly when the plugin is unloaded, which in the end prevents the plugin ClassLoader from being garbage collected. When the plugin gets loaded again, a new ClassLoader gets created, causing two of them to co-exist, which can introduce problems like ClassCastExceptions. We have searched extensively for this in the current code base, but this time, we've pretty much ruled out this to be the cause of the problems that we're experiencing today.

One thing that caught our eye is that the cluster node that is experiencing this problem is the recipient of Callable instances that are sent to it regularly (every 3 seconds) by other nodes in the cluster. The nodes are using this API to submit the Callable to the local node:

com.hazelcast.core.IExecutorService#submitToMember(java.util.concurrent.Callable<T>, com.hazelcast.core.Member)

The Callable implementation is part of the plugin (the other nodes run the same plugin).

When we shut down the other cluster nodes, the rogue ClassLoader on the local node seems to eventually (probably after a garbage collect) disappear!

Can receiving a Callable through this mechanism be preventing the class on the recipient side from being unloaded? If so, how can this be improved?

0 Answers
Related