FLX_BACKEND_HAZELCAST_005 Alert
A critical Alert is raised when the node has seen less than 1 or 2 members of the flex-enterprise Hazelcast instance for five minutes."
With unstable hazelcast cluster members, the functionality of flex will be affected. The flex core services like master and job services depend on the hazelcast cluster. Failed hazelcast cluster will affect the working of master and job services.
Log in to the servers which are hosting the hazelcast services. Usually hazelcast run on the service node.
The hazelcast cluster comprises 3 hazelcast members.
$ ssh SERVER
Note the health state of the affected service.
$ docker ps | grep <service name>
Get the logs of the services.
$ docker logs <container_name>
You can give try to restart one service at a time and wait for it to become healthy. Check the logs again and verify the error repeats.
$ docker compose stop <service name>
$ docker compose start <service name>
If the error repeats even after restart, The issue should be escalated to the customer support team.
You can check the logs of the hazelcast cluster
$ docker logs <container_name>