FLX_BACKEND_REDIS_005 Alert
A critical alert is raised when Redis cluster leadership changes have been detected for 5 minutes.
This alert is triggered when the Redis cluster detects a change in cluster leadership. The cluster leadership can change if the MASTER Redis node fail and the leadership is passed to one of the Slave nodes. This change in leadership triggers the alarm.
We should also check the status of the cluster when this alert is raised. Since the usual suspect is a failed master Redis node.
Login to any one of the services nodes which are hosting the Redis.
Redis runs on three node cluster where it has one MASTER (leader) and two SLAVE.
To check the Redis cluster from the Redis console.
Command to login to Redis console.
$ docker compose exec redis redis-cli -h <server_ip> -p 26379
once logged in to the Redis console, first check the status of MASTER Redis.
redis> SENTINEL master flex-redis-master
you should be able to see the number of connected slaves from the above command.
eg:
31) "num-slaves"
32) "2"
Then double-check the status of SLAVE by running the below command.
redis> SENTINEL slaves flex-redis-master
If you happen to find out a missing SECONDARY node from the above output then it is required to bring up the failed Redis service on the node.
To bring up the failed Redis service on the failed node it requires the usual service start steps.
Log in to the failed Redis service node.
$ ssh SERVER
Note the health state of the affected Redis service.
$ docker ps | grep redis
Get the logs of the services.
$ docker logs <container_name>
You can give try to restart the Redis service and wait for it to become healthy. Check the logs again and verify if any replication error occurs.
$ docker compose stop <service name>
$ docker compose start <service name>
Check the status of the Redis health check in xymon.
Once again use the same Redis console command to check the status of the slaves and master Redis.