FLX_BACKEND_REDIS_006 Alert
This alert is triggered when the Redis slave replication lags. In a flex environment, Redis runs on three node cluster where it has one MASTER (leader) and two SLAVE. A constant replication process happens between slave and master. Slave periodically tries to resync with master db and always be in the same state as master.
Redis cluster with two failed slave nodes means that we have no failover in case of an incident on master Redis. A replication lag on the Redis cluster can cause issues with other dependent application services and can cause an unstable Redis cluster. It's crucial to have a healthy Redis cluster with proper replication.
Login to any one of the service nodes where Redis is running.
Check the Redis cluster from the Redis console.
Command to login to Redis console.
$ docker compose exec redis redis-cli -h <server_ip> -p 26379
once logged in to the Redis console, first check the status of MASTER Redis.
redis> SENTINEL master flex-redis-master
you should be able to see the number of connected slaves from the above command.
eg:
31) "num-slaves"
32) "2"
Then double-check the status of SLAVE by running the below command.
redis> SENTINEL slaves flex-redis-master
To bring up the failed Redis service on the failed node it requires the usual service start steps.
$ ssh SERVER
Note the health state of the affected Redis service.
$ docker ps | grep redis
Get the logs of the services.
$ docker logs <container_name>
You can give try to restart the Redis service and wait for it to become healthy. Check the logs again and verify if any replication error occurs.
$ docker compose stop <service name>
$ docker compose start <service name>
Check the status of the Redis health check in xymon.