FLX_BACKEND_REDIS_004 Alert

Description

A critical alert is raised when the number of connected Redis slaves is less than 2 for 5 minutes.

Severity

This alert is flagged as Critical.

Customer Impact

This alert is triggered when there are 1 or 2 failed slaves in a Redis cluster. In a flex environment, Redis runs on three node cluster where it has one MASTER (leader) and two SLAVE.

Redis cluster with two failed slave nodes means that we have no failover in case of an incident on master Redis. A failed Redis cluster can cause issues with other dependent application services. It's crucial to have a healthy Redis cluster.

Operational Remediation Process

Login to the services nodes where the failed Redis slave is reported.

To bring up the failed Redis service on the failed node it requires the usual service start steps.

$ ssh SERVER

Note the health state of the affected Redis service.

$ docker ps | grep redis

Get the logs of the services.

$ docker logs <container_name>

You can give try to restart the Redis service and wait for it to become healthy. Check the logs again and verify if any replication error occurs.

$ cd /flex
$ docker compose stop <service name>
$ docker compose start <service name>

Check the status of the Redis health check in xymon.

To check the Redis cluster from the Redis console.

Command to login to Redis console.

$ docker compose exec redis redis-cli -h <server_ip> -p 26379

once logged in to the Redis console, first check the status of MASTER Redis.

redis> SENTINEL master flex-redis-master

you should be able to see the number of connected slaves from the above command.

		eg:
		31) "num-slaves"
		32) "2"

Then double-check the status of SLAVE by running the below command.

redis> SENTINEL slaves flex-redis-master