FLX_BACKEND_MONGODB_005 Alert

Description

A critical Alert is raised when mongodb cluster leadership changes have been detected for 5 minutes.

Severity

This alert is flagged as Warning.

Customer Impact

This alert is triggered when the mongodb cluster detects a change in cluster leadership. The cluster leadership can change if one of the PRIMARY mongodb nodes fails and the leadership is passed to one of the secondary. This change in leadership triggers the alarm.

We should also check the status of the cluster when this alert is raised.

Operational Remediation Process

Login to any one of the services nodes which are hosting the mongodb.

MongoDB runs on three node clusters where it has one PRIMARY (leader) and two SECONDARY(slave).

Command to login to mongodb console.

$ docker compose exec mongo mongo \
      --username root --password <mongodb_root_password> \
      --authenticationDatabase admin

once logged in to mongodb console, first, check the cluster status by running the below command

mongodb> rs.status()

This command should show the PRIMARY and SECONDARY node information. It should also show the duration of how long a node has been assigned with the role of PRIMARY.

If you happen to find out a missing SECONDARY node in the rs.status() output then it is required to bring up the failed mongodb service on the node.

To bring up the failed mongodb service on the failed node, requires the usual service start steps. Login to the failed mongodb service node.

$ ssh SERVER

Note the health state of the affected mongodb service.

$ docker ps | grep mongo

Get the logs of the services.

$ docker logs <container_name>

You can give try to restart the mongodb service and wait for it to become healthy. Check the logs again and verify if any replication error occurs.

$ docker compose stop <service name>
$ docker compose start <service name>