FLX_BACKEND_ARANGODB_001 Alert

Description

A critical alert is raised when there are fewer than 3 members in the arangoDB cluster.

Severity

This alert is flagged as Critical.

Customer Impact

This alert is triggered when there is a broken ArangoDB cluster. ArangoDB requires at least 3 Arango coordinator services and 3 Arango server services to run on the services node. If any one of the coordinators or ArangoDB service fails, this alert will be triggered. Failed Arangodb cluster will impact various service which connects to ArangoDB, eg metadata, and secret services. This will overall impact flex functionality.

Operational Remediation Process

Login to the services node which has a failed ArangoDB node.

The ArangoDB cluster consists of 3 members. The cluster should have healthy coordinators and ArangoDB services running on the services node.

$ ssh SERVER

Note the health state of the affected ArangoDB service.

$ docker ps | grep arango

You need to restart the ArangoDB service if the health status is 'unhealthy'

$ docker compose stop <service name>
$ docker compose start <service name>

Check the health status after restarting.

$ docker ps | grep <service name>

Check the status of the service on the Xymon page too.