FLX_APP_SPRINGBOOT_001 Alert

Description

This alert is raised when there is a constant increase in the number of errors in a flex service within 5 minutes.

Severity

This alert is flagged as Warning.

Customer Impact

A constant increase in the number of errors in a flex service within 5 minutes is unusual behavior and should be noted as a warning.

Eg: Repeated Errors on the flex service could be caused due to the failure of its dependant flex service.

Eg: Repeated errors on the index services could be caused due to any incompatible metadata change which in turn can damage the search functionality of Flex.

  • best case: Increase in error on non-critical services shouldn't cause any impact on the flex server. Eg: data-aggregator, panel, operational dashboard, global header, etc.
  • worst case: Error on critical services like job, and master will impact flex functionality.

Operational Remediation Process

Operational guideline:

$ ssh SERVER

Note the health state of the affected service.

$ docker ps | grep <service name>

Get the logs of the services.

$ docker logs <container_name>

You can give try to restart one service at a time and wait for it to become healthy. Check the logs again and verify the error repeats.

$ docker compose stop <service name>
$ docker compose start <service name>

If the error repeats even after restart, The issue should be escalated to the customer support team.