FLX_INFRA_CADVISOR_004 Alert
This alert is raised when a node has seen more than two restarts of the flex service in the last two days.
Any flex service which is unstable and restarts often needs to be addressed immediately. A constant restart of service could be caused by the service running out of memory or the service itself having multiple errors causing the restart. Any constant impact on an individual flex service can cause an additional impact on other flex services. Flex functionality will break if there are more restart on a service and the impact can be seen in flex UI.
Depending on the service failure, the flex application will be impacted. For example, failed services like data aggregators won't cause any outages. A failure of the data aggregator service will impact only the data aggregator dashboard and will not harm flex functionality as a whole. However, If a job node sees more than two restarts of job services, then that would impact the flex running jobs and the overall usage of flex.
Login to the server.
$ ssh SERVER
Note the health state of the affected service. Make a note of the running period of the container.
$ docker ps | grep <service name>
To check if the container is getting killed due to out of memory, run
$ dmesg -T
If the container is getting killed due to being out of memory, it is required to allocate enough memory to the service.
To check if there are repeated errors occurring within the container, run
$ docker logs <container_name>
More metrics on memory and CPU usage can be found on the Grafana page.