FLX_INFRA_OS_001 Alert
This alert indicates that there is a sudden high CPU usage that is persistent for an hour. Flex applications (or even the whole system) may become extremely slow and start to lag. It can also lead to a system freeze and the system stop responding. When all containers in a particular node(say service node) are not available will cause an application outage.
- best case: The CPU usage can be temporary when the flex is highly utilized and cpu will recover after a few hours.
- worst case: The system might freeze and all the containers inside the server will be unavailable.
Try to log in to the server which is being reported in the alert.
Different commands help monitor the system’s load over different periods. Usually, a smaller number is better, as a higher number indicates an overloaded machine.
The uptime command is also useful for viewing the load average of the system. This command displays the current system time, the uptime of the machine, the number of users currently logged into the system, and the load averages for the last 1, 5, and 15-minute durations.
The top command displays the dynamic statistics of a running Linux system in real-time. Try to find out the process which is most CPU-consuming.
The ps command is a flexible and widely used tool for identifying the processes running in the system and the number of resources they’re using to run.
The first thing to do when the CPU becomes overloaded is to identify any processes of this kind and terminate or restart them.
If it happens to be a container process that is consuming the CPU, then we need to take a decision of restarting the container or not.
Usually restarting the container would clear the CPU-consuming process but we don't want a Flex critical process to be cleared. Chances are that the load on Flex is huge and the container is trying its best to process all the queued jobs. Hence if the job container on the job node is consuming a High CPU, then we might need to get more detail from the customer on the actual usage of the flex system before restarting anything on the server. In the backend, each flex application service is running on two containers HA clusters on two different hosts. Eg a job service should have two job containers running on two different hosts. Hence we sometimes have the option to restart containers on the High CPU load machine when the respective container on the other machine is normal.
Before restarting the container get the logs of the services.
$ docker logs <container_name>
You can give try to restart one service at a time and wait for it to become healthy.
$ docker compose stop <service name>
$ docker compose start <service name>
Again use the CPU commands like top to identify the load.
Reboot the system: If nothing else seems to work and you can afford it, rebooting the system may solve the problem.