FLX_INFRA_OS_005 Alert

Description

A critical alert is raised when a process gets killed due to an out-of-memory condition in a host.

Severity

This alert is flagged as Critical.

Customer Impact

This alert is triggered when a process gets killed due to an out-of-memory condition in a host. Usually, a process or container might get killed because they used more memory than allowed. This alert is triggered as soon as the OOM error is reported in the kernel log. The impact depends on the process/container which got killed due to OOM. The impact will be huge if containers like Job or master service get killed. Hence it's crucial to address the issue as soon as possible.

Operational Remediation Process

Login to any one of the nodes which have reported the OOM.

You can find out the process (or Container) which got killed due to OOM from the kernel buffer log.

$ dmesg -T

As soon as you identify the process (container) from the logs, you are required to restart the failed process (or container)

Let's say we have an OOM killed container, in that case, you are required to bring up the killed container.

$ cd /flex
$ docker compose up -d  <service-name>

Wait for the container to become healthy and monitor the memory usage on the server in a timely manner.

Sometimes due to extensive usage of flex, there can be a high mem usage happening on the flex services. Eg: more number of jobs triggered on the flex UI can cause the job service containers to get killed due to OOM. Hence in this case we are required to ramp up the memory limit of the job node.

NOTE: This is just a timely workaround.

Edit the docker-compose.yml file to update the mem_limit to a higher value.

$ cd /flex
$ vim docker-compose.yml

Search for the service paragraph and change the Mem_limit to a bigger value, e.g: mem_limit: 2G increase it to mem_limit: 4G.

After changing the memory limit, it's required to restart the container.

$ cd /flex
$ docker compose stop <service name>
$ docker compose start <service name>