FLX_INFRA_OS_011 Alert
A critical alert is raised when the node has got too many sockets in 'mem' state which result in TCP out-of-memory errors.
This alert is triggered when the number of TCP sockets in 'mem' state in a node drasticicaly increase within 5 minutes.
This indicates that one of the backend service is using more number of TCP sockets in 'mem' state than usual and causing TCP out-of-memory errors on the node. The service causing the TCP out-of-memory error would be healthy but will not be active to do it job due to exhausted TCP connection. This makes the service non-functional and causing the dependant services with additional errors.
Login to any one of the server nodes which is reported in the alert.
Check the kernel ring buffer logs to see if the TCP out-of-memory error is reported.
$ dmsg -T
</div>
</div>
Check grafana to get the acutal TCP socket usage on each node.
Login to grafana, browse "Flex - Infrastructure" and then click on "TCP / socket stats" to get the graph.
Identify the service which is causing the TCP out-of-memory issue.
```sh
$ ss -ntmp
$ nstat -a | egrep -i "(lost|retransmit)"
Once the service is identified, we need to restart the container to fix the issue.
$ docker ps | grep <service name>
Get the logs of the services.
$ docker logs <container_name>
You can restart the service one node at a time.
$ docker compose stop <service name>
$ docker compose start <service name>
Check again the kernel ring buffer to see if tcp out of memory error stoped.