FLX_INFRA_OS_004 Alert
High memory usage will cause the server to slow down. As server will have less memory for caches making it more inefficient to execute new processes.
Flex memory usage depends on the number and usage of the different microservices so it can be a difficult issue to pinpoint. The memory spike might be due to a sudden high usage of flex UI which reflect in the server or the server itself isn't having enough memory resources to run the services.
Depending on the impacted server, multiple flex services might get impacted due to a shortage of memory. Eg: if the master node gets a high memory usage alert then the impact can be seen in the flex UI being slow to respond.
Login to xymon and see the server memory usage graph.
See if there is a memory usage patter over time.
- Log in to the server in question and for which we got alert
$ ssh SERVER
-
Execute command
free -mto see the free memory -
Get the top memory consuming process in the server using below command.
$ top -b -o +%MEM | head -n 22
- We also have docker tools that can simplify
$ docker stats
The memory alerts might be caused by the following issues:
- services nodes - arangodb-server /Rabbimq/Redis/Mongo leak memory
- fix: restart the affected containers
- non-services nodes - flex service memory leak
- fix: restart the affected flex service. If the issue persists even after restarting, it requires further investigation from the developer team for a memory leak.
- none of the above - the server might simply not have enough memory to run all the apps. Review the memory usage of other nodes and consider either of following a) Moving applications around to different servers where memory is available b) Deploying new instances and moving applications. c) Increasing the size of instances.