FLX_INFRA_CADVISOR_003 Alert
High memory usage can cause issues on the server, as the server will have less memory for caches making it more inefficient to process new jobs.
Flex memory usage depends on the number and usage of the different microservices so it can be a difficult issue to pinpoint.
Depending on the impacted server, multiple flex services might get impacted due to a shortage of memory.
Chances are there will out of memory occurring on containers and getting killed to insufficient memory.
Login to Xymon and see the server memory usage pattern.
- Log in to the server in question and for which we got alert
$ ssh SERVER
- Execute command
free -mto see the free memory and swap space
$ top -b -o +%MEM | head -n 22
- For Flex we also have docker tools that can simplify
docker statsYou can usedocker psto translate from the container id to the service name
The memory alerts might be caused by the following issues:
- services nodes - arangodb-server /Rabbimq/Redis/Mongo leak memory
- fix: restart the affected containers
- non-services nodes - flex service memory leak
- fix: restart the affected flex service. If the issue persists even after restarting, it requires further investigation from the developer team for a memory leak.
- none of the above - the server might simply not have enough memory to run all the apps. Review the memory usage of other nodes and consider either of following a) Moving applications around to different servers where memory is available b) Deploying new instances and moving applications. c) Increasing the size of instances and moving applications