FLX_INFRA_OS_009 Alert
A critical Alert is raised when a specific server's disk/mount is likely to become full within the next 4 hours.
This alert warns the user about the chances that the disk/mount is likely to become full within 4 hours. Hence it's required to address the issue before the disk usage hits 100%, which can cause a flex outage.
Depending on the impacted disk, multiple impacts can be seen on flex application services or flex backend services.
- best case: Flex upload or jobs might fail due to insufficient disk available.
- worst case: Some core-Flex services can be completely down if the disk hits 100% usage. Eg: mysql will stop its writing if the disk is full.
We need to identify the file or folder which is consuming most of the space in the server.
Possible reasons can be:
- /tmp storing old files or uploaded files.
- /home hoarding disk space. User storing huge data in /home directory.
- /var/log/ log files are not being rotated and not deleted.
- /var/lib/docker consuming more space. Chances that huge files are getting stored in the container itself, hence consuming the docker volume directory.
- /flex/Too many flex application logs getting written in this directory.
Operational guidelines are:
- log in to the server
$ ssh SERVER
- Change to the directory which is showing the disk usage alert.
$ cd /flex
- Find the file/directory consuming the space.
$ du -shx *
- To remove the file.
$ rm -rf <file_name>