FLX_INFRA_OS_008 Alert
A critical Alert is raised when the disk on the server is experiencing high Disk IOPS (read and write) value for the last 15 minutes.
This alert indicates that there is a sudden high IOPS on both read and write happening on the disk on a node.
IOPS is an important performance metric for storage disks, as it can impact the speed and responsiveness of an application or system that relies on that storage. For example, in our environment, it is required to make sure that the IOPS is normal to ensure that database queries are processed quickly and efficiently.
- best case: The high IOPS can be temporary during any assets upload or multiple uploads. The usage return to normal once the upload finishes.
- worst case: The high IOPS can be causing service slowness like DB queries getting slow causing the UI to work slow.
We need to identify the process or container which is causing the high IOps (both read and write)
The best command to check disk IOPS is using the iostat command. It provides statistics on IOPS for all storage devices on the system. It is used to monitor disk workload in real-time.
Eg: To display disk I/O statistics every 2 seconds for 5 times, the following command can be used:
$ iostat -xd 2 5
If I need to check the IOPS for a specific device, such as /dev/sda, I will run the command like this. Here the interval is 3s.
$ iostat -xd 3 /dev/sda
To identify which process might be causing high IOPs usage, Execute
$ pidstat -dl 20
$ iotop