FLX_INFRA_OS_010 Alert

Description

A critical alert is raised when the system time is either too fast or too slow compared to a reference clock time.

Severity

This alert is flagged as Critical.

Customer Impact

This alert is triggered when the NTP drift second is greater than 0.05 seconds for more than 5 minutes.

All the flex servers should have the same time. The servers are configured to get the time from an external NTP server.

A time difference between servers can cause application errors on the services. For example, jobs running in flex can go to a pending state because there is a time difference between the servers which are hosting the job service and the server hosting other flex services. This alert is enabled to check if there is any NTP drift occurring on the system.

Operational Remediation Process

Login to any one of the nodes which have reported the NTP time drift.

By default, the NTP server configuration can be found in /etc/chrony.config

use the below command to get more information on the actual NTP drift

$ chronyc sources -v

Note that the value of MS other than (*) usually indicates a problem with the NTP server. Eg: if the value of MS is "^?" Then it means the servers are unreachable.

$ chronyc sourcestats

Check the chrony service status and see if there are any error reported.

$ systemctl status chronyd
$ journalctl -u chrony

Chrony logs can be found in /var/log/chrony/ directory. This can give more insight into the drift.

All the servers should have the identical NTP configuration to get the time. Hence we can compare the configuration from healthy servers.