FLX_BACKEND_ELASTIC_003 Alert

Description

A critical Alert is raised when the Elasticsearch cluster is unhealthy.

Severity

This alert is flagged as Critical.

Customer Impact

This alert indicates that the one or many node which makes the elasticsearch cluster is unhealthy. In the flex backend, the elasticsearch is configured to run on 3 node cluster. This Alert will be raised if the number of elasitcsearch services in a cluster is less than 2, Indicating an unstable elastic search cluster. The elasticsearch services can go unhealthy if there is an internal error within the elasticsearch app itself or when the elasticsearch service goes out of memory if there is heavy usage due to reindexing.

Unhealthy elasticsearch can happen for 3 reasons.

  1. Unhealthy elasticsearch service nodes.
  2. UNASSIGNED shards.
  3. Disk usage on the service node is more than 85%.

Operational Remediation Process

Elasticsearch services usually run on the three service nodes.

We can get the overall health of the elasticsearch cluster by calling the elasticsearch URL inside the service node.

$ curl -XGET http://<elastichost>:9200/_cluster/health?pretty

A healthy cluster should have this output

	{
	  "cluster_name" : "nrl",
	  "status" : "green",
	  "timed_out" : false,
	  "number_of_nodes" : 3,
	  "number_of_data_nodes" : 3,
	  "active_primary_shards" : 325,
	  "active_shards" : 650,
	  "relocating_shards" : 0,
	  "initializing_shards" : 0,
	  "unassigned_shards" : 0,
	  "delayed_unassigned_shards" : 0,
	  "number_of_pending_tasks" : 0,
	  "number_of_in_flight_fetch" : 0,
	  "task_max_waiting_in_queue_millis" : 0,
	  "active_shards_percent_as_number" : 100.0
	}

In the above output, you can see that the number_of_nodes is 3. This shows that the cluster is healthy with 3 nodes running.

Also, the status should be green

we need to identify the failed elasticsearch service node if the number_of_nodes is less than 3.

  1. Unhealthy elasticsearch service nodes.

Log in to each service node and check the health status of the elasticsearch service.

$ ssh SERVER
get the elastic search container name and the health status.
$ docker ps | grep <service name>

If the health status is unhealthy, we are required to restart the elasticsearch container and check again the health status.

$ docker compose stop <service name>
$ docker compose start <service name>

You can get the logs of the elasticsearch service using

$ docker logs <service name>

You need to query the elasticsearch URL again to confirm if the failed elasticsearch node rejoined the cluster.

$ curl -XGET http://<elastichost>:9200/_cluster/health?pretty
  1. UNASSIGNED shards.

Sometime the elasticsearch cluster status goes to yellow if there are unassigned shard. you can check the shards status by querying below curl

$ curl -XGET http://<elastichost>:9200/_cat/shards | grep UNASSIGNED

See reference

  1. Disk usage on the service node is more than 85%.

Also, we need to make sure the disk space of /flex should not exceed 85% disk usage. Elasitcsearch cluster can go into an unhealthy state if the disk usage is more than the allocation.disk.watermark which is usually 85%. To find out the watermark allocation execute

$ curl -s '<elastichost>:9200/_cat/allocation?v'

The only way to resolve this issue is by increasing the disk space or by housekeeping /flex folder to make the over overall disk usage below 85%.