FLX_BACKEND_RABBITMQ_003 Alert
A critical alert is raised when a single or more RabbitMQ queue depth values are getting big over time.
This alert is triggered when a single or more RabbitMQ queue depth value is greater than 1000 messages for 10 minutes. There are multiple reasons for the messages to get queued up and doesn't get processed. This can be caused due to:
-
No consumer is available in the queue to process the message.
-
Heavy load on the application causes more messages in the queue which the consumer cant handle.
-
Less memory on consumer service which can't handle the incoming number of messages.
-
When a reindex is triggered, usually, more messages get queued in indexelastic queues for a short period of time.
Log in to one of the RabbitMQ server and also enable local port forwarding to the RabbitMQ port 15672
$ ssh SERVER -L15672:localhost:15672
Open the web browser and log in to rabbitmq GUI by keying http://localhost:15672
Check the overall messaging status by seeing the number of messages in ready and total queue in the rabbitmq UI.
Click the Queue tab and search for the queue which is having high queue depth (more messages queued).
Click the queue and check the number of consumers assigned to the queue. Usually, two consumer IPs should be reported under the consumer details with multiple connections. If there aren't two consumers reported in queue detail, then probably one of the service containers related to the queue is unhealthy or down.
$ ssh SERVER
Note the health state of the service (affected queue service).
$ docker ps | grep <service name>
Check for errors in the service logs
$ docker logs <service name>
You need to restart the service if the health status is 'unhealthy'
$ docker compose stop <service name>
$ docker compose start <service name>
Check the health status after restarting.
$ docker ps | grep <service name>
Check the status of the queue. See if you can see two consumer IPs in the consumer detail.
Check if there is enough memory on the service container(consumer) with is related to the queue.
$ grep -A20 <service_name> /flex/docker-compose.yml
$ docker stats
Also, check the service logs for errors.
$ docker logs <service name>