FLX_BACKEND_RABBITMQ_005 Alert

Description

A critical alert is raised when a Split brain condition is detected in the RabbitMQ cluster.

Severity

This alert is flagged as Critical.

Customer Impact

This alert is triggered when the Split-brain in the RabbitMQ cluster has been detected. While a RabbitMQ partition is in place, the two (or more!) sides of the cluster can evolve independently, with both sides thinking the other has crashed. This scenario is known as split-brain. This can cause two independent clusters to run in parallel. The application can be still functional but it would be processing messages via a single node cluster.

Operational Remediation Process

Log in to one of the RabbitMQ servers and also enable local port forwarding to the rabbitmq port 15672

$ ssh SERVER -L15672:localhost:15672

Open the web browser and log in to RabbitMQ GUI by keying http://localhost:15672

Check the rabbitmq cluster status on the RabbitMQ UI.

For a healthy cluster status, you should see three network partition RabbitMQ nodes running. Also, the queue shouldn't have a huge number of messages piled up.

In the case of a split brain, you might see two node cluster or a single-node cluster. Chances are that the RabbitMQ on another service node is running on its own cluster.

The only approach to the split-brain issue in RabbitMQ is to recreate the entire RabbitMQ service on the node which is not part of the RabbitMQ cluster.

  1. Stop and delete RabbitMQ container
$ docker compose -f docker-compose.yml stop rabbitmq
$ docker compose -f docker-compose.yml rm -f rabbitmq
  1. Deregister RabbitMQ service
  2. Move or delete /flex/rabbitmq/storage
$ sudo mv /flex/rabbitmq/rabbitmq /flex/rabbitmq/rabbitmq.old
  1. One by one start RabbitMQ Recreate the container
$ docker compose up -d rabbitmq
  1. Check consul UI to make sure it registered
  2. Ensure that the second and third nodes are joining the cluster.
$ docker compose -f docker-compose.yml exec rabbitmq rabbitmqctl cluster_status
  1. Use the RabbitMQ UI to confirm the cluster node status.