FLX_BACKEND_JAVA_002 Alert

Description

A critical Alert is raised when the node has spent too much time on garbage collection over five minutes.

Severity

This alert is flagged as Warning.

Customer Impact

This alert usually occurs when there is a memory leak within the application service or the application in the node is aggressively using the heapspace.

This alert needs further investigation on the overall heap space usage of the node hosting the affected service.

A Memory Leak is a situation where there are objects present in the heap that are no longer used, but the garbage collector is unable to remove them from memory. The alert is raised when the node has spent too much time on garbage collection(indicating high heapspace usage). This usually impacts the flex functionality and degrades system performance over time. If not dealt with, the application will eventually exhaust its resources, finally terminating with a fatal java.lang.OutOfMemoryError.

Operational Remediation Process

Identify if there is a memory leak in the system.

  1. Login to the environment's Grafana page to review the heapspace usage.
  2. Click the search Icon on the left of grafana page and search for Flex Enterprise Dashboard.
  3. Click the Flex Enterprise Dashboard to view the overall CPU, memory , heapspace, garbage_collection, and other metrics of Flex.
  4. In this dashboard, you will be able to see two metric heapspace and garbage collection
  5. Select the timeframe on which you want to see the metric graphs.
  6. If you happen to find an unusual high heapspace usage on the node, chances are that there is a memory leak or the application is aggressively using heapspace.

One way to workaround this issue is to increase the heapspace and memory limit of the affected service.

If the heap usage is not decreasing, it's recommended to stop the affected service, and increase the mem_limit and Heap_space of the service in docker-compose.yml. Then start the service again with the new heap space configuration.

You can stop the service.

$ docker compose stop <container name>

Edit the docker-compose.yml file to update the mem_limit and HEAPSPACE to a higher value.

$ cd /flex
$ vim docker-compose.yml

Search for the service paragraph and change the Mem_limit to a bigger value, e.g: `mem_limit: 2G` increase it to `mem_limit: 4G`.

Similarly, change the `HEAPSPACE` to a bigger value to match it up with `mem_limit`, e.g: `HEAPSPACE=1g` increase it to `HEAPSPACE=2g`.

After changing the memory limit and heap space, it's required to restart the container.

```sh
$ cd /flex
$ docker compose up -d <service name>

Repeat the same steps on the remaining services if a heap space issue occurs on those hosts.

If the heap space usage doesn't lower after increasing the limit. We can try to rampup the memory and heapspace limit again provided there are resources(memory) available in the node. If the issue still happens after increasing the limit for the second time, then we should check for errors that might be indicating memory leaks.

In case of memory leak scenario, you are required to gather heapdump and threaddump of the affected service and raise a ticket with Dalet's R&D team for further investigation.

Capturing thread dumps

In order to capture thread or heap dump we need the jattach tool to be installed inside the container.

Login to the container using the docker command.

$ docker exec -it <container_name> /bin/sh

Execute the below one liner command within the container:

$ echo "APT::Default-Release \"$(awk -F"[)(]+" '/VERSION=/ {print $2}' /etc/os-release)\";" > /etc/apt/apt.conf.d/default-release && echo "deb http://deb.debian.org/debian testing main" >> /etc/apt/sources.list && apt update && apt-get install -y jattach

This should install the jattach tool inside the container.

To collect the thread dump run this command within the container:

$ jattach 1 threaddump

You'll probably want to save this output to a file - the easiest thing to do is to use one of the already mapped volumes (there's usually one for logs, for example!) and then you'll be able to access the file on the host and upload / copy it wherever necessary.

Capturing heap dumps

Heap dumps will usually take quite a lot of space - up to the amount of heap allocated to the JVM you're inspecting (eg. 2G+ for enterprise!). Make sure you have enough space available and save the dump straight to a path that is mounted from the host already (eg. the one used for logs) to ensure you're not leaving massive files in a container (which end up on the root volume which is usually smaller than the volume where we keep logs!)

Run this command within the container:

$ jattach 1 dumpheap <filename>