Metrics Collection

Dalet Flex uses Prometheus protocol to collect various system and application metrics and sotre them into a TimeSeries database (TSDB), allowing operators to get an eagle-eye retrospective on the platform and helps with monitoring and occasionally, troubleshooting.

By default, both logs and metrics collection is performed on every single instance and shipped to the monitor instance, holding data for long duration (defaults to one month).

Info

Collecting logs and metrics consumes a bit of CPU resources (few %). While insignificant enough, that's still a loss of resources.

Should you be willing to completely disable logs and metrics collection, this can be done by setting the appropriate variable:

dalet_baseos_monitoring_enabled: false

This is however not recommended practice, as it would turn Flex in some kind of a blackbox, with no proper troubleshooting insights.

Collection Frequency

Metrics collection is processed by taking instant snapshots of the system. By default, these snapshots happen every 15s, which seems to be a good compromise. More frequent collection leads to more resources consumption (CPU, but also more data to be stored), less frequent, implies that you'd miss some events. For instance, if you take a snapshot of CPU usage every 5 minutes or so, any unusual ephemeral load spike during this timeframe will be ignored.

Sampling frequency can be accomodated by setting:

dalet_flex_metrics_scrape_interval: 15s

where format: is <VALUE><UNIT> where support units are: y, w, d, h, m, s, ms.

Data Retention

Every server instance is responsible for exposing its own data, being collected by a local observability agent (Grafana Alloy, before being sent to persistent TSDB engine(s).

By default, the monitor instance use VictoriaMetrics as its Prometheus-compatible TSDB. VictoriaMetrics provides the same features than regular Prometheus server implementation while being less memory-hungry.

Should there be a specific compatibility need for, tThe TSDB implementation can be switched back to Prometheus by setting:

dalet_flex_metrics_subsystem: prometheus

and data retention can be adjusted through:

dalet_flex_metrics_retention_period: 30d

Metrics older than specified threshold will be automatically pruned and increasing retention period will lead to increased disk usage.

External Status

Aside from collecting system and application metrics, it sometimes might be useful to get live status of service's state, from an end-user perspective (sometimes called RUM, for Real-User-Monitoring).

It is for example possible to target HTTP(s) endpoint and do periodic queries, just to ensure the service's up. Status will then be collected and formatted in Prometheus syntax, ready to be queries in PromQL.

Tip

Keep in mind that any collected metric can be further queried and either displayed as a graph (or equivalent widget) in Grafana dashboards, or be used to trigger an automated alert to on-all personnel (e.g. raise an incident if the public HTTP endpoint is not reachable for the past 30s).

Monitored web targets can be easily extended by defining them as:

  - name: string          # Target's name
    enabled: bool         # Define if the target should be monitored (defaults to true).
    url: string           # Optional, URL of the Web resource to be monitored
    kv: string            # Optional, key from service registry to retrieve URL from.
                          # At least one of 'url' or 'kv' field must be defined.
    url_path: string      # Optional: specific path to append the URL to be monitored.
    expect: string|int    # Optional: expected HTTP code(s) for monitor probe to consider successful.
                          # Must be either an integer (e.g. 200) or a pipe-separated string for multiple codes
                          # (e.g. '302|401'). Defaults to 2xx codes if unspecified.

and declaring a list of those in the following variable:

dalet_flex_metrics_scrape_web_targets_custom: []

Grafana Access

Dalet Flex comes with pre-bundled and configured Grafana instance and packaged dashboards, ready to be consumed.

You may however require several user accounts to be created for your operators team. It's easy to push for additional Grafana user accounts (Viewer privilege, i.e. read-only rights), allowing them to access metrics and dashboards and perform queries:

Accounts are to be formatted as:

  - username: USER_NAME
    password: USER_PASSWORD
    email:    USER_EMAIL

and added back to the following variable:

dalet_flex_metrics_grafana_extra_users: []