Added info on health configuration, with a separate page for Charts, Dimensions, Alarms, Contexts (#4895)
Chris Akritidis committed
Dec 3, 2018 at 04:47 UTC
0fad9bf5b9b4c4bbe7f6eae19e66f2d9a5fa3b92
3 files changed
+56
-23
docs/Charts.md
new
+25
@@ -0,0 +1,25 @@
1
+# Charts, contexts, families
2
+
3
+Before configuring an alarm or writing a collector, it's important to understand how Netdata organizes collected metrics into charts.
4
+
5
+## Charts
6
+
7
+Each chart that you see on the netdata dashboard contains one or more dimensions, one for each collected or calculated metric.
8
+
9
+The chart name or chart id is what you see in parentheses at the top left corner of the chart you are interested in. For example, if you go to the system cpu chart: `http://your.netdata.ip:19999/#menu_system_submenu_cpu`, you will see at the top left of the chart the label "Total CPU utilization (system.cpu)". In this case, the chart name is `system.cpu`.
10
+
11
+## Dimensions
12
+
13
+Most charts depict more than one dimensions. The dimensions of a chart are called "series" in some applications. You can see these dimensions on the right side of a chart, right under the date and time. For the system.cpu example we used, you will see the dimensions softirq, irq, user etc. Note that these are not always simple metrics (raw data). They could be calculated values (percentages, aggregates and more).
14
+
15
+## Families
16
+
17
+When you have several instances of a monitored hardware or software (e.g. network interfaces, mysql instances etc.), you need to be able to identify each one separately. Netdata uses "families" to identify such instances. For example, if I have the network interfaces `eth0` and `eth1`, `eth0` will be one family, and `eth1` will be another.
18
+
19
+The reasoning behind calling these instances "families" is that different charts for the same instance can and many times are related (relatives, family, you get it). The family of a chart is usually the name of the netdata dashboard submenu that you see selected on the right navigation pane, when you are looking at a chart. For the example of the two network interfaces, you would see a submenu `eth0` and a submenu `eth1` under the "Network Interfaces" menu on the right navigation pane.
20
+
21
+## Contexts
22
+
23
+A context is a grouping of identical charts, for each instance of the hardware or software monitored. For example, `health/health.d/net.conf` refers to four contexts: `net.drops`, `net.fifo`, `net.net`, `net.packets`. You can see the context of a chart if you hover over the date right above the dimensions of the chart. The line that appears shows you two things: the collector that produces the chart and the chart context.
24
+
25
+For example, let's take the `net.packets` context. You will see on the dashboard as many charts with context net.packets as you have network interfaces (families). These charts will be named `net_packets.[family]`. For the example of the two interfaces `eth0` and `eth1`, you will see charts named `net_packets.eth0` and `net_packets.eth1`. Both of these charts show the exact same dimensions, but for different instances of a network interface.
docs/generator/buildyaml.sh
+2
-1
@@ -139,7 +139,8 @@ echo -ne "- Running netdata:
139
"
140
navpart 2 daemon
141
navpart 2 daemon/config
142
-
142
+echo -ne " - 'docs/Charts.md'
143
+"
144
navpart 2 web/server "" "Web server"
145
navpart 3 web/server "" "" 2 excludefirstlevel
146
echo -ne " - Running behind another web server:
health/README.md
+29
-22
@@ -9,8 +9,8 @@ netdata, since many charts are dynamically created during runtime (for example,
9
chart tracking network interface packet drops, is automatically created on the first
10
packet dropped).
11
12
-Netdata also supports alarm **templates**, so that an alarm can be attached to all
13
-the charts of the same context (i.e. all network interfaces, or all disks, or all mysql servers, etc.)
12
+Netdata also supports alarm **templates**, so that an alarm can be attached to all the charts of the same context (i.e. all network interfaces, or all disks, or all mysql servers, etc.).
13
+
14
15
Each alarm can execute a single query to the database using statistical algorithms against past data,
16
but alarms can be combined. So, if you need 2 queries in the database, you can combine
@@ -145,7 +145,7 @@ This is useful when you centralize metrics from multiple hosts, to one netdata.
145
This line is only used in alarm templates. It filters the charts. So, if you need to create
146
an alarm template for a few of a kind of chart (a few of your disks, or a few of your network
147
interfaces, or a few your mysql servers, etc), you can create an alarm template that would
148
-normally be applied to all of them, and filter them by family.
148
+normally be applied to all of them, and filter them by [family](../docs/Charts.md#families).
149
150
The format is:
151
@@ -153,14 +153,7 @@ The format is:
153
families: SIMPLE PATTERN LIST
154
```
155
156
-Simple patterns list is a lists of space separated patterns. Use ` * ` as wildcard and ` ! `
157
-for a negative match. Processing is left to right, and on the first hit (positive or negative),
158
-processing stops.
159
-
160
-So. `families: *` means, match anything, while `families: !bad*pattern* *` means anything
161
-except `bad*pattern*` (where `*` is a wildcard to match any sequence of characters).
162
-
163
-The family of a chart is usually the submenu of the netdata dashboard it appears.
156
+The simple pattern syntax and operation is explained in [simple patterns](../libnetdata/simple_pattern/).
157
158
---
159
@@ -349,6 +342,16 @@ delay: [[[up U] [down D] multiplier M] max X]
342
their matching one) and a delay is in place.
343
- All are reset to their defaults when the alarm switches state without a delay in place.
344
345
+#### Alarm line `option`
346
+
347
+The only possible value for the `option` line is
348
+
349
+```
350
+option: no-clear-notification
351
+```
352
+
353
+For some alarms we need compare two time-frames, to detect anomalies. For example, `health.d/httpcheck.conf` has an alarm template called `web_service_slow` that compares the average http call response time over the last 3 minutes, compared to the average over the last hour. It triggers a warning alarm when the average of the last 3 minutes is twice the average of the last hour. In such cases, it is easy to trigger the alarm, but difficult to tell when the alarm is cleared. As time passes, the newest window moves into the older, so the average response time of the last hour will keep increasing. Eventually, the comparison will find the averages in the two time-frames close enough to clear the alarm. However, the issue was not resolved, it's just a matter of the newer data "polluting" the old. For such alarms, it's a good idea to tell Netdata to not clear the notification, by using the `no-clear-notification` option.
354
+
355
---
356
357
### Expressions
@@ -419,10 +422,19 @@ Which in turn, results in the following behavior:
422
423
### Variables
424
422
-netdata supports 3 new internal indexes for variables that will be used in health monitoring:
425
+You can find all the variables that can be used for a given chart, using
426
+`http://your.netdata.ip:19999/api/v1/alarm_variables?chart=CHART_NAME`
427
+Example: [variables for the `system.cpu` chart of the registry](https://registry.my-netdata.io/api/v1/alarm_variables?chart=system.cpu).
428
+
429
+_Hint: If you don't know how to find the CHART_NAME, you can read about it [here](../docs/Charts.md#charts)._
430
+
431
424
- - **chart local variables**. All the dimensions of the chart are exposed as local variables.
425
- All chart alarms names are exposed as variables too.
432
+Netdata supports 3 internal indexes for variables that will be used in health monitoring.
433
+<details markdown="1"><summary>The variables below can be used in both chart alarms and context templates.</summary>
434
+Although the `alarm_variables` link shows you variables for a particular chart, the same variables can also be used in templates for charts belonging to the same [context](../docs/Charts.md#contexts). The reason is that all charts of a given contexts are essentially identical, with the only difference being the [family](../docs/Charts.md#families) that identifies a particular hardware or software instance. Charts and templates do not apply to specific families anyway, unless if you explicitly limit an alarm with the [alarm line `families`](#alarm-line-families).
435
+</details>
436
+
437
+ - **chart local variables**. All the dimensions of the chart are exposed as local variables. The value of $this for the other configured alarms of the chart also appears, under the name of each configured alarm.
438
439
Charts also define a few special variables:
440
@@ -448,20 +460,15 @@ netdata supports 3 new internal indexes for variables that will be used in healt
460
461
- **special variables*** are:
462
451
- - `this`, which is resolved to the value of the current alarm.
463
+ - `$this`, which is resolved to the value of the current alarm.
464
453
- - `status`, which is resolved to the current status of the alarm (the current = the last
465
+ - `$status`, which is resolved to the current status of the alarm (the current = the last
466
status, i.e. before the current database lookup and the evaluation of the `calc` line).
467
This values can be compared with `$REMOVED`, `$UNINITIALIZED`, `$UNDEFINED`, `$CLEAR`,
468
`$WARNING`, `$CRITICAL`. These values are incremental, ie. `$status > $CLEAL` works as
469
expected.
470
459
- - `now`, which is resolved to current unix timestamp.
460
-
461
-You can find all the variables that can be used for a given chart, using
462
-`http://your.netdata.ip:19999/api/v1/alarm_variables?chart=NAME`.
463
-This will dump all the indexes from the chart's perspective.
464
-Example: [variables for the `system.cpu` chart of the registry](https://registry.my-netdata.io/api/v1/alarm_variables?chart=system.cpu).
471
+ - `$now`, which is resolved to current unix timestamp.
472
473
## Alarm Statuses
474