@cryptotaxi247 / netdata-1 / commits / 0cf35dc90

Update docs health monitoring and health management api (#6435)

* Update docs health monitoring and health management api * Update docs health monitoring and health management api

Jelger Haanstra committed Jul 17, 2019 at 12:29 UTC 0cf35dc906a97d8efc841abe331433066503a925
2 files changed +42 -43
health/README.md
+20 -23
@@ -65,7 +65,7 @@ This line starts an alarm or alarm template.
65 alarm: NAME
66 ```
67
68 -or
68 +or
69
70 ```
71 template: NAME
@@ -161,7 +161,7 @@ The simple pattern syntax and operation is explained in [simple patterns](../lib
161 This line makes a database lookup to find a value. This result of this lookup is available as `$this`.
162
163 The format is:
164 -
164 +
165 ```
166 lookup: METHOD AFTER [at BEFORE] [every DURATION] [OPTIONS] [of DIMENSIONS]
167 ```
@@ -311,15 +311,15 @@ delay: [[[up U] [down D] multiplier M] max X]
311 notification for this event will be sent 10 seconds after the actual event. This is used in
312 hope the alarm will get back to its previous state within the duration given. The default `U`
313 is zero.
314 -
314 +
315 - `down D` defines the delay to be applied to a notification for an alarm that moves to lower
316 state (i.e. CRITICAL to WARNING, CRITICAL to CLEAR, WARNING to CLEAR). For example, `down 1m`
317 will delay the notification by 1 minute. This is used to prevent notifications for flapping
318 alarms. The default `D` is zero.
319 -
319 +
320 - `mutliplier M` multiplies `U` and `D` when an alarm changes state, while a notification is
321 delayed. The default multiplier is `1.0`.
322 -
322 +
323 - `max X` defines the maximum absolute notification delay an alarm may get. The default `X`
324 is `max(U * M, D * M)` (i.e. the max duration of `U` or `D` multiplied once with `M`).
325
@@ -361,13 +361,13 @@ repeat: [off] [warning DURATION] [critical DURATION]
361
362 #### Alarm line `option`
363
364 -The only possible value for the `option` line is
364 +The only possible value for the `option` line is
365
366 ```
367 option: no-clear-notification
368 ```
369
370 -For some alarms we need compare two time-frames, to detect anomalies. For example, `health.d/httpcheck.conf` has an alarm template called `web_service_slow` that compares the average http call response time over the last 3 minutes, compared to the average over the last hour. It triggers a warning alarm when the average of the last 3 minutes is twice the average of the last hour. In such cases, it is easy to trigger the alarm, but difficult to tell when the alarm is cleared. As time passes, the newest window moves into the older, so the average response time of the last hour will keep increasing. Eventually, the comparison will find the averages in the two time-frames close enough to clear the alarm. However, the issue was not resolved, it's just a matter of the newer data "polluting" the old. For such alarms, it's a good idea to tell Netdata to not clear the notification, by using the `no-clear-notification` option.
370 +For some alarms we need compare two time-frames, to detect anomalies. For example, `health.d/httpcheck.conf` has an alarm template called `web_service_slow` that compares the average http call response time over the last 3 minutes, compared to the average over the last hour. It triggers a warning alarm when the average of the last 3 minutes is twice the average of the last hour. In such cases, it is easy to trigger the alarm, but difficult to tell when the alarm is cleared. As time passes, the newest window moves into the older, so the average response time of the last hour will keep increasing. Eventually, the comparison will find the averages in the two time-frames close enough to clear the alarm. However, the issue was not resolved, it's just a matter of the newer data "polluting" the old. For such alarms, it's a good idea to tell Netdata to not clear the notification, by using the `no-clear-notification` option.
371
372 ---
373
@@ -417,14 +417,14 @@ crit: $this > (($status == $CRITICAL) ? (85) : (95))
417 The above say:
418 * If the alarm is currently a warning, then the threshold for being considered a warning
419 is 75, otherwise it's 85.
420 -
420 +
421 * If the alarm is currently critical, then the threshold for being considered critical
422 is 85, otherwise it's 95.
423
424 Which in turn, results in the following behavior:
425 * While the value is rising, it will trigger a warning when it exceeds 85, and a critical
426 alert when it exceeds 95.
427 -
427 +
428 * While the value is falling, it will return to a warning state when it goes below 85,
429 and a normal state when it goes below 75.
430
@@ -442,13 +442,13 @@ Which in turn, results in the following behavior:
442 You can find all the variables that can be used for a given chart, using
443 `http://your.netdata.ip:19999/api/v1/alarm_variables?chart=CHART_NAME`
444 Example: [variables for the `system.cpu` chart of the registry](https://registry.my-netdata.io/api/v1/alarm_variables?chart=system.cpu).
445 -
445 +
446 _Hint: If you don't know how to find the CHART_NAME, you can read about it [here](../docs/Charts.md#charts)._
447
448
449 -Netdata supports 3 internal indexes for variables that will be used in health monitoring.
449 +Netdata supports 3 internal indexes for variables that will be used in health monitoring.
450 <details markdown="1"><summary>The variables below can be used in both chart alarms and context templates.</summary>
451 -Although the `alarm_variables` link shows you variables for a particular chart, the same variables can also be used in templates for charts belonging to the same [context](../docs/Charts.md#contexts). The reason is that all charts of a given contexts are essentially identical, with the only difference being the [family](../docs/Charts.md#families) that identifies a particular hardware or software instance. Charts and templates do not apply to specific families anyway, unless if you explicitly limit an alarm with the [alarm line `families`](#alarm-line-families).
451 +Although the `alarm_variables` link shows you variables for a particular chart, the same variables can also be used in templates for charts belonging to the same [context](../docs/Charts.md#contexts). The reason is that all charts of a given contexts are essentially identical, with the only difference being the [family](../docs/Charts.md#families) that identifies a particular hardware or software instance. Charts and templates do not apply to specific families anyway, unless if you explicitly limit an alarm with the [alarm line `families`](#alarm-line-families).
452 </details>
453
454 - **chart local variables**. All the dimensions of the chart are exposed as local variables. The value of $this for the other configured alarms of the chart also appears, under the name of each configured alarm.
@@ -478,13 +478,13 @@ Although the `alarm_variables` link shows you variables for a particular chart,
478 - **special variables*** are:
479
480 - `$this`, which is resolved to the value of the current alarm.
481 -
481 +
482 - `$status`, which is resolved to the current status of the alarm (the current = the last
483 status, i.e. before the current database lookup and the evaluation of the `calc` line).
484 This values can be compared with `$REMOVED`, `$UNINITIALIZED`, `$UNDEFINED`, `$CLEAR`,
485 `$WARNING`, `$CRITICAL`. These values are incremental, ie. `$status > $CLEAR` works as
486 expected.
487 -
487 +
488 - `$now`, which is resolved to current unix timestamp.
489
490 ## Alarm Statuses
@@ -493,16 +493,16 @@ Alarms can have the following statuses:
493
494 - `REMOVED` - the alarm has been deleted (this happens when a SIGUSR2 is sent to netdata
495 to reload health configuration)
496 -
496 +
497 - `UNINITIALIZED` - the alarm is not initialized yet
498 -
498 +
499 - `UNDEFINED` - the alarm failed to be calculated (i.e. the database lookup failed,
500 a division by zero occurred, etc)
501 -
501 +
502 - `CLEAR` - the alarm is not armed / raised (i.e. is OK)
503 -
503 +
504 - `WARNING` - the warning expression resulted in true or non-zero
505 -
505 +
506 - `CRITICAL` - the critical expression resulted in true or non-zero
507
508 The external script will be called for all status changes.
@@ -675,9 +675,6 @@ You can find how netdata interpreted the expressions by examining the alarm at `
675
676 ## Disabling health checks or silencing notifications at runtime
677
678 -The health checks can be controlled at runtime via the [health management api](../web/api/health/#health-management-api).
678 +It's currently not possible to schedule notifications from within the alarm template. For those scenarios where you need to temporary disable notifications (for instance when running backups triggers a disk alert) you can disable or silence notifications are runtime. The health checks can be controlled at runtime via the [health management api](../web/api/health/#health-management-api).
679
680 [![analytics](https://www.google-analytics.com/collect?v=1&aip=1&t=pageview&_s=1&ds=github&dr=https%3A%2F%2Fgithub.com%2Fnetdata%2Fnetdata&dl=https%3A%2F%2Fmy-netdata.io%2Fgithub%2Fhealth%2FREADME&_u=MAC~&cid=5792dfd7-8dc4-476b-af31-da2fdb9f93d2&tid=UA-64295674-3)]()
681 -
682 -
683 -
web/api/health/README.md
+22 -20
@@ -50,7 +50,7 @@ From Netdata v1.16.0 and beyond, the configuration controlled via the API comman
50 Specifically, the API allows you to:
51 - Disable health checks completely. Alarm conditions will not be evaluated at all and no entries will be added to the alarm log.
52 - Silence alarm notifications. Alarm conditions will be evaluated, the alarms will appear in the log and the netdata UI will show the alarms as active, but no notifications will be sent.
53 - - Disable or Silence specific alarms that match selectors on alarm/template name, chart, context, host and family.
53 + - Disable or Silence specific alarms that match selectors on alarm/template name, chart, context, host and family.
54
55 The API is available by default, but it is protected by an `api authorization token` that is stored in the file you will see in the following entry of `http://localhost:19999/netdata.conf`:
56
@@ -59,13 +59,15 @@ The API is available by default, but it is protected by an `api authorization to
59 # netdata management api key file = /var/lib/netdata/netdata.api.key
60 ```
61
62 -You can access the API via GET requests, by adding the bearer token to an `Authorization` http header, like this:
62 +You can access the API via GET requests, by adding the bearer token to an `Authorization` http header, like this:
63
64 ```
65 -curl "http://myserver/api/v1/manage/health?cmd=RESET" -H "X-Auth-Token: Mytoken"
65 +curl "http://myserver/api/v1/manage/health?cmd=RESET" -H "X-Auth-Token: Mytoken"
66 ```
67
68 -The command `RESET` just returns netdata to the default operation, with all health checks and notifications enabled.
68 +By default access to the health management API is only allowed from `localhost`. Accessing the API from anything else will return a 403 error with the message `You are not allowed to access this resource.`. You can change permissions by editing the `allow management from` variable in netdata.conf within the [web] section. See [web server access lists](../../server/#access-lists) for more information.
69 +
70 +The command `RESET` just returns netdata to the default operation, with all health checks and notifications enabled.
71 If you've configured and entered your token correclty, you should see the plain text response `All health checks and notifications are enabled`.
72
73 ### Disable or silence all alarms
@@ -73,14 +75,14 @@ If you've configured and entered your token correclty, you should see the plain
75 If all you need is temporarily disable all health checks, then you issue the following before your maintenance period starts:
76
77 ```
76 -curl "http://myserver/api/v1/manage/health?cmd=DISABLE ALL" -H "X-Auth-Token: Mytoken"
78 +curl "http://myserver/api/v1/manage/health?cmd=DISABLE ALL" -H "X-Auth-Token: Mytoken"
79 ```
80
81 The effect of disabling health checks is that the alarm criteria are not evaluated at all and nothing is written in the alarm log.
82 If you want the health checks to be running but to not receive any notifications during your maintenance period, you can instead use this:
83
84 ```
83 -curl "http://myserver/api/v1/manage/health?cmd=SILENCE ALL" -H "X-Auth-Token: Mytoken"
85 +curl "http://myserver/api/v1/manage/health?cmd=SILENCE ALL" -H "X-Auth-Token: Mytoken"
86 ```
87
88 Alarms may then still be raised and logged in netdata, so you'll be able to see them via the UI.
@@ -88,44 +90,44 @@ Alarms may then still be raised and logged in netdata, so you'll be able to see
90 Regardless of the option you choose, at the end of your maintenance period you revert to the normal state via the RESET command.
91
92 ```
91 - curl "http://myserver/api/v1/manage/health?cmd=RESET" -H "X-Auth-Token: Mytoken"
93 + curl "http://myserver/api/v1/manage/health?cmd=RESET" -H "X-Auth-Token: Mytoken"
94 ```
95
96 ### Disable or silence specific alarms
97
96 -If you do not wish to disable/silence all alarms, then the `DISABLE ALL` and `SILENCE ALL` commands can't be used.
98 +If you do not wish to disable/silence all alarms, then the `DISABLE ALL` and `SILENCE ALL` commands can't be used.
99 Instead, the following commands expect that one or more alarm selectors will be added, so that only alarms that match the selectors are disabled or silenced.
98 -- `DISABLE` : Set the mode to disable health checks.
99 -- `SILENCE` : Set the mode to silence notifications.
100 +- `DISABLE` : Set the mode to disable health checks.
101 +- `SILENCE` : Set the mode to silence notifications.
102
101 -You will normally put one of these commands in the same request with your first alarm selector, but it's possible to issue them separately as well.
102 -You will get a warning in the response, if a selector was added without a SILENCE/DISABLE command, or vice versa.
103 +You will normally put one of these commands in the same request with your first alarm selector, but it's possible to issue them separately as well.
104 +You will get a warning in the response, if a selector was added without a SILENCE/DISABLE command, or vice versa.
105
104 -Each request can specify a single alarm `selector`, with one or more `selection criteria`.
105 -A single alarm will match a `selector` if all selection criteria match the alarm.
106 +Each request can specify a single alarm `selector`, with one or more `selection criteria`.
107 +A single alarm will match a `selector` if all selection criteria match the alarm.
108 You can add as many selectors as you like.
109 In essence, the rule is: IF (alarm matches all the criteria in selector1 OR all the criteria in selector2 OR ...) THEN apply the DISABLE or SILENCE command.
110
111 To clear all selectors and reset the mode to default, use the `RESET` command.
112
111 -The following example silences notifications for all the alarms with context=load:
113 +The following example silences notifications for all the alarms with context=load:
114
115 ```
114 -curl "http://myserver/api/v1/manage/health?cmd=SILENCE&context=load" -H "X-Auth-Token: Mytoken"
116 +curl "http://myserver/api/v1/manage/health?cmd=SILENCE&context=load" -H "X-Auth-Token: Mytoken"
117 ```
118
117 -#### Selection criteria
119 +#### Selection criteria
120
119 -The `selection criteria` are key/value pairs, in the format `key : value`, where value is a netdata [simple pattern](../../../libnetdata/simple_pattern/). This means that you can create very powerful selectors (you will rarely need more than one or two).
121 +The `selection criteria` are key/value pairs, in the format `key : value`, where value is a netdata [simple pattern](../../../libnetdata/simple_pattern/). This means that you can create very powerful selectors (you will rarely need more than one or two).
122
123 The accepted keys for the `selection criteria` are the following:
122 -- `alarm` : The expression provided will match both `alarm` and `template` names.
124 +- `alarm` : The expression provided will match both `alarm` and `template` names.
125 - `chart` : Chart ids/names, as shown on the dashboard. These will match the `on` entry of a configured `alarm`.
126 - `context` : Chart context, as shown on the dashboard. These will match the `on` entry of a configured `template`.
127 - `hosts` : The hostnames that will need to match.
128 - `families` : The alarm families.
129
128 -You can add any of the selection criteria you need on the request, to ensure that only the alarms you are interested in are matched and disabled/silenced. e.g. there is no reason to add `hosts: *`, if you want the criteria to be applied to alarms for all hosts.
130 +You can add any of the selection criteria you need on the request, to ensure that only the alarms you are interested in are matched and disabled/silenced. e.g. there is no reason to add `hosts: *`, if you want the criteria to be applied to alarms for all hosts.
131
132 Example 1: Disable all health checks for context = `random`
133