moved related wiki pages into the repo (#4428)
* moved related wiki pages into the repo * updated web server docs * fixed typos
Costa Tsaousis committed
Oct 18, 2018 at 17:31 UTC
e76aac74e69c7dd03060e800e206eee777661a0c
70 files changed
+4082
-56
.gitignore
+1
-1
@@ -87,7 +87,7 @@ system/netdata.plist
87
system/netdata-freebsd
88
system/edit-config
89
90
-health/alarm-notify.sh
90
+health/notifications/alarm-notify.sh
91
collectors/cgroups.plugin/cgroup-name.sh
92
collectors/tc.plugin/tc-qos-helper.sh
93
collectors/charts.d.plugin/charts.d.plugin
CMakeLists.txt
+2
-2
@@ -368,8 +368,8 @@ set(API_PLUGIN_FILES
368
web/api/rrd2json.h
369
web/api/web_api_v1.c
370
web/api/web_api_v1.h
371
- web/api/web_buffer_svg.c
372
- web/api/web_buffer_svg.h
371
+ web/api/badges/web_buffer_svg.c
372
+ web/api/badges/web_buffer_svg.h
373
)
374
375
set(STREAMING_PLUGIN_FILES
Makefile.am
+2
-2
@@ -287,8 +287,8 @@ API_PLUGIN_FILES = \
287
web/api/rrd2json.h \
288
web/api/web_api_v1.c \
289
web/api/web_api_v1.h \
290
- web/api/web_buffer_svg.c \
291
- web/api/web_buffer_svg.h \
290
+ web/api/badges/web_buffer_svg.c \
291
+ web/api/badges/web_buffer_svg.h \
292
$(NULL)
293
294
STREAMING_PLUGIN_FILES = \
collectors/apps.plugin/README.md
+165
-17
@@ -22,22 +22,6 @@ utilization of exit processes. Their utilization is accounted at their currently
22
So, `apps.plugin` is perfectly able to measure the resources used by shell scripts and other processes
23
that fork/spawn other short lived processes hundreds of times per second.
24
25
-For example, ssh to a server running netdata and execute this:
26
-
27
-```sh
28
-while true; do ls -l /var/run >/dev/null; done
29
-```
30
-
31
-All the console tools will report that a a CPU core is 100% used, but they will fail to identify which
32
-process is using all that CPU (because there is no single process using it - thousands of `ls` per second
33
-are using it). Netdata however, will be able to identify that `ssh` is using it
34
-(`ssh` is the parent process group defined in its [default config](apps_groups.conf)):
35
-
36
-
37
-
38
-This feature makes `apps.plugin` unique in narrowing down the list of offending processes that may be
39
-responsible for slow downs, or abusing system resources.
40
-
25
## Charts
26
27
`apps.plugin` provides charts for 3 sections:
@@ -221,4 +205,168 @@ Examples below for process group `sql`:
205
- Open Sockets 
206
207
224
-For more information about badges check [Generating Badges](https://github.com/netdata/netdata/wiki/Generating-Badges)
\ No newline at end of file
208
+For more information about badges check [Generating Badges](../../web/api/badges)
209
+
210
+## Comparison with console tools
211
+
212
+Ssh to a server running netdata and execute this:
213
+
214
+```sh
215
+while true; do ls -l /var/run >/dev/null; done
216
+```
217
+
218
+In most systems `/var/run` is a `tmpfs` device, so there is nothing that can stop this command
219
+from consuming entirely one of the CPU cores of the machine.
220
+
221
+As we will see below, **none** of the console performance monitoring tools can report that this
222
+command is using 100% CPU. They do report of course that the CPU is busy, but **they fail to
223
+identify the process that consumes so much CPU**.
224
+
225
+Here is what common Linux console monitoring tools report:
226
+
227
+#### top
228
+
229
+`top` reports that `bash` is using just 14%.
230
+
231
+If you check the total system CPU utilization, it says there is no idle CPU at all, but `top`
232
+fails to provide a breakdown of the CPU consumption in the system. The sum of the CPU utilization
233
+of all processes reported by `top`, is 15.6%.
234
+
235
+```
236
+top - 18:46:28 up 3 days, 20:14, 2 users, load average: 0.22, 0.05, 0.02
237
+Tasks: 76 total, 2 running, 74 sleeping, 0 stopped, 0 zombie
238
+%Cpu(s): 32.8 us, 65.6 sy, 0.0 ni, 0.0 id, 0.0 wa, 1.3 hi, 0.3 si, 0.0 st
239
+KiB Mem : 1016576 total, 244112 free, 52012 used, 720452 buff/cache
240
+KiB Swap: 0 total, 0 free, 0 used. 753712 avail Mem
241
+
242
+ PID USER PR NI VIRT RES SHR S %CPU %MEM TIME+ COMMAND
243
+12789 root 20 0 14980 4180 3020 S 14.0 0.4 0:02.82 bash
244
+ 9 root 20 0 0 0 0 S 1.0 0.0 0:22.36 rcuos/0
245
+ 642 netdata 20 0 132024 20112 2660 S 0.3 2.0 14:26.29 netdata
246
+12522 netdata 20 0 9508 2476 1828 S 0.3 0.2 0:02.26 apps.plugin
247
+ 1 root 20 0 67196 10216 7500 S 0.0 1.0 0:04.83 systemd
248
+ 2 root 20 0 0 0 0 S 0.0 0.0 0:00.00 kthreadd
249
+```
250
+
251
+#### htop
252
+
253
+Exactly like `top`, `htop` is providing an incomplete breakdown of the system CPU utilization.
254
+
255
+```
256
+ CPU[||||||||||||||||||||||||100.0%] Tasks: 27, 11 thr; 2 running
257
+ Mem[||||||||||||||||||||85.4M/993M] Load average: 1.16 0.88 0.90
258
+ Swp[ 0K/0K] Uptime: 3 days, 21:37:03
259
+
260
+ PID USER PRI NI VIRT RES SHR S CPU% MEM% TIME+ Command
261
+12789 root 20 0 15104 4484 3208 S 14.0 0.4 10:57.15 -bash
262
+ 7024 netdata 20 0 9544 2480 1744 S 0.7 0.2 0:00.88 /usr/libexec/netd
263
+ 7009 netdata 20 0 138M 21016 2712 S 0.7 2.1 0:00.89 /usr/sbin/netdata
264
+ 7012 netdata 20 0 138M 21016 2712 S 0.0 2.1 0:00.31 /usr/sbin/netdata
265
+ 563 root 20 0 308M 202M 202M S 0.0 20.4 1:00.81 /usr/lib/systemd/
266
+ 7019 netdata 20 0 138M 21016 2712 S 0.0 2.1 0:00.14 /usr/sbin/netdata
267
+```
268
+
269
+#### atop
270
+
271
+`atop` also fails to break down CPU usage.
272
+
273
+```
274
+ATOP - localhost 2016/12/10 20:11:27 ----------- 10s elapsed
275
+PRC | sys 1.13s | user 0.43s | #proc 75 | #zombie 0 | #exit 5383 |
276
+CPU | sys 67% | user 31% | irq 2% | idle 0% | wait 0% |
277
+CPL | avg1 1.34 | avg5 1.05 | avg15 0.96 | csw 51346 | intr 10508 |
278
+MEM | tot 992.8M | free 211.5M | cache 470.0M | buff 87.2M | slab 164.7M |
279
+SWP | tot 0.0M | free 0.0M | | vmcom 207.6M | vmlim 496.4M |
280
+DSK | vda | busy 0% | read 0 | write 4 | avio 1.50 ms |
281
+NET | transport | tcpi 16 | tcpo 15 | udpi 0 | udpo 0 |
282
+NET | network | ipi 16 | ipo 15 | ipfrw 0 | deliv 16 |
283
+NET | eth0 ---- | pcki 16 | pcko 15 | si 1 Kbps | so 4 Kbps |
284
+
285
+ PID SYSCPU USRCPU VGROW RGROW RDDSK WRDSK ST EXC S CPU CMD 1/600
286
+12789 0.98s 0.40s 0K 0K 0K 336K -- - S 14% bash
287
+ 9 0.08s 0.00s 0K 0K 0K 0K -- - S 1% rcuos/0
288
+ 7024 0.03s 0.00s 0K 0K 0K 0K -- - S 0% apps.plugin
289
+ 7009 0.01s 0.01s 0K 0K 0K 4K -- - S 0% netdata
290
+```
291
+
292
+#### glances
293
+
294
+And the same is true for `glances`. The system runs at 100%, but `glances` reports only 17%
295
+per process utilization.
296
+
297
+Note also, that being a `python` program, `glances` uses 1.6% CPU while it runs.
298
+
299
+
300
+```
301
+localhost Uptime: 3 days, 21:42:00
302
+
303
+CPU [100.0%] CPU 100.0% MEM 23.7% SWAP 0.0% LOAD 1-core
304
+MEM [ 23.7%] user: 30.9% total: 993M total: 0 1 min: 1.18
305
+SWAP [ 0.0%] system: 67.8% used: 236M used: 0 5 min: 1.08
306
+ idle: 0.0% free: 757M free: 0 15 min: 1.00
307
+
308
+NETWORK Rx/s Tx/s TASKS 75 (90 thr), 1 run, 74 slp, 0 oth
309
+eth0 168b 2Kb
310
+eth1 0b 0b CPU% MEM% PID USER NI S Command
311
+lo 0b 0b 13.5 0.4 12789 root 0 S -bash
312
+ 1.6 2.2 7025 root 0 R /usr/bin/python /u
313
+DISK I/O R/s W/s 1.0 0.0 9 root 0 S rcuos/0
314
+vda1 0 4K 0.3 0.2 7024 netdata 0 S /usr/libexec/netda
315
+ 0.3 0.0 7 root 0 S rcu_sched
316
+FILE SYS Used Total 0.3 2.1 7009 netdata 0 S /usr/sbin/netdata
317
+/ (vda1) 1.56G 29.5G 0.0 0.0 17 root 0 S oom_reaper
318
+```
319
+
320
+#### why this happens?
321
+
322
+All the console tools report usage based on the processes found running *at the moment they
323
+examine the process tree*. So, they see just one `ls` command, which is actually very quick
324
+with minor CPU utilization. But the shell, is spawning hundreds of them, one after another
325
+(much like shell scripts do).
326
+
327
+#### what netdata reports?
328
+
329
+The total CPU utilization of the system:
330
+
331
+
332
+<br/>_**Figure 1**: The system overview section at netdata, just a few seconds after the command was run_
333
+
334
+And at the applications `apps.plugin` breaks down CPU usage per application:
335
+
336
+
337
+<br/>_**Figure 2**: The Applications section at netdata, just a few seconds after the command was run_
338
+
339
+So, the `ssh` session is using 95% CPU time.
340
+
341
+Why `ssh`?
342
+
343
+`apps.plugin` groups all processes based on its configuration file
344
+[`/etc/netdata/apps_groups.conf`](apps_groups.conf)
345
+(to edit it on your system run `/etc/netdata/edit-config apps_groups.conf`).
346
+The default configuration has nothing for `bash`, but it has for `sshd`, so netdata accumulates
347
+all ssh sessions to a dimension on the charts, called `ssh`. This includes all the processes in
348
+the process tree of `sshd`, **including the exited children**.
349
+
350
+> Distributions based on `systemd`, provide another way to get cpu utilization per user session
351
+> or service running: control groups, or cgroups, commonly used as part of containers
352
+> `apps.plugin` does not use these mechanisms. The process grouping made by `apps.plugin` works
353
+> on any Linux, `systemd` based or not.
354
+
355
+#### a more technical description of how netdata works
356
+
357
+netdata reads `/proc/<pid>/stat` for all processes, once per second and extracts `utime` and
358
+`stime` (user and system cpu utilization), much like all the console tools do.
359
+
360
+But it [also extracts `cutime` and `cstime`](https://github.com/netdata/netdata/blob/62596cc6b906b1564657510ca9135c08f6d4cdda/src/apps_plugin.c#L636-L642)
361
+that account the user and system time of the exit children of each process. By keeping a map in
362
+memory of the whole process tree, it is capable of assigning the right time to every process,
363
+taking into account all its exited children.
364
+
365
+It is tricky, since a process may be running for 1 hour and once it exits, its parent should not
366
+receive the whole 1 hour of cpu time in just 1 second - you have to subtract the cpu time that has
367
+been reported for it prior to this iteration.
368
+
369
+It is even trickier, because walking through the entire process tree takes some time itself. So,
370
+if you sum the CPU utilization of all processes, you might have more CPU time than the reported
371
+total cpu time of the system. netdata solves this, by adapting the per process cpu utilization to
372
+the total of the system. [Netdata adds charts that document this normalization](https://london.my-netdata.io/default.html#menu_netdata_submenu_apps_plugin).
collectors/cgroups.plugin/Makefile.am
+1
@@ -17,4 +17,5 @@ dist_plugins_SCRIPTS = \
17
18
dist_noinst_DATA = \
19
cgroup-name.sh.in \
20
+ README.md \
21
$(NULL)
collectors/cgroups.plugin/README.md
new
+187
@@ -0,0 +1,187 @@
1
+# cgroups.plugin
2
+
3
+You can monitor containers and virtual machines using **cgroups**.
4
+
5
+cgroups (or control groups), are a Linux kernel feature that provides accounting and resource usage limiting for processes. When cgroups are bundled with namespaces (i.e. isolation), they form what we usually call **containers**.
6
+
7
+cgroups are hierarchical, meaning that cgroups can contain child cgroups, which can contain more cgroups, etc. All accounting is reported (and resource usage limits are applied) also in a hierarchical way.
8
+
9
+To visualize cgroup metrics netdata provides configuration for cherry picking the cgroups of interest. By default (without any configuration) netdata should pick **systemd services**, all kinds of **containers** (lxc, docker, etc) and **virtual machines** spawn by managers that register them with cgroups (qemu, libvirt, etc).
10
+
11
+## configuring netdata for cgroups
12
+
13
+For each cgroup available in the system, netdata provides this configuration:
14
+
15
+```
16
+[plugin:cgroups]
17
+ enable cgroup XXX = yes | no
18
+```
19
+
20
+But it also provides a few patterns to provide a sane default (`yes` or `no`).
21
+
22
+Below we see, how this works.
23
+
24
+### how netdata finds the available cgroups
25
+
26
+Linux exposes resource usage reporting and provides dynamic configuration for cgroups, using virtual files (usually) under `/sys/fs/cgroup`. netdata reads `/proc/self/mountinfo` to detect the exact mount point of cgroups. netdata also allows manual configuration of this mount point, using these settings:
27
+
28
+```
29
+[plugin:cgroups]
30
+ check for new cgroups every = 10
31
+ path to /sys/fs/cgroup/cpuacct = /sys/fs/cgroup/cpuacct
32
+ path to /sys/fs/cgroup/blkio = /sys/fs/cgroup/blkio
33
+ path to /sys/fs/cgroup/memory = /sys/fs/cgroup/memory
34
+ path to /sys/fs/cgroup/devices = /sys/fs/cgroup/devices
35
+```
36
+
37
+netdata rescans these directories for added or removed cgroups every `check for new cgroups every` seconds.
38
+
39
+### hierarchical search for cgroups
40
+
41
+Since cgroups are hierarchical, for each of the directories shown above, netdata walks through the subdirectories recursively searching for cgroups (each subdirectory is another cgroup).
42
+
43
+For each of the directories found, netdata provides a configuration variable:
44
+
45
+```
46
+[plugin:cgroups]
47
+ search for cgroups under PATH = yes | no
48
+```
49
+
50
+To provide a sane default for this setting, netdata uses the following pattern list (patterns starting with `!` give a negative match and their order is important: the first matching a path will be used):
51
+
52
+```
53
+[plugin:cgroups]
54
+ search for cgroups in subpaths matching = !*/init.scope !*-qemu !/init.scope !/system !/systemd !/user !/user.slice *
55
+```
56
+
57
+So, we disable checking for **child cgroups** in systemd internal cgroups ([systemd services are monitored by netdata](https://github.com/netdata/netdata/wiki/monitoring-systemd-services)), user cgroups (normally used for desktop and remote user sessions), qemu virtual machines (child cgroups of virtual machines) and `init.scope`. All others are enabled.
58
+
59
+
60
+### enabled cgroups
61
+
62
+To check if the cgroup is enabled, netdata uses this setting:
63
+
64
+```
65
+[plugin:cgroups]
66
+ enable cgroup NAME = yes | no
67
+```
68
+
69
+To provide a sane default, netdata uses the following pattern list (it checks the pattern against the path of the cgroup):
70
+
71
+```
72
+[plugin:cgroups]
73
+ enable by default cgroups matching = !*/init.scope *.scope !*/vcpu* !*/emulator !*.mount !*.partition !*.service !*.slice !*.swap !*.user !/ !/docker !/libvirt !/lxc !/lxc/*/ns !/lxc/*/ns/* !/machine !/qemu !/system !/systemd !/user *
74
+```
75
+
76
+The above provides the default `yes` or `no` setting for the cgroup. However, there is an additional step. In many cases the cgroups found in the `/sys/fs/cgroup` hierarchy are just random numbers and in many cases these numbers are ephemeral: they change across reboots or sessions.
77
+
78
+So, we need to somehow map the paths of the cgroups to names, to provide consistent netdata configuration (i.e. there is no point to say `enable cgroup 1234 = yes | no`, if `1234` is a random number that changes over time - we need a name for the cgroup first, so that `enable cgroup NAME = yes | no` will be consistent).
79
+
80
+For this mapping netdata provides 2 configuration options:
81
+
82
+```
83
+[plugin:cgroups]
84
+ run script to rename cgroups matching = *.scope *docker* *lxc* *qemu* !/ !*.mount !*.partition !*.service !*.slice !*.swap !*.user *
85
+ script to get cgroup names = /usr/libexec/netdata/plugins.d/cgroup-name.sh
86
+```
87
+
88
+The whole point for the additional pattern list, is to limit the number of times the script will be called. Without this pattern list, the script might be called thousands of times, depending on the number of cgroups available in the system.
89
+
90
+The above pattern list is matched against the path of the cgroup. For matched cgroups, netdata calls the script [cgroup-name.sh](https://github.com/netdata/netdata/blob/master/collectors/cgroups.plugin/cgroup-name.sh.in) to get its name. This script queries `docker`, or applies heuristics to find give a name for the cgroup.
91
+
92
+## Monitoring systemd services
93
+
94
+netdata monitors **systemd services**. Example:
95
+
96
+
97
+
98
+Support per distribution:
99
+
100
+system|systemd services<br/>charts shown|`tree`<br/>`/sys/fs/cgroup`|comments
101
+:-------:|:-------:|:-------:|:------------
102
+Arch Linux|YES| |
103
+Gentoo|NO| |can be enabled, see below
104
+Ubuntu 16.04 LTS|YES| |
105
+Ubuntu 16.10|YES|[here](http://pastebin.com/PiWbQEXy)|
106
+Fedora 25|YES|[here](http://pastebin.com/ax0373wF)|
107
+Debian 8|NO| |can be enabled, see below
108
+AMI|NO|[here](http://pastebin.com/FrxmptjL)|not a systemd system
109
+Centos 7.3.1611|NO|[here](http://pastebin.com/SpzgezAg)|can be enabled, see below
110
+
111
+#### how to enable cgroup accounting on systemd systems that is by default disabled
112
+
113
+You can verify there is no accounting enabled, by running `systemd-cgtop`. The program will show only resources for cgroup ` / `, but all services will show nothing.
114
+
115
+To enable cgroup accounting, execute this:
116
+
117
+```sh
118
+sed -e 's|^#Default\(.*\)Accounting=.*$|Default\1Accounting=yes|g' /etc/systemd/system.conf >/tmp/system.conf
119
+```
120
+
121
+To see the changes it made, run this:
122
+
123
+```
124
+# diff /etc/systemd/system.conf /tmp/system.conf
125
+40,44c40,44
126
+< #DefaultCPUAccounting=no
127
+< #DefaultIOAccounting=no
128
+< #DefaultBlockIOAccounting=no
129
+< #DefaultMemoryAccounting=no
130
+< #DefaultTasksAccounting=yes
131
+---
132
+> DefaultCPUAccounting=yes
133
+> DefaultIOAccounting=yes
134
+> DefaultBlockIOAccounting=yes
135
+> DefaultMemoryAccounting=yes
136
+> DefaultTasksAccounting=yes
137
+```
138
+
139
+If you are happy with the changes, run:
140
+
141
+```sh
142
+# copy the file to the right location
143
+sudo cp /tmp/system.conf /etc/systemd/system.conf
144
+
145
+# restart systemd to take it into account
146
+sudo systemctl daemon-reexec
147
+```
148
+
149
+(`systemctl daemon-reload` does not reload the configuration of the server - so you have to execute `systemctl daemon-reexec`).
150
+
151
+Now, when you run `systemd-cgtop`, services will start reporting usage (if it does not, restart a service - any service - to wake it up). Refresh your netdata dashboard, and you will have the charts too.
152
+
153
+In case memory accounting is missing, you will need to enable it at your kernel, by appending the following kernel boot options and rebooting:
154
+
155
+```
156
+cgroup_enable=memory swapaccount=1
157
+```
158
+
159
+You can add the above, directly at the `linux` line in your `/boot/grub/grub.cfg` or appending them to the `GRUB_CMDLINE_LINUX` in `/etc/default/grub` (in which case you will have to run `update-grub` before rebooting). On DigitalOcean debian images you may have to set it at `/etc/default/grub.d/50-cloudimg-settings.cfg`.
160
+
161
+---
162
+
163
+## Monitoring ephemeral containers
164
+
165
+netdata monitors containers automatically when it is installed at the host, or when it is installed in a container that has access to the `/proc` and `/sys` filesystems of the host.
166
+
167
+netdata prior to v1.6 had 2 issues when such containers were monitored:
168
+
169
+1. network interface alarms where triggering when containers were stopped
170
+
171
+2. charts were never cleaned up, so after some time dozens of containers were showing up on the dashboard, and they were occupying memory.
172
+
173
+
174
+### the current netdata
175
+
176
+network interfaces and cgroups (containers) are now self-cleaned.
177
+
178
+So, when a network interface or container stops, netdata might log a few errors in error.log complaining about files it cannot find, but immediately:
179
+
180
+1. it will detect this is a removed container or network interface
181
+2. it will freeze/pause all alarms for them
182
+3. it will mark their charts as obsolete
183
+4. obsolete charts are not be offered on new dashboard sessions (so hit F5 and the charts are gone)
184
+5. existing dashboard sessions will continue to see them, but of course they will not refresh
185
+6. obsolete charts will be removed from memory, 1 hour after the last user viewed them (configurable with `[global].cleanup obsolete charts after seconds = 3600` (at netdata.conf).
186
+7. when obsolete charts are removed from memory they are also deleted from disk (configurable with `[global].delete obsolete charts files = yes`)
187
+
collectors/diskspace.plugin/README.md
+1
-1
@@ -1,4 +1,4 @@
1
-> for disks performance monitoring, see the `proc` plugin, [here](../linux-proc.plugin/#monitoring-disks-performance-with-netdata)
1
+> for disks performance monitoring, see the `proc` plugin, [here](../proc.plugin/#monitoring-disks-performance-with-netdata)
2
3
# diskspace.plugin
4
collectors/plugins.d/README.md
+1
-11
@@ -3,17 +3,6 @@
3
`plugins.d` is the netdata internal plugin that collects metrics
4
from external processes, thus allowing netdata to use **external plugins**.
5
6
-netdata supports plugins written in **any language**. The only requirement netdata has from its plugins, is to be able to print data at their output.
7
-
8
-Plugins can be written in the appropriate language for their job. For example:
9
-
10
-- You can collect data from JMX, using a java application
11
-- You can collect data from a REST API, using a node.js application
12
-- You can collect data from a system command, using a shell script
13
-- etc.
14
-
15
-Many of these languages can run their code efficiently, but they require a lot of resources when they are initialized. netdata suggests that plugins will be **initialized once and run forever** (until stopped by netdata). This way, the expensive part of their execution, their initialization, is eliminated.
16
-
6
## Provided External Plugins
7
8
plugin|language|O/S|description
@@ -33,6 +22,7 @@ Each of these modular plugins has each own methods for defining modules. Please
22
This plugin allows netdata to use **external plugins** for data collection:
23
24
1. external data collection plugins may be written in any computer language.
25
+
26
2. external data collection plugins may use O/S capabilities or `setuid` to
27
run with escalated privileges (compared to the netdata daemon).
28
The communication between the external plugin and netdata is unidirectional
collectors/proc.plugin/README.md
+2
-2
@@ -25,9 +25,9 @@
25
26
---
27
28
-# Monitoring Disks' Performance with netdata
28
+# Monitoring Disks
29
30
-> Live demo of disk monitoring at: **[http://london.netdata.rocks](http://london.netdata.rocks/#disk)**
30
+> Live demo of disk monitoring at: **[http://london.netdata.rocks](https://registry.my-netdata.io/#menu_disk)**
31
32
Performance monitoring for Linux disks is quite complicated. The main reason is the plethora of disk technologies available. There are many different hardware disk technologies, but there are even more **virtual disk** technologies that can provide additional storage features.
33
collectors/tc.plugin/README.md
+177
-3
@@ -1,9 +1,183 @@
1
## tc.plugin
2
3
+Live demo - **[see it in action here](https://registry.my-netdata.io/#menu_tc)** !
4
+
5
+
6
+
7
Netdata monitors `tc` QoS classes for all interfaces.
8
5
-If you also use [FireQOS](http://firehol.org/tutorial/fireqos-new-user/)) it will collect interface and class names.
9
+If you also use [FireQOS](http://firehol.org/tutorial/fireqos-new-user/) it will collect
10
+interface and class names.
11
+
12
+There is a [shell helper](tc-qos-helper.sh.in) for this (all parsing is done by the plugin
13
+in `C` code - this shell script is just a configuration for the command to run to get `tc` output).
14
+
15
+The source of the tc plugin is [here](plugin_tc.c). It is somewhat complex, because a state
16
+machine was needed to keep track of all the `tc` classes, including the pseudo classes tc
17
+dynamically creates.
18
+
19
+## Motivation
20
+
21
+One category of metrics missing in Linux monitoring, is bandwidth consumption for each open
22
+socket (inbound and outbound traffic). So, you cannot tell how much bandwidth your web server,
23
+your database server, your backup, your ssh sessions, etc are using.
24
+
25
+To solve this problem, the most *adventurous* Linux monitoring tools install kernel modules to
26
+capture all traffic, analyze it and provide reports per application. A lot of work, CPU intensive
27
+and with a great degree of risk (due to the kernel modules involved which might affect the
28
+stability of the whole system). Not to mention that such solutions are probably better suited
29
+for a core linux router in your network.
30
+
31
+Others use NFACCT, the netfilter accounting module which is already part of the Linux firewall.
32
+However, this would require configuring a firewall on every system you want to measure bandwidth.
33
+
34
+QoS monitoring attempts to solve this in a much cleaner way.
35
+
36
+## Introduction to QoS
37
+
38
+One of the features the Linux kernel has, but it is rarely used, is its ability to
39
+**apply QoS on traffic**. Even most interesting is that it can apply QoS to **both inbound and
40
+outbound traffic**.
41
+
42
+QoS is about 2 features:
43
+
44
+1. **Classify traffic**
45
+
46
+ Classification is the process of organizing traffic in groups, called **classes**.
47
+ Classification can evaluate every aspect of network packets, like source and destination ports,
48
+ source and destination IPs, netfilter marks, etc.
49
+
50
+ When you classify traffic, you just assign a label to it. For example **I call `web server`
51
+ traffic, the traffic from my server's tcp/80, tcp/443 and to my server's tcp/80, tcp/443,
52
+ while I call `web surfing` all other tcp/80 and tcp/443 traffic**. You can use any combinations
53
+ you like. There is no limit.
54
+
55
+2. **Apply traffic shaping rules to these classes**
56
+
57
+ Traffic shaping is used to control how network interface bandwidth should be shared among the
58
+ classes. Of course we are not interested for this feature to just monitor the traffic.
59
+ Classification will be enough for monitoring everything.
60
+
61
+The key reasons of applying QoS on all servers (even cloud ones) are:
62
+
63
+ - **ensure administrative tasks (like ssh, dns, etc) will always have a small but guaranteed
64
+ bandwidth.** QoS can guarantee that services like ssh, dns, ntp, etc will always have a small
65
+ supply of bandwidth. So, no matter what happens, you will be able to ssh to your server and
66
+ DNS will always work.
67
+
68
+ - **ensure other administrative tasks will not monopolize all the available bandwidth.**
69
+ Services like backups, file copies, database dumps, etc can easily monopolize all the
70
+ available bandwidth. It is common for example a nightly backup, or a huge file transfer
71
+ to negatively influence the end-user experience. QoS can fix that.
72
+
73
+ - **ensure each end-user connection will get a fair cut of the available bandwidth.**
74
+ Several QoS queuing disciplines in Linux do this automatically, without any configuration from you.
75
+ The result is that new sockets are favored over older ones, so that users will get a snappier
76
+ experience, while others are transferring large amounts of traffic.
77
+
78
+ - **protect the servers from DDoS attacks.**
79
+ When your system is under a DDoS attack, it will get a lot more bandwidth compared to the one it
80
+ can handle and probably your applications will crash. Setting a limit on the inbound traffic using
81
+ QoS, will protect your servers (throttle the requests) and depending on the size of the attack may
82
+ allow your legitimate users to access the server, while the attack is taking place.
83
+
84
+
85
+Once **traffic classification** is applied, netdata can visualize the bandwidth consumption per
86
+class in real-time (no configuration is needed for netdata - it will figure it out).
87
+
88
+QoS, is extremely light. You will configure it once, and this is it. It will not bother you again
89
+and it will not use any noticeable CPU resources, especially on application and database servers.
90
+
91
+## QoS in Linux? Have you lost your mind?
92
+
93
+Yes I know... but no, I have not!
94
+
95
+Of course, `tc` is probably **the most undocumented, complicated and unfriendly** command in Linux.
96
+
97
+For example, for matching a simple port range in `tc`, e.g. all the high ports, from 1025 to 65535
98
+inclusive, you have to match these:
99
+
100
+```
101
+1025/0xffff 1026/0xfffe 1028/0xfffc 1032/0xfff8 1040/0xfff0
102
+1056/0xffe0 1088/0xffc0 1152/0xff80 1280/0xff00 1536/0xfe00
103
+2048/0xf800 4096/0xf000 8192/0xe000 16384/0xc000 32768/0x8000
104
+```
105
+
106
+I know what you are thinking right now! **And I agree!**
107
+
108
+This is why I wrote **[FireQOS](https://firehol.org/tutorial/fireqos-new-user/)**, a tool to
109
+simplify QoS management in Linux.
110
+
111
+The **[FireHOL](https://firehol.org/)** package already distributes **[FireQOS](https://firehol.org/tutorial/fireqos-new-user/)**.
112
+Check the **[FireQOS tutorial](https://firehol.org/tutorial/fireqos-new-user/)**
113
+to learn how to write your own QoS configuration.
114
+
115
+With **[FireQOS](https://firehol.org/tutorial/fireqos-new-user/)**, it is **really simple for everyone
116
+to use QoS in Linux**. Just install the package `firehol`. It should already be available for your
117
+distribution. If not, check the **[FireHOL Installation Guide](https://firehol.org/installing/)**.
118
+After that, you will have the `fireqos` command.
119
+
120
+This is the file `/etc/firehol/fireqos.conf` we use at the netdata demo site:
121
+
122
+```sh
123
+ # configure the netdata ports
124
+ server_netdata_ports="tcp/19999"
125
+
126
+ interface eth0 world bidirectional ethernet balanced rate 50Mbit
127
+ class arp
128
+ match arp
129
+
130
+ class icmp
131
+ match icmp
132
+
133
+ class dns commit 1Mbit
134
+ server dns
135
+ client dns
136
+
137
+ class ntp
138
+ server ntp
139
+ client ntp
140
+
141
+ class ssh commit 2Mbit
142
+ server ssh
143
+ client ssh
144
+
145
+ class rsync commit 2Mbit max 10Mbit
146
+ server rsync
147
+ client rsync
148
+
149
+ class web_server commit 40Mbit
150
+ server http
151
+ server netdata
152
+
153
+ class client
154
+ client surfing
155
+
156
+ class nms commit 1Mbit
157
+ match input src 10.2.3.5
158
+```
159
+
160
+Nothing more is needed. You just run `fireqos start` to apply this configuration, restart netdata
161
+and you have real-time visualization of the bandwidth consumption of your applications. FireQOS is
162
+not a daemon. It will just convert the configuration to `tc` commands. It will run them and it will
163
+exit.
164
+
165
+**IMPORTANT**: If you copy this configuration to apply it to your system, please adapt the
166
+speeds - experiment in non-production environments to learn the tool, before applying it on
167
+your servers.
168
+
169
+And this is what you are going to get:
170
+
171
+
172
+
173
+---
174
+
175
+## More examples:
176
+
177
+This is QoS from a linux router. Check these features:
178
+
179
+1. It is real-time (per second updates)
180
+2. QoS really works in Linux - check that the `background` traffic is squeezed when `surfing` needs it.
181
7
-There is a [shell helper](tc-qos-helper.sh.in) for this (all parsing is done by the plugin in `C` code - this shell script is just a configuration for the command to run to get `tc` output).
182
+
183
9
-The source of the tc plugin is [here](plugin_tc.c). It is somewhat complex, because a state machine was needed to keep track of all the `tc` classes, including the pseudo classes tc dynamically creates.
configure.ac
+2
@@ -577,6 +577,7 @@ AC_CONFIG_FILES([
577
database/Makefile
578
diagrams/Makefile
579
health/Makefile
580
+ health/notifications/Makefile
581
libnetdata/Makefile
582
libnetdata/adaptive_resortable_list/Makefile
583
libnetdata/avl/Makefile
@@ -602,6 +603,7 @@ AC_CONFIG_FILES([
603
tests/Makefile
604
web/Makefile
605
web/api/Makefile
606
+ web/api/badges/Makefile
607
web/gui/Makefile
608
web/server/Makefile
609
web/server/single/Makefile
daemon/README.md
+445
@@ -0,0 +1,445 @@
1
+# Netdata daemon
2
+
3
+
4
+
5
+## Command line options
6
+
7
+Normally you don't need to supply any command line arguments to netdata.
8
+
9
+If you do though, they override the configuration equivalent options.
10
+
11
+To get a list of all command line parameters supported, run:
12
+
13
+```sh
14
+netdata -h
15
+```
16
+
17
+The program will print the supported command line parameters.
18
+
19
+The command line options of the netdata 1.10.0 version are the following:
20
+```
21
+
22
+ ^
23
+ |.-. .-. .-. .-. . netdata
24
+ | '-' '-' '-' '-' real-time performance monitoring, done right!
25
+ +----+-----+-----+-----+-----+-----+-----+-----+-----+-----+-----+-----+--->
26
+
27
+ Copyright (C) 2016-2017, Costa Tsaousis <costa@tsaousis.gr>
28
+ Released under GNU General Public License v3 or later.
29
+ All rights reserved.
30
+
31
+ Home Page : https://my-netdata.io
32
+ Source Code: https://github.com/netdata/netdata
33
+ Wiki / Docs: https://github.com/netdata/netdata/wiki
34
+ Support : https://github.com/netdata/netdata/issues
35
+ License : https://github.com/netdata/netdata/blob/master/LICENSE.md
36
+
37
+ Twitter : https://twitter.com/linuxnetdata
38
+ Facebook : https://www.facebook.com/linuxnetdata/
39
+
40
+
41
+ SYNOPSIS: netdata [options]
42
+
43
+ Options:
44
+
45
+ -c filename Configuration file to load.
46
+ Default: /etc/netdata/netdata.conf
47
+
48
+ -D Do not fork. Run in the foreground.
49
+ Default: run in the background
50
+
51
+ -h Display this help message.
52
+
53
+ -P filename File to save a pid while running.
54
+ Default: do not save pid to a file
55
+
56
+ -i IP The IP address to listen to.
57
+ Default: all IP addresses IPv4 and IPv6
58
+
59
+ -p port API/Web port to use.
60
+ Default: 19999
61
+
62
+ -s path Prefix for /proc and /sys (for containers).
63
+ Default: no prefix
64
+
65
+ -t seconds The internal clock of netdata.
66
+ Default: 1
67
+
68
+ -u username Run as user.
69
+ Default: netdata
70
+
71
+ -v Print netdata version and exit.
72
+
73
+ -V Print netdata version and exit.
74
+
75
+ -W options See Advanced options below.
76
+
77
+
78
+ Advanced options:
79
+
80
+ -W stacksize=N Set the stacksize (in bytes).
81
+
82
+ -W debug_flags=N Set runtime tracing to debug.log.
83
+
84
+ -W unittest Run internal unittests and exit.
85
+
86
+ -W set section option value
87
+ set netdata.conf option from the command line.
88
+
89
+ -W simple-pattern pattern string
90
+ Check if string matches pattern and exit.
91
+
92
+
93
+ Signals netdata handles:
94
+
95
+ - HUP Close and reopen log files.
96
+ - USR1 Save internal DB to disk.
97
+ - USR2 Reload health configuration.
98
+```
99
+
100
+## Log files
101
+
102
+netdata uses 3 log files:
103
+
104
+1. `error.log`
105
+2. `access.log`
106
+3. `debug.log`
107
+
108
+Any of them can be disabled by setting it to `/dev/null` or `none` in `netdata.conf`.
109
+By default `error.log` and `access.log` are enabled. `debug.log` is only enabled if
110
+debugging/tracing is also enabled (netdata needs to be compiled with debugging enabled).
111
+
112
+Log files are stored in `/var/log/netdata/` by default.
113
+
114
+#### error.log
115
+
116
+The `error.log` is the `stderr` of the netdata daemon and all external plugins run by netdata.
117
+
118
+So if any process, in the netdata process tree, writes anything to its standard error,
119
+it will appear in `error.log`.
120
+
121
+For most netdata programs (including standard external plugins shipped by netdata), the
122
+following lines may appear:
123
+
124
+tag|description
125
+:---:|:----
126
+`INFO`|Something important the user should know.
127
+`ERROR`|Something that might disable a part of netdata.<br/>The log line includes `errno` (if it is not zero).
128
+`FATAL`|Something prevented a program from running.<br/>The log line includes `errno` (if it is not zero) and the program exited.
129
+
130
+So, when auto-detection of data collection fail, `ERROR` lines are logged and the relevant modules
131
+are disabled, but the program continues to run.
132
+
133
+When a netdata program cannot run at all, a `FATAL` line is logged.
134
+
135
+#### access.log
136
+
137
+The `access.log` logs web requests. The format is:
138
+
139
+```txt
140
+DATE: ID: (sent/all = SENT_BYTES/ALL_BYTES bytes PERCENT_COMPRESSION%, prep/sent/total PREP_TIME/SENT_TIME/TOTAL_TIME ms): ACTION CODE URL
141
+```
142
+
143
+where:
144
+
145
+ - `ID` is the client ID. Client IDs are auto-incremented every time a client connects to netdata.
146
+ - `SENT_BYTES` is the number of bytes sent to the client, without the HTTP response header.
147
+ - `ALL_BYTES` is the number of bytes of the response, before compression.
148
+ - `PERCENT_COMPRESSION` is the percentage of traffic saved due to compression.
149
+ - `PREP_TIME` is the time in milliseconds needed to prepared the response.
150
+ - `SENT_TIME` is the time in milliseconds needed to sent the response to the client.
151
+ - `TOTAL_TIME` is the total time the request was inside netdata (from the first byte of the request to the last byte of the response).
152
+ - `ACTION` can be `filecopy`, `options` (used in CORS), `data` (API call).
153
+
154
+
155
+#### debug.log
156
+
157
+See [debugging](#debugging).
158
+
159
+
160
+## OOM Score
161
+
162
+netdata runs with `OOMScore = 1000`. This means netdata will be the first to be killed when your
163
+server runs out of memory.
164
+
165
+You can set netdata OOMScore in `netdata.conf`, like this:
166
+
167
+```
168
+[global]
169
+ OOM score = 1000
170
+```
171
+
172
+netdata logs its OOM score when it starts:
173
+
174
+```sh
175
+# grep OOM /var/log/netdata/error.log
176
+2017-10-15 03:47:31: netdata INFO : Adjusted my Out-Of-Memory (OOM) score from 0 to 1000.
177
+```
178
+
179
+#### OOM score and systemd
180
+
181
+netdata will not be able to lower its OOM Score below zero, when it is started as the `netdata`
182
+user (systemd case).
183
+
184
+To allow netdata control its OOM Score in such cases, you will need to edit
185
+`netdata.service` and set:
186
+
187
+```
188
+[Service]
189
+# The minimum netdata Out-Of-Memory (OOM) score.
190
+# netdata (via [global].OOM score in netdata.conf) can only increase the value set here.
191
+# To decrease it, set the minimum here and set the same or a higher value in netdata.conf.
192
+# Valid values: -1000 (never kill netdata) to 1000 (always kill netdata).
193
+OOMScoreAdjust=-1000
194
+```
195
+
196
+Run `systemctl daemon-reload` to reload these changes.
197
+
198
+The above, sets and OOMScore for netdata to `-1000`, so that netdata can increase it via
199
+`netdata.conf`.
200
+
201
+If you want to control it entirely via systemd, you can set in `netdata.conf`:
202
+
203
+```
204
+[global]
205
+ OOM score = keep
206
+```
207
+
208
+Using the above, whatever OOM Score you have set at `netdata.service` will be maintained by netdata.
209
+
210
+
211
+## netdata process scheduling policy
212
+
213
+By default netdata runs with the `idle` process scheduling policy, so that it uses CPU resources, only when there is idle CPU to spare. On very busy servers (or weak servers), this can lead to gaps on the charts.
214
+
215
+You can set netdata scheduling policy in `netdata.conf`, like this:
216
+
217
+```
218
+[global]
219
+ process scheduling policy = idle
220
+```
221
+
222
+You can use the following:
223
+
224
+policy|description
225
+:-----:|:--------
226
+`idle`|use CPU only when there is spare - this is lower than nice 19 - it is the default for netdata and it is so low that netdata will run in "slow motion" under extreme system load, resulting in short (1-2 seconds) gaps at the charts.
227
+`other`<br/>or<br/>`nice`|this is the default policy for all processes under Linux. It provides dynamic priorities based on the `nice` level of each process. Check below for setting this `nice` level for netdata.
228
+`batch`|This policy is similar to `other` in that it schedules the thread according to its dynamic priority (based on the `nice` value). The difference is that this policy will cause the scheduler to always assume that the thread is CPU-intensive. Consequently, the scheduler will apply a small scheduling penalty with respect to wake-up behavior, so that this thread is mildly disfavored in scheduling decisions.
229
+`fifo`|`fifo` can be used only with static priorities higher than 0, which means that when a `fifo` threads becomes runnable, it will always immediately preempt any currently running `other`, `batch`, or `idle` thread. `fifo` is a simple scheduling algorithm without time slicing.
230
+`rr`|a simple enhancement of `fifo`. Everything described above for `fifo` also applies to `rr`, except that each thread is allowed to run only for a maximum time quantum.
231
+`keep`<br/>or<br/>`none`|do not set scheduling policy, priority or nice level - i.e. keep running with whatever it is set already (e.g. by systemd).
232
+
233
+For more information see `man sched`.
234
+
235
+### scheduling priority for `rr` and `fifo`
236
+
237
+Once the policy is set to one of `rr` or `fifo`, the following will appear:
238
+
239
+```
240
+[global]
241
+ process scheduling priority = 0
242
+```
243
+
244
+These priorities are usually from 0 to 99. Higher numbers make the process more important.
245
+
246
+### nice level for policies `other` or `batch`
247
+
248
+When the policy is set to `other`, `nice`, or `batch`, the following will appear:
249
+
250
+```
251
+[global]
252
+ process nice level = 19
253
+```
254
+
255
+## scheduling settings and systemd
256
+
257
+netdata will not be able to set its scheduling policy and priority to more important values when it is started as the `netdata` user (systemd case).
258
+
259
+You can set these settings at `/etc/systemd/system/netdata.service`:
260
+
261
+```
262
+[Service]
263
+# By default netdata switches to scheduling policy idle, which makes it use CPU, only
264
+# when there is spare available.
265
+# Valid policies: other (the system default) | batch | idle | fifo | rr
266
+#CPUSchedulingPolicy=other
267
+
268
+# This sets the maximum scheduling priority netdata can set (for policies: rr and fifo).
269
+# netdata (via [global].process scheduling priority in netdata.conf) can only lower this value.
270
+# Priority gets values 1 (lowest) to 99 (highest).
271
+#CPUSchedulingPriority=1
272
+
273
+# For scheduling policy 'other' and 'batch', this sets the lowest niceness of netdata.
274
+# netdata (via [global].process nice level in netdata.conf) can only increase the value set here.
275
+#Nice=0
276
+```
277
+
278
+Run `systemctl daemon-reload` to reload these changes.
279
+
280
+Now, tell netdata to keep these settings, as set by systemd, by editing `netdata.conf` and setting:
281
+
282
+```
283
+[global]
284
+ process scheduling policy = keep
285
+```
286
+
287
+Using the above, whatever scheduling settings you have set at `netdata.service` will be maintained by netdata.
288
+
289
+
290
+#### Example 1: netdata with nice -1 on non-systemd systems
291
+
292
+On a system that is not based on systemd, to make netdata run with nice level -1 (a little bit higher to the default for all programs), edit netdata.conf and set:
293
+
294
+```
295
+[global]
296
+ process scheduling policy = other
297
+ process nice level = -1
298
+```
299
+
300
+then execute this to restart netdata:
301
+
302
+```sh
303
+sudo service netdata restart
304
+```
305
+
306
+
307
+#### Example 2: netdata with nice -1 on systemd systems
308
+
309
+On a system that is based on systemd, to make netdata run with nice level -1 (a little bit higher to the default for all programs), edit netdata.conf and set:
310
+
311
+```
312
+[global]
313
+ process scheduling policy = keep
314
+```
315
+
316
+edit /etc/systemd/system/netdata.service and set:
317
+
318
+```
319
+[Service]
320
+CPUSchedulingPolicy=other
321
+Nice=-1
322
+```
323
+
324
+then execute:
325
+
326
+```sh
327
+sudo systemctl daemon-reload
328
+sudo systemctl restart netdata
329
+```
330
+
331
+## virtual memory
332
+
333
+You may notice that netdata's virtual memory size, as reported by `ps` or `/proc/pid/status`
334
+(or even netdata's applications virtual memory chart) is unrealistically high.
335
+
336
+For example, it may be reported to be 150+MB, even if the resident memory size is just 25MB.
337
+Similar values may be reported for netdata plugins too.
338
+
339
+Check this for example: A netdata installation with default settings on Ubuntu 16.04LTS.
340
+The top chart is **real memory used**, while the bottom one is **virtual memory**:
341
+
342
+
343
+
344
+#### why this happens?
345
+
346
+The system memory allocator allocates virtual memory arenas, per thread running.
347
+On Linux systems this defaults to 16MB per thread on 64 bit machines. So, if you get the
348
+difference between real and virtual memory and divide it by 16MB you will roughly get the
349
+number of threads running.
350
+
351
+The system does this for speed. Having a separate memory arena for each thread, allows the
352
+threads to run in parallel in multi-core systems, without any locks between them.
353
+
354
+This behaviour is system specific. For example, the chart above when running netdata on alpine
355
+linux (that uses **musl** instead of **glibc**) is this:
356
+
357
+
358
+
359
+#### can we do anything to lower it?
360
+
361
+Since netdata already uses minimal memory allocations while it runs (i.e. it adapts its memory
362
+on start, so that while repeatedly collects data it does not do memory allocations), it already
363
+instructs the system memory allocator to minimize the memory arenas for each thread. We have also
364
+added [2 configuration options](https://github.com/netdata/netdata/blob/5645b1ee35248d94e6931b64a8688f7f0d865ec6/src/main.c#L410-L418)
365
+to allow you tweak these settings.
366
+
367
+However, even if we instructed the memory allocator to use just one arena, it seems it allocates
368
+an arena per thread.
369
+
370
+netdata also supports `jemalloc` and `tcmalloc`, however both behave exactly the same to the
371
+glibc memory allocator in this aspect.
372
+
373
+#### Is this a problem?
374
+
375
+No, it is not.
376
+
377
+Linux reserves real memory (physical RAM) in pages (on x86 machines pages are 4KB each).
378
+So even if the system memory allocator is allocating huge amounts of virtual memory,
379
+only the 4KB pages that are actually used are reserving physical RAM. The **real memory** chart
380
+on netdata application section, shows the amount of physical memory these pages occupy(it
381
+accounts the whole pages, even if parts of them are actually used).
382
+
383
+
384
+## Debugging
385
+
386
+When you compile netdata with debugging:
387
+
388
+1. compiler optimizations for your CPU are disabled (netdata will run somewhat slower)
389
+
390
+2. a lot of code is added all over netdata, to log debug messages to `/var/log/netdata/debug.log`. However, nothing is printed by default. netdata allows you to select which sections of netdata you want to trace. Tracing is activated via the config option `debug flags`. It accepts a hex number, to enable or disable specific sections. You can find the options supported at [log.h](https://github.com/netdata/netdata/blob/master/libnetdata/log/log.h). They are the `D_*` defines. The value `0xffffffffffffffff` will enable all possible debug flags.
391
+
392
+Once netdata is compiled with debugging and tracing is enabled for a few sections, the file `/var/log/netdata/debug.log` will contain the messages.
393
+
394
+> Do not forget to disable tracing (`debug flags = 0`) when you are done tracing. The file `debug.log` can grow too fast.
395
+
396
+#### compiling netdata with debugging
397
+
398
+To compile netdata with debugging, use this:
399
+
400
+```sh
401
+# step into the netdata source directory
402
+cd /usr/src/netdata.git
403
+
404
+# run the installer with debugging enabled
405
+CFLAGS="-O1 -ggdb -DNETDATA_INTERNAL_CHECKS=1" ./netdata-installer.sh
406
+```
407
+
408
+The above will compile and install netdata with debugging info embedded. You can now use `debug flags` to set the section(s) you need to trace.
409
+
410
+#### debugging crashes
411
+
412
+We have made the most to make netdata crash free. If however, netdata crashes on your system, it would be very helpful to provide stack traces of the crash. Without them, is will be almost impossible to find the issue (the code base is quite large to find such an issue by just objerving it).
413
+
414
+To provide stack traces, **you need to have netdata compiled with debugging**. There is no need to enable any tracing (`debug flags`).
415
+
416
+Then you need to be in one of the following 2 cases:
417
+
418
+1. netdata crashes and you have a core dump
419
+
420
+2. you can reproduce the crash
421
+
422
+If you are not on these cases, you need to find a way to be (i.e. if your system does not produce core dumps, check your distro documentation to enable them).
423
+
424
+#### netdata crashes and you have a core dump
425
+
426
+> you need to have netdata compiled with debugging info for this to work (check above)
427
+
428
+Run the following command and post the output on a github issue.
429
+
430
+```sh
431
+gdb $(which netdata) /path/to/core/dump
432
+```
433
+
434
+#### you can reproduce a netdata crash on your system
435
+
436
+> you need to have netdata compiled with debugging info for this to work (check above)
437
+
438
+Install the package `valgrind` and run:
439
+
440
+```sh
441
+valgrind $(which netdata) -D
442
+```
443
+
444
+netdata will start and it will be a lot slower. Now reproduce the crash and `valgrind` will dump on your console the stack trace. Open a new github issue and post the output.
445
+
database/README.md
+206
@@ -0,0 +1,206 @@
1
+# netdata database
2
+
3
+Although `netdata` does all its calculations using `long double`, it stores all values using
4
+a [custom-made 32-bit number](../libnetdata/storage_number/).
5
+
6
+So, for each dimension of a chart, netdata will need: `4 bytes for the value * the entries
7
+of its history`. It will not store any other data for each value in the time series database.
8
+Since all its values are stored in a time series with fixed step, the time each value
9
+corresponds can be calculated at run time, using the position of a value in the round robin database.
10
+
11
+The default history is 3.600 entries, thus it will need 14.4KB for each chart dimension.
12
+If you need 1.000 dimensions, they will occupy just 14.4MB.
13
+
14
+Of course, 3.600 entries is a very short history, especially if data collection frequency is set
15
+to 1 second. You will have just one hour of data.
16
+
17
+For a day of data and 1.000 dimensions, you will need: 86.400 seconds * 4 bytes * 1.000
18
+dimensions = 345MB of RAM.
19
+
20
+Currently the only option you have to lower this number is to use
21
+**[Memory Deduplication - Kernel Same Page Merging - KSM](#ksm)**.
22
+
23
+## Memory modes
24
+
25
+Currently netdata supports 5 memory modes:
26
+
27
+1. `ram`, data are purely in memory. Data are never saved on disk. This mode uses `mmap()` and
28
+ supports [KSM](#ksm).
29
+
30
+2. `save`, (the default) data are only in RAM while netdata runs and are saved to / loaded from
31
+ disk on netdata restart. It also uses `mmap()` and supports [KSM](#ksm).
32
+
33
+3. `map`, data are in memory mapped files. This works like the swap. Keep in mind though, this
34
+ will have a constant write on your disk. When netdata writes data on its memory, the Linux kernel
35
+ marks the related memory pages as dirty and automatically starts updating them on disk.
36
+ Unfortunately we cannot control how frequently this works. The Linux kernel uses exactly the
37
+ same algorithm it uses for its swap memory. Check below for additional information on running a
38
+ dedicated central netdata server. This mode uses `mmap()` but does not support [KSM](#ksm).
39
+
40
+4. `none`, without a database (collected metrics can only be streamed to another netdata).
41
+
42
+5. `alloc`, like `ram` but it uses `calloc()` and does not support [KSM](#ksm). This mode is the
43
+ fallback for all others except `none`.
44
+
45
+You can select the memory mode by editing netdata.conf and setting:
46
+
47
+```
48
+[global]
49
+ # ram, save (the default, save on exit, load on start), map (swap like)
50
+ memory mode = save
51
+
52
+ # the directory where data are saved
53
+ cache directory = /var/cache/netdata
54
+```
55
+
56
+## Running netdata in embedded devices
57
+
58
+Embedded devices usually have very limited RAM resources available.
59
+
60
+There are 2 settings for you to tweak:
61
+
62
+1. `update every`, which controls the data collection frequency
63
+2. `history`, which controls the size of the database in RAM
64
+
65
+By default `update every = 1` and `history = 3600`. This gives you an hour of data with per
66
+second updates.
67
+
68
+If you set `update every = 2` and `history = 1800`, you will still have an hour of data, but
69
+collected once every 2 seconds. This will **cut in half** both CPU and RAM resources consumed
70
+by netdata. Of course experiment a bit. On very weak devices you might have to use
71
+`update every = 5` and `history = 720` (still 1 hour of data, but 1/5 of the CPU and RAM resources).
72
+
73
+You can also disable [data collection plugins](../collectors) you don't need.
74
+Disabling such plugins will also free both CPU and RAM resources.
75
+
76
+## running a dedicated central netdata server
77
+
78
+netdata allows streaming data between netdata nodes. This allows us to have a central netdata
79
+server that will maintain the entire database for all nodes, and will also run health checks/alarms
80
+for all nodes.
81
+
82
+For this central netdata, memory size can be a problem. Fortunately, netdata supports several
83
+memory modes. What is interesting for this setup is `memory mode = map`.
84
+
85
+In this mode, the database of netdata is stored in memory mapped files. netdata continues to read
86
+and write the database in memory, but the kernel automatically loads and saves memory pages from/to
87
+disk.
88
+
89
+**We suggest _not_ to use this mode on nodes that run other applications.** There will always be
90
+dirty memory to be synced and this syncing process may influence the way other applications work.
91
+This mode however is ideal when we need a central netdata server that would normally need huge
92
+amounts of memory. Using memory mode `map` we can overcome all memory restrictions.
93
+
94
+There are a few kernel options that provide finer control on the way this syncing works. But before
95
+explaining them, a brief introduction of how netdata database works is needed.
96
+
97
+For each chart, netdata maps the following files:
98
+
99
+1. `chart/main.db`, this is the file that maintains chart information. Every time data are collected
100
+ for a chart, this is updated.
101
+
102
+2. `chart/dimension_name.db`, this is the file for each dimension. At its beginning there is a
103
+ header, followed by the round robin database where metrics are stored.
104
+
105
+So, every time netdata collects data, the following pages will become dirty:
106
+
107
+1. the chart file
108
+2. the header part of all dimension files
109
+3. if the collected metrics are stored far enough in the dimension file, another page will
110
+ become dirty, for each dimension
111
+
112
+Each page in Linux is 4KB. So, with 200 charts and 1000 dimensions, there will be 1200 to 2200 4KB
113
+pages dirty pages every second. Of course 1200 of them will always be dirty (the chart header and
114
+the dimensions headers) and 1000 will be dirty for about 1000 seconds (4 bytes per metric, 4KB per
115
+page, so 1000 seconds, or 16 minutes per page).
116
+
117
+Hopefully, the Linux kernel does not sync all these data every second. The frequency they are
118
+synced is controlled by `/proc/sys/vm/dirty_expire_centisecs` or the
119
+`sysctl` `vm.dirty_expire_centisecs`. The default on most systems is 3000 (30 seconds).
120
+
121
+On a busy server centralizing metrics from 20+ servers you will experience this:
122
+
123
+
124
+
125
+As you can see, there is quite some stress (this is `iowait`) every 30 seconds.
126
+
127
+A simple solution is to increase this time to 10 minutes (60000). This is the same system
128
+with this setting in 10 minutes:
129
+
130
+
131
+
132
+Of course, setting this to 10 minutes means that data on disk might be up to 10 minutes old if you
133
+get an abnormal shutdown.
134
+
135
+There are 2 more options to tweak:
136
+
137
+1. `dirty_background_ratio`, by default `10`.
138
+2. `dirty_ratio`, by default `20`.
139
+
140
+These control the amount of memory that should be dirty for disk syncing to be triggered.
141
+On dedicated netdata servers, you can use: `80` and `90` respectively, so that all RAM is given
142
+to netdata.
143
+
144
+With these settings, you can expect a little `iowait` spike once every 10 minutes and in case
145
+of system crash, data on disk will be up to 10 minutes old.
146
+
147
+
148
+
149
+To have these settings automatically applied on boot, create the file `/etc/sysctl.d/netdata-memory.conf` with these contents:
150
+
151
+```
152
+vm.dirty_expire_centisecs = 60000
153
+vm.dirty_background_ratio = 80
154
+vm.dirty_ratio = 90
155
+vm.dirty_writeback_centisecs = 0
156
+```
157
+
158
+## KSM
159
+
160
+Netdata offers all its round robin database to kernel for deduplication.
161
+
162
+In the past KSM has been criticized for consuming a lot of CPU resources.
163
+Although this is true when KSM is used for deduplicating certain applications, it is not true with
164
+netdata, since the netdata memory is written very infrequently (if you have 24 hours of metrics in
165
+netdata, each byte at the in-memory database will be updated just once per day).
166
+
167
+KSM is a solution that will provide 60+% memory savings to netdata.
168
+
169
+#### Enable KSM in kernel
170
+
171
+You need to run a kernel compiled with:
172
+
173
+```sh
174
+CONFIG_KSM=y
175
+```
176
+
177
+When KSM is enabled at the kernel is just available for the user to enable it.
178
+
179
+So, if you build a kernel with `CONFIG_KSM=y` you will just get a few files in `/sys/kernel/mm/ksm`. Nothing else happens. There is no performance penalty (apart I guess from the memory this code occupies into the kernel).
180
+
181
+The files that `CONFIG_KSM=y` offers include:
182
+
183
+- `/sys/kernel/mm/ksm/run` by default `0`. You have to set this to `1` for the kernel to spawn `ksmd`.
184
+- `/sys/kernel/mm/ksm/sleep_millisecs`, by default `20`. The frequency ksmd should evaluate memory for deduplication.
185
+- `/sys/kernel/mm/ksm/pages_to_scan`, by default `100`. The amount of pages ksmd will evaluate on each run.
186
+
187
+So, by default `ksmd` is just disabled. It will not harm performance and the user/admin can control the CPU resources he/she is willing `ksmd` to use.
188
+
189
+#### Run `ksmd` kernel daemon
190
+
191
+To activate / run `ksmd` you need to run:
192
+
193
+```sh
194
+echo 1 >/sys/kernel/mm/ksm/run
195
+echo 1000 >/sys/kernel/mm/ksm/sleep_millisecs
196
+```
197
+
198
+With these settings ksmd does not even appear in the running process list (it will run once per second and evaluate 100 pages for de-duplication).
199
+
200
+Put the above lines in your boot sequence (`/etc/rc.local` or equivalent) to have `ksmd` run at boot.
201
+
202
+## Monitoring Kernel Memory de-duplication performance
203
+
204
+Netdata will create charts for kernel memory de-duplication performance, like this:
205
+
206
+
health/Makefile.am
+3
-15
@@ -3,26 +3,14 @@
3
AUTOMAKE_OPTIONS = subdir-objects
4
MAINTAINERCLEANFILES = $(srcdir)/Makefile.in
5
6
-CLEANFILES = \
7
- alarm-notify.sh \
8
- $(NULL)
9
-
10
-include $(top_srcdir)/build/subst.inc
11
-SUFFIXES = .in
12
-
13
-dist_libconfig_DATA = \
14
- health_alarm_notify.conf \
15
- health_email_recipients.conf \
6
+SUBDIRS = \
7
+ notifications \
8
$(NULL)
9
18
-dist_plugins_SCRIPTS = \
19
- alarm-notify.sh \
20
- alarm-email.sh \
21
- alarm-test.sh \
10
+CLEANFILES = \
11
$(NULL)
12
13
dist_noinst_DATA = \
25
- alarm-notify.sh.in \
14
README.md \
15
$(NULL)
16
health/README.md
+657
@@ -0,0 +1,657 @@
1
+
2
+# Health monitoring
3
+
4
+Each netdata node runs an independent thread evaluating health monitoring checks.
5
+This thread has lock free access to the database, so that it can operate as a watchdog.
6
+
7
+Health checks (alarms) are attached to netdata charts, allowing netdata to automatically
8
+activate an alarm as soon as a chart is created. This is very important for
9
+netdata, since many charts are dynamically created during runtime (for example, the
10
+chart tracking network interface packet drops, is automatically created on the first
11
+packet dropped).
12
+
13
+Netdata also supports alarm **templates**, so that an alarm can be attached to all
14
+the charts of the same context (i.e. all network interfaces, or all disks, or all mysql servers, etc.)
15
+
16
+Each alarm can execute a single query to the database using statistical algorithms against past data,
17
+but alarms can be combined. So, if you need 2 queries in the database, you can combine
18
+2 alarms together (both will run a query to the database, and the results can be combined).
19
+
20
+Each alarm has unlimited access to all the metrics collected. So, a single alarm can
21
+use expressions combining the latest value of any number of metrics.
22
+
23
+## Health configuration reference
24
+
25
+Stock netdata health configuration is in `/usr/lib/netdata/conf.d/health.d`.
26
+These files can be overwritten by copying them and editing them in `/etc/netdata/health.d`
27
+(run `/etc/netdata/edit-config` to edit them).
28
+
29
+In `/etc/netdata/health.d` you can also put any number of files (in any number of sub-directories)
30
+with a suffix `.conf` to have them processed by netdata.
31
+
32
+Health configuration can be reloaded at any time, without restarting netdata.
33
+Just send netdata the SIGUSR2 signal, like this:
34
+
35
+```sh
36
+killall -USR2 netdata
37
+```
38
+
39
+### Entities in the health files
40
+
41
+There are 2 entities:
42
+
43
+1. **alarms**, which are attached to specific charts, and
44
+
45
+2. **templates**, which define rules that should be applied to all charts having a
46
+ specific `context`. You can use this feature to apply **alarms** to all disks,
47
+ all network interfaces, all mysql databases, all nginx web servers, etc.
48
+
49
+Both of these entities have exactly the same format and feature set.
50
+The only difference is the label `alarm` or `template`.
51
+
52
+netdata supports overriding **templates** with **alarms**.
53
+For example, when a template is defined for a set of charts, an alarm with exactly the
54
+same name attached to the same chart the template matches, will have higher precedence
55
+(i.e. netdata will use the alarm on this chart and prevent the template from being applied
56
+to it).
57
+
58
+### The format
59
+
60
+The following lines are parsed.
61
+
62
+#### alarm line `alarm` or `template`
63
+
64
+This line starts an alarm or alarm template.
65
+
66
+```
67
+alarm: NAME
68
+```
69
+
70
+or
71
+
72
+```
73
+template: NAME
74
+```
75
+
76
+This line has to be first on each alarm or template.
77
+`NAME` is anything you would like to name it (the only symbols allowed are `.` and `_`).
78
+
79
+---
80
+
81
+#### alarm line `on`
82
+
83
+This line defines the data the alarm should be attached to.
84
+
85
+For alarms:
86
+
87
+```
88
+on: CHART
89
+```
90
+
91
+For `CHART` you can use a chart `id` or `name` of the chart, as shown on the dashboard.
92
+
93
+For alarm templates:
94
+
95
+```
96
+on: CONTEXT
97
+```
98
+
99
+`CONTEXT` is the template of a chart. For example the charts `mysql_local.net` and
100
+`mysql_server2.net` have the same context: `mysql.net`. So, you can use this to apply
101
+alarms to all `mysql.net` charts.
102
+
103
+To find the `CONTEXT` of a chart hover over its date, above the legend. A tooltip will
104
+appear with this format `plugin:nodule, context`. For example, the bandwidth chart of
105
+a network interface says:
106
+
107
+```
108
+proc:/proc/dev/dev, net.net
109
+```
110
+
111
+So, `plugin = proc`, `module = /proc/net/dev` and `context = net.net`.
112
+
113
+---
114
+
115
+#### alarm line `os`
116
+
117
+This alarm or template will be used only if the O/S of the host loading it, matches this
118
+pattern list. The value is a space separated list of simple patterns (use `*` as wildcard,
119
+prefix with `!` for a negative match, order is important).
120
+
121
+```
122
+os: linux freebsd macos
123
+```
124
+
125
+---
126
+
127
+#### alarm line `hosts`
128
+
129
+This alarm or template will be used only if the hostname of the host loading it, matches
130
+this pattern list. The value is a space separated list of simple patterns (use `*` as wildcard,
131
+prefix with `!` for a negative match, order is important).
132
+
133
+```
134
+hosts: server1 server2 database* !redis3 redis*
135
+```
136
+
137
+The above says: use this alarm on all hosts named `server1`, `server2`, `database*`, and
138
+all `redis*` except `redis3`.
139
+
140
+This is useful when you centralize metrics from multiple hosts, to one netdata.
141
+
142
+---
143
+
144
+#### alarm line `families`
145
+
146
+This line is only used in alarm templates. It filters the charts. So, if you need to create
147
+an alarm template for a few of a kind of chart (a few of your disks, or a few of your network
148
+interfaces, or a few your mysql servers, etc), you can create an alarm template that would
149
+normally be applied to all of them, and filter them by family.
150
+
151
+The format is:
152
+
153
+```
154
+families: SIMPLE PATTERN LIST
155
+```
156
+
157
+Simple patterns list is a lists of space separated patterns. Use ` * ` as wildcard and ` ! `
158
+for a negative match. Processing is left to right, and on the first hit (positive or negative),
159
+processing stops.
160
+
161
+So. `families: *` means, match anything, while `families: !bad*pattern* *` means anything
162
+except `bad*pattern*` (where `*` is a wildcard to match any sequence of characters).
163
+
164
+The family of a chart is usually the submenu of the netdata dashboard it appears.
165
+
166
+---
167
+
168
+#### alarm line `lookup`
169
+
170
+This lines makes a database lookup to find a value. This result of this lookup is available as `$this`.
171
+
172
+The format is:
173
+
174
+```
175
+lookup: METHOD AFTER [at BEFORE] [every DURATION] [OPTIONS] [of DIMENSIONS]
176
+```
177
+
178
+Everything is the same with [badges](../web/api/badges/). In short:
179
+
180
+- `METHOD` is one of `average`, `min`, `max`, `sum`, `incremental-sum`.
181
+ This is required.
182
+
183
+- `AFTER` is a relative number of seconds, but it also accepts a single letter for changing
184
+ the units, like `-1s` = 1 second in the past, `-1m` = 1 minute in the past, `-1h` = 1 hour
185
+ in the past, `-1d` = 1 day in the past. You need a negative number (i.e. how far in the past
186
+ to look for the value). **This is required**.
187
+
188
+- `at BEFORE` is by default 0 and is not required. Using this you can define the end of the
189
+ lookup. So data will be evaluated between `AFTER` and `BEFORE`.
190
+
191
+- `every DURATION` sets the updated frequency of the lookup (supports single letter units as
192
+ above too).
193
+
194
+- `OPTIONS` is a space separated list of `percentage`, `absolute`, `min2max`, `unaligned`,
195
+ `match-ids`, `match-names`. Check the badges documentation for more info.
196
+
197
+- `of DIMENSIONS` is optional and has to be the last parameter. Dimensions have to be separated
198
+ by `,` or `|`. The space characters found in dimensions will be kept as-is (a few dimensions
199
+ have spaces in their names). This accepts netdata simple patterns and the `match-ids` and
200
+ `match-names` options affect the searches for dimensions.
201
+
202
+The result of the lookup will be available as `$this` and `$NAME` in expressions.
203
+The timestamps of the timeframe evaluated by the database lookup is available as variables
204
+`$after` and `$before` (both are unix timestamps).
205
+
206
+---
207
+
208
+#### alarm line `calc`
209
+
210
+This expression is evaluated just after the `lookup` (if any). Its purpose is to apply some
211
+calculation before using the value looked up from the db.
212
+
213
+You can also have an expression without a lookup, using other variables that are available.
214
+
215
+The result of the calculation will be available as `$this` in warning and critical expressions
216
+(overwriting the `lookup` one).
217
+
218
+Format:
219
+
220
+```
221
+calc: EXPRESSION
222
+```
223
+
224
+Check [Expressions](#expressions) for more information.
225
+
226
+---
227
+
228
+#### alarm line `every`
229
+
230
+Sets the update frequency of this alarm. This is the same to the `every DURATION` given
231
+in the `lookup` lines.
232
+
233
+Format:
234
+
235
+```
236
+every: DURATION
237
+```
238
+
239
+`DURATION` accepts `s` for seconds, `m` is minutes, `h` for hours, `d` for days.
240
+
241
+---
242
+
243
+#### alarm lines `green` and `red`
244
+
245
+Set the green and red thresholds of a chart. Both are available as `$green` and `$red` in
246
+expressions. If multiple alarms define different thresholds, the ones defined by the first
247
+alarm will be used. These will eventually visualized on the dashboard, so only one set of
248
+them is allowed. If you need multiple sets of them in different alarms, use absolute numbers
249
+instead of `$red` and `$green`.
250
+
251
+Format:
252
+
253
+```
254
+green: NUMBER
255
+red: NUMBER
256
+```
257
+
258
+---
259
+
260
+#### alarm lines `warn` and `crit`
261
+
262
+These expressions should evaluate to true or false (alternatively non-zero or zero).
263
+They trigger the alarm. Both are optional.
264
+
265
+Format:
266
+
267
+```
268
+warn: EXPRESSION
269
+crit: EXPRESSION
270
+```
271
+Check [Expressions](#expressions) for more information.
272
+
273
+---
274
+
275
+#### alarm line `to`
276
+
277
+This will be the first parameter of the script to be executed when the alarm switches status.
278
+Its meaning is left up to the `exec` script.
279
+
280
+The default `exec` script, `alarm-notify.sh`, uses this field as a space separated list of roles,
281
+which are then consulted to find the exact recipients per notification method.
282
+
283
+Format:
284
+
285
+```
286
+to: ROLE1 ROLE2 ROLE3 ...
287
+```
288
+
289
+---
290
+
291
+#### alarm line `exec`
292
+
293
+The script that will be executed when the alarm changes status.
294
+
295
+Format:
296
+
297
+```
298
+exec: SCRIPT
299
+```
300
+
301
+The default `SCRIPT` is netdata's `alarm-notify.sh`, which supports all the notifications
302
+methods netdata supports, including custom hooks.
303
+
304
+---
305
+
306
+#### alarm line `delay`
307
+
308
+This is used to provide optional hysteresis settings for the notifications, to defend
309
+against notification floods. These settings do not affect the actual alarm - only the time
310
+the `exec` script is executed.
311
+
312
+Format:
313
+
314
+```
315
+delay: [[[up U] [down D] multiplier M] max X]
316
+```
317
+
318
+- `up U` defines the delay to be applied to a notification for an alarm that raised its status
319
+ (i.e. CLEAR to WARNING, CLEAR to CRITICAL, WARNING to CRITICAL). For example, `up 10s`, the
320
+ notification for this event will be sent 10 seconds after the actual event. This is used in
321
+ hope the alarm will get back to its previous state within the duration given. The default `U`
322
+ is zero.
323
+
324
+- `down D` defines the delay to be applied to a notification for an alarm that moves to lower
325
+ state (i.e. CRITICAL to WARNING, CRITICAL to CLEAR, WARNING to CLEAR). For example, `down 1m`
326
+ will delay the notification by 1 minute. This is used to prevent notifications for flapping
327
+ alarms. The default `D` is zero.
328
+
329
+- `mutliplier M` multiplies `U` and `D` when an alarm changes state, while a notification is
330
+ delayed. The default multiplier is `1.0`.
331
+
332
+- `max X` defines the maximum absolute notification delay an alarm may get. The default `X`
333
+ is `max(U * M, D * M)` (i.e. the max duration of `U` or `D` multiplied once with `M`).
334
+
335
+ Example:
336
+
337
+ `delay: up 10s down 15m multiplier 2 max 1h`
338
+
339
+ The time is `00:00:00` and the status of the alarm is CLEAR.
340
+
341
+ time of event|new status|delay|notification will be sent|why
342
+ -------------|----------|:---:|-------------------------|---
343
+ 00:00:01 | WARNING | `up 10s` | 00:00:11 |first state switch
344
+ 00:00:05 | CLEAR | `down 15m x2`| 00:30:05 |the alarm changes state while a notification is delayed, so it was multiplied
345
+ 00:00:06 | WARNING | `up 10s x2 x2` | 00:00:26 |multiplied twice
346
+ 00:00:07|CLEAR|`down 15m x2 x2 x2`|00:45:07|multiplied 3 times.
347
+
348
+ So:
349
+ - `U` and `D` are multiplied by `M` every time the alarm changes state (any state, not just
350
+ their matching one) and a delay is in place.
351
+ - All are reset to their defaults when the alarm switches state without a delay in place.
352
+
353
+---
354
+
355
+### Expressions
356
+
357
+netdata has an internal [infix expression parser](../libnetdata/eval).
358
+This parses expressions and creates an internal structure that allows fast execution of them.
359
+
360
+These operators are supported `+`, `-`, `*`, `/`, `<`, `<=`, `<>`, `!=`, `>`, `>=`, `&&`, `||`,
361
+`!`, `AND`, `OR`, `NOT`. Boolean operators result in either `1` (true) or `0` (false).
362
+
363
+The conditional evaluation operator `?` is supported too. Using this operator IF-THEN-ELSE
364
+conditional statements can be specified. The format is: `(condition) ? (true expression) :
365
+(false expression)`. So, netdata will first evaluate the `condition` and based on the result
366
+will either evaluate `true expression` or `false expression`.
367
+Example: `($this > 0) ? ($avail * 2) : ($used / 2)`.
368
+Nested such expressions are also supported (i.e. `true expression` and `false expression` can
369
+contain conditional evaluations).
370
+
371
+Expressions also support the `abs()` function.
372
+
373
+Expressions can have variables. Variables start with `$`. Check below for more information.
374
+
375
+There are two special values you can use:
376
+
377
+ - `nan`, for example `$this != nan` will check if the variable `this` is available.
378
+ A variable can be `nan` if the database lookup failed. All calculations (i.e. addition,
379
+ multiplication, etc) with a `nan` result in a `nan`.
380
+
381
+ - `inf`, for example `$this != inf` will check if `this` is not infinite. A value or
382
+ variable can be infinite if divided by zero. All calculations (i.e. addition,
383
+ multiplication, etc) with a `inf` result in a `inf`.
384
+
385
+---
386
+
387
+### Special use of the conditional operator
388
+
389
+A common (but not necessarily obvious) use of the conditional evaluation operator is
390
+to provide [hysteresis](https://en.wikipedia.org/wiki/Hysteresis) around the critical
391
+or warning thresholds. This usage helps to avoid bogus messages resulting from small
392
+variations in the value when it is varying regularly but staying close to the threshold
393
+value, without needing to delay sending messages at all.
394
+
395
+An example of such usage from the default CPU usage alarms bundled with netdata is:
396
+
397
+```
398
+warn: $this > (($status >= $WARNING) ? (75) : (85))
399
+crit: $this > (($status == $CRITICAL) ? (85) : (95))
400
+```
401
+
402
+The above say:
403
+* If the alarm is currently a warning, then the threshold for being considered a warning
404
+ is 75, otherwise it's 85.
405
+
406
+* If the alarm is currently critical, then the threshold for being considered critical
407
+ is 85, otherwise it's 95.
408
+
409
+Which in turn, results in the following behavior:
410
+* While the value is rising, it will trigger a warning when it exceeds 85, and a critical
411
+ alert when it exceeds 95.
412
+
413
+* While the value is falling, it will return to a warning state when it goes below 85,
414
+ and a normal state when it goes below 75.
415
+
416
+* If the value is constantly varying between 80 and 90, then it will trigger a warning the
417
+ first time it goes above 85, but will remain a warning until it goes below 75 (or goes above 85).
418
+
419
+* If the value is constantly varying between 90 and 100, then it will trigger a critical alert
420
+ the first time it goes above 95, but will remain a critical alert goes below 85 (at which
421
+ point it will return to being a warning).
422
+
423
+---
424
+
425
+### Variables
426
+
427
+netdata supports 3 new internal indexes for variables that will be used in health monitoring:
428
+
429
+ - **chart local variables**. All the dimensions of the chart are exposed as local variables.
430
+ All chart alarms names are exposed as variables too.
431
+
432
+ Charts also define a few special variables:
433
+
434
+ - `$last_collected_t` is the unix timestamp of the last data collection
435
+ - `$collected_total_raw` is the sum of all the dimensions (their last collected values)
436
+ - `$update_every` is the update frequency of the chart
437
+ - `$green` and `$red` the threshold defined in alarms (these are per chart - the charts
438
+ inherits them from the the first alarm that defined them)
439
+
440
+ Chart dimensions define their last calculated (i.e. interpolated) value, exactly as
441
+ shown on the charts, but also a variable with their name and suffix `_raw` that resolves
442
+ to the last collected value - as collected and another with suffix `_last_collected_t`
443
+ that resolves to unix timestamp the dimension was last collected (there may be dimensions
444
+ that fail to be collected while others continue normally).
445
+
446
+ - **family variables**. Families are used to group charts together. For example all `eth0`
447
+ charts, have `family = eth0`. This index includes all local variables, but if there are
448
+ overlapping variables, only the first are exposed.
449
+
450
+ - **host variables**. All the dimensions of all charts, including all alarms, in fullname.
451
+ Fullname is `CHART.VARIABLE`, where `CHART` is either the chart id or the chart name (both
452
+ are supported).
453
+
454
+ - **special variables*** are:
455
+
456
+ - `this`, which is resolved to the value of the current alarm.
457
+
458
+ - `status`, which is resolved to the current status of the alarm (the current = the last
459
+ status, i.e. before the current database lookup and the evaluation of the `calc` line).
460
+ This values can be compared with `$REMOVED`, `$UNINITIALIZED`, `$UNDEFINED`, `$CLEAR`,
461
+ `$WARNING`, `$CRITICAL`. These values are incremental, ie. `$status > $CLEAL` works as
462
+ expected.
463
+
464
+ - `now`, which is resolved to current unix timestamp.
465
+
466
+You can find all the variables that can be used for a given chart, using
467
+`http://your.netdata.ip:19999/api/v1/alarm_variables?chart=NAME`.
468
+This will dump all the indexes from the chart's perspective.
469
+Example: [variables for the `system.cpu` chart of the registry](https://registry.my-netdata.io/api/v1/alarm_variables?chart=system.cpu).
470
+
471
+## Alarm Statuses
472
+
473
+Alarms can have the following statuses:
474
+
475
+ - `REMOVED` - the alarm has been deleted (this happens when a SIGUSR2 is sent to netdata
476
+ to reload health configuration)
477
+
478
+ - `UNINITIALIZED` - the alarm is not initialized yet
479
+
480
+ - `UNDEFINED` - the alarm failed to be calculated (i.e. the database lookup failed,
481
+ a division by zero occurred, etc)
482
+
483
+ - `CLEAR` - the alarm is not armed / raised (i.e. is OK)
484
+
485
+ - `WARNING` - the warning expression resulted in true or non-zero
486
+
487
+ - `CRITICAL` - the critical expression resulted in true or non-zero
488
+
489
+The external script will be called for all status changes.
490
+
491
+## Examples
492
+
493
+
494
+Check the **[health.d directory](health.d)** for all alarms shipped with netdata.
495
+
496
+Here are a few examples:
497
+
498
+### Example 1
499
+
500
+A simple check if an apache server is alive:
501
+
502
+```
503
+template: apache_last_collected_secs
504
+ on: apache.requests
505
+ calc: $now - $last_collected_t
506
+ every: 10s
507
+ warn: $this > ( 5 * $update_every)
508
+ crit: $this > (10 * $update_every)
509
+```
510
+
511
+The above checks that netdata is able to collect data from apache. In detail:
512
+
513
+```
514
+template: apache_last_collected_secs
515
+```
516
+
517
+The above defines a **template** named `apache_last_collected_secs`.
518
+The name is important since `$apache_last_collected_secs` resolves to the `calc` line.
519
+So, try to give something descriptive.
520
+
521
+```
522
+ on: apache.requests
523
+```
524
+
525
+The above applies the **template** to all charts that have `context = apache.requests`
526
+(i.e. all your apache servers).
527
+
528
+```
529
+ calc: $now - $last_collected_t
530
+```
531
+
532
+- `$now` is a standard variable that resolves to the current timestamp.
533
+
534
+- `$last_collected_t` is the last data collection timestamp of the chart.
535
+ So this calculation gives the number of seconds passed since the last data collection.
536
+
537
+```
538
+ every: 10s
539
+```
540
+
541
+The alarm will be evaluated every 10 seconds.
542
+
543
+```
544
+ warn: $this > ( 5 * $update_every)
545
+ crit: $this > (10 * $update_every)
546
+```
547
+
548
+If these result in non-zero or true, they trigger the alarm.
549
+
550
+- `$this` refers to the value of this alarm (i.e. the result of the `calc` line.
551
+ We could also use `$apache_last_collected_secs`.
552
+
553
+`$update_every` is the update frequency of the chart, in seconds.
554
+
555
+So, the warning condition checks if we have not collected data from apache for 5
556
+iterations and the critical condition checks for 10 iterations.
557
+
558
+### Example 2
559
+
560
+Check if any of the disks is critically low on disk space:
561
+
562
+```
563
+template: disk_full_percent
564
+ on: disk.space
565
+ calc: $used * 100 / ($avail + $used)
566
+ every: 1m
567
+ warn: $this > 80
568
+ crit: $this > 95
569
+```
570
+
571
+`$used` and `$avail` are the `used` and `avail` chart dimensions as shown on the dashboard.
572
+
573
+So, the `calc` line finds the percentage of used space. `$this` resolves to this percentage.
574
+
575
+### Example 3
576
+
577
+Predict if any disk will run out of space in the near future.
578
+
579
+We do this in 2 steps:
580
+
581
+Calculate the disk fill rate:
582
+
583
+```
584
+ template: disk_fill_rate
585
+ on: disk.space
586
+ lookup: max -1s at -30m unaligned of avail
587
+ calc: ($this - $avail) / (30 * 60)
588
+ every: 15s
589
+```
590
+
591
+In the `calc` line: `$this` is the result of the `lookup` line (i.e. the free space 30 minutes
592
+ago) and `$avail` is the current disk free space. So the `calc` line will either have a positive
593
+number of GB/second if the disk if filling up, or a negative number of GB/second if the disk is
594
+freeing up space.
595
+
596
+There is no `warn` or `crit` lines here. So, this template will just do the calculation and
597
+nothing more.
598
+
599
+Predict the hours after which the disk will run out of space:
600
+
601
+```
602
+ template: disk_full_after_hours
603
+ on: disk.space
604
+ calc: $avail / $disk_fill_rate / 3600
605
+ every: 10s
606
+ warn: $this > 0 and $this < 48
607
+ crit: $this > 0 and $this < 24
608
+```
609
+
610
+The `calc` line estimates the time in hours, we will run out of disk space. Of course, only
611
+positive values are interesting for this check, so the warning and critical conditions check
612
+for positive values and that we have enough free space for 48 and 24 hours respectively.
613
+
614
+Once this alarm triggers we will receive an email like this:
615
+
616
+
617
+
618
+### Example 4
619
+
620
+Check if any network interface is dropping packets:
621
+
622
+```
623
+template: 30min_packet_drops
624
+ on: net.drops
625
+ lookup: sum -30m unaligned absolute
626
+ every: 10s
627
+ crit: $this > 0
628
+```
629
+
630
+The `lookup` line will calculate the sum of the all dropped packets in the last 30 minutes.
631
+
632
+The `crit` line will issue a critical alarm if even a single packet has been dropped.
633
+
634
+Note that the drops chart does not exist if a network interface has never dropped a single packet.
635
+When netdata detects a dropped packet, it will add the chart and it will automatically attach this
636
+alarm to it.
637
+
638
+## Troubleshooting
639
+
640
+You can compile netdata with [debugging](../daemon#debugging) and then set in `netdata.conf`:
641
+
642
+```
643
+[global]
644
+ debug flags = 0x0000000000800000
645
+```
646
+
647
+Then check your `/var/log/netdata/debug.log`. It will show you how it works.
648
+Important: this will generate a lot of output in debug.log.
649
+
650
+You can find the context of charts by looking up the chart in either
651
+`http://your.netdata:19999/netdata.conf` or `http://your.netdata:19999/api/v1/charts`.
652
+
653
+You can find how netdata interpreted the expressions by examining the alarm at
654
+`http://your.netdata:19999/api/v1/alarms?all`. For each expression, netdata will return the
655
+expression as given in its config file, and the same expression with additional parentheses
656
+added to indicate the evaluation flow of the expression.
657
+
health/notifications/Makefile.am
new
+45
@@ -0,0 +1,45 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+AUTOMAKE_OPTIONS = subdir-objects
4
+MAINTAINERCLEANFILES = $(srcdir)/Makefile.in
5
+
6
+CLEANFILES = \
7
+ alarm-notify.sh \
8
+ $(NULL)
9
+
10
+include $(top_srcdir)/build/subst.inc
11
+SUFFIXES = .in
12
+
13
+dist_libconfig_DATA = \
14
+ health_alarm_notify.conf \
15
+ health_email_recipients.conf \
16
+ $(NULL)
17
+
18
+dist_plugins_SCRIPTS = \
19
+ alarm-notify.sh \
20
+ alarm-email.sh \
21
+ alarm-test.sh \
22
+ $(NULL)
23
+
24
+dist_noinst_DATA = \
25
+ alarm-notify.sh.in \
26
+ README.md \
27
+ $(NULL)
28
+
29
+include alerta/Makefile.inc
30
+include awssns/Makefile.inc
31
+include discord/Makefile.inc
32
+include email/Makefile.inc
33
+include flock/Makefile.inc
34
+include irc/Makefile.inc
35
+include kavenegar/Makefile.inc
36
+include messagebird/Makefile.inc
37
+include pagerduty/Makefile.inc
38
+include pushbullet/Makefile.inc
39
+include pushover/Makefile.inc
40
+include rocketchat/Makefile.inc
41
+include slack/Makefile.inc
42
+include syslog/Makefile.inc
43
+include telegram/Makefile.inc
44
+include twilio/Makefile.inc
45
+include web/Makefile.inc
health/notifications/README.md
new
+60
@@ -0,0 +1,60 @@
1
+# Netdata alarm notifications
2
+
3
+The `exec` line in health configuration defines an external script that will be called once
4
+the alarm is triggered. The default script is **[alarm-notify.sh](alarm-notify.sh.in)**.
5
+
6
+You can change the default script globally by editing `/etc/netdata/netdata.conf`.
7
+
8
+`alarm-notify.sh` is capable of sending notifications:
9
+
10
+- to multiple recipients
11
+- using multiple notification methods
12
+- filtering severity per recipient
13
+
14
+It uses **roles**. For example `sysadmin`, `webmaster`, `dba`, etc.
15
+
16
+Each alarm is assigned to one or more roles, using the `to` line of the alarm configuration.
17
+Then `alarm-notify.sh` uses its own configuration file `/etc/netdata/health_alarm_notify.conf`
18
+the default is [here](health_alarm_notify.conf)
19
+(to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`)
20
+to find the destination address of the notification for each method.
21
+
22
+Each role may have one or more destinations.
23
+
24
+So, for example the `sysadmin` role may send:
25
+
26
+1. emails to admin1@example.com and admin2@example.com
27
+2. pushover.net notifications to USERTOKENS `A`, `B` and `C`.
28
+3. pushbullet.com push notifications to admin1@example.com and admin2@example.com
29
+4. messages to slack.com channel `#alarms` and `#systems`.
30
+5. messages to Discord channels `#alarms` and `#systems`.
31
+
32
+## Configuration
33
+
34
+Edit [`/etc/netdata/health_alarm_notify.conf`](health_alarm_notify.conf)
35
+by running `/etc/netdata/edit-config health_alarm_notify.conf`:
36
+
37
+- settings per notification method:
38
+
39
+ all notification methods except email, require some configuration
40
+ (i.e. API keys, tokens, destination rooms, channels, etc).
41
+
42
+2. **recipients** per **role** per **notification method**
43
+
44
+## Testing Notifications
45
+
46
+You can run the following command by hand, to test alarms configuration:
47
+
48
+```sh
49
+# become user netdata
50
+su -s /bin/bash netdata
51
+
52
+# enable debugging info on the console
53
+export NETDATA_ALARM_NOTIFY_DEBUG=1
54
+
55
+# send test alarms to sysadmin
56
+/usr/libexec/netdata/plugins.d/alarm-notify.sh test
57
+
58
+# send test alarms to any role
59
+/usr/libexec/netdata/plugins.d/alarm-notify.sh test "ROLE"
60
+```
health/notifications/alarm-email.sh
renamed
health/notifications/alarm-notify.sh.in
renamed
health/notifications/alarm-test.sh
renamed
health/notifications/alerta/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ alerta/README.md \
10
+ alerta/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/alerta/README.md
new
+236
@@ -0,0 +1,236 @@
1
+# alerta.io notifications
2
+
3
+The alerta monitoring system is a tool used to consolidate and de-duplicate alerts from multiple sources for quick ‘at-a-glance’ visualisation. With just one system you can monitor alerts from many other monitoring tools on a single screen.
4
+
5
+
6
+
7
+When receiving alerts from multiple sources you can quickly become overwhelmed. With Alerta any alert with the same environment and resource is considered a duplicate if it has the same severity. If it has a different severity it is correlated so that you only see the most recent one. Awesome.
8
+
9
+main site http://www.alerta.io
10
+
11
+We can send Netadata alarms to Alerta so yo can see in one place alerts coming from many Netdata hosts or also from a multihost Netadata configuration.\
12
+The big advantage over other notifications method is that you have in a main view all active alarms with only las state, but you can also search history.
13
+
14
+## Setting up an Alerta server with Ubuntu 16.04
15
+
16
+Here we will set a basic Alerta server to test it with Netdata alerts.\
17
+More advanced configurations are out os scope of this tutorial.
18
+
19
+source: http://alerta.readthedocs.io/en/latest/gettingstarted/tutorial-1-deploy-alerta.html
20
+
21
+I recommend to set up the server in a separated server, VM or container.\
22
+If you have other Nginx or Apache server in your organization, I recommend to proxy to this new server.
23
+
24
+Set us as root for easiest working
25
+```
26
+sudo su
27
+cd
28
+```
29
+
30
+Install Mongodb https://docs.mongodb.com/manual/tutorial/install-mongodb-on-ubuntu/
31
+```
32
+apt-key adv --keyserver hkp://keyserver.ubuntu.com:80 --recv 2930ADAE8CAF5059EE73BB4B58712A2291FA4AD5
33
+echo "deb [ arch=amd64,arm64 ] https://repo.mongodb.org/apt/ubuntu xenial/mongodb-org/3.6 multiverse" | tee /etc/apt/sources.list.d/mongodb-org-3.6.list
34
+apt-get update
35
+apt-get install -y mongodb-org
36
+systemctl enable mongod
37
+systemctl start mongod
38
+systemctl status mongod
39
+```
40
+
41
+Install Nginx and Alerta uwsgi
42
+```
43
+apt-get install -y python-pip python-dev nginx
44
+pip install alerta-server uwsgi
45
+```
46
+
47
+Install web console
48
+```
49
+cd /var/www/html
50
+mkdir alerta
51
+cd alerta
52
+wget -q -O - https://github.com/alerta/angular-alerta-webui/tarball/master | tar zxf -
53
+mv alerta*/app/* .
54
+cd
55
+```
56
+## Services configuration
57
+
58
+Create a wsgi python file
59
+```
60
+nano /var/www/wsgi.py
61
+```
62
+fill with
63
+```
64
+from alerta import app
65
+```
66
+Create uWsgi configuration file
67
+```
68
+nano /etc/uwsgi.ini
69
+```
70
+fill with
71
+```
72
+[uwsgi]
73
+chdir = /var/www
74
+mount = /alerta/api=wsgi.py
75
+callable = app
76
+manage-script-name = true
77
+
78
+master = true
79
+processes = 5
80
+logger = syslog:alertad
81
+
82
+socket = /tmp/uwsgi.sock
83
+chmod-socket = 664
84
+uid = www-data
85
+gid = www-data
86
+vacuum = true
87
+
88
+die-on-term = true
89
+```
90
+Create a systemd configuration file
91
+```
92
+nano /etc/systemd/system/uwsgi.service
93
+```
94
+fill with
95
+```
96
+[Unit]
97
+Description=uWSGI service
98
+
99
+[Service]
100
+ExecStart=/usr/local/bin/uwsgi --ini /etc/uwsgi.ini
101
+
102
+[Install]
103
+WantedBy=multi-user.target
104
+```
105
+enable service
106
+```
107
+systemctl start uwsgi
108
+systemctl status uwsgi
109
+systemctl enable uwsgi
110
+```
111
+Configure nginx to serve Alerta as a uWsgi application on /alerta/api
112
+```
113
+nano /etc/nginx/sites-enabled/default
114
+```
115
+fill with
116
+```
117
+server {
118
+ listen 80 default_server;
119
+ listen [::]:80 default_server;
120
+
121
+ location /alerta/api { try_files $uri @alerta/api; }
122
+ location @alerta/api {
123
+ include uwsgi_params;
124
+ uwsgi_pass unix:/tmp/uwsgi.sock;
125
+ proxy_set_header Host $host:$server_port;
126
+ proxy_set_header X-Real-IP $remote_addr;
127
+ proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
128
+ }
129
+
130
+ location / {
131
+ root /var/www/html;
132
+ }
133
+}
134
+```
135
+restart nginx
136
+```
137
+service nginx restart
138
+```
139
+## Config web console
140
+```
141
+nano /var/www/html/config.js
142
+```
143
+fill with
144
+```
145
+'use strict';
146
+
147
+angular.module('config', [])
148
+ .constant('config', {
149
+ 'endpoint' : "/alerta/api",
150
+ 'provider' : "basic",
151
+ 'colors' : {},
152
+ 'severity' : {},
153
+ 'audio' : {}
154
+ });
155
+```
156
+
157
+## Config Alerta server
158
+
159
+source: http://alerta.readthedocs.io/en/latest/configuration.html
160
+
161
+Create a random string to use as SECRET_KEY
162
+```
163
+cat /dev/urandom | tr -dc A-Za-z0-9_\!\@\#\$\%\^\&\*\(\)-+= | head -c 32 && echo
164
+```
165
+will output something like
166
+```
167
+0pv8Bw7VKfW6avDAz_TqzYPme_fYV%7g
168
+```
169
+Edit alertad.conf
170
+```
171
+nano /etc/alertad.conf
172
+```
173
+fill with (take care about all single quotes)
174
+```
175
+BASE_URL='/alerta/api'
176
+AUTH_REQUIRED=True
177
+SECRET_KEY='0pv8Bw7VKfW6avDAz_TqzYPme_fYV%7g'
178
+ADMIN_USERS=['<here put you email for future login>']
179
+```
180
+
181
+restart
182
+```
183
+systemctl restart uwsgi
184
+```
185
+
186
+* go to console to http://yourserver/alerta/
187
+* go to Login -> Create an account
188
+* use your email for login so and administrative account will be created
189
+
190
+## create an API KEY
191
+
192
+You need an API KEY to send messages from any source.\
193
+To create an API KEY go to Configuration -> Api Keys\
194
+Then create a API KEY with write permisions.
195
+
196
+## configure Netdata to send alarms to Alerta
197
+
198
+On your system run:
199
+
200
+```
201
+/etc/netdata/edit-config health_alarm_notify.conf
202
+```
203
+
204
+and set
205
+
206
+```
207
+# enable/disable sending alerta notifications
208
+SEND_ALERTA="YES"
209
+
210
+# here set your alerta server API url
211
+# this is the API url you defined when installed Alerta server,
212
+# it is the same for all users. Do not include last slash.
213
+ALERTA_WEBHOOK_URL="http://yourserver/alerta/api"
214
+
215
+# Login with an administrative user to you Alerta server and create an API KEY
216
+# with write permissions.
217
+ALERTA_API_KEY="you last created API KEY"
218
+
219
+# you can define environments in /etc/alertad.conf option ALLOWED_ENVIRONMENTS
220
+# standard environments are Production and Development
221
+# if a role's recipients are not configured, a notification will be send to
222
+# this Environment (empty = do not send a notification for unconfigured roles):
223
+DEFAULT_RECIPIENT_ALERTA="Production"
224
+```
225
+
226
+## Test alarms
227
+
228
+We can test alarms with standard
229
+```
230
+sudo su -s /bin/bash netdata
231
+/opt/netdata/netdata-plugins/plugins.d/alarm-notify.sh test
232
+exit
233
+```
234
+But the problem is that Netdata will send 3 alarms, and because last alarm is "CLEAR" you will not se them in main Alerta page, you need to select to see "closed" alarma in top-right lookup.
235
+
236
+A little change in alarm-notify.sh that let us test each state one by one will be useful.
\ No newline at end of file
health/notifications/awssns/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ awssns/README.md \
10
+ awssns/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/awssns/README.md
new
+31
@@ -0,0 +1,31 @@
1
+# Amazon SNS notifications
2
+
3
+As part of it's AWS suite, Amazon provides a notification broker service called 'Simple Notification Service' or SNS. Amazon SNS works kind of similarly to Netdata's own notification system, allowing dispatch of a single notification to multiple subscribers of different types. Among other things, SNS supports sending notifications to:
4
+
5
+* Email addresses.
6
+* Mobile Phones via SMS.
7
+* HTTP or HTTPS web hooks.
8
+* AWS Lambda functions.
9
+* AWS SQS queues.
10
+* Mobile applications via push notifications.
11
+
12
+To get this working, you will need:
13
+
14
+* The Amazon Web Services CLI tools. Most distributions provide these with the package name `awscli`.
15
+* An actual home directory for the user you run Netdata as, instead of just using `/` as a home directory. Setup of this is distribution specific. `/var/lib/netdata` is the recommended directory (because the permissions will already be correct) if you are using a dedicated user (which is how most distributions work).
16
+* An Amazon SNS topic to send notifications to with one or more subscribers. The [Getting Started](https://docs.aws.amazon.com/sns/latest/dg/GettingStarted.html) section of the Amazon SNS documentation covers the basics of how to set this up. Make note of the Topic ARN when you create the topic.
17
+* While not mandatory, it is highly recommended to create a dedicated IAM user on your account for netdata to send notifications. This user needs to have programmatic access, and should only allow access to SNS. If you're really paranoid, you can create one for each system or group of systems.
18
+
19
+Once you have all the above, run the follwing command as the user netdata runs under:
20
+
21
+ aws configure
22
+
23
+THis will prompt you for the access key and secret key for accessing Amazon SNS (as well as the default region and output format, but you can leave those blank because we don't use them).
24
+
25
+Once that's done, you're ready to go and can specify the desired topic ARN as a recipient.
26
+
27
+Notes:
28
+
29
+ * Netdata's native email notification support is far better in almost all respects than it's support through Amazon SNS. If you want email notifications, use the native support, not SNS.
30
+ * If you need to change the notification format for SNS notifications, you can do so by specifying the format in `AWSSNS_MESSAGE_FORMAT` in the configuration. This variable supports all the same vairiables you can use in custom notifications.
31
+ * While Amazon SNS supports sending differently formatted messages for different delivery methods, netdata does not currently support this functionality.
\ No newline at end of file
health/notifications/discord/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ discord/README.md \
10
+ discord/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/discord/README.md
new
+44
@@ -0,0 +1,44 @@
1
+# Discordapp.com notifications
2
+
3
+This is what you will get:
4
+
5
+
6
+
7
+You need:
8
+
9
+1. The **incoming webhook URL** as given by Discord. Create a webhook by following the official [Discord documentation](https://support.discordapp.com/hc/en-us/articles/228383668-Intro-to-Webhooks). You can use the same on all your netdata servers (or you can have multiple if you like - your decision).
10
+2. One or more Discord channels to post the messages to.
11
+
12
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
13
+
14
+```
15
+###############################################################################
16
+# sending discord notifications
17
+
18
+# note: multiple recipients can be given like this:
19
+# "CHANNEL1 CHANNEL2 ..."
20
+
21
+# enable/disable sending discord notifications
22
+SEND_DISCORD="YES"
23
+
24
+# Create a webhook by following the official documentation -
25
+# https://support.discordapp.com/hc/en-us/articles/228383668-Intro-to-Webhooks
26
+DISCORD_WEBHOOK_URL="https://discordapp.com/api/webhooks/XXXXXXXXXXXXX/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
27
+
28
+# if a role's recipients are not configured, a notification will be send to
29
+# this discord channel (empty = do not send a notification for unconfigured
30
+# roles):
31
+DEFAULT_RECIPIENT_DISCORD="alarms"
32
+
33
+```
34
+
35
+You can define multiple channels like this: `alarms systems`.
36
+You can give different channels per **role** using these (at the same file):
37
+
38
+```
39
+role_recipients_discord[sysadmin]="systems"
40
+role_recipients_discord[dba]="databases systems"
41
+role_recipients_discord[webmaster]="marketing development"
42
+```
43
+
44
+The keywords `systems`, `databases`, `marketing`, `development` are discordapp.com channels (they should already exist within your discord server).
health/notifications/email/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ email/README.md \
10
+ email/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/email/README.md
new
+31
@@ -0,0 +1,31 @@
1
+# email notifications
2
+
3
+You need a working `sendmail` command for email alerts to work. Almost all MTAs provide a `sendmail` interface.
4
+
5
+netdata sends all emails as user `netdata`, so make sure your `sendmail` works for local users.
6
+
7
+email notifications look like this:
8
+
9
+
10
+
11
+## configuration
12
+
13
+To edit `health_alarm_notify.conf` on your system run `/etc/netdata/edit-config health_alarm_notify.conf`.
14
+
15
+You can configure recipients in [`/etc/netdata/health_alarm_notify.conf`](https://github.com/netdata/netdata/blob/99d44b7d0c4e006b11318a28ba4a7e7d3f9b3bae/conf.d/health_alarm_notify.conf#L101).
16
+
17
+You can also configure per role recipients [in the same file, a few lines below](https://github.com/netdata/netdata/blob/99d44b7d0c4e006b11318a28ba4a7e7d3f9b3bae/conf.d/health_alarm_notify.conf#L313).
18
+
19
+Changes to this file do not require netdata restart.
20
+
21
+You can test your configuration by issuing the commands:
22
+
23
+```sh
24
+# become user netdata
25
+sudo su -s /bin/bash netdata
26
+
27
+# send a test alarm
28
+/usr/libexec/netdata/plugins.d/alarm-notify.sh test [ROLE]
29
+```
30
+
31
+Where `[ROLE]` is the role you want to test. The default (if you don't give a `[ROLE]`) is `sysadmin`.
\ No newline at end of file
health/notifications/flock/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ flock/README.md \
10
+ flock/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/flock/README.md
new
+31
@@ -0,0 +1,31 @@
1
+# flock.com notifications
2
+
3
+This is what you will get:
4
+
5
+
6
+
7
+
8
+You need:
9
+
10
+The **incoming webhook URL** as given by flock.com. You can use the same on all your netdata servers (or you can have multiple if you like - your decision).
11
+
12
+Get them here: https://admin.flock.com/webhooks
13
+
14
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
15
+
16
+```
17
+###############################################################################
18
+# sending flock notifications
19
+
20
+# enable/disable sending pushover notifications
21
+SEND_FLOCK="YES"
22
+
23
+# Login to flock.com and create an incoming webhook.
24
+# You need only one for all your netdata servers.
25
+# Without it, netdata cannot send flock notifications.
26
+FLOCK_WEBHOOK_URL="https://api.flock.com/hooks/sendMessage/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
27
+
28
+# if a role recipient is not configured, no notification will be sent
29
+DEFAULT_RECIPIENT_FLOCK="alarms"
30
+
31
+```
\ No newline at end of file
health/notifications/health_alarm_notify.conf
renamed
health/notifications/health_email_recipients.conf
renamed
health/notifications/irc/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ irc/README.md \
10
+ irc/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/irc/README.md
new
+73
@@ -0,0 +1,73 @@
1
+# IRC notifications
2
+
3
+This is what you will get:
4
+
5
+IRCCloud web client:
6
+
7
+
8
+Irssi terminal client:
9
+
10
+
11
+
12
+You need:
13
+1. The `nc` utility. If you do not set the path, netdata will search for it in your system `$PATH`.
14
+
15
+Set the path for `nc` in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
16
+
17
+```
18
+#------------------------------------------------------------------------------
19
+# external commands
20
+#
21
+# The full path of the nc command.
22
+# If empty, the system $PATH will be searched for it.
23
+# If not found, irc notifications will be silently disabled.
24
+nc="/usr/bin/nc"
25
+
26
+```
27
+
28
+2. Αn `IRC_NETWORK` to which your preffered channels belong to.
29
+3. One or more channels ( `DEFAULT_RECIPIENT_IRC` ) to post the messages to.
30
+4. An `IRC_NICKNAME` and an `IRC_REALNAME` to identify in IRC.
31
+
32
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
33
+
34
+```
35
+#------------------------------------------------------------------------------
36
+# irc notification options
37
+#
38
+# irc notifications require only the nc utility to be installed.
39
+
40
+# multiple recipients can be given like this:
41
+# "<irc_channel_1> <irc_channel_2> ..."
42
+
43
+# enable/disable sending irc notifications
44
+SEND_IRC="YES"
45
+
46
+# if a role's recipients are not configured, a notification will not be sent.
47
+# (empty = do not send a notification for unconfigured roles):
48
+DEFAULT_RECIPIENT_IRC="#system-alarms"
49
+
50
+# The irc network to which the recipients belong. It must be the full network.
51
+IRC_NETWORK="irc.freenode.net"
52
+
53
+# The irc nickname which is required to send the notification. It must not be
54
+# an already registered name as the connection's MODE is defined as a 'guest'.
55
+IRC_NICKNAME="netdata-alarm-user"
56
+
57
+# The irc realname which is required in order to make the connection and is an
58
+# extra identifier.
59
+IRC_REALNAME="netdata-user"
60
+
61
+```
62
+
63
+You can define multiple channels like this: `#system-alarms #networking-alarms`.
64
+You can also filter the notifications like this: `#system-alarms|critical`.
65
+You can give different channels per **role** using these (at the same file):
66
+
67
+```
68
+role_recipients_irc[sysadmin]="#user-alarms #networking-alarms #system-alarms"
69
+role_recipients_irc[dba]="#databases-alarms"
70
+role_recipients_irc[webmaster]="#networking-alarms"
71
+```
72
+
73
+The keywords `#user-alarms`, `#networking-alarms`, `#system-alarms`, `#databases-alarms` are irc channels which belong to the specified IRC network.
\ No newline at end of file
health/notifications/kavenegar/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ kavenegar/README.md \
10
+ kavenegar/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/kavenegar/README.md
new
+39
@@ -0,0 +1,39 @@
1
+# Kavenegar notifications
2
+
3
+[Kavenegar](https://www.kavenegar.com/) as service for software developers, based in Iran, provides send and receive SMS, calling voice by using its APIs.
4
+
5
+Will look like this on your Android device:
6
+
7
+
8
+
9
+You will need:
10
+
11
+1. Signup and Login to kavenegar.com
12
+2. Get your APIKEY and Sender from http://panel.kavenegar.com/client/setting/account
13
+3. Fill in KAVENEGAR_API_KEY="" KAVENEGAR_SENDER=""
14
+4. Add the recipient phone numbers to DEFAULT_RECIPIENT_KAVENEGAR=""
15
+
16
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
17
+
18
+```
19
+###############################################################################
20
+# Kavenegar (kavenegar.com) SMS options
21
+
22
+# multiple recipients can be given like this:
23
+# "09155555555 09177777777"
24
+
25
+# enable/disable sending kavenegar SMS
26
+SEND_KAVENEGAR="YES"
27
+
28
+# to get an access key, after selecting and purchasing your desired service
29
+# at http://kavenegar.com/pricing.html
30
+# login to your account, go to your dashboard and my account are
31
+# https://panel.kavenegar.com/Client/setting/account from API Key
32
+# copy your api key. You can generate new API Key too.
33
+# You can find and select kevenegar sender number from this place.
34
+
35
+# Without an API key, netdata cannot send KAVENEGAR text messages.
36
+KAVENEGAR_API_KEY=""
37
+KAVENEGAR_SENDER=""
38
+DEFAULT_RECIPIENT_KAVENEGAR=""
39
+```
\ No newline at end of file
health/notifications/messagebird/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ messagebird/README.md \
10
+ messagebird/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/messagebird/README.md
new
+38
@@ -0,0 +1,38 @@
1
+
2
+Will look like this on your Android device:
3
+
4
+
5
+
6
+You will need:
7
+
8
+1. Signup and Login to messagebird.com
9
+2. Pick an SMS capable number after sign up to get some free credits
10
+3. Go to <https://www.messagebird.com/app/settings/developers/access>
11
+4. Create a new access key under 'API ACCESS (REST)' (you will want a live key)
12
+3. Fill in MESSAGEBIRD_ACCESS_KEY="XXXXXXXX" MESSAGEBIRD_NUMBER="+XXXXXXXXXXX"
13
+4. Add the recipient phone numbers to DEFAULT_RECIPIENT_MESSAGEBIRD="+XXXXXXXXXXX"
14
+
15
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
16
+
17
+```
18
+#------------------------------------------------------------------------------
19
+# Messagebird (messagebird.com) SMS options
20
+
21
+# multiple recipients can be given like this:
22
+# "+15555555555 +17777777777"
23
+
24
+# enable/disable sending messagebird SMS
25
+SEND_MESSAGEBIRD="YES"
26
+
27
+# to get an access key, create a free account at https://www.messagebird.com
28
+# verify and activate the account (no CC info needed)
29
+# login to your account and enter your phonenumber to get some free credits
30
+# to get the API key, click on 'API' in the sidebar, then 'API Access (REST)'
31
+# click 'Add access key' and fill in data (you want a live key to send SMS)
32
+
33
+# Without an access key, netdata cannot send Messagebird text messages.
34
+MESSAGEBIRD_ACCESS_KEY="XXXXXXXX"
35
+MESSAGEBIRD_NUMBER="XXXXXXX"
36
+DEFAULT_RECIPIENT_MESSAGEBIRD="XXXXXXX"
37
+
38
+```
health/notifications/pagerduty/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ pagerduty/README.md \
10
+ pagerduty/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/pagerduty/README.md
new
+34
@@ -0,0 +1,34 @@
1
+
2
+[PagerDuty](https://www.pagerduty.com/company/) is the enterprise incident resolution service that integrates with ITOps and DevOps monitoring stacks to improve operational reliability and agility. From enriching and aggregating events to correlating them into incidents, PagerDuty streamlines the incident management process by reducing alert noise and resolution times.
3
+
4
+Here is an example of a PagerDuty dashboard with netdata notifications:
5
+
6
+
7
+
8
+To have netdata send notifications to PagerDuty, you'll first need to set up a PagerDuty `Generic API` service and install the PagerDuty agent on the host running netdata. See the following guide for details:
9
+
10
+https://www.pagerduty.com/docs/guides/agent-install-guide/
11
+
12
+During the setup of the `Generic API` PagerDuty service, you'll obtain a `pagerduty service key`. Keep this **service key** handy.
13
+
14
+Once the PagerDuty agent is installed on your host and can send notifications from your host to your `Generic API` service on PagerDuty, add the **service key** to `DEFAULT_RECIPIENT_PD` in `health_alarm_notify.conf`:
15
+
16
+```
17
+#------------------------------------------------------------------------------
18
+# pagerduty.com notification options
19
+#
20
+# pagerduty.com notifications require the pagerduty agent to be installed and
21
+# a "Generic API" pagerduty service.
22
+# https://www.pagerduty.com/docs/guides/agent-install-guide/
23
+
24
+# multiple recipients can be given like this:
25
+# "<pd_service_key_1> <pd_service_key_2> ..."
26
+
27
+# enable/disable sending pagerduty notifications
28
+SEND_PD="YES"
29
+
30
+# if a role's recipients are not configured, a notification will be sent to
31
+# the "General API" pagerduty.com service that uses this service key.
32
+# (empty = do not send a notification for unconfigured roles):
33
+DEFAULT_RECIPIENT_PD="<service key>"
34
+```
health/notifications/pushbullet/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ pushbullet/README.md \
10
+ pushbullet/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/pushbullet/README.md
new
+42
@@ -0,0 +1,42 @@
1
+$ PushBullet notifications
2
+
3
+Will look like this on your browser:
4
+
5
+
6
+And like this on your Android device:
7
+
8
+
9
+
10
+
11
+You will need:
12
+
13
+1. Signup and Login to pushbullet.com
14
+2. Get your Access Token, go to https://www.pushbullet.com/#settings/account and create a new one
15
+3. Fill in the PUSHBULLET_ACCESS_TOKEN with that value
16
+4. Add the recipient emails to DEFAULT_RECIPIENT_PUSHBULLET
17
+!!PLEASE NOTE THAT IF THE RECIPIENT DOES NOT HAVE A PUSHBULLET ACCOUNT, PUSHBULLET SERVICE WILL SEND AN EMAIL!!
18
+
19
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
20
+
21
+```
22
+###############################################################################
23
+# pushbullet (pushbullet.com) push notification options
24
+
25
+# multiple recipients can be given like this:
26
+# "user1@email.com user2@mail.com"
27
+
28
+# enable/disable sending pushbullet notifications
29
+SEND_PUSHBULLET="YES"
30
+
31
+# Signup and Login to pushbullet.com
32
+# To get your Access Token, go to https://www.pushbullet.com/#settings/account
33
+# And create a new access token
34
+# Then just set the recipients emails
35
+# Please note that the if the email in the DEFAULT_RECIPIENT_PUSHBULLET does
36
+# not have a pushbullet account, the pushbullet service will send an email
37
+# to that address instead
38
+
39
+# Without an access token, netdata cannot send pushbullet notifications.
40
+PUSHBULLET_ACCESS_TOKEN="o.Sometokenhere"
41
+DEFAULT_RECIPIENT_PUSHBULLET="admin1@example.com admin3@somemail.com"
42
+```
\ No newline at end of file
health/notifications/pushover/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ pushover/README.md \
10
+ pushover/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/pushover/README.md
new
+17
@@ -0,0 +1,17 @@
1
+
2
+# PushOver notifications
3
+
4
+pushover.net allows you to receive push notifications on your mobile phone. The service seems free for up to 7.500 messages per month.
5
+
6
+netdata will send warning messages with priority `0` and critical messages with priority `1`. pushover.net allows you to select do-not-disturb hours. The way this is configured, critical notifications will ring and vibrate your phone, even during the do-not-disturb-hours. All other notifications will be delivered silently.
7
+
8
+You need:
9
+
10
+1. APP TOKEN. You can use the same on all your netdata servers.
11
+2. USER TOKEN for each user you are going to send notifications to. This is the actual recipient of the notification.
12
+
13
+The configuration is like above (slack messages).
14
+
15
+pushover.net notifications look like this:
16
+
17
+
\ No newline at end of file
health/notifications/rocketchat/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ rocketchat/README.md \
10
+ rocketchat/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/rocketchat/README.md
new
+46
@@ -0,0 +1,46 @@
1
+# Rocket.Chat notifications
2
+
3
+This is what you will get:
4
+
5
+You need:
6
+
7
+1. The **incoming webhook URL** as given by RocketChat. You can use the same on all your netdata servers (or you can have multiple if you like - your decision).
8
+2. One or more channels to post the messages to.
9
+
10
+Get them here: https://rocket.chat/docs/administrator-guides/integrations/index.html#how-to-create-a-new-incoming-webhook
11
+
12
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
13
+
14
+```
15
+#------------------------------------------------------------------------------
16
+# rocketchat (rocket.chat) global notification options
17
+
18
+# multiple recipients can be given like this:
19
+# "CHANNEL1 CHANNEL2 ..."
20
+
21
+# enable/disable sending rocketchat notifications
22
+SEND_ROCKETCHAT="YES"
23
+
24
+# Login to rocket.chat and create an incoming webhook. You need only one for all
25
+# your netdata servers (or you can have one for each of your netdata).
26
+# Without it, netdata cannot send rocketchat notifications.
27
+ROCKETCHAT_WEBHOOK_URL="<your_incoming_webhook_url>"
28
+
29
+# if a role's recipients are not configured, a notification will be send to
30
+# this rocketchat channel (empty = do not send a notification for unconfigured
31
+# roles).
32
+DEFAULT_RECIPIENT_ROCKETCHAT="monitoring_alarms"
33
+
34
+```
35
+
36
+You can define multiple channels like this: `alarms systems`.
37
+You can give different channels per **role** using these (at the same file):
38
+
39
+```
40
+role_recipients_rocketchat[sysadmin]="systems"
41
+role_recipients_rocketchat[dba]="databases systems"
42
+role_recipients_rocketchat[webmaster]="marketing development"
43
+```
44
+
45
+The keywords `systems`, `databases`, `marketing`, `development` are RocketChat channels (they should already exist).
46
+Both public and private channels can be used, even if they differ from the channel configured in yout RocketChat incomming webhook.
health/notifications/slack/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ slack/README.md \
10
+ slack/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/slack/README.md
new
+45
@@ -0,0 +1,45 @@
1
+# Slack.com notifications
2
+
3
+This is what you will get:
4
+
5
+
6
+You need:
7
+
8
+1. The **incoming webhook URL** as given by slack.com. You can use the same on all your netdata servers (or you can have multiple if you like - your decision).
9
+2. One or more channels to post the messages to.
10
+
11
+Get them here: https://api.slack.com/incoming-webhooks
12
+
13
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
14
+
15
+```
16
+###############################################################################
17
+# sending slack notifications
18
+
19
+# note: multiple recipients can be given like this:
20
+# "CHANNEL1 CHANNEL2 ..."
21
+
22
+# enable/disable sending pushover notifications
23
+SEND_SLACK="YES"
24
+
25
+# Login to slack.com and create an incoming webhook.
26
+# You need only one for all your netdata servers.
27
+# Without it, netdata cannot send slack notifications.
28
+SLACK_WEBHOOK_URL="https://hooks.slack.com/services/XXXXXXXX/XXXXXXXX/XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX"
29
+
30
+# if a role recipient is not configured, a notification will be send to
31
+# this slack channel:
32
+DEFAULT_RECIPIENT_SLACK="alarms"
33
+
34
+```
35
+
36
+You can define multiple channels like this: `alarms systems`.
37
+You can give different channels per **role** using these (at the same file):
38
+
39
+```
40
+role_recipients_slack[sysadmin]="systems"
41
+role_recipients_slack[dba]="databases systems"
42
+role_recipients_slack[webmaster]="marketing development"
43
+```
44
+
45
+The keywords `systems`, `databases`, `marketing`, `development` are slack.com channels (they should already exist in slack).
health/notifications/syslog/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ syslog/README.md \
10
+ syslog/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/syslog/README.md
new
+23
@@ -0,0 +1,23 @@
1
+# syslog notifications
2
+
3
+You need a working `logger` command for this to work. This is the case on pretty much every Linux system in existence, and most BSD systems.
4
+
5
+Logged messages will look like this:
6
+
7
+ netdata WARNING on hostname at Tue Apr 3 09:00:00 EDT 2018: disk_space._ out of disk space time = 5h
8
+
9
+## configuration
10
+
11
+System log targets are configured as recipients in [`/etc/netdata/health_alarm_notify.conf`](https://github.com/netdata/netdata/blob/36bedc044584dea791fd29455bdcd287c3306cb2/conf.d/health_alarm_notify.conf#L534) (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`).
12
+
13
+You can als configure per-role targets in the same file a bit further down.
14
+
15
+Targets are defined as follows:
16
+
17
+ [[facility.level][@host[:port]]/]prefix
18
+
19
+`prefix` defines what the log messages are prefixed with. By default, all lines are prefixed with 'netdata'.
20
+
21
+The `facility` and `level` are the standard syslog facility and level options, for more info on them see your local `logger` and `syslog` documentation. By default, netdata will log to the `local6` facility, with a log level dependent on the type of message (`crit` for CRITICAL, `warning` for WARNING, and `info` for everything else).
22
+
23
+You can configure sending directly to remote log servers by specifying a host (and optionally a port). However, this has a somewhat high overhead, so it is much preferred to use your local syslog daemon to handle the forwarding of messages to remote systems (pretty much all of them allow at least simple forwarding, and most of the really popular ones support complex queueing and routing of messages to remote log servers).
health/notifications/telegram/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ telegram/README.md \
10
+ telegram/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/telegram/README.md
new
+19
@@ -0,0 +1,19 @@
1
+# Telegram.org notifications
2
+
3
+[Telegram](https://telegram.org/) is a messaging app with a focus on speed and security, it’s super-fast, simple and free. You can use Telegram on all your devices at the same time — your messages sync seamlessly across any number of your phones, tablets or computers.
4
+
5
+With Telegram, you can send messages, photos, videos and files of any type (doc, zip, mp3, etc), as well as create groups for up to 30,000 people or channels for broadcasting to unlimited audiences. You can write to your phone contacts and find people by their usernames. As a result, Telegram is like SMS and email combined — and can take care of all your personal or business messaging needs.
6
+
7
+netdata will send warning messages without vibration.
8
+
9
+You need:
10
+
11
+1. A bot token. To get one, contact the [@BotFather](https://t.me/BotFather) bot and send the command `/newbot`. Follow the instructions.
12
+2. A chat id for every chat you want to send messages to. Contact the [@myidbot](https://t.me/myidbot) bot and send the command `/getid` to get your personal chat id or invite him into a group and issue the same command to get the group chat id.
13
+3. Start a conversation with your bot or invite him into a group you want to sent messages to.
14
+
15
+See slack for configuration.
16
+
17
+Telegram messages look like this:
18
+
19
+
health/notifications/twilio/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ twilio/README.md \
10
+ twilio/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/twilio/README.md
new
+40
@@ -0,0 +1,40 @@
1
+# Twilio.com notifications
2
+
3
+Will look like this on your Android device:
4
+
5
+
6
+
7
+You will need:
8
+
9
+1. Signup and Login to twilio.com
10
+2. Pick an SMS capable number during sign up.
11
+3. Get your SID, and Token from <https://www.twilio.com/console>
12
+3. Fill in TWILIO_ACCOUNT_SID="XXXXXXXX" TWILIO_ACCOUNT_TOKEN="XXXXXXXXX" TWILIO_NUMBER="+XXXXXXXXXXX"
13
+4. Add the recipient phone numbers to DEFAULT_RECIPIENT_TWILIO="+XXXXXXXXXXX"
14
+
15
+!!PLEASE NOTE THAT IF YOUR ACCOUNT IS A TRIAL ACCOUNT YOU WILL ONLY BE ABLE TO SEND NOTIFICATIONS TO THE NUMBER YOU SIGNED UP WITH
16
+
17
+Set them in `/etc/netdata/health_alarm_notify.conf` (to edit it on your system run `/etc/netdata/edit-config health_alarm_notify.conf`), like this:
18
+
19
+```
20
+###############################################################################
21
+# Twilio (twilio.com) SMS options
22
+
23
+# multiple recipients can be given like this:
24
+# "+15555555555 +17777777777"
25
+
26
+# enable/disable sending twilio SMS
27
+SEND_TWILIO="YES"
28
+
29
+# Signup for free trial and select a SMS capable Twilio Number
30
+# To get your Account SID and Token, go to https://www.twilio.com/console
31
+# Place your sid, token and number below.
32
+# Then just set the recipients' phone numbers.
33
+# The trial account is only allowed to use the number specified when set up.
34
+
35
+# Without an account sid and token, netdata cannot send Twilio text messages.
36
+TWILIO_ACCOUNT_SID="xxxxxxxxx"
37
+TWILIO_ACCOUNT_TOKEN="xxxxxxxxxx"
38
+TWILIO_NUMBER="xxxxxxxxxxx"
39
+DEFAULT_RECIPIENT_TWILIO="+15555555555"
40
+```
health/notifications/web/Makefile.inc
new
+12
@@ -0,0 +1,12 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+# THIS IS NOT A COMPLETE Makefile
4
+# IT IS INCLUDED BY ITS PARENT'S Makefile.am
5
+# IT IS REQUIRED TO REFERENCE ALL FILES RELATIVE TO THE PARENT
6
+
7
+# install these files
8
+dist_noinst_DATA += \
9
+ web/README.md \
10
+ web/Makefile.inc \
11
+ $(NULL)
12
+
health/notifications/web/README.md
new
+6
@@ -0,0 +1,6 @@
1
+# Dashboard notifications
2
+
3
+The netdata dashboard shows HTML notifications, when it is open.
4
+
5
+Such web notifications look like this:
6
+
libnetdata/storage_number/README.md
+10
@@ -0,0 +1,10 @@
1
+# netdata storage number
2
+
3
+Although `netdata` does all its calculations using `long double`, it stores all values using
4
+a **custom-made 32-bit number**.
5
+
6
+This custom-made number can store in 29 bits values from `-167772150000000.0` to `167772150000000.0`
7
+with a precision of 0.00001 (yes, it's a floating point number, meaning that higher integer values
8
+have less decimal precision) and 3 bits for flags.
9
+
10
+This provides an extremely optimized memory footprint with just 0.0001% max accuracy loss.
registry/README.md
+152
@@ -0,0 +1,152 @@
1
+# netdata registry
2
+
3
+Netdata registry implements the `my-netdata` menu on netdata dashboards.
4
+The `my-netdata` menu lists the netdata servers you have visited.
5
+
6
+## Why?
7
+
8
+netdata provides distributed monitoring.
9
+
10
+Traditional monitoring solutions centralize all the data to provide unified dashboards across all servers. Before netdata, this was the standard practice. However it has a few issues:
11
+
12
+1. due to the resources required, the number of metrics collected is limited.
13
+2. for the same reason, the data collection frequency is not that high, at best it will be once every 10 or 15 seconds, at worst every 5 or 10 mins.
14
+2. the central monitoring solution needs dedicated resources, thus becoming "another bottleneck" in the whole ecosystem. It also requires maintenance, administration, etc.
15
+3. most centralized monitoring solutions are usually only good for presenting *statistics of past performance* (i.e. cannot be used for real-time performance troubleshooting).
16
+
17
+Netdata has a different approach:
18
+
19
+1. data collection happens per second
20
+2. thousands of metrics per server are collected
21
+3. data do not leave the server where they are collected
22
+4. netdata servers do not talk to each other
23
+5. your browser connects all the netdata servers
24
+
25
+Using netdata, your monitoring infrastructure is embedded on each server, limiting significantly the need of additional resources. netdata is blazingly fast, very resource efficient and utilizes server resources that already exist and are spare (on each server). This allows **scaling out** the monitoring infrastructure.
26
+
27
+However, the netdata approach introduces a few new issues that need to be addressed, one being **the list of netdata we have installed**, i.e. the URLs our netdata servers are listening.
28
+
29
+To solve this, netdata utilizes a **central registry**. This registry, together with certain browser features, allow netdata to provide unified cross server dashboards. For example, when you jump from server to server using the `my-netdata` menu, several session settings (like the currently viewed charts, the current zoom and pan operations on the charts, etc) are propagated to the new server, so that the new dashboard will come with exactly the same view.
30
+
31
+## What is the registry?
32
+
33
+The registry keeps track of 3 entities:
34
+
35
+1. **machines**: i.e. the netdata installations (a random GUID generated by each netdata the first time it starts; we call this **machine_guid**)
36
+
37
+ For each netdata installation (each `machine_guid`) the registry keeps track of the different URLs it is accessed.
38
+
39
+2. **persons**: i.e. the web browsers accessing the netdata installations (a random GUID generated by the registry the first time it sees a new web browser; we call this **person_guid**)
40
+
41
+ For each person, the registry keeps track of the netdata installations it has accessed and their URLs.
42
+
43
+3. **URLs** of netdata installations (as seen by the web browsers)
44
+
45
+ For each URL, the registry keeps the URL and nothing more. Each URL is linked to *persons* and *machines*. The only way to find a URL is to know its **machine_guid** or have a **person_guid** it is linked to it.
46
+
47
+## Who talks to the registry?
48
+
49
+Your web browser **only**! Check here if this is against your policies: [how to not send any information to a thirdparty server](https://github.com/netdata/netdata/wiki/netdata-security#registry-or-how-to-not-send-any-information-to-a-thirdparty-server)
50
+
51
+Your netdata servers do not talk to the registry. This is a UML diagram of its operation:
52
+
53
+
54
+
55
+## What data the registry maintains?
56
+
57
+Its database contains:
58
+
59
+- **random person GUIDs** (generated by the registry as a browser cookie)
60
+- **random machine GUIDs** (generated by each netdata server on its first run), including the hostname of the server netdata is running (without the domain)
61
+- **URLs** (the base URL for accessing a netdata server, as seen by the web browser)
62
+
63
+For *persons* and *machines*, the registry keeps links to *URLs*, each link with 2 timestamps (first time seen, last time seen) and a counter (number of times it has been seen).
64
+
65
+## Which is the default registry?
66
+
67
+`https://registry.my-netdata.io`, which is currently served by `https://london.my-netdata.io`. This registry listens to both HTTP and HTTPS requests but the default is HTTPS.
68
+
69
+#### Can this registry handle the global load of netdata installations?
70
+
71
+Yeap! The registry can handle 50.000 - 100.000 requests **per second per core** (depending on the type of CPU, the computer's memory bandwidth, etc). 50.000 is on J1900 (celeron 2Ghz).
72
+
73
+We believe, it can do it...
74
+
75
+## Every netdata can be a registry
76
+
77
+Yes, you read correct, **every netdata can be a registry**. Just pick one and configure it.
78
+
79
+**To turn any netdata into a registry**, edit `/etc/netdata/netdata.conf` and set:
80
+
81
+```
82
+[registry]
83
+ enabled = yes
84
+ registry to announce = http://your.registry:19999
85
+```
86
+
87
+Restart your netdata to activate it.
88
+
89
+Then, you need to tell **all your other netdata servers to advertise your registry**, instead of the default. To do this, on each of your netdata servers, edit `/etc/netdata/netdata.conf` and set:
90
+
91
+```
92
+[registry]
93
+ enabled = no
94
+ registry to announce = http://your.registry:19999
95
+```
96
+
97
+Note that we have not enabled the registry on the other servers. Only one netdata (the registry) needs `[registry].enabled = yes`.
98
+
99
+This is it. You have your registry now.
100
+
101
+You may also want to give your server different names under the **my-netdata** menu (i.e. to have them sorted / grouped). You can change its registry name, by setting on each netdata server:
102
+
103
+```
104
+[registry]
105
+ registry hostname = Group1 - Master DB
106
+```
107
+
108
+So this server will appear in **my-netdata** as `Group1 - Master DB`. The max name length is 50 characters.
109
+
110
+#### limiting access to the registry
111
+
112
+netdata v1.9+ support limiting access to the registry from given IPs, like this:
113
+```
114
+[registry]
115
+ allow from = *
116
+```
117
+
118
+`allow from` settings are [netdata simple patterns](https://github.com/netdata/netdata/wiki/Configuration#netdata-simple-patterns): string matches that use `*` as wildcard (any number of times) and a `!` prefix for a negative match. So: `allow from = !10.1.2.3 10.*` will allow all IPs in `10.*` except `10.1.2.3`. The order is important: left to right, the first positive or negative match is used.
119
+
120
+Keep in mind that connections to netdata API ports are filtered by `[web].allow connections from`. So, IPs allowed by `[registry].allow from` should also be allowed by `[web].allow connection from`.
121
+
122
+#### Where is the registry database stored?
123
+
124
+`/var/lib/netdata/registry/*.db`
125
+
126
+There can be up to 2 files:
127
+
128
+- `registry-log.db`, the transaction log
129
+
130
+ all incoming requests that affect the registry are saved in this file in real-time.
131
+
132
+- `registry.db`, the database
133
+
134
+ every `[registry].registry save db every new entries` entries in `registry-log.db`, netdata will save its database to `registry.db` and empty `registry-log.db`.
135
+
136
+Both files are machine readable text files.
137
+
138
+## The future
139
+
140
+The registry opens a whole world of new possibilities for netdata. Check here what we think: https://github.com/netdata/netdata/issues/416
141
+
142
+## Troubleshooting the registry
143
+
144
+The registry URL should be set to the URL of a netdata dashboard. This server has to have `[registry].enabled = yes`. So, accessing the registry URL directly with your web browser, should present the dashboard of the netdata operating the registry.
145
+
146
+To use the registry, your web browser needs to support **third party cookies**, since the cookies are set by the registry while you are browsing the dashboard of another netdata server. The registry, the first time it sees a new web browser it tries to figure if the web browser has cookies enabled or not. It does this by setting a cookie and redirecting the browser back to itself hoping that it will receive the cookie. If it does not receive the cookie, the registry will keep redirecting your web browser back to itself, which after a few redirects will fail with an error like this:
147
+
148
+```
149
+ERROR 409: Cannot ACCESS netdata registry: https://registry.my-netdata.io responded with: {"status":"redirect","registry":"https://registry.my-netdata.io"}
150
+```
151
+
152
+This error is printed on your web browser console (press F12 on your browser to show it).
streaming/README.md
+413
@@ -0,0 +1,413 @@
1
+# Metrics streaming
2
+
3
+Each netdata is able to replicate/mirror its database to another netdata, by streaming collected
4
+metrics, in real-time to it. This is quite different to [data archiving to third party time-series
5
+databases](../backends).
6
+
7
+When a netdata streams metrics to another netdata, the receiving one is able to perform everything
8
+a netdata performs:
9
+
10
+- visualize them with a dashboard
11
+- run health checks that trigger alarms and send alarm notifications
12
+- archive metrics to a backend time-series database
13
+
14
+The following configurations are supported:
15
+
16
+#### netdata without a database or web API (headless collector)
17
+
18
+Local netdata (`slave`), **without any database or alarms**, collects metrics and sends them to
19
+another netdata (`master`).
20
+
21
+The user can take the full functionality of the `slave` netdata at
22
+http://master.ip:19999/host/slave.hostname/. Alarms for the `slave` are served by the `master`.
23
+
24
+In this mode the `slave` is just a plain data collector.
25
+It runs with... **5MB** of RAM (yes, you read correct), spawns all external plugins, but instead
26
+of maintaining a local database and accepting dashboard requests, it streams all metrics to the
27
+`master`.
28
+
29
+The same `master` can collect data for any number of `slaves`.
30
+
31
+#### database replication
32
+
33
+Local netdata (`slave`), **with a local database (and possibly alarms)**, collects metrics and
34
+sends them to another netdata (`master`).
35
+
36
+The user can use all the functions **at both** http://slave.ip:19999/ and
37
+http://master.ip:19999/host/slave.hostname/.
38
+
39
+The `slave` and the `master` may have different data retention policies for the same metrics.
40
+
41
+Alarms for the `slave` are triggered by **both** the `slave` and the `master` (and actually
42
+each can have different alarms configurations or have alarms disabled).
43
+
44
+#### netdata proxies
45
+
46
+Local netdata (`slave`), with or without a database, collects metrics and sends them to another
47
+netdata (`proxy`), which may or may not maintain a database, which forwards them to another
48
+netdata (`master`).
49
+
50
+Alarms for the slave can be triggered by any of the involved hosts that maintains a database.
51
+
52
+Any number of daisy chaining netdata servers are supported, each with or without a database and
53
+with or without alarms for the `slave` metrics.
54
+
55
+#### mix and match with backends
56
+
57
+All nodes that maintain a database can also send their data to a backend database.
58
+This allows quite complex setups.
59
+
60
+Example:
61
+
62
+1. netdata `A`, `B` do not maintain a database and stream metrics to netdata `C`(live streaming functionality, i.e. this PR)
63
+2. netdata `C` maintains a database for `A`, `B`, `C` and archives all metrics to `graphite` with 10 second detail (backends functionality)
64
+3. netdata `C` also streams data for `A`, `B`, `C` to netdata `D`, which also collects data from `E`, `F` and `G` from another DMZ (live streaming functionality, i.e. this PR)
65
+4. netdata `D` is just a proxy, without a database, that streams all data to a remote site at netdata `H`
66
+5. netdata `H` maintains a database for `A`, `B`, `C`, `D`, `E`, `F`, `G`, `H` and sends all data to `opentsdb` with 5 seconds detail (backends functionality)
67
+6. alarms are triggered by `H` for all hosts
68
+7. users can use all the netdata that maintain a database to view metrics (i.e. at `H` all hosts can be viewed).
69
+
70
+#### netdata.conf configuration
71
+
72
+These are options that affect the operation of netdata in this area:
73
+
74
+```
75
+[global]
76
+ memory mode = none | ram | save | map
77
+```
78
+
79
+`[global].memory mode = none` disables the database at this host. This also disables health
80
+monitoring (there cannot be health monitoring without a database).
81
+
82
+```
83
+[web]
84
+ mode = none | static-threaded | single-threaded | multi-threaded
85
+```
86
+
87
+`[web].mode = none` disables the API (netdata will not listen to any ports).
88
+This also disables the registry (there cannot be a registry without an API).
89
+
90
+```
91
+[backend]
92
+ enabled = yes | no
93
+ type = graphite | opentsdb
94
+ destination = IP:PORT ...
95
+ update every = 10
96
+```
97
+
98
+`[backend]` configures data archiving to a backend (it archives all databases maintained on
99
+this host).
100
+
101
+#### streaming configuration
102
+
103
+A new file is introduced: [stream.conf](stream.conf) (to edit it on your system run
104
+`/etc/netdata/edit-config stream.conf`). This file holds streaming configuration for both the
105
+sending and the receiving netdata.
106
+
107
+API keys are used to authorize the communication of a pair of sending-receiving netdata.
108
+Once the communication is authorized, the sending netdata can push metrics for any number of hosts.
109
+
110
+You can generate an API key with the command `uuidgen`. API keys are just random GUIDs.
111
+You can use the same API key on all your netdata, or use a different API key for any pair of
112
+sending-receiving netdata.
113
+
114
+##### options for the sending node
115
+
116
+This is the section for the sending netdata. On the receiving node, `[stream].enabled` can be `no`.
117
+If it is `yes`, the receiving node will also stream the metrics to another node (i.e. it will be
118
+a `proxy`).
119
+
120
+```
121
+[stream]
122
+ enabled = yes | no
123
+ destination = IP:PORT ...
124
+ api key = XXXXXXXXXXX
125
+```
126
+
127
+This is an overview of how these options can be combined:
128
+
129
+target | memory<br/>mode | web<br/>mode | stream<br/>enabled | backend | alarms | dashboard
130
+-------|:-----------:|:---:|:------:|:-------:|:---------:|:----:
131
+headless collector|`none`|`none`|`yes`|only for `data source = as collected`|not possible|no
132
+headless proxy|`none`|not `none`|`yes`|only for `data source = as collected`|not possible|no
133
+proxy with db|not `none`|not `none`|`yes`|possible|possible|yes
134
+central netdata|not `none`|not `none`|`no`|possible|possible|yes
135
+
136
+##### options for the receiving node
137
+
138
+`stream.conf` looks like this:
139
+
140
+```sh
141
+# replace API_KEY with your uuidgen generated GUID
142
+[API_KEY]
143
+ enabled = yes
144
+ default history = 3600
145
+ default memory mode = save
146
+ health enabled by default = auto
147
+ allow from = *
148
+```
149
+
150
+You can add many such sections, one for each API key. The above are used as default values for
151
+all hosts pushed with this API key.
152
+
153
+You can also add sections like this:
154
+
155
+```sh
156
+# replace MACHINE_GUID with the slave /var/lib/netdata/registry/netdata.public.unique.id
157
+[MACHINE_GUID]
158
+ enabled = yes
159
+ history = 3600
160
+ memory mode = save
161
+ health enabled = yes
162
+ allow from = *
163
+```
164
+
165
+The above is the receiver configuration of a single host, at the receiver end. `MACHINE_GUID` is
166
+the unique id the netdata generating the metrics (i.e. the netdata that originally collects
167
+them `/var/lib/netdata/registry/netdata.unique.id`). So, metrics for netdata `A` that pass through
168
+any number of other netdata, will have the same `MACHINE_GUID`.
169
+
170
+####### allow from
171
+
172
+`allow from` settings are [netdata simple patterns](../libnetdata/simple_pattern): string matches
173
+that use `*` as wildcard (any number of times) and a `!` prefix for a negative match.
174
+So: `allow from = !10.1.2.3 10.*` will allow all IPs in `10.*` except `10.1.2.3`. The order is
175
+important: left to right, the first positive or negative match is used.
176
+
177
+`allow from` is available in netdata v1.9+
178
+
179
+#### tracing
180
+
181
+When a `slave` is trying to push metrics to a `master` or `proxy`, it logs entries like these:
182
+
183
+```
184
+2017-02-25 01:57:44: netdata: ERROR: Failed to connect to '10.11.12.1', port '19999' (errno 111, Connection refused)
185
+2017-02-25 01:57:44: netdata: ERROR: STREAM costa-pc [send to 10.11.12.1:19999]: failed to connect
186
+2017-02-25 01:58:04: netdata: INFO : STREAM costa-pc [send to 10.11.12.1:19999]: initializing communication...
187
+2017-02-25 01:58:04: netdata: INFO : STREAM costa-pc [send to 10.11.12.1:19999]: waiting response from remote netdata...
188
+2017-02-25 01:58:14: netdata: INFO : STREAM costa-pc [send to 10.11.12.1:19999]: established communication - sending metrics...
189
+2017-02-25 01:58:14: netdata: ERROR: STREAM costa-pc [send]: discarding 1900 bytes of metrics already in the buffer.
190
+2017-02-25 01:58:14: netdata: INFO : STREAM costa-pc [send]: ready - sending metrics...
191
+```
192
+
193
+The receiving end (`proxy` or `master`) logs entries like these:
194
+
195
+```
196
+2017-02-25 01:58:04: netdata: INFO : STREAM [receive from [10.11.12.11]:33554]: new client connection.
197
+2017-02-25 01:58:04: netdata: INFO : STREAM costa-pc [10.11.12.11]:33554: receive thread created (task id 7698)
198
+2017-02-25 01:58:14: netdata: INFO : Host 'costa-pc' with guid '12345678-b5a6-11e6-8a50-00508db7e9c9' initialized, os: linux, update every: 1, memory mode: ram, history entries: 3600, streaming: disabled, health: enabled, cache_dir: '/var/cache/netdata/12345678-b5a6-11e6-8a50-00508db7e9c9', varlib_dir: '/var/lib/netdata/12345678-b5a6-11e6-8a50-00508db7e9c9', health_log: '/var/lib/netdata/12345678-b5a6-11e6-8a50-00508db7e9c9/health/health-log.db', alarms default handler: '/usr/libexec/netdata/plugins.d/alarm-notify.sh', alarms default recipient: 'root'
199
+2017-02-25 01:58:14: netdata: INFO : STREAM costa-pc [receive from [10.11.12.11]:33554]: initializing communication...
200
+2017-02-25 01:58:14: netdata: INFO : STREAM costa-pc [receive from [10.11.12.11]:33554]: receiving metrics...
201
+```
202
+
203
+For netdata v1.9+, streaming can also be monitored via `access.log`.
204
+
205
+
206
+#### Viewing remote host dashboards, using mirrored databases
207
+
208
+On any receiving netdata, that maintains remote databases and has its web server enabled,
209
+`my-netdata` menu will include a list of the mirrored databases.
210
+
211
+
212
+
213
+Selecting any of these, the server will offer a dashboard using the mirrored metrics.
214
+
215
+
216
+## Monitoring ephemeral nodes
217
+
218
+Auto-scaling is probably the most trendy service deployment strategy these days.
219
+
220
+Auto-scaling detects the need for additional resources and boots VMs on demand, based on a template. Soon after they start running the applications, a load balancer starts distributing traffic to them, allowing the service to grow horizontally to the scale needed to handle the load. When demands falls, auto-scaling starts shutting down VMs that are no longer needed.
221
+
222
+<p align="center">
223
+<img src="https://cloud.githubusercontent.com/assets/2662304/23627426/65a9074a-02b9-11e7-9664-cd8f258a00af.png"/>
224
+</p>
225
+
226
+What a fantastic feature for controlling infrastructure costs! Pay only for what you need for the time you need it!
227
+
228
+In auto-scaling, all servers are ephemeral, they live for just a few hours. Every VM is a brand new instance of the application, that was automatically created based on a template.
229
+
230
+So, how can we monitor them? How can we be sure that everything is working as expected on all of them?
231
+
232
+### The netdata way
233
+
234
+We recently made a significant improvement at the core of netdata to support monitoring such setups.
235
+
236
+Following the netdata way of monitoring, we wanted:
237
+
238
+1. **real-time performance monitoring**, collecting **_thousands of metrics per server per second_**, visualized in interactive, automatically created dashboards.
239
+2. **real-time alarms**, for all nodes.
240
+3. **zero configuration**, all ephemeral servers should have exactly the same configuration, and nothing should be configured at any system for each of the ephemeral nodes. We shouldn't care if 10 or 100 servers are spawned to handle the load.
241
+4. **self-cleanup**, so that nothing needs to be done for cleaning up the monitoring infrastructure from the hundreds of nodes that may have been monitored through time.
242
+
243
+#### How it works
244
+
245
+All monitoring solutions, including netdata, work like this:
246
+
247
+1. `collect metrics`, from the system and the running applications
248
+2. `store metrics`, in a time-series database
249
+3. `examine metrics` periodically, for triggering alarms and sending alarm notifications
250
+4. `visualize metrics`, so that users can see what exactly is happening
251
+
252
+netdata used to be self-contained, so that all these functions were handled entirely by each server. The changes we made, allow each netdata to be configured independently for each function. So, each netdata can now act as:
253
+
254
+- a `self contained system`, much like it used to be.
255
+- a `data collector`, that collects metrics from a host and pushes them to another netdata (with or without a local database and alarms).
256
+- a `proxy`, that receives metrics from other hosts and pushes them immediately to other netdata servers. netdata proxies can also be `store and forward proxies` meaning that they are able to maintain a local database for all metrics passing through them (with or without alarms).
257
+- a `time-series database` node, where data are kept, alarms are run and queries are served to visualise the metrics.
258
+
259
+### Configuring an auto-scaling setup
260
+
261
+<p align="center">
262
+<img src="https://cloud.githubusercontent.com/assets/2662304/23627468/96daf7ba-02b9-11e7-95ac-1f767dd8dab8.png"/>
263
+</p>
264
+
265
+You need a netdata `master`. This node should not be ephemeral. It will be the node where all ephemeral nodes (let's call them `slaves`) will be sending their metrics.
266
+
267
+The master will need to authorize the slaves for accepting their metrics. This is done with an API key.
268
+
269
+#### API keys
270
+
271
+API keys are just random GUIDs. Use the Linux command `uuidgen` to generate one. You can use the same API key for all your `slaves`, or you can configure one API for each of them. This is entirely your decision.
272
+
273
+We suggest to use the same API key for each ephemeral node template you have, so that all replicas of the same ephemeral node will have exactly the same configuration.
274
+
275
+I will use this API_KEY: `11111111-2222-3333-4444-555555555555`. Replace it with your own.
276
+
277
+#### Configuring the `master`
278
+
279
+On the master, edit `/etc/netdata/stream.conf` (to edit it on your system run `/etc/netdata/edit-config stream.conf`) and set these:
280
+
281
+```bash
282
+[11111111-2222-3333-4444-555555555555]
283
+ # enable/disable this API key
284
+ enabled = yes
285
+
286
+ # one hour of data for each of the slaves
287
+ default history = 3600
288
+
289
+ # do not save slave metrics on disk
290
+ default memory = ram
291
+
292
+ # alarms checks, only while the slave is connected
293
+ health enabled by default = auto
294
+```
295
+*`stream.conf` on master, to enable receiving metrics from slaves using the API key.*
296
+
297
+If you used many API keys, you can add one such section for each API key.
298
+
299
+When done, restart netdata on the `master` node. It is now ready to receive metrics.
300
+
301
+#### Configuring the `slaves`
302
+
303
+On each of the slaves, edit `/etc/netdata/stream.conf` (to edit it on your system run `/etc/netdata/edit-config stream.conf`) and set these:
304
+
305
+```bash
306
+[stream]
307
+ # stream metrics to another netdata
308
+ enabled = yes
309
+
310
+ # the IP and PORT of the master
311
+ destination = 10.11.12.13:19999
312
+
313
+ # the API key to use
314
+ api key = 11111111-2222-3333-4444-555555555555
315
+```
316
+*`stream.conf` on slaves, to enable pushing metrics to master at `10.11.12.13:19999`.*
317
+
318
+Using just the above configuration, the `slaves` will be pushing their metrics to the `master` netdata, but they will still maintain a local database of the metrics and run health checks. To disable them, edit `/etc/netdata/netdata.conf` and set:
319
+
320
+```bash
321
+[global]
322
+ # disable the local database
323
+ memory mode = none
324
+
325
+[health]
326
+ # disable health checks
327
+ enabled = no
328
+```
329
+*`netdata.conf` configuration on slaves, to disable the local database and health checks.*
330
+
331
+Keep in mind that setting `memory mode = none` will also force `[health].enabled = no` (health checks require access to a local database). But you can keep the database and disable health checks if you need to. You are however sending all the metrics to the master server, which can handle the health checking (`[health].enabled = yes`)
332
+
333
+#### netdata unique id
334
+
335
+The file `/var/lib/netdata/registry/netdata.public.unique.id` contains a random GUID that **uniquely identifies each netdata**. This file is automatically generated, by netdata, the first time it is started and remains unaltaired forever.
336
+
337
+> If you are building an image to be used for automated provisioning of autoscaled VMs, it important to delete that file from the image, so that each instance of your image will generate its own.
338
+
339
+#### Troubleshooting metrics streaming
340
+
341
+Both the sender and the receiver of metrics log information at `/var/log/netdata/error.log`.
342
+
343
+
344
+On both master and slave do this:
345
+
346
+```
347
+tail -f /var/log/netdata/error.log | grep STREAM
348
+```
349
+
350
+If the slave manages to connect to the master you will see something like (on the master):
351
+
352
+```
353
+2017-03-09 09:38:52: netdata: INFO : STREAM [receive from [10.11.12.86]:38564]: new client connection.
354
+2017-03-09 09:38:52: netdata: INFO : STREAM xxx [10.11.12.86]:38564: receive thread created (task id 27721)
355
+2017-03-09 09:38:52: netdata: INFO : STREAM xxx [receive from [10.11.12.86]:38564]: client willing to stream metrics for host 'xxx' with machine_guid '1234567-1976-11e6-ae19-7cdd9077342a': update every = 1, history = 3600, memory mode = ram, health auto
356
+2017-03-09 09:38:52: netdata: INFO : STREAM xxx [receive from [10.11.12.86]:38564]: initializing communication...
357
+2017-03-09 09:38:52: netdata: INFO : STREAM xxx [receive from [10.11.12.86]:38564]: receiving metrics...
358
+```
359
+
360
+and something like this on the slave:
361
+
362
+```
363
+2017-03-09 09:38:28: netdata: INFO : STREAM xxx [send to box:19999]: connecting...
364
+2017-03-09 09:38:28: netdata: INFO : STREAM xxx [send to box:19999]: initializing communication...
365
+2017-03-09 09:38:28: netdata: INFO : STREAM xxx [send to box:19999]: waiting response from remote netdata...
366
+2017-03-09 09:38:28: netdata: INFO : STREAM xxx [send to box:19999]: established communication - sending metrics...
367
+```
368
+
369
+### Archiving to a time-series database
370
+
371
+The `master` netdata node can also archive metrics, for all `slaves`, to a time-series database. At the time of this writing, netdata supports:
372
+
373
+- graphite
374
+- opentsdb
375
+- prometheus
376
+- json document DBs
377
+- all the compatibles to the above (e.g. kairosdb, influxdb, etc)
378
+
379
+Check the netdata [backends documentation](../backends) for configuring this.
380
+
381
+This is how such a solution will work:
382
+
383
+<p align="center">
384
+<img src="https://cloud.githubusercontent.com/assets/2662304/23627295/e3569adc-02b8-11e7-9d55-4014bf98c1b3.png"/>
385
+</p>
386
+
387
+### An advanced setup
388
+
389
+netdata also supports `proxies` with and without a local database, and data retention can be different between all nodes.
390
+
391
+This means a setup like the following is also possible:
392
+
393
+<p align="center">
394
+<img src="https://cloud.githubusercontent.com/assets/2662304/23629551/bb1fd9c2-02c0-11e7-90f5-cab5a3ed4c53.png"/>
395
+</p>
396
+
397
+
398
+## proxies
399
+
400
+A proxy is a netdata that is receiving metrics from a netdata, and streams them to another netdata.
401
+
402
+netdata proxies may or may not maintain a database for the metrics passing through them.
403
+When they maintain a database, they can also run health checks (alarms and notifications)
404
+for the remote host that is streaming the metrics.
405
+
406
+To configure a proxy, configure it as a receiving and a sending netdata at the same time,
407
+using [stream.conf](stream.conf).
408
+
409
+The sending side of a netdata proxy, connects and disconnects to the final destination of the
410
+metrics, following the same pattern of the receiving side.
411
+
412
+For a practical example see [Monitoring ephemeral nodes](#monitoring-ephemeral-nodes).
413
+
web/api/Makefile.am
+4
@@ -3,6 +3,10 @@
3
AUTOMAKE_OPTIONS = subdir-objects
4
MAINTAINERCLEANFILES = $(srcdir)/Makefile.in
5
6
+SUBDIR = \
7
+ badges \
8
+ $(NULL)
9
+
10
dist_noinst_DATA = \
11
README.md \
12
$(NULL)
web/api/README.md
+83
@@ -0,0 +1,83 @@
1
+# netdata REST API
2
+
3
+The complete documentation of the netdata API is available at the **[Swagger Editor](https://editor.swagger.io/?url=https://raw.githubusercontent.com/netdata/netdata/master/web/netdata-swagger.yaml)**.
4
+
5
+If your prefer it over the Swagger Editor, you can also use **[Swagger UI](https://registry.my-netdata.io/swagger/#!/default/get_data)**. This however does not provide all the information available.
6
+
7
+## google charts
8
+
9
+netdata is a [Google Visualization API datatable and datasource provider](https://developers.google.com/chart/interactive/docs/reference), so it can directly be used with [Google Charts](https://developers.google.com/chart/interactive/docs/).
10
+
11
+Check this [single chart, jsfiddle example](https://jsfiddle.net/ktsaou/ensu4uws/9/):
12
+
13
+
14
+
15
+and this [multi chart, jsfiddle example](https://jsfiddle.net/ktsaou/L5y2eqp2/):
16
+
17
+
18
+
19
+
20
+## using the api from shell scripts
21
+
22
+Shell scripts can now query netdata easily:
23
+
24
+```sh
25
+eval "$(curl -s 'http://localhost:19999/api/v1/allmetrics')"
26
+```
27
+
28
+after this command, all the netdata metrics are exposed to shell. Check:
29
+
30
+```sh
31
+# source the metrics
32
+eval "$(curl -s 'http://localhost:19999/api/v1/allmetrics')"
33
+
34
+# let's see if there are variables exposed by netdata for system.cpu
35
+set | grep "^NETDATA_SYSTEM_CPU"
36
+
37
+NETDATA_SYSTEM_CPU_GUEST=0
38
+NETDATA_SYSTEM_CPU_GUEST_NICE=0
39
+NETDATA_SYSTEM_CPU_IDLE=95
40
+NETDATA_SYSTEM_CPU_IOWAIT=0
41
+NETDATA_SYSTEM_CPU_IRQ=0
42
+NETDATA_SYSTEM_CPU_NICE=0
43
+NETDATA_SYSTEM_CPU_SOFTIRQ=0
44
+NETDATA_SYSTEM_CPU_STEAL=0
45
+NETDATA_SYSTEM_CPU_SYSTEM=1
46
+NETDATA_SYSTEM_CPU_USER=4
47
+NETDATA_SYSTEM_CPU_VISIBLETOTAL=5
48
+
49
+# let's see the total cpu utilization of the system
50
+echo ${NETDATA_SYSTEM_CPU_VISIBLETOTAL}
51
+5
52
+
53
+# what about alarms?
54
+set | grep "^NETDATA_ALARM_SYSTEM_SWAP_"
55
+NETDATA_ALARM_SYSTEM_SWAP_RAM_IN_SWAP_STATUS=CRITICAL
56
+NETDATA_ALARM_SYSTEM_SWAP_RAM_IN_SWAP_VALUE=53
57
+NETDATA_ALARM_SYSTEM_SWAP_USED_SWAP_STATUS=CLEAR
58
+NETDATA_ALARM_SYSTEM_SWAP_USED_SWAP_VALUE=51
59
+
60
+# let's get the current status of the alarm 'ram in swap'
61
+echo ${NETDATA_ALARM_SYSTEM_SWAP_RAM_IN_SWAP_STATUS}
62
+CRITICAL
63
+
64
+# is it fast?
65
+time curl -s 'http://localhost:19999/api/v1/allmetrics' >/dev/null
66
+
67
+real 0m0,070s
68
+user 0m0,000s
69
+sys 0m0,007s
70
+
71
+# it is...
72
+# 0.07 seconds for curl to be loaded, connect to netdata and fetch the response back...
73
+```
74
+
75
+The `_VISIBLETOTAL` variable sums up all the dimensions of each chart.
76
+
77
+The format of the variables is:
78
+
79
+```sh
80
+NETDATA_${chart_id^^}_${dimension_id^^}="${value}"
81
+```
82
+
83
+The value is rounded to the closest integer, since shell script cannot process decimal numbers.
web/api/badges/Makefile.am
new
+8
@@ -0,0 +1,8 @@
1
+# SPDX-License-Identifier: GPL-3.0-or-later
2
+
3
+AUTOMAKE_OPTIONS = subdir-objects
4
+MAINTAINERCLEANFILES = $(srcdir)/Makefile.in
5
+
6
+dist_noinst_DATA = \
7
+ README.md \
8
+ $(NULL)
web/api/badges/README.md
new
+324
@@ -0,0 +1,324 @@
1
+# Netdata badges
2
+
3
+**Badges are cool!**
4
+
5
+Netdata can generate badges for any chart and any dimension at any time-frame. Badges come in `SVG` and can be added to any web page using an `<IMG>` HTML tag.
6
+
7
+**Netdata badges are powerful**!
8
+
9
+Given that netdata collects from **1.000** to **5.000** metrics per server (depending on the number of network interfaces, disks, cpu cores, applications running, users logged in, containers running, etc) and that netdata already has data reduction/aggregation functions embedded, the badges can be quite powerful.
10
+
11
+For each metric/dimension and for arbitrary time-frames badges can show **min**, **max** or **average** value, but also **sum** or **incremental-sum** to have their **volume**.
12
+
13
+For example, there is [a chart in netdata that shows the current requests/s of nginx](http://london.my-netdata.io/#nginx_local_nginx). Using this chart alone we can show the following badges (we could add more time-frames, like **today**, **yesterday**, etc):
14
+
15
+<a href="https://registry.my-netdata.io/#nginx_local_nginx"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=nginx_local.connections&dimensions=active&value_color=grey:null%7Cblue&label=nginx%20active%20connections%20now&units=null&precision=0"/></a> <a href="https://registry.my-netdata.io/#nginx_local_nginx"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=nginx_local.connections&dimensions=active&after=-3600&value_color=orange&label=last%20hour%20average&units=null&options=unaligned&precision=0"/></a> <a href="https://registry.my-netdata.io/#nginx_local_nginx"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=nginx_local.connections&dimensions=active&group=max&after=-3600&value_color=red&label=last%20hour%20max&units=null&options=unaligned&precision=0"/></a>
16
+
17
+Similarly, there is [a chart that shows outbound bandwidth per class](http://london.my-netdata.io/#tc_eth0), using QoS data. So it shows `kilobits/s` per class. Using this chart we can show:
18
+
19
+<a href="https://registry.my-netdata.io/#tc_eth0"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=tc.world_out&dimensions=web_server&value_color=green&label=web%20server%20sends%20now&units=kbps"/></a> <a href="https://registry.my-netdata.io/#tc_eth0"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=tc.world_out&dimensions=web_server&after=-86400&options=unaligned&group=sum÷=8388608&value_color=blue&label=web%20server%20sent%20today&units=GB"/></a>
20
+
21
+The right one is a **volume** calculation. Netdata calculated the total of the last 86.400 seconds (a day) which gives `kilobits`, then divided it by 8 to make it KB, then by 1024 to make it MB and then by 1024 to make it GB. Calculations like this are quite accurate, since for every value collected, every second, netdata interpolates it to second boundary using microsecond calculations.
22
+
23
+Let's see a few more badge examples (they come from the [netdata registry](https://github.com/netdata/netdata/wiki/mynetdata-menu-item)):
24
+
25
+- **cpu usage of user `root`** (you can pick any user; 100% = 1 core). This will be `green <10%`, `yellow <20%`, `orange <50%`, `blue <100%` (1 core), `red` otherwise (you define thresholds and colors on the URL).
26
+
27
+ <a href="https://registry.my-netdata.io/#apps_cpu"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=users.cpu&dimensions=root&value_color=grey:null%7Cgreen%3C10%7Cyellow%3C20%7Corange%3C50%7Cblue%3C100%7Cred&label=root%20user%20cpu%20now&units=%25"></img></a> <a href="https://registry.my-netdata.io/#apps_cpu"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=users.cpu&dimensions=root&after=-3600&value_color=grey:null%7Cgreen%3C10%7Cyellow%3C20%7Corange%3C50%7Cblue%3C100%7Cred&label=root%20user%20average%20cpu%20last%20hour&units=%25"></img></a>
28
+
29
+- **mysql queries per second**
30
+
31
+ <a href="https://registry.my-netdata.io/#mysql_local"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=mysql_local.queries&dimensions=questions&label=mysql%20queries%20now&value_color=red&units=%5Cs"></img></a> <a href="https://registry.my-netdata.io/#mysql_local"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=mysql_local.queries&dimensions=questions&after=-3600&options=unaligned&group=sum&label=mysql%20queries%20this%20hour&value_color=green&units=null"></img></a> <a href="https://registry.my-netdata.io/#mysql_local"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=mysql_local.queries&dimensions=questions&after=-86400&options=unaligned&group=sum&label=mysql%20queries%20today&value_color=blue&units=null"></img></a>
32
+
33
+ niche ones: **mysql SELECT statements with JOIN, which did full table scans**:
34
+
35
+ <a href="https://registry.my-netdata.io/#mysql_local_issues"><img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=mysql_local.join_issues&dimensions=scan&after=-3600&label=full%20table%20scans%20the%20last%20hour&value_color=orange&group=sum&units=null"></img></a>
36
+
37
+---
38
+
39
+> So, every single line on the charts of a [netdata dashboard](http://london.my-netdata.io/), can become a badge and this badge can calculate **average**, **min**, **max**, or **volume** for any time-frame! And you can also vary the badge color using conditions on the calculated value.
40
+
41
+---
42
+
43
+## How to create badges
44
+
45
+The basic URL is `http://your.netdata:19999/api/v1/badge.svg?option1&option2&option3&...`.
46
+
47
+Here is what you can put for `options` (these are standard netdata API options):
48
+
49
+- `chart=CHART.NAME`
50
+
51
+ The chart to get the values from.
52
+
53
+ **This is the only parameter required** and with just this parameter, netdata will return the sum of the latest values of all chart dimensions.
54
+
55
+ Example:
56
+
57
+```html
58
+ <a href="#">
59
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu"></img>
60
+ </a>
61
+```
62
+
63
+ Which produces this:
64
+
65
+ <a href="#">
66
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu"></img>
67
+ </a>
68
+
69
+- `alarm=NAME`
70
+
71
+ Render the current value and status of an alarm linked to the chart. This option can be ignored if the badge to be generated is not related to an alarm.
72
+
73
+ The current value of the alarm will be rendered. The color of the badge will indicate the status of the alarm.
74
+
75
+ For alarm badges, **both `chart` and `alarm` parameters are required**.
76
+
77
+- `dimensions=DIMENSION1|DIMENSION2|...`
78
+
79
+ The dimensions of the chart to use. If you don't set any dimension, all will be used. When multiple dimensions are used, netdata will sum their values. You can append `options=absolute` if you want this sum to convert all values to positive before adding them.
80
+
81
+ Pipes in HTML have to escaped with `%7C`.
82
+
83
+ Example:
84
+
85
+```html
86
+ <a href="#">
87
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&dimensions=system%7Cnice"></img>
88
+ </a>
89
+```
90
+
91
+ Which produces this:
92
+
93
+ <a href="#">
94
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&dimensions=system%7Cnice"></img>
95
+ </a>
96
+
97
+- `before=SECONDS` and `after=SECONDS`
98
+
99
+ The timeframe. These can be absolute unix timestamps, or relative to now, number of seconds. By default `before=0` and `after=-1` (1 second in the past).
100
+
101
+ To get the last minute set `after=-60`. This will give the average of the last complete minute (XX:XX:00 - XX:XX:59).
102
+
103
+ To get the max of the last hour set `after=-3600&group=max`. This will give the maximum value of the last complete hour (XX:00:00 - XX:59:59)
104
+
105
+ Example:
106
+
107
+```html
108
+ <a href="#">
109
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60"></img>
110
+ </a>
111
+```
112
+
113
+ Which produces the average of last complete minute (XX:XX:00 - XX:XX:59):
114
+
115
+ <a href="#">
116
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60"></img>
117
+ </a>
118
+
119
+ While this is the previous minute (one minute before the last one, again aligned XX:XX:00 - XX:XX:59):
120
+
121
+```html
122
+ <a href="#">
123
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&before=-60&after=-60"></img>
124
+ </a>
125
+```
126
+
127
+ It produces this:
128
+
129
+ <a href="#">
130
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&before=-60&after=-60"></img>
131
+ </a>
132
+
133
+- `group=min` or `group=max` or `group=average` (the default) or `group=sum` or `group=incremental-sum`
134
+
135
+ If netdata will have to reduce (aggregate) the data to calculate the value, which aggregation method to use.
136
+
137
+ - `max` will find the max value for the timeframe. This works on both positive and negative dimensions. It will find the most extreme value.
138
+
139
+ - `min` will find the min value for the timeframe. This works on both positive and negative dimensions. It will find the number closest to zero.
140
+
141
+ - `average` will calculate the average value for the timeframe.
142
+
143
+ - `sum` will sum all the values for the timeframe. This is nice for finding the volume of dimensions for a timeframe. So if you have a dimension that reports `X per second`, you can find the volume of the dimension in a timeframe, by adding its values in that timeframe.
144
+
145
+ - `incremental-sum` will sum the difference of each value to its next. Let's assume you have a dimension that does not measure the rate of something, but the absolute value of it. So it has values like this "1, 5, 3, 7, 4". `incremental-sum` will calculate the difference of adjacent values. In this example, they will be `(5 - 1) + (3 - 5) + (7 - 3) + (4 - 7) = 3` (which is equal to the last value minus the first = 4 - 1).
146
+
147
+- `options=opt1|opt2|opt3|...`
148
+
149
+ These fine tune various options of the API. Here is what you can use for badges (the API has more option, but only these are useful for badges):
150
+
151
+ - `percentage`, instead of returning the value, calculate the percentage of the sum of the selected dimensions, versus the sum of all the dimensions of the chart. This also sets the units to `%`.
152
+
153
+ - `absolute` or `abs`, turn all values positive and then sum them.
154
+
155
+ - `display_absolute` or `display-absolute`, to use the signed value during color calculation, but display the absolute value on the badge.
156
+
157
+ - `min2max`, when multiple dimensions are given, do not sum them, but take their `max - min`.
158
+
159
+ - `unaligned`, when data are reduced / aggregated (e.g. the request is about the average of the last minute, or hour), netdata by default aligns them so that the charts will have a constant shape (so average per minute returns always XX:XX:00 - XX:XX:59). Setting the `unaligned` option, netdata will aggregate data without any alignment, so if the request is for 60 seconds, it will aggregate the latest 60 seconds of collected data.
160
+
161
+These are options dedicated to badges:
162
+
163
+- `label=TEXT`
164
+
165
+ The label of the badge.
166
+
167
+- `units=TEXT`
168
+
169
+ The units of the badge. If you want to put a `/`, please put a `\`. This is because netdata allows badges parameters to be given as path in URL, instead of query string. You can also use `null` or `empty` to show it without any units.
170
+
171
+ The units `seconds`, `minutes` and `hours` trigger special formatting. The value has to be in this unit, and netdata will automatically change it to show a more pretty duration.
172
+
173
+- `multiply=NUMBER`
174
+
175
+ Multiply the value with this number. The default is `1`.
176
+
177
+- `divide=NUMBER`
178
+
179
+ Divide the value with this number. The default is `1`.
180
+
181
+- `label_color=COLOR`
182
+
183
+ The color of the label (the left part). You can use any HTML color, include `#NNN` and `#NNNNNN`. The following colors are defined in netdata (and you can use them by name): `green`, `brightgreen`, `yellow`, `yellowgreen`, `orange`, `red`, `blue`, `grey`, `gray`, `lightgrey`, `lightgray`. These are taken from https://github.com/badges/shields so they are compatible with standard badges.
184
+
185
+- `value_color=COLOR:null|COLOR<VALUE|COLOR>VALUE|COLOR>=VALUE|COLOR<=VALUE|...`
186
+
187
+ You can add a pipe delimited list of conditions to pick the color. The first matching (left to right) will be used.
188
+
189
+ Example: `value_color=grey:null|green<10|yellow<100|orange<1000|blue<10000|red`
190
+
191
+ The above will set `grey` if no value exists (not collected within the `gap when lost iterations above` in netdata.conf for the chart), `green` if the value is less than 10, `yellow` if the value is less than 100, etc up to `red` which will be used if no other conditions match.
192
+
193
+ The supported operators are `<`, `>`, `<=`, `>=`, `=` (or `:`) and `!=` (or `<>`).
194
+
195
+- `precision=NUMBER`
196
+
197
+ The number of decimal digits of the value. By default netdata will add:
198
+
199
+ - no decimal digits for values > 1000
200
+ - 1 decimal digit for values > 100
201
+ - 2 decimal digits for values > 1
202
+ - 3 decimal digits for values > 0.1
203
+ - 4 decimal digits for values <= 0.1
204
+
205
+ Using the `precision=NUMBER` you can set your preference per badge.
206
+
207
+- `scale=XXX`
208
+
209
+ This option scales the svg image. It accepts values above or equal to 100 (100% is the default scale). For example, lets get a few different sizes:
210
+
211
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60&scale=100"></img> original<br/>
212
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60&scale=125"></img> `scale=125`<br/>
213
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60&scale=150"></img> `scale=150`<br/>
214
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60&scale=175"></img> `scale=175`<br/>
215
+ <img src="http://registry.my-netdata.io/api/v1/badge.svg?chart=system.cpu&after=-60&scale=200"></img> `scale=200`
216
+
217
+
218
+- `refresh=auto` or `refresh=SECONDS`
219
+
220
+ This option enables auto-refreshing of images. netdata will send the HTTP header `Refresh: SECONDS` to the web browser, thus requesting automatic refresh of the images at regular intervals.
221
+
222
+ `auto` will calculate the proper `SECONDS` to avoid unnecessary refreshes. If `SECONDS` is zero, this feature is disabled (it is also disabled by default).
223
+
224
+ Auto-refreshing like this, works only if you access the badge directly. So, you may have to put it an `embed` or `iframe` for it to be auto-refreshed. Use something like this:
225
+
226
+```html
227
+<embed src="BADGE_URL" type="image/svg+xml" height="20" />
228
+```
229
+
230
+ Another way is to use javascript to auto-refresh them. You can auto-refresh all the netdata badges on a page using javascript. You have to add a class to all the netdata badges, like this `<img class="netdata-badge" src="..."/>`. Then add this javascript code to your page (it requires jquery):
231
+
232
+```html
233
+<script>
234
+ var NETDATA_BADGES_AUTOREFRESH_SECONDS = 5;
235
+ function refreshNetdataBadges() {
236
+ var now = new Date().getTime().toString();
237
+ $('.netdata-badge').each(function() {
238
+ this.src = this.src.replace(/\&_=\d*/, '') + '&_=' + now;
239
+ });
240
+ setTimeout(refreshNetdataBadges, NETDATA_BADGES_AUTOREFRESH_SECONDS * 1000);
241
+ }
242
+ setTimeout(refreshNetdataBadges, NETDATA_BADGES_AUTOREFRESH_SECONDS * 1000);
243
+</script>
244
+```
245
+
246
+A more advanced badges refresh method is to include `http://your.netdata.ip:19999/refresh-badges.js` in your page. For more information and use example, [check this](https://github.com/netdata/netdata/blob/master/web/gui/refresh-badges.js).
247
+
248
+---
249
+
250
+## Escaping URLs
251
+
252
+Keep in mind that if you add badge URLs to your HTML pages you have to escape the special characters:
253
+
254
+character|name|escape sequence
255
+:-------:|:--:|:-------------:
256
+` `|space (in labels and units)|`%20`
257
+` # `|hash (for colors)|`%23`
258
+` % `|percent (in units)|`%25`
259
+` < `|less than|`%3C`
260
+` > `|greater than|`%3E`
261
+` \ `|backslash (when you need a `/`)|`%5C`
262
+` \| `|pipe (delimiting parameters)|`%7C`
263
+
264
+---
265
+
266
+## Using the path instead of the query string
267
+
268
+The badges can also be generated using the URL path for passing parameters. The format is exactly the same.
269
+
270
+So instead of:
271
+
272
+ `http://your.netdata:19999/api/v1/badge.svg?option1&option2&option3&...`
273
+
274
+you can write:
275
+
276
+ `http://your.netdata:19999/api/v1/badge.svg/option1/option2/option3/...`
277
+
278
+You can also append anything else you like, like this:
279
+
280
+ `http://your.netdata:19999/api/v1/badge.svg/option1/option2/option3/my-super-badge.svg`
281
+
282
+## FAQ
283
+
284
+#### Is it fast?
285
+On modern hardware, netdata can generate about **2.000 badges per second per core**, before noticing any delays. It generates a badge in about half a millisecond!
286
+
287
+Of course these timing are for badges that use recent data. If you need badges that do calculations over long durations (a day, or more), timing will differ. netdata logs its timings at its `access.log`, so take a look there before adding a heavy badge on a busy web site. Of course, you can cache such badges or have a cron job get them from netdata and save them at your web server at regular intervals.
288
+
289
+
290
+#### Embedding badges in github
291
+
292
+You have 2 options a) SVG images with markdown and b) SVG images with HTML (directly in .md files).
293
+
294
+For example, this is the cpu badge shown above:
295
+
296
+- Markdown example:
297
+
298
+```md
299
+[](https://registry.my-netdata.io/#apps_cpu)
300
+```
301
+
302
+- HTML example:
303
+
304
+```html
305
+<a href="https://registry.my-netdata.io/#apps_cpu">
306
+ <img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=users.cpu&dimensions=root&value_color=grey:null%7Cgreen%3C10%7Cyellow%3C20%7Corange%3C50%7Cblue%3C100%7Cred&label=root%20user%20cpu%20now&units=%25"></img>
307
+</a>
308
+```
309
+
310
+Both produce this:
311
+
312
+<a href="https://registry.my-netdata.io/#apps_cpu">
313
+ <img src="https://registry.my-netdata.io/api/v1/badge.svg?chart=users.cpu&dimensions=root&value_color=grey:null%7Cgreen%3C10%7Cyellow%3C20%7Corange%3C50%7Cblue%3C100%7Cred&label=root%20user%20cpu%20now&units=%25"></img>
314
+</a>
315
+
316
+#### auto-refreshing badges in github
317
+
318
+Unfortunately it cannot be done. Github fetches all the images using a proxy and rewrites all the URLs to be served by the proxy.
319
+
320
+You can refresh them from your browser console though. Press F12 to open the web browser console (switch to the console too), paste the following and press enter. They will refresh:
321
+
322
+```js
323
+var len = document.images.length; while(len--) { document.images[len].src = document.images[len].src.replace(/\?cacheBuster=\d*/, "") + "?cacheBuster=" + new Date().getTime().toString(); };
324
+```
\ No newline at end of file
web/api/badges/web_buffer_svg.c
renamed
web/api/badges/web_buffer_svg.h
renamed
+1
-1
@@ -3,7 +3,7 @@
3
#ifndef NETDATA_WEB_BUFFER_SVG_H
4
#define NETDATA_WEB_BUFFER_SVG_H 1
5
6
-#include "web_api_v1.h"
6
+#include "web/api/web_api_v1.h"
7
8
extern void buffer_svg(BUFFER *wb, const char *label, calculated_number value, const char *units, const char *label_color, const char *value_color, int precision, int scale, uint32_t options);
9
extern char *format_value_and_unit(char *value_string, size_t value_string_len, calculated_number value, const char *units, int precision);
web/api/web_api_v1.h
+1
-1
@@ -4,7 +4,7 @@
4
#define NETDATA_WEB_API_V1_H 1
5
6
#include "daemon/common.h"
7
-#include "web_buffer_svg.h"
7
+#include "web/api/badges/web_buffer_svg.h"
8
#include "rrd2json.h"
9
10
extern int web_client_api_request_v1_data_group(char *name, int def);
web/server/README.md
+107
@@ -0,0 +1,107 @@
1
+# netdata web server
2
+
3
+netdata supports 3 implementation of its internal web server:
4
+
5
+- `static-threaded` is a web server with a fix (configured number of threads)
6
+- `single-threaded` is a simple web server running with a single thread
7
+- `multi-threaded` is a web server that spawns a thread for each client connection
8
+- `none` to disable the web server
9
+
10
+We suggest to use the `static-threaded` one. It is the most efficient.
11
+
12
+All versions of the web servers use non-blocking I/O.
13
+
14
+All web servers respect the `keep-alive` HTTP header to serve multiple HTTP requests via the same connection.
15
+
16
+
17
+## Configuration
18
+
19
+#### selecting the web server
20
+
21
+You can select the web server implementation by editing `netdata.conf` and setting:
22
+
23
+```
24
+[web]
25
+ mode = none | single-threaded | multi-threaded | static-threaded
26
+```
27
+
28
+The `static` web server supports also these settings:
29
+
30
+```
31
+[web]
32
+ mode = static-threaded
33
+ web server threads = 4
34
+ web server max sockets = 512
35
+```
36
+
37
+The default number of processor threads is `min(cpu cores, 6)`.
38
+
39
+The `web server max sockets` setting is automatically adjusted to 50% of the max number of open files
40
+netdata is allowed to use (via `/etc/security/limits.conf` or systemd), to allow enough file descriptors
41
+to be available for data collection.
42
+
43
+#### binding netdata to multiple ports
44
+
45
+netdata can bind to multiple IPs and ports. Up to 100 sockets can be used
46
+(you can increase it at compile time with `CFLAGS="-DMAX_LISTEN_FDS=200" ./netdata-installer.sh ...`).
47
+
48
+The ports to bind are controlled via `[web].bind to`, like this:
49
+
50
+```
51
+[web]
52
+ default port = 19999
53
+ bind to = 127.0.0.1 10.1.1.1:19998 hostname:19997 [::]:19996 localhost:19995 *:http unix:/tmp/netdata.sock
54
+```
55
+
56
+Using the above, netdata will bind to:
57
+ - IPv4 127.0.0.1 at port 19999 (port was used from `default port`)
58
+ - IPv4 10.1.1.1 at port 19998
59
+ - All the IPs `hostname` resolves to (both IPv4 and IPv6 depending on the resolved IPs) at port 19997
60
+ - All IPv6 IPs at port 19996
61
+ - All the IPs `localhost` resolves to (both IPv4 and IPv6 depending the resolved IPs) at port 19996
62
+ - All IPv4 and IPv6 IPs at port `http` as set in `/etc/services`
63
+ - Unix domain socket `/tmp/netdata.sock`
64
+
65
+The option `[web].default port` is used when an entries in `[web].bind to` do not specify a port.
66
+
67
+#### access lists
68
+
69
+Netdata supports access lists in `netdata.conf`:
70
+
71
+```
72
+[web]
73
+ allow connections from = localhost *
74
+ allow dashboard from = localhost *
75
+ allow badges from = *
76
+ allow streaming from = *
77
+ allow netdata.conf from = localhost fd* 10.* 192.168.* 172.16.* 172.17.* 172.18.* 172.19.* 172.20.* 172.21.* 172.22.* 172.23.* 172.24.* 172.25.* 172.26.* 172.27.* 172.28.* 172.29.* 172.30.* 172.31.*
78
+```
79
+
80
+`*` does string matches on the IPs of the clients.
81
+
82
+- `allow connections from` matches anyone that connects on the netdata port(s).
83
+ So, if someone is not allowed, it will be connected and disconnected immediately, without reading even
84
+ a single byte from its connection. This is a global settings with higher priority to any of the ones below.
85
+
86
+- `allow dashboard from` receives the request and examines if it is a static dashboard file or an API call the
87
+ dashboards do.
88
+
89
+- `allow badges from` checks if the API request is for a badge. Badges are not matched by `allow dashboard from`.
90
+
91
+- `allow streaming from` checks if the slave willing to stream metrics to this netdata is allowed.
92
+ This can be controlled per API KEY and MACHINE GUID in [stream.conf](../../streaming/stream.conf).
93
+ The setting in `netdata.conf` is checked before the ones in [stream.conf](../../streaming/stream.conf).
94
+
95
+- `allow netdata.conf from` checks the IP to allow `http://netdata.host:19999/netdata.conf`.
96
+ By default it allows only private lans.
97
+
98
+## DDoS protection
99
+
100
+If you publish your netdata to the internet, you may want to apply some protection against DDoS:
101
+
102
+1. Use the `static-threaded` web server (it is the default)
103
+2. Use reasonable `[web].web server max sockets` (the default is)
104
+3. Don't use all your cpu cores for netdata (lower `[web].web server threads`)
105
+4. Run netdata with a low process scheduling priority (the default is the lowest)
106
+5. If possible, proxy netdata via a full featured web server (nginx, apache, etc)
107
+
web/server/multi/README.md
+8
@@ -0,0 +1,8 @@
1
+# `multi-threaded` web server
2
+
3
+The `multi-threaded` web server spawns a thread for each connection it receives.
4
+
5
+Each thread uses non-blocking I/O so it can serve any number of web requests in parallel,
6
+though this is not supported by HTTP, so in practice each thread serves all the requests sequentially.
7
+
8
+Each thread respects the `keep-alive` HTTP header to serve multiple HTTP requests via the same connection.
\ No newline at end of file
web/server/single/README.md
+6
@@ -0,0 +1,6 @@
1
+# `single-threaded` web server
2
+
3
+The `single-threaded` web server runs as a single thread inside netdata.
4
+It uses non-blocking I/O so it can serve any number of web requests in parallel.
5
+
6
+This web server respects the `keep-alive` HTTP header to serve multiple HTTP requests via the same connection.
\ No newline at end of file
web/server/static/README.md
+9
@@ -0,0 +1,9 @@
1
+# `static-threaded` web server
2
+
3
+The `static-threaded` web server spawns a fixed number of threads.
4
+All the threads are concurrently listening for web requests on the same sockets.
5
+The kernel distributes the incoming requests to them.
6
+
7
+Each thread uses non-blocking I/O so it can serve any number of web requests in parallel.
8
+
9
+This web server respects the `keep-alive` HTTP header to serve multiple HTTP requests via the same connection.
\ No newline at end of file