Minor documentation improvements (#4566)
* Formatting updates in database/README.md * More formarring in README.md files * README.md formatting * Minor formatting change * Minor changes in registy/README.md * Minor formatting * Minor formatting change
George Moschovitis committed
Nov 8, 2018 at 00:53 UTC
112924d4e7ec9680a1761dca2bf965d8fc31b847
5 files changed
+49
-55
backends/prometheus/README.md
+14
-11
@@ -1,18 +1,20 @@
1
-> IMPORTANT: the format netdata sends metrics to prometheus has changed since netdata v1.7. The new prometheus backend for netdata supports a lot more features and is aligned to the development of the rest of the netdata backends.
2
-
1
# Using netdata with Prometheus
2
3
+> IMPORTANT: the format netdata sends metrics to prometheus has changed since netdata v1.7. The new prometheus backend for netdata supports a lot more features and is aligned to the development of the rest of the netdata backends.
4
+
5
Prometheus is a distributed monitoring system which offers a very simple setup along with a robust data model. Recently netdata added support for Prometheus. I'm going to quickly show you how to install both netdata and prometheus on the same server. We can then use grafana pointed at Prometheus to obtain long term metrics netdata offers. I'm assuming we are starting at a fresh ubuntu shell (whether you'd like to follow along in a VM or a cloud instance is up to you).
6
7
## Installing netdata and prometheus
8
9
### Installing netdata
10
+
11
There are number of ways to install netdata according to [Installation](https://github.com/netdata/netdata/wiki/Installation)
12
The suggested way of installing the latest netdata and keep it upgrade automatically. Using one line installation:
13
14
```
15
bash <(curl -Ss https://my-netdata.io/kickstart.sh)
16
```
17
+
18
At this point we should have netdata listening on port 19999. Attempt to take your browser here:
19
20
```
@@ -22,15 +24,16 @@ http://your.netdata.ip:19999
24
*(replace `your.netdata.ip` with the IP or hostname of the server running netdata)*
25
26
### Installing Prometheus
27
+
28
In order to install prometheus we are going to introduce our own systemd startup script along with an example of prometheus.yaml configuration. Prometheus needs to be pointed to your server at a specific target url for it to scrape netdata's api. Prometheus is always a pull model meaning netdata is the passive client within this architecture. Prometheus always initiates the connection with netdata.
29
27
-##### Download Prometheus
30
+#### Download Prometheus
31
32
```sh
33
wget -O /tmp/prometheus-2.3.2.linux-amd64.tar.gz https://github.com/prometheus/prometheus/releases/download/v2.3.2/prometheus-2.3.2.linux-amd64.tar.gz
34
```
35
33
-##### Create prometheus system user
36
+#### Create prometheus system user
37
38
```sh
39
sudo useradd -r prometheus
@@ -104,6 +107,7 @@ scrape_configs:
107
static_configs:
108
- targets: ['{your.netdata.ip}:19999']
109
```
110
+
111
#### Install nodes.yml
112
113
The following is completely optional, it will enable Prometheus to generate alerts from some NetData sources. Tweak the values to your own needs. We will use the following `nodes.yml` file below. Save it at `/opt/prometheus/nodes.yml`, and add a *- "nodes.yml"* entry under the *rule_files:* section in the example prometheus.yml file above.
@@ -166,7 +170,6 @@ ExecStop=/bin/kill -SIGINT $MAINPID
170
[Install]
171
WantedBy=multi-user.target
172
```
169
-
173
##### Start Prometheus
174
175
```
@@ -180,7 +183,7 @@ If everything is working correctly when you fetch `http://your.prometheus.ip:909
183
184
---
185
183
-## netdata support for prometheus
186
+## Netdata support for prometheus
187
188
> IMPORTANT: the format netdata sends metrics to prometheus has changed since netdata v1.6. The new format allows easier queries for metrics and supports both `as collected` and normalized metrics.
189
@@ -208,7 +211,7 @@ Then each netdata chart contains metrics called `dimensions`. All the dimensions
211
212
### netdata data source
213
211
-netdata can send metrics to prometheus from 3 data sources:
214
+Netdata can send metrics to prometheus from 3 data sources:
215
216
- `as collected` or `raw` - this data source sends the metrics to prometheus as they are collected. No conversion is done by netdata. The latest value for each metric is just given to prometheus. This is the most preferred method by prometheus, but it is also the harder to work with. To work with this data source, you will need to understand how to get meaningful values out of them.
217
@@ -231,7 +234,6 @@ netdata can send metrics to prometheus from 3 data sources:
234
235
Keep in mind that early versions of netdata were sending the metrics as: `CHART_DIMENSION{}`.
236
234
-
237
### Querying Metrics
238
239
Fetch with your web browser this URL:
@@ -298,6 +300,7 @@ netdata_system_cpu_total{chart="system.cpu",family="cpu",dimension="iowait"} 233
300
# COMMENT netdata_system_cpu_total: chart "system.cpu", context "system.cpu", family "cpu", dimension "idle", value * 1 / 1 delta gives percentage (counter)
301
netdata_system_cpu_total{chart="system.cpu",family="cpu",dimension="idle"} 918470 1500066716438
302
```
303
+
304
*(netdata response for `system.cpu` with source=`as-collected`)*
305
306
For more information check prometheus documentation.
@@ -315,11 +318,11 @@ The `format=prometheus` parameter only exports the host's netdata metrics. If y
318
319
This will report all upstream host data, and `honor_labels` will make Prometheus take note of the instance names provided.
320
318
-### timestamps
321
+### Timestamps
322
323
To pass the metrics through prometheus pushgateway, netdata supports the option `×tamps=no` to send the metrics without timestamps.
324
322
-## netdata host variables
325
+## Netdata host variables
326
327
netdata collects various system configuration metrics, like the max number of TCP sockets supported, the max number of files allowed system-wide, various IPC sizes, etc. These metrics are not exposed to prometheus by default.
328
@@ -369,7 +372,7 @@ netdata sends all metrics prefixed with `netdata_`. You can change this in `netd
372
373
It can also be changed from the URL, by appending `&prefix=netdata`.
374
372
-### accuracy of `average` and `sum` data sources
375
+### Accuracy of `average` and `sum` data sources
376
377
When the data source is set to `average` or `sum`, netdata remembers the last access of each client accessing prometheus metrics and uses this last access time to respond with the `average` or `sum` of all the entries in the database since that. This means that prometheus servers are not losing data when they access netdata with data source = `average` or `sum`.
378
database/README.md
+8
-8
@@ -1,4 +1,4 @@
1
-# netdata database
1
+# Netdata database
2
3
Although `netdata` does all its calculations using `long double`, it stores all values using
4
a [custom-made 32-bit number](../libnetdata/storage_number/).
@@ -26,17 +26,17 @@ Currently netdata supports 5 memory modes:
26
27
1. `ram`, data are purely in memory. Data are never saved on disk. This mode uses `mmap()` and
28
supports [KSM](#ksm).
29
-
29
+
30
2. `save`, (the default) data are only in RAM while netdata runs and are saved to / loaded from
31
disk on netdata restart. It also uses `mmap()` and supports [KSM](#ksm).
32
-
32
+
33
3. `map`, data are in memory mapped files. This works like the swap. Keep in mind though, this
34
will have a constant write on your disk. When netdata writes data on its memory, the Linux kernel
35
marks the related memory pages as dirty and automatically starts updating them on disk.
36
Unfortunately we cannot control how frequently this works. The Linux kernel uses exactly the
37
same algorithm it uses for its swap memory. Check below for additional information on running a
38
dedicated central netdata server. This mode uses `mmap()` but does not support [KSM](#ksm).
39
-
39
+
40
4. `none`, without a database (collected metrics can only be streamed to another netdata).
41
42
5. `alloc`, like `ram` but it uses `calloc()` and does not support [KSM](#ksm). This mode is the
@@ -73,9 +73,9 @@ by netdata. Of course experiment a bit. On very weak devices you might have to u
73
You can also disable [data collection plugins](../collectors) you don't need.
74
Disabling such plugins will also free both CPU and RAM resources.
75
76
-## running a dedicated central netdata server
76
+## Running a dedicated central netdata server
77
78
-netdata allows streaming data between netdata nodes. This allows us to have a central netdata
78
+Netdata allows streaming data between netdata nodes. This allows us to have a central netdata
79
server that will maintain the entire database for all nodes, and will also run health checks/alarms
80
for all nodes.
81
@@ -166,7 +166,7 @@ netdata, each byte at the in-memory database will be updated just once per day).
166
167
KSM is a solution that will provide 60+% memory savings to netdata.
168
169
-#### Enable KSM in kernel
169
+### Enable KSM in kernel
170
171
You need to run a kernel compiled with:
172
@@ -186,7 +186,7 @@ The files that `CONFIG_KSM=y` offers include:
186
187
So, by default `ksmd` is just disabled. It will not harm performance and the user/admin can control the CPU resources he/she is willing `ksmd` to use.
188
189
-#### Run `ksmd` kernel daemon
189
+### Run `ksmd` kernel daemon
190
191
To activate / run `ksmd` you need to run:
192
health/README.md
+21
-27
@@ -1,4 +1,3 @@
1
-
1
# Health monitoring
2
3
Each netdata node runs an independent thread evaluating health monitoring checks.
@@ -40,16 +39,16 @@ killall -USR2 netdata
39
40
There are 2 entities:
41
43
-1. **alarms**, which are attached to specific charts, and
42
+1. **alarms**, which are attached to specific charts, and
43
45
-2. **templates**, which define rules that should be applied to all charts having a
44
+1. **templates**, which define rules that should be applied to all charts having a
45
specific `context`. You can use this feature to apply **alarms** to all disks,
46
all network interfaces, all mysql databases, all nginx web servers, etc.
47
48
Both of these entities have exactly the same format and feature set.
49
The only difference is the label `alarm` or `template`.
50
52
-netdata supports overriding **templates** with **alarms**.
51
+Netdata supports overriding **templates** with **alarms**.
52
For example, when a template is defined for a set of charts, an alarm with exactly the
53
same name attached to the same chart the template matches, will have higher precedence
54
(i.e. netdata will use the alarm on this chart and prevent the template from being applied
@@ -59,7 +58,7 @@ to it).
58
59
The following lines are parsed.
60
62
-#### alarm line `alarm` or `template`
61
+#### Alarm line `alarm` or `template`
62
63
This line starts an alarm or alarm template.
64
@@ -78,7 +77,7 @@ This line has to be first on each alarm or template.
77
78
---
79
81
-#### alarm line `on`
80
+#### Alarm line `on`
81
82
This line defines the data the alarm should be attached to.
83
@@ -112,7 +111,7 @@ So, `plugin = proc`, `module = /proc/net/dev` and `context = net.net`.
111
112
---
113
115
-#### alarm line `os`
114
+#### Alarm line `os`
115
116
This alarm or template will be used only if the O/S of the host loading it, matches this
117
pattern list. The value is a space separated list of simple patterns (use `*` as wildcard,
@@ -124,7 +123,7 @@ os: linux freebsd macos
123
124
---
125
127
-#### alarm line `hosts`
126
+#### Alarm line `hosts`
127
128
This alarm or template will be used only if the hostname of the host loading it, matches
129
this pattern list. The value is a space separated list of simple patterns (use `*` as wildcard,
@@ -141,7 +140,7 @@ This is useful when you centralize metrics from multiple hosts, to one netdata.
140
141
---
142
144
-#### alarm line `families`
143
+#### Alarm line `families`
144
145
This line is only used in alarm templates. It filters the charts. So, if you need to create
146
an alarm template for a few of a kind of chart (a few of your disks, or a few of your network
@@ -165,7 +164,7 @@ The family of a chart is usually the submenu of the netdata dashboard it appears
164
165
---
166
168
-#### alarm line `lookup`
167
+#### Alarm line `lookup`
168
169
This lines makes a database lookup to find a value. This result of this lookup is available as `$this`.
170
@@ -205,7 +204,7 @@ The timestamps of the timeframe evaluated by the database lookup is available as
204
205
---
206
208
-#### alarm line `calc`
207
+#### Alarm line `calc`
208
209
This expression is evaluated just after the `lookup` (if any). Its purpose is to apply some
210
calculation before using the value looked up from the db.
@@ -225,7 +224,7 @@ Check [Expressions](#expressions) for more information.
224
225
---
226
228
-#### alarm line `every`
227
+#### Alarm line `every`
228
229
Sets the update frequency of this alarm. This is the same to the `every DURATION` given
230
in the `lookup` lines.
@@ -240,7 +239,7 @@ every: DURATION
239
240
---
241
243
-#### alarm lines `green` and `red`
242
+#### Alarm lines `green` and `red`
243
244
Set the green and red thresholds of a chart. Both are available as `$green` and `$red` in
245
expressions. If multiple alarms define different thresholds, the ones defined by the first
@@ -257,7 +256,7 @@ red: NUMBER
256
257
---
258
260
-#### alarm lines `warn` and `crit`
259
+#### Alarm lines `warn` and `crit`
260
261
These expressions should evaluate to true or false (alternatively non-zero or zero).
262
They trigger the alarm. Both are optional.
@@ -272,7 +271,7 @@ Check [Expressions](#expressions) for more information.
271
272
---
273
275
-#### alarm line `to`
274
+#### Alarm line `to`
275
276
This will be the first parameter of the script to be executed when the alarm switches status.
277
Its meaning is left up to the `exec` script.
@@ -288,7 +287,7 @@ to: ROLE1 ROLE2 ROLE3 ...
287
288
---
289
291
-#### alarm line `exec`
290
+#### Alarm line `exec`
291
292
The script that will be executed when the alarm changes status.
293
@@ -303,7 +302,7 @@ methods netdata supports, including custom hooks.
302
303
---
304
306
-#### alarm line `delay`
305
+#### Alarm line `delay`
306
307
This is used to provide optional hysteresis settings for the notifications, to defend
308
against notification floods. These settings do not affect the actual alarm - only the time
@@ -374,13 +373,9 @@ Expressions can have variables. Variables start with `$`. Check below for more i
373
374
There are two special values you can use:
375
377
- - `nan`, for example `$this != nan` will check if the variable `this` is available.
378
- A variable can be `nan` if the database lookup failed. All calculations (i.e. addition,
379
- multiplication, etc) with a `nan` result in a `nan`.
376
+- `nan`, for example `$this != nan` will check if the variable `this` is available. A variable can be `nan` if the database lookup failed. All calculations (i.e. addition, multiplication, etc) with a `nan` result in a `nan`.
377
381
- - `inf`, for example `$this != inf` will check if `this` is not infinite. A value or
382
- variable can be infinite if divided by zero. All calculations (i.e. addition,
383
- multiplication, etc) with a `inf` result in a `inf`.
378
+- `inf`, for example `$this != inf` will check if `this` is not infinite. A value or variable can be infinite if divided by zero. All calculations (i.e. addition, multiplication, etc) with a `inf` result in a `inf`.
379
380
---
381
@@ -412,10 +407,10 @@ Which in turn, results in the following behavior:
407
408
* While the value is falling, it will return to a warning state when it goes below 85,
409
and a normal state when it goes below 75.
415
-
410
+
411
* If the value is constantly varying between 80 and 90, then it will trigger a warning the
412
first time it goes above 85, but will remain a warning until it goes below 75 (or goes above 85).
418
-
413
+
414
* If the value is constantly varying between 90 and 100, then it will trigger a critical alert
415
the first time it goes above 95, but will remain a critical alert goes below 85 (at which
416
point it will return to being a warning).
@@ -653,5 +648,4 @@ You can find the context of charts by looking up the chart in either
648
You can find how netdata interpreted the expressions by examining the alarm at
649
`http://your.netdata:19999/api/v1/alarms?all`. For each expression, netdata will return the
650
expression as given in its config file, and the same expression with additional parentheses
656
-added to indicate the evaluation flow of the expression.
657
-
651
+added to indicate the evaluation flow of the expression.
registry/README.md
+2
-2
@@ -36,11 +36,11 @@ The registry keeps track of 3 entities:
36
37
For each netdata installation (each `machine_guid`) the registry keeps track of the different URLs it is accessed.
38
39
-2. **persons**: i.e. the web browsers accessing the netdata installations (a random GUID generated by the registry the first time it sees a new web browser; we call this **person_guid**)
39
+1. **persons**: i.e. the web browsers accessing the netdata installations (a random GUID generated by the registry the first time it sees a new web browser; we call this **person_guid**)
40
41
For each person, the registry keeps track of the netdata installations it has accessed and their URLs.
42
43
-3. **URLs** of netdata installations (as seen by the web browsers)
43
+1. **URLs** of netdata installations (as seen by the web browsers)
44
45
For each URL, the registry keeps the URL and nothing more. Each URL is linked to *persons* and *machines*. The only way to find a URL is to know its **machine_guid** or have a **person_guid** it is linked to it.
46
streaming/README.md
+4
-7
@@ -13,7 +13,7 @@ a netdata performs:
13
14
The following configurations are supported:
15
16
-#### netdata without a database or web API (headless collector)
16
+#### Netdata without a database or web API (headless collector)
17
18
Local netdata (`slave`), **without any database or alarms**, collects metrics and sends them to
19
another netdata (`master`).
@@ -28,7 +28,7 @@ of maintaining a local database and accepting dashboard requests, it streams all
28
29
The same `master` can collect data for any number of `slaves`.
30
31
-#### database replication
31
+#### Database replication
32
33
Local netdata (`slave`), **with a local database (and possibly alarms)**, collects metrics and
34
sends them to another netdata (`master`).
@@ -306,10 +306,10 @@ On each of the slaves, edit `/etc/netdata/stream.conf` (to edit it on your syste
306
[stream]
307
# stream metrics to another netdata
308
enabled = yes
309
-
309
+
310
# the IP and PORT of the master
311
destination = 10.11.12.13:19999
312
-
312
+
313
# the API key to use
314
api key = 11111111-2222-3333-4444-555555555555
315
```
@@ -340,7 +340,6 @@ The file `/var/lib/netdata/registry/netdata.public.unique.id` contains a random
340
341
Both the sender and the receiver of metrics log information at `/var/log/netdata/error.log`.
342
343
-
343
On both master and slave do this:
344
345
```
@@ -394,7 +393,6 @@ This means a setup like the following is also possible:
393
<img src="https://cloud.githubusercontent.com/assets/2662304/23629551/bb1fd9c2-02c0-11e7-90f5-cab5a3ed4c53.png"/>
394
</p>
395
397
-
396
## proxies
397
398
A proxy is a netdata that is receiving metrics from a netdata, and streams them to another netdata.
@@ -410,4 +408,3 @@ The sending side of a netdata proxy, connects and disconnects to the final desti
408
metrics, following the same pattern of the receiving side.
409
410
For a practical example see [Monitoring ephemeral nodes](#monitoring-ephemeral-nodes).
413
-