@cryptotaxi247 / netdata-1 / commits / 112924d4e

Minor documentation improvements (#4566)

* Formatting updates in database/README.md * More formarring in README.md files * README.md formatting * Minor formatting change * Minor changes in registy/README.md * Minor formatting * Minor formatting change

George Moschovitis committed Nov 8, 2018 at 00:53 UTC 112924d4e7ec9680a1761dca2bf965d8fc31b847
5 files changed +49 -55
backends/prometheus/README.md
+14 -11
@@ -1,18 +1,20 @@
1 -> IMPORTANT: the format netdata sends metrics to prometheus has changed since netdata v1.7. The new prometheus backend for netdata supports a lot more features and is aligned to the development of the rest of the netdata backends.
2 -
1 # Using netdata with Prometheus
2
3 +> IMPORTANT: the format netdata sends metrics to prometheus has changed since netdata v1.7. The new prometheus backend for netdata supports a lot more features and is aligned to the development of the rest of the netdata backends.
4 +
5 Prometheus is a distributed monitoring system which offers a very simple setup along with a robust data model. Recently netdata added support for Prometheus. I'm going to quickly show you how to install both netdata and prometheus on the same server. We can then use grafana pointed at Prometheus to obtain long term metrics netdata offers. I'm assuming we are starting at a fresh ubuntu shell (whether you'd like to follow along in a VM or a cloud instance is up to you).
6
7 ## Installing netdata and prometheus
8
9 ### Installing netdata
10 +
11 There are number of ways to install netdata according to [Installation](https://github.com/netdata/netdata/wiki/Installation)
12 The suggested way of installing the latest netdata and keep it upgrade automatically. Using one line installation:
13
14 ```
15 bash <(curl -Ss https://my-netdata.io/kickstart.sh)
16 ```
17 +
18 At this point we should have netdata listening on port 19999. Attempt to take your browser here:
19
20 ```
@@ -22,15 +24,16 @@ http://your.netdata.ip:19999
24 *(replace `your.netdata.ip` with the IP or hostname of the server running netdata)*
25
26 ### Installing Prometheus
27 +
28 In order to install prometheus we are going to introduce our own systemd startup script along with an example of prometheus.yaml configuration. Prometheus needs to be pointed to your server at a specific target url for it to scrape netdata's api. Prometheus is always a pull model meaning netdata is the passive client within this architecture. Prometheus always initiates the connection with netdata.
29
27 -##### Download Prometheus
30 +#### Download Prometheus
31
32 ```sh
33 wget -O /tmp/prometheus-2.3.2.linux-amd64.tar.gz https://github.com/prometheus/prometheus/releases/download/v2.3.2/prometheus-2.3.2.linux-amd64.tar.gz
34 ```
35
33 -##### Create prometheus system user
36 +#### Create prometheus system user
37
38 ```sh
39 sudo useradd -r prometheus
@@ -104,6 +107,7 @@ scrape_configs:
107 static_configs:
108 - targets: ['{your.netdata.ip}:19999']
109 ```
110 +
111 #### Install nodes.yml
112
113 The following is completely optional, it will enable Prometheus to generate alerts from some NetData sources. Tweak the values to your own needs. We will use the following `nodes.yml` file below. Save it at `/opt/prometheus/nodes.yml`, and add a *- "nodes.yml"* entry under the *rule_files:* section in the example prometheus.yml file above.
@@ -166,7 +170,6 @@ ExecStop=/bin/kill -SIGINT $MAINPID
170 [Install]
171 WantedBy=multi-user.target
172 ```
169 -
173 ##### Start Prometheus
174
175 ```
@@ -180,7 +183,7 @@ If everything is working correctly when you fetch `http://your.prometheus.ip:909
183
184 ---
185
183 -## netdata support for prometheus
186 +## Netdata support for prometheus
187
188 > IMPORTANT: the format netdata sends metrics to prometheus has changed since netdata v1.6. The new format allows easier queries for metrics and supports both `as collected` and normalized metrics.
189
@@ -208,7 +211,7 @@ Then each netdata chart contains metrics called `dimensions`. All the dimensions
211
212 ### netdata data source
213
211 -netdata can send metrics to prometheus from 3 data sources:
214 +Netdata can send metrics to prometheus from 3 data sources:
215
216 - `as collected` or `raw` - this data source sends the metrics to prometheus as they are collected. No conversion is done by netdata. The latest value for each metric is just given to prometheus. This is the most preferred method by prometheus, but it is also the harder to work with. To work with this data source, you will need to understand how to get meaningful values out of them.
217
@@ -231,7 +234,6 @@ netdata can send metrics to prometheus from 3 data sources:
234
235 Keep in mind that early versions of netdata were sending the metrics as: `CHART_DIMENSION{}`.
236
234 -
237 ### Querying Metrics
238
239 Fetch with your web browser this URL:
@@ -298,6 +300,7 @@ netdata_system_cpu_total{chart="system.cpu",family="cpu",dimension="iowait"} 233
300 # COMMENT netdata_system_cpu_total: chart "system.cpu", context "system.cpu", family "cpu", dimension "idle", value * 1 / 1 delta gives percentage (counter)
301 netdata_system_cpu_total{chart="system.cpu",family="cpu",dimension="idle"} 918470 1500066716438
302 ```
303 +
304 *(netdata response for `system.cpu` with source=`as-collected`)*
305
306 For more information check prometheus documentation.
@@ -315,11 +318,11 @@ The `format=prometheus` parameter only exports the host's netdata metrics. If y
318
319 This will report all upstream host data, and `honor_labels` will make Prometheus take note of the instance names provided.
320
318 -### timestamps
321 +### Timestamps
322
323 To pass the metrics through prometheus pushgateway, netdata supports the option `&timestamps=no` to send the metrics without timestamps.
324
322 -## netdata host variables
325 +## Netdata host variables
326
327 netdata collects various system configuration metrics, like the max number of TCP sockets supported, the max number of files allowed system-wide, various IPC sizes, etc. These metrics are not exposed to prometheus by default.
328
@@ -369,7 +372,7 @@ netdata sends all metrics prefixed with `netdata_`. You can change this in `netd
372
373 It can also be changed from the URL, by appending `&prefix=netdata`.
374
372 -### accuracy of `average` and `sum` data sources
375 +### Accuracy of `average` and `sum` data sources
376
377 When the data source is set to `average` or `sum`, netdata remembers the last access of each client accessing prometheus metrics and uses this last access time to respond with the `average` or `sum` of all the entries in the database since that. This means that prometheus servers are not losing data when they access netdata with data source = `average` or `sum`.
378
database/README.md
+8 -8
@@ -1,4 +1,4 @@
1 -# netdata database
1 +# Netdata database
2
3 Although `netdata` does all its calculations using `long double`, it stores all values using
4 a [custom-made 32-bit number](../libnetdata/storage_number/).
@@ -26,17 +26,17 @@ Currently netdata supports 5 memory modes:
26
27 1. `ram`, data are purely in memory. Data are never saved on disk. This mode uses `mmap()` and
28 supports [KSM](#ksm).
29 -
29 +
30 2. `save`, (the default) data are only in RAM while netdata runs and are saved to / loaded from
31 disk on netdata restart. It also uses `mmap()` and supports [KSM](#ksm).
32 -
32 +
33 3. `map`, data are in memory mapped files. This works like the swap. Keep in mind though, this
34 will have a constant write on your disk. When netdata writes data on its memory, the Linux kernel
35 marks the related memory pages as dirty and automatically starts updating them on disk.
36 Unfortunately we cannot control how frequently this works. The Linux kernel uses exactly the
37 same algorithm it uses for its swap memory. Check below for additional information on running a
38 dedicated central netdata server. This mode uses `mmap()` but does not support [KSM](#ksm).
39 -
39 +
40 4. `none`, without a database (collected metrics can only be streamed to another netdata).
41
42 5. `alloc`, like `ram` but it uses `calloc()` and does not support [KSM](#ksm). This mode is the
@@ -73,9 +73,9 @@ by netdata. Of course experiment a bit. On very weak devices you might have to u
73 You can also disable [data collection plugins](../collectors) you don't need.
74 Disabling such plugins will also free both CPU and RAM resources.
75
76 -## running a dedicated central netdata server
76 +## Running a dedicated central netdata server
77
78 -netdata allows streaming data between netdata nodes. This allows us to have a central netdata
78 +Netdata allows streaming data between netdata nodes. This allows us to have a central netdata
79 server that will maintain the entire database for all nodes, and will also run health checks/alarms
80 for all nodes.
81
@@ -166,7 +166,7 @@ netdata, each byte at the in-memory database will be updated just once per day).
166
167 KSM is a solution that will provide 60+% memory savings to netdata.
168
169 -#### Enable KSM in kernel
169 +### Enable KSM in kernel
170
171 You need to run a kernel compiled with:
172
@@ -186,7 +186,7 @@ The files that `CONFIG_KSM=y` offers include:
186
187 So, by default `ksmd` is just disabled. It will not harm performance and the user/admin can control the CPU resources he/she is willing `ksmd` to use.
188
189 -#### Run `ksmd` kernel daemon
189 +### Run `ksmd` kernel daemon
190
191 To activate / run `ksmd` you need to run:
192
health/README.md
+21 -27
@@ -1,4 +1,3 @@
1 -
1 # Health monitoring
2
3 Each netdata node runs an independent thread evaluating health monitoring checks.
@@ -40,16 +39,16 @@ killall -USR2 netdata
39
40 There are 2 entities:
41
43 -1. **alarms**, which are attached to specific charts, and
42 +1. **alarms**, which are attached to specific charts, and
43
45 -2. **templates**, which define rules that should be applied to all charts having a
44 +1. **templates**, which define rules that should be applied to all charts having a
45 specific `context`. You can use this feature to apply **alarms** to all disks,
46 all network interfaces, all mysql databases, all nginx web servers, etc.
47
48 Both of these entities have exactly the same format and feature set.
49 The only difference is the label `alarm` or `template`.
50
52 -netdata supports overriding **templates** with **alarms**.
51 +Netdata supports overriding **templates** with **alarms**.
52 For example, when a template is defined for a set of charts, an alarm with exactly the
53 same name attached to the same chart the template matches, will have higher precedence
54 (i.e. netdata will use the alarm on this chart and prevent the template from being applied
@@ -59,7 +58,7 @@ to it).
58
59 The following lines are parsed.
60
62 -#### alarm line `alarm` or `template`
61 +#### Alarm line `alarm` or `template`
62
63 This line starts an alarm or alarm template.
64
@@ -78,7 +77,7 @@ This line has to be first on each alarm or template.
77
78 ---
79
81 -#### alarm line `on`
80 +#### Alarm line `on`
81
82 This line defines the data the alarm should be attached to.
83
@@ -112,7 +111,7 @@ So, `plugin = proc`, `module = /proc/net/dev` and `context = net.net`.
111
112 ---
113
115 -#### alarm line `os`
114 +#### Alarm line `os`
115
116 This alarm or template will be used only if the O/S of the host loading it, matches this
117 pattern list. The value is a space separated list of simple patterns (use `*` as wildcard,
@@ -124,7 +123,7 @@ os: linux freebsd macos
123
124 ---
125
127 -#### alarm line `hosts`
126 +#### Alarm line `hosts`
127
128 This alarm or template will be used only if the hostname of the host loading it, matches
129 this pattern list. The value is a space separated list of simple patterns (use `*` as wildcard,
@@ -141,7 +140,7 @@ This is useful when you centralize metrics from multiple hosts, to one netdata.
140
141 ---
142
144 -#### alarm line `families`
143 +#### Alarm line `families`
144
145 This line is only used in alarm templates. It filters the charts. So, if you need to create
146 an alarm template for a few of a kind of chart (a few of your disks, or a few of your network
@@ -165,7 +164,7 @@ The family of a chart is usually the submenu of the netdata dashboard it appears
164
165 ---
166
168 -#### alarm line `lookup`
167 +#### Alarm line `lookup`
168
169 This lines makes a database lookup to find a value. This result of this lookup is available as `$this`.
170
@@ -205,7 +204,7 @@ The timestamps of the timeframe evaluated by the database lookup is available as
204
205 ---
206
208 -#### alarm line `calc`
207 +#### Alarm line `calc`
208
209 This expression is evaluated just after the `lookup` (if any). Its purpose is to apply some
210 calculation before using the value looked up from the db.
@@ -225,7 +224,7 @@ Check [Expressions](#expressions) for more information.
224
225 ---
226
228 -#### alarm line `every`
227 +#### Alarm line `every`
228
229 Sets the update frequency of this alarm. This is the same to the `every DURATION` given
230 in the `lookup` lines.
@@ -240,7 +239,7 @@ every: DURATION
239
240 ---
241
243 -#### alarm lines `green` and `red`
242 +#### Alarm lines `green` and `red`
243
244 Set the green and red thresholds of a chart. Both are available as `$green` and `$red` in
245 expressions. If multiple alarms define different thresholds, the ones defined by the first
@@ -257,7 +256,7 @@ red: NUMBER
256
257 ---
258
260 -#### alarm lines `warn` and `crit`
259 +#### Alarm lines `warn` and `crit`
260
261 These expressions should evaluate to true or false (alternatively non-zero or zero).
262 They trigger the alarm. Both are optional.
@@ -272,7 +271,7 @@ Check [Expressions](#expressions) for more information.
271
272 ---
273
275 -#### alarm line `to`
274 +#### Alarm line `to`
275
276 This will be the first parameter of the script to be executed when the alarm switches status.
277 Its meaning is left up to the `exec` script.
@@ -288,7 +287,7 @@ to: ROLE1 ROLE2 ROLE3 ...
287
288 ---
289
291 -#### alarm line `exec`
290 +#### Alarm line `exec`
291
292 The script that will be executed when the alarm changes status.
293
@@ -303,7 +302,7 @@ methods netdata supports, including custom hooks.
302
303 ---
304
306 -#### alarm line `delay`
305 +#### Alarm line `delay`
306
307 This is used to provide optional hysteresis settings for the notifications, to defend
308 against notification floods. These settings do not affect the actual alarm - only the time
@@ -374,13 +373,9 @@ Expressions can have variables. Variables start with `$`. Check below for more i
373
374 There are two special values you can use:
375
377 - - `nan`, for example `$this != nan` will check if the variable `this` is available.
378 - A variable can be `nan` if the database lookup failed. All calculations (i.e. addition,
379 - multiplication, etc) with a `nan` result in a `nan`.
376 +- `nan`, for example `$this != nan` will check if the variable `this` is available. A variable can be `nan` if the database lookup failed. All calculations (i.e. addition, multiplication, etc) with a `nan` result in a `nan`.
377
381 - - `inf`, for example `$this != inf` will check if `this` is not infinite. A value or
382 - variable can be infinite if divided by zero. All calculations (i.e. addition,
383 - multiplication, etc) with a `inf` result in a `inf`.
378 +- `inf`, for example `$this != inf` will check if `this` is not infinite. A value or variable can be infinite if divided by zero. All calculations (i.e. addition, multiplication, etc) with a `inf` result in a `inf`.
379
380 ---
381
@@ -412,10 +407,10 @@ Which in turn, results in the following behavior:
407
408 * While the value is falling, it will return to a warning state when it goes below 85,
409 and a normal state when it goes below 75.
415 -
410 +
411 * If the value is constantly varying between 80 and 90, then it will trigger a warning the
412 first time it goes above 85, but will remain a warning until it goes below 75 (or goes above 85).
418 -
413 +
414 * If the value is constantly varying between 90 and 100, then it will trigger a critical alert
415 the first time it goes above 95, but will remain a critical alert goes below 85 (at which
416 point it will return to being a warning).
@@ -653,5 +648,4 @@ You can find the context of charts by looking up the chart in either
648 You can find how netdata interpreted the expressions by examining the alarm at
649 `http://your.netdata:19999/api/v1/alarms?all`. For each expression, netdata will return the
650 expression as given in its config file, and the same expression with additional parentheses
656 -added to indicate the evaluation flow of the expression.
657 -
651 +added to indicate the evaluation flow of the expression.
registry/README.md
+2 -2
@@ -36,11 +36,11 @@ The registry keeps track of 3 entities:
36
37 For each netdata installation (each `machine_guid`) the registry keeps track of the different URLs it is accessed.
38
39 -2. **persons**: i.e. the web browsers accessing the netdata installations (a random GUID generated by the registry the first time it sees a new web browser; we call this **person_guid**)
39 +1. **persons**: i.e. the web browsers accessing the netdata installations (a random GUID generated by the registry the first time it sees a new web browser; we call this **person_guid**)
40
41 For each person, the registry keeps track of the netdata installations it has accessed and their URLs.
42
43 -3. **URLs** of netdata installations (as seen by the web browsers)
43 +1. **URLs** of netdata installations (as seen by the web browsers)
44
45 For each URL, the registry keeps the URL and nothing more. Each URL is linked to *persons* and *machines*. The only way to find a URL is to know its **machine_guid** or have a **person_guid** it is linked to it.
46
streaming/README.md
+4 -7
@@ -13,7 +13,7 @@ a netdata performs:
13
14 The following configurations are supported:
15
16 -#### netdata without a database or web API (headless collector)
16 +#### Netdata without a database or web API (headless collector)
17
18 Local netdata (`slave`), **without any database or alarms**, collects metrics and sends them to
19 another netdata (`master`).
@@ -28,7 +28,7 @@ of maintaining a local database and accepting dashboard requests, it streams all
28
29 The same `master` can collect data for any number of `slaves`.
30
31 -#### database replication
31 +#### Database replication
32
33 Local netdata (`slave`), **with a local database (and possibly alarms)**, collects metrics and
34 sends them to another netdata (`master`).
@@ -306,10 +306,10 @@ On each of the slaves, edit `/etc/netdata/stream.conf` (to edit it on your syste
306 [stream]
307 # stream metrics to another netdata
308 enabled = yes
309 -
309 +
310 # the IP and PORT of the master
311 destination = 10.11.12.13:19999
312 -
312 +
313 # the API key to use
314 api key = 11111111-2222-3333-4444-555555555555
315 ```
@@ -340,7 +340,6 @@ The file `/var/lib/netdata/registry/netdata.public.unique.id` contains a random
340
341 Both the sender and the receiver of metrics log information at `/var/log/netdata/error.log`.
342
343 -
343 On both master and slave do this:
344
345 ```
@@ -394,7 +393,6 @@ This means a setup like the following is also possible:
393 <img src="https://cloud.githubusercontent.com/assets/2662304/23629551/bb1fd9c2-02c0-11e7-90f5-cab5a3ed4c53.png"/>
394 </p>
395
397 -
396 ## proxies
397
398 A proxy is a netdata that is receiving metrics from a netdata, and streams them to another netdata.
@@ -410,4 +408,3 @@ The sending side of a netdata proxy, connects and disconnects to the final desti
408 metrics, following the same pattern of the receiving side.
409
410 For a practical example see [Monitoring ephemeral nodes](#monitoring-ephemeral-nodes).
413 -