Fix Remark Lint Warnings for Backends (#6917)
* fix remark warnings for AWS Kinesis doc * fix remark warnings for MongoDB * fix remark warnings on OpenTSDB doc * fix lint warnings on prometheus docs * format prometheus page * fix main backend readme * fix remark warnings except local links * remove slash in prometheus doc * remove link to header to fix lint error * make character limit to 120 not 80
Promise Akpan committed
Oct 3, 2019 at 18:21 UTC
5d7b0043d5a0196e2eeb79bd68a65f9ff7a54630
7 files changed
+358
-332
backends/README.md
+93
-94
@@ -1,26 +1,25 @@
1
# Metrics long term archiving
2
3
-Netdata supports backends for archiving the metrics, or providing long term dashboards,
4
-using Grafana or other tools, like this:
3
+Netdata supports backends for archiving the metrics, or providing long term dashboards, using Grafana or other tools,
4
+like this:
5
6

7
8
-Since Netdata collects thousands of metrics per server per second, which would easily congest any backend
9
-server when several Netdata servers are sending data to it, Netdata allows sending metrics at a lower
10
-frequency, by resampling them.
8
+Since Netdata collects thousands of metrics per server per second, which would easily congest any backend server when
9
+several Netdata servers are sending data to it, Netdata allows sending metrics at a lower frequency, by resampling them.
10
12
-So, although Netdata collects metrics every second, it can send to the backend servers averages or sums every
13
-X seconds (though, it can send them per second if you need it to).
11
+So, although Netdata collects metrics every second, it can send to the backend servers averages or sums every X seconds
12
+(though, it can send them per second if you need it to).
13
14
## features
15
16
1. Supported backends
17
19
- - **graphite** (`plaintext interface`, used by **Graphite**, **InfluxDB**, **KairosDB**,
20
- **Blueflood**, **ElasticSearch** via logstash tcp input and the graphite codec, etc)
18
+ - **graphite** (`plaintext interface`, used by **Graphite**, **InfluxDB**, **KairosDB**, **Blueflood**,
19
+ **ElasticSearch** via logstash tcp input and the graphite codec, etc)
20
22
- metrics are sent to the backend server as `prefix.hostname.chart.dimension`. `prefix` is
23
- configured below, `hostname` is the hostname of the machine (can also be configured).
21
+ metrics are sent to the backend server as `prefix.hostname.chart.dimension`. `prefix` is configured below,
22
+ `hostname` is the hostname of the machine (can also be configured).
23
24
- **opentsdb** (`telnet or HTTP interfaces`, used by **OpenTSDB**, **InfluxDB**, **KairosDB**, etc)
25
@@ -33,12 +32,12 @@ X seconds (though, it can send them per second if you need it to).
32
- **prometheus** is described at [prometheus page](prometheus/) since it pulls data from Netdata.
33
34
- **prometheus remote write** (a binary snappy-compressed protocol buffer encoding over HTTP used by
36
- **Elasticsearch**, **Gnocchi**, **Graphite**, **InfluxDB**, **Kafka**, **OpenTSDB**,
37
- **PostgreSQL/TimescaleDB**, **Splunk**, **VictoriaMetrics**,
38
- and a lot of other [storage providers](https://prometheus.io/docs/operating/integrations/#remote-endpoints-and-storage))
35
+ **Elasticsearch**, **Gnocchi**, **Graphite**, **InfluxDB**, **Kafka**, **OpenTSDB**, **PostgreSQL/TimescaleDB**,
36
+ **Splunk**, **VictoriaMetrics**, and a lot of other [storage
37
+ providers](https://prometheus.io/docs/operating/integrations/#remote-endpoints-and-storage))
38
40
- metrics are labeled in the format, which is used by Netdata for the [plaintext prometheus protocol](prometheus/).
41
- Notes on using the remote write backend are [here](prometheus/remote_write/).
39
+ metrics are labeled in the format, which is used by Netdata for the [plaintext prometheus
40
+ protocol](prometheus/). Notes on using the remote write backend are [here](prometheus/remote_write/).
41
42
- **AWS Kinesis Data Streams**
43
@@ -54,32 +53,37 @@ X seconds (though, it can send them per second if you need it to).
53
54
4. Netdata supports three modes of operation for all backends:
55
57
- - `as-collected` sends to backends the metrics as they are collected, in the units they are collected.
58
- So, counters are sent as counters and gauges are sent as gauges, much like all data collectors do.
59
- For example, to calculate CPU utilization in this format, you need to know how to convert kernel ticks to percentage.
56
+ - `as-collected` sends to backends the metrics as they are collected, in the units they are collected. So,
57
+ counters are sent as counters and gauges are sent as gauges, much like all data collectors do. For example, to
58
+ calculate CPU utilization in this format, you need to know how to convert kernel ticks to percentage.
59
61
- - `average` sends to backends normalized metrics from the Netdata database.
62
- In this mode, all metrics are sent as gauges, in the units Netdata uses. This abstracts data collection
63
- and simplifies visualization, but you will not be able to copy and paste queries from other sources to convert units.
64
- For example, CPU utilization percentage is calculated by Netdata, so Netdata will convert ticks to percentage and
65
- send the average percentage to the backend.
60
+ - `average` sends to backends normalized metrics from the Netdata database. In this mode, all metrics are sent as
61
+ gauges, in the units Netdata uses. This abstracts data collection and simplifies visualization, but you will not
62
+ be able to copy and paste queries from other sources to convert units. For example, CPU utilization percentage
63
+ is calculated by Netdata, so Netdata will convert ticks to percentage and send the average percentage to the
64
+ backend.
65
67
- - `sum` or `volume`: the sum of the interpolated values shown on the Netdata graphs is sent to the backend.
68
- So, if Netdata is configured to send data to the backend every 10 seconds, the sum of the 10 values shown on the
66
+ - `sum` or `volume`: the sum of the interpolated values shown on the Netdata graphs is sent to the backend. So, if
67
+ Netdata is configured to send data to the backend every 10 seconds, the sum of the 10 values shown on the
68
Netdata charts will be used.
69
71
-Time-series databases suggest to collect the raw values (`as-collected`). If you plan to invest on building your monitoring around a time-series database and you already know (or you will invest in learning) how to convert units and normalize the metrics in Grafana or other visualization tools, we suggest to use `as-collected`.
70
+ Time-series databases suggest to collect the raw values (`as-collected`). If you plan to invest on building your
71
+ monitoring around a time-series database and you already know (or you will invest in learning) how to convert units
72
+ and normalize the metrics in Grafana or other visualization tools, we suggest to use `as-collected`.
73
73
-If, on the other hand, you just need long term archiving of Netdata metrics and you plan to mainly work with Netdata, we suggest to use `average`. It decouples visualization from data collection, so it will generally be a lot simpler. Furthermore, if you use `average`, the charts shown in the back-end will match exactly what you see in Netdata, which is not necessarily true for the other modes of operation.
74
+ If, on the other hand, you just need long term archiving of Netdata metrics and you plan to mainly work with
75
+ Netdata, we suggest to use `average`. It decouples visualization from data collection, so it will generally be a lot
76
+ simpler. Furthermore, if you use `average`, the charts shown in the back-end will match exactly what you see in
77
+ Netdata, which is not necessarily true for the other modes of operation.
78
79
5. This code is smart enough, not to slow down Netdata, independently of the speed of the backend server.
80
81
## configuration
82
79
-In `/etc/netdata/netdata.conf` you should have something like this (if not download the latest version
80
-of `netdata.conf` from your Netdata):
83
+In `/etc/netdata/netdata.conf` you should have something like this (if not download the latest version of `netdata.conf`
84
+from your Netdata):
85
82
-```
86
+```conf
87
[backend]
88
enabled = yes | no
89
type = graphite | opentsdb:telnet | opentsdb:http | opentsdb:https | prometheus_remote_write | json | kinesis | mongodb
@@ -98,92 +102,87 @@ of `netdata.conf` from your Netdata):
102
103
- `enabled = yes | no`, enables or disables sending data to a backend
104
101
-- `type = graphite | opentsdb:telnet | opentsdb:http | opentsdb:https | json | kinesis | mongodb`, selects the backend type
105
+- `type = graphite | opentsdb:telnet | opentsdb:http | opentsdb:https | json | kinesis | mongodb`, selects the backend
106
+ type
107
103
-- `destination = host1 host2 host3 ...`, accepts **a space separated list** of hostnames,
104
- IPs (IPv4 and IPv6) and ports to connect to.
105
- Netdata will use the **first available** to send the metrics.
108
+- `destination = host1 host2 host3 ...`, accepts **a space separated list** of hostnames, IPs (IPv4 and IPv6) and
109
+ ports to connect to. Netdata will use the **first available** to send the metrics.
110
111
The format of each item in this list, is: `[PROTOCOL:]IP[:PORT]`.
112
113
`PROTOCOL` can be `udp` or `tcp`. `tcp` is the default and only supported by the current backends.
114
111
- `IP` can be `XX.XX.XX.XX` (IPv4), or `[XX:XX...XX:XX]` (IPv6).
112
- For IPv6 you can to enclose the IP in `[]` to separate it from the port.
115
+ `IP` can be `XX.XX.XX.XX` (IPv4), or `[XX:XX...XX:XX]` (IPv6). For IPv6 you can to enclose the IP in `[]` to
116
+ separate it from the port.
117
118
`PORT` can be a number of a service name. If omitted, the default port for the backend will be used
119
(graphite = 2003, opentsdb = 4242).
120
121
Example IPv4:
122
119
-```
123
+```conf
124
destination = 10.11.14.2:4242 10.11.14.3:4242 10.11.14.4:4242
125
```
126
127
Example IPv6 and IPv4 together:
128
125
-```
129
+```conf
130
destination = [ffff:...:0001]:2003 10.11.12.1:2003
131
```
132
129
- When multiple servers are defined, Netdata will try the next one when the first one fails. This allows
130
- you to load-balance different servers: give your backend servers in different order on each Netdata.
133
+ When multiple servers are defined, Netdata will try the next one when the first one fails. This allows you to
134
+ load-balance different servers: give your backend servers in different order on each Netdata.
135
132
- Netdata also ships [`nc-backend.sh`](nc-backend.sh),
133
- a script that can be used as a fallback backend to save the metrics to disk and push them to the
134
- time-series database when it becomes available again. It can also be used to monitor / trace / debug
135
- the metrics Netdata generates.
136
+ Netdata also ships [`nc-backend.sh`](nc-backend.sh), a script that can be used as a fallback backend to save the
137
+ metrics to disk and push them to the time-series database when it becomes available again. It can also be used to
138
+ monitor / trace / debug the metrics Netdata generates.
139
140
For kinesis backend `destination` should be set to an AWS region (for example, `us-east-1`).
141
142
The MongoDB backend doesn't use the `destination` option for its configuration. It uses the `mongodb.conf`
143
[configuration file](../backends/mongodb/) instead.
144
142
-- `data source = as collected`, or `data source = average`, or `data source = sum`, selects the kind of
143
- data that will be sent to the backend.
145
+- `data source = as collected`, or `data source = average`, or `data source = sum`, selects the kind of data that will
146
+ be sent to the backend.
147
145
-- `hostname = my-name`, is the hostname to be used for sending data to the backend server. By default
146
- this is `[global].hostname`.
148
+- `hostname = my-name`, is the hostname to be used for sending data to the backend server. By default this is
149
+ `[global].hostname`.
150
151
- `prefix = Netdata`, is the prefix to add to all metrics.
152
150
-- `update every = 10`, is the number of seconds between sending data to the backend. Netdata will add
151
- some randomness to this number, to prevent stressing the backend server when many Netdata servers send
152
- data to the same backend. This randomness does not affect the quality of the data, only the time they
153
- are sent.
154
-
155
-- `buffer on failures = 10`, is the number of iterations (each iteration is `[backend].update every` seconds)
156
- to buffer data, when the backend is not available. If the backend fails to receive the data after that
157
- many failures, data loss on the backend is expected (Netdata will also log it).
158
-
159
-- `timeout ms = 20000`, is the timeout in milliseconds to wait for the backend server to process the data.
160
- By default this is `2 * update_every * 1000`.
161
-
162
-- `send hosts matching = localhost *` includes one or more space separated patterns, using `*` as wildcard
163
- (any number of times within each pattern). The patterns are checked against the hostname (the localhost
164
- is always checked as `localhost`), allowing us to filter which hosts will be sent to the backend when
165
- this Netdata is a central Netdata aggregating multiple hosts. A pattern starting with `!` gives a
166
- negative match. So to match all hosts named `*db*` except hosts containing `*slave*`, use
167
- `!*slave* *db*` (so, the order is important: the first pattern matching the hostname will be used - positive
168
- or negative).
169
-
170
-- `send charts matching = *` includes one or more space separated patterns, using `*` as wildcard (any
171
- number of times within each pattern). The patterns are checked against both chart id and chart name.
172
- A pattern starting with `!` gives a negative match. So to match all charts named `apps.*`
173
- except charts ending in `*reads`, use `!*reads apps.*` (so, the order is important: the first pattern
174
- matching the chart id or the chart name will be used - positive or negative).
175
-
176
-- `send names instead of ids = yes | no` controls the metric names Netdata should send to backend.
177
- Netdata supports names and IDs for charts and dimensions. Usually IDs are unique identifiers as read
178
- by the system and names are human friendly labels (also unique). Most charts and metrics have the same
179
- ID and name, but in several cases they are different: disks with device-mapper, interrupts, QoS classes,
180
- statsd synthetic charts, etc.
181
-
182
-- `host tags = list of TAG=VALUE` defines tags that should be appended on all metrics for the given host.
183
- These are currently only sent to opentsdb and prometheus. Please use the appropriate format for each
184
- time-series db. For example opentsdb likes them like `TAG1=VALUE1 TAG2=VALUE2`, but prometheus like
185
- `tag1="value1",tag2="value2"`. Host tags are mirrored with database replication (streaming of metrics
186
- between Netdata servers).
153
+- `update every = 10`, is the number of seconds between sending data to the backend. Netdata will add some randomness
154
+ to this number, to prevent stressing the backend server when many Netdata servers send data to the same backend.
155
+ This randomness does not affect the quality of the data, only the time they are sent.
156
+
157
+- `buffer on failures = 10`, is the number of iterations (each iteration is `[backend].update every` seconds) to
158
+ buffer data, when the backend is not available. If the backend fails to receive the data after that many failures,
159
+ data loss on the backend is expected (Netdata will also log it).
160
+
161
+- `timeout ms = 20000`, is the timeout in milliseconds to wait for the backend server to process the data. By default
162
+ this is `2 * update_every * 1000`.
163
+
164
+- `send hosts matching = localhost *` includes one or more space separated patterns, using `*` as wildcard (any number
165
+ of times within each pattern). The patterns are checked against the hostname (the localhost is always checked as
166
+ `localhost`), allowing us to filter which hosts will be sent to the backend when this Netdata is a central Netdata
167
+ aggregating multiple hosts. A pattern starting with `!` gives a negative match. So to match all hosts named `*db*`
168
+ except hosts containing `*slave*`, use `!*slave* *db*` (so, the order is important: the first pattern matching the
169
+ hostname will be used - positive or negative).
170
+
171
+- `send charts matching = *` includes one or more space separated patterns, using `*` as wildcard (any number of times
172
+ within each pattern). The patterns are checked against both chart id and chart name. A pattern starting with `!`
173
+ gives a negative match. So to match all charts named `apps.*` except charts ending in `*reads`, use `!*reads
174
+ apps.*` (so, the order is important: the first pattern matching the chart id or the chart name will be used -
175
+ positive or negative).
176
+
177
+- `send names instead of ids = yes | no` controls the metric names Netdata should send to backend. Netdata supports
178
+ names and IDs for charts and dimensions. Usually IDs are unique identifiers as read by the system and names are
179
+ human friendly labels (also unique). Most charts and metrics have the same ID and name, but in several cases they
180
+ are different: disks with device-mapper, interrupts, QoS classes, statsd synthetic charts, etc.
181
+
182
+- `host tags = list of TAG=VALUE` defines tags that should be appended on all metrics for the given host. These are
183
+ currently only sent to opentsdb and prometheus. Please use the appropriate format for each time-series db. For
184
+ example opentsdb likes them like `TAG1=VALUE1 TAG2=VALUE2`, but prometheus like `tag1="value1",tag2="value2"`. Host
185
+ tags are mirrored with database replication (streaming of metrics between Netdata servers).
186
187
## monitoring operation
188
@@ -194,16 +193,15 @@ Netdata provides 5 charts:
193
194
2. **Buffered data size**, the amount of data (in KB) Netdata added the buffer.
195
197
-3. ~~**Backend latency**, the time the backend server needed to process the data Netdata sent.
198
- If there was a re-connection involved, this includes the connection time.~~
199
- (this chart has been removed, because it only measures the time Netdata needs to give the data
200
- to the O/S - since the backend servers do not ack the reception, Netdata does not have any means
201
- to measure this properly).
196
+3. ~~**Backend latency**, the time the backend server needed to process the data Netdata sent. If there was a
197
+ re-connection involved, this includes the connection time.~~ (this chart has been removed, because it only measures
198
+ the time Netdata needs to give the data to the O/S - since the backend servers do not ack the reception, Netdata
199
+ does not have any means to measure this properly).
200
201
4. **Backend operations**, the number of operations performed by Netdata.
202
205
-5. **Backend thread CPU usage**, the CPU resources consumed by the Netdata thread, that is responsible
206
- for sending the metrics to the backend server.
203
+5. **Backend thread CPU usage**, the CPU resources consumed by the Netdata thread, that is responsible for sending the
204
+ metrics to the backend server.
205
206

207
@@ -216,7 +214,8 @@ Netdata adds 4 alarms:
214
1. `backend_last_buffering`, number of seconds since the last successful buffering of backend data
215
2. `backend_metrics_sent`, percentage of metrics sent to the backend server
216
3. `backend_metrics_lost`, number of metrics lost due to repeating failures to contact the backend server
219
-4. ~~`backend_slow`, the percentage of time between iterations needed by the backend time to process the data sent by Netdata~~ (this was misleading and has been removed).
217
+4. ~~`backend_slow`, the percentage of time between iterations needed by the backend time to process the data sent by
218
+ Netdata~~ (this was misleading and has been removed).
219
220

221
backends/WALKTHROUGH.md
+122
-171
@@ -2,123 +2,102 @@
2
3
## Intro
4
5
-In this article I will walk you through the basics of getting Netdata,
6
-Prometheus and Grafana all working together and monitoring your application
7
-servers. This article will be using docker on your local workstation. We will be
8
-working with docker in an ad-hoc way, launching containers that run ‘/bin/bash’
9
-and attaching a TTY to them. I use docker here in a purely academic fashion and
10
-do not condone running Netdata in a container. I pick this method so individuals
11
-without cloud accounts or access to VMs can try this out and for it’s speed of
12
-deployment.
5
+In this article I will walk you through the basics of getting Netdata, Prometheus and Grafana all working together and
6
+monitoring your application servers. This article will be using docker on your local workstation. We will be working
7
+with docker in an ad-hoc way, launching containers that run ‘/bin/bash’ and attaching a TTY to them. I use docker here
8
+in a purely academic fashion and do not condone running Netdata in a container. I pick this method so individuals
9
+without cloud accounts or access to VMs can try this out and for it’s speed of deployment.
10
11
## Why Netdata, Prometheus, and Grafana
12
16
-Some time ago I was introduced to Netdata by a coworker. We were attempting to
17
-troubleshoot python code which seemed to be bottlenecked. I was instantly
18
-impressed by the amount of metrics Netdata exposes to you. I quickly added
19
-Netdata to my set of go-to tools when troubleshooting systems performance.
20
-
21
-Some time ago, even later, I was introduced to Prometheus. Prometheus is a
22
-monitoring application which flips the normal architecture around and polls
23
-rest endpoints for its metrics. This architectural change greatly simplifies
24
-and decreases the time necessary to begin monitoring your applications.
25
-Compared to current monitoring solutions the time spent on designing the
26
-infrastructure is greatly reduced. Running a single Prometheus server per
27
-application becomes feasible with the help of Grafana.
28
-
29
-Grafana has been the go to graphing tool for… some time now. It’s awesome,
30
-anyone that has used it knows it’s awesome. We can point Grafana at Prometheus
31
-and use Prometheus as a data source. This allows a pretty simple overall
32
-monitoring architecture: Install Netdata on your application servers, point
33
-Prometheus at Netdata, and then point Grafana at Prometheus.
34
-
35
-I’m omitting an important ingredient in this stack in order to keep this tutorial
36
-simple and that is service discovery. My personal preference is to use Consul.
37
-Prometheus can plug into consul and automatically begin to scrape new hosts that
38
-register a Netdata client with Consul.
39
-
40
-At the end of this tutorial you will understand how each technology fits
41
-together to create a modern monitoring stack. This stack will offer you
42
-visibility into your application and systems performance.
13
+Some time ago I was introduced to Netdata by a coworker. We were attempting to troubleshoot python code which seemed to
14
+be bottlenecked. I was instantly impressed by the amount of metrics Netdata exposes to you. I quickly added Netdata to
15
+my set of go-to tools when troubleshooting systems performance.
16
+
17
+Some time ago, even later, I was introduced to Prometheus. Prometheus is a monitoring application which flips the normal
18
+architecture around and polls rest endpoints for its metrics. This architectural change greatly simplifies and decreases
19
+the time necessary to begin monitoring your applications. Compared to current monitoring solutions the time spent on
20
+designing the infrastructure is greatly reduced. Running a single Prometheus server per application becomes feasible
21
+with the help of Grafana.
22
+
23
+Grafana has been the go to graphing tool for… some time now. It’s awesome, anyone that has used it knows it’s awesome.
24
+We can point Grafana at Prometheus and use Prometheus as a data source. This allows a pretty simple overall monitoring
25
+architecture: Install Netdata on your application servers, point Prometheus at Netdata, and then point Grafana at
26
+Prometheus.
27
+
28
+I’m omitting an important ingredient in this stack in order to keep this tutorial simple and that is service discovery.
29
+My personal preference is to use Consul. Prometheus can plug into consul and automatically begin to scrape new hosts
30
+that register a Netdata client with Consul.
31
+
32
+At the end of this tutorial you will understand how each technology fits together to create a modern monitoring stack.
33
+This stack will offer you visibility into your application and systems performance.
34
35
## Getting Started - Netdata
36
46
-To begin let’s create our container which we will install Netdata on. We need
47
-to run a container, forward the necessary port that Netdata listens on, and
48
-attach a tty so we can interact with the bash shell on the container. But
49
-before we do this we want name resolution between the two containers to work.
50
-In order to accomplish this we will create a user-defined network and attach
51
-both containers to this network. The first command we should run is:
37
+To begin let’s create our container which we will install Netdata on. We need to run a container, forward the necessary
38
+port that Netdata listens on, and attach a tty so we can interact with the bash shell on the container. But before we do
39
+this we want name resolution between the two containers to work. In order to accomplish this we will create a
40
+user-defined network and attach both containers to this network. The first command we should run is:
41
42
```sh
43
docker network create --driver bridge netdata-tutorial
44
```
45
57
-With this user-defined network created we can now launch our container we will
58
-install Netdata on and point it to this network.
46
+With this user-defined network created we can now launch our container we will install Netdata on and point it to this
47
+network.
48
49
```sh
50
docker run -it --name netdata --hostname netdata --network=netdata-tutorial -p 19999:19999 centos:latest '/bin/bash'
51
```
52
64
-This command creates an interactive tty session (-it), gives the container both
65
-a name in relation to the docker daemon and a hostname (this is so you know what
66
-container is which when working in the shells and docker maps hostname
67
-resolution to this container), forwards the local port 19999 to the container’s
68
-port 19999 (-p 19999:19999), sets the command to run (/bin/bash) and then
69
-chooses the base container images (centos:latest). After running this you should
70
-be sitting inside the shell of the container.
53
+This command creates an interactive tty session (-it), gives the container both a name in relation to the docker daemon
54
+and a hostname (this is so you know what container is which when working in the shells and docker maps hostname
55
+resolution to this container), forwards the local port 19999 to the container’s port 19999 (-p 19999:19999), sets the
56
+command to run (/bin/bash) and then chooses the base container images (centos:latest). After running this you should be
57
+sitting inside the shell of the container.
58
72
-After we have entered the shell we can install Netdata. This process could not
73
-be easier. If you take a look at [this link](../packaging/installer/#installation), the Netdata devs give us
74
-several one-liners to install Netdata. I have not had any issues with these one
75
-liners and their bootstrapping scripts so far (If you guys run into anything do
76
-share). Run the following command in your container.
59
+After we have entered the shell we can install Netdata. This process could not be easier. If you take a look at [this
60
+link](../packaging/installer/#installation), the Netdata devs give us several one-liners to install Netdata. I have not
61
+had any issues with these one liners and their bootstrapping scripts so far (If you guys run into anything do share).
62
+Run the following command in your container.
63
64
```sh
65
bash <(curl -Ss https://my-netdata.io/kickstart.sh) --dont-wait
66
```
67
82
-After the install completes you should be able to hit the Netdata dashboard at
83
-<http://localhost:19999/> (replace localhost if you’re doing this on a VM or have
84
-the docker container hosted on a machine not on your local system). If this is
85
-your first time using Netdata I suggest you take a look around. The amount of
86
-time I’ve spent digging through /proc and calculating my own metrics has been
87
-greatly reduced by this tool. Take it all in.
68
+After the install completes you should be able to hit the Netdata dashboard at <http://localhost:19999/> (replace
69
+localhost if you’re doing this on a VM or have the docker container hosted on a machine not on your local system). If
70
+this is your first time using Netdata I suggest you take a look around. The amount of time I’ve spent digging through
71
+/proc and calculating my own metrics has been greatly reduced by this tool. Take it all in.
72
73
Next I want to draw your attention to a particular endpoint. Navigate to
90
-<http://localhost:19999/api/v1/allmetrics?format=prometheus&help=yes> In your
91
-browser. This is the endpoint which publishes all the metrics in a format which
92
-Prometheus understands. Let’s take a look at one of these metrics.
93
-`netdata_system_cpu_percentage_average{chart="system.cpu",family="cpu",dimension="system"}
94
-0.0831255 1501271696000` This metric is representing several things which I will
95
-go in more details in the section on prometheus. For now understand that this
96
-metric: `netdata_system_cpu_percentage_average` has several labels: [chart,
97
-family, dimension]. This corresponds with the first cpu chart you see on the
98
-Netdata dashboard.
74
+<http://localhost:19999/api/v1/allmetrics?format=prometheus&help=yes> In your browser. This is the endpoint which
75
+publishes all the metrics in a format which Prometheus understands. Let’s take a look at one of these metrics.
76
+`netdata_system_cpu_percentage_average{chart="system.cpu",family="cpu",dimension="system"} 0.0831255 1501271696000` This
77
+metric is representing several things which I will go in more details in the section on prometheus. For now understand
78
+that this metric: `netdata_system_cpu_percentage_average` has several labels: (chart, family, dimension). This
79
+corresponds with the first cpu chart you see on the Netdata dashboard.
80
81

82
102
-This CHART is called ‘system.cpu’, The FAMILY is cpu, and the DIMENSION we are
103
-observing is “system”. You can begin to draw links between the charts in Netdata
104
-to the prometheus metrics format in this manner.
83
+This CHART is called ‘system.cpu’, The FAMILY is cpu, and the DIMENSION we are observing is “system”. You can begin to
84
+draw links between the charts in Netdata to the prometheus metrics format in this manner.
85
86
## Prometheus
87
108
-We will be installing prometheus in a container for purpose of demonstration.
109
-While prometheus does have an official container I would like to walk through
110
-the install process and setup on a fresh container. This will allow anyone
88
+We will be installing prometheus in a container for purpose of demonstration. While prometheus does have an official
89
+container I would like to walk through the install process and setup on a fresh container. This will allow anyone
90
reading to migrate this tutorial to a VM or Server of any sort.
91
113
-Let’s start another container in the same fashion as we did the Netdata
114
-container.
92
+Let’s start another container in the same fashion as we did the Netdata container.
93
94
```sh
95
docker run -it --name prometheus --hostname prometheus
96
--network=netdata-tutorial -p 9090:9090 centos:latest '/bin/bash'
97
```
98
121
-This should drop you into a shell once again. Once there quickly install your favorite editor as we will be editing files later in this tutorial.
99
+This should drop you into a shell once again. Once there quickly install your favorite editor as we will be editing
100
+files later in this tutorial.
101
102
```sh
103
yum install vim -y
@@ -139,39 +118,33 @@ mkdir /opt/prometheus
118
sudo tar -xvf /tmp/prometheus-*linux-amd64.tar.gz -C /opt/prometheus --strip=1
119
```
120
142
-This should get prometheus installed into the container. Let’s test that we can run prometheus and connect to it’s web interface.
121
+This should get prometheus installed into the container. Let’s test that we can run prometheus and connect to it’s web
122
+interface.
123
124
```sh
125
/opt/prometheus/prometheus
126
```
127
148
-Now attempt to go to <http://localhost:9090/>. You should be presented with the
149
-prometheus homepage. This is a good point to talk about Prometheus’s data model
150
-which can be viewed here: <https://prometheus.io/docs/concepts/data_model/> As
151
-explained we have two key elements in Prometheus metrics. We have the ‘metric’
152
-and its ‘labels’. Labels allow for granularity between metrics. Let’s use our
153
-previous example to further explain.
128
+Now attempt to go to <http://localhost:9090/>. You should be presented with the prometheus homepage. This is a good
129
+point to talk about Prometheus’s data model which can be viewed here: <https://prometheus.io/docs/concepts/data_model/>
130
+As explained we have two key elements in Prometheus metrics. We have the ‘metric’ and its ‘labels’. Labels allow for
131
+granularity between metrics. Let’s use our previous example to further explain.
132
155
-```
133
+```conf
134
netdata_system_cpu_percentage_average{chart="system.cpu",family="cpu",dimension="system"} 0.0831255 1501271696000
135
```
136
159
-Here our metric is
160
-‘netdata_system_cpu_percentage_average’ and our labels are ‘chart’, ‘family’,
161
-and ‘dimension. The last two values constitute the actual metric value for the
162
-metric type (gauge, counter, etc…). We can begin graphing system metrics with
163
-this information, but first we need to hook up Prometheus to poll Netdata stats.
164
-
165
-Let’s move our attention to Prometheus’s configuration. Prometheus gets it
166
-config from the file located (in our example) at
167
-`/opt/prometheus/prometheus.yml`. I won’t spend an extensive amount of time
168
-going over the configuration values documented here:
169
-<https://prometheus.io/docs/operating/configuration/>. We will be adding a new
170
-“job” under the “scrape_configs”. Let’s make the “scrape_configs” section look
171
-like this (we can use the dns name Netdata due to the custom user-defined
172
-network we created in docker beforehand).
173
-
174
-```yml
137
+Here our metric is ‘netdata_system_cpu_percentage_average’ and our labels are ‘chart’, ‘family’, and ‘dimension. The
138
+last two values constitute the actual metric value for the metric type (gauge, counter, etc…). We can begin graphing
139
+system metrics with this information, but first we need to hook up Prometheus to poll Netdata stats.
140
+
141
+Let’s move our attention to Prometheus’s configuration. Prometheus gets it config from the file located (in our example)
142
+at `/opt/prometheus/prometheus.yml`. I won’t spend an extensive amount of time going over the configuration values
143
+documented here: <https://prometheus.io/docs/operating/configuration/>. We will be adding a new“job” under the
144
+“scrape_configs”. Let’s make the “scrape_configs” section look like this (we can use the dns name Netdata due to the
145
+custom user-defined network we created in docker beforehand).
146
+
147
+```yaml
148
scrape_configs:
149
# The job name is added as a label `job=<job_name>` to any timeseries scraped from this config.
150
- job_name: 'prometheus'
@@ -192,84 +165,66 @@ scrape_configs:
165
- targets: ['netdata:19999']
166
```
167
195
-Let’s start prometheus once again by running `/opt/prometheus/prometheus`. If we
196
-
197
-now navigate to prometheus at ‘<http://localhost:9090/targets’> we should see our
198
-
199
-target being successfully scraped. If we now go back to the Prometheus’s
200
-homepage and begin to type ‘netdata\_’ Prometheus should auto complete metrics
201
-it is now scraping.
168
+Let’s start prometheus once again by running `/opt/prometheus/prometheus`. If we now navigate to prometheus at
169
+‘<http://localhost:9090/targets’> we should see our target being successfully scraped. If we now go back to the
170
+Prometheus’s homepage and begin to type ‘netdata\_’ Prometheus should auto complete metrics it is now scraping.
171
172

173
205
-Let’s now start exploring how we can graph some metrics. Back in our NetData
206
-container lets get the CPU spinning with a pointless busy loop. On the shell do
207
-the following:
174
+Let’s now start exploring how we can graph some metrics. Back in our NetData container lets get the CPU spinning with a
175
+pointless busy loop. On the shell do the following:
176
209
-```
177
+```sh
178
[root@netdata /]# while true; do echo "HOT HOT HOT CPU"; done
179
```
180
213
-Our NetData cpu graph should be showing some activity. Let’s represent this in
214
-Prometheus. In order to do this let’s keep our metrics page open for reference:
215
-<http://localhost:19999/api/v1/allmetrics?format=prometheus&help=yes> We are
216
-setting out to graph the data in the CPU chart so let’s search for “system.cpu”
217
-in the metrics page above. We come across a section of metrics with the first
218
-comments `# COMMENT homogeneous chart "system.cpu", context "system.cpu", family
219
-"cpu", units "percentage"` Followed by the metrics. This is a good start now let
220
-us drill down to the specific metric we would like to graph.
181
+Our NetData cpu graph should be showing some activity. Let’s represent this in Prometheus. In order to do this let’s
182
+keep our metrics page open for reference: <http://localhost:19999/api/v1/allmetrics?format=prometheus&help=yes> We are
183
+setting out to graph the data in the CPU chart so let’s search for “system.cpu”in the metrics page above. We come across
184
+a section of metrics with the first comments `# COMMENT homogeneous chart "system.cpu", context "system.cpu", family
185
+"cpu", units "percentage"` Followed by the metrics. This is a good start now let us drill down to the specific metric we
186
+would like to graph.
187
222
-```
188
+```conf
189
# COMMENT
190
netdata_system_cpu_percentage_average: dimension "system", value is percentage, gauge, dt 1501275951 to 1501275951 inclusive
191
netdata_system_cpu_percentage_average{chart="system.cpu",family="cpu",dimension="system"} 0.0000000 1501275951000
192
```
193
228
-Here we learn that the metric name we care about is
229
-‘netdata_system_cpu_percentage_average’ so throw this into Prometheus and see
230
-what we get. We should see something similar to this (I shut off my busy loop)
194
+Here we learn that the metric name we care about is‘netdata_system_cpu_percentage_average’ so throw this into Prometheus
195
+and see what we get. We should see something similar to this (I shut off my busy loop)
196
197

198
234
-This is a good step toward what we want. Also make note that Prometheus will tag
235
-on an ‘instance’ label for us which corresponds to our statically defined job in
236
-the configuration file. This allows us to tailor our queries to specific
237
-instances. Now we need to isolate the dimension we want in our query. To do this
238
-let us refine the query slightly. Let’s query the dimension also. Place this
239
-into our query text box.
240
-`netdata_system_cpu_percentage_average{dimension="system"}` We now wind up with
241
-the following graph.
199
+This is a good step toward what we want. Also make note that Prometheus will tag on an ‘instance’ label for us which
200
+corresponds to our statically defined job in the configuration file. This allows us to tailor our queries to specific
201
+instances. Now we need to isolate the dimension we want in our query. To do this let us refine the query slightly. Let’s
202
+query the dimension also. Place this into our query text box.
203
+`netdata_system_cpu_percentage_average{dimension="system"}` We now wind up with the following graph.
204
205

206
245
-Awesome, this is exactly what we wanted. If you haven’t caught on yet we can
246
-emulate entire charts from NetData by using the `chart` dimension. If you’d like
247
-you can combine the ‘chart’ and ‘instance’ dimension to create per-instance
248
-charts. Let’s give this a try:
249
-`netdata_system_cpu_percentage_average{chart="system.cpu", instance="netdata:19999"}`
250
-
251
-This is the basics of using Prometheus to query NetData. I’d advise everyone at
252
-this point to read [this page](../backends/prometheus/#using-netdata-with-prometheus).
253
-The key point here is that NetData can export metrics from its internal DB or
254
-can send metrics “as-collected” by specifying the ‘source=as-collected’ url
255
-parameter like so.
256
-<http://localhost:19999/api/v1/allmetrics?format=prometheus&help=yes&types=yes&source=as-collected>
257
-If you choose to use this method you will need to use Prometheus's set of
258
-functions here: <https://prometheus.io/docs/querying/functions/> to obtain useful
259
-metrics as you are now dealing with raw counters from the system. For example
260
-you will have to use the `irate()` function over a counter to get that metric's
261
-rate per second. If your graphing needs are met by using the metrics returned by
262
-NetData's internal database (not specifying any source= url parameter) then use
263
-that. If you find limitations then consider re-writing your queries using the
264
-raw data and using Prometheus functions to get the desired chart.
207
+Awesome, this is exactly what we wanted. If you haven’t caught on yet we can emulate entire charts from NetData by using
208
+the `chart` dimension. If you’d like you can combine the ‘chart’ and ‘instance’ dimension to create per-instance charts.
209
+Let’s give this a try: `netdata_system_cpu_percentage_average{chart="system.cpu", instance="netdata:19999"}`
210
+
211
+This is the basics of using Prometheus to query NetData. I’d advise everyone at this point to read [this
212
+page](../backends/prometheus/#using-netdata-with-prometheus). The key point here is that NetData can export metrics from
213
+its internal DB or can send metrics “as-collected” by specifying the ‘source=as-collected’ url parameter like so.
214
+<http://localhost:19999/api/v1/allmetrics?format=prometheus&help=yes&types=yes&source=as-collected> If you choose to use
215
+this method you will need to use Prometheus's set of functions here: <https://prometheus.io/docs/querying/functions/> to
216
+obtain useful metrics as you are now dealing with raw counters from the system. For example you will have to use the
217
+`irate()` function over a counter to get that metric's rate per second. If your graphing needs are met by using the
218
+metrics returned by NetData's internal database (not specifying any source= url parameter) then use that. If you find
219
+limitations then consider re-writing your queries using the raw data and using Prometheus functions to get the desired
220
+chart.
221
222
## Grafana
223
268
-Finally we make it to grafana. This is the easiest part in my opinion. This time
269
-we will actually run the official grafana docker container as all configuration
270
-we need to do is done via the GUI. Let’s run the following command:
224
+Finally we make it to grafana. This is the easiest part in my opinion. This time we will actually run the official
225
+grafana docker container as all configuration we need to do is done via the GUI. Let’s run the following command:
226
272
-```
227
+```sh
228
docker run -i -p 3000:3000 --network=netdata-tutorial grafana/grafana
229
```
230
@@ -277,26 +232,22 @@ This will get grafana running at ‘<http://localhost:3000/’> Let’s go there
232
233
login using the credentials Admin:Admin.
234
280
-The first thing we want to do is click ‘Add data source’. Let’s make it look
281
-like the following screenshot
235
+The first thing we want to do is click ‘Add data source’. Let’s make it look like the following screenshot
236
237

238
285
-With this completed let’s graph! Create a new Dashboard by clicking on the top
286
-left Grafana Icon and create a new graph in that dashboard. Fill in the query
287
-like we did above and save.
239
+With this completed let’s graph! Create a new Dashboard by clicking on the top left Grafana Icon and create a new graph
240
+in that dashboard. Fill in the query like we did above and save.
241
242

243
244
## Conclusion
245
293
-There you have it, a complete systems monitoring stack which is very easy to
294
-deploy. From here I would begin to understand how Prometheus and a service
295
-discovery mechanism such as Consul can play together nicely. My current prod
296
-deployments automatically register Netdata services into Consul and Prometheus
297
-automatically begins to scrape them. Once achieved you do not have to think
298
-about the monitoring system until Prometheus cannot keep up with your scale.
299
-Once this happens there are options presented in the Prometheus documentation
300
-for solving this. Hope this was helpful, happy monitoring.
246
+There you have it, a complete systems monitoring stack which is very easy to deploy. From here I would begin to
247
+understand how Prometheus and a service discovery mechanism such as Consul can play together nicely. My current prod
248
+deployments automatically register Netdata services into Consul and Prometheus automatically begins to scrape them. Once
249
+achieved you do not have to think about the monitoring system until Prometheus cannot keep up with your scale. Once this
250
+happens there are options presented in the Prometheus documentation for solving this. Hope this was helpful, happy
251
+monitoring.
252
253
[](<>)
backends/aws_kinesis/README.md
+14
-7
@@ -2,11 +2,17 @@
2
3
## Prerequisites
4
5
-To use AWS Kinesis as a backend AWS SDK for C++ should be [installed](https://docs.aws.amazon.com/en_us/sdk-for-cpp/v1/developer-guide/setup.html) first. `libcrypto`, `libssl`, and `libcurl` are also required to compile Netdata with Kinesis support enabled. Next, Netdata should be re-installed from the source. The installer will detect that the required libraries are now available.
5
+To use AWS Kinesis as a backend AWS SDK for C++ should be
6
+[installed](https://docs.aws.amazon.com/en_us/sdk-for-cpp/v1/developer-guide/setup.html) first. `libcrypto`, `libssl`,
7
+and `libcurl` are also required to compile Netdata with Kinesis support enabled. Next, Netdata should be re-installed
8
+from the source. The installer will detect that the required libraries are now available.
9
7
-If the AWS SDK for C++ is being installed from source, it is useful to set `-DBUILD_ONLY="kinesis"`. Otherwise, the building process could take a very long time. Take a note, that the default installation path for the libraries is `/usr/local/lib64`. Many Linux distributions don't include this path as the default one for a library search, so it is advisable to use the following options to `cmake` while building the AWS SDK:
10
+If the AWS SDK for C++ is being installed from source, it is useful to set `-DBUILD_ONLY="kinesis"`. Otherwise, the
11
+building process could take a very long time. Take a note, that the default installation path for the libraries is
12
+`/usr/local/lib64`. Many Linux distributions don't include this path as the default one for a library search, so it is
13
+advisable to use the following options to `cmake` while building the AWS SDK:
14
9
-```
15
+```sh
16
cmake -DCMAKE_INSTALL_LIBDIR=/usr/lib -DCMAKE_INSTALL_INCLUDEDIR=/usr/include -DBUILD_SHARED_LIBS=OFF -DBUILD_ONLY=kinesis <aws-sdk-cpp sources>
17
```
18
@@ -14,7 +20,7 @@ cmake -DCMAKE_INSTALL_LIBDIR=/usr/lib -DCMAKE_INSTALL_INCLUDEDIR=/usr/include -D
20
21
To enable data sending to the kinesis backend set the following options in `netdata.conf`:
22
17
-```
23
+```conf
24
[backend]
25
enabled = yes
26
type = kinesis
@@ -25,7 +31,7 @@ set the `destination` option to an AWS region.
31
32
In the Netdata configuration directory run `./edit-config aws_kinesis.conf` and set AWS credentials and stream name:
33
28
-```
34
+```yaml
35
# AWS credentials
36
aws_access_key_id = your_access_key_id
37
aws_secret_access_key = your_secret_access_key
@@ -34,8 +40,9 @@ aws_secret_access_key = your_secret_access_key
40
stream name = your_stream_name
41
```
42
37
-Alternatively, AWS credentials can be set for the *netdata* user using AWS SDK for C++ [standard methods](https://docs.aws.amazon.com/sdk-for-cpp/v1/developer-guide/credentials.html).
43
+Alternatively, AWS credentials can be set for the `netdata` user using AWS SDK for C++ [standard methods](https://docs.aws.amazon.com/sdk-for-cpp/v1/developer-guide/credentials.html).
44
39
-A partition key for every record is computed automatically by Netdata with the purpose to distribute records across available shards evenly.
45
+A partition key for every record is computed automatically by Netdata with the purpose to distribute records across
46
+available shards evenly.
47
48
[](<>)
backends/mongodb/README.md
+9
-5
@@ -2,21 +2,24 @@
2
3
## Prerequisites
4
5
-To use MongoDB as a backend, `libmongoc` 1.7.0 or higher should be [installed](http://mongoc.org/libmongoc/current/installing.html) first. Next, Netdata should be re-installed from the source. The installer will detect that the required libraries are now available.
5
+To use MongoDB as a backend, `libmongoc` 1.7.0 or higher should be
6
+[installed](http://mongoc.org/libmongoc/current/installing.html) first. Next, Netdata should be re-installed from the
7
+source. The installer will detect that the required libraries are now available.
8
9
## Configuration
10
11
To enable data sending to the MongoDB backend set the following options in `netdata.conf`:
12
11
-```
13
+```conf
14
[backend]
15
enabled = yes
16
type = mongodb
17
```
18
17
-In the Netdata configuration directory run `./edit-config mongodb.conf` and set [MongoDB URI](https://docs.mongodb.com/manual/reference/connection-string/), database name, and collection name:
19
+In the Netdata configuration directory run `./edit-config mongodb.conf` and set [MongoDB
20
+URI](https://docs.mongodb.com/manual/reference/connection-string/), database name, and collection name:
21
19
-```
22
+```yaml
23
# URI
24
uri = mongodb://<hostname>
25
@@ -27,6 +30,7 @@ database = your_database_name
30
collection = your_collection_name
31
```
32
30
-The default socket timeout depends on the backend update interval. The timeout is 500 ms shorter than the interval (but not less than 1000 ms). You can alter the timeout using the `sockettimeoutms` MongoDB URI option.
33
+The default socket timeout depends on the backend update interval. The timeout is 500 ms shorter than the interval (but
34
+not less than 1000 ms). You can alter the timeout using the `sockettimeoutms` MongoDB URI option.
35
36
[](<>)
backends/opentsdb/README.md
+12
-6
@@ -1,25 +1,31 @@
1
# OpenTSDB with HTTP
2
3
-Netdata can easily communicate with OpenTSDB using HTTP API. To enable this channel, set the following options in your `netdata.conf`:
3
+Netdata can easily communicate with OpenTSDB using HTTP API. To enable this channel, set the following options in your
4
+`netdata.conf`:
5
5
-```
6
+```conf
7
[backend]
8
type = opentsdb:http
9
destination = localhost:4242
10
```
11
11
-In this example, OpenTSDB is running with its default port, which is `4242`. If you run OpenTSDB on a different port, change the `destination = localhost:4242` line accordingly.
12
+In this example, OpenTSDB is running with its default port, which is `4242`. If you run OpenTSDB on a different port,
13
+change the `destination = localhost:4242` line accordingly.
14
15
## HTTPS
16
15
-As of [v1.16.0](https://github.com/netdata/netdata/releases/tag/v1.16.0), Netdata can send metrics to OpenTSDB using TLS/SSL. Unfortunately, OpenTDSB does not support encrypted connections, so you will have to configure a reverse proxy to enable HTTPS communication between Netdata and OpenTSBD. You can set up a reverse proxy with [Nginx](../../docs/Running-behind-nginx.md).
17
+As of [v1.16.0](https://github.com/netdata/netdata/releases/tag/v1.16.0), Netdata can send metrics to OpenTSDB using
18
+TLS/SSL. Unfortunately, OpenTDSB does not support encrypted connections, so you will have to configure a reverse proxy
19
+to enable HTTPS communication between Netdata and OpenTSBD. You can set up a reverse proxy with
20
+[Nginx](../../docs/Running-behind-nginx.md).
21
22
After your proxy is configured, make the following changes to `netdata.conf`:
23
19
-```
24
+```conf
25
[backend]
26
type = opentsdb:https
27
destination = localhost:8082
28
```
29
25
-In this example, we used the port `8082` for our reverse proxy. If your reverse proxy listens on a different port, change the `destination = localhost:8082` line accordingly.
30
+In this example, we used the port `8082` for our reverse proxy. If your reverse proxy listens on a different port,
31
+change the `destination = localhost:8082` line accordingly.
backends/prometheus/README.md
+97
-44
@@ -1,15 +1,19 @@
1
# Using Netdata with Prometheus
2
3
-> IMPORTANT: the format Netdata sends metrics to prometheus has changed since Netdata v1.7. The new prometheus backend for Netdata supports a lot more features and is aligned to the development of the rest of the Netdata backends.
3
+> IMPORTANT: the format Netdata sends metrics to prometheus has changed since Netdata v1.7. The new prometheus backend
4
+> for Netdata supports a lot more features and is aligned to the development of the rest of the Netdata backends.
5
5
-Prometheus is a distributed monitoring system which offers a very simple setup along with a robust data model. Recently Netdata added support for Prometheus. I'm going to quickly show you how to install both Netdata and prometheus on the same server. We can then use grafana pointed at Prometheus to obtain long term metrics Netdata offers. I'm assuming we are starting at a fresh ubuntu shell (whether you'd like to follow along in a VM or a cloud instance is up to you).
6
+Prometheus is a distributed monitoring system which offers a very simple setup along with a robust data model. Recently
7
+Netdata added support for Prometheus. I'm going to quickly show you how to install both Netdata and prometheus on the
8
+same server. We can then use grafana pointed at Prometheus to obtain long term metrics Netdata offers. I'm assuming we
9
+are starting at a fresh ubuntu shell (whether you'd like to follow along in a VM or a cloud instance is up to you).
10
11
## Installing Netdata and prometheus
12
13
### Installing Netdata
14
11
-There are number of ways to install Netdata according to [Installation](../../packaging/installer/#installation)\
12
-The suggested way of installing the latest Netdata and keep it upgrade automatically. Using one line installation:
15
+There are number of ways to install Netdata according to [Installation](../../packaging/installer/). The suggested way
16
+of installing the latest Netdata and keep it upgrade automatically. Using one line installation:
17
18
```sh
19
bash <(curl -Ss https://my-netdata.io/kickstart.sh)
@@ -17,7 +21,7 @@ bash <(curl -Ss https://my-netdata.io/kickstart.sh)
21
22
At this point we should have Netdata listening on port 19999. Attempt to take your browser here:
23
20
-```
24
+```sh
25
http://your.netdata.ip:19999
26
```
27
@@ -25,7 +29,10 @@ _(replace `your.netdata.ip` with the IP or hostname of the server running Netdat
29
30
### Installing Prometheus
31
28
-In order to install prometheus we are going to introduce our own systemd startup script along with an example of prometheus.yaml configuration. Prometheus needs to be pointed to your server at a specific target url for it to scrape Netdata's api. Prometheus is always a pull model meaning Netdata is the passive client within this architecture. Prometheus always initiates the connection with Netdata.
32
+In order to install prometheus we are going to introduce our own systemd startup script along with an example of
33
+prometheus.yaml configuration. Prometheus needs to be pointed to your server at a specific target url for it to scrape
34
+Netdata's api. Prometheus is always a pull model meaning Netdata is the passive client within this architecture.
35
+Prometheus always initiates the connection with Netdata.
36
37
#### Download Prometheus
38
@@ -113,7 +120,10 @@ scrape_configs:
120
121
#### Install nodes.yml
122
116
-The following is completely optional, it will enable Prometheus to generate alerts from some NetData sources. Tweak the values to your own needs. We will use the following `nodes.yml` file below. Save it at `/opt/prometheus/nodes.yml`, and add a _- "nodes.yml"_ entry under the _rule_files:_ section in the example prometheus.yml file above.
123
+The following is completely optional, it will enable Prometheus to generate alerts from some NetData sources. Tweak the
124
+values to your own needs. We will use the following `nodes.yml` file below. Save it at `/opt/prometheus/nodes.yml`, and
125
+add a _- "nodes.yml"_ entry under the _rule_files:_ section in the example prometheus.yml file above.
126
+
127
```yaml
128
groups:
129
- name: nodes
@@ -156,7 +166,7 @@ groups:
166
167
Save this service file as `/etc/systemd/system/prometheus.service`:
168
159
-```
169
+```sh
170
[Unit]
171
Description=Prometheus Server
172
AssertPathExists=/opt/prometheus
@@ -183,13 +193,15 @@ sudo systemctl enable prometheus
193
194
Prometheus should now start and listen on port 9090. Attempt to head there with your browser.
195
186
-If everything is working correctly when you fetch `http://your.prometheus.ip:9090` you will see a 'Status' tab. Click this and click on 'targets' We should see the Netdata host as a scraped target.
196
+If everything is working correctly when you fetch `http://your.prometheus.ip:9090` you will see a 'Status' tab. Click
197
+this and click on 'targets' We should see the Netdata host as a scraped target.
198
188
-- - -
199
+---
200
201
## Netdata support for prometheus
202
192
-> IMPORTANT: the format Netdata sends metrics to prometheus has changed since Netdata v1.6. The new format allows easier queries for metrics and supports both `as collected` and normalized metrics.
203
+> IMPORTANT: the format Netdata sends metrics to prometheus has changed since Netdata v1.6. The new format allows easier
204
+> queries for metrics and supports both `as collected` and normalized metrics.
205
206
Before explaining the changes, we have to understand the key differences between Netdata and prometheus.
207
@@ -203,7 +215,8 @@ Each chart in Netdata has several properties (common to all its metrics):
215
216
- `chart_name` - a more human friendly name for `chart_id`, also unique.
217
206
-- `context` - this is the template of the chart. All disk I/O charts have the same context, all mysql requests charts have the same context, etc. This is used for alarm templates to match all the charts they should be attached to.
218
+- `context` - this is the template of the chart. All disk I/O charts have the same context, all mysql requests charts
219
+ have the same context, etc. This is used for alarm templates to match all the charts they should be attached to.
220
221
- `family` groups a set of charts together. It is used as the submenu of the dashboard.
222
@@ -211,35 +224,52 @@ Each chart in Netdata has several properties (common to all its metrics):
224
225
#### dimensions
226
214
-Then each Netdata chart contains metrics called `dimensions`. All the dimensions of a chart have the same units of measurement, and are contextually in the same category (ie. the metrics for disk bandwidth are `read` and `write` and they are both in the same chart).
227
+Then each Netdata chart contains metrics called `dimensions`. All the dimensions of a chart have the same units of
228
+measurement, and are contextually in the same category (ie. the metrics for disk bandwidth are `read` and `write` and
229
+they are both in the same chart).
230
231
### Netdata data source
232
233
Netdata can send metrics to prometheus from 3 data sources:
234
220
-- `as collected` or `raw` - this data source sends the metrics to prometheus as they are collected. No conversion is done by Netdata. The latest value for each metric is just given to prometheus. This is the most preferred method by prometheus, but it is also the harder to work with. To work with this data source, you will need to understand how to get meaningful values out of them.
235
+- `as collected` or `raw` - this data source sends the metrics to prometheus as they are collected. No conversion is
236
+ done by Netdata. The latest value for each metric is just given to prometheus. This is the most preferred method by
237
+ prometheus, but it is also the harder to work with. To work with this data source, you will need to understand how
238
+ to get meaningful values out of them.
239
222
- The format of the metrics is: `CONTEXT{chart="CHART",family="FAMILY",dimension="DIMENSION"}`.
240
+ The format of the metrics is: `CONTEXT{chart="CHART",family="FAMILY",dimension="DIMENSION"}`.
241
224
- If the metric is a counter (`incremental` in Netdata lingo), `_total` is appended the context.
242
+ If the metric is a counter (`incremental` in Netdata lingo), `_total` is appended the context.
243
226
- Unlike prometheus, Netdata allows each dimension of a chart to have a different algorithm and conversion constants (`multiplier` and `divisor`). In this case, that the dimensions of a charts are heterogeneous, Netdata will use this format: `CONTEXT_DIMENSION{chart="CHART",family="FAMILY"}`
244
+ Unlike prometheus, Netdata allows each dimension of a chart to have a different algorithm and conversion constants
245
+ (`multiplier` and `divisor`). In this case, that the dimensions of a charts are heterogeneous, Netdata will use this
246
+ format: `CONTEXT_DIMENSION{chart="CHART",family="FAMILY"}`
247
228
-- `average` - this data source uses the Netdata database to send the metrics to prometheus as they are presented on the Netdata dashboard. So, all the metrics are sent as gauges, at the units they are presented in the Netdata dashboard charts. This is the easiest to work with.
248
+- `average` - this data source uses the Netdata database to send the metrics to prometheus as they are presented on
249
+ the Netdata dashboard. So, all the metrics are sent as gauges, at the units they are presented in the Netdata
250
+ dashboard charts. This is the easiest to work with.
251
230
- The format of the metrics is: `CONTEXT_UNITS_average{chart="CHART",family="FAMILY",dimension="DIMENSION"}`.
252
+ The format of the metrics is: `CONTEXT_UNITS_average{chart="CHART",family="FAMILY",dimension="DIMENSION"}`.
253
232
- When this source is used, Netdata keeps track of the last access time for each prometheus server fetching the metrics. This last access time is used at the subsequent queries of the same prometheus server to identify the time-frame the `average` will be calculated. So, no matter how frequently prometheus scrapes Netdata, it will get all the database data. To identify each prometheus server, Netdata uses by default the IP of the client fetching the metrics. If there are multiple prometheus servers fetching data from the same Netdata, using the same IP, each prometheus server can append `server=NAME` to the URL. Netdata will use this `NAME` to uniquely identify the prometheus server.
254
+ When this source is used, Netdata keeps track of the last access time for each prometheus server fetching the
255
+ metrics. This last access time is used at the subsequent queries of the same prometheus server to identify the
256
+ time-frame the `average` will be calculated.
257
+
258
+ So, no matter how frequently prometheus scrapes Netdata, it will get all the database data.
259
+ To identify each prometheus server, Netdata uses by default the IP of the client fetching the metrics.
260
+
261
+ If there are multiple prometheus servers fetching data from the same Netdata, using the same IP, each prometheus
262
+ server can append `server=NAME` to the URL. Netdata will use this `NAME` to uniquely identify the prometheus server.
263
264
- `sum` or `volume`, is like `average` but instead of averaging the values, it sums them.
265
236
- The format of the metrics is: `CONTEXT_UNITS_sum{chart="CHART",family="FAMILY",dimension="DIMENSION"}`.
237
- All the other operations are the same with `average`.
266
+ The format of the metrics is: `CONTEXT_UNITS_sum{chart="CHART",family="FAMILY",dimension="DIMENSION"}`. All the
267
+ other operations are the same with `average`.
268
239
-To change the data source to `sum` or `as-collected` you need to provide the `source` parameter in the request URL.
240
-e.g.: `http://your.netdata.ip:19999/api/v1/allmetrics?format=prometheus&help=yes&source=as-collected`
269
+ To change the data source to `sum` or `as-collected` you need to provide the `source` parameter in the request URL.
270
+ e.g.: `http://your.netdata.ip:19999/api/v1/allmetrics?format=prometheus&help=yes&source=as-collected`
271
242
-Keep in mind that early versions of Netdata were sending the metrics as: `CHART_DIMENSION{}`.
272
+ Keep in mind that early versions of Netdata were sending the metrics as: `CHART_DIMENSION{}`.
273
274
### Querying Metrics
275
@@ -251,7 +281,9 @@ _(replace `your.netdata.ip` with the ip or hostname of your Netdata server)_
281
282
Netdata will respond with all the metrics it sends to prometheus.
283
254
-If you search that page for `"system.cpu"` you will find all the metrics Netdata is exporting to prometheus for this chart. `system.cpu` is the chart name on the Netdata dashboard (on the Netdata dashboard all charts have a text heading such as : `Total CPU utilization (system.cpu)`. What we are interested here in the chart name: `system.cpu`).
284
+If you search that page for `"system.cpu"` you will find all the metrics Netdata is exporting to prometheus for this
285
+chart. `system.cpu` is the chart name on the Netdata dashboard (on the Netdata dashboard all charts have a text heading
286
+such as : `Total CPU utilization (system.cpu)`. What we are interested here in the chart name: `system.cpu`).
287
288
Searching for `"system.cpu"` reveals:
289
@@ -281,7 +313,9 @@ netdata_system_cpu_percentage_average{chart="system.cpu",family="cpu",dimension=
313
314
_(Netdata response for `system.cpu` with source=`average`)_
315
284
-In `average` or `sum` data sources, all values are normalized and are reported to prometheus as gauges. Now, use the 'expression' text form in prometheus. Begin to type the metrics we are looking for: `netdata_system_cpu`. You should see that the text form begins to auto-fill as prometheus knows about this metric.
316
+In `average` or `sum` data sources, all values are normalized and are reported to prometheus as gauges. Now, use the
317
+'expression' text form in prometheus. Begin to type the metrics we are looking for: `netdata_system_cpu`. You should see
318
+that the text form begins to auto-fill as prometheus knows about this metric.
319
320
If the data source was `as collected`, the response would be:
321
@@ -315,7 +349,9 @@ For more information check prometheus documentation.
349
350
### Streaming data from upstream hosts
351
318
-The `format=prometheus` parameter only exports the host's Netdata metrics. If you are using the master/slave functionality of Netdata this ignores any upstream hosts - so you should consider using the below in your **prometheus.yml**:
352
+The `format=prometheus` parameter only exports the host's Netdata metrics. If you are using the master/slave
353
+functionality of Netdata this ignores any upstream hosts - so you should consider using the below in your
354
+**prometheus.yml**:
355
356
```yaml
357
metrics_path: '/api/v1/allmetrics'
@@ -324,31 +360,38 @@ The `format=prometheus` parameter only exports the host's Netdata metrics. If y
360
honor_labels: true
361
```
362
327
-This will report all upstream host data, and `honor_labels` will make Prometheus take note of the instance names provided.
363
+This will report all upstream host data, and `honor_labels` will make Prometheus take note of the instance names
364
+provided.
365
366
### Timestamps
367
331
-To pass the metrics through prometheus pushgateway, Netdata supports the option `×tamps=no` to send the metrics without timestamps.
368
+To pass the metrics through prometheus pushgateway, Netdata supports the option `×tamps=no` to send the metrics
369
+without timestamps.
370
371
## Netdata host variables
372
335
-Netdata collects various system configuration metrics, like the max number of TCP sockets supported, the max number of files allowed system-wide, various IPC sizes, etc. These metrics are not exposed to prometheus by default.
373
+Netdata collects various system configuration metrics, like the max number of TCP sockets supported, the max number of
374
+files allowed system-wide, various IPC sizes, etc. These metrics are not exposed to prometheus by default.
375
376
To expose them, append `variables=yes` to the Netdata URL.
377
378
### TYPE and HELP
379
341
-To save bandwidth, and because prometheus does not use them anyway, `# TYPE` and `# HELP` lines are suppressed. If wanted they can be re-enabled via `types=yes` and `help=yes`, e.g. `/api/v1/allmetrics?format=prometheus&types=yes&help=yes`
380
+To save bandwidth, and because prometheus does not use them anyway, `# TYPE` and `# HELP` lines are suppressed. If
381
+wanted they can be re-enabled via `types=yes` and `help=yes`, e.g.
382
+`/api/v1/allmetrics?format=prometheus&types=yes&help=yes`
383
384
### Names and IDs
385
345
-Netdata supports names and IDs for charts and dimensions. Usually IDs are unique identifiers as read by the system and names are human friendly labels (also unique).
386
+Netdata supports names and IDs for charts and dimensions. Usually IDs are unique identifiers as read by the system and
387
+names are human friendly labels (also unique).
388
347
-Most charts and metrics have the same ID and name, but in several cases they are different: disks with device-mapper, interrupts, QoS classes, statsd synthetic charts, etc.
389
+Most charts and metrics have the same ID and name, but in several cases they are different: disks with device-mapper,
390
+interrupts, QoS classes, statsd synthetic charts, etc.
391
392
The default is controlled in `netdata.conf`:
393
351
-```
394
+```conf
395
[backend]
396
send names instead of ids = yes | no
397
```
@@ -362,18 +405,21 @@ You can overwrite it from prometheus, by appending to the URL:
405
406
Netdata can filter the metrics it sends to prometheus with this setting:
407
365
-```
408
+```conf
409
[backend]
410
send charts matching = *
411
```
412
370
-This settings accepts a space separated list of patterns to match the **charts** to be sent to prometheus. Each pattern can use `*` as wildcard, any number of times (e.g `*a*b*c*` is valid). Patterns starting with `!` give a negative match (e.g `!*.bad users.* groups.*` will send all the users and groups except `bad` user and `bad` group). The order is important: the first match (positive or negative) left to right, is used.
413
+This settings accepts a space separated list of patterns to match the **charts** to be sent to prometheus. Each pattern
414
+can use `*` as wildcard, any number of times (e.g `*a*b*c*` is valid). Patterns starting with `!` give a negative match
415
+(e.g `!*.bad users.* groups.*` will send all the users and groups except `bad` user and `bad` group). The order is
416
+important: the first match (positive or negative) left to right, is used.
417
418
### Changing the prefix of Netdata metrics
419
420
Netdata sends all metrics prefixed with `netdata_`. You can change this in `netdata.conf`, like this:
421
376
-```
422
+```conf
423
[backend]
424
prefix = netdata
425
```
@@ -382,16 +428,23 @@ It can also be changed from the URL, by appending `&prefix=netdata`.
428
429
### Metric Units
430
385
-The default source `average` adds the unit of measurement to the name of each metric (e.g. `_KiB_persec`).
386
-To hide the units and get the same metric names as with the other sources, append to the URL `&hideunits=yes`.
431
+The default source `average` adds the unit of measurement to the name of each metric (e.g. `_KiB_persec`). To hide the
432
+units and get the same metric names as with the other sources, append to the URL `&hideunits=yes`.
433
388
-The units were standardized in v1.12, with the effect of changing the metric names.
389
-To get the metric names as they were before v1.12, append to the URL `&oldunits=yes`
434
+The units were standardized in v1.12, with the effect of changing the metric names. To get the metric names as they were
435
+before v1.12, append to the URL `&oldunits=yes`
436
437
### Accuracy of `average` and `sum` data sources
438
393
-When the data source is set to `average` or `sum`, Netdata remembers the last access of each client accessing prometheus metrics and uses this last access time to respond with the `average` or `sum` of all the entries in the database since that. This means that prometheus servers are not losing data when they access Netdata with data source = `average` or `sum`.
439
+When the data source is set to `average` or `sum`, Netdata remembers the last access of each client accessing prometheus
440
+metrics and uses this last access time to respond with the `average` or `sum` of all the entries in the database since
441
+that. This means that prometheus servers are not losing data when they access Netdata with data source = `average` or
442
+`sum`.
443
395
-To uniquely identify each prometheus server, Netdata uses the IP of the client accessing the metrics. If however the IP is not good enough for identifying a single prometheus server (e.g. when prometheus servers are accessing Netdata through a web proxy, or when multiple prometheus servers are NATed to a single IP), each prometheus may append `&server=NAME` to the URL. This `NAME` is used by Netdata to uniquely identify each prometheus server and keep track of its last access time.
444
+To uniquely identify each prometheus server, Netdata uses the IP of the client accessing the metrics. If however the IP
445
+is not good enough for identifying a single prometheus server (e.g. when prometheus servers are accessing Netdata
446
+through a web proxy, or when multiple prometheus servers are NATed to a single IP), each prometheus may append
447
+`&server=NAME` to the URL. This `NAME` is used by Netdata to uniquely identify each prometheus server and keep track of
448
+its last access time.
449
450
[](<>)
backends/prometheus/remote_write/README.md
+11
-5
@@ -2,26 +2,32 @@
2
3
## Prerequisites
4
5
-To use the prometheus remote write API with [storage providers](https://prometheus.io/docs/operating/integrations/#remote-endpoints-and-storage) [protobuf](https://developers.google.com/protocol-buffers/) and [snappy](https://github.com/google/snappy) libraries should be installed first. Next, Netdata should be re-installed from the source. The installer will detect that the required libraries and utilities are now available.
5
+To use the prometheus remote write API with [storage
6
+providers](https://prometheus.io/docs/operating/integrations/#remote-endpoints-and-storage)
7
+[protobuf](https://developers.google.com/protocol-buffers/) and [snappy](https://github.com/google/snappy) libraries
8
+should be installed first. Next, Netdata should be re-installed from the source. The installer will detect that the
9
+required libraries and utilities are now available.
10
11
## Configuration
12
13
An additional option in the backend configuration section is available for the remote write backend:
14
11
-```
15
+```conf
16
[backend]
17
remote write URL path = /receive
18
```
19
16
-The default value is `/receive`. `remote write URL path` is used to set an endpoint path for the remote write protocol. For example, if your endpoint is `http://example.domain:example_port/storage/read` you should set
20
+The default value is `/receive`. `remote write URL path` is used to set an endpoint path for the remote write protocol.
21
+For example, if your endpoint is `http://example.domain:example_port/storage/read` you should set
22
18
-```
23
+```conf
24
[backend]
25
destination = example.domain:example_port
26
remote write URL path = /storage/read
27
```
28
24
-`buffered` and `lost` dimensions in the Netdata Backend Data Size operation monitoring chart estimate uncompressed buffer size on failures.
29
+`buffered` and `lost` dimensions in the Netdata Backend Data Size operation monitoring chart estimate uncompressed
30
+buffer size on failures.
31
32
## Notes
33