Add documentation for network interfaces (#5381)
* Add documentation for network interfaces * Minor fix * Format chart names * Add an example
Vladimir Kobal committed
Feb 14, 2019 at 11:24 UTC
7db34d77071e86a3760c9d4921661ac72b46ccb6
2 files changed
+82
-12
collectors/proc.plugin/README.md
+79
-12
@@ -37,33 +37,33 @@ Hopefully, the Linux kernel provides many metrics that can provide deep insights
37
38
### Monitored disk metrics
39
40
-- I/O bandwidth/s (kb/s)
40
+- **I/O bandwidth/s (kb/s)**
41
The amount of data transferred from and to the disk.
42
-- I/O operations/s
42
+- **I/O operations/s**
43
The number of I/O operations completed.
44
-- Queued I/O operations
44
+- **Queued I/O operations**
45
The number of currently queued I/O operations. For traditional disks that execute commands one after another, one of them is being run by the disk and the rest are just waiting in a queue.
46
-- Backlog size (time in ms)
46
+- **Backlog size (time in ms)**
47
The expected duration of the currently queued I/O operations.
48
-- Utilization (time percentage)
48
+- **Utilization (time percentage)**
49
The percentage of time the disk was busy with something. This is a very interesting metric, since for most disks, that execute commands sequentially, **this is the key indication of congestion**. A sequential disk that is 100% of the available time busy, has no time to do anything more, so even if the bandwidth or the number of operations executed by the disk is low, its capacity has been reached.
50
Of course, for newer disk technologies (like fusion cards) that are capable to execute multiple commands in parallel, this metric is just meaningless.
51
-- Average I/O operation time (ms)
51
+- **Average I/O operation time (ms)**
52
The average time for I/O requests issued to the device to be served. This includes the time spent by the requests in queue and the time spent servicing them.
53
-- Average I/O operation size (kb)
53
+- **Average I/O operation size (kb)**
54
The average amount of data of the completed I/O operations.
55
-- Average Service Time (ms)
55
+- **Average Service Time (ms)**
56
The average service time for completed I/O operations. This metric is calculated using the total busy time of the disk and the number of completed operations. If the disk is able to execute multiple parallel operations the reporting average service time will be misleading.
57
-- Merged I/O operations/s
57
+- **Merged I/O operations/s**
58
The Linux kernel is capable of merging I/O operations. So, if two requests to read data from the disk are adjacent, the Linux kernel may merge them to one before giving them to disk. This metric measures the number of operations that have been merged by the Linux kernel.
59
-- Total I/O time
59
+- **Total I/O time**
60
The sum of the duration of all completed I/O operations. This number can exceed the interval if the disk is able to execute multiple I/O operations in parallel.
61
-- Space usage
61
+- **Space usage**
62
For mounted disks, netdata will provide a chart for their space, with 3 dimensions:
63
1. free
64
2. used
65
3. reserved for root
66
-- inode usage
66
+- **inode usage**
67
For mounted disks, netdata will provide a chart for their inodes (number of file and directories), with 3 dimensions:
68
1. free
69
2. used
@@ -250,6 +250,73 @@ each state.
250
251
`schedstat filename to monitor`, `cpuidle name filename to monitor`, and `cpuidle time filename to monitor` in the `[plugin:proc:/proc/stat]` configuration section
252
253
+## Monitoring Network Interfaces
254
+
255
+### Monitored network interface metrics
256
+
257
+- **Physical Network Interfaces Aggregated Bandwidth (kilobits/s)**
258
+ The amount of data received and sent through all physical interfaces in the system. This is the source of data for the Net Inbound and Net Outbound dials in the System Overview section.
259
+
260
+- **Bandwidth (kilobits/s)**
261
+ The amount of data received and sent through the interface.
262
+- **Packets (packets/s)**
263
+ The number of packets received, packets sent, and multicast packets transmitted through the interface.
264
+
265
+- **Interface Errors (errors/s)**
266
+ The number of errors for the inbound and outbound traffic on the interface.
267
+- **Interface Drops (drops/s)**
268
+ The number of packets dropped for the inbound and outbound traffic on the interface.
269
+- **Interface FIFO Buffer Errors (errors/s)**
270
+ The number of FIFO buffer errors encountered while receiving and transmitting data through the interface.
271
+- **Compressed Packets (packets/s)**
272
+ The number of compressed packets transmitted or received by the device driver.
273
+- **Network Interface Events (events/s)**
274
+ The number of packet framing errors, collisions detected on the interface, and carrier losses detected by the device driver.
275
+
276
+By default netdata will enable monitoring metrics only when they are not zero. If they are constantly zero they are ignored. Metrics that will start having values, after netdata is started, will be detected and charts will be automatically added to the dashboard (a refresh of the dashboard is needed for them to appear though).
277
+
278
+#### alarms
279
+
280
+There are several alarms defined in `health.d/net.conf`.
281
+
282
+The tricky ones are `inbound packets dropped` and `inbound packets dropped ratio`. They have quite a strict policy so that they warn users about possible issues. These alarms can be annoying for some network configurations. It is especially true for some bonding configurations if an interface is a slave or a bonding interface itself. If it is expected to have a certain number of drops on an interface for a certain network configuration, a separate alarm with different triggering thresholds can be created or the existing one can be disabled for this specific interface. It can be done with the help of the [families](../../health/#alarm-line-families) line in the alarm configuration. For example, if you want to disable the `inbound packets dropped` alarm for `eth0`, set `families: !eth0 *` in the alarm definition for `template: inbound_packets_dropped`.
283
+
284
+#### configuration
285
+
286
+Module configuration:
287
+
288
+```
289
+[plugin:proc:/proc/net/dev]
290
+ # filename to monitor = /proc/net/dev
291
+ # path to get virtual interfaces = /sys/devices/virtual/net/%s
292
+ # path to get net device speed = /sys/class/net/%s/speed
293
+ # enable new interfaces detected at runtime = auto
294
+ # bandwidth for all interfaces = auto
295
+ # packets for all interfaces = auto
296
+ # errors for all interfaces = auto
297
+ # drops for all interfaces = auto
298
+ # fifo for all interfaces = auto
299
+ # compressed packets for all interfaces = auto
300
+ # frames, collisions, carrier counters for all interfaces = auto
301
+ # disable by default interfaces matching = lo fireqos* *-ifb
302
+ # refresh interface speed every seconds = 10
303
+```
304
+
305
+Per interface configuration:
306
+
307
+```
308
+[plugin:proc:/proc/net/dev:enp0s3]
309
+ # enabled = yes
310
+ # virtual = no
311
+ # bandwidth = auto
312
+ # packets = auto
313
+ # errors = auto
314
+ # drops = auto
315
+ # fifo = auto
316
+ # compressed = auto
317
+ # events = auto
318
+```
319
+
320
## Linux Anti-DDoS
321
322

health/health.d/net.conf
+3
@@ -50,6 +50,9 @@
50
# check if an interface is dropping packets
51
# the alarm is checked every 1 minute
52
# and examines the last 10 minutes of data
53
+#
54
+# it is possible to have expected packet drops on an interface for some network configurations
55
+# look at the Monitoring Network Interfaces section in the proc.plugin documentation for more information
56
57
template: inbound_packets_dropped
58
on: net.drops