health: put guides into subdirs (#16358)
Ilya Mashchenko committed
Nov 8, 2023 at 10:59 UTC
fa86868728a1cbb2b1454951fa7f1f66bf62b8f3
307 files changed
+26
-156
health/guides/adaptec_raid/adaptec_raid_ld_status.md
renamed
health/guides/adaptec_raid/adaptec_raid_pd_state.md
renamed
health/guides/anomalies/anomalies_anomaly_flags.md
renamed
health/guides/anomalies/anomalies_anomaly_probabilities.md
renamed
health/guides/apcupsd/apcupsd_10min_ups_load.md
renamed
health/guides/apcupsd/apcupsd_last_collected_secs.md
renamed
health/guides/apcupsd/apcupsd_ups_charge.md
renamed
health/guides/beanstalk/beanstalk_number_of_tubes.md
renamed
health/guides/beanstalk/beanstalk_server_buried_jobs.md
renamed
health/guides/beanstalk/beanstalk_tube_buried_jobs.md
renamed
health/guides/bind_rndc_stats_file_size.md
deleted
-55
@@ -1,55 +0,0 @@
1
-### Understand the alert
2
-
3
-This alert is related to the `BIND` DNS server and its statistics file. If you receive this alert, it means that the file size has crossed a predetermined threshold (warning state at 512 MB and critical state at 1024 MB). This can negatively impact the performance of your DNS server.
4
-
5
-### What is BIND?
6
-
7
-BIND (Berkeley Internet Name Domain) is an open-source DNS server that provides DNS services on Linux servers. It is widely used to implement DNS services on the internet.
8
-
9
-### What is the BIND statistics file?
10
-
11
-BIND keeps track of various metrics and statistics in a `*.stats` file, which is defined by the `statistics-file` option in the BIND configuration. This file can grow in size over time as it accumulates more data.
12
-
13
-### Troubleshoot the alert
14
-
15
-If you receive this alert, it's time to take action to reduce the size of the BIND statistics file. You can do this by:
16
-
17
-1. Review and determine which statistics in the file are necessary:
18
-
19
- - You can check the content of the statistics file to identify the metrics collected.
20
- - Consult the documentation or support forums of the relevant network services to determine which statistics are important for your use case.
21
-
22
-2. Update your BIND configuration to reduce the collection of unnecessary statistics:
23
-
24
- - Edit your BIND configuration file (usually `/etc/bind/named.conf` or `/etc/named.conf`) to remove or comment out the unnecessary statistics options.
25
- - If you're not sure which options to remove, search the BIND documentation or seek assistance from administrators who have experience with BIND configuration.
26
-
27
-3. Restart the BIND service to apply your changes:
28
-
29
- ```
30
- sudo systemctl restart bind9
31
- ```
32
-
33
- (replace `bind9` with the service name of your BIND installation if different)
34
-
35
-4. Manually delete or rotate the large statistics file:
36
-
37
- - To delete the file, use this command:
38
-
39
- ```
40
- sudo rm /path/to/your/stats/file
41
- ```
42
-
43
- - To rotate the file, you can use the `logrotate` utility:
44
-
45
- ```
46
- sudo logrotate --force /etc/logrotate.d/bind
47
- ```
48
-
49
- (update the configuration file path if your set up uses a different location)
50
-
51
-After completing these steps, monitor the size of the BIND statistics file to ensure it doesn't grow beyond your desired threshold.
52
-
53
-### Useful resources
54
-
55
-1. [BIND Documentation](https://bind9.readthedocs.io/en/latest/)
health/guides/boinc/boinc_active_tasks.md
renamed
health/guides/boinc/boinc_compute_errors.md
renamed
health/guides/boinc/boinc_total_tasks.md
renamed
health/guides/boinc/boinc_upload_errors.md
renamed
health/guides/btrfs/btrfs_allocated.md
renamed
health/guides/btrfs/btrfs_data.md
renamed
health/guides/btrfs/btrfs_device_corruption_errors.md
renamed
health/guides/btrfs/btrfs_device_flush_errors.md
renamed
health/guides/btrfs/btrfs_device_generation_errors.md
renamed
health/guides/btrfs/btrfs_device_read_errors.md
renamed
health/guides/btrfs/btrfs_device_write_errors.md
renamed
health/guides/btrfs/btrfs_metadata.md
renamed
health/guides/btrfs/btrfs_system.md
renamed
health/guides/ceph/ceph_cluster_space_usage.md
renamed
health/guides/cgroups/cgroup_10min_cpu_usage.md
renamed
health/guides/cgroups/cgroup_10s_received_packets_storm.md
renamed
health/guides/cgroups/cgroup_1m_received_packets_rate.md
renamed
health/guides/cgroups/cgroup_ram_in_use.md
renamed
health/guides/cgroups/k8s_cgroup_10min_cpu_usage.md
renamed
health/guides/cgroups/k8s_cgroup_10s_received_packets_storm.md
renamed
health/guides/cgroups/k8s_cgroup_1m_received_packets_rate.md
renamed
health/guides/cgroups/k8s_cgroup_ram_in_use.md
renamed
health/guides/cockroachdb/cockroachdb_open_file_descriptors_limit.md
renamed
health/guides/cockroachdb/cockroachdb_unavailable_ranges.md
renamed
health/guides/cockroachdb/cockroachdb_underreplicated_ranges.md
renamed
health/guides/cockroachdb/cockroachdb_used_storage_capacity.md
renamed
health/guides/cockroachdb/cockroachdb_used_usable_storage_capacity.md
renamed
health/guides/consul/consul_autopilot_health_status.md
renamed
health/guides/consul/consul_autopilot_server_health_status.md
renamed
health/guides/consul/consul_client_rpc_requests_exceeded.md
renamed
health/guides/consul/consul_client_rpc_requests_failed.md
renamed
health/guides/consul/consul_gc_pause_time.md
renamed
health/guides/consul/consul_license_expiration_time.md
renamed
health/guides/consul/consul_node_health_check_status.md
renamed
health/guides/consul/consul_raft_leader_last_contact_time.md
renamed
health/guides/consul/consul_raft_leadership_transitions.md
renamed
health/guides/consul/consul_raft_thread_fsm_saturation.md
renamed
health/guides/consul/consul_raft_thread_main_saturation.md
renamed
health/guides/consul/consul_service_health_check_status.md
renamed
health/guides/cpu/10min_cpu_iowait.md
renamed
health/guides/cpu/10min_cpu_usage.md
renamed
health/guides/cpu/20min_steal_cpu.md
renamed
health/guides/dbengine/10min_dbengine_global_flushing_errors.md
renamed
health/guides/dbengine/10min_dbengine_global_flushing_warnings.md
renamed
health/guides/dbengine/10min_dbengine_global_fs_errors.md
renamed
health/guides/dbengine/10min_dbengine_global_io_errors.md
renamed
health/guides/disks/10min_disk_backlog.md
renamed
health/guides/disks/10min_disk_utilization.md
renamed
health/guides/disks/bcache_cache_dirty.md
renamed
health/guides/disks/bcache_cache_errors.md
renamed
health/guides/disks/disk_inode_usage.md
renamed
health/guides/disks/disk_space_usage.md
renamed
health/guides/dns_query/dns_query_query_status.md
renamed
health/guides/dnsmasq/dnsmasq_dhcp_dhcp_range_utilization.md
renamed
health/guides/docker/docker_container_unhealthy.md
renamed
health/guides/elasticsearch/elasticsearch_cluster_health_status_red.md
renamed
health/guides/elasticsearch/elasticsearch_cluster_health_status_yellow.md
renamed
health/guides/elasticsearch/elasticsearch_node_index_health_red.md
renamed
health/guides/elasticsearch/elasticsearch_node_indices_search_time_fetch.md
renamed
health/guides/elasticsearch/elasticsearch_node_indices_search_time_query.md
renamed
health/guides/entropy/lowest_entropy.md
renamed
health/guides/exporting/exporting_last_buffering.md
renamed
health/guides/exporting/exporting_metrics_sent.md
renamed
health/guides/fping/fping_host_latency.md
renamed
health/guides/fping/fping_host_reachable.md
renamed
health/guides/gearman/gearman_workers_queued.md
renamed
health/guides/geth/geth_chainhead_diff_between_header_block.md
renamed
health/guides/go.d_job_last_collected_secs.md
deleted
-37
@@ -1,37 +0,0 @@
1
-### Understand the alert
2
-
3
-The Netdata Agent also monitors itself, so this is an alert about the Netdata go.d plugin. The Netdata Agent keeps track of the number of seconds since the last successful data collection for each data collection job.
4
-This alert indicates that a particular job has failed to collect metrics for several consecutive attempts.
5
-
6
-You can see all the modules that are orchestrated by the go.d.plugin in our [go.d.plugin documentation page](https://learn.netdata.cloud/docs/agent/collectors/go.d.plugin#available-modules)
7
-
8
-### Troubleshoot the alert
9
-
10
-- Check the Netdata logs
11
-
12
-You need to identify why the Agent cannot collect metrics for a specific job. Inspect the Agent logs for this specific job.
13
-
14
-Host machine:
15
-
16
- ```
17
- tail -f /var/log/netdata/error.log | grep <job> OR <module_name>
18
- ```
19
-
20
-Docker:
21
-
22
- ```
23
- docker logs <netdata_container> 2>&1 | grep <job> OR <module_name>
24
- ```
25
-
26
-Kubernetes:
27
- 1. Find the pod name of the node which produced the alert.
28
-
29
- ```
30
- kubectl -n <namespace> get pod -o wide -l app=netdata | grep <node_name>
31
- ```
32
- 2. Inspect it's logs
33
-
34
- ```
35
- kubectl logs -n <namespace> <pod_name> | grep <job> OR <module_name>
36
- ```
37
-
\ No newline at end of file
health/guides/haproxy/haproxy_backend_server_status.md
renamed
health/guides/haproxy/haproxy_backend_status.md
renamed
health/guides/hdfs/hdfs_capacity_usage.md
renamed
health/guides/hdfs/hdfs_dead_nodes.md
renamed
health/guides/hdfs/hdfs_missing_blocks.md
renamed
health/guides/hdfs/hdfs_num_failed_volumes.md
renamed
health/guides/hdfs/hdfs_stale_nodes.md
renamed
health/guides/httpcheck/httpcheck_web_service_bad_content.md
renamed
health/guides/httpcheck/httpcheck_web_service_bad_status.md
renamed
health/guides/httpcheck/httpcheck_web_service_no_connection.md
renamed
health/guides/httpcheck/httpcheck_web_service_slow.md
renamed
health/guides/httpcheck/httpcheck_web_service_timeouts.md
renamed
health/guides/httpcheck/httpcheck_web_service_unreachable.md
renamed
health/guides/httpcheck/httpcheck_web_service_up.md
renamed
health/guides/ioping/ioping_disk_latency.md
renamed
health/guides/ipc/semaphore_arrays_used.md
renamed
health/guides/ipc/semaphores_used.md
renamed
health/guides/ipfs/ipfs_datastore_usage.md
renamed
health/guides/ipmi/ipmi_events.md
renamed
health/guides/ipmi/ipmi_sensors_states.md
renamed
health/guides/kubelet/kubelet_10s_pleg_relist_latency_quantile_05.md
renamed
health/guides/kubelet/kubelet_10s_pleg_relist_latency_quantile_09.md
renamed
health/guides/kubelet/kubelet_10s_pleg_relist_latency_quantile_099.md
renamed
health/guides/kubelet/kubelet_1m_pleg_relist_latency_quantile_05.md
renamed
health/guides/kubelet/kubelet_1m_pleg_relist_latency_quantile_09.md
renamed
health/guides/kubelet/kubelet_1m_pleg_relist_latency_quantile_099.md
renamed
health/guides/kubelet/kubelet_node_config_error.md
renamed
health/guides/kubelet/kubelet_operations_error.md
renamed
health/guides/kubelet/kubelet_token_requests.md
renamed
health/guides/linux_power_supply/linux_power_supply_capacity.md
renamed
health/guides/load/load_average_1.md
renamed
health/guides/load/load_average_15.md
renamed
health/guides/load/load_average_5.md
renamed
health/guides/load/load_cpu_number.md
renamed
health/guides/mdstat/mdstat_disks.md
renamed
health/guides/mdstat/mdstat_last_collected.md
renamed
health/guides/mdstat/mdstat_mismatch_cnt.md
renamed
health/guides/mdstat/mdstat_nonredundant_last_collected.md
renamed
health/guides/megacli/megacli_adapter_state.md
renamed
health/guides/megacli/megacli_bbu_cycle_count.md
renamed
health/guides/megacli/megacli_bbu_relative_charge.md
renamed
health/guides/megacli/megacli_pd_media_errors.md
renamed
health/guides/megacli/megacli_pd_predictive_failures.md
renamed
health/guides/memcached/memcached_cache_fill_rate.md
renamed
health/guides/memcached/memcached_cache_memory_usage.md
renamed
health/guides/memcached/memcached_out_of_cache_space_time.md
renamed
health/guides/memory/1hour_ecc_memory_correctable.md
renamed
health/guides/memory/1hour_ecc_memory_uncorrectable.md
renamed
health/guides/memory/1hour_memory_hw_corrupted.md
renamed
health/guides/ml/ml_1min_node_ar.md
renamed
+26
-26
@@ -1,26 +1,26 @@
1
-### Understand the alert
2
-
3
-This alert is triggered when the [node anomaly rate](https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection#node-anomaly-rate) exceeds the threshold defined in the [alert configuration](https://github.com/netdata/netdata/blob/master/health/health.d/ml.conf) over the most recent 1 minute window evaluated.
4
-
5
-For example, with the default of `warn: $this > 1`, this means that 1% or more of the metrics collected on the node have across the most recent 1 minute window been flagged as [anomalous](https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection) by Netdata.
6
-
7
-### Troubleshoot the alert
8
-
9
-This alert is a signal that some significant percentage of metrics within your infrastructure have been flagged as anomalous accoring to the ML based anomaly detection models the Netdata agent continually trains and re-trains for each metric. This tells us something somewhere might look strange in some way. THe next step is to try drill in and see what metrics are actually driving this.
10
-
11
-1. **Filter for the node or nodes relevant**: First we need to reduce as much noise as possible by filtering for just those nodes that have the elevated node anomaly rate. Look at the `anomaly_detection.anomaly_rate` chart and group by `node` to see which nodes have an elevated anomaly rate. Filter for just those nodes since this will reduce any noise as much as possible.
12
-
13
-2. **Highlight the area of interest**: Highlight the timeframne of interest where you see an elevated anomaly rate.
14
-
15
-3. **Check the anomalies tab**: Check the [Anomaly Advisor](https://learn.netdata.cloud/docs/ml-and-troubleshooting/anomaly-advisor) ("Anomalies" tab) to see an ordered list of what metrics were most anomalous in the highlighted window.
16
-
17
-4. **Press the AR% button on Overview**: You can also press the "[AR%](https://blog.netdata.cloud/anomaly-rates-in-the-menu/)" button on the Overview or single node dashboard to see what parts of the menu have the highest chart anomaly rates. Pressing the AR% button should add some "pills" to each menu item and if you hover over it you will see that chart within each menu section that was most anomalous during the highlighted timeframe.
18
-
19
-5. **Use Metric Correlations**: Use [metric correlations](https://learn.netdata.cloud/docs/ml-and-troubleshooting/metric-correlations) to see what metrics may have changed most significantly comparing before to the highlighted timeframe.
20
-
21
-### Useful resources
22
-
23
-1. [Machine learning (ML) powered anomaly detection](https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection)
24
-2. [Anomaly Advisor](https://learn.netdata.cloud/docs/ml-and-troubleshooting/anomaly-advisor)
25
-3. [Metric Correlations](https://learn.netdata.cloud/docs/ml-and-troubleshooting/metric-correlations)
26
-4. [Anomaly Rates in the Menu!](https://blog.netdata.cloud/anomaly-rates-in-the-menu/)
1
+### Understand the alert
2
+
3
+This alert is triggered when the [node anomaly rate](https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection#node-anomaly-rate) exceeds the threshold defined in the [alert configuration](https://github.com/netdata/netdata/blob/master/health/health.d/ml.conf) over the most recent 1 minute window evaluated.
4
+
5
+For example, with the default of `warn: $this > 1`, this means that 1% or more of the metrics collected on the node have across the most recent 1 minute window been flagged as [anomalous](https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection) by Netdata.
6
+
7
+### Troubleshoot the alert
8
+
9
+This alert is a signal that some significant percentage of metrics within your infrastructure have been flagged as anomalous accoring to the ML based anomaly detection models the Netdata agent continually trains and re-trains for each metric. This tells us something somewhere might look strange in some way. THe next step is to try drill in and see what metrics are actually driving this.
10
+
11
+1. **Filter for the node or nodes relevant**: First we need to reduce as much noise as possible by filtering for just those nodes that have the elevated node anomaly rate. Look at the `anomaly_detection.anomaly_rate` chart and group by `node` to see which nodes have an elevated anomaly rate. Filter for just those nodes since this will reduce any noise as much as possible.
12
+
13
+2. **Highlight the area of interest**: Highlight the timeframne of interest where you see an elevated anomaly rate.
14
+
15
+3. **Check the anomalies tab**: Check the [Anomaly Advisor](https://learn.netdata.cloud/docs/ml-and-troubleshooting/anomaly-advisor) ("Anomalies" tab) to see an ordered list of what metrics were most anomalous in the highlighted window.
16
+
17
+4. **Press the AR% button on Overview**: You can also press the "[AR%](https://blog.netdata.cloud/anomaly-rates-in-the-menu/)" button on the Overview or single node dashboard to see what parts of the menu have the highest chart anomaly rates. Pressing the AR% button should add some "pills" to each menu item and if you hover over it you will see that chart within each menu section that was most anomalous during the highlighted timeframe.
18
+
19
+5. **Use Metric Correlations**: Use [metric correlations](https://learn.netdata.cloud/docs/ml-and-troubleshooting/metric-correlations) to see what metrics may have changed most significantly comparing before to the highlighted timeframe.
20
+
21
+### Useful resources
22
+
23
+1. [Machine learning (ML) powered anomaly detection](https://learn.netdata.cloud/docs/ml-and-troubleshooting/machine-learning-ml-powered-anomaly-detection)
24
+2. [Anomaly Advisor](https://learn.netdata.cloud/docs/ml-and-troubleshooting/anomaly-advisor)
25
+3. [Metric Correlations](https://learn.netdata.cloud/docs/ml-and-troubleshooting/metric-correlations)
26
+4. [Anomaly Rates in the Menu!](https://blog.netdata.cloud/anomaly-rates-in-the-menu/)
health/guides/mysql/mysql_10s_slow_queries.md
renamed
health/guides/mysql/mysql_10s_table_locks_immediate.md
renamed
health/guides/mysql/mysql_10s_table_locks_waited.md
renamed
health/guides/mysql/mysql_10s_waited_locks_ratio.md
renamed
health/guides/mysql/mysql_connections.md
renamed
health/guides/mysql/mysql_galera_cluster_size.md
renamed
health/guides/mysql/mysql_galera_cluster_size_max_2m.md
renamed
health/guides/mysql/mysql_galera_cluster_state_crit.md
renamed
health/guides/mysql/mysql_galera_cluster_state_warn.md
renamed
health/guides/mysql/mysql_galera_cluster_status.md
renamed
health/guides/mysql/mysql_replication.md
renamed
health/guides/mysql/mysql_replication_lag.md
renamed
health/guides/net/10min_fifo_errors.md
renamed
health/guides/net/10min_netisr_backlog_exceeded.md
renamed
health/guides/net/10s_received_packets_storm.md
renamed
health/guides/net/1m_received_packets_rate.md
renamed
health/guides/net/1m_received_traffic_overflow.md
renamed
health/guides/net/1m_sent_traffic_overflow.md
renamed
health/guides/net/inbound_packets_dropped.md
renamed
health/guides/net/inbound_packets_dropped_ratio.md
renamed
health/guides/net/interface_inbound_errors.md
renamed
health/guides/net/interface_outbound_errors.md
renamed
health/guides/net/interface_speed.md
renamed
health/guides/net/outbound_packets_dropped.md
renamed
health/guides/net/outbound_packets_dropped_ratio.md
renamed
health/guides/netdev/1min_netdev_backlog_exceeded.md
renamed
health/guides/netdev/1min_netdev_budget_ran_outs.md
renamed
health/guides/netfilter/netfilter_conntrack_full.md
renamed
-3
@@ -37,9 +37,6 @@ sysctl -w net.netfilter.nf_conntrack_max=<YOUR DESIRED LIMIT HERE>
37
echo "net.netfilter.nf_conntrack_max=<YOUR DESIRED LIMIT HERE>" >> /etc/sysctl.conf
38
```
39
40
-</details>
41
-
42
-
40
### Useful resources
41
42
1. [Netfilter](https://en.wikipedia.org/wiki/Netfilter)
health/guides/nut/nut_10min_ups_load.md
renamed
health/guides/nut/nut_last_collected_secs.md
renamed
health/guides/nut/nut_ups_charge.md
renamed
health/guides/nvme/nvme_device_critical_warnings_state.md
renamed
health/guides/pihole/pihole_blocklist_last_update.md
renamed
health/guides/pihole/pihole_status.md
renamed
health/guides/ping/ping_host_latency.md
renamed
health/guides/ping/ping_host_reachable.md
renamed
health/guides/ping/ping_packet_loss.md
renamed
health/guides/portcheck/portcheck_connection_fails.md
renamed
health/guides/portcheck/portcheck_connection_timeouts.md
renamed
health/guides/portcheck/portcheck_service_reachable.md
renamed
health/guides/postgres/postgres_acquired_locks_utilization.md
renamed
health/guides/postgres/postgres_db_cache_io_ratio.md
renamed
health/guides/postgres/postgres_db_deadlocks_rate.md
renamed
health/guides/postgres/postgres_db_transactions_rollback_ratio.md
renamed
health/guides/postgres/postgres_index_bloat_size_perc.md
renamed
health/guides/postgres/postgres_table_bloat_size_perc.md
renamed
health/guides/postgres/postgres_table_cache_io_ratio.md
renamed
health/guides/postgres/postgres_table_index_cache_io_ratio.md
renamed
health/guides/postgres/postgres_table_last_autoanalyze_time.md
renamed
health/guides/postgres/postgres_table_last_autovacuum_time.md
renamed
health/guides/postgres/postgres_table_toast_cache_io_ratio.md
renamed
health/guides/postgres/postgres_table_toast_index_cache_io_ratio.md
renamed
health/guides/postgres/postgres_total_connection_utilization.md
renamed
health/guides/postgres/postgres_txid_exhaustion_perc.md
renamed
health/guides/processes/active_processes.md
renamed
health/guides/python.d_job_last_collected_secs.md
deleted
-35
@@ -1,35 +0,0 @@
1
-### Understand the alert
2
-
3
-The Netdata Agent also monitors itself, so this is an alert about the Netdata Python plugin. The Netdata Agent monitors the number of seconds since the last successful data collection for each of python plugin modules. This alert indicates that a specific module cannot reach the component it monitors to collect metrics from it.
4
-
5
-You can see all the modules that are orchestrated by the python.d.plugin in [our GitHub repo](https://github.com/netdata/netdata/tree/master/collectors/python.d.plugin)
6
-
7
-### Troubleshoot the alert
8
-
9
-- Check the Netdata logs
10
-
11
-You need to identify why the Agent cannot collect metrics for a specific job. Inspect the Agent logs for this specific job.
12
-
13
-Host machine:
14
-
15
- ```
16
- tail -f /var/log/netdata/error.log | grep <job> OR <module_name>
17
- ```
18
-
19
-Docker:
20
-
21
- ```
22
- docker logs <netdata_container> 2>&1 | grep <job> OR <module_name>
23
- ```
24
-
25
-Kubernetes:
26
- 1. Find the pod name of the node which produced the alert.
27
-
28
- ```
29
- kubectl -n <namespace> get pod -o wide -l app=netdata | grep <node_name>
30
- ```
31
- 2. Inspect it's logs
32
-
33
- ```
34
- kubectl logs -n <namespace> <pod_name> | grep <job> OR <module_name>
35
- ```