Regenerate integrations docs (#21998)
Co-authored-by: ilyam8 <22274335+ilyam8@users.noreply.github.com>
Netdata bot committed
Mar 21, 2026 at 00:03 UTC
5919f18c9259065cea6de0aeb632a91da1fb4cfd
44 files changed
+18084
-5
src/collectors/COLLECTORS.md
+40
@@ -314,6 +314,46 @@ Need a dedicated integration? [Submit a feature request](https://github.com/netd
314
|-------------|-------------|
315
| [AWS EC2 Compute instances](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/aws_ec2_compute_instances.md) | Track AWS EC2 instances key metrics for optimized performance and cost management. |
316
| [AWS Quota](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/aws_quota.md) | Monitor AWS service quotas for effective resource usage and cost management. |
317
+| [Azure API Management](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_api_management.md) | Monitor API Management gateway performance including request throughput, response status codes, gateway and backend response times, failed request counts, capacity utilization, event hub events, websocket message counts, and network connection status. |
318
+| [Azure App Service](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_app_service.md) | Monitor App Service web applications including HTTP request rates and response status codes, response times, CPU and memory usage, network throughput, file IO operations, .NET runtime statistics (threads, GC, assemblies), Azure Functions execution counts and units, and Flex Consumption plan metrics. |
319
+| [Azure Application Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_gateway.md) | Monitor Application Gateway performance including throughput and traffic volume, request rates and response status codes, backend health and latency breakdown (connect, first byte, last byte), client latency, current and new connections, WebSocket sessions, capacity and compute units, CPU utilization, TLS connections, and WAF security events including rule matches, challenges, and penalty box activity. |
320
+| [Azure Application Insights](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_insights.md) | Monitor application performance through Application Insights including availability test results and duration, server request rates and response times, dependency call tracking and failures, exception rates by source, browser page load timing breakdown, process CPU and memory usage, IO rates, HTTP request queue depth, page views, and trace volume. |
321
+| [Azure Cache for Redis](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cache_for_redis.md) | Monitor Azure Cache for Redis including cache hit and miss rates, read and write throughput, server load and CPU utilization, memory usage, connected clients, operations per second, command processing rates, latency percentiles, key eviction and expiration, and geo-replication health and sync status. |
322
+| [Azure Cognitive Services](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cognitive_services.md) | Monitor Azure AI and Cognitive Services including API call volume, success and client error rates, response latency, token processing rates for language models, content safety filtering, fine-tuning operations, provisioned throughput utilization, rate-limiting events, active inference connections, and context token cache performance. |
323
+| [Azure Container Apps](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_apps.md) | Monitor Container Apps including CPU and memory usage, network traffic, replica counts, request processing rates, response times, restart frequency, and resource reservation utilization. |
324
+| [Azure Container Instances](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_instances.md) | Monitor Container Instance groups including CPU and memory usage and network bytes transferred in and out. |
325
+| [Azure Container Registry](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_registry.md) | Monitor Container Registry including storage usage, successful and failed pull and push operation counts, and task run duration. |
326
+| [Azure Cosmos DB Account](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cosmos_db_account.md) | Monitor Cosmos DB accounts including request unit consumption and throttling, document counts and storage, data and index sizes, replication latency, availability percentages, provisioned throughput utilization, and normalized RU consumption per partition. |
327
+| [Azure Data Explorer Cluster](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_explorer_cluster.md) | Monitor Azure Data Explorer (Kusto) clusters including ingestion latency, volume, and success rates, query performance and concurrency, cache utilization, CPU and memory usage, export operations, streaming ingest throughput, materialized view health, instance counts, and follower lag. |
328
+| [Azure Data Factory](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_factory.md) | Monitor Data Factory including pipeline, activity, and trigger run success and failure counts, integration runtime CPU and memory utilization, available capacity and queue lengths, SSIS package execution rates, copy operations throughput, data flow processing metrics, and overall factory resource utilization. |
329
+| [Azure Event Grid Topic](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_grid_topic.md) | Monitor Event Grid topics including publish success and failure counts, publish latency, event delivery and routing rates, delivery success and failure counts, dead-lettered events, and matched event routing. |
330
+| [Azure Event Hubs Namespace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_hubs_namespace.md) | Monitor Event Hubs namespaces including incoming and outgoing message rates, byte throughput, captured messages and bytes, throttled and quota-exceeded request counts, active connections, and total connection counts. |
331
+| [Azure ExpressRoute Circuit](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_circuit.md) | Monitor ExpressRoute circuits including bits per second in and out, ARP and BGP availability percentages, packet drops, and QoS bit rate throughput. |
332
+| [Azure ExpressRoute Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_gateway.md) | Monitor ExpressRoute gateways including bits and packets per second for ingress and egress, connection counts, CPU utilization, active flow counts, and gateway scale unit counts. |
333
+| [Azure Firewall](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_firewall.md) | Monitor Azure Firewall including data processed, throughput, application and network rule hit counts, SNAT port utilization, health state percentage, and latency probes. |
334
+| [Azure Front Door](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_front_door.md) | Monitor Azure Front Door including request counts and rates, response sizes, total latency, origin health probe percentages, origin request counts, origin latency, WAF request counts by action and rule, and WebSocket connection metrics. |
335
+| [Azure Functions](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_functions.md) | Monitor Azure Functions execution including function invocation counts, execution units (MB-milliseconds), HTTP request rates and response codes, CPU and memory consumption, and Flex Consumption plan metrics for always-ready and on-demand instances. |
336
+| [Azure IoT Hub](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_iot_hub.md) | Monitor IoT Hub including device telemetry message rates and quota usage, routing delivery and latency, device twin read and write operations, direct method invocations, cloud-to-device messaging and feedback, job completion rates, device connection and authentication events, and event grid publish status. |
337
+| [Azure Key Vault](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_key_vault.md) | Monitor Key Vault including overall vault availability, API saturation approaching service limits, and service API hit and latency metrics. |
338
+| [Azure Kubernetes Service Cluster](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_kubernetes_service_cluster.md) | Monitor AKS cluster health including API server and etcd resource usage, pod scheduling status and readiness, node capacity and conditions, cluster autoscaler behavior, and per-node CPU, memory, disk, and network utilization. |
339
+| [Azure Load Balancer](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_load_balancer.md) | Monitor Azure Load Balancer health and throughput including data path and health probe availability, SYN and SNAT connection counts, byte and packet throughput, allocated and used SNAT ports, and connection attempt rates. |
340
+| [Azure Log Analytics Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_log_analytics_workspace.md) | Monitor Log Analytics workspaces including ingestion volume and latency, query execution counts and volume, available storage capacity, and per-table breakdowns of ingestion rates and billing volume. |
341
+| [Azure Logic Apps Workflow](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_logic_apps_workflow.md) | Monitor Logic Apps workflow execution including run completions and failures, action execution counts, trigger firing rates, run and action latency, billable executions, and action-level success and failure breakdowns. |
342
+| [Azure Machine Learning Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_machine_learning_workspace.md) | Monitor Azure Machine Learning workspaces including active model deployments and registered models, pipeline run completions and failures, compute node utilization and preemptions, quota usage, managed endpoint request latency and rates, estimated GPU utilization, and storage utilization. |
343
+| [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) | This collector monitors Azure resources through the Azure Monitor Metrics API. |
344
+| [Azure MySQL Flexible Server](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_mysql_flexible_server.md) | Monitor MySQL Flexible Server including active connections, aborted connections, query rates, replication lag, storage utilization, CPU and memory usage, IO operations, InnoDB buffer pool efficiency, network throughput, and HA replication status. |
345
+| [Azure NAT Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_nat_gateway.md) | Monitor NAT Gateway including byte and packet counts, connection counts, dropped packets, total SNAT connection counts, and datapath availability. |
346
+| [Azure PostgreSQL Flexible Server](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_postgresql_flexible_server.md) | Monitor PostgreSQL Flexible Server including active connections, transaction rates, replication lag, storage and backup utilization, CPU and memory usage, IO throughput, autovacuum activity, PgBouncer connection pooling, database sessions, and burstable instance CPU credits. |
347
+| [Azure Service Bus Namespace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_service_bus_namespace.md) | Monitor Service Bus namespaces including incoming and outgoing message rates, active connections, active and dead-lettered message counts, scheduled message counts, completed and abandoned requests, server errors, throttled requests, CPU and memory utilization, and pending checkpoint operations. |
348
+| [Azure SQL Database](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_database.md) | Monitor SQL Database performance including CPU and DTU utilization, storage consumption, active sessions and workers, deadlocks, IO rates, tempdb usage, in-memory OLTP storage, and serverless auto-pause and billing metrics. |
349
+| [Azure SQL Elastic Pool](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_elastic_pool.md) | Monitor SQL Elastic Pool resource consumption including eDTU and CPU utilization, storage usage, active sessions and workers, IO rates, tempdb usage, and in-memory OLTP storage across all databases in the pool. |
350
+| [Azure SQL Managed Instance](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_managed_instance.md) | Monitor SQL Managed Instance performance including virtual core CPU utilization, storage consumption, IO throughput, and average request wait times. |
351
+| [Azure Storage Account](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_storage_account.md) | Monitor Azure Storage Account operations including transaction counts, availability percentages, success and end-to-end latency, ingress and egress throughput, and used capacity. |
352
+| [Azure Stream Analytics Job](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_stream_analytics_job.md) | Monitor Stream Analytics jobs including input and output event counts, streaming unit utilization, watermark delay, backlogged input events, runtime and data conversion errors, out-of-order events, and late input events. |
353
+| [Azure Synapse Analytics Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_synapse_analytics_workspace.md) | Monitor Synapse Analytics workspaces including pipeline and activity run metrics, SQL request counts and data processing volumes, data flow activity execution, integration runtime CPU and memory utilization, and link table event processing. |
354
+| [Azure Virtual Machine](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine.md) | Monitor Azure Virtual Machines including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates. |
355
+| [Azure Virtual Machine Scale Set](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine_scale_set.md) | Monitor Virtual Machine Scale Sets including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates across all instances in the scale set. |
356
+| [Azure VPN Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_vpn_gateway.md) | Monitor VPN Gateway including site-to-site bandwidth and BGP peer status, point-to-site connection counts and bandwidth, per-tunnel ingress and egress traffic with packet counts and drops, IPsec security association counts, route table sizes, NAT flow counts and packet translations, and gateway-level bandwidth utilization. |
357
| [BOSH](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/bosh.md) | Keep an eye on BOSH deployment metrics for improved cloud orchestration and resource management. |
358
| [Cloud Foundry](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/cloud_foundry.md) | Track Cloud Foundry platform metrics for optimized application deployment and management. |
359
| [Cloud Foundry Firehose](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/cloud_foundry_firehose.md) | Monitor Cloud Foundry Firehose metrics for comprehensive platform diagnostics and management. |
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_api_management.md
new
+441
@@ -0,0 +1,441 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_api_management.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure API Management"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'api', 'management', 'gateway', 'apim']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure API Management
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor API Management gateway performance including request throughput, response status codes, gateway and backend response times, failed request counts, capacity utilization, event hub events, websocket message counts, and network connection status.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.api_management.capacity | capacity | percentage |
88
+| azure_monitor.api_management.gateway_cpu | cpu | percentage |
89
+| azure_monitor.api_management.gateway_memory | memory | percentage |
90
+| azure_monitor.api_management.requests | requests | requests/s |
91
+| azure_monitor.api_management.request_duration | overall, backend | milliseconds |
92
+| azure_monitor.api_management.eventhub_events | total, successful, failed, dropped, rejected, throttled, timed_out | events/s |
93
+| azure_monitor.api_management.eventhub_bytes | sent | bytes/s |
94
+| azure_monitor.api_management.websocket_connections | connection_attempts | attempts/s |
95
+| azure_monitor.api_management.websocket_messages | messages | messages/s |
96
+| azure_monitor.api_management.network_connectivity | connectivity | status |
97
+
98
+
99
+
100
+## Alerts
101
+
102
+
103
+The following alerts are available:
104
+
105
+| Alert name | On metric | Description |
106
+|:------------|:----------|:------------|
107
+| [ am_api_management_capacity ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.capacity | APIM capacity on ${label:resource_name} |
108
+| [ am_api_management_gateway_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.gateway_cpu | APIM gateway CPU on ${label:resource_name} |
109
+| [ am_api_management_gateway_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.gateway_memory | APIM gateway memory on ${label:resource_name} |
110
+| [ am_api_management_request_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.request_duration | APIM request duration on ${label:resource_name} |
111
+| [ am_api_management_backend_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.request_duration | APIM backend duration on ${label:resource_name} |
112
+| [ am_api_management_network_connectivity ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.network_connectivity | APIM network connectivity on ${label:resource_name} |
113
+| [ am_api_management_eventhub_failed_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.eventhub_events | APIM EventHub failed events on ${label:resource_name} |
114
+| [ am_api_management_eventhub_dropped_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.eventhub_events | APIM EventHub dropped events on ${label:resource_name} |
115
+| [ am_api_management_eventhub_rejected_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.eventhub_events | APIM EventHub rejected events on ${label:resource_name} |
116
+| [ am_api_management_eventhub_throttled_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.eventhub_events | APIM EventHub throttled events on ${label:resource_name} |
117
+| [ am_api_management_eventhub_timedout_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_api_management.conf) | azure_monitor.api_management.eventhub_events | APIM EventHub timed out events on ${label:resource_name} |
118
+
119
+
120
+## Setup
121
+
122
+
123
+You can configure the **azure_monitor** collector in two ways:
124
+
125
+| Method | Best for | How to |
126
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
127
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
128
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
129
+
130
+:::important
131
+
132
+UI configuration requires paid Netdata Cloud plan.
133
+
134
+:::
135
+
136
+
137
+### Prerequisites
138
+
139
+#### Create an Azure monitoring principal
140
+
141
+Create a service principal or use a managed identity with the following permissions:
142
+
143
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
144
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
145
+
146
+For service principal authentication:
147
+```bash
148
+# Create the service principal
149
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
150
+ --scopes /subscriptions/<subscription-id>
151
+
152
+# Note the appId (client_id), password (client_secret), and tenant
153
+```
154
+
155
+For managed identity (on Azure VMs, VMSS, or AKS):
156
+```bash
157
+# Assign Monitoring Reader role to the VM's managed identity
158
+az role assignment create --assignee <managed-identity-principal-id> \
159
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
160
+```
161
+
162
+
163
+
164
+### Configuration
165
+
166
+#### Options
167
+
168
+The following options can be defined globally: update_every, autodetection_retry.
169
+
170
+Profile files are loaded from:
171
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
172
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
173
+
174
+User profile files with the same filename override stock profiles.
175
+
176
+
177
+<details open><summary>Config options</summary>
178
+
179
+
180
+
181
+| Group | Option | Description | Default | Required |
182
+|:------|:-----|:------------|:--------|:---------:|
183
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
184
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
185
+| **Target** | subscription_id | Azure subscription ID. | | yes |
186
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
187
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
188
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
189
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
190
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
191
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
192
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
193
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
194
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
195
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
196
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
197
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
198
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
199
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
200
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
+
202
+
203
+</details>
204
+
205
+
206
+#### via UI
207
+
208
+Configure the **azure_monitor** collector from the Netdata web interface:
209
+
210
+1. Go to **Nodes**.
211
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
212
+3. The **Collectors → Jobs** view opens by default.
213
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
214
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
215
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
216
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
217
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
218
+
219
+
220
+#### via File
221
+
222
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
223
+
224
+The file format is YAML. Generally, the structure is:
225
+
226
+```yaml
227
+update_every: 1
228
+autodetection_retry: 0
229
+jobs:
230
+ - name: some_name1
231
+ - name: some_name2
232
+```
233
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
234
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
235
+
236
+```bash
237
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
238
+sudo ./edit-config go.d/azure_monitor.conf
239
+```
240
+
241
+##### Examples
242
+
243
+###### Service principal (auto-discover all resources)
244
+
245
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
246
+
247
+```yaml
248
+jobs:
249
+ - name: prod
250
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
+ auth:
252
+ mode: service_principal
253
+ mode_service_principal:
254
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
255
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
256
+ client_secret: "your-client-secret"
257
+
258
+```
259
+###### Managed identity (Azure VM/VMSS/AKS)
260
+
261
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
262
+
263
+<details open><summary>Config</summary>
264
+
265
+```yaml
266
+jobs:
267
+ - name: prod
268
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
269
+ auth:
270
+ mode: managed_identity
271
+
272
+```
273
+</details>
274
+
275
+###### Specific profiles only
276
+
277
+Monitor only specific Azure services instead of auto-discovering all resource types.
278
+
279
+<details open><summary>Config</summary>
280
+
281
+```yaml
282
+jobs:
283
+ - name: databases
284
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
285
+ profiles:
286
+ - sql_database
287
+ - postgres_flexible
288
+ - redis_cache
289
+ auth:
290
+ mode: service_principal
291
+ mode_service_principal:
292
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
293
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
294
+ client_secret: "your-client-secret"
295
+
296
+```
297
+</details>
298
+
299
+###### Filter by resource group
300
+
301
+Only monitor resources in specific resource groups.
302
+
303
+<details open><summary>Config</summary>
304
+
305
+```yaml
306
+jobs:
307
+ - name: prod-rg
308
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
309
+ resource_groups:
310
+ - production-rg
311
+ - staging-rg
312
+ auth:
313
+ mode: default
314
+
315
+```
316
+</details>
317
+
318
+###### Azure Government cloud
319
+
320
+Connect to Azure Government cloud environment.
321
+
322
+<details open><summary>Config</summary>
323
+
324
+```yaml
325
+jobs:
326
+ - name: gov
327
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
328
+ cloud: government
329
+ auth:
330
+ mode: service_principal
331
+ mode_service_principal:
332
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
333
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
334
+ client_secret: "your-client-secret"
335
+
336
+```
337
+</details>
338
+
339
+
340
+
341
+## Troubleshooting
342
+
343
+### Debug Mode
344
+
345
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
346
+
347
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
348
+should give you clues as to why the collector isn't working.
349
+
350
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
351
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
352
+
353
+ ```bash
354
+ cd /usr/libexec/netdata/plugins.d/
355
+ ```
356
+
357
+- Switch to the `netdata` user.
358
+
359
+ ```bash
360
+ sudo -u netdata -s
361
+ ```
362
+
363
+- Run the `go.d.plugin` to debug the collector:
364
+
365
+ ```bash
366
+ ./go.d.plugin -d -m azure_monitor
367
+ ```
368
+
369
+ To debug a specific job:
370
+
371
+ ```bash
372
+ ./go.d.plugin -d -m azure_monitor -j jobName
373
+ ```
374
+
375
+### Getting Logs
376
+
377
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
378
+
379
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
380
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
381
+
382
+#### System with systemd
383
+
384
+Use the following command to view logs generated since the last Netdata service restart:
385
+
386
+```bash
387
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
388
+```
389
+
390
+#### System without systemd
391
+
392
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
393
+
394
+```bash
395
+grep azure_monitor /var/log/netdata/collector.log
396
+```
397
+
398
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
399
+
400
+#### Docker Container
401
+
402
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
403
+
404
+```bash
405
+docker logs netdata 2>&1 | grep azure_monitor
406
+```
407
+
408
+### No metrics are collected
409
+
410
+Verify the following:
411
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
412
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
413
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
414
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
415
+
416
+
417
+### Missing metrics for some resource types
418
+
419
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
420
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
421
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
422
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
423
+
424
+
425
+### Metrics appear delayed
426
+
427
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
428
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
429
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
430
+
431
+
432
+### Authentication errors in sovereign clouds
433
+
434
+For Azure Government or Azure China clouds, set the `cloud` parameter:
435
+- Azure Government: `cloud: government`
436
+- Azure China (21Vianet): `cloud: china`
437
+
438
+Ensure the service principal is registered in the correct cloud tenant.
439
+
440
+
441
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_app_service.md
new
+454
@@ -0,0 +1,454 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_app_service.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure App Service"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'app', 'service', 'web', 'webapp', 'paas']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure App Service
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor App Service web applications including HTTP request rates and response status codes, response times, CPU and memory usage, network throughput, file IO operations, .NET runtime statistics (threads, GC, assemblies), Azure Functions execution counts and units, and Flex Consumption plan metrics.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.app_service.requests | requests | requests/s |
88
+| azure_monitor.app_service.response_time | average | seconds |
89
+| azure_monitor.app_service.http_status | 2xx, 3xx, 4xx, 5xx | responses/s |
90
+| azure_monitor.app_service.http_error_detail | 101_websocket, 401_unauthorized, 403_forbidden, 404_not_found, 406_not_acceptable | responses/s |
91
+| azure_monitor.app_service.cpu | average | percentage |
92
+| azure_monitor.app_service.cpu_time | total | seconds |
93
+| azure_monitor.app_service.memory_usage | average_working_set, working_set, private | bytes |
94
+| azure_monitor.app_service.health | average | percentage |
95
+| azure_monitor.app_service.network_traffic | received, sent | bytes/s |
96
+| azure_monitor.app_service.connections | average | connections |
97
+| azure_monitor.app_service.io_throughput | read, write, other | bytes/s |
98
+| azure_monitor.app_service.io_operations | read, write, other | operations/s |
99
+| azure_monitor.app_service.threads | average | threads |
100
+| azure_monitor.app_service.handles | average | handles |
101
+| azure_monitor.app_service.gc_collections | gen0, gen1, gen2 | collections/s |
102
+| azure_monitor.app_service.function_executions | total | executions/s |
103
+| azure_monitor.app_service.function_execution_units | total | MB-milliseconds/s |
104
+| azure_monitor.app_service.always_ready_function_executions | total | executions/s |
105
+| azure_monitor.app_service.always_ready_function_execution_units | total | MB-milliseconds/s |
106
+| azure_monitor.app_service.always_ready_units | total | units |
107
+| azure_monitor.app_service.on_demand_function_executions | total | executions/s |
108
+| azure_monitor.app_service.on_demand_function_execution_units | total | MB-milliseconds/s |
109
+| azure_monitor.app_service.request_queue | queued | requests |
110
+| azure_monitor.app_service.instances | running | instances |
111
+| azure_monitor.app_service.assemblies | loaded | assemblies |
112
+| azure_monitor.app_service.app_domains | loaded, unloaded | domains |
113
+
114
+
115
+
116
+## Alerts
117
+
118
+
119
+The following alerts are available:
120
+
121
+| Alert name | On metric | Description |
122
+|:------------|:----------|:------------|
123
+| [ am_app_service_health_check ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.health | App Service health on ${label:resource_name} |
124
+| [ am_app_service_http_5xx_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.http_status | App Service 5xx errors on ${label:resource_name} |
125
+| [ am_app_service_response_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.response_time | App Service response time on ${label:resource_name} |
126
+| [ am_app_service_cpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.cpu | App Service CPU on ${label:resource_name} |
127
+| [ am_app_service_request_queue ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.request_queue | App Service request queue on ${label:resource_name} |
128
+| [ am_app_service_http_4xx_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.http_status | App Service 4xx errors on ${label:resource_name} |
129
+| [ am_app_service_http_403_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.http_error_detail | App Service 403 forbidden on ${label:resource_name} |
130
+| [ am_app_service_http_401_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_app_service.conf) | azure_monitor.app_service.http_error_detail | App Service 401 unauthorized on ${label:resource_name} |
131
+
132
+
133
+## Setup
134
+
135
+
136
+You can configure the **azure_monitor** collector in two ways:
137
+
138
+| Method | Best for | How to |
139
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
140
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
141
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
142
+
143
+:::important
144
+
145
+UI configuration requires paid Netdata Cloud plan.
146
+
147
+:::
148
+
149
+
150
+### Prerequisites
151
+
152
+#### Create an Azure monitoring principal
153
+
154
+Create a service principal or use a managed identity with the following permissions:
155
+
156
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
157
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
158
+
159
+For service principal authentication:
160
+```bash
161
+# Create the service principal
162
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
163
+ --scopes /subscriptions/<subscription-id>
164
+
165
+# Note the appId (client_id), password (client_secret), and tenant
166
+```
167
+
168
+For managed identity (on Azure VMs, VMSS, or AKS):
169
+```bash
170
+# Assign Monitoring Reader role to the VM's managed identity
171
+az role assignment create --assignee <managed-identity-principal-id> \
172
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
173
+```
174
+
175
+
176
+
177
+### Configuration
178
+
179
+#### Options
180
+
181
+The following options can be defined globally: update_every, autodetection_retry.
182
+
183
+Profile files are loaded from:
184
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
185
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
186
+
187
+User profile files with the same filename override stock profiles.
188
+
189
+
190
+<details open><summary>Config options</summary>
191
+
192
+
193
+
194
+| Group | Option | Description | Default | Required |
195
+|:------|:-----|:------------|:--------|:---------:|
196
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
197
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
198
+| **Target** | subscription_id | Azure subscription ID. | | yes |
199
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
200
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
201
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
202
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
203
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
204
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
205
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
206
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
207
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
208
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
209
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
210
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
211
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
212
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
213
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
214
+
215
+
216
+</details>
217
+
218
+
219
+#### via UI
220
+
221
+Configure the **azure_monitor** collector from the Netdata web interface:
222
+
223
+1. Go to **Nodes**.
224
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
225
+3. The **Collectors → Jobs** view opens by default.
226
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
227
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
228
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
229
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
230
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
231
+
232
+
233
+#### via File
234
+
235
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
236
+
237
+The file format is YAML. Generally, the structure is:
238
+
239
+```yaml
240
+update_every: 1
241
+autodetection_retry: 0
242
+jobs:
243
+ - name: some_name1
244
+ - name: some_name2
245
+```
246
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
247
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
248
+
249
+```bash
250
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
251
+sudo ./edit-config go.d/azure_monitor.conf
252
+```
253
+
254
+##### Examples
255
+
256
+###### Service principal (auto-discover all resources)
257
+
258
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
259
+
260
+```yaml
261
+jobs:
262
+ - name: prod
263
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
264
+ auth:
265
+ mode: service_principal
266
+ mode_service_principal:
267
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
268
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
269
+ client_secret: "your-client-secret"
270
+
271
+```
272
+###### Managed identity (Azure VM/VMSS/AKS)
273
+
274
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
275
+
276
+<details open><summary>Config</summary>
277
+
278
+```yaml
279
+jobs:
280
+ - name: prod
281
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
282
+ auth:
283
+ mode: managed_identity
284
+
285
+```
286
+</details>
287
+
288
+###### Specific profiles only
289
+
290
+Monitor only specific Azure services instead of auto-discovering all resource types.
291
+
292
+<details open><summary>Config</summary>
293
+
294
+```yaml
295
+jobs:
296
+ - name: databases
297
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
298
+ profiles:
299
+ - sql_database
300
+ - postgres_flexible
301
+ - redis_cache
302
+ auth:
303
+ mode: service_principal
304
+ mode_service_principal:
305
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
306
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
307
+ client_secret: "your-client-secret"
308
+
309
+```
310
+</details>
311
+
312
+###### Filter by resource group
313
+
314
+Only monitor resources in specific resource groups.
315
+
316
+<details open><summary>Config</summary>
317
+
318
+```yaml
319
+jobs:
320
+ - name: prod-rg
321
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ resource_groups:
323
+ - production-rg
324
+ - staging-rg
325
+ auth:
326
+ mode: default
327
+
328
+```
329
+</details>
330
+
331
+###### Azure Government cloud
332
+
333
+Connect to Azure Government cloud environment.
334
+
335
+<details open><summary>Config</summary>
336
+
337
+```yaml
338
+jobs:
339
+ - name: gov
340
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
341
+ cloud: government
342
+ auth:
343
+ mode: service_principal
344
+ mode_service_principal:
345
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
346
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
347
+ client_secret: "your-client-secret"
348
+
349
+```
350
+</details>
351
+
352
+
353
+
354
+## Troubleshooting
355
+
356
+### Debug Mode
357
+
358
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
359
+
360
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
361
+should give you clues as to why the collector isn't working.
362
+
363
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
364
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
365
+
366
+ ```bash
367
+ cd /usr/libexec/netdata/plugins.d/
368
+ ```
369
+
370
+- Switch to the `netdata` user.
371
+
372
+ ```bash
373
+ sudo -u netdata -s
374
+ ```
375
+
376
+- Run the `go.d.plugin` to debug the collector:
377
+
378
+ ```bash
379
+ ./go.d.plugin -d -m azure_monitor
380
+ ```
381
+
382
+ To debug a specific job:
383
+
384
+ ```bash
385
+ ./go.d.plugin -d -m azure_monitor -j jobName
386
+ ```
387
+
388
+### Getting Logs
389
+
390
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
391
+
392
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
393
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
394
+
395
+#### System with systemd
396
+
397
+Use the following command to view logs generated since the last Netdata service restart:
398
+
399
+```bash
400
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
401
+```
402
+
403
+#### System without systemd
404
+
405
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
406
+
407
+```bash
408
+grep azure_monitor /var/log/netdata/collector.log
409
+```
410
+
411
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
412
+
413
+#### Docker Container
414
+
415
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
416
+
417
+```bash
418
+docker logs netdata 2>&1 | grep azure_monitor
419
+```
420
+
421
+### No metrics are collected
422
+
423
+Verify the following:
424
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
425
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
426
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
427
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
428
+
429
+
430
+### Missing metrics for some resource types
431
+
432
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
433
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
434
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
435
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
436
+
437
+
438
+### Metrics appear delayed
439
+
440
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
441
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
442
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
443
+
444
+
445
+### Authentication errors in sovereign clouds
446
+
447
+For Azure Government or Azure China clouds, set the `cloud` parameter:
448
+- Azure Government: `cloud: government`
449
+- Azure China (21Vianet): `cloud: china`
450
+
451
+Ensure the service principal is registered in the correct cloud tenant.
452
+
453
+
454
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_gateway.md
new
+447
@@ -0,0 +1,447 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_gateway.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Application Gateway"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'application', 'gateway', 'load', 'balancer', 'waf']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Application Gateway
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Application Gateway performance including throughput and traffic volume, request rates and response status codes, backend health and latency breakdown (connect, first byte, last byte), client latency, current and new connections, WebSocket sessions, capacity and compute units, CPU utilization, TLS connections, and WAF security events including rule matches, challenges, and penalty box activity.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.application_gateway.throughput | average | bytes/s |
88
+| azure_monitor.application_gateway.traffic_volume | received, sent | bytes/s |
89
+| azure_monitor.application_gateway.requests | total, failed | requests/s |
90
+| azure_monitor.application_gateway.response_status | gateway, backend | responses/s |
91
+| azure_monitor.application_gateway.backend_health | healthy, unhealthy | hosts |
92
+| azure_monitor.application_gateway.backend_request_load | per_healthy_host | requests |
93
+| azure_monitor.application_gateway.backend_latency | connect, first_byte, last_byte | milliseconds |
94
+| azure_monitor.application_gateway.client_latency | total_time, client_rtt | milliseconds |
95
+| azure_monitor.application_gateway.current_connections | current | connections |
96
+| azure_monitor.application_gateway.new_connections | average | connections/s |
97
+| azure_monitor.application_gateway.websocket_connections | active | connections |
98
+| azure_monitor.application_gateway.websocket_close_codes | total | connections/s |
99
+| azure_monitor.application_gateway.capacity | capacity, compute, billed, fixed_billed | units |
100
+| azure_monitor.application_gateway.cpu | average | percentage |
101
+| azure_monitor.application_gateway.tls_connections | total | connections/s |
102
+| azure_monitor.application_gateway.waf_requests | total, blocked, matched | requests/s |
103
+| azure_monitor.application_gateway.waf_rule_matches | managed, custom, bot | matches/s |
104
+| azure_monitor.application_gateway.waf_challenges | captcha, js_challenge | requests/s |
105
+| azure_monitor.application_gateway.waf_penalty_box | size | IPs |
106
+| azure_monitor.application_gateway.waf_penalty_box_hits | total | hits/s |
107
+
108
+
109
+
110
+## Alerts
111
+
112
+
113
+The following alerts are available:
114
+
115
+| Alert name | On metric | Description |
116
+|:------------|:----------|:------------|
117
+| [ am_appgw_failed_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.requests | App Gateway failed requests on ${label:resource_name} |
118
+| [ am_appgw_unhealthy_hosts ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.backend_health | App Gateway unhealthy backends on ${label:resource_name} |
119
+| [ am_appgw_backend_connect_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.backend_latency | App Gateway backend connect time on ${label:resource_name} |
120
+| [ am_appgw_backend_first_byte ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.backend_latency | App Gateway backend TTFB on ${label:resource_name} |
121
+| [ am_appgw_total_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.client_latency | App Gateway total request time on ${label:resource_name} |
122
+| [ am_appgw_cpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.cpu | App Gateway CPU utilization on ${label:resource_name} |
123
+| [ am_appgw_waf_blocked_ratio ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_gateway.conf) | azure_monitor.application_gateway.waf_requests | App Gateway WAF block ratio on ${label:resource_name} |
124
+
125
+
126
+## Setup
127
+
128
+
129
+You can configure the **azure_monitor** collector in two ways:
130
+
131
+| Method | Best for | How to |
132
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
133
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
134
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
135
+
136
+:::important
137
+
138
+UI configuration requires paid Netdata Cloud plan.
139
+
140
+:::
141
+
142
+
143
+### Prerequisites
144
+
145
+#### Create an Azure monitoring principal
146
+
147
+Create a service principal or use a managed identity with the following permissions:
148
+
149
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
150
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
151
+
152
+For service principal authentication:
153
+```bash
154
+# Create the service principal
155
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
156
+ --scopes /subscriptions/<subscription-id>
157
+
158
+# Note the appId (client_id), password (client_secret), and tenant
159
+```
160
+
161
+For managed identity (on Azure VMs, VMSS, or AKS):
162
+```bash
163
+# Assign Monitoring Reader role to the VM's managed identity
164
+az role assignment create --assignee <managed-identity-principal-id> \
165
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
166
+```
167
+
168
+
169
+
170
+### Configuration
171
+
172
+#### Options
173
+
174
+The following options can be defined globally: update_every, autodetection_retry.
175
+
176
+Profile files are loaded from:
177
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
178
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
179
+
180
+User profile files with the same filename override stock profiles.
181
+
182
+
183
+<details open><summary>Config options</summary>
184
+
185
+
186
+
187
+| Group | Option | Description | Default | Required |
188
+|:------|:-----|:------------|:--------|:---------:|
189
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
190
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
191
+| **Target** | subscription_id | Azure subscription ID. | | yes |
192
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
193
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
194
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
195
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
196
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
197
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
198
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
199
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
200
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
201
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
202
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
203
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
204
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
205
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
206
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
207
+
208
+
209
+</details>
210
+
211
+
212
+#### via UI
213
+
214
+Configure the **azure_monitor** collector from the Netdata web interface:
215
+
216
+1. Go to **Nodes**.
217
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
218
+3. The **Collectors → Jobs** view opens by default.
219
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
220
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
221
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
222
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
223
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
224
+
225
+
226
+#### via File
227
+
228
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
229
+
230
+The file format is YAML. Generally, the structure is:
231
+
232
+```yaml
233
+update_every: 1
234
+autodetection_retry: 0
235
+jobs:
236
+ - name: some_name1
237
+ - name: some_name2
238
+```
239
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
240
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
241
+
242
+```bash
243
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
244
+sudo ./edit-config go.d/azure_monitor.conf
245
+```
246
+
247
+##### Examples
248
+
249
+###### Service principal (auto-discover all resources)
250
+
251
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
252
+
253
+```yaml
254
+jobs:
255
+ - name: prod
256
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
257
+ auth:
258
+ mode: service_principal
259
+ mode_service_principal:
260
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
261
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
262
+ client_secret: "your-client-secret"
263
+
264
+```
265
+###### Managed identity (Azure VM/VMSS/AKS)
266
+
267
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
268
+
269
+<details open><summary>Config</summary>
270
+
271
+```yaml
272
+jobs:
273
+ - name: prod
274
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
275
+ auth:
276
+ mode: managed_identity
277
+
278
+```
279
+</details>
280
+
281
+###### Specific profiles only
282
+
283
+Monitor only specific Azure services instead of auto-discovering all resource types.
284
+
285
+<details open><summary>Config</summary>
286
+
287
+```yaml
288
+jobs:
289
+ - name: databases
290
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
291
+ profiles:
292
+ - sql_database
293
+ - postgres_flexible
294
+ - redis_cache
295
+ auth:
296
+ mode: service_principal
297
+ mode_service_principal:
298
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
300
+ client_secret: "your-client-secret"
301
+
302
+```
303
+</details>
304
+
305
+###### Filter by resource group
306
+
307
+Only monitor resources in specific resource groups.
308
+
309
+<details open><summary>Config</summary>
310
+
311
+```yaml
312
+jobs:
313
+ - name: prod-rg
314
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
315
+ resource_groups:
316
+ - production-rg
317
+ - staging-rg
318
+ auth:
319
+ mode: default
320
+
321
+```
322
+</details>
323
+
324
+###### Azure Government cloud
325
+
326
+Connect to Azure Government cloud environment.
327
+
328
+<details open><summary>Config</summary>
329
+
330
+```yaml
331
+jobs:
332
+ - name: gov
333
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
334
+ cloud: government
335
+ auth:
336
+ mode: service_principal
337
+ mode_service_principal:
338
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
339
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
340
+ client_secret: "your-client-secret"
341
+
342
+```
343
+</details>
344
+
345
+
346
+
347
+## Troubleshooting
348
+
349
+### Debug Mode
350
+
351
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
352
+
353
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
354
+should give you clues as to why the collector isn't working.
355
+
356
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
357
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
358
+
359
+ ```bash
360
+ cd /usr/libexec/netdata/plugins.d/
361
+ ```
362
+
363
+- Switch to the `netdata` user.
364
+
365
+ ```bash
366
+ sudo -u netdata -s
367
+ ```
368
+
369
+- Run the `go.d.plugin` to debug the collector:
370
+
371
+ ```bash
372
+ ./go.d.plugin -d -m azure_monitor
373
+ ```
374
+
375
+ To debug a specific job:
376
+
377
+ ```bash
378
+ ./go.d.plugin -d -m azure_monitor -j jobName
379
+ ```
380
+
381
+### Getting Logs
382
+
383
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
384
+
385
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
386
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
387
+
388
+#### System with systemd
389
+
390
+Use the following command to view logs generated since the last Netdata service restart:
391
+
392
+```bash
393
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
394
+```
395
+
396
+#### System without systemd
397
+
398
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
399
+
400
+```bash
401
+grep azure_monitor /var/log/netdata/collector.log
402
+```
403
+
404
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
405
+
406
+#### Docker Container
407
+
408
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
409
+
410
+```bash
411
+docker logs netdata 2>&1 | grep azure_monitor
412
+```
413
+
414
+### No metrics are collected
415
+
416
+Verify the following:
417
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
418
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
419
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
420
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
421
+
422
+
423
+### Missing metrics for some resource types
424
+
425
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
426
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
427
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
428
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
429
+
430
+
431
+### Metrics appear delayed
432
+
433
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
434
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
435
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
436
+
437
+
438
+### Authentication errors in sovereign clouds
439
+
440
+For Azure Government or Azure China clouds, set the `cloud` parameter:
441
+- Azure Government: `cloud: government`
442
+- Azure China (21Vianet): `cloud: china`
443
+
444
+Ensure the service principal is registered in the correct cloud tenant.
445
+
446
+
447
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_insights.md
new
+455
@@ -0,0 +1,455 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_insights.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Application Insights"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'application', 'insights', 'apm', 'monitoring', 'telemetry']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Application Insights
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor application performance through Application Insights including availability test results and duration, server request rates and response times, dependency call tracking and failures, exception rates by source, browser page load timing breakdown, process CPU and memory usage, IO rates, HTTP request queue depth, page views, and trace volume.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.application_insights.availability_percentage | average | percentage |
88
+| azure_monitor.application_insights.availability_tests | tests | tests/s |
89
+| azure_monitor.application_insights.availability_duration | average | milliseconds |
90
+| azure_monitor.application_insights.server_requests | total, failed | requests/s |
91
+| azure_monitor.application_insights.server_request_rate | average | requests/s |
92
+| azure_monitor.application_insights.server_response_time | average | milliseconds |
93
+| azure_monitor.application_insights.dependency_calls | total, failed | calls/s |
94
+| azure_monitor.application_insights.dependency_duration | average | milliseconds |
95
+| azure_monitor.application_insights.exceptions | total, browser, server | exceptions/s |
96
+| azure_monitor.application_insights.browser_page_load_time | total | milliseconds |
97
+| azure_monitor.application_insights.browser_timing_breakdown | network, send, receive, processing | milliseconds |
98
+| azure_monitor.application_insights.cpu_utilization | process, processor | percentage |
99
+| azure_monitor.application_insights.memory | available, private | bytes |
100
+| azure_monitor.application_insights.process_io_rate | average | bytes/s |
101
+| azure_monitor.application_insights.exception_rate | average | exceptions/s |
102
+| azure_monitor.application_insights.http_request_execution_time | average | milliseconds |
103
+| azure_monitor.application_insights.http_request_queue | queued | requests |
104
+| azure_monitor.application_insights.http_request_rate | average | requests/s |
105
+| azure_monitor.application_insights.page_views | views | views/s |
106
+| azure_monitor.application_insights.page_view_load_time | average | milliseconds |
107
+| azure_monitor.application_insights.traces | traces | traces/s |
108
+
109
+
110
+
111
+## Alerts
112
+
113
+
114
+The following alerts are available:
115
+
116
+| Alert name | On metric | Description |
117
+|:------------|:----------|:------------|
118
+| [ am_appinsights_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.availability_percentage | App Insights availability on ${label:resource_name} |
119
+| [ am_appinsights_availability_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.availability_duration | App Insights availability test duration on ${label:resource_name} |
120
+| [ am_appinsights_failed_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.server_requests | App Insights failed requests on ${label:resource_name} |
121
+| [ am_appinsights_response_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.server_response_time | App Insights server response time on ${label:resource_name} |
122
+| [ am_appinsights_failed_dependencies ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.dependency_calls | App Insights failed dependencies on ${label:resource_name} |
123
+| [ am_appinsights_dependency_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.dependency_duration | App Insights dependency duration on ${label:resource_name} |
124
+| [ am_appinsights_exception_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.exception_rate | App Insights exception rate on ${label:resource_name} |
125
+| [ am_appinsights_server_exceptions ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.exceptions | App Insights server exceptions on ${label:resource_name} |
126
+| [ am_appinsights_browser_page_load_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.browser_page_load_time | App Insights browser page load time on ${label:resource_name} |
127
+| [ am_appinsights_process_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.cpu_utilization | App Insights process CPU on ${label:resource_name} |
128
+| [ am_appinsights_processor_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.cpu_utilization | App Insights processor CPU on ${label:resource_name} |
129
+| [ am_appinsights_http_execution_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.http_request_execution_time | App Insights HTTP execution time on ${label:resource_name} |
130
+| [ am_appinsights_http_queue_length ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.http_request_queue | App Insights HTTP request queue on ${label:resource_name} |
131
+| [ am_appinsights_page_view_load_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_application_insights.conf) | azure_monitor.application_insights.page_view_load_time | App Insights page view load time on ${label:resource_name} |
132
+
133
+
134
+## Setup
135
+
136
+
137
+You can configure the **azure_monitor** collector in two ways:
138
+
139
+| Method | Best for | How to |
140
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
141
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
142
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
143
+
144
+:::important
145
+
146
+UI configuration requires paid Netdata Cloud plan.
147
+
148
+:::
149
+
150
+
151
+### Prerequisites
152
+
153
+#### Create an Azure monitoring principal
154
+
155
+Create a service principal or use a managed identity with the following permissions:
156
+
157
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
158
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
159
+
160
+For service principal authentication:
161
+```bash
162
+# Create the service principal
163
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
164
+ --scopes /subscriptions/<subscription-id>
165
+
166
+# Note the appId (client_id), password (client_secret), and tenant
167
+```
168
+
169
+For managed identity (on Azure VMs, VMSS, or AKS):
170
+```bash
171
+# Assign Monitoring Reader role to the VM's managed identity
172
+az role assignment create --assignee <managed-identity-principal-id> \
173
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
174
+```
175
+
176
+
177
+
178
+### Configuration
179
+
180
+#### Options
181
+
182
+The following options can be defined globally: update_every, autodetection_retry.
183
+
184
+Profile files are loaded from:
185
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
186
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
187
+
188
+User profile files with the same filename override stock profiles.
189
+
190
+
191
+<details open><summary>Config options</summary>
192
+
193
+
194
+
195
+| Group | Option | Description | Default | Required |
196
+|:------|:-----|:------------|:--------|:---------:|
197
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
198
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
199
+| **Target** | subscription_id | Azure subscription ID. | | yes |
200
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
201
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
202
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
203
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
204
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
205
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
206
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
207
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
208
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
209
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
210
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
211
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
212
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
213
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
214
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
215
+
216
+
217
+</details>
218
+
219
+
220
+#### via UI
221
+
222
+Configure the **azure_monitor** collector from the Netdata web interface:
223
+
224
+1. Go to **Nodes**.
225
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
226
+3. The **Collectors → Jobs** view opens by default.
227
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
228
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
229
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
230
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
231
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
232
+
233
+
234
+#### via File
235
+
236
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
237
+
238
+The file format is YAML. Generally, the structure is:
239
+
240
+```yaml
241
+update_every: 1
242
+autodetection_retry: 0
243
+jobs:
244
+ - name: some_name1
245
+ - name: some_name2
246
+```
247
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
248
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
249
+
250
+```bash
251
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
252
+sudo ./edit-config go.d/azure_monitor.conf
253
+```
254
+
255
+##### Examples
256
+
257
+###### Service principal (auto-discover all resources)
258
+
259
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
260
+
261
+```yaml
262
+jobs:
263
+ - name: prod
264
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
265
+ auth:
266
+ mode: service_principal
267
+ mode_service_principal:
268
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
269
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
270
+ client_secret: "your-client-secret"
271
+
272
+```
273
+###### Managed identity (Azure VM/VMSS/AKS)
274
+
275
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
276
+
277
+<details open><summary>Config</summary>
278
+
279
+```yaml
280
+jobs:
281
+ - name: prod
282
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ auth:
284
+ mode: managed_identity
285
+
286
+```
287
+</details>
288
+
289
+###### Specific profiles only
290
+
291
+Monitor only specific Azure services instead of auto-discovering all resource types.
292
+
293
+<details open><summary>Config</summary>
294
+
295
+```yaml
296
+jobs:
297
+ - name: databases
298
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ profiles:
300
+ - sql_database
301
+ - postgres_flexible
302
+ - redis_cache
303
+ auth:
304
+ mode: service_principal
305
+ mode_service_principal:
306
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
307
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
308
+ client_secret: "your-client-secret"
309
+
310
+```
311
+</details>
312
+
313
+###### Filter by resource group
314
+
315
+Only monitor resources in specific resource groups.
316
+
317
+<details open><summary>Config</summary>
318
+
319
+```yaml
320
+jobs:
321
+ - name: prod-rg
322
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ resource_groups:
324
+ - production-rg
325
+ - staging-rg
326
+ auth:
327
+ mode: default
328
+
329
+```
330
+</details>
331
+
332
+###### Azure Government cloud
333
+
334
+Connect to Azure Government cloud environment.
335
+
336
+<details open><summary>Config</summary>
337
+
338
+```yaml
339
+jobs:
340
+ - name: gov
341
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
342
+ cloud: government
343
+ auth:
344
+ mode: service_principal
345
+ mode_service_principal:
346
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
347
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
348
+ client_secret: "your-client-secret"
349
+
350
+```
351
+</details>
352
+
353
+
354
+
355
+## Troubleshooting
356
+
357
+### Debug Mode
358
+
359
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
360
+
361
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
362
+should give you clues as to why the collector isn't working.
363
+
364
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
365
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
366
+
367
+ ```bash
368
+ cd /usr/libexec/netdata/plugins.d/
369
+ ```
370
+
371
+- Switch to the `netdata` user.
372
+
373
+ ```bash
374
+ sudo -u netdata -s
375
+ ```
376
+
377
+- Run the `go.d.plugin` to debug the collector:
378
+
379
+ ```bash
380
+ ./go.d.plugin -d -m azure_monitor
381
+ ```
382
+
383
+ To debug a specific job:
384
+
385
+ ```bash
386
+ ./go.d.plugin -d -m azure_monitor -j jobName
387
+ ```
388
+
389
+### Getting Logs
390
+
391
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
392
+
393
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
394
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
395
+
396
+#### System with systemd
397
+
398
+Use the following command to view logs generated since the last Netdata service restart:
399
+
400
+```bash
401
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
402
+```
403
+
404
+#### System without systemd
405
+
406
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
407
+
408
+```bash
409
+grep azure_monitor /var/log/netdata/collector.log
410
+```
411
+
412
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
413
+
414
+#### Docker Container
415
+
416
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
417
+
418
+```bash
419
+docker logs netdata 2>&1 | grep azure_monitor
420
+```
421
+
422
+### No metrics are collected
423
+
424
+Verify the following:
425
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
426
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
427
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
428
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
429
+
430
+
431
+### Missing metrics for some resource types
432
+
433
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
434
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
435
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
436
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
437
+
438
+
439
+### Metrics appear delayed
440
+
441
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
442
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
443
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
444
+
445
+
446
+### Authentication errors in sovereign clouds
447
+
448
+For Azure Government or Azure China clouds, set the `cloud` parameter:
449
+- Azure Government: `cloud: government`
450
+- Azure China (21Vianet): `cloud: china`
451
+
452
+Ensure the service principal is registered in the correct cloud tenant.
453
+
454
+
455
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cache_for_redis.md
new
+463
@@ -0,0 +1,463 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cache_for_redis.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Cache for Redis"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'redis', 'cache', 'memory', 'nosql']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Cache for Redis
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Cache for Redis including cache hit and miss rates, read and write throughput, server load and CPU utilization, memory usage, connected clients, operations per second, command processing rates, latency percentiles, key eviction and expiration, and geo-replication health and sync status. Provides per-shard breakdowns for clustered deployments.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.redis_cache.cache_hits | hits, misses | operations/s |
88
+| azure_monitor.redis_cache.throughput | read, write | bytes/s |
89
+| azure_monitor.redis_cache.server_load | maximum | percentage |
90
+| azure_monitor.redis_cache.cpu | maximum | percentage |
91
+| azure_monitor.redis_cache.memory_usage | used, rss | bytes |
92
+| azure_monitor.redis_cache.memory_utilization | maximum | percentage |
93
+| azure_monitor.redis_cache.clients | maximum | clients |
94
+| azure_monitor.redis_cache.operations | maximum | operations/s |
95
+| azure_monitor.redis_cache.commands | total | commands/s |
96
+| azure_monitor.redis_cache.command_types | get, set | commands/s |
97
+| azure_monitor.redis_cache.latency | average | microseconds |
98
+| azure_monitor.redis_cache.key_events | evicted, expired | keys/s |
99
+| azure_monitor.redis_cache.total_keys | maximum | keys |
100
+| azure_monitor.redis_cache.errors | maximum | errors |
101
+| azure_monitor.redis_cache.miss_rate | miss_rate | percentage |
102
+| azure_monitor.redis_cache.latency_p99 | p99 | microseconds |
103
+| azure_monitor.redis_cache.connection_rate | created, closed | connections/s |
104
+| azure_monitor.redis_cache.aad_clients | maximum | clients |
105
+| azure_monitor.redis_cache.instance_cache_hits | hits, misses | operations/s |
106
+| azure_monitor.redis_cache.instance_throughput | read, write | bytes/s |
107
+| azure_monitor.redis_cache.instance_server_load | server_load, cpu | percentage |
108
+| azure_monitor.redis_cache.instance_memory_usage | used, rss | bytes |
109
+| azure_monitor.redis_cache.instance_memory_utilization | maximum | percentage |
110
+| azure_monitor.redis_cache.instance_clients | maximum | clients |
111
+| azure_monitor.redis_cache.instance_operations | maximum | operations/s |
112
+| azure_monitor.redis_cache.instance_commands | total, get, set | commands/s |
113
+| azure_monitor.redis_cache.instance_total_keys | maximum | keys |
114
+| azure_monitor.redis_cache.instance_key_events | evicted, expired | keys/s |
115
+| azure_monitor.redis_cache.geo_replication_lag | average | seconds |
116
+| azure_monitor.redis_cache.geo_replication_sync_offset | average | bytes |
117
+| azure_monitor.redis_cache.geo_replication_sync_events | started, finished | events/s |
118
+| azure_monitor.redis_cache.geo_replication_health | average | status |
119
+
120
+
121
+
122
+## Alerts
123
+
124
+
125
+The following alerts are available:
126
+
127
+| Alert name | On metric | Description |
128
+|:------------|:----------|:------------|
129
+| [ am_redis_cache_server_load ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.server_load | Redis server load on ${label:resource_name} |
130
+| [ am_redis_cache_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.cpu | Redis CPU on ${label:resource_name} |
131
+| [ am_redis_cache_memory_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.memory_utilization | Redis memory utilization on ${label:resource_name} |
132
+| [ am_redis_cache_miss_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.miss_rate | Redis cache miss rate on ${label:resource_name} |
133
+| [ am_redis_cache_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.errors | Redis errors on ${label:resource_name} |
134
+| [ am_redis_cache_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.latency | Redis average latency on ${label:resource_name} |
135
+| [ am_redis_cache_latency_p99 ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.latency_p99 | Redis P99 latency on ${label:resource_name} |
136
+| [ am_redis_cache_evicted_keys ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.key_events | Redis key evictions on ${label:resource_name} |
137
+| [ am_redis_cache_geo_replication_health ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.geo_replication_health | Redis geo-replication health on ${label:resource_name} |
138
+| [ am_redis_cache_geo_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.geo_replication_lag | Redis geo-replication lag on ${label:resource_name} |
139
+| [ am_redis_cache_geo_replication_sync_offset ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_redis_cache.conf) | azure_monitor.redis_cache.geo_replication_sync_offset | Redis geo-replication sync offset on ${label:resource_name} |
140
+
141
+
142
+## Setup
143
+
144
+
145
+You can configure the **azure_monitor** collector in two ways:
146
+
147
+| Method | Best for | How to |
148
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
149
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
150
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
151
+
152
+:::important
153
+
154
+UI configuration requires paid Netdata Cloud plan.
155
+
156
+:::
157
+
158
+
159
+### Prerequisites
160
+
161
+#### Create an Azure monitoring principal
162
+
163
+Create a service principal or use a managed identity with the following permissions:
164
+
165
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
166
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
167
+
168
+For service principal authentication:
169
+```bash
170
+# Create the service principal
171
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
172
+ --scopes /subscriptions/<subscription-id>
173
+
174
+# Note the appId (client_id), password (client_secret), and tenant
175
+```
176
+
177
+For managed identity (on Azure VMs, VMSS, or AKS):
178
+```bash
179
+# Assign Monitoring Reader role to the VM's managed identity
180
+az role assignment create --assignee <managed-identity-principal-id> \
181
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
182
+```
183
+
184
+
185
+
186
+### Configuration
187
+
188
+#### Options
189
+
190
+The following options can be defined globally: update_every, autodetection_retry.
191
+
192
+Profile files are loaded from:
193
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
194
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
195
+
196
+User profile files with the same filename override stock profiles.
197
+
198
+
199
+<details open><summary>Config options</summary>
200
+
201
+
202
+
203
+| Group | Option | Description | Default | Required |
204
+|:------|:-----|:------------|:--------|:---------:|
205
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
206
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
207
+| **Target** | subscription_id | Azure subscription ID. | | yes |
208
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
209
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
210
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
211
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
212
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
213
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
214
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
215
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
216
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
217
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
218
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
219
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
220
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
221
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
222
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
223
+
224
+
225
+</details>
226
+
227
+
228
+#### via UI
229
+
230
+Configure the **azure_monitor** collector from the Netdata web interface:
231
+
232
+1. Go to **Nodes**.
233
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
234
+3. The **Collectors → Jobs** view opens by default.
235
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
236
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
237
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
238
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
239
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
240
+
241
+
242
+#### via File
243
+
244
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
245
+
246
+The file format is YAML. Generally, the structure is:
247
+
248
+```yaml
249
+update_every: 1
250
+autodetection_retry: 0
251
+jobs:
252
+ - name: some_name1
253
+ - name: some_name2
254
+```
255
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
256
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
257
+
258
+```bash
259
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
260
+sudo ./edit-config go.d/azure_monitor.conf
261
+```
262
+
263
+##### Examples
264
+
265
+###### Service principal (auto-discover all resources)
266
+
267
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
268
+
269
+```yaml
270
+jobs:
271
+ - name: prod
272
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
273
+ auth:
274
+ mode: service_principal
275
+ mode_service_principal:
276
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
277
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
278
+ client_secret: "your-client-secret"
279
+
280
+```
281
+###### Managed identity (Azure VM/VMSS/AKS)
282
+
283
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
284
+
285
+<details open><summary>Config</summary>
286
+
287
+```yaml
288
+jobs:
289
+ - name: prod
290
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
291
+ auth:
292
+ mode: managed_identity
293
+
294
+```
295
+</details>
296
+
297
+###### Specific profiles only
298
+
299
+Monitor only specific Azure services instead of auto-discovering all resource types.
300
+
301
+<details open><summary>Config</summary>
302
+
303
+```yaml
304
+jobs:
305
+ - name: databases
306
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
307
+ profiles:
308
+ - sql_database
309
+ - postgres_flexible
310
+ - redis_cache
311
+ auth:
312
+ mode: service_principal
313
+ mode_service_principal:
314
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
315
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
316
+ client_secret: "your-client-secret"
317
+
318
+```
319
+</details>
320
+
321
+###### Filter by resource group
322
+
323
+Only monitor resources in specific resource groups.
324
+
325
+<details open><summary>Config</summary>
326
+
327
+```yaml
328
+jobs:
329
+ - name: prod-rg
330
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
331
+ resource_groups:
332
+ - production-rg
333
+ - staging-rg
334
+ auth:
335
+ mode: default
336
+
337
+```
338
+</details>
339
+
340
+###### Azure Government cloud
341
+
342
+Connect to Azure Government cloud environment.
343
+
344
+<details open><summary>Config</summary>
345
+
346
+```yaml
347
+jobs:
348
+ - name: gov
349
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
350
+ cloud: government
351
+ auth:
352
+ mode: service_principal
353
+ mode_service_principal:
354
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ client_secret: "your-client-secret"
357
+
358
+```
359
+</details>
360
+
361
+
362
+
363
+## Troubleshooting
364
+
365
+### Debug Mode
366
+
367
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
368
+
369
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
370
+should give you clues as to why the collector isn't working.
371
+
372
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
373
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
374
+
375
+ ```bash
376
+ cd /usr/libexec/netdata/plugins.d/
377
+ ```
378
+
379
+- Switch to the `netdata` user.
380
+
381
+ ```bash
382
+ sudo -u netdata -s
383
+ ```
384
+
385
+- Run the `go.d.plugin` to debug the collector:
386
+
387
+ ```bash
388
+ ./go.d.plugin -d -m azure_monitor
389
+ ```
390
+
391
+ To debug a specific job:
392
+
393
+ ```bash
394
+ ./go.d.plugin -d -m azure_monitor -j jobName
395
+ ```
396
+
397
+### Getting Logs
398
+
399
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
400
+
401
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
402
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
403
+
404
+#### System with systemd
405
+
406
+Use the following command to view logs generated since the last Netdata service restart:
407
+
408
+```bash
409
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
410
+```
411
+
412
+#### System without systemd
413
+
414
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
415
+
416
+```bash
417
+grep azure_monitor /var/log/netdata/collector.log
418
+```
419
+
420
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
421
+
422
+#### Docker Container
423
+
424
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
425
+
426
+```bash
427
+docker logs netdata 2>&1 | grep azure_monitor
428
+```
429
+
430
+### No metrics are collected
431
+
432
+Verify the following:
433
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
434
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
435
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
436
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
437
+
438
+
439
+### Missing metrics for some resource types
440
+
441
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
442
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
443
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
444
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
445
+
446
+
447
+### Metrics appear delayed
448
+
449
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
450
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
451
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
452
+
453
+
454
+### Authentication errors in sovereign clouds
455
+
456
+For Azure Government or Azure China clouds, set the `cloud` parameter:
457
+- Azure Government: `cloud: government`
458
+- Azure China (21Vianet): `cloud: china`
459
+
460
+Ensure the service principal is registered in the correct cloud tenant.
461
+
462
+
463
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cognitive_services.md
new
+510
@@ -0,0 +1,510 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cognitive_services.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Cognitive Services"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'cognitive', 'ai', 'openai', 'gpt', 'ml']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Cognitive Services
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure AI and Cognitive Services including API call volume, success and client error rates, response latency, token processing rates for language models, content safety filtering, fine-tuning operations, provisioned throughput utilization, rate-limiting events, active inference connections, and context token cache performance.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.cognitive_services.availability | availability | percentage |
88
+| azure_monitor.cognitive_services.calls | total, successful, blocked, token | calls/s |
89
+| azure_monitor.cognitive_services.errors | total, client, server | errors/s |
90
+| azure_monitor.cognitive_services.latency | average | milliseconds |
91
+| azure_monitor.cognitive_services.data_transfer | in, out | bytes/s |
92
+| azure_monitor.cognitive_services.rate_limit | rate_limit | requests/s |
93
+| azure_monitor.cognitive_services.openai_availability | availability | percentage |
94
+| azure_monitor.cognitive_services.openai_requests | requests | requests/s |
95
+| azure_monitor.cognitive_services.openai_latency | time_to_response, time_to_first_token, time_between_tokens, time_to_last_byte | milliseconds |
96
+| azure_monitor.cognitive_services.openai_generation_speed | tokens_per_second | tokens/s |
97
+| azure_monitor.cognitive_services.openai_token_usage | total, prompt, generated, active | tokens/s |
98
+| azure_monitor.cognitive_services.openai_audio_tokens | prompt, completion | tokens/s |
99
+| azure_monitor.cognitive_services.openai_cache_match_rate | cache_match_rate | percentage |
100
+| azure_monitor.cognitive_services.openai_provisioned_utilization | utilization | percentage |
101
+| azure_monitor.cognitive_services.openai_finetuning | training_hours | hours/s |
102
+| azure_monitor.cognitive_services.openai_realtime_usage | seconds_used | seconds/s |
103
+| azure_monitor.cognitive_services.model_availability | availability | percentage |
104
+| azure_monitor.cognitive_services.model_requests | requests | requests/s |
105
+| azure_monitor.cognitive_services.model_latency | time_to_response, time_to_first_token, time_between_tokens, time_to_last_byte | milliseconds |
106
+| azure_monitor.cognitive_services.model_generation_speed | tokens_per_second | tokens/s |
107
+| azure_monitor.cognitive_services.model_token_usage | total, input, output | tokens/s |
108
+| azure_monitor.cognitive_services.model_audio_tokens | input, output | tokens/s |
109
+| azure_monitor.cognitive_services.model_cache_tokens | cache_read, cache_write_1h, cache_write_5m | tokens/s |
110
+| azure_monitor.cognitive_services.model_provisioned_utilization | utilization | percentage |
111
+| azure_monitor.cognitive_services.model_pages | total, annotated | pages/s |
112
+| azure_monitor.cognitive_services.model_generated_images | generated | images/s |
113
+| azure_monitor.cognitive_services.content_safety_requests | total, harmful, blocked | requests/s |
114
+| azure_monitor.cognitive_services.content_safety_abusive_users | abusive_users | users/s |
115
+| azure_monitor.cognitive_services.content_safety_system_events | events | events |
116
+| azure_monitor.cognitive_services.content_safety_moderation | text, image | requests/s |
117
+| azure_monitor.cognitive_services.job_duration | average | milliseconds |
118
+| azure_monitor.cognitive_services.personalizer_actions | occurrences | occurrences/s |
119
+| azure_monitor.cognitive_services.personalizer_actions_per_event | average | actions |
120
+| azure_monitor.cognitive_services.personalizer_feature_occurrences | action, context, slot | occurrences/s |
121
+| azure_monitor.cognitive_services.personalizer_feature_cardinality | action, context, slot | features |
122
+| azure_monitor.cognitive_services.personalizer_features_per_event | action, context, slot | features |
123
+| azure_monitor.cognitive_services.personalizer_namespaces_per_event | action, context, slot | namespaces |
124
+| azure_monitor.cognitive_services.personalizer_rewards | average, slot | reward |
125
+| azure_monitor.cognitive_services.personalizer_estimator_rewards | online, baseline, baseline_random | reward |
126
+| azure_monitor.cognitive_services.personalizer_estimator_slot_rewards | online, baseline, baseline_random | reward |
127
+| azure_monitor.cognitive_services.personalizer_slots | average | slots |
128
+| azure_monitor.cognitive_services.personalizer_slot_occurrences | occurrences | occurrences/s |
129
+| azure_monitor.cognitive_services.personalizer_event_counts | online, baseline_random, user_baseline | events/s |
130
+| azure_monitor.cognitive_services.personalizer_estimated_rewards | online, baseline_random, user_baseline | reward/s |
131
+| azure_monitor.cognitive_services.speech_transcription | realtime, batch, batch_whisper, fast, fast_whisper | seconds/s |
132
+| azure_monitor.cognitive_services.speech_translation | translated | seconds/s |
133
+| azure_monitor.cognitive_services.speech_synthesis | synthesized | characters/s |
134
+| azure_monitor.cognitive_services.speech_video_synthesis | synthesized | seconds/s |
135
+| azure_monitor.cognitive_services.speech_avatar | hosting, training | seconds/s |
136
+| azure_monitor.cognitive_services.speech_speaker_recognition | transactions | transactions/s |
137
+| azure_monitor.cognitive_services.speech_speaker_profiles | profiles | profiles/s |
138
+| azure_monitor.cognitive_services.speech_model_hosting | hosting | hours/s |
139
+| azure_monitor.cognitive_services.speech_voice_live_tokens | audio_input, audio_output, cached_audio_input, text_input, text_output, cached_text_input | tokens/s |
140
+| azure_monitor.cognitive_services.speech_voice_model | hosting | hours/s |
141
+| azure_monitor.cognitive_services.speech_voice_training | training | minutes/s |
142
+| azure_monitor.cognitive_services.translator_text | standard, custom, trained | characters/s |
143
+| azure_monitor.cognitive_services.translator_document | standard, custom | characters/s |
144
+| azure_monitor.cognitive_services.translator_document_sync | standard, custom | characters/s |
145
+| azure_monitor.cognitive_services.translator_pro_app | seconds | seconds/s |
146
+| azure_monitor.cognitive_services.vision_transactions | computer_vision, custom_vision | transactions/s |
147
+| azure_monitor.cognitive_services.vision_training | training_time | seconds/s |
148
+| azure_monitor.cognitive_services.vision_images_stored | stored | images/s |
149
+| azure_monitor.cognitive_services.face_transactions | transactions | transactions/s |
150
+| azure_monitor.cognitive_services.face_images_trained | trained | images/s |
151
+| azure_monitor.cognitive_services.faces_stored | stored | faces/s |
152
+| azure_monitor.cognitive_services.luis_requests | speech, text | requests/s |
153
+| azure_monitor.cognitive_services.text_processing | text, health_text, question_answering | records/s |
154
+| azure_monitor.cognitive_services.processed_characters | characters | characters/s |
155
+| azure_monitor.cognitive_services.processed_images | images | images/s |
156
+| azure_monitor.cognitive_services.processed_pages | pages | pages/s |
157
+| azure_monitor.cognitive_services.carnegie_inference | inferences | inferences/s |
158
+| azure_monitor.cognitive_services.personalizer_events | total, learned, non_activated | events/s |
159
+| azure_monitor.cognitive_services.personalizer_reward_tracking | matched, observed | rewards/s |
160
+
161
+
162
+
163
+## Alerts
164
+
165
+
166
+The following alerts are available:
167
+
168
+| Alert name | On metric | Description |
169
+|:------------|:----------|:------------|
170
+| [ am_cognitive_services_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.availability | Cognitive Services availability on ${label:resource_name} |
171
+| [ am_cognitive_services_server_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.errors | Cognitive Services server errors on ${label:resource_name} |
172
+| [ am_cognitive_services_client_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.errors | Cognitive Services client errors on ${label:resource_name} |
173
+| [ am_cognitive_services_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.latency | Cognitive Services latency on ${label:resource_name} |
174
+| [ am_cognitive_services_rate_limit ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.rate_limit | Cognitive Services rate limiting on ${label:resource_name} |
175
+| [ am_cognitive_services_blocked_calls ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.calls | Cognitive Services blocked calls on ${label:resource_name} |
176
+| [ am_cognitive_services_openai_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.openai_availability | Azure OpenAI availability on ${label:resource_name} |
177
+| [ am_cognitive_services_openai_time_to_response ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.openai_latency | Azure OpenAI time to response on ${label:resource_name} |
178
+| [ am_cognitive_services_openai_time_to_first_token ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.openai_latency | Azure OpenAI time to first token on ${label:resource_name} |
179
+| [ am_cognitive_services_openai_provisioned_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.openai_provisioned_utilization | Azure OpenAI provisioned utilization on ${label:resource_name} |
180
+| [ am_cognitive_services_model_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.model_availability | Model availability on ${label:resource_name} |
181
+| [ am_cognitive_services_model_time_to_response ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.model_latency | Model time to response on ${label:resource_name} |
182
+| [ am_cognitive_services_model_time_to_first_token ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.model_latency | Model time to first token on ${label:resource_name} |
183
+| [ am_cognitive_services_model_provisioned_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.model_provisioned_utilization | Model provisioned utilization on ${label:resource_name} |
184
+| [ am_cognitive_services_harmful_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.content_safety_requests | Harmful content requests on ${label:resource_name} |
185
+| [ am_cognitive_services_blocked_content_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.content_safety_requests | Blocked content requests on ${label:resource_name} |
186
+| [ am_cognitive_services_abusive_users ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cognitive_services.conf) | azure_monitor.cognitive_services.content_safety_abusive_users | Abusive users detected on ${label:resource_name} |
187
+
188
+
189
+## Setup
190
+
191
+
192
+You can configure the **azure_monitor** collector in two ways:
193
+
194
+| Method | Best for | How to |
195
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
196
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
197
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
198
+
199
+:::important
200
+
201
+UI configuration requires paid Netdata Cloud plan.
202
+
203
+:::
204
+
205
+
206
+### Prerequisites
207
+
208
+#### Create an Azure monitoring principal
209
+
210
+Create a service principal or use a managed identity with the following permissions:
211
+
212
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
213
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
214
+
215
+For service principal authentication:
216
+```bash
217
+# Create the service principal
218
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
219
+ --scopes /subscriptions/<subscription-id>
220
+
221
+# Note the appId (client_id), password (client_secret), and tenant
222
+```
223
+
224
+For managed identity (on Azure VMs, VMSS, or AKS):
225
+```bash
226
+# Assign Monitoring Reader role to the VM's managed identity
227
+az role assignment create --assignee <managed-identity-principal-id> \
228
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
229
+```
230
+
231
+
232
+
233
+### Configuration
234
+
235
+#### Options
236
+
237
+The following options can be defined globally: update_every, autodetection_retry.
238
+
239
+Profile files are loaded from:
240
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
241
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
242
+
243
+User profile files with the same filename override stock profiles.
244
+
245
+
246
+<details open><summary>Config options</summary>
247
+
248
+
249
+
250
+| Group | Option | Description | Default | Required |
251
+|:------|:-----|:------------|:--------|:---------:|
252
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
253
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
254
+| **Target** | subscription_id | Azure subscription ID. | | yes |
255
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
256
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
257
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
258
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
259
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
260
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
261
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
262
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
263
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
264
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
265
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
266
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
267
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
268
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
269
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
270
+
271
+
272
+</details>
273
+
274
+
275
+#### via UI
276
+
277
+Configure the **azure_monitor** collector from the Netdata web interface:
278
+
279
+1. Go to **Nodes**.
280
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
281
+3. The **Collectors → Jobs** view opens by default.
282
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
283
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
284
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
285
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
286
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
287
+
288
+
289
+#### via File
290
+
291
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
292
+
293
+The file format is YAML. Generally, the structure is:
294
+
295
+```yaml
296
+update_every: 1
297
+autodetection_retry: 0
298
+jobs:
299
+ - name: some_name1
300
+ - name: some_name2
301
+```
302
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
303
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
304
+
305
+```bash
306
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
307
+sudo ./edit-config go.d/azure_monitor.conf
308
+```
309
+
310
+##### Examples
311
+
312
+###### Service principal (auto-discover all resources)
313
+
314
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
315
+
316
+```yaml
317
+jobs:
318
+ - name: prod
319
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ auth:
321
+ mode: service_principal
322
+ mode_service_principal:
323
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ client_secret: "your-client-secret"
326
+
327
+```
328
+###### Managed identity (Azure VM/VMSS/AKS)
329
+
330
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
331
+
332
+<details open><summary>Config</summary>
333
+
334
+```yaml
335
+jobs:
336
+ - name: prod
337
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
338
+ auth:
339
+ mode: managed_identity
340
+
341
+```
342
+</details>
343
+
344
+###### Specific profiles only
345
+
346
+Monitor only specific Azure services instead of auto-discovering all resource types.
347
+
348
+<details open><summary>Config</summary>
349
+
350
+```yaml
351
+jobs:
352
+ - name: databases
353
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
+ profiles:
355
+ - sql_database
356
+ - postgres_flexible
357
+ - redis_cache
358
+ auth:
359
+ mode: service_principal
360
+ mode_service_principal:
361
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
362
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
363
+ client_secret: "your-client-secret"
364
+
365
+```
366
+</details>
367
+
368
+###### Filter by resource group
369
+
370
+Only monitor resources in specific resource groups.
371
+
372
+<details open><summary>Config</summary>
373
+
374
+```yaml
375
+jobs:
376
+ - name: prod-rg
377
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ resource_groups:
379
+ - production-rg
380
+ - staging-rg
381
+ auth:
382
+ mode: default
383
+
384
+```
385
+</details>
386
+
387
+###### Azure Government cloud
388
+
389
+Connect to Azure Government cloud environment.
390
+
391
+<details open><summary>Config</summary>
392
+
393
+```yaml
394
+jobs:
395
+ - name: gov
396
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
397
+ cloud: government
398
+ auth:
399
+ mode: service_principal
400
+ mode_service_principal:
401
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
402
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
+ client_secret: "your-client-secret"
404
+
405
+```
406
+</details>
407
+
408
+
409
+
410
+## Troubleshooting
411
+
412
+### Debug Mode
413
+
414
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
415
+
416
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
417
+should give you clues as to why the collector isn't working.
418
+
419
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
420
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
421
+
422
+ ```bash
423
+ cd /usr/libexec/netdata/plugins.d/
424
+ ```
425
+
426
+- Switch to the `netdata` user.
427
+
428
+ ```bash
429
+ sudo -u netdata -s
430
+ ```
431
+
432
+- Run the `go.d.plugin` to debug the collector:
433
+
434
+ ```bash
435
+ ./go.d.plugin -d -m azure_monitor
436
+ ```
437
+
438
+ To debug a specific job:
439
+
440
+ ```bash
441
+ ./go.d.plugin -d -m azure_monitor -j jobName
442
+ ```
443
+
444
+### Getting Logs
445
+
446
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
447
+
448
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
449
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
450
+
451
+#### System with systemd
452
+
453
+Use the following command to view logs generated since the last Netdata service restart:
454
+
455
+```bash
456
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
457
+```
458
+
459
+#### System without systemd
460
+
461
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
462
+
463
+```bash
464
+grep azure_monitor /var/log/netdata/collector.log
465
+```
466
+
467
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
468
+
469
+#### Docker Container
470
+
471
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
472
+
473
+```bash
474
+docker logs netdata 2>&1 | grep azure_monitor
475
+```
476
+
477
+### No metrics are collected
478
+
479
+Verify the following:
480
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
481
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
482
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
483
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
484
+
485
+
486
+### Missing metrics for some resource types
487
+
488
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
489
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
490
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
491
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
492
+
493
+
494
+### Metrics appear delayed
495
+
496
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
497
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
498
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
499
+
500
+
501
+### Authentication errors in sovereign clouds
502
+
503
+For Azure Government or Azure China clouds, set the `cloud` parameter:
504
+- Azure Government: `cloud: government`
505
+- Azure China (21Vianet): `cloud: china`
506
+
507
+Ensure the service principal is registered in the correct cloud tenant.
508
+
509
+
510
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_apps.md
new
+453
@@ -0,0 +1,453 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_apps.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Container Apps"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'container', 'apps', 'serverless', 'microservices']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Container Apps
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Container Apps including CPU and memory usage, network traffic, replica counts, request processing rates, response times, restart frequency, and resource reservation utilization.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.container_apps.cpu_usage | average, maximum | nanocores |
88
+| azure_monitor.container_apps.cpu_percentage | average, maximum | percentage |
89
+| azure_monitor.container_apps.memory_working_set | average, maximum | bytes |
90
+| azure_monitor.container_apps.memory_percentage | average, maximum | percentage |
91
+| azure_monitor.container_apps.replicas | average | replicas |
92
+| azure_monitor.container_apps.reserved_cores | per_revision, total | cores |
93
+| azure_monitor.container_apps.restart_count | restarts | restarts/s |
94
+| azure_monitor.container_apps.requests | requests | requests/s |
95
+| azure_monitor.container_apps.response_time | average | milliseconds |
96
+| azure_monitor.container_apps.network | received, sent | bytes/s |
97
+| azure_monitor.container_apps.resiliency_timeouts | connection, request | timeouts/s |
98
+| azure_monitor.container_apps.resiliency_retries | retries | retries/s |
99
+| azure_monitor.container_apps.resiliency_pending_connections | pending | requests/s |
100
+| azure_monitor.container_apps.resiliency_ejections | ejected, aborted | ejections/s |
101
+| azure_monitor.container_apps.gpu_utilization | average, maximum | percentage |
102
+| azure_monitor.container_apps.jvm_buffer_count | average | buffers |
103
+| azure_monitor.container_apps.jvm_buffer_memory | used, limit | bytes |
104
+| azure_monitor.container_apps.jvm_gc_count | collections | collections/s |
105
+| azure_monitor.container_apps.jvm_gc_duration | duration | milliseconds/s |
106
+| azure_monitor.container_apps.jvm_memory_pool | used, committed, limit | bytes |
107
+| azure_monitor.container_apps.jvm_memory_total | used, committed, limit | bytes |
108
+| azure_monitor.container_apps.jvm_thread_count | average | threads |
109
+
110
+
111
+
112
+## Alerts
113
+
114
+
115
+The following alerts are available:
116
+
117
+| Alert name | On metric | Description |
118
+|:------------|:----------|:------------|
119
+| [ am_container_apps_cpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.cpu_percentage | Container Apps CPU on ${label:resource_name} |
120
+| [ am_container_apps_memory_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.memory_percentage | Container Apps memory on ${label:resource_name} |
121
+| [ am_container_apps_restarts ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.restart_count | Container Apps restarts on ${label:resource_name} |
122
+| [ am_container_apps_response_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.response_time | Container Apps response time on ${label:resource_name} |
123
+| [ am_container_apps_resiliency_timeouts ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.resiliency_timeouts | Container Apps resiliency timeouts on ${label:resource_name} |
124
+| [ am_container_apps_resiliency_retries ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.resiliency_retries | Container Apps resiliency retries on ${label:resource_name} |
125
+| [ am_container_apps_pending_connections ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.resiliency_pending_connections | Container Apps pending connections on ${label:resource_name} |
126
+| [ am_container_apps_host_ejections ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.resiliency_ejections | Container Apps host ejections on ${label:resource_name} |
127
+| [ am_container_apps_replica_count ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.replicas | Container Apps replicas on ${label:resource_name} |
128
+| [ am_container_apps_gpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.gpu_utilization | Container Apps GPU on ${label:resource_name} |
129
+| [ am_container_apps_jvm_gc_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_apps.conf) | azure_monitor.container_apps.jvm_gc_duration | Container Apps JVM GC duration on ${label:resource_name} |
130
+
131
+
132
+## Setup
133
+
134
+
135
+You can configure the **azure_monitor** collector in two ways:
136
+
137
+| Method | Best for | How to |
138
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
139
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
140
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
141
+
142
+:::important
143
+
144
+UI configuration requires paid Netdata Cloud plan.
145
+
146
+:::
147
+
148
+
149
+### Prerequisites
150
+
151
+#### Create an Azure monitoring principal
152
+
153
+Create a service principal or use a managed identity with the following permissions:
154
+
155
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
156
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
157
+
158
+For service principal authentication:
159
+```bash
160
+# Create the service principal
161
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
162
+ --scopes /subscriptions/<subscription-id>
163
+
164
+# Note the appId (client_id), password (client_secret), and tenant
165
+```
166
+
167
+For managed identity (on Azure VMs, VMSS, or AKS):
168
+```bash
169
+# Assign Monitoring Reader role to the VM's managed identity
170
+az role assignment create --assignee <managed-identity-principal-id> \
171
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
172
+```
173
+
174
+
175
+
176
+### Configuration
177
+
178
+#### Options
179
+
180
+The following options can be defined globally: update_every, autodetection_retry.
181
+
182
+Profile files are loaded from:
183
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
184
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
185
+
186
+User profile files with the same filename override stock profiles.
187
+
188
+
189
+<details open><summary>Config options</summary>
190
+
191
+
192
+
193
+| Group | Option | Description | Default | Required |
194
+|:------|:-----|:------------|:--------|:---------:|
195
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
196
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
197
+| **Target** | subscription_id | Azure subscription ID. | | yes |
198
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
199
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
200
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
201
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
202
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
203
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
204
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
205
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
206
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
207
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
208
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
209
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
210
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
211
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
212
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
213
+
214
+
215
+</details>
216
+
217
+
218
+#### via UI
219
+
220
+Configure the **azure_monitor** collector from the Netdata web interface:
221
+
222
+1. Go to **Nodes**.
223
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
224
+3. The **Collectors → Jobs** view opens by default.
225
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
226
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
227
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
228
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
229
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
230
+
231
+
232
+#### via File
233
+
234
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
235
+
236
+The file format is YAML. Generally, the structure is:
237
+
238
+```yaml
239
+update_every: 1
240
+autodetection_retry: 0
241
+jobs:
242
+ - name: some_name1
243
+ - name: some_name2
244
+```
245
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
246
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
247
+
248
+```bash
249
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
250
+sudo ./edit-config go.d/azure_monitor.conf
251
+```
252
+
253
+##### Examples
254
+
255
+###### Service principal (auto-discover all resources)
256
+
257
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
258
+
259
+```yaml
260
+jobs:
261
+ - name: prod
262
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
263
+ auth:
264
+ mode: service_principal
265
+ mode_service_principal:
266
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
267
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
268
+ client_secret: "your-client-secret"
269
+
270
+```
271
+###### Managed identity (Azure VM/VMSS/AKS)
272
+
273
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
274
+
275
+<details open><summary>Config</summary>
276
+
277
+```yaml
278
+jobs:
279
+ - name: prod
280
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
281
+ auth:
282
+ mode: managed_identity
283
+
284
+```
285
+</details>
286
+
287
+###### Specific profiles only
288
+
289
+Monitor only specific Azure services instead of auto-discovering all resource types.
290
+
291
+<details open><summary>Config</summary>
292
+
293
+```yaml
294
+jobs:
295
+ - name: databases
296
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
297
+ profiles:
298
+ - sql_database
299
+ - postgres_flexible
300
+ - redis_cache
301
+ auth:
302
+ mode: service_principal
303
+ mode_service_principal:
304
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
305
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
306
+ client_secret: "your-client-secret"
307
+
308
+```
309
+</details>
310
+
311
+###### Filter by resource group
312
+
313
+Only monitor resources in specific resource groups.
314
+
315
+<details open><summary>Config</summary>
316
+
317
+```yaml
318
+jobs:
319
+ - name: prod-rg
320
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ resource_groups:
322
+ - production-rg
323
+ - staging-rg
324
+ auth:
325
+ mode: default
326
+
327
+```
328
+</details>
329
+
330
+###### Azure Government cloud
331
+
332
+Connect to Azure Government cloud environment.
333
+
334
+<details open><summary>Config</summary>
335
+
336
+```yaml
337
+jobs:
338
+ - name: gov
339
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
340
+ cloud: government
341
+ auth:
342
+ mode: service_principal
343
+ mode_service_principal:
344
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
345
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
346
+ client_secret: "your-client-secret"
347
+
348
+```
349
+</details>
350
+
351
+
352
+
353
+## Troubleshooting
354
+
355
+### Debug Mode
356
+
357
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
358
+
359
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
360
+should give you clues as to why the collector isn't working.
361
+
362
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
363
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
364
+
365
+ ```bash
366
+ cd /usr/libexec/netdata/plugins.d/
367
+ ```
368
+
369
+- Switch to the `netdata` user.
370
+
371
+ ```bash
372
+ sudo -u netdata -s
373
+ ```
374
+
375
+- Run the `go.d.plugin` to debug the collector:
376
+
377
+ ```bash
378
+ ./go.d.plugin -d -m azure_monitor
379
+ ```
380
+
381
+ To debug a specific job:
382
+
383
+ ```bash
384
+ ./go.d.plugin -d -m azure_monitor -j jobName
385
+ ```
386
+
387
+### Getting Logs
388
+
389
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
390
+
391
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
392
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
393
+
394
+#### System with systemd
395
+
396
+Use the following command to view logs generated since the last Netdata service restart:
397
+
398
+```bash
399
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
400
+```
401
+
402
+#### System without systemd
403
+
404
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
405
+
406
+```bash
407
+grep azure_monitor /var/log/netdata/collector.log
408
+```
409
+
410
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
411
+
412
+#### Docker Container
413
+
414
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
415
+
416
+```bash
417
+docker logs netdata 2>&1 | grep azure_monitor
418
+```
419
+
420
+### No metrics are collected
421
+
422
+Verify the following:
423
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
424
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
425
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
426
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
427
+
428
+
429
+### Missing metrics for some resource types
430
+
431
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
432
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
433
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
434
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
435
+
436
+
437
+### Metrics appear delayed
438
+
439
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
440
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
441
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
442
+
443
+
444
+### Authentication errors in sovereign clouds
445
+
446
+For Azure Government or Azure China clouds, set the `cloud` parameter:
447
+- Azure Government: `cloud: government`
448
+- Azure China (21Vianet): `cloud: china`
449
+
450
+Ensure the service principal is registered in the correct cloud tenant.
451
+
452
+
453
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_instances.md
new
+427
@@ -0,0 +1,427 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_instances.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Container Instances"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'container', 'instances', 'aci', 'docker']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Container Instances
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Container Instance groups including CPU and memory usage and network bytes transferred in and out.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.container_instances.cpu_usage | average, maximum | millicores |
88
+| azure_monitor.container_instances.memory_usage | average, maximum | bytes |
89
+| azure_monitor.container_instances.network | received, sent | bytes/s |
90
+
91
+
92
+
93
+## Alerts
94
+
95
+
96
+The following alerts are available:
97
+
98
+| Alert name | On metric | Description |
99
+|:------------|:----------|:------------|
100
+| [ am_container_instances_cpu_usage ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_instances.conf) | azure_monitor.container_instances.cpu_usage | Container CPU on ${label:resource_name} |
101
+| [ am_container_instances_memory_usage ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_instances.conf) | azure_monitor.container_instances.memory_usage | Container memory on ${label:resource_name} |
102
+| [ am_container_instances_network_rx ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_instances.conf) | azure_monitor.container_instances.network | Container inbound traffic on ${label:resource_name} |
103
+| [ am_container_instances_network_tx ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_instances.conf) | azure_monitor.container_instances.network | Container outbound traffic on ${label:resource_name} |
104
+
105
+
106
+## Setup
107
+
108
+
109
+You can configure the **azure_monitor** collector in two ways:
110
+
111
+| Method | Best for | How to |
112
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
113
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
114
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
115
+
116
+:::important
117
+
118
+UI configuration requires paid Netdata Cloud plan.
119
+
120
+:::
121
+
122
+
123
+### Prerequisites
124
+
125
+#### Create an Azure monitoring principal
126
+
127
+Create a service principal or use a managed identity with the following permissions:
128
+
129
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
130
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
131
+
132
+For service principal authentication:
133
+```bash
134
+# Create the service principal
135
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
136
+ --scopes /subscriptions/<subscription-id>
137
+
138
+# Note the appId (client_id), password (client_secret), and tenant
139
+```
140
+
141
+For managed identity (on Azure VMs, VMSS, or AKS):
142
+```bash
143
+# Assign Monitoring Reader role to the VM's managed identity
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+```
147
+
148
+
149
+
150
+### Configuration
151
+
152
+#### Options
153
+
154
+The following options can be defined globally: update_every, autodetection_retry.
155
+
156
+Profile files are loaded from:
157
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
158
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+
160
+User profile files with the same filename override stock profiles.
161
+
162
+
163
+<details open><summary>Config options</summary>
164
+
165
+
166
+
167
+| Group | Option | Description | Default | Required |
168
+|:------|:-----|:------------|:--------|:---------:|
169
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
170
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
171
+| **Target** | subscription_id | Azure subscription ID. | | yes |
172
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
173
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
174
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
175
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
176
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
177
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
178
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
179
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
180
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
181
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
187
+
188
+
189
+</details>
190
+
191
+
192
+#### via UI
193
+
194
+Configure the **azure_monitor** collector from the Netdata web interface:
195
+
196
+1. Go to **Nodes**.
197
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
198
+3. The **Collectors → Jobs** view opens by default.
199
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
200
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
201
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
202
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
203
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
204
+
205
+
206
+#### via File
207
+
208
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
209
+
210
+The file format is YAML. Generally, the structure is:
211
+
212
+```yaml
213
+update_every: 1
214
+autodetection_retry: 0
215
+jobs:
216
+ - name: some_name1
217
+ - name: some_name2
218
+```
219
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
220
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
221
+
222
+```bash
223
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
224
+sudo ./edit-config go.d/azure_monitor.conf
225
+```
226
+
227
+##### Examples
228
+
229
+###### Service principal (auto-discover all resources)
230
+
231
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
232
+
233
+```yaml
234
+jobs:
235
+ - name: prod
236
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
237
+ auth:
238
+ mode: service_principal
239
+ mode_service_principal:
240
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
241
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ client_secret: "your-client-secret"
243
+
244
+```
245
+###### Managed identity (Azure VM/VMSS/AKS)
246
+
247
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
248
+
249
+<details open><summary>Config</summary>
250
+
251
+```yaml
252
+jobs:
253
+ - name: prod
254
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
255
+ auth:
256
+ mode: managed_identity
257
+
258
+```
259
+</details>
260
+
261
+###### Specific profiles only
262
+
263
+Monitor only specific Azure services instead of auto-discovering all resource types.
264
+
265
+<details open><summary>Config</summary>
266
+
267
+```yaml
268
+jobs:
269
+ - name: databases
270
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
271
+ profiles:
272
+ - sql_database
273
+ - postgres_flexible
274
+ - redis_cache
275
+ auth:
276
+ mode: service_principal
277
+ mode_service_principal:
278
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
279
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
280
+ client_secret: "your-client-secret"
281
+
282
+```
283
+</details>
284
+
285
+###### Filter by resource group
286
+
287
+Only monitor resources in specific resource groups.
288
+
289
+<details open><summary>Config</summary>
290
+
291
+```yaml
292
+jobs:
293
+ - name: prod-rg
294
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ resource_groups:
296
+ - production-rg
297
+ - staging-rg
298
+ auth:
299
+ mode: default
300
+
301
+```
302
+</details>
303
+
304
+###### Azure Government cloud
305
+
306
+Connect to Azure Government cloud environment.
307
+
308
+<details open><summary>Config</summary>
309
+
310
+```yaml
311
+jobs:
312
+ - name: gov
313
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
314
+ cloud: government
315
+ auth:
316
+ mode: service_principal
317
+ mode_service_principal:
318
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
319
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ client_secret: "your-client-secret"
321
+
322
+```
323
+</details>
324
+
325
+
326
+
327
+## Troubleshooting
328
+
329
+### Debug Mode
330
+
331
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
332
+
333
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
334
+should give you clues as to why the collector isn't working.
335
+
336
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
337
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
338
+
339
+ ```bash
340
+ cd /usr/libexec/netdata/plugins.d/
341
+ ```
342
+
343
+- Switch to the `netdata` user.
344
+
345
+ ```bash
346
+ sudo -u netdata -s
347
+ ```
348
+
349
+- Run the `go.d.plugin` to debug the collector:
350
+
351
+ ```bash
352
+ ./go.d.plugin -d -m azure_monitor
353
+ ```
354
+
355
+ To debug a specific job:
356
+
357
+ ```bash
358
+ ./go.d.plugin -d -m azure_monitor -j jobName
359
+ ```
360
+
361
+### Getting Logs
362
+
363
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
364
+
365
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
366
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
367
+
368
+#### System with systemd
369
+
370
+Use the following command to view logs generated since the last Netdata service restart:
371
+
372
+```bash
373
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
374
+```
375
+
376
+#### System without systemd
377
+
378
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
379
+
380
+```bash
381
+grep azure_monitor /var/log/netdata/collector.log
382
+```
383
+
384
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
385
+
386
+#### Docker Container
387
+
388
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
389
+
390
+```bash
391
+docker logs netdata 2>&1 | grep azure_monitor
392
+```
393
+
394
+### No metrics are collected
395
+
396
+Verify the following:
397
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
398
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
399
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
400
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
401
+
402
+
403
+### Missing metrics for some resource types
404
+
405
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
406
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
407
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
408
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
409
+
410
+
411
+### Metrics appear delayed
412
+
413
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
414
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
415
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
416
+
417
+
418
+### Authentication errors in sovereign clouds
419
+
420
+For Azure Government or Azure China clouds, set the `cloud` parameter:
421
+- Azure Government: `cloud: government`
422
+- Azure China (21Vianet): `cloud: china`
423
+
424
+Ensure the service principal is registered in the correct cloud tenant.
425
+
426
+
427
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_registry.md
new
+427
@@ -0,0 +1,427 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_registry.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Container Registry"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'container', 'registry', 'acr', 'docker', 'images']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Container Registry
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Container Registry including storage usage, successful and failed pull and push operation counts, and task run duration.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.container_registry.storage_used | used | bytes |
88
+| azure_monitor.container_registry.pull_count | successful, total | pulls/s |
89
+| azure_monitor.container_registry.push_count | successful, total | pushes/s |
90
+| azure_monitor.container_registry.agentpool_cpu_time | cpu_time | seconds/s |
91
+| azure_monitor.container_registry.run_duration | duration | milliseconds/s |
92
+
93
+
94
+
95
+## Alerts
96
+
97
+
98
+The following alerts are available:
99
+
100
+| Alert name | On metric | Description |
101
+|:------------|:----------|:------------|
102
+| [ am_container_registry_pull_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_registry.conf) | azure_monitor.container_registry.pull_count | ACR pull failures on ${label:resource_name} |
103
+| [ am_container_registry_push_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_container_registry.conf) | azure_monitor.container_registry.push_count | ACR push failures on ${label:resource_name} |
104
+
105
+
106
+## Setup
107
+
108
+
109
+You can configure the **azure_monitor** collector in two ways:
110
+
111
+| Method | Best for | How to |
112
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
113
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
114
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
115
+
116
+:::important
117
+
118
+UI configuration requires paid Netdata Cloud plan.
119
+
120
+:::
121
+
122
+
123
+### Prerequisites
124
+
125
+#### Create an Azure monitoring principal
126
+
127
+Create a service principal or use a managed identity with the following permissions:
128
+
129
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
130
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
131
+
132
+For service principal authentication:
133
+```bash
134
+# Create the service principal
135
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
136
+ --scopes /subscriptions/<subscription-id>
137
+
138
+# Note the appId (client_id), password (client_secret), and tenant
139
+```
140
+
141
+For managed identity (on Azure VMs, VMSS, or AKS):
142
+```bash
143
+# Assign Monitoring Reader role to the VM's managed identity
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+```
147
+
148
+
149
+
150
+### Configuration
151
+
152
+#### Options
153
+
154
+The following options can be defined globally: update_every, autodetection_retry.
155
+
156
+Profile files are loaded from:
157
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
158
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+
160
+User profile files with the same filename override stock profiles.
161
+
162
+
163
+<details open><summary>Config options</summary>
164
+
165
+
166
+
167
+| Group | Option | Description | Default | Required |
168
+|:------|:-----|:------------|:--------|:---------:|
169
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
170
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
171
+| **Target** | subscription_id | Azure subscription ID. | | yes |
172
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
173
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
174
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
175
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
176
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
177
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
178
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
179
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
180
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
181
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
187
+
188
+
189
+</details>
190
+
191
+
192
+#### via UI
193
+
194
+Configure the **azure_monitor** collector from the Netdata web interface:
195
+
196
+1. Go to **Nodes**.
197
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
198
+3. The **Collectors → Jobs** view opens by default.
199
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
200
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
201
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
202
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
203
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
204
+
205
+
206
+#### via File
207
+
208
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
209
+
210
+The file format is YAML. Generally, the structure is:
211
+
212
+```yaml
213
+update_every: 1
214
+autodetection_retry: 0
215
+jobs:
216
+ - name: some_name1
217
+ - name: some_name2
218
+```
219
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
220
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
221
+
222
+```bash
223
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
224
+sudo ./edit-config go.d/azure_monitor.conf
225
+```
226
+
227
+##### Examples
228
+
229
+###### Service principal (auto-discover all resources)
230
+
231
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
232
+
233
+```yaml
234
+jobs:
235
+ - name: prod
236
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
237
+ auth:
238
+ mode: service_principal
239
+ mode_service_principal:
240
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
241
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ client_secret: "your-client-secret"
243
+
244
+```
245
+###### Managed identity (Azure VM/VMSS/AKS)
246
+
247
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
248
+
249
+<details open><summary>Config</summary>
250
+
251
+```yaml
252
+jobs:
253
+ - name: prod
254
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
255
+ auth:
256
+ mode: managed_identity
257
+
258
+```
259
+</details>
260
+
261
+###### Specific profiles only
262
+
263
+Monitor only specific Azure services instead of auto-discovering all resource types.
264
+
265
+<details open><summary>Config</summary>
266
+
267
+```yaml
268
+jobs:
269
+ - name: databases
270
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
271
+ profiles:
272
+ - sql_database
273
+ - postgres_flexible
274
+ - redis_cache
275
+ auth:
276
+ mode: service_principal
277
+ mode_service_principal:
278
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
279
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
280
+ client_secret: "your-client-secret"
281
+
282
+```
283
+</details>
284
+
285
+###### Filter by resource group
286
+
287
+Only monitor resources in specific resource groups.
288
+
289
+<details open><summary>Config</summary>
290
+
291
+```yaml
292
+jobs:
293
+ - name: prod-rg
294
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ resource_groups:
296
+ - production-rg
297
+ - staging-rg
298
+ auth:
299
+ mode: default
300
+
301
+```
302
+</details>
303
+
304
+###### Azure Government cloud
305
+
306
+Connect to Azure Government cloud environment.
307
+
308
+<details open><summary>Config</summary>
309
+
310
+```yaml
311
+jobs:
312
+ - name: gov
313
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
314
+ cloud: government
315
+ auth:
316
+ mode: service_principal
317
+ mode_service_principal:
318
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
319
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ client_secret: "your-client-secret"
321
+
322
+```
323
+</details>
324
+
325
+
326
+
327
+## Troubleshooting
328
+
329
+### Debug Mode
330
+
331
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
332
+
333
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
334
+should give you clues as to why the collector isn't working.
335
+
336
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
337
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
338
+
339
+ ```bash
340
+ cd /usr/libexec/netdata/plugins.d/
341
+ ```
342
+
343
+- Switch to the `netdata` user.
344
+
345
+ ```bash
346
+ sudo -u netdata -s
347
+ ```
348
+
349
+- Run the `go.d.plugin` to debug the collector:
350
+
351
+ ```bash
352
+ ./go.d.plugin -d -m azure_monitor
353
+ ```
354
+
355
+ To debug a specific job:
356
+
357
+ ```bash
358
+ ./go.d.plugin -d -m azure_monitor -j jobName
359
+ ```
360
+
361
+### Getting Logs
362
+
363
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
364
+
365
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
366
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
367
+
368
+#### System with systemd
369
+
370
+Use the following command to view logs generated since the last Netdata service restart:
371
+
372
+```bash
373
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
374
+```
375
+
376
+#### System without systemd
377
+
378
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
379
+
380
+```bash
381
+grep azure_monitor /var/log/netdata/collector.log
382
+```
383
+
384
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
385
+
386
+#### Docker Container
387
+
388
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
389
+
390
+```bash
391
+docker logs netdata 2>&1 | grep azure_monitor
392
+```
393
+
394
+### No metrics are collected
395
+
396
+Verify the following:
397
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
398
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
399
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
400
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
401
+
402
+
403
+### Missing metrics for some resource types
404
+
405
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
406
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
407
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
408
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
409
+
410
+
411
+### Metrics appear delayed
412
+
413
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
414
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
415
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
416
+
417
+
418
+### Authentication errors in sovereign clouds
419
+
420
+For Azure Government or Azure China clouds, set the `cloud` parameter:
421
+- Azure Government: `cloud: government`
422
+- Azure China (21Vianet): `cloud: china`
423
+
424
+Ensure the service principal is registered in the correct cloud tenant.
425
+
426
+
427
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cosmos_db_account.md
new
+450
@@ -0,0 +1,450 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cosmos_db_account.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Cosmos DB Account"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'cosmos', 'cosmosdb', 'nosql', 'database', 'documentdb']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Cosmos DB Account
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Cosmos DB accounts including request unit consumption and throttling, document counts and storage, data and index sizes, replication latency, availability percentages, provisioned throughput utilization, and normalized RU consumption per partition.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.cosmos_db.total_requests | count | requests/s |
88
+| azure_monitor.cosmos_db.request_units | total | RU/s |
89
+| azure_monitor.cosmos_db.normalized_ru_consumption | maximum | percentage |
90
+| azure_monitor.cosmos_db.server_side_latency | direct, gateway | milliseconds |
91
+| azure_monitor.cosmos_db.replication_latency | average | milliseconds |
92
+| azure_monitor.cosmos_db.storage | data, index, quota | bytes |
93
+| azure_monitor.cosmos_db.document_count | average | documents |
94
+| azure_monitor.cosmos_db.partition_size | maximum | bytes |
95
+| azure_monitor.cosmos_db.partition_count | maximum | partitions |
96
+| azure_monitor.cosmos_db.provisioned_throughput | provisioned, autoscale_max, autoscaled | RU/s |
97
+| azure_monitor.cosmos_db.partition_throughput | maximum | RU/s |
98
+| azure_monitor.cosmos_db.availability | average | percentage |
99
+| azure_monitor.cosmos_db.metadata_requests | count | requests/s |
100
+| azure_monitor.cosmos_db.api_requests | mongo, cassandra, gremlin | requests/s |
101
+| azure_monitor.cosmos_db.api_request_charges | mongo, cassandra, gremlin | RU/s |
102
+| azure_monitor.cosmos_db.cassandra_connections | total | connections/s |
103
+| azure_monitor.cosmos_db.dedicated_gateway_cpu | average | percentage |
104
+| azure_monitor.cosmos_db.dedicated_gateway_memory | average | bytes |
105
+| azure_monitor.cosmos_db.dedicated_gateway_requests | count | requests/s |
106
+| azure_monitor.cosmos_db.integrated_cache_hit_rate | item, query | percentage |
107
+
108
+
109
+
110
+## Alerts
111
+
112
+
113
+The following alerts are available:
114
+
115
+| Alert name | On metric | Description |
116
+|:------------|:----------|:------------|
117
+| [ am_cosmos_db_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.availability | Cosmos DB availability on ${label:resource_name} |
118
+| [ am_cosmos_db_normalized_ru_consumption ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.normalized_ru_consumption | Cosmos DB RU consumption on ${label:resource_name} |
119
+| [ am_cosmos_db_storage_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.storage | Cosmos DB storage utilization on ${label:resource_name} |
120
+| [ am_cosmos_db_server_side_latency_direct ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.server_side_latency | Cosmos DB direct latency on ${label:resource_name} |
121
+| [ am_cosmos_db_server_side_latency_gateway ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.server_side_latency | Cosmos DB gateway latency on ${label:resource_name} |
122
+| [ am_cosmos_db_replication_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.replication_latency | Cosmos DB replication latency on ${label:resource_name} |
123
+| [ am_cosmos_db_dedicated_gateway_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.dedicated_gateway_cpu | Cosmos DB dedicated gateway CPU on ${label:resource_name} |
124
+| [ am_cosmos_db_integrated_cache_item_hit_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.integrated_cache_hit_rate | Cosmos DB cache item hit rate on ${label:resource_name} |
125
+| [ am_cosmos_db_integrated_cache_query_hit_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.integrated_cache_hit_rate | Cosmos DB cache query hit rate on ${label:resource_name} |
126
+| [ am_cosmos_db_cassandra_connection_closures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_cosmos_db.conf) | azure_monitor.cosmos_db.cassandra_connections | Cosmos DB Cassandra connection closures on ${label:resource_name} |
127
+
128
+
129
+## Setup
130
+
131
+
132
+You can configure the **azure_monitor** collector in two ways:
133
+
134
+| Method | Best for | How to |
135
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
136
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
137
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
138
+
139
+:::important
140
+
141
+UI configuration requires paid Netdata Cloud plan.
142
+
143
+:::
144
+
145
+
146
+### Prerequisites
147
+
148
+#### Create an Azure monitoring principal
149
+
150
+Create a service principal or use a managed identity with the following permissions:
151
+
152
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
153
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
154
+
155
+For service principal authentication:
156
+```bash
157
+# Create the service principal
158
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
159
+ --scopes /subscriptions/<subscription-id>
160
+
161
+# Note the appId (client_id), password (client_secret), and tenant
162
+```
163
+
164
+For managed identity (on Azure VMs, VMSS, or AKS):
165
+```bash
166
+# Assign Monitoring Reader role to the VM's managed identity
167
+az role assignment create --assignee <managed-identity-principal-id> \
168
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
169
+```
170
+
171
+
172
+
173
+### Configuration
174
+
175
+#### Options
176
+
177
+The following options can be defined globally: update_every, autodetection_retry.
178
+
179
+Profile files are loaded from:
180
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
181
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
182
+
183
+User profile files with the same filename override stock profiles.
184
+
185
+
186
+<details open><summary>Config options</summary>
187
+
188
+
189
+
190
+| Group | Option | Description | Default | Required |
191
+|:------|:-----|:------------|:--------|:---------:|
192
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
193
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
194
+| **Target** | subscription_id | Azure subscription ID. | | yes |
195
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
196
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
197
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
198
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
199
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
203
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
204
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
205
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
206
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
207
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
208
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
209
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
210
+
211
+
212
+</details>
213
+
214
+
215
+#### via UI
216
+
217
+Configure the **azure_monitor** collector from the Netdata web interface:
218
+
219
+1. Go to **Nodes**.
220
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
221
+3. The **Collectors → Jobs** view opens by default.
222
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
223
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
224
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
225
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
226
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
227
+
228
+
229
+#### via File
230
+
231
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
232
+
233
+The file format is YAML. Generally, the structure is:
234
+
235
+```yaml
236
+update_every: 1
237
+autodetection_retry: 0
238
+jobs:
239
+ - name: some_name1
240
+ - name: some_name2
241
+```
242
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
243
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
244
+
245
+```bash
246
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
247
+sudo ./edit-config go.d/azure_monitor.conf
248
+```
249
+
250
+##### Examples
251
+
252
+###### Service principal (auto-discover all resources)
253
+
254
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
255
+
256
+```yaml
257
+jobs:
258
+ - name: prod
259
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
260
+ auth:
261
+ mode: service_principal
262
+ mode_service_principal:
263
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
264
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
265
+ client_secret: "your-client-secret"
266
+
267
+```
268
+###### Managed identity (Azure VM/VMSS/AKS)
269
+
270
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
271
+
272
+<details open><summary>Config</summary>
273
+
274
+```yaml
275
+jobs:
276
+ - name: prod
277
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
278
+ auth:
279
+ mode: managed_identity
280
+
281
+```
282
+</details>
283
+
284
+###### Specific profiles only
285
+
286
+Monitor only specific Azure services instead of auto-discovering all resource types.
287
+
288
+<details open><summary>Config</summary>
289
+
290
+```yaml
291
+jobs:
292
+ - name: databases
293
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
294
+ profiles:
295
+ - sql_database
296
+ - postgres_flexible
297
+ - redis_cache
298
+ auth:
299
+ mode: service_principal
300
+ mode_service_principal:
301
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
302
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
303
+ client_secret: "your-client-secret"
304
+
305
+```
306
+</details>
307
+
308
+###### Filter by resource group
309
+
310
+Only monitor resources in specific resource groups.
311
+
312
+<details open><summary>Config</summary>
313
+
314
+```yaml
315
+jobs:
316
+ - name: prod-rg
317
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
318
+ resource_groups:
319
+ - production-rg
320
+ - staging-rg
321
+ auth:
322
+ mode: default
323
+
324
+```
325
+</details>
326
+
327
+###### Azure Government cloud
328
+
329
+Connect to Azure Government cloud environment.
330
+
331
+<details open><summary>Config</summary>
332
+
333
+```yaml
334
+jobs:
335
+ - name: gov
336
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
337
+ cloud: government
338
+ auth:
339
+ mode: service_principal
340
+ mode_service_principal:
341
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
342
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
343
+ client_secret: "your-client-secret"
344
+
345
+```
346
+</details>
347
+
348
+
349
+
350
+## Troubleshooting
351
+
352
+### Debug Mode
353
+
354
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
355
+
356
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
357
+should give you clues as to why the collector isn't working.
358
+
359
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
360
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
361
+
362
+ ```bash
363
+ cd /usr/libexec/netdata/plugins.d/
364
+ ```
365
+
366
+- Switch to the `netdata` user.
367
+
368
+ ```bash
369
+ sudo -u netdata -s
370
+ ```
371
+
372
+- Run the `go.d.plugin` to debug the collector:
373
+
374
+ ```bash
375
+ ./go.d.plugin -d -m azure_monitor
376
+ ```
377
+
378
+ To debug a specific job:
379
+
380
+ ```bash
381
+ ./go.d.plugin -d -m azure_monitor -j jobName
382
+ ```
383
+
384
+### Getting Logs
385
+
386
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
387
+
388
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
389
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
390
+
391
+#### System with systemd
392
+
393
+Use the following command to view logs generated since the last Netdata service restart:
394
+
395
+```bash
396
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
397
+```
398
+
399
+#### System without systemd
400
+
401
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
402
+
403
+```bash
404
+grep azure_monitor /var/log/netdata/collector.log
405
+```
406
+
407
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
408
+
409
+#### Docker Container
410
+
411
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
412
+
413
+```bash
414
+docker logs netdata 2>&1 | grep azure_monitor
415
+```
416
+
417
+### No metrics are collected
418
+
419
+Verify the following:
420
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
421
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
422
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
423
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
424
+
425
+
426
+### Missing metrics for some resource types
427
+
428
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
429
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
430
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
431
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
432
+
433
+
434
+### Metrics appear delayed
435
+
436
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
437
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
438
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
439
+
440
+
441
+### Authentication errors in sovereign clouds
442
+
443
+For Azure Government or Azure China clouds, set the `cloud` parameter:
444
+- Azure Government: `cloud: government`
445
+- Azure China (21Vianet): `cloud: china`
446
+
447
+Ensure the service principal is registered in the correct cloud tenant.
448
+
449
+
450
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_explorer_cluster.md
new
+485
@@ -0,0 +1,485 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_explorer_cluster.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Data Explorer Cluster"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'data', 'explorer', 'kusto', 'adx', 'analytics']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Data Explorer Cluster
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Data Explorer (Kusto) clusters including ingestion latency, volume, and success rates, query performance and concurrency, cache utilization, CPU and memory usage, export operations, streaming ingest throughput, materialized view health, instance counts, and follower lag.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.data_explorer.cpu_utilization | average | percentage |
88
+| azure_monitor.data_explorer.utilization | ingestion, cache | percentage |
89
+| azure_monitor.data_explorer.keep_alive | average | count |
90
+| azure_monitor.data_explorer.instance_count | average, maximum, minimum | instances |
91
+| azure_monitor.data_explorer.extents | average | extents |
92
+| azure_monitor.data_explorer.throttled_commands | total | commands/s |
93
+| azure_monitor.data_explorer.follower_latency | average | milliseconds |
94
+| azure_monitor.data_explorer.query_duration | average | milliseconds |
95
+| azure_monitor.data_explorer.query_count | count | queries/s |
96
+| azure_monitor.data_explorer.concurrent_queries | average, maximum | queries |
97
+| azure_monitor.data_explorer.throttled_queries | total | queries/s |
98
+| azure_monitor.data_explorer.weak_consistency_latency | average | seconds |
99
+| azure_monitor.data_explorer.ingestion_result | total | sources/s |
100
+| azure_monitor.data_explorer.ingestion_volume | total | bytes/s |
101
+| azure_monitor.data_explorer.ingestion_latency | average | seconds |
102
+| azure_monitor.data_explorer.events | received, processed, dropped | events/s |
103
+| azure_monitor.data_explorer.blobs | received, processed, dropped | blobs/s |
104
+| azure_monitor.data_explorer.discovery_latency | average | seconds |
105
+| azure_monitor.data_explorer.stage_latency | average | seconds |
106
+| azure_monitor.data_explorer.ingestion_queue | length | messages |
107
+| azure_monitor.data_explorer.queue_oldest_message | age | seconds |
108
+| azure_monitor.data_explorer.received_data_size | total | bytes/s |
109
+| azure_monitor.data_explorer.batches_processed | total | batches/s |
110
+| azure_monitor.data_explorer.batch_blob_count | average | blobs |
111
+| azure_monitor.data_explorer.batch_size | average | bytes |
112
+| azure_monitor.data_explorer.batch_duration | average | seconds |
113
+| azure_monitor.data_explorer.export_utilization | maximum | percentage |
114
+| azure_monitor.data_explorer.continuous_export_records | total | records/s |
115
+| azure_monitor.data_explorer.continuous_export_result | count | results/s |
116
+| azure_monitor.data_explorer.continuous_export_pending | maximum | jobs |
117
+| azure_monitor.data_explorer.continuous_export_lateness | maximum | minutes |
118
+| azure_monitor.data_explorer.streaming_ingest_data_rate | average | bytes/s |
119
+| azure_monitor.data_explorer.streaming_ingest_duration | average | milliseconds |
120
+| azure_monitor.data_explorer.streaming_ingest_result | count | results/s |
121
+| azure_monitor.data_explorer.streaming_ingest_utilization | average | percentage |
122
+| azure_monitor.data_explorer.materialized_view_health | health | status |
123
+| azure_monitor.data_explorer.materialized_view_age | minutes | minutes |
124
+| azure_monitor.data_explorer.materialized_view_age_seconds | average | seconds |
125
+| azure_monitor.data_explorer.materialized_view_records_in_delta | average | records |
126
+| azure_monitor.data_explorer.materialized_view_extents_rebuild | average | extents |
127
+| azure_monitor.data_explorer.materialized_view_data_loss | maximum | status |
128
+| azure_monitor.data_explorer.materialized_view_result | average | status |
129
+| azure_monitor.data_explorer.partitioning_percentage | total, hot | percentage |
130
+| azure_monitor.data_explorer.partitioned_records | average | records |
131
+
132
+
133
+
134
+## Alerts
135
+
136
+
137
+The following alerts are available:
138
+
139
+| Alert name | On metric | Description |
140
+|:------------|:----------|:------------|
141
+| [ am_data_explorer_keep_alive ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.keep_alive | Data Explorer keep alive on ${label:resource_name} |
142
+| [ am_data_explorer_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.cpu_utilization | Data Explorer CPU on ${label:resource_name} |
143
+| [ am_data_explorer_ingestion_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.utilization | Data Explorer ingestion utilization on ${label:resource_name} |
144
+| [ am_data_explorer_cache_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.utilization | Data Explorer cache utilization on ${label:resource_name} |
145
+| [ am_data_explorer_throttled_commands ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.throttled_commands | Data Explorer throttled commands on ${label:resource_name} |
146
+| [ am_data_explorer_query_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.query_duration | Data Explorer query duration on ${label:resource_name} |
147
+| [ am_data_explorer_throttled_queries ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.throttled_queries | Data Explorer throttled queries on ${label:resource_name} |
148
+| [ am_data_explorer_ingestion_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.ingestion_latency | Data Explorer ingestion latency on ${label:resource_name} |
149
+| [ am_data_explorer_events_dropped ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.events | Data Explorer events dropped on ${label:resource_name} |
150
+| [ am_data_explorer_blobs_dropped ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.blobs | Data Explorer blobs dropped on ${label:resource_name} |
151
+| [ am_data_explorer_ingestion_queue_length ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.ingestion_queue | Data Explorer ingestion queue on ${label:resource_name} |
152
+| [ am_data_explorer_queue_oldest_message ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.queue_oldest_message | Data Explorer queue oldest message age on ${label:resource_name} |
153
+| [ am_data_explorer_export_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.export_utilization | Data Explorer export utilization on ${label:resource_name} |
154
+| [ am_data_explorer_continuous_export_pending ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.continuous_export_pending | Data Explorer continuous export pending on ${label:resource_name} |
155
+| [ am_data_explorer_continuous_export_lateness ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.continuous_export_lateness | Data Explorer continuous export lateness on ${label:resource_name} |
156
+| [ am_data_explorer_streaming_ingest_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.streaming_ingest_utilization | Data Explorer streaming ingest utilization on ${label:resource_name} |
157
+| [ am_data_explorer_streaming_ingest_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.streaming_ingest_duration | Data Explorer streaming ingest duration on ${label:resource_name} |
158
+| [ am_data_explorer_materialized_view_health ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.materialized_view_health | Data Explorer materialized view health on ${label:resource_name} |
159
+| [ am_data_explorer_materialized_view_age ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.materialized_view_age | Data Explorer materialized view age on ${label:resource_name} |
160
+| [ am_data_explorer_materialized_view_data_loss ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.materialized_view_data_loss | Data Explorer materialized view data loss on ${label:resource_name} |
161
+| [ am_data_explorer_follower_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_explorer.conf) | azure_monitor.data_explorer.follower_latency | Data Explorer follower latency on ${label:resource_name} |
162
+
163
+
164
+## Setup
165
+
166
+
167
+You can configure the **azure_monitor** collector in two ways:
168
+
169
+| Method | Best for | How to |
170
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
171
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
172
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
173
+
174
+:::important
175
+
176
+UI configuration requires paid Netdata Cloud plan.
177
+
178
+:::
179
+
180
+
181
+### Prerequisites
182
+
183
+#### Create an Azure monitoring principal
184
+
185
+Create a service principal or use a managed identity with the following permissions:
186
+
187
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
188
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
189
+
190
+For service principal authentication:
191
+```bash
192
+# Create the service principal
193
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
194
+ --scopes /subscriptions/<subscription-id>
195
+
196
+# Note the appId (client_id), password (client_secret), and tenant
197
+```
198
+
199
+For managed identity (on Azure VMs, VMSS, or AKS):
200
+```bash
201
+# Assign Monitoring Reader role to the VM's managed identity
202
+az role assignment create --assignee <managed-identity-principal-id> \
203
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
204
+```
205
+
206
+
207
+
208
+### Configuration
209
+
210
+#### Options
211
+
212
+The following options can be defined globally: update_every, autodetection_retry.
213
+
214
+Profile files are loaded from:
215
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
216
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
217
+
218
+User profile files with the same filename override stock profiles.
219
+
220
+
221
+<details open><summary>Config options</summary>
222
+
223
+
224
+
225
+| Group | Option | Description | Default | Required |
226
+|:------|:-----|:------------|:--------|:---------:|
227
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
228
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
229
+| **Target** | subscription_id | Azure subscription ID. | | yes |
230
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
231
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
232
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
233
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
234
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
235
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
236
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
237
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
238
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
239
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
240
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
241
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
242
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
243
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
244
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
245
+
246
+
247
+</details>
248
+
249
+
250
+#### via UI
251
+
252
+Configure the **azure_monitor** collector from the Netdata web interface:
253
+
254
+1. Go to **Nodes**.
255
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
256
+3. The **Collectors → Jobs** view opens by default.
257
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
258
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
259
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
260
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
261
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
262
+
263
+
264
+#### via File
265
+
266
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
267
+
268
+The file format is YAML. Generally, the structure is:
269
+
270
+```yaml
271
+update_every: 1
272
+autodetection_retry: 0
273
+jobs:
274
+ - name: some_name1
275
+ - name: some_name2
276
+```
277
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
278
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
279
+
280
+```bash
281
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
282
+sudo ./edit-config go.d/azure_monitor.conf
283
+```
284
+
285
+##### Examples
286
+
287
+###### Service principal (auto-discover all resources)
288
+
289
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
290
+
291
+```yaml
292
+jobs:
293
+ - name: prod
294
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ auth:
296
+ mode: service_principal
297
+ mode_service_principal:
298
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
300
+ client_secret: "your-client-secret"
301
+
302
+```
303
+###### Managed identity (Azure VM/VMSS/AKS)
304
+
305
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
306
+
307
+<details open><summary>Config</summary>
308
+
309
+```yaml
310
+jobs:
311
+ - name: prod
312
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
313
+ auth:
314
+ mode: managed_identity
315
+
316
+```
317
+</details>
318
+
319
+###### Specific profiles only
320
+
321
+Monitor only specific Azure services instead of auto-discovering all resource types.
322
+
323
+<details open><summary>Config</summary>
324
+
325
+```yaml
326
+jobs:
327
+ - name: databases
328
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
329
+ profiles:
330
+ - sql_database
331
+ - postgres_flexible
332
+ - redis_cache
333
+ auth:
334
+ mode: service_principal
335
+ mode_service_principal:
336
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
337
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
338
+ client_secret: "your-client-secret"
339
+
340
+```
341
+</details>
342
+
343
+###### Filter by resource group
344
+
345
+Only monitor resources in specific resource groups.
346
+
347
+<details open><summary>Config</summary>
348
+
349
+```yaml
350
+jobs:
351
+ - name: prod-rg
352
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ resource_groups:
354
+ - production-rg
355
+ - staging-rg
356
+ auth:
357
+ mode: default
358
+
359
+```
360
+</details>
361
+
362
+###### Azure Government cloud
363
+
364
+Connect to Azure Government cloud environment.
365
+
366
+<details open><summary>Config</summary>
367
+
368
+```yaml
369
+jobs:
370
+ - name: gov
371
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
372
+ cloud: government
373
+ auth:
374
+ mode: service_principal
375
+ mode_service_principal:
376
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
377
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ client_secret: "your-client-secret"
379
+
380
+```
381
+</details>
382
+
383
+
384
+
385
+## Troubleshooting
386
+
387
+### Debug Mode
388
+
389
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
390
+
391
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
392
+should give you clues as to why the collector isn't working.
393
+
394
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
395
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
396
+
397
+ ```bash
398
+ cd /usr/libexec/netdata/plugins.d/
399
+ ```
400
+
401
+- Switch to the `netdata` user.
402
+
403
+ ```bash
404
+ sudo -u netdata -s
405
+ ```
406
+
407
+- Run the `go.d.plugin` to debug the collector:
408
+
409
+ ```bash
410
+ ./go.d.plugin -d -m azure_monitor
411
+ ```
412
+
413
+ To debug a specific job:
414
+
415
+ ```bash
416
+ ./go.d.plugin -d -m azure_monitor -j jobName
417
+ ```
418
+
419
+### Getting Logs
420
+
421
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
422
+
423
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
424
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
425
+
426
+#### System with systemd
427
+
428
+Use the following command to view logs generated since the last Netdata service restart:
429
+
430
+```bash
431
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
432
+```
433
+
434
+#### System without systemd
435
+
436
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
437
+
438
+```bash
439
+grep azure_monitor /var/log/netdata/collector.log
440
+```
441
+
442
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
443
+
444
+#### Docker Container
445
+
446
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
447
+
448
+```bash
449
+docker logs netdata 2>&1 | grep azure_monitor
450
+```
451
+
452
+### No metrics are collected
453
+
454
+Verify the following:
455
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
456
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
457
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
458
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
459
+
460
+
461
+### Missing metrics for some resource types
462
+
463
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
464
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
465
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
466
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
467
+
468
+
469
+### Metrics appear delayed
470
+
471
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
472
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
473
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
474
+
475
+
476
+### Authentication errors in sovereign clouds
477
+
478
+For Azure Government or Azure China clouds, set the `cloud` parameter:
479
+- Azure Government: `cloud: government`
480
+- Azure China (21Vianet): `cloud: china`
481
+
482
+Ensure the service principal is registered in the correct cloud tenant.
483
+
484
+
485
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_factory.md
new
+494
@@ -0,0 +1,494 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_factory.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Data Factory"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'data', 'factory', 'adf', 'etl', 'pipeline']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Data Factory
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Data Factory including pipeline, activity, and trigger run success and failure counts, integration runtime CPU and memory utilization, available capacity and queue lengths, SSIS package execution rates, copy operations throughput, data flow processing metrics, and overall factory resource utilization.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.data_factory.pipeline_runs | succeeded, failed, cancelled | runs/s |
88
+| azure_monitor.data_factory.pipeline_elapsed_time | elapsed_time | runs/s |
89
+| azure_monitor.data_factory.activity_runs | succeeded, failed, cancelled | runs/s |
90
+| azure_monitor.data_factory.trigger_runs | succeeded, failed, cancelled | runs/s |
91
+| azure_monitor.data_factory.ssis_ir_starts | succeeded, failed, cancelled | runs/s |
92
+| azure_monitor.data_factory.ssis_ir_stops | succeeded, stuck | runs/s |
93
+| azure_monitor.data_factory.ssis_package_executions | succeeded, failed, cancelled | executions/s |
94
+| azure_monitor.data_factory.ir_cpu | average | percentage |
95
+| azure_monitor.data_factory.ir_memory | available | bytes |
96
+| azure_monitor.data_factory.ir_nodes | available | nodes |
97
+| azure_monitor.data_factory.ir_queue | queue_length | tasks |
98
+| azure_monitor.data_factory.ir_task_pickup_delay | average | seconds |
99
+| azure_monitor.data_factory.factory_size | current, max_allowed | GiB |
100
+| azure_monitor.data_factory.entity_count | current, max_allowed | entities |
101
+| azure_monitor.data_factory.mvnet_ir_copy_capacity | utilization, available | percentage |
102
+| azure_monitor.data_factory.mvnet_ir_copy_queue | waiting | tasks |
103
+| azure_monitor.data_factory.mvnet_ir_external_capacity | utilization, available | percentage |
104
+| azure_monitor.data_factory.mvnet_ir_external_queue | waiting | tasks |
105
+| azure_monitor.data_factory.mvnet_ir_pipeline_capacity | utilization, available | percentage |
106
+| azure_monitor.data_factory.mvnet_ir_pipeline_queue | waiting | tasks |
107
+| azure_monitor.data_factory.airflow_ir_cpu | percentage | percentage |
108
+| azure_monitor.data_factory.airflow_ir_cpu_usage | average | millicores |
109
+| azure_monitor.data_factory.airflow_ir_memory | percentage | percentage |
110
+| azure_monitor.data_factory.airflow_ir_memory_usage | average | bytes |
111
+| azure_monitor.data_factory.airflow_ir_nodes | average | nodes |
112
+| azure_monitor.data_factory.airflow_ir_dag_collection | average | milliseconds |
113
+| azure_monitor.data_factory.airflow_ir_dag_bag | total | dags/s |
114
+| azure_monitor.data_factory.airflow_ir_dag_errors | callback_exceptions, file_refresh, import | errors/s |
115
+| azure_monitor.data_factory.airflow_ir_dag_processing_duration | last_duration | milliseconds |
116
+| azure_monitor.data_factory.airflow_ir_dag_processing_lag | last_run_seconds_ago, processor_timeouts, total_parse_time | seconds |
117
+| azure_monitor.data_factory.airflow_ir_dag_processing_activity | manager_stalls, processes | events/s |
118
+| azure_monitor.data_factory.airflow_ir_dag_run_duration | success, failed | milliseconds |
119
+| azure_monitor.data_factory.airflow_ir_dag_run_scheduling | dependency_check, first_task_delay, schedule_delay | milliseconds |
120
+| azure_monitor.data_factory.airflow_ir_executor | open_slots, queued_tasks, running_tasks | slots/s |
121
+| azure_monitor.data_factory.airflow_ir_jobs | started, ended, heartbeat_failures | jobs/s |
122
+| azure_monitor.data_factory.airflow_ir_operators | successes, failures | operations/s |
123
+| azure_monitor.data_factory.airflow_ir_pool_slots | open, queued, running | slots/s |
124
+| azure_monitor.data_factory.airflow_ir_pool_starving | starving | tasks/s |
125
+| azure_monitor.data_factory.airflow_ir_scheduler_critical_section | duration | milliseconds |
126
+| azure_monitor.data_factory.airflow_ir_scheduler_activity | critical_section_busy, heartbeats, failed_sla_emails | events/s |
127
+| azure_monitor.data_factory.airflow_ir_scheduler_orphaned_tasks | adopted, cleared | tasks/s |
128
+| azure_monitor.data_factory.airflow_ir_scheduler_tasks | executable, running, starving, killed_externally | tasks/s |
129
+| azure_monitor.data_factory.airflow_ir_task_instances | started, succeeded, failed, finished, previously_succeeded, created_via_operator | instances/s |
130
+| azure_monitor.data_factory.airflow_ir_task_instance_duration | average | milliseconds |
131
+| azure_monitor.data_factory.airflow_ir_task_dag_changes | removed, restored | tasks/s |
132
+| azure_monitor.data_factory.airflow_ir_triggers | succeeded, failed, running | triggers/s |
133
+| azure_monitor.data_factory.airflow_ir_trigger_issues | blocked_main_thread, celery_timeout_errors | events/s |
134
+| azure_monitor.data_factory.airflow_ir_zombies | killed | tasks/s |
135
+
136
+
137
+
138
+## Alerts
139
+
140
+
141
+The following alerts are available:
142
+
143
+| Alert name | On metric | Description |
144
+|:------------|:----------|:------------|
145
+| [ am_data_factory_pipeline_failed_runs ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.pipeline_runs | Data Factory pipeline failures on ${label:resource_name} |
146
+| [ am_data_factory_pipeline_cancelled_runs ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.pipeline_runs | Data Factory pipeline cancellations on ${label:resource_name} |
147
+| [ am_data_factory_activity_failed_runs ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.activity_runs | Data Factory activity failures on ${label:resource_name} |
148
+| [ am_data_factory_trigger_failed_runs ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.trigger_runs | Data Factory trigger failures on ${label:resource_name} |
149
+| [ am_data_factory_ssis_ir_start_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.ssis_ir_starts | Data Factory SSIS IR start failures on ${label:resource_name} |
150
+| [ am_data_factory_ssis_ir_stop_stuck ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.ssis_ir_stops | Data Factory SSIS IR stuck stops on ${label:resource_name} |
151
+| [ am_data_factory_ssis_package_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.ssis_package_executions | Data Factory SSIS package failures on ${label:resource_name} |
152
+| [ am_data_factory_ir_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.ir_cpu | Data Factory IR CPU on ${label:resource_name} |
153
+| [ am_data_factory_ir_queue_length ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.ir_queue | Data Factory IR queue depth on ${label:resource_name} |
154
+| [ am_data_factory_ir_task_pickup_delay ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.ir_task_pickup_delay | Data Factory IR task pickup delay on ${label:resource_name} |
155
+| [ am_data_factory_size_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.factory_size | Data Factory size utilization on ${label:resource_name} |
156
+| [ am_data_factory_entity_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.entity_count | Data Factory entity count on ${label:resource_name} |
157
+| [ am_data_factory_mvnet_ir_copy_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.mvnet_ir_copy_capacity | Data Factory MVNet IR copy utilization on ${label:resource_name} |
158
+| [ am_data_factory_mvnet_ir_external_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.mvnet_ir_external_capacity | Data Factory MVNet IR external utilization on ${label:resource_name} |
159
+| [ am_data_factory_mvnet_ir_pipeline_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.mvnet_ir_pipeline_capacity | Data Factory MVNet IR pipeline utilization on ${label:resource_name} |
160
+| [ am_data_factory_airflow_ir_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_cpu | Data Factory Airflow IR CPU on ${label:resource_name} |
161
+| [ am_data_factory_airflow_ir_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_memory | Data Factory Airflow IR memory on ${label:resource_name} |
162
+| [ am_data_factory_airflow_ir_dag_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_dag_errors | Data Factory Airflow DAG errors on ${label:resource_name} |
163
+| [ am_data_factory_airflow_ir_operator_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_operators | Data Factory Airflow operator failures on ${label:resource_name} |
164
+| [ am_data_factory_airflow_ir_job_heartbeat_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_jobs | Data Factory Airflow job heartbeat failures on ${label:resource_name} |
165
+| [ am_data_factory_airflow_ir_pool_starving ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_pool_starving | Data Factory Airflow pool starvation on ${label:resource_name} |
166
+| [ am_data_factory_airflow_ir_tasks_killed_externally ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_scheduler_tasks | Data Factory Airflow tasks killed externally on ${label:resource_name} |
167
+| [ am_data_factory_airflow_ir_tasks_starving ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_scheduler_tasks | Data Factory Airflow scheduler starving tasks on ${label:resource_name} |
168
+| [ am_data_factory_airflow_ir_task_instance_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_task_instances | Data Factory Airflow task failures on ${label:resource_name} |
169
+| [ am_data_factory_airflow_ir_trigger_issues ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_trigger_issues | Data Factory Airflow trigger issues on ${label:resource_name} |
170
+| [ am_data_factory_airflow_ir_zombies ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_data_factory.conf) | azure_monitor.data_factory.airflow_ir_zombies | Data Factory Airflow zombie tasks on ${label:resource_name} |
171
+
172
+
173
+## Setup
174
+
175
+
176
+You can configure the **azure_monitor** collector in two ways:
177
+
178
+| Method | Best for | How to |
179
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
180
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
181
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
182
+
183
+:::important
184
+
185
+UI configuration requires paid Netdata Cloud plan.
186
+
187
+:::
188
+
189
+
190
+### Prerequisites
191
+
192
+#### Create an Azure monitoring principal
193
+
194
+Create a service principal or use a managed identity with the following permissions:
195
+
196
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
197
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
198
+
199
+For service principal authentication:
200
+```bash
201
+# Create the service principal
202
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
203
+ --scopes /subscriptions/<subscription-id>
204
+
205
+# Note the appId (client_id), password (client_secret), and tenant
206
+```
207
+
208
+For managed identity (on Azure VMs, VMSS, or AKS):
209
+```bash
210
+# Assign Monitoring Reader role to the VM's managed identity
211
+az role assignment create --assignee <managed-identity-principal-id> \
212
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
213
+```
214
+
215
+
216
+
217
+### Configuration
218
+
219
+#### Options
220
+
221
+The following options can be defined globally: update_every, autodetection_retry.
222
+
223
+Profile files are loaded from:
224
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
225
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
226
+
227
+User profile files with the same filename override stock profiles.
228
+
229
+
230
+<details open><summary>Config options</summary>
231
+
232
+
233
+
234
+| Group | Option | Description | Default | Required |
235
+|:------|:-----|:------------|:--------|:---------:|
236
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
237
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
238
+| **Target** | subscription_id | Azure subscription ID. | | yes |
239
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
240
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
241
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
242
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
243
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
244
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
245
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
246
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
247
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
248
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
249
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
250
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
251
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
252
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
253
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
254
+
255
+
256
+</details>
257
+
258
+
259
+#### via UI
260
+
261
+Configure the **azure_monitor** collector from the Netdata web interface:
262
+
263
+1. Go to **Nodes**.
264
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
265
+3. The **Collectors → Jobs** view opens by default.
266
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
267
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
268
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
269
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
270
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
271
+
272
+
273
+#### via File
274
+
275
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
276
+
277
+The file format is YAML. Generally, the structure is:
278
+
279
+```yaml
280
+update_every: 1
281
+autodetection_retry: 0
282
+jobs:
283
+ - name: some_name1
284
+ - name: some_name2
285
+```
286
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
287
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
288
+
289
+```bash
290
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
291
+sudo ./edit-config go.d/azure_monitor.conf
292
+```
293
+
294
+##### Examples
295
+
296
+###### Service principal (auto-discover all resources)
297
+
298
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
299
+
300
+```yaml
301
+jobs:
302
+ - name: prod
303
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
304
+ auth:
305
+ mode: service_principal
306
+ mode_service_principal:
307
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
308
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
309
+ client_secret: "your-client-secret"
310
+
311
+```
312
+###### Managed identity (Azure VM/VMSS/AKS)
313
+
314
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
315
+
316
+<details open><summary>Config</summary>
317
+
318
+```yaml
319
+jobs:
320
+ - name: prod
321
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ auth:
323
+ mode: managed_identity
324
+
325
+```
326
+</details>
327
+
328
+###### Specific profiles only
329
+
330
+Monitor only specific Azure services instead of auto-discovering all resource types.
331
+
332
+<details open><summary>Config</summary>
333
+
334
+```yaml
335
+jobs:
336
+ - name: databases
337
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
338
+ profiles:
339
+ - sql_database
340
+ - postgres_flexible
341
+ - redis_cache
342
+ auth:
343
+ mode: service_principal
344
+ mode_service_principal:
345
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
346
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
347
+ client_secret: "your-client-secret"
348
+
349
+```
350
+</details>
351
+
352
+###### Filter by resource group
353
+
354
+Only monitor resources in specific resource groups.
355
+
356
+<details open><summary>Config</summary>
357
+
358
+```yaml
359
+jobs:
360
+ - name: prod-rg
361
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
362
+ resource_groups:
363
+ - production-rg
364
+ - staging-rg
365
+ auth:
366
+ mode: default
367
+
368
+```
369
+</details>
370
+
371
+###### Azure Government cloud
372
+
373
+Connect to Azure Government cloud environment.
374
+
375
+<details open><summary>Config</summary>
376
+
377
+```yaml
378
+jobs:
379
+ - name: gov
380
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ cloud: government
382
+ auth:
383
+ mode: service_principal
384
+ mode_service_principal:
385
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
386
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
387
+ client_secret: "your-client-secret"
388
+
389
+```
390
+</details>
391
+
392
+
393
+
394
+## Troubleshooting
395
+
396
+### Debug Mode
397
+
398
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
399
+
400
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
401
+should give you clues as to why the collector isn't working.
402
+
403
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
404
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
405
+
406
+ ```bash
407
+ cd /usr/libexec/netdata/plugins.d/
408
+ ```
409
+
410
+- Switch to the `netdata` user.
411
+
412
+ ```bash
413
+ sudo -u netdata -s
414
+ ```
415
+
416
+- Run the `go.d.plugin` to debug the collector:
417
+
418
+ ```bash
419
+ ./go.d.plugin -d -m azure_monitor
420
+ ```
421
+
422
+ To debug a specific job:
423
+
424
+ ```bash
425
+ ./go.d.plugin -d -m azure_monitor -j jobName
426
+ ```
427
+
428
+### Getting Logs
429
+
430
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
431
+
432
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
433
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
434
+
435
+#### System with systemd
436
+
437
+Use the following command to view logs generated since the last Netdata service restart:
438
+
439
+```bash
440
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
441
+```
442
+
443
+#### System without systemd
444
+
445
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
446
+
447
+```bash
448
+grep azure_monitor /var/log/netdata/collector.log
449
+```
450
+
451
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
452
+
453
+#### Docker Container
454
+
455
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
456
+
457
+```bash
458
+docker logs netdata 2>&1 | grep azure_monitor
459
+```
460
+
461
+### No metrics are collected
462
+
463
+Verify the following:
464
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
465
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
466
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
467
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
468
+
469
+
470
+### Missing metrics for some resource types
471
+
472
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
473
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
474
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
475
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
476
+
477
+
478
+### Metrics appear delayed
479
+
480
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
481
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
482
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
483
+
484
+
485
+### Authentication errors in sovereign clouds
486
+
487
+For Azure Government or Azure China clouds, set the `cloud` parameter:
488
+- Azure Government: `cloud: government`
489
+- Azure China (21Vianet): `cloud: china`
490
+
491
+Ensure the service principal is registered in the correct cloud tenant.
492
+
493
+
494
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_grid_topic.md
new
+432
@@ -0,0 +1,432 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_grid_topic.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Event Grid Topic"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'event', 'grid', 'messaging', 'events', 'serverless']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Event Grid Topic
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Event Grid topics including publish success and failure counts, publish latency, event delivery and routing rates, delivery success and failure counts, dead-lettered events, and matched event routing.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.event_grid.publish_rate | success, failed | events/s |
88
+| azure_monitor.event_grid.publish_latency | latency | milliseconds |
89
+| azure_monitor.event_grid.routing | matched, unmatched | events/s |
90
+| azure_monitor.event_grid.advanced_filter_evaluations | evaluations | evaluations/s |
91
+| azure_monitor.event_grid.delivery | delivered, failed, dropped, dead_lettered | events/s |
92
+| azure_monitor.event_grid.destination_processing_duration | average | milliseconds |
93
+
94
+
95
+
96
+## Alerts
97
+
98
+
99
+The following alerts are available:
100
+
101
+| Alert name | On metric | Description |
102
+|:------------|:----------|:------------|
103
+| [ am_event_grid_publish_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_grid.conf) | azure_monitor.event_grid.publish_rate | Event Grid publish failures on ${label:resource_name} |
104
+| [ am_event_grid_delivery_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_grid.conf) | azure_monitor.event_grid.delivery | Event Grid delivery failures on ${label:resource_name} |
105
+| [ am_event_grid_dropped_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_grid.conf) | azure_monitor.event_grid.delivery | Event Grid dropped events on ${label:resource_name} |
106
+| [ am_event_grid_dead_lettered_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_grid.conf) | azure_monitor.event_grid.delivery | Event Grid dead-lettered events on ${label:resource_name} |
107
+| [ am_event_grid_unmatched_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_grid.conf) | azure_monitor.event_grid.routing | Event Grid unmatched events on ${label:resource_name} |
108
+| [ am_event_grid_destination_processing_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_grid.conf) | azure_monitor.event_grid.destination_processing_duration | Event Grid destination processing duration on ${label:resource_name} |
109
+
110
+
111
+## Setup
112
+
113
+
114
+You can configure the **azure_monitor** collector in two ways:
115
+
116
+| Method | Best for | How to |
117
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
118
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
119
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
120
+
121
+:::important
122
+
123
+UI configuration requires paid Netdata Cloud plan.
124
+
125
+:::
126
+
127
+
128
+### Prerequisites
129
+
130
+#### Create an Azure monitoring principal
131
+
132
+Create a service principal or use a managed identity with the following permissions:
133
+
134
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
135
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
136
+
137
+For service principal authentication:
138
+```bash
139
+# Create the service principal
140
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
141
+ --scopes /subscriptions/<subscription-id>
142
+
143
+# Note the appId (client_id), password (client_secret), and tenant
144
+```
145
+
146
+For managed identity (on Azure VMs, VMSS, or AKS):
147
+```bash
148
+# Assign Monitoring Reader role to the VM's managed identity
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
151
+```
152
+
153
+
154
+
155
+### Configuration
156
+
157
+#### Options
158
+
159
+The following options can be defined globally: update_every, autodetection_retry.
160
+
161
+Profile files are loaded from:
162
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
163
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
164
+
165
+User profile files with the same filename override stock profiles.
166
+
167
+
168
+<details open><summary>Config options</summary>
169
+
170
+
171
+
172
+| Group | Option | Description | Default | Required |
173
+|:------|:-----|:------------|:--------|:---------:|
174
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
176
+| **Target** | subscription_id | Azure subscription ID. | | yes |
177
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
178
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
179
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
181
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
182
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
183
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
184
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
185
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
186
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
187
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
188
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
189
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
190
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
191
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
192
+
193
+
194
+</details>
195
+
196
+
197
+#### via UI
198
+
199
+Configure the **azure_monitor** collector from the Netdata web interface:
200
+
201
+1. Go to **Nodes**.
202
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
203
+3. The **Collectors → Jobs** view opens by default.
204
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
205
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
206
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
207
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
208
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
209
+
210
+
211
+#### via File
212
+
213
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
214
+
215
+The file format is YAML. Generally, the structure is:
216
+
217
+```yaml
218
+update_every: 1
219
+autodetection_retry: 0
220
+jobs:
221
+ - name: some_name1
222
+ - name: some_name2
223
+```
224
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
225
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
226
+
227
+```bash
228
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
229
+sudo ./edit-config go.d/azure_monitor.conf
230
+```
231
+
232
+##### Examples
233
+
234
+###### Service principal (auto-discover all resources)
235
+
236
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
237
+
238
+```yaml
239
+jobs:
240
+ - name: prod
241
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ auth:
243
+ mode: service_principal
244
+ mode_service_principal:
245
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
246
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
247
+ client_secret: "your-client-secret"
248
+
249
+```
250
+###### Managed identity (Azure VM/VMSS/AKS)
251
+
252
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
253
+
254
+<details open><summary>Config</summary>
255
+
256
+```yaml
257
+jobs:
258
+ - name: prod
259
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
260
+ auth:
261
+ mode: managed_identity
262
+
263
+```
264
+</details>
265
+
266
+###### Specific profiles only
267
+
268
+Monitor only specific Azure services instead of auto-discovering all resource types.
269
+
270
+<details open><summary>Config</summary>
271
+
272
+```yaml
273
+jobs:
274
+ - name: databases
275
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
276
+ profiles:
277
+ - sql_database
278
+ - postgres_flexible
279
+ - redis_cache
280
+ auth:
281
+ mode: service_principal
282
+ mode_service_principal:
283
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
284
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
285
+ client_secret: "your-client-secret"
286
+
287
+```
288
+</details>
289
+
290
+###### Filter by resource group
291
+
292
+Only monitor resources in specific resource groups.
293
+
294
+<details open><summary>Config</summary>
295
+
296
+```yaml
297
+jobs:
298
+ - name: prod-rg
299
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
300
+ resource_groups:
301
+ - production-rg
302
+ - staging-rg
303
+ auth:
304
+ mode: default
305
+
306
+```
307
+</details>
308
+
309
+###### Azure Government cloud
310
+
311
+Connect to Azure Government cloud environment.
312
+
313
+<details open><summary>Config</summary>
314
+
315
+```yaml
316
+jobs:
317
+ - name: gov
318
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
319
+ cloud: government
320
+ auth:
321
+ mode: service_principal
322
+ mode_service_principal:
323
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ client_secret: "your-client-secret"
326
+
327
+```
328
+</details>
329
+
330
+
331
+
332
+## Troubleshooting
333
+
334
+### Debug Mode
335
+
336
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
337
+
338
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
339
+should give you clues as to why the collector isn't working.
340
+
341
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
342
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
343
+
344
+ ```bash
345
+ cd /usr/libexec/netdata/plugins.d/
346
+ ```
347
+
348
+- Switch to the `netdata` user.
349
+
350
+ ```bash
351
+ sudo -u netdata -s
352
+ ```
353
+
354
+- Run the `go.d.plugin` to debug the collector:
355
+
356
+ ```bash
357
+ ./go.d.plugin -d -m azure_monitor
358
+ ```
359
+
360
+ To debug a specific job:
361
+
362
+ ```bash
363
+ ./go.d.plugin -d -m azure_monitor -j jobName
364
+ ```
365
+
366
+### Getting Logs
367
+
368
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
369
+
370
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
371
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
372
+
373
+#### System with systemd
374
+
375
+Use the following command to view logs generated since the last Netdata service restart:
376
+
377
+```bash
378
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
379
+```
380
+
381
+#### System without systemd
382
+
383
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
384
+
385
+```bash
386
+grep azure_monitor /var/log/netdata/collector.log
387
+```
388
+
389
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
390
+
391
+#### Docker Container
392
+
393
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
394
+
395
+```bash
396
+docker logs netdata 2>&1 | grep azure_monitor
397
+```
398
+
399
+### No metrics are collected
400
+
401
+Verify the following:
402
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
403
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
404
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
405
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
406
+
407
+
408
+### Missing metrics for some resource types
409
+
410
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
411
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
412
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
413
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
414
+
415
+
416
+### Metrics appear delayed
417
+
418
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
419
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
420
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
421
+
422
+
423
+### Authentication errors in sovereign clouds
424
+
425
+For Azure Government or Azure China clouds, set the `cloud` parameter:
426
+- Azure Government: `cloud: government`
427
+- Azure China (21Vianet): `cloud: china`
428
+
429
+Ensure the service principal is registered in the correct cloud tenant.
430
+
431
+
432
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_hubs_namespace.md
new
+443
@@ -0,0 +1,443 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_hubs_namespace.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Event Hubs Namespace"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'event', 'hubs', 'streaming', 'kafka', 'messaging']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Event Hubs Namespace
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Event Hubs namespaces including incoming and outgoing message rates, byte throughput, captured messages and bytes, throttled and quota-exceeded request counts, active connections, and total connection counts.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.event_hubs.message_flow | in, out | messages/s |
88
+| azure_monitor.event_hubs.data_throughput | in, out | bytes/s |
89
+| azure_monitor.event_hubs.requests | incoming, successful | requests/s |
90
+| azure_monitor.event_hubs.errors | server, user, throttled, quota_exceeded | errors/s |
91
+| azure_monitor.event_hubs.connections | active | connections |
92
+| azure_monitor.event_hubs.connection_events | opened, closed | connections |
93
+| azure_monitor.event_hubs.captured_messages | total | messages/s |
94
+| azure_monitor.event_hubs.captured_data | total | bytes/s |
95
+| azure_monitor.event_hubs.namespace_size | average | bytes |
96
+| azure_monitor.event_hubs.capture_backlog | backlog | messages |
97
+| azure_monitor.event_hubs.namespace_resources | cpu, memory | percentage |
98
+| azure_monitor.event_hubs.replication_lag | messages | messages |
99
+| azure_monitor.event_hubs.replication_lag_duration | duration | seconds |
100
+
101
+
102
+
103
+## Alerts
104
+
105
+
106
+The following alerts are available:
107
+
108
+| Alert name | On metric | Description |
109
+|:------------|:----------|:------------|
110
+| [ am_event_hubs_server_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.errors | Event Hubs server errors on ${label:resource_name} |
111
+| [ am_event_hubs_throttled_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.errors | Event Hubs throttled requests on ${label:resource_name} |
112
+| [ am_event_hubs_quota_exceeded ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.errors | Event Hubs quota exceeded on ${label:resource_name} |
113
+| [ am_event_hubs_user_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.errors | Event Hubs user errors on ${label:resource_name} |
114
+| [ am_event_hubs_success_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.requests | Event Hubs request success rate on ${label:resource_name} |
115
+| [ am_event_hubs_namespace_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.namespace_resources | Event Hubs namespace CPU on ${label:resource_name} |
116
+| [ am_event_hubs_namespace_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.namespace_resources | Event Hubs namespace memory on ${label:resource_name} |
117
+| [ am_event_hubs_capture_backlog ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.capture_backlog | Event Hubs capture backlog on ${label:resource_name} |
118
+| [ am_event_hubs_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.replication_lag | Event Hubs replication lag on ${label:resource_name} |
119
+| [ am_event_hubs_replication_lag_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_event_hubs.conf) | azure_monitor.event_hubs.replication_lag_duration | Event Hubs replication lag duration on ${label:resource_name} |
120
+
121
+
122
+## Setup
123
+
124
+
125
+You can configure the **azure_monitor** collector in two ways:
126
+
127
+| Method | Best for | How to |
128
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
129
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
130
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
131
+
132
+:::important
133
+
134
+UI configuration requires paid Netdata Cloud plan.
135
+
136
+:::
137
+
138
+
139
+### Prerequisites
140
+
141
+#### Create an Azure monitoring principal
142
+
143
+Create a service principal or use a managed identity with the following permissions:
144
+
145
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
146
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
147
+
148
+For service principal authentication:
149
+```bash
150
+# Create the service principal
151
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
152
+ --scopes /subscriptions/<subscription-id>
153
+
154
+# Note the appId (client_id), password (client_secret), and tenant
155
+```
156
+
157
+For managed identity (on Azure VMs, VMSS, or AKS):
158
+```bash
159
+# Assign Monitoring Reader role to the VM's managed identity
160
+az role assignment create --assignee <managed-identity-principal-id> \
161
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
162
+```
163
+
164
+
165
+
166
+### Configuration
167
+
168
+#### Options
169
+
170
+The following options can be defined globally: update_every, autodetection_retry.
171
+
172
+Profile files are loaded from:
173
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
174
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
175
+
176
+User profile files with the same filename override stock profiles.
177
+
178
+
179
+<details open><summary>Config options</summary>
180
+
181
+
182
+
183
+| Group | Option | Description | Default | Required |
184
+|:------|:-----|:------------|:--------|:---------:|
185
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
186
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
187
+| **Target** | subscription_id | Azure subscription ID. | | yes |
188
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
189
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
190
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
191
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
192
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
193
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
194
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
195
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
196
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
197
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
198
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
199
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
200
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
201
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
202
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
+
204
+
205
+</details>
206
+
207
+
208
+#### via UI
209
+
210
+Configure the **azure_monitor** collector from the Netdata web interface:
211
+
212
+1. Go to **Nodes**.
213
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
214
+3. The **Collectors → Jobs** view opens by default.
215
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
216
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
217
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
218
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
219
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
220
+
221
+
222
+#### via File
223
+
224
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
225
+
226
+The file format is YAML. Generally, the structure is:
227
+
228
+```yaml
229
+update_every: 1
230
+autodetection_retry: 0
231
+jobs:
232
+ - name: some_name1
233
+ - name: some_name2
234
+```
235
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
236
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
237
+
238
+```bash
239
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
240
+sudo ./edit-config go.d/azure_monitor.conf
241
+```
242
+
243
+##### Examples
244
+
245
+###### Service principal (auto-discover all resources)
246
+
247
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
248
+
249
+```yaml
250
+jobs:
251
+ - name: prod
252
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
253
+ auth:
254
+ mode: service_principal
255
+ mode_service_principal:
256
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
257
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
258
+ client_secret: "your-client-secret"
259
+
260
+```
261
+###### Managed identity (Azure VM/VMSS/AKS)
262
+
263
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
264
+
265
+<details open><summary>Config</summary>
266
+
267
+```yaml
268
+jobs:
269
+ - name: prod
270
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
271
+ auth:
272
+ mode: managed_identity
273
+
274
+```
275
+</details>
276
+
277
+###### Specific profiles only
278
+
279
+Monitor only specific Azure services instead of auto-discovering all resource types.
280
+
281
+<details open><summary>Config</summary>
282
+
283
+```yaml
284
+jobs:
285
+ - name: databases
286
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
287
+ profiles:
288
+ - sql_database
289
+ - postgres_flexible
290
+ - redis_cache
291
+ auth:
292
+ mode: service_principal
293
+ mode_service_principal:
294
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
296
+ client_secret: "your-client-secret"
297
+
298
+```
299
+</details>
300
+
301
+###### Filter by resource group
302
+
303
+Only monitor resources in specific resource groups.
304
+
305
+<details open><summary>Config</summary>
306
+
307
+```yaml
308
+jobs:
309
+ - name: prod-rg
310
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
311
+ resource_groups:
312
+ - production-rg
313
+ - staging-rg
314
+ auth:
315
+ mode: default
316
+
317
+```
318
+</details>
319
+
320
+###### Azure Government cloud
321
+
322
+Connect to Azure Government cloud environment.
323
+
324
+<details open><summary>Config</summary>
325
+
326
+```yaml
327
+jobs:
328
+ - name: gov
329
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
330
+ cloud: government
331
+ auth:
332
+ mode: service_principal
333
+ mode_service_principal:
334
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
335
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
336
+ client_secret: "your-client-secret"
337
+
338
+```
339
+</details>
340
+
341
+
342
+
343
+## Troubleshooting
344
+
345
+### Debug Mode
346
+
347
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
348
+
349
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
350
+should give you clues as to why the collector isn't working.
351
+
352
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
353
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
354
+
355
+ ```bash
356
+ cd /usr/libexec/netdata/plugins.d/
357
+ ```
358
+
359
+- Switch to the `netdata` user.
360
+
361
+ ```bash
362
+ sudo -u netdata -s
363
+ ```
364
+
365
+- Run the `go.d.plugin` to debug the collector:
366
+
367
+ ```bash
368
+ ./go.d.plugin -d -m azure_monitor
369
+ ```
370
+
371
+ To debug a specific job:
372
+
373
+ ```bash
374
+ ./go.d.plugin -d -m azure_monitor -j jobName
375
+ ```
376
+
377
+### Getting Logs
378
+
379
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
380
+
381
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
382
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
383
+
384
+#### System with systemd
385
+
386
+Use the following command to view logs generated since the last Netdata service restart:
387
+
388
+```bash
389
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
390
+```
391
+
392
+#### System without systemd
393
+
394
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
395
+
396
+```bash
397
+grep azure_monitor /var/log/netdata/collector.log
398
+```
399
+
400
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
401
+
402
+#### Docker Container
403
+
404
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
405
+
406
+```bash
407
+docker logs netdata 2>&1 | grep azure_monitor
408
+```
409
+
410
+### No metrics are collected
411
+
412
+Verify the following:
413
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
414
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
415
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
416
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
417
+
418
+
419
+### Missing metrics for some resource types
420
+
421
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
422
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
423
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
424
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
425
+
426
+
427
+### Metrics appear delayed
428
+
429
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
430
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
431
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
432
+
433
+
434
+### Authentication errors in sovereign clouds
435
+
436
+For Azure Government or Azure China clouds, set the `cloud` parameter:
437
+- Azure Government: `cloud: government`
438
+- Azure China (21Vianet): `cloud: china`
439
+
440
+Ensure the service principal is registered in the correct cloud tenant.
441
+
442
+
443
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_circuit.md
new
+433
@@ -0,0 +1,433 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_circuit.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure ExpressRoute Circuit"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'expressroute', 'circuit', 'networking', 'hybrid']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure ExpressRoute Circuit
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor ExpressRoute circuits including bits per second in and out, ARP and BGP availability percentages, packet drops, and QoS bit rate throughput.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.express_route_circuit.arp_availability | average | percentage |
88
+| azure_monitor.express_route_circuit.bgp_availability | average | percentage |
89
+| azure_monitor.express_route_circuit.throughput | in, out | bits/s |
90
+| azure_monitor.express_route_circuit.bandwidth_utilization | ingress, egress | percentage |
91
+| azure_monitor.express_route_circuit.qos_dropped_bits | in, out | bits/s |
92
+| azure_monitor.express_route_circuit.fastpath_routes | maximum | routes |
93
+| azure_monitor.express_route_circuit.global_reach_throughput | in, out | bits/s |
94
+
95
+
96
+
97
+## Alerts
98
+
99
+
100
+The following alerts are available:
101
+
102
+| Alert name | On metric | Description |
103
+|:------------|:----------|:------------|
104
+| [ am_express_route_circuit_arp_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_circuit.conf) | azure_monitor.express_route_circuit.arp_availability | ExpressRoute ARP availability on ${label:resource_name} |
105
+| [ am_express_route_circuit_bgp_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_circuit.conf) | azure_monitor.express_route_circuit.bgp_availability | ExpressRoute BGP availability on ${label:resource_name} |
106
+| [ am_express_route_circuit_ingress_bandwidth_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_circuit.conf) | azure_monitor.express_route_circuit.bandwidth_utilization | ExpressRoute ingress bandwidth on ${label:resource_name} |
107
+| [ am_express_route_circuit_egress_bandwidth_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_circuit.conf) | azure_monitor.express_route_circuit.bandwidth_utilization | ExpressRoute egress bandwidth on ${label:resource_name} |
108
+| [ am_express_route_circuit_qos_drop_in ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_circuit.conf) | azure_monitor.express_route_circuit.qos_dropped_bits | ExpressRoute QoS ingress drops on ${label:resource_name} |
109
+| [ am_express_route_circuit_qos_drop_out ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_circuit.conf) | azure_monitor.express_route_circuit.qos_dropped_bits | ExpressRoute QoS egress drops on ${label:resource_name} |
110
+
111
+
112
+## Setup
113
+
114
+
115
+You can configure the **azure_monitor** collector in two ways:
116
+
117
+| Method | Best for | How to |
118
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
119
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
120
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
121
+
122
+:::important
123
+
124
+UI configuration requires paid Netdata Cloud plan.
125
+
126
+:::
127
+
128
+
129
+### Prerequisites
130
+
131
+#### Create an Azure monitoring principal
132
+
133
+Create a service principal or use a managed identity with the following permissions:
134
+
135
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
136
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
137
+
138
+For service principal authentication:
139
+```bash
140
+# Create the service principal
141
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
142
+ --scopes /subscriptions/<subscription-id>
143
+
144
+# Note the appId (client_id), password (client_secret), and tenant
145
+```
146
+
147
+For managed identity (on Azure VMs, VMSS, or AKS):
148
+```bash
149
+# Assign Monitoring Reader role to the VM's managed identity
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
152
+```
153
+
154
+
155
+
156
+### Configuration
157
+
158
+#### Options
159
+
160
+The following options can be defined globally: update_every, autodetection_retry.
161
+
162
+Profile files are loaded from:
163
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
164
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
165
+
166
+User profile files with the same filename override stock profiles.
167
+
168
+
169
+<details open><summary>Config options</summary>
170
+
171
+
172
+
173
+| Group | Option | Description | Default | Required |
174
+|:------|:-----|:------------|:--------|:---------:|
175
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
176
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
177
+| **Target** | subscription_id | Azure subscription ID. | | yes |
178
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
179
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
180
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
182
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
183
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
184
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
185
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
186
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
187
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
188
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
189
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
190
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
191
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
192
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
193
+
194
+
195
+</details>
196
+
197
+
198
+#### via UI
199
+
200
+Configure the **azure_monitor** collector from the Netdata web interface:
201
+
202
+1. Go to **Nodes**.
203
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
204
+3. The **Collectors → Jobs** view opens by default.
205
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
206
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
207
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
208
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
209
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
210
+
211
+
212
+#### via File
213
+
214
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
215
+
216
+The file format is YAML. Generally, the structure is:
217
+
218
+```yaml
219
+update_every: 1
220
+autodetection_retry: 0
221
+jobs:
222
+ - name: some_name1
223
+ - name: some_name2
224
+```
225
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
226
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
227
+
228
+```bash
229
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
230
+sudo ./edit-config go.d/azure_monitor.conf
231
+```
232
+
233
+##### Examples
234
+
235
+###### Service principal (auto-discover all resources)
236
+
237
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
238
+
239
+```yaml
240
+jobs:
241
+ - name: prod
242
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
243
+ auth:
244
+ mode: service_principal
245
+ mode_service_principal:
246
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
247
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
248
+ client_secret: "your-client-secret"
249
+
250
+```
251
+###### Managed identity (Azure VM/VMSS/AKS)
252
+
253
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
254
+
255
+<details open><summary>Config</summary>
256
+
257
+```yaml
258
+jobs:
259
+ - name: prod
260
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
261
+ auth:
262
+ mode: managed_identity
263
+
264
+```
265
+</details>
266
+
267
+###### Specific profiles only
268
+
269
+Monitor only specific Azure services instead of auto-discovering all resource types.
270
+
271
+<details open><summary>Config</summary>
272
+
273
+```yaml
274
+jobs:
275
+ - name: databases
276
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
277
+ profiles:
278
+ - sql_database
279
+ - postgres_flexible
280
+ - redis_cache
281
+ auth:
282
+ mode: service_principal
283
+ mode_service_principal:
284
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
285
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
286
+ client_secret: "your-client-secret"
287
+
288
+```
289
+</details>
290
+
291
+###### Filter by resource group
292
+
293
+Only monitor resources in specific resource groups.
294
+
295
+<details open><summary>Config</summary>
296
+
297
+```yaml
298
+jobs:
299
+ - name: prod-rg
300
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
301
+ resource_groups:
302
+ - production-rg
303
+ - staging-rg
304
+ auth:
305
+ mode: default
306
+
307
+```
308
+</details>
309
+
310
+###### Azure Government cloud
311
+
312
+Connect to Azure Government cloud environment.
313
+
314
+<details open><summary>Config</summary>
315
+
316
+```yaml
317
+jobs:
318
+ - name: gov
319
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ cloud: government
321
+ auth:
322
+ mode: service_principal
323
+ mode_service_principal:
324
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ client_secret: "your-client-secret"
327
+
328
+```
329
+</details>
330
+
331
+
332
+
333
+## Troubleshooting
334
+
335
+### Debug Mode
336
+
337
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
338
+
339
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
340
+should give you clues as to why the collector isn't working.
341
+
342
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
343
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
344
+
345
+ ```bash
346
+ cd /usr/libexec/netdata/plugins.d/
347
+ ```
348
+
349
+- Switch to the `netdata` user.
350
+
351
+ ```bash
352
+ sudo -u netdata -s
353
+ ```
354
+
355
+- Run the `go.d.plugin` to debug the collector:
356
+
357
+ ```bash
358
+ ./go.d.plugin -d -m azure_monitor
359
+ ```
360
+
361
+ To debug a specific job:
362
+
363
+ ```bash
364
+ ./go.d.plugin -d -m azure_monitor -j jobName
365
+ ```
366
+
367
+### Getting Logs
368
+
369
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
370
+
371
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
372
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
373
+
374
+#### System with systemd
375
+
376
+Use the following command to view logs generated since the last Netdata service restart:
377
+
378
+```bash
379
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
380
+```
381
+
382
+#### System without systemd
383
+
384
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
385
+
386
+```bash
387
+grep azure_monitor /var/log/netdata/collector.log
388
+```
389
+
390
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
391
+
392
+#### Docker Container
393
+
394
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
395
+
396
+```bash
397
+docker logs netdata 2>&1 | grep azure_monitor
398
+```
399
+
400
+### No metrics are collected
401
+
402
+Verify the following:
403
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
404
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
405
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
406
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
407
+
408
+
409
+### Missing metrics for some resource types
410
+
411
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
412
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
413
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
414
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
415
+
416
+
417
+### Metrics appear delayed
418
+
419
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
420
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
421
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
422
+
423
+
424
+### Authentication errors in sovereign clouds
425
+
426
+For Azure Government or Azure China clouds, set the `cloud` parameter:
427
+- Azure Government: `cloud: government`
428
+- Azure China (21Vianet): `cloud: china`
429
+
430
+Ensure the service principal is registered in the correct cloud tenant.
431
+
432
+
433
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_gateway.md
new
+435
@@ -0,0 +1,435 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_gateway.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure ExpressRoute Gateway"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'expressroute', 'gateway', 'networking', 'hybrid']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure ExpressRoute Gateway
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor ExpressRoute gateways including bits and packets per second for ingress and egress, connection counts, CPU utilization, active flow counts, and gateway scale unit counts.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.express_route_gateway.connection_throughput | in, out | bits/s |
88
+| azure_monitor.express_route_gateway.gateway_throughput | average | bits/s |
89
+| azure_monitor.express_route_gateway.cpu_utilization | average | percentage |
90
+| azure_monitor.express_route_gateway.packets | average | packets/s |
91
+| azure_monitor.express_route_gateway.active_flows | average | flows |
92
+| azure_monitor.express_route_gateway.routes_advertised | maximum | routes |
93
+| azure_monitor.express_route_gateway.routes_learned | maximum | routes |
94
+| azure_monitor.express_route_gateway.route_changes | total | changes/s |
95
+| azure_monitor.express_route_gateway.max_flows_creation_rate | maximum | flows/s |
96
+| azure_monitor.express_route_gateway.vm_count | maximum | VMs |
97
+
98
+
99
+
100
+## Alerts
101
+
102
+
103
+The following alerts are available:
104
+
105
+| Alert name | On metric | Description |
106
+|:------------|:----------|:------------|
107
+| [ am_express_route_gateway_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_gateway.conf) | azure_monitor.express_route_gateway.cpu_utilization | ExpressRoute GW CPU on ${label:resource_name} |
108
+| [ am_express_route_gateway_active_flows ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_gateway.conf) | azure_monitor.express_route_gateway.active_flows | ExpressRoute GW active flows on ${label:resource_name} |
109
+| [ am_express_route_gateway_route_changes ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_gateway.conf) | azure_monitor.express_route_gateway.route_changes | ExpressRoute GW route churn on ${label:resource_name} |
110
+| [ am_express_route_gateway_routes_advertised ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_gateway.conf) | azure_monitor.express_route_gateway.routes_advertised | ExpressRoute GW routes advertised on ${label:resource_name} |
111
+| [ am_express_route_gateway_routes_learned ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_express_route_gateway.conf) | azure_monitor.express_route_gateway.routes_learned | ExpressRoute GW routes learned on ${label:resource_name} |
112
+
113
+
114
+## Setup
115
+
116
+
117
+You can configure the **azure_monitor** collector in two ways:
118
+
119
+| Method | Best for | How to |
120
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
121
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
122
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
123
+
124
+:::important
125
+
126
+UI configuration requires paid Netdata Cloud plan.
127
+
128
+:::
129
+
130
+
131
+### Prerequisites
132
+
133
+#### Create an Azure monitoring principal
134
+
135
+Create a service principal or use a managed identity with the following permissions:
136
+
137
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
138
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
139
+
140
+For service principal authentication:
141
+```bash
142
+# Create the service principal
143
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
144
+ --scopes /subscriptions/<subscription-id>
145
+
146
+# Note the appId (client_id), password (client_secret), and tenant
147
+```
148
+
149
+For managed identity (on Azure VMs, VMSS, or AKS):
150
+```bash
151
+# Assign Monitoring Reader role to the VM's managed identity
152
+az role assignment create --assignee <managed-identity-principal-id> \
153
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
154
+```
155
+
156
+
157
+
158
+### Configuration
159
+
160
+#### Options
161
+
162
+The following options can be defined globally: update_every, autodetection_retry.
163
+
164
+Profile files are loaded from:
165
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
166
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
167
+
168
+User profile files with the same filename override stock profiles.
169
+
170
+
171
+<details open><summary>Config options</summary>
172
+
173
+
174
+
175
+| Group | Option | Description | Default | Required |
176
+|:------|:-----|:------------|:--------|:---------:|
177
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
179
+| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
181
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
182
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
184
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
185
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
186
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
187
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
188
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
189
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
190
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
191
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
192
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
193
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
194
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
195
+
196
+
197
+</details>
198
+
199
+
200
+#### via UI
201
+
202
+Configure the **azure_monitor** collector from the Netdata web interface:
203
+
204
+1. Go to **Nodes**.
205
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
206
+3. The **Collectors → Jobs** view opens by default.
207
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
208
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
209
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
210
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
211
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
212
+
213
+
214
+#### via File
215
+
216
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
217
+
218
+The file format is YAML. Generally, the structure is:
219
+
220
+```yaml
221
+update_every: 1
222
+autodetection_retry: 0
223
+jobs:
224
+ - name: some_name1
225
+ - name: some_name2
226
+```
227
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
228
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
229
+
230
+```bash
231
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
232
+sudo ./edit-config go.d/azure_monitor.conf
233
+```
234
+
235
+##### Examples
236
+
237
+###### Service principal (auto-discover all resources)
238
+
239
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
240
+
241
+```yaml
242
+jobs:
243
+ - name: prod
244
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
245
+ auth:
246
+ mode: service_principal
247
+ mode_service_principal:
248
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
249
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
250
+ client_secret: "your-client-secret"
251
+
252
+```
253
+###### Managed identity (Azure VM/VMSS/AKS)
254
+
255
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
256
+
257
+<details open><summary>Config</summary>
258
+
259
+```yaml
260
+jobs:
261
+ - name: prod
262
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
263
+ auth:
264
+ mode: managed_identity
265
+
266
+```
267
+</details>
268
+
269
+###### Specific profiles only
270
+
271
+Monitor only specific Azure services instead of auto-discovering all resource types.
272
+
273
+<details open><summary>Config</summary>
274
+
275
+```yaml
276
+jobs:
277
+ - name: databases
278
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
279
+ profiles:
280
+ - sql_database
281
+ - postgres_flexible
282
+ - redis_cache
283
+ auth:
284
+ mode: service_principal
285
+ mode_service_principal:
286
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
287
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
288
+ client_secret: "your-client-secret"
289
+
290
+```
291
+</details>
292
+
293
+###### Filter by resource group
294
+
295
+Only monitor resources in specific resource groups.
296
+
297
+<details open><summary>Config</summary>
298
+
299
+```yaml
300
+jobs:
301
+ - name: prod-rg
302
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
303
+ resource_groups:
304
+ - production-rg
305
+ - staging-rg
306
+ auth:
307
+ mode: default
308
+
309
+```
310
+</details>
311
+
312
+###### Azure Government cloud
313
+
314
+Connect to Azure Government cloud environment.
315
+
316
+<details open><summary>Config</summary>
317
+
318
+```yaml
319
+jobs:
320
+ - name: gov
321
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ cloud: government
323
+ auth:
324
+ mode: service_principal
325
+ mode_service_principal:
326
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
327
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
328
+ client_secret: "your-client-secret"
329
+
330
+```
331
+</details>
332
+
333
+
334
+
335
+## Troubleshooting
336
+
337
+### Debug Mode
338
+
339
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
340
+
341
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
342
+should give you clues as to why the collector isn't working.
343
+
344
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
345
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
346
+
347
+ ```bash
348
+ cd /usr/libexec/netdata/plugins.d/
349
+ ```
350
+
351
+- Switch to the `netdata` user.
352
+
353
+ ```bash
354
+ sudo -u netdata -s
355
+ ```
356
+
357
+- Run the `go.d.plugin` to debug the collector:
358
+
359
+ ```bash
360
+ ./go.d.plugin -d -m azure_monitor
361
+ ```
362
+
363
+ To debug a specific job:
364
+
365
+ ```bash
366
+ ./go.d.plugin -d -m azure_monitor -j jobName
367
+ ```
368
+
369
+### Getting Logs
370
+
371
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
372
+
373
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
374
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
375
+
376
+#### System with systemd
377
+
378
+Use the following command to view logs generated since the last Netdata service restart:
379
+
380
+```bash
381
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
382
+```
383
+
384
+#### System without systemd
385
+
386
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
387
+
388
+```bash
389
+grep azure_monitor /var/log/netdata/collector.log
390
+```
391
+
392
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
393
+
394
+#### Docker Container
395
+
396
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
397
+
398
+```bash
399
+docker logs netdata 2>&1 | grep azure_monitor
400
+```
401
+
402
+### No metrics are collected
403
+
404
+Verify the following:
405
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
406
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
407
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
408
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
409
+
410
+
411
+### Missing metrics for some resource types
412
+
413
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
414
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
415
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
416
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
417
+
418
+
419
+### Metrics appear delayed
420
+
421
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
422
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
423
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
424
+
425
+
426
+### Authentication errors in sovereign clouds
427
+
428
+For Azure Government or Azure China clouds, set the `cloud` parameter:
429
+- Azure Government: `cloud: government`
430
+- Azure China (21Vianet): `cloud: china`
431
+
432
+Ensure the service principal is registered in the correct cloud tenant.
433
+
434
+
435
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_firewall.md
new
+430
@@ -0,0 +1,430 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_firewall.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Firewall"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'firewall', 'network', 'security', 'nsg']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Firewall
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Firewall including data processed, throughput, application and network rule hit counts, SNAT port utilization, health state percentage, and latency probes.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.firewall.health | average | percentage |
88
+| azure_monitor.firewall.latency | average | milliseconds |
89
+| azure_monitor.firewall.throughput | average | bits/s |
90
+| azure_monitor.firewall.data_processed | total | bytes/s |
91
+| azure_monitor.firewall.rule_hits | application, network | hits/s |
92
+| azure_monitor.firewall.snat_port_utilization | average | percentage |
93
+| azure_monitor.firewall.capacity | average | units |
94
+
95
+
96
+
97
+## Alerts
98
+
99
+
100
+The following alerts are available:
101
+
102
+| Alert name | On metric | Description |
103
+|:------------|:----------|:------------|
104
+| [ am_firewall_health ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_firewall.conf) | azure_monitor.firewall.health | Firewall health on ${label:resource_name} |
105
+| [ am_firewall_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_firewall.conf) | azure_monitor.firewall.latency | Firewall latency on ${label:resource_name} |
106
+| [ am_firewall_snat_port_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_firewall.conf) | azure_monitor.firewall.snat_port_utilization | Firewall SNAT port utilization on ${label:resource_name} |
107
+
108
+
109
+## Setup
110
+
111
+
112
+You can configure the **azure_monitor** collector in two ways:
113
+
114
+| Method | Best for | How to |
115
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
116
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
117
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
118
+
119
+:::important
120
+
121
+UI configuration requires paid Netdata Cloud plan.
122
+
123
+:::
124
+
125
+
126
+### Prerequisites
127
+
128
+#### Create an Azure monitoring principal
129
+
130
+Create a service principal or use a managed identity with the following permissions:
131
+
132
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
133
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
134
+
135
+For service principal authentication:
136
+```bash
137
+# Create the service principal
138
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
139
+ --scopes /subscriptions/<subscription-id>
140
+
141
+# Note the appId (client_id), password (client_secret), and tenant
142
+```
143
+
144
+For managed identity (on Azure VMs, VMSS, or AKS):
145
+```bash
146
+# Assign Monitoring Reader role to the VM's managed identity
147
+az role assignment create --assignee <managed-identity-principal-id> \
148
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+```
150
+
151
+
152
+
153
+### Configuration
154
+
155
+#### Options
156
+
157
+The following options can be defined globally: update_every, autodetection_retry.
158
+
159
+Profile files are loaded from:
160
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
161
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+
163
+User profile files with the same filename override stock profiles.
164
+
165
+
166
+<details open><summary>Config options</summary>
167
+
168
+
169
+
170
+| Group | Option | Description | Default | Required |
171
+|:------|:-----|:------------|:--------|:---------:|
172
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
173
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
174
+| **Target** | subscription_id | Azure subscription ID. | | yes |
175
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
176
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
177
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
179
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
180
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
181
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
182
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
183
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
184
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
190
+
191
+
192
+</details>
193
+
194
+
195
+#### via UI
196
+
197
+Configure the **azure_monitor** collector from the Netdata web interface:
198
+
199
+1. Go to **Nodes**.
200
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
201
+3. The **Collectors → Jobs** view opens by default.
202
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
203
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
204
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
205
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
206
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
207
+
208
+
209
+#### via File
210
+
211
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
212
+
213
+The file format is YAML. Generally, the structure is:
214
+
215
+```yaml
216
+update_every: 1
217
+autodetection_retry: 0
218
+jobs:
219
+ - name: some_name1
220
+ - name: some_name2
221
+```
222
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
223
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
224
+
225
+```bash
226
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
227
+sudo ./edit-config go.d/azure_monitor.conf
228
+```
229
+
230
+##### Examples
231
+
232
+###### Service principal (auto-discover all resources)
233
+
234
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
235
+
236
+```yaml
237
+jobs:
238
+ - name: prod
239
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
240
+ auth:
241
+ mode: service_principal
242
+ mode_service_principal:
243
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
244
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
245
+ client_secret: "your-client-secret"
246
+
247
+```
248
+###### Managed identity (Azure VM/VMSS/AKS)
249
+
250
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
251
+
252
+<details open><summary>Config</summary>
253
+
254
+```yaml
255
+jobs:
256
+ - name: prod
257
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
258
+ auth:
259
+ mode: managed_identity
260
+
261
+```
262
+</details>
263
+
264
+###### Specific profiles only
265
+
266
+Monitor only specific Azure services instead of auto-discovering all resource types.
267
+
268
+<details open><summary>Config</summary>
269
+
270
+```yaml
271
+jobs:
272
+ - name: databases
273
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
274
+ profiles:
275
+ - sql_database
276
+ - postgres_flexible
277
+ - redis_cache
278
+ auth:
279
+ mode: service_principal
280
+ mode_service_principal:
281
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
282
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ client_secret: "your-client-secret"
284
+
285
+```
286
+</details>
287
+
288
+###### Filter by resource group
289
+
290
+Only monitor resources in specific resource groups.
291
+
292
+<details open><summary>Config</summary>
293
+
294
+```yaml
295
+jobs:
296
+ - name: prod-rg
297
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
298
+ resource_groups:
299
+ - production-rg
300
+ - staging-rg
301
+ auth:
302
+ mode: default
303
+
304
+```
305
+</details>
306
+
307
+###### Azure Government cloud
308
+
309
+Connect to Azure Government cloud environment.
310
+
311
+<details open><summary>Config</summary>
312
+
313
+```yaml
314
+jobs:
315
+ - name: gov
316
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
317
+ cloud: government
318
+ auth:
319
+ mode: service_principal
320
+ mode_service_principal:
321
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ client_secret: "your-client-secret"
324
+
325
+```
326
+</details>
327
+
328
+
329
+
330
+## Troubleshooting
331
+
332
+### Debug Mode
333
+
334
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
335
+
336
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
337
+should give you clues as to why the collector isn't working.
338
+
339
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
340
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
341
+
342
+ ```bash
343
+ cd /usr/libexec/netdata/plugins.d/
344
+ ```
345
+
346
+- Switch to the `netdata` user.
347
+
348
+ ```bash
349
+ sudo -u netdata -s
350
+ ```
351
+
352
+- Run the `go.d.plugin` to debug the collector:
353
+
354
+ ```bash
355
+ ./go.d.plugin -d -m azure_monitor
356
+ ```
357
+
358
+ To debug a specific job:
359
+
360
+ ```bash
361
+ ./go.d.plugin -d -m azure_monitor -j jobName
362
+ ```
363
+
364
+### Getting Logs
365
+
366
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
367
+
368
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
369
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
370
+
371
+#### System with systemd
372
+
373
+Use the following command to view logs generated since the last Netdata service restart:
374
+
375
+```bash
376
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
377
+```
378
+
379
+#### System without systemd
380
+
381
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
382
+
383
+```bash
384
+grep azure_monitor /var/log/netdata/collector.log
385
+```
386
+
387
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
388
+
389
+#### Docker Container
390
+
391
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
392
+
393
+```bash
394
+docker logs netdata 2>&1 | grep azure_monitor
395
+```
396
+
397
+### No metrics are collected
398
+
399
+Verify the following:
400
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
401
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
402
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
403
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
404
+
405
+
406
+### Missing metrics for some resource types
407
+
408
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
409
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
410
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
411
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
412
+
413
+
414
+### Metrics appear delayed
415
+
416
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
417
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
418
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
419
+
420
+
421
+### Authentication errors in sovereign clouds
422
+
423
+For Azure Government or Azure China clouds, set the `cloud` parameter:
424
+- Azure Government: `cloud: government`
425
+- Azure China (21Vianet): `cloud: china`
426
+
427
+Ensure the service principal is registered in the correct cloud tenant.
428
+
429
+
430
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_front_door.md
new
+439
@@ -0,0 +1,439 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_front_door.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Front Door"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'front', 'door', 'cdn', 'waf', 'networking']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Front Door
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Front Door including request counts and rates, response sizes, total latency, origin health probe percentages, origin request counts, origin latency, WAF request counts by action and rule, and WebSocket connection metrics.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.front_door.latency | total, origin | milliseconds |
88
+| azure_monitor.front_door.origin_health | health | percentage |
89
+| azure_monitor.front_door.byte_hit_ratio | hit_ratio | percentage |
90
+| azure_monitor.front_door.error_rate | 4xx, 5xx | percentage |
91
+| azure_monitor.front_door.requests | client, origin | requests/s |
92
+| azure_monitor.front_door.data_transfer | request, response | bytes/s |
93
+| azure_monitor.front_door.origin_shield_requests | to_shield, to_origin, rate_limited | requests/s |
94
+| azure_monitor.front_door.origin_shield_data_transfer | request | bytes/s |
95
+| azure_monitor.front_door.waf_requests | total | requests/s |
96
+| azure_monitor.front_door.waf_challenges | captcha, js_challenge | requests/s |
97
+| azure_monitor.front_door.websocket_connections | requested, active | connections/s |
98
+| azure_monitor.front_door.websocket_duration | average | milliseconds |
99
+
100
+
101
+
102
+## Alerts
103
+
104
+
105
+The following alerts are available:
106
+
107
+| Alert name | On metric | Description |
108
+|:------------|:----------|:------------|
109
+| [ am_front_door_origin_health ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.origin_health | Front Door origin health on ${label:resource_name} |
110
+| [ am_front_door_total_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.latency | Front Door total latency on ${label:resource_name} |
111
+| [ am_front_door_origin_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.latency | Front Door origin latency on ${label:resource_name} |
112
+| [ am_front_door_5xx_error_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.error_rate | Front Door 5xx error rate on ${label:resource_name} |
113
+| [ am_front_door_4xx_error_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.error_rate | Front Door 4xx error rate on ${label:resource_name} |
114
+| [ am_front_door_byte_hit_ratio ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.byte_hit_ratio | Front Door cache hit ratio on ${label:resource_name} |
115
+| [ am_front_door_waf_rate_limited ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_front_door.conf) | azure_monitor.front_door.origin_shield_requests | Front Door origin shield rate limiting on ${label:resource_name} |
116
+
117
+
118
+## Setup
119
+
120
+
121
+You can configure the **azure_monitor** collector in two ways:
122
+
123
+| Method | Best for | How to |
124
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
125
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
126
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
127
+
128
+:::important
129
+
130
+UI configuration requires paid Netdata Cloud plan.
131
+
132
+:::
133
+
134
+
135
+### Prerequisites
136
+
137
+#### Create an Azure monitoring principal
138
+
139
+Create a service principal or use a managed identity with the following permissions:
140
+
141
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
142
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
143
+
144
+For service principal authentication:
145
+```bash
146
+# Create the service principal
147
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
148
+ --scopes /subscriptions/<subscription-id>
149
+
150
+# Note the appId (client_id), password (client_secret), and tenant
151
+```
152
+
153
+For managed identity (on Azure VMs, VMSS, or AKS):
154
+```bash
155
+# Assign Monitoring Reader role to the VM's managed identity
156
+az role assignment create --assignee <managed-identity-principal-id> \
157
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
158
+```
159
+
160
+
161
+
162
+### Configuration
163
+
164
+#### Options
165
+
166
+The following options can be defined globally: update_every, autodetection_retry.
167
+
168
+Profile files are loaded from:
169
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
170
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
171
+
172
+User profile files with the same filename override stock profiles.
173
+
174
+
175
+<details open><summary>Config options</summary>
176
+
177
+
178
+
179
+| Group | Option | Description | Default | Required |
180
+|:------|:-----|:------------|:--------|:---------:|
181
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
182
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
183
+| **Target** | subscription_id | Azure subscription ID. | | yes |
184
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
185
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
186
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
187
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
188
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
189
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
190
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
191
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
192
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
193
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
194
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
195
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
196
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
197
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
198
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
199
+
200
+
201
+</details>
202
+
203
+
204
+#### via UI
205
+
206
+Configure the **azure_monitor** collector from the Netdata web interface:
207
+
208
+1. Go to **Nodes**.
209
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
210
+3. The **Collectors → Jobs** view opens by default.
211
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
212
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
213
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
214
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
215
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
216
+
217
+
218
+#### via File
219
+
220
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
221
+
222
+The file format is YAML. Generally, the structure is:
223
+
224
+```yaml
225
+update_every: 1
226
+autodetection_retry: 0
227
+jobs:
228
+ - name: some_name1
229
+ - name: some_name2
230
+```
231
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
232
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
233
+
234
+```bash
235
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
236
+sudo ./edit-config go.d/azure_monitor.conf
237
+```
238
+
239
+##### Examples
240
+
241
+###### Service principal (auto-discover all resources)
242
+
243
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
244
+
245
+```yaml
246
+jobs:
247
+ - name: prod
248
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
249
+ auth:
250
+ mode: service_principal
251
+ mode_service_principal:
252
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
253
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
254
+ client_secret: "your-client-secret"
255
+
256
+```
257
+###### Managed identity (Azure VM/VMSS/AKS)
258
+
259
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
260
+
261
+<details open><summary>Config</summary>
262
+
263
+```yaml
264
+jobs:
265
+ - name: prod
266
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
267
+ auth:
268
+ mode: managed_identity
269
+
270
+```
271
+</details>
272
+
273
+###### Specific profiles only
274
+
275
+Monitor only specific Azure services instead of auto-discovering all resource types.
276
+
277
+<details open><summary>Config</summary>
278
+
279
+```yaml
280
+jobs:
281
+ - name: databases
282
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ profiles:
284
+ - sql_database
285
+ - postgres_flexible
286
+ - redis_cache
287
+ auth:
288
+ mode: service_principal
289
+ mode_service_principal:
290
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
291
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
292
+ client_secret: "your-client-secret"
293
+
294
+```
295
+</details>
296
+
297
+###### Filter by resource group
298
+
299
+Only monitor resources in specific resource groups.
300
+
301
+<details open><summary>Config</summary>
302
+
303
+```yaml
304
+jobs:
305
+ - name: prod-rg
306
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
307
+ resource_groups:
308
+ - production-rg
309
+ - staging-rg
310
+ auth:
311
+ mode: default
312
+
313
+```
314
+</details>
315
+
316
+###### Azure Government cloud
317
+
318
+Connect to Azure Government cloud environment.
319
+
320
+<details open><summary>Config</summary>
321
+
322
+```yaml
323
+jobs:
324
+ - name: gov
325
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ cloud: government
327
+ auth:
328
+ mode: service_principal
329
+ mode_service_principal:
330
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
331
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
332
+ client_secret: "your-client-secret"
333
+
334
+```
335
+</details>
336
+
337
+
338
+
339
+## Troubleshooting
340
+
341
+### Debug Mode
342
+
343
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
344
+
345
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
346
+should give you clues as to why the collector isn't working.
347
+
348
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
349
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
350
+
351
+ ```bash
352
+ cd /usr/libexec/netdata/plugins.d/
353
+ ```
354
+
355
+- Switch to the `netdata` user.
356
+
357
+ ```bash
358
+ sudo -u netdata -s
359
+ ```
360
+
361
+- Run the `go.d.plugin` to debug the collector:
362
+
363
+ ```bash
364
+ ./go.d.plugin -d -m azure_monitor
365
+ ```
366
+
367
+ To debug a specific job:
368
+
369
+ ```bash
370
+ ./go.d.plugin -d -m azure_monitor -j jobName
371
+ ```
372
+
373
+### Getting Logs
374
+
375
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
376
+
377
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
378
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
379
+
380
+#### System with systemd
381
+
382
+Use the following command to view logs generated since the last Netdata service restart:
383
+
384
+```bash
385
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
386
+```
387
+
388
+#### System without systemd
389
+
390
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
391
+
392
+```bash
393
+grep azure_monitor /var/log/netdata/collector.log
394
+```
395
+
396
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
397
+
398
+#### Docker Container
399
+
400
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
401
+
402
+```bash
403
+docker logs netdata 2>&1 | grep azure_monitor
404
+```
405
+
406
+### No metrics are collected
407
+
408
+Verify the following:
409
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
410
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
411
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
412
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
413
+
414
+
415
+### Missing metrics for some resource types
416
+
417
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
418
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
419
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
420
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
421
+
422
+
423
+### Metrics appear delayed
424
+
425
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
426
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
427
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
428
+
429
+
430
+### Authentication errors in sovereign clouds
431
+
432
+For Azure Government or Azure China clouds, set the `cloud` parameter:
433
+- Azure Government: `cloud: government`
434
+- Azure China (21Vianet): `cloud: china`
435
+
436
+Ensure the service principal is registered in the correct cloud tenant.
437
+
438
+
439
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_functions.md
new
+442
@@ -0,0 +1,442 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_functions.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Functions"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'functions', 'serverless', 'faas', 'lambda']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Functions
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Functions execution including function invocation counts, execution units (MB-milliseconds), HTTP request rates and response codes, CPU and memory consumption, and Flex Consumption plan metrics for always-ready and on-demand instances. Uses the same underlying metrics as App Service since Azure Functions runs on the App Service platform.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.app_service.requests | requests | requests/s |
88
+| azure_monitor.app_service.response_time | average | seconds |
89
+| azure_monitor.app_service.http_status | 2xx, 3xx, 4xx, 5xx | responses/s |
90
+| azure_monitor.app_service.http_error_detail | 101_websocket, 401_unauthorized, 403_forbidden, 404_not_found, 406_not_acceptable | responses/s |
91
+| azure_monitor.app_service.cpu | average | percentage |
92
+| azure_monitor.app_service.cpu_time | total | seconds |
93
+| azure_monitor.app_service.memory_usage | average_working_set, working_set, private | bytes |
94
+| azure_monitor.app_service.health | average | percentage |
95
+| azure_monitor.app_service.network_traffic | received, sent | bytes/s |
96
+| azure_monitor.app_service.connections | average | connections |
97
+| azure_monitor.app_service.io_throughput | read, write, other | bytes/s |
98
+| azure_monitor.app_service.io_operations | read, write, other | operations/s |
99
+| azure_monitor.app_service.threads | average | threads |
100
+| azure_monitor.app_service.handles | average | handles |
101
+| azure_monitor.app_service.gc_collections | gen0, gen1, gen2 | collections/s |
102
+| azure_monitor.app_service.function_executions | total | executions/s |
103
+| azure_monitor.app_service.function_execution_units | total | MB-milliseconds/s |
104
+| azure_monitor.app_service.always_ready_function_executions | total | executions/s |
105
+| azure_monitor.app_service.always_ready_function_execution_units | total | MB-milliseconds/s |
106
+| azure_monitor.app_service.always_ready_units | total | units |
107
+| azure_monitor.app_service.on_demand_function_executions | total | executions/s |
108
+| azure_monitor.app_service.on_demand_function_execution_units | total | MB-milliseconds/s |
109
+| azure_monitor.app_service.request_queue | queued | requests |
110
+| azure_monitor.app_service.instances | running | instances |
111
+| azure_monitor.app_service.assemblies | loaded | assemblies |
112
+| azure_monitor.app_service.app_domains | loaded, unloaded | domains |
113
+
114
+
115
+
116
+## Alerts
117
+
118
+There are no alerts configured by default for this integration.
119
+
120
+
121
+## Setup
122
+
123
+
124
+You can configure the **azure_monitor** collector in two ways:
125
+
126
+| Method | Best for | How to |
127
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
128
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
129
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
130
+
131
+:::important
132
+
133
+UI configuration requires paid Netdata Cloud plan.
134
+
135
+:::
136
+
137
+
138
+### Prerequisites
139
+
140
+#### Create an Azure monitoring principal
141
+
142
+Create a service principal or use a managed identity with the following permissions:
143
+
144
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
145
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
146
+
147
+For service principal authentication:
148
+```bash
149
+# Create the service principal
150
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
151
+ --scopes /subscriptions/<subscription-id>
152
+
153
+# Note the appId (client_id), password (client_secret), and tenant
154
+```
155
+
156
+For managed identity (on Azure VMs, VMSS, or AKS):
157
+```bash
158
+# Assign Monitoring Reader role to the VM's managed identity
159
+az role assignment create --assignee <managed-identity-principal-id> \
160
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
161
+```
162
+
163
+
164
+
165
+### Configuration
166
+
167
+#### Options
168
+
169
+The following options can be defined globally: update_every, autodetection_retry.
170
+
171
+Profile files are loaded from:
172
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
173
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
174
+
175
+User profile files with the same filename override stock profiles.
176
+
177
+
178
+<details open><summary>Config options</summary>
179
+
180
+
181
+
182
+| Group | Option | Description | Default | Required |
183
+|:------|:-----|:------------|:--------|:---------:|
184
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
185
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
186
+| **Target** | subscription_id | Azure subscription ID. | | yes |
187
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
188
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
189
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
190
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
191
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
192
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
193
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
194
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
195
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
196
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
197
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
198
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
199
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
200
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
201
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
+
203
+
204
+</details>
205
+
206
+
207
+#### via UI
208
+
209
+Configure the **azure_monitor** collector from the Netdata web interface:
210
+
211
+1. Go to **Nodes**.
212
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
213
+3. The **Collectors → Jobs** view opens by default.
214
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
215
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
216
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
217
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
218
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
219
+
220
+
221
+#### via File
222
+
223
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
224
+
225
+The file format is YAML. Generally, the structure is:
226
+
227
+```yaml
228
+update_every: 1
229
+autodetection_retry: 0
230
+jobs:
231
+ - name: some_name1
232
+ - name: some_name2
233
+```
234
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
235
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
236
+
237
+```bash
238
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
239
+sudo ./edit-config go.d/azure_monitor.conf
240
+```
241
+
242
+##### Examples
243
+
244
+###### Service principal (auto-discover all resources)
245
+
246
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
247
+
248
+```yaml
249
+jobs:
250
+ - name: prod
251
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
252
+ auth:
253
+ mode: service_principal
254
+ mode_service_principal:
255
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
256
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
257
+ client_secret: "your-client-secret"
258
+
259
+```
260
+###### Managed identity (Azure VM/VMSS/AKS)
261
+
262
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
263
+
264
+<details open><summary>Config</summary>
265
+
266
+```yaml
267
+jobs:
268
+ - name: prod
269
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
270
+ auth:
271
+ mode: managed_identity
272
+
273
+```
274
+</details>
275
+
276
+###### Specific profiles only
277
+
278
+Monitor only specific Azure services instead of auto-discovering all resource types.
279
+
280
+<details open><summary>Config</summary>
281
+
282
+```yaml
283
+jobs:
284
+ - name: databases
285
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
286
+ profiles:
287
+ - sql_database
288
+ - postgres_flexible
289
+ - redis_cache
290
+ auth:
291
+ mode: service_principal
292
+ mode_service_principal:
293
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
294
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ client_secret: "your-client-secret"
296
+
297
+```
298
+</details>
299
+
300
+###### Filter by resource group
301
+
302
+Only monitor resources in specific resource groups.
303
+
304
+<details open><summary>Config</summary>
305
+
306
+```yaml
307
+jobs:
308
+ - name: prod-rg
309
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
310
+ resource_groups:
311
+ - production-rg
312
+ - staging-rg
313
+ auth:
314
+ mode: default
315
+
316
+```
317
+</details>
318
+
319
+###### Azure Government cloud
320
+
321
+Connect to Azure Government cloud environment.
322
+
323
+<details open><summary>Config</summary>
324
+
325
+```yaml
326
+jobs:
327
+ - name: gov
328
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
329
+ cloud: government
330
+ auth:
331
+ mode: service_principal
332
+ mode_service_principal:
333
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
334
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
335
+ client_secret: "your-client-secret"
336
+
337
+```
338
+</details>
339
+
340
+
341
+
342
+## Troubleshooting
343
+
344
+### Debug Mode
345
+
346
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
347
+
348
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
349
+should give you clues as to why the collector isn't working.
350
+
351
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
352
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
353
+
354
+ ```bash
355
+ cd /usr/libexec/netdata/plugins.d/
356
+ ```
357
+
358
+- Switch to the `netdata` user.
359
+
360
+ ```bash
361
+ sudo -u netdata -s
362
+ ```
363
+
364
+- Run the `go.d.plugin` to debug the collector:
365
+
366
+ ```bash
367
+ ./go.d.plugin -d -m azure_monitor
368
+ ```
369
+
370
+ To debug a specific job:
371
+
372
+ ```bash
373
+ ./go.d.plugin -d -m azure_monitor -j jobName
374
+ ```
375
+
376
+### Getting Logs
377
+
378
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
379
+
380
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
381
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
382
+
383
+#### System with systemd
384
+
385
+Use the following command to view logs generated since the last Netdata service restart:
386
+
387
+```bash
388
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
389
+```
390
+
391
+#### System without systemd
392
+
393
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
394
+
395
+```bash
396
+grep azure_monitor /var/log/netdata/collector.log
397
+```
398
+
399
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
400
+
401
+#### Docker Container
402
+
403
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
404
+
405
+```bash
406
+docker logs netdata 2>&1 | grep azure_monitor
407
+```
408
+
409
+### No metrics are collected
410
+
411
+Verify the following:
412
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
413
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
414
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
415
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
416
+
417
+
418
+### Missing metrics for some resource types
419
+
420
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
421
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
422
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
423
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
424
+
425
+
426
+### Metrics appear delayed
427
+
428
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
429
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
430
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
431
+
432
+
433
+### Authentication errors in sovereign clouds
434
+
435
+For Azure Government or Azure China clouds, set the `cloud` parameter:
436
+- Azure Government: `cloud: government`
437
+- Azure China (21Vianet): `cloud: china`
438
+
439
+Ensure the service principal is registered in the correct cloud tenant.
440
+
441
+
442
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_iot_hub.md
new
+476
@@ -0,0 +1,476 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_iot_hub.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure IoT Hub"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'iot', 'hub', 'devices', 'telemetry', 'mqtt']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure IoT Hub
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor IoT Hub including device telemetry message rates and quota usage, routing delivery and latency, device twin read and write operations, direct method invocations, cloud-to-device messaging and feedback, job completion rates, device connection and authentication events, and event grid publish status.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.iot_hub.c2d_commands | completed, abandoned, rejected | messages/s |
88
+| azure_monitor.iot_hub.c2d_messages_expired | expired | messages/s |
89
+| azure_monitor.iot_hub.c2d_methods | successful, failed | invocations/s |
90
+| azure_monitor.iot_hub.c2d_methods_request_size | average | bytes |
91
+| azure_monitor.iot_hub.c2d_methods_response_size | average | bytes |
92
+| azure_monitor.iot_hub.c2d_twin_reads | successful, failed | operations/s |
93
+| azure_monitor.iot_hub.c2d_twin_read_size | average | bytes |
94
+| azure_monitor.iot_hub.c2d_twin_updates | successful, failed | operations/s |
95
+| azure_monitor.iot_hub.c2d_twin_update_size | average | bytes |
96
+| azure_monitor.iot_hub.d2c_telemetry | attempted, sent | messages/s |
97
+| azure_monitor.iot_hub.d2c_telemetry_throttle | throttled | errors/s |
98
+| azure_monitor.iot_hub.d2c_twin_reads | successful, failed | operations/s |
99
+| azure_monitor.iot_hub.d2c_twin_read_size | average | bytes |
100
+| azure_monitor.iot_hub.d2c_twin_updates | successful, failed | operations/s |
101
+| azure_monitor.iot_hub.d2c_twin_update_size | average | bytes |
102
+| azure_monitor.iot_hub.routing_deliveries | delivered, dropped, orphaned, invalid, fallback | messages/s |
103
+| azure_monitor.iot_hub.routing_delivery_by_endpoint | builtin_events, event_hubs, service_bus_queues, service_bus_topics, storage | messages/s |
104
+| azure_monitor.iot_hub.routing_storage_blobs | blobs | blobs/s |
105
+| azure_monitor.iot_hub.routing_storage_data | bytes | bytes/s |
106
+| azure_monitor.iot_hub.routing_latency | builtin_events, event_hubs, service_bus_queues, service_bus_topics, storage | milliseconds |
107
+| azure_monitor.iot_hub.routing_deliveries_preview | deliveries | deliveries/s |
108
+| azure_monitor.iot_hub.routing_delivery_latency_preview | average | milliseconds |
109
+| azure_monitor.iot_hub.routing_data_size_preview | bytes | bytes/s |
110
+| azure_monitor.iot_hub.connections | successful | connections/s |
111
+| azure_monitor.iot_hub.connected_devices | connected | devices |
112
+| azure_monitor.iot_hub.total_devices | total | devices |
113
+| azure_monitor.iot_hub.daily_message_quota | used | messages |
114
+| azure_monitor.iot_hub.data_usage | data_usage, data_usage_v2 | bytes/s |
115
+| azure_monitor.iot_hub.configurations | operations | operations/s |
116
+| azure_monitor.iot_hub.event_grid_deliveries | deliveries | deliveries/s |
117
+| azure_monitor.iot_hub.event_grid_latency | average | milliseconds |
118
+| azure_monitor.iot_hub.jobs_status | completed, failed | operations/s |
119
+| azure_monitor.iot_hub.jobs_cancel | successful, failed | operations/s |
120
+| azure_monitor.iot_hub.jobs_create_method | successful, failed | operations/s |
121
+| azure_monitor.iot_hub.jobs_create_twin_update | successful, failed | operations/s |
122
+| azure_monitor.iot_hub.jobs_list | successful, failed | operations/s |
123
+| azure_monitor.iot_hub.jobs_query | successful, failed | operations/s |
124
+| azure_monitor.iot_hub.twin_queries | successful, failed | queries/s |
125
+| azure_monitor.iot_hub.twin_queries_result_size | average | bytes |
126
+
127
+
128
+
129
+## Alerts
130
+
131
+
132
+The following alerts are available:
133
+
134
+| Alert name | On metric | Description |
135
+|:------------|:----------|:------------|
136
+| [ am_iot_hub_d2c_telemetry_throttle ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.d2c_telemetry_throttle | IoT Hub telemetry throttling on ${label:resource_name} |
137
+| [ am_iot_hub_c2d_messages_expired ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.c2d_messages_expired | IoT Hub C2D messages expiring on ${label:resource_name} |
138
+| [ am_iot_hub_c2d_methods_failed ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.c2d_methods | IoT Hub direct method failures on ${label:resource_name} |
139
+| [ am_iot_hub_c2d_twin_read_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.c2d_twin_reads | IoT Hub backend twin read failures on ${label:resource_name} |
140
+| [ am_iot_hub_c2d_twin_update_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.c2d_twin_updates | IoT Hub backend twin update failures on ${label:resource_name} |
141
+| [ am_iot_hub_d2c_twin_read_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.d2c_twin_reads | IoT Hub device twin read failures on ${label:resource_name} |
142
+| [ am_iot_hub_d2c_twin_update_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.d2c_twin_updates | IoT Hub device twin update failures on ${label:resource_name} |
143
+| [ am_iot_hub_routing_dropped ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.routing_deliveries | IoT Hub routing dropped messages on ${label:resource_name} |
144
+| [ am_iot_hub_routing_orphaned ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.routing_deliveries | IoT Hub routing orphaned messages on ${label:resource_name} |
145
+| [ am_iot_hub_routing_invalid ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.routing_deliveries | IoT Hub routing invalid messages on ${label:resource_name} |
146
+| [ am_iot_hub_routing_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.routing_latency | IoT Hub routing latency on ${label:resource_name} |
147
+| [ am_iot_hub_event_grid_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.event_grid_latency | IoT Hub Event Grid latency on ${label:resource_name} |
148
+| [ am_iot_hub_jobs_failed ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.jobs_status | IoT Hub job failures on ${label:resource_name} |
149
+| [ am_iot_hub_twin_query_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.twin_queries | IoT Hub twin query failures on ${label:resource_name} |
150
+| [ am_iot_hub_c2d_commands_abandoned ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.c2d_commands | IoT Hub C2D commands abandoned on ${label:resource_name} |
151
+| [ am_iot_hub_c2d_commands_rejected ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.c2d_commands | IoT Hub C2D commands rejected on ${label:resource_name} |
152
+| [ am_iot_hub_routing_delivery_latency_preview ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_iot_hub.conf) | azure_monitor.iot_hub.routing_delivery_latency_preview | IoT Hub routing delivery latency on ${label:resource_name} |
153
+
154
+
155
+## Setup
156
+
157
+
158
+You can configure the **azure_monitor** collector in two ways:
159
+
160
+| Method | Best for | How to |
161
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
162
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
163
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
164
+
165
+:::important
166
+
167
+UI configuration requires paid Netdata Cloud plan.
168
+
169
+:::
170
+
171
+
172
+### Prerequisites
173
+
174
+#### Create an Azure monitoring principal
175
+
176
+Create a service principal or use a managed identity with the following permissions:
177
+
178
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
179
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
180
+
181
+For service principal authentication:
182
+```bash
183
+# Create the service principal
184
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
185
+ --scopes /subscriptions/<subscription-id>
186
+
187
+# Note the appId (client_id), password (client_secret), and tenant
188
+```
189
+
190
+For managed identity (on Azure VMs, VMSS, or AKS):
191
+```bash
192
+# Assign Monitoring Reader role to the VM's managed identity
193
+az role assignment create --assignee <managed-identity-principal-id> \
194
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
195
+```
196
+
197
+
198
+
199
+### Configuration
200
+
201
+#### Options
202
+
203
+The following options can be defined globally: update_every, autodetection_retry.
204
+
205
+Profile files are loaded from:
206
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
207
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
208
+
209
+User profile files with the same filename override stock profiles.
210
+
211
+
212
+<details open><summary>Config options</summary>
213
+
214
+
215
+
216
+| Group | Option | Description | Default | Required |
217
+|:------|:-----|:------------|:--------|:---------:|
218
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
219
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
220
+| **Target** | subscription_id | Azure subscription ID. | | yes |
221
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
222
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
223
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
224
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
225
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
226
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
227
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
228
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
229
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
230
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
231
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
232
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
233
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
234
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
235
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
236
+
237
+
238
+</details>
239
+
240
+
241
+#### via UI
242
+
243
+Configure the **azure_monitor** collector from the Netdata web interface:
244
+
245
+1. Go to **Nodes**.
246
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
247
+3. The **Collectors → Jobs** view opens by default.
248
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
249
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
250
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
251
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
252
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
253
+
254
+
255
+#### via File
256
+
257
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
258
+
259
+The file format is YAML. Generally, the structure is:
260
+
261
+```yaml
262
+update_every: 1
263
+autodetection_retry: 0
264
+jobs:
265
+ - name: some_name1
266
+ - name: some_name2
267
+```
268
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
269
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
270
+
271
+```bash
272
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
273
+sudo ./edit-config go.d/azure_monitor.conf
274
+```
275
+
276
+##### Examples
277
+
278
+###### Service principal (auto-discover all resources)
279
+
280
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
281
+
282
+```yaml
283
+jobs:
284
+ - name: prod
285
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
286
+ auth:
287
+ mode: service_principal
288
+ mode_service_principal:
289
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
290
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
291
+ client_secret: "your-client-secret"
292
+
293
+```
294
+###### Managed identity (Azure VM/VMSS/AKS)
295
+
296
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
297
+
298
+<details open><summary>Config</summary>
299
+
300
+```yaml
301
+jobs:
302
+ - name: prod
303
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
304
+ auth:
305
+ mode: managed_identity
306
+
307
+```
308
+</details>
309
+
310
+###### Specific profiles only
311
+
312
+Monitor only specific Azure services instead of auto-discovering all resource types.
313
+
314
+<details open><summary>Config</summary>
315
+
316
+```yaml
317
+jobs:
318
+ - name: databases
319
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ profiles:
321
+ - sql_database
322
+ - postgres_flexible
323
+ - redis_cache
324
+ auth:
325
+ mode: service_principal
326
+ mode_service_principal:
327
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
328
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
329
+ client_secret: "your-client-secret"
330
+
331
+```
332
+</details>
333
+
334
+###### Filter by resource group
335
+
336
+Only monitor resources in specific resource groups.
337
+
338
+<details open><summary>Config</summary>
339
+
340
+```yaml
341
+jobs:
342
+ - name: prod-rg
343
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
344
+ resource_groups:
345
+ - production-rg
346
+ - staging-rg
347
+ auth:
348
+ mode: default
349
+
350
+```
351
+</details>
352
+
353
+###### Azure Government cloud
354
+
355
+Connect to Azure Government cloud environment.
356
+
357
+<details open><summary>Config</summary>
358
+
359
+```yaml
360
+jobs:
361
+ - name: gov
362
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
363
+ cloud: government
364
+ auth:
365
+ mode: service_principal
366
+ mode_service_principal:
367
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
368
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
369
+ client_secret: "your-client-secret"
370
+
371
+```
372
+</details>
373
+
374
+
375
+
376
+## Troubleshooting
377
+
378
+### Debug Mode
379
+
380
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
381
+
382
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
383
+should give you clues as to why the collector isn't working.
384
+
385
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
386
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
387
+
388
+ ```bash
389
+ cd /usr/libexec/netdata/plugins.d/
390
+ ```
391
+
392
+- Switch to the `netdata` user.
393
+
394
+ ```bash
395
+ sudo -u netdata -s
396
+ ```
397
+
398
+- Run the `go.d.plugin` to debug the collector:
399
+
400
+ ```bash
401
+ ./go.d.plugin -d -m azure_monitor
402
+ ```
403
+
404
+ To debug a specific job:
405
+
406
+ ```bash
407
+ ./go.d.plugin -d -m azure_monitor -j jobName
408
+ ```
409
+
410
+### Getting Logs
411
+
412
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
413
+
414
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
415
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
416
+
417
+#### System with systemd
418
+
419
+Use the following command to view logs generated since the last Netdata service restart:
420
+
421
+```bash
422
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
423
+```
424
+
425
+#### System without systemd
426
+
427
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
428
+
429
+```bash
430
+grep azure_monitor /var/log/netdata/collector.log
431
+```
432
+
433
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
434
+
435
+#### Docker Container
436
+
437
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
438
+
439
+```bash
440
+docker logs netdata 2>&1 | grep azure_monitor
441
+```
442
+
443
+### No metrics are collected
444
+
445
+Verify the following:
446
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
447
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
448
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
449
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
450
+
451
+
452
+### Missing metrics for some resource types
453
+
454
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
455
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
456
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
457
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
458
+
459
+
460
+### Metrics appear delayed
461
+
462
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
463
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
464
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
465
+
466
+
467
+### Authentication errors in sovereign clouds
468
+
469
+For Azure Government or Azure China clouds, set the `cloud` parameter:
470
+- Azure Government: `cloud: government`
471
+- Azure China (21Vianet): `cloud: china`
472
+
473
+Ensure the service principal is registered in the correct cloud tenant.
474
+
475
+
476
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_key_vault.md
new
+427
@@ -0,0 +1,427 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_key_vault.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Key Vault"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'key', 'vault', 'secrets', 'certificates', 'security']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Key Vault
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Key Vault including overall vault availability, API saturation approaching service limits, and service API hit and latency metrics.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.key_vault.availability | average | percentage |
88
+| azure_monitor.key_vault.api_latency | average | milliseconds |
89
+| azure_monitor.key_vault.saturation | average | percentage |
90
+| azure_monitor.key_vault.api_activity | hits, results | operations/s |
91
+
92
+
93
+
94
+## Alerts
95
+
96
+
97
+The following alerts are available:
98
+
99
+| Alert name | On metric | Description |
100
+|:------------|:----------|:------------|
101
+| [ am_key_vault_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_key_vault.conf) | azure_monitor.key_vault.availability | Key Vault availability on ${label:resource_name} |
102
+| [ am_key_vault_api_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_key_vault.conf) | azure_monitor.key_vault.api_latency | Key Vault API latency on ${label:resource_name} |
103
+| [ am_key_vault_saturation ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_key_vault.conf) | azure_monitor.key_vault.saturation | Key Vault saturation on ${label:resource_name} |
104
+
105
+
106
+## Setup
107
+
108
+
109
+You can configure the **azure_monitor** collector in two ways:
110
+
111
+| Method | Best for | How to |
112
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
113
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
114
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
115
+
116
+:::important
117
+
118
+UI configuration requires paid Netdata Cloud plan.
119
+
120
+:::
121
+
122
+
123
+### Prerequisites
124
+
125
+#### Create an Azure monitoring principal
126
+
127
+Create a service principal or use a managed identity with the following permissions:
128
+
129
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
130
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
131
+
132
+For service principal authentication:
133
+```bash
134
+# Create the service principal
135
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
136
+ --scopes /subscriptions/<subscription-id>
137
+
138
+# Note the appId (client_id), password (client_secret), and tenant
139
+```
140
+
141
+For managed identity (on Azure VMs, VMSS, or AKS):
142
+```bash
143
+# Assign Monitoring Reader role to the VM's managed identity
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+```
147
+
148
+
149
+
150
+### Configuration
151
+
152
+#### Options
153
+
154
+The following options can be defined globally: update_every, autodetection_retry.
155
+
156
+Profile files are loaded from:
157
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
158
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+
160
+User profile files with the same filename override stock profiles.
161
+
162
+
163
+<details open><summary>Config options</summary>
164
+
165
+
166
+
167
+| Group | Option | Description | Default | Required |
168
+|:------|:-----|:------------|:--------|:---------:|
169
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
170
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
171
+| **Target** | subscription_id | Azure subscription ID. | | yes |
172
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
173
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
174
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
175
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
176
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
177
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
178
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
179
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
180
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
181
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
187
+
188
+
189
+</details>
190
+
191
+
192
+#### via UI
193
+
194
+Configure the **azure_monitor** collector from the Netdata web interface:
195
+
196
+1. Go to **Nodes**.
197
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
198
+3. The **Collectors → Jobs** view opens by default.
199
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
200
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
201
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
202
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
203
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
204
+
205
+
206
+#### via File
207
+
208
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
209
+
210
+The file format is YAML. Generally, the structure is:
211
+
212
+```yaml
213
+update_every: 1
214
+autodetection_retry: 0
215
+jobs:
216
+ - name: some_name1
217
+ - name: some_name2
218
+```
219
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
220
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
221
+
222
+```bash
223
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
224
+sudo ./edit-config go.d/azure_monitor.conf
225
+```
226
+
227
+##### Examples
228
+
229
+###### Service principal (auto-discover all resources)
230
+
231
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
232
+
233
+```yaml
234
+jobs:
235
+ - name: prod
236
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
237
+ auth:
238
+ mode: service_principal
239
+ mode_service_principal:
240
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
241
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ client_secret: "your-client-secret"
243
+
244
+```
245
+###### Managed identity (Azure VM/VMSS/AKS)
246
+
247
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
248
+
249
+<details open><summary>Config</summary>
250
+
251
+```yaml
252
+jobs:
253
+ - name: prod
254
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
255
+ auth:
256
+ mode: managed_identity
257
+
258
+```
259
+</details>
260
+
261
+###### Specific profiles only
262
+
263
+Monitor only specific Azure services instead of auto-discovering all resource types.
264
+
265
+<details open><summary>Config</summary>
266
+
267
+```yaml
268
+jobs:
269
+ - name: databases
270
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
271
+ profiles:
272
+ - sql_database
273
+ - postgres_flexible
274
+ - redis_cache
275
+ auth:
276
+ mode: service_principal
277
+ mode_service_principal:
278
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
279
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
280
+ client_secret: "your-client-secret"
281
+
282
+```
283
+</details>
284
+
285
+###### Filter by resource group
286
+
287
+Only monitor resources in specific resource groups.
288
+
289
+<details open><summary>Config</summary>
290
+
291
+```yaml
292
+jobs:
293
+ - name: prod-rg
294
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ resource_groups:
296
+ - production-rg
297
+ - staging-rg
298
+ auth:
299
+ mode: default
300
+
301
+```
302
+</details>
303
+
304
+###### Azure Government cloud
305
+
306
+Connect to Azure Government cloud environment.
307
+
308
+<details open><summary>Config</summary>
309
+
310
+```yaml
311
+jobs:
312
+ - name: gov
313
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
314
+ cloud: government
315
+ auth:
316
+ mode: service_principal
317
+ mode_service_principal:
318
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
319
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ client_secret: "your-client-secret"
321
+
322
+```
323
+</details>
324
+
325
+
326
+
327
+## Troubleshooting
328
+
329
+### Debug Mode
330
+
331
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
332
+
333
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
334
+should give you clues as to why the collector isn't working.
335
+
336
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
337
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
338
+
339
+ ```bash
340
+ cd /usr/libexec/netdata/plugins.d/
341
+ ```
342
+
343
+- Switch to the `netdata` user.
344
+
345
+ ```bash
346
+ sudo -u netdata -s
347
+ ```
348
+
349
+- Run the `go.d.plugin` to debug the collector:
350
+
351
+ ```bash
352
+ ./go.d.plugin -d -m azure_monitor
353
+ ```
354
+
355
+ To debug a specific job:
356
+
357
+ ```bash
358
+ ./go.d.plugin -d -m azure_monitor -j jobName
359
+ ```
360
+
361
+### Getting Logs
362
+
363
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
364
+
365
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
366
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
367
+
368
+#### System with systemd
369
+
370
+Use the following command to view logs generated since the last Netdata service restart:
371
+
372
+```bash
373
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
374
+```
375
+
376
+#### System without systemd
377
+
378
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
379
+
380
+```bash
381
+grep azure_monitor /var/log/netdata/collector.log
382
+```
383
+
384
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
385
+
386
+#### Docker Container
387
+
388
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
389
+
390
+```bash
391
+docker logs netdata 2>&1 | grep azure_monitor
392
+```
393
+
394
+### No metrics are collected
395
+
396
+Verify the following:
397
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
398
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
399
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
400
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
401
+
402
+
403
+### Missing metrics for some resource types
404
+
405
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
406
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
407
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
408
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
409
+
410
+
411
+### Metrics appear delayed
412
+
413
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
414
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
415
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
416
+
417
+
418
+### Authentication errors in sovereign clouds
419
+
420
+For Azure Government or Azure China clouds, set the `cloud` parameter:
421
+- Azure Government: `cloud: government`
422
+- Azure China (21Vianet): `cloud: china`
423
+
424
+Ensure the service principal is registered in the correct cloud tenant.
425
+
426
+
427
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_kubernetes_service_cluster.md
new
+455
@@ -0,0 +1,455 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_kubernetes_service_cluster.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Kubernetes Service Cluster"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'kubernetes', 'aks', 'k8s', 'containers', 'cluster']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Kubernetes Service Cluster
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor AKS cluster health including API server and etcd resource usage, pod scheduling status and readiness, node capacity and conditions, cluster autoscaler behavior, and per-node CPU, memory, disk, and network utilization.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.aks.apiserver_cpu | average, maximum | percentage |
88
+| azure_monitor.aks.apiserver_memory | average, maximum | percentage |
89
+| azure_monitor.aks.apiserver_inflight_requests | average | requests |
90
+| azure_monitor.aks.etcd_cpu | average, maximum | percentage |
91
+| azure_monitor.aks.etcd_memory | average, maximum | percentage |
92
+| azure_monitor.aks.etcd_database | average, maximum | percentage |
93
+| azure_monitor.aks.pod_status_phase | average | pods |
94
+| azure_monitor.aks.pod_status_ready | average | pods |
95
+| azure_monitor.aks.allocatable_cpu | average | cores |
96
+| azure_monitor.aks.allocatable_memory | average | bytes |
97
+| azure_monitor.aks.node_conditions | average | nodes |
98
+| azure_monitor.aks.autoscaler_health | safe_to_autoscale, cooldown | state |
99
+| azure_monitor.aks.autoscaler_unschedulable_pods | average | pods |
100
+| azure_monitor.aks.autoscaler_unneeded_nodes | average | nodes |
101
+| azure_monitor.aks.node_cpu_millicores | average, maximum | millicores |
102
+| azure_monitor.aks.node_cpu_percentage | average, maximum | percentage |
103
+| azure_monitor.aks.node_memory_working_set | average, maximum | bytes |
104
+| azure_monitor.aks.node_memory_working_set_percentage | average, maximum | percentage |
105
+| azure_monitor.aks.node_memory_rss | average, maximum | bytes |
106
+| azure_monitor.aks.node_memory_rss_percentage | average, maximum | percentage |
107
+| azure_monitor.aks.node_disk_usage | average, maximum | bytes |
108
+| azure_monitor.aks.node_disk_percentage | average, maximum | percentage |
109
+| azure_monitor.aks.node_network | in, out | bytes |
110
+
111
+
112
+
113
+## Alerts
114
+
115
+
116
+The following alerts are available:
117
+
118
+| Alert name | On metric | Description |
119
+|:------------|:----------|:------------|
120
+| [ am_aks_apiserver_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.apiserver_cpu | AKS API server CPU on ${label:resource_name} |
121
+| [ am_aks_apiserver_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.apiserver_memory | AKS API server memory on ${label:resource_name} |
122
+| [ am_aks_apiserver_inflight_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.apiserver_inflight_requests | AKS API server inflight requests on ${label:resource_name} |
123
+| [ am_aks_etcd_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.etcd_cpu | AKS etcd CPU on ${label:resource_name} |
124
+| [ am_aks_etcd_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.etcd_memory | AKS etcd memory on ${label:resource_name} |
125
+| [ am_aks_etcd_database ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.etcd_database | AKS etcd database usage on ${label:resource_name} |
126
+| [ am_aks_autoscaler_safe_to_autoscale ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.autoscaler_health | AKS autoscaler unsafe on ${label:resource_name} |
127
+| [ am_aks_autoscaler_unschedulable_pods ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.autoscaler_unschedulable_pods | AKS unschedulable pods on ${label:resource_name} |
128
+| [ am_aks_node_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.node_cpu_percentage | AKS node CPU on ${label:resource_name} |
129
+| [ am_aks_node_memory_working_set ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.node_memory_working_set_percentage | AKS node memory working set on ${label:resource_name} |
130
+| [ am_aks_node_memory_rss ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.node_memory_rss_percentage | AKS node memory RSS on ${label:resource_name} |
131
+| [ am_aks_node_disk ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_aks.conf) | azure_monitor.aks.node_disk_percentage | AKS node disk usage on ${label:resource_name} |
132
+
133
+
134
+## Setup
135
+
136
+
137
+You can configure the **azure_monitor** collector in two ways:
138
+
139
+| Method | Best for | How to |
140
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
141
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
142
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
143
+
144
+:::important
145
+
146
+UI configuration requires paid Netdata Cloud plan.
147
+
148
+:::
149
+
150
+
151
+### Prerequisites
152
+
153
+#### Create an Azure monitoring principal
154
+
155
+Create a service principal or use a managed identity with the following permissions:
156
+
157
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
158
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
159
+
160
+For service principal authentication:
161
+```bash
162
+# Create the service principal
163
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
164
+ --scopes /subscriptions/<subscription-id>
165
+
166
+# Note the appId (client_id), password (client_secret), and tenant
167
+```
168
+
169
+For managed identity (on Azure VMs, VMSS, or AKS):
170
+```bash
171
+# Assign Monitoring Reader role to the VM's managed identity
172
+az role assignment create --assignee <managed-identity-principal-id> \
173
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
174
+```
175
+
176
+
177
+
178
+### Configuration
179
+
180
+#### Options
181
+
182
+The following options can be defined globally: update_every, autodetection_retry.
183
+
184
+Profile files are loaded from:
185
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
186
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
187
+
188
+User profile files with the same filename override stock profiles.
189
+
190
+
191
+<details open><summary>Config options</summary>
192
+
193
+
194
+
195
+| Group | Option | Description | Default | Required |
196
+|:------|:-----|:------------|:--------|:---------:|
197
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
198
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
199
+| **Target** | subscription_id | Azure subscription ID. | | yes |
200
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
201
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
202
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
203
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
204
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
205
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
206
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
207
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
208
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
209
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
210
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
211
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
212
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
213
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
214
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
215
+
216
+
217
+</details>
218
+
219
+
220
+#### via UI
221
+
222
+Configure the **azure_monitor** collector from the Netdata web interface:
223
+
224
+1. Go to **Nodes**.
225
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
226
+3. The **Collectors → Jobs** view opens by default.
227
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
228
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
229
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
230
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
231
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
232
+
233
+
234
+#### via File
235
+
236
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
237
+
238
+The file format is YAML. Generally, the structure is:
239
+
240
+```yaml
241
+update_every: 1
242
+autodetection_retry: 0
243
+jobs:
244
+ - name: some_name1
245
+ - name: some_name2
246
+```
247
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
248
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
249
+
250
+```bash
251
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
252
+sudo ./edit-config go.d/azure_monitor.conf
253
+```
254
+
255
+##### Examples
256
+
257
+###### Service principal (auto-discover all resources)
258
+
259
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
260
+
261
+```yaml
262
+jobs:
263
+ - name: prod
264
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
265
+ auth:
266
+ mode: service_principal
267
+ mode_service_principal:
268
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
269
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
270
+ client_secret: "your-client-secret"
271
+
272
+```
273
+###### Managed identity (Azure VM/VMSS/AKS)
274
+
275
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
276
+
277
+<details open><summary>Config</summary>
278
+
279
+```yaml
280
+jobs:
281
+ - name: prod
282
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ auth:
284
+ mode: managed_identity
285
+
286
+```
287
+</details>
288
+
289
+###### Specific profiles only
290
+
291
+Monitor only specific Azure services instead of auto-discovering all resource types.
292
+
293
+<details open><summary>Config</summary>
294
+
295
+```yaml
296
+jobs:
297
+ - name: databases
298
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ profiles:
300
+ - sql_database
301
+ - postgres_flexible
302
+ - redis_cache
303
+ auth:
304
+ mode: service_principal
305
+ mode_service_principal:
306
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
307
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
308
+ client_secret: "your-client-secret"
309
+
310
+```
311
+</details>
312
+
313
+###### Filter by resource group
314
+
315
+Only monitor resources in specific resource groups.
316
+
317
+<details open><summary>Config</summary>
318
+
319
+```yaml
320
+jobs:
321
+ - name: prod-rg
322
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ resource_groups:
324
+ - production-rg
325
+ - staging-rg
326
+ auth:
327
+ mode: default
328
+
329
+```
330
+</details>
331
+
332
+###### Azure Government cloud
333
+
334
+Connect to Azure Government cloud environment.
335
+
336
+<details open><summary>Config</summary>
337
+
338
+```yaml
339
+jobs:
340
+ - name: gov
341
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
342
+ cloud: government
343
+ auth:
344
+ mode: service_principal
345
+ mode_service_principal:
346
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
347
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
348
+ client_secret: "your-client-secret"
349
+
350
+```
351
+</details>
352
+
353
+
354
+
355
+## Troubleshooting
356
+
357
+### Debug Mode
358
+
359
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
360
+
361
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
362
+should give you clues as to why the collector isn't working.
363
+
364
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
365
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
366
+
367
+ ```bash
368
+ cd /usr/libexec/netdata/plugins.d/
369
+ ```
370
+
371
+- Switch to the `netdata` user.
372
+
373
+ ```bash
374
+ sudo -u netdata -s
375
+ ```
376
+
377
+- Run the `go.d.plugin` to debug the collector:
378
+
379
+ ```bash
380
+ ./go.d.plugin -d -m azure_monitor
381
+ ```
382
+
383
+ To debug a specific job:
384
+
385
+ ```bash
386
+ ./go.d.plugin -d -m azure_monitor -j jobName
387
+ ```
388
+
389
+### Getting Logs
390
+
391
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
392
+
393
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
394
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
395
+
396
+#### System with systemd
397
+
398
+Use the following command to view logs generated since the last Netdata service restart:
399
+
400
+```bash
401
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
402
+```
403
+
404
+#### System without systemd
405
+
406
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
407
+
408
+```bash
409
+grep azure_monitor /var/log/netdata/collector.log
410
+```
411
+
412
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
413
+
414
+#### Docker Container
415
+
416
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
417
+
418
+```bash
419
+docker logs netdata 2>&1 | grep azure_monitor
420
+```
421
+
422
+### No metrics are collected
423
+
424
+Verify the following:
425
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
426
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
427
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
428
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
429
+
430
+
431
+### Missing metrics for some resource types
432
+
433
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
434
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
435
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
436
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
437
+
438
+
439
+### Metrics appear delayed
440
+
441
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
442
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
443
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
444
+
445
+
446
+### Authentication errors in sovereign clouds
447
+
448
+For Azure Government or Azure China clouds, set the `cloud` parameter:
449
+- Azure Government: `cloud: government`
450
+- Azure China (21Vianet): `cloud: china`
451
+
452
+Ensure the service principal is registered in the correct cloud tenant.
453
+
454
+
455
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_load_balancer.md
new
+432
@@ -0,0 +1,432 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_load_balancer.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Load Balancer"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'load', 'balancer', 'networking', 'lb']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Load Balancer
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Load Balancer health and throughput including data path and health probe availability, SYN and SNAT connection counts, byte and packet throughput, allocated and used SNAT ports, and connection attempt rates.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.load_balancers.vip_availability | average | percentage |
88
+| azure_monitor.load_balancers.dip_availability | average | percentage |
89
+| azure_monitor.load_balancers.global_backend_availability | average | percentage |
90
+| azure_monitor.load_balancers.byte_throughput | total | bytes/s |
91
+| azure_monitor.load_balancers.packet_throughput | total | packets/s |
92
+| azure_monitor.load_balancers.syn_count | total | packets/s |
93
+| azure_monitor.load_balancers.snat_connections | total | connections/s |
94
+| azure_monitor.load_balancers.snat_ports | allocated, used | ports |
95
+
96
+
97
+
98
+## Alerts
99
+
100
+
101
+The following alerts are available:
102
+
103
+| Alert name | On metric | Description |
104
+|:------------|:----------|:------------|
105
+| [ am_load_balancers_vip_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_load_balancers.conf) | azure_monitor.load_balancers.vip_availability | LB data path availability on ${label:resource_name} |
106
+| [ am_load_balancers_dip_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_load_balancers.conf) | azure_monitor.load_balancers.dip_availability | LB health probe status on ${label:resource_name} |
107
+| [ am_load_balancers_global_backend_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_load_balancers.conf) | azure_monitor.load_balancers.global_backend_availability | LB global backend availability on ${label:resource_name} |
108
+| [ am_load_balancers_snat_port_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_load_balancers.conf) | azure_monitor.load_balancers.snat_ports | LB SNAT port utilization on ${label:resource_name} |
109
+
110
+
111
+## Setup
112
+
113
+
114
+You can configure the **azure_monitor** collector in two ways:
115
+
116
+| Method | Best for | How to |
117
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
118
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
119
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
120
+
121
+:::important
122
+
123
+UI configuration requires paid Netdata Cloud plan.
124
+
125
+:::
126
+
127
+
128
+### Prerequisites
129
+
130
+#### Create an Azure monitoring principal
131
+
132
+Create a service principal or use a managed identity with the following permissions:
133
+
134
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
135
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
136
+
137
+For service principal authentication:
138
+```bash
139
+# Create the service principal
140
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
141
+ --scopes /subscriptions/<subscription-id>
142
+
143
+# Note the appId (client_id), password (client_secret), and tenant
144
+```
145
+
146
+For managed identity (on Azure VMs, VMSS, or AKS):
147
+```bash
148
+# Assign Monitoring Reader role to the VM's managed identity
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
151
+```
152
+
153
+
154
+
155
+### Configuration
156
+
157
+#### Options
158
+
159
+The following options can be defined globally: update_every, autodetection_retry.
160
+
161
+Profile files are loaded from:
162
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
163
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
164
+
165
+User profile files with the same filename override stock profiles.
166
+
167
+
168
+<details open><summary>Config options</summary>
169
+
170
+
171
+
172
+| Group | Option | Description | Default | Required |
173
+|:------|:-----|:------------|:--------|:---------:|
174
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
176
+| **Target** | subscription_id | Azure subscription ID. | | yes |
177
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
178
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
179
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
181
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
182
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
183
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
184
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
185
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
186
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
187
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
188
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
189
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
190
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
191
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
192
+
193
+
194
+</details>
195
+
196
+
197
+#### via UI
198
+
199
+Configure the **azure_monitor** collector from the Netdata web interface:
200
+
201
+1. Go to **Nodes**.
202
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
203
+3. The **Collectors → Jobs** view opens by default.
204
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
205
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
206
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
207
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
208
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
209
+
210
+
211
+#### via File
212
+
213
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
214
+
215
+The file format is YAML. Generally, the structure is:
216
+
217
+```yaml
218
+update_every: 1
219
+autodetection_retry: 0
220
+jobs:
221
+ - name: some_name1
222
+ - name: some_name2
223
+```
224
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
225
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
226
+
227
+```bash
228
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
229
+sudo ./edit-config go.d/azure_monitor.conf
230
+```
231
+
232
+##### Examples
233
+
234
+###### Service principal (auto-discover all resources)
235
+
236
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
237
+
238
+```yaml
239
+jobs:
240
+ - name: prod
241
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ auth:
243
+ mode: service_principal
244
+ mode_service_principal:
245
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
246
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
247
+ client_secret: "your-client-secret"
248
+
249
+```
250
+###### Managed identity (Azure VM/VMSS/AKS)
251
+
252
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
253
+
254
+<details open><summary>Config</summary>
255
+
256
+```yaml
257
+jobs:
258
+ - name: prod
259
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
260
+ auth:
261
+ mode: managed_identity
262
+
263
+```
264
+</details>
265
+
266
+###### Specific profiles only
267
+
268
+Monitor only specific Azure services instead of auto-discovering all resource types.
269
+
270
+<details open><summary>Config</summary>
271
+
272
+```yaml
273
+jobs:
274
+ - name: databases
275
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
276
+ profiles:
277
+ - sql_database
278
+ - postgres_flexible
279
+ - redis_cache
280
+ auth:
281
+ mode: service_principal
282
+ mode_service_principal:
283
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
284
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
285
+ client_secret: "your-client-secret"
286
+
287
+```
288
+</details>
289
+
290
+###### Filter by resource group
291
+
292
+Only monitor resources in specific resource groups.
293
+
294
+<details open><summary>Config</summary>
295
+
296
+```yaml
297
+jobs:
298
+ - name: prod-rg
299
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
300
+ resource_groups:
301
+ - production-rg
302
+ - staging-rg
303
+ auth:
304
+ mode: default
305
+
306
+```
307
+</details>
308
+
309
+###### Azure Government cloud
310
+
311
+Connect to Azure Government cloud environment.
312
+
313
+<details open><summary>Config</summary>
314
+
315
+```yaml
316
+jobs:
317
+ - name: gov
318
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
319
+ cloud: government
320
+ auth:
321
+ mode: service_principal
322
+ mode_service_principal:
323
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ client_secret: "your-client-secret"
326
+
327
+```
328
+</details>
329
+
330
+
331
+
332
+## Troubleshooting
333
+
334
+### Debug Mode
335
+
336
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
337
+
338
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
339
+should give you clues as to why the collector isn't working.
340
+
341
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
342
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
343
+
344
+ ```bash
345
+ cd /usr/libexec/netdata/plugins.d/
346
+ ```
347
+
348
+- Switch to the `netdata` user.
349
+
350
+ ```bash
351
+ sudo -u netdata -s
352
+ ```
353
+
354
+- Run the `go.d.plugin` to debug the collector:
355
+
356
+ ```bash
357
+ ./go.d.plugin -d -m azure_monitor
358
+ ```
359
+
360
+ To debug a specific job:
361
+
362
+ ```bash
363
+ ./go.d.plugin -d -m azure_monitor -j jobName
364
+ ```
365
+
366
+### Getting Logs
367
+
368
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
369
+
370
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
371
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
372
+
373
+#### System with systemd
374
+
375
+Use the following command to view logs generated since the last Netdata service restart:
376
+
377
+```bash
378
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
379
+```
380
+
381
+#### System without systemd
382
+
383
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
384
+
385
+```bash
386
+grep azure_monitor /var/log/netdata/collector.log
387
+```
388
+
389
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
390
+
391
+#### Docker Container
392
+
393
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
394
+
395
+```bash
396
+docker logs netdata 2>&1 | grep azure_monitor
397
+```
398
+
399
+### No metrics are collected
400
+
401
+Verify the following:
402
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
403
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
404
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
405
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
406
+
407
+
408
+### Missing metrics for some resource types
409
+
410
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
411
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
412
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
413
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
414
+
415
+
416
+### Metrics appear delayed
417
+
418
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
419
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
420
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
421
+
422
+
423
+### Authentication errors in sovereign clouds
424
+
425
+For Azure Government or Azure China clouds, set the `cloud` parameter:
426
+- Azure Government: `cloud: government`
427
+- Azure China (21Vianet): `cloud: china`
428
+
429
+Ensure the service principal is registered in the correct cloud tenant.
430
+
431
+
432
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_log_analytics_workspace.md
new
+471
@@ -0,0 +1,471 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_log_analytics_workspace.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Log Analytics Workspace"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'log', 'analytics', 'monitoring', 'logs', 'kusto']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Log Analytics Workspace
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Log Analytics workspaces including ingestion volume and latency, query execution counts and volume, available storage capacity, and per-table breakdowns of ingestion rates and billing volume.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.log_analytics.query_availability | availability | percentage |
88
+| azure_monitor.log_analytics.ingestion_latency | average, maximum, minimum | seconds |
89
+| azure_monitor.log_analytics.ingestion_volume | records | records/s |
90
+| azure_monitor.log_analytics.queries | total, failed | queries/s |
91
+| azure_monitor.log_analytics.export_data | exported | bytes/s |
92
+| azure_monitor.log_analytics.export_records | exported | records/s |
93
+| azure_monitor.log_analytics.export_failures | failed | exports/s |
94
+| azure_monitor.log_analytics.legacy_cpu_utilization | processor, privileged, user, idle | percentage |
95
+| azure_monitor.log_analytics.legacy_cpu_detailed | dpc, interrupt, io_wait, nice | percentage |
96
+| azure_monitor.log_analytics.legacy_cpu_pct | privileged, user | percentage |
97
+| azure_monitor.log_analytics.legacy_memory_utilization | available, used, committed | percentage |
98
+| azure_monitor.log_analytics.legacy_memory_available | available_mbytes, available_mbytes_memory | megabytes |
99
+| azure_monitor.log_analytics.legacy_memory_used_kbytes | used | kilobytes |
100
+| azure_monitor.log_analytics.legacy_memory_used_mbytes | used | megabytes |
101
+| azure_monitor.log_analytics.legacy_memory_free | physical, virtual, shared | megabytes |
102
+| azure_monitor.log_analytics.legacy_swap_utilization | available, used | percentage |
103
+| azure_monitor.log_analytics.legacy_swap_size | available, used | megabytes |
104
+| azure_monitor.log_analytics.legacy_disk_space_utilization | free, used | percentage |
105
+| azure_monitor.log_analytics.legacy_disk_inodes | free, used | percentage |
106
+| azure_monitor.log_analytics.legacy_disk_free | free | megabytes |
107
+| azure_monitor.log_analytics.legacy_disk_io_throughput | read, write, logical, physical | bytes/s |
108
+| azure_monitor.log_analytics.legacy_disk_io_operations | reads, writes, transfers | operations/s |
109
+| azure_monitor.log_analytics.legacy_disk_io_latency | read, write, transfer | seconds |
110
+| azure_monitor.log_analytics.legacy_disk_queue | queue_length | operations |
111
+| azure_monitor.log_analytics.legacy_paging | reads, writes, total | pages/s |
112
+| azure_monitor.log_analytics.legacy_paging_files | free, stored | megabytes |
113
+| azure_monitor.log_analytics.legacy_network_throughput_win | received, sent, total | bytes/s |
114
+| azure_monitor.log_analytics.legacy_network_traffic | total, received, transmitted | bytes |
115
+| azure_monitor.log_analytics.legacy_network_packets | received, transmitted | packets |
116
+| azure_monitor.log_analytics.legacy_network_errors | rx, tx, collisions | errors |
117
+| azure_monitor.log_analytics.legacy_system | processes | processes |
118
+| azure_monitor.log_analytics.legacy_processor_queue | queue_length | threads |
119
+| azure_monitor.log_analytics.legacy_uptime | uptime | seconds |
120
+| azure_monitor.log_analytics.legacy_users | users | users |
121
+| azure_monitor.log_analytics.legacy_events | events | events/s |
122
+| azure_monitor.log_analytics.legacy_heartbeat | heartbeats | heartbeats/s |
123
+| azure_monitor.log_analytics.legacy_updates | updates | updates |
124
+
125
+
126
+
127
+## Alerts
128
+
129
+
130
+The following alerts are available:
131
+
132
+| Alert name | On metric | Description |
133
+|:------------|:----------|:------------|
134
+| [ am_log_analytics_query_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.query_availability | Log Analytics query availability on ${label:resource_name} |
135
+| [ am_log_analytics_ingestion_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.ingestion_latency | Log Analytics ingestion latency on ${label:resource_name} |
136
+| [ am_log_analytics_query_failure_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.queries | Log Analytics query failures on ${label:resource_name} |
137
+| [ am_log_analytics_export_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.export_failures | Log Analytics export failures on ${label:resource_name} |
138
+| [ am_log_analytics_legacy_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_cpu_utilization | Legacy agent CPU on ${label:resource_name} |
139
+| [ am_log_analytics_legacy_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_memory_utilization | Legacy agent memory on ${label:resource_name} |
140
+| [ am_log_analytics_legacy_swap ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_swap_utilization | Legacy agent swap usage on ${label:resource_name} |
141
+| [ am_log_analytics_legacy_disk_space ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_disk_space_utilization | Legacy agent disk space on ${label:resource_name} |
142
+| [ am_log_analytics_legacy_disk_inodes ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_disk_inodes | Legacy agent inode usage on ${label:resource_name} |
143
+| [ am_log_analytics_legacy_disk_read_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_disk_io_latency | Legacy agent disk read latency on ${label:resource_name} |
144
+| [ am_log_analytics_legacy_disk_write_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_disk_io_latency | Legacy agent disk write latency on ${label:resource_name} |
145
+| [ am_log_analytics_legacy_disk_queue ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_disk_queue | Legacy agent disk queue on ${label:resource_name} |
146
+| [ am_log_analytics_legacy_network_rx_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_network_errors | Legacy agent network RX errors on ${label:resource_name} |
147
+| [ am_log_analytics_legacy_network_tx_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_log_analytics.conf) | azure_monitor.log_analytics.legacy_network_errors | Legacy agent network TX errors on ${label:resource_name} |
148
+
149
+
150
+## Setup
151
+
152
+
153
+You can configure the **azure_monitor** collector in two ways:
154
+
155
+| Method | Best for | How to |
156
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
157
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
158
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
159
+
160
+:::important
161
+
162
+UI configuration requires paid Netdata Cloud plan.
163
+
164
+:::
165
+
166
+
167
+### Prerequisites
168
+
169
+#### Create an Azure monitoring principal
170
+
171
+Create a service principal or use a managed identity with the following permissions:
172
+
173
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
174
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
175
+
176
+For service principal authentication:
177
+```bash
178
+# Create the service principal
179
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
180
+ --scopes /subscriptions/<subscription-id>
181
+
182
+# Note the appId (client_id), password (client_secret), and tenant
183
+```
184
+
185
+For managed identity (on Azure VMs, VMSS, or AKS):
186
+```bash
187
+# Assign Monitoring Reader role to the VM's managed identity
188
+az role assignment create --assignee <managed-identity-principal-id> \
189
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
190
+```
191
+
192
+
193
+
194
+### Configuration
195
+
196
+#### Options
197
+
198
+The following options can be defined globally: update_every, autodetection_retry.
199
+
200
+Profile files are loaded from:
201
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
202
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
203
+
204
+User profile files with the same filename override stock profiles.
205
+
206
+
207
+<details open><summary>Config options</summary>
208
+
209
+
210
+
211
+| Group | Option | Description | Default | Required |
212
+|:------|:-----|:------------|:--------|:---------:|
213
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
214
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
215
+| **Target** | subscription_id | Azure subscription ID. | | yes |
216
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
217
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
218
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
219
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
220
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
221
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
222
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
223
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
224
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
225
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
226
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
227
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
228
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
229
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
230
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
231
+
232
+
233
+</details>
234
+
235
+
236
+#### via UI
237
+
238
+Configure the **azure_monitor** collector from the Netdata web interface:
239
+
240
+1. Go to **Nodes**.
241
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
242
+3. The **Collectors → Jobs** view opens by default.
243
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
244
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
245
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
246
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
247
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
248
+
249
+
250
+#### via File
251
+
252
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
253
+
254
+The file format is YAML. Generally, the structure is:
255
+
256
+```yaml
257
+update_every: 1
258
+autodetection_retry: 0
259
+jobs:
260
+ - name: some_name1
261
+ - name: some_name2
262
+```
263
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
264
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
265
+
266
+```bash
267
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
268
+sudo ./edit-config go.d/azure_monitor.conf
269
+```
270
+
271
+##### Examples
272
+
273
+###### Service principal (auto-discover all resources)
274
+
275
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
276
+
277
+```yaml
278
+jobs:
279
+ - name: prod
280
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
281
+ auth:
282
+ mode: service_principal
283
+ mode_service_principal:
284
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
285
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
286
+ client_secret: "your-client-secret"
287
+
288
+```
289
+###### Managed identity (Azure VM/VMSS/AKS)
290
+
291
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
292
+
293
+<details open><summary>Config</summary>
294
+
295
+```yaml
296
+jobs:
297
+ - name: prod
298
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ auth:
300
+ mode: managed_identity
301
+
302
+```
303
+</details>
304
+
305
+###### Specific profiles only
306
+
307
+Monitor only specific Azure services instead of auto-discovering all resource types.
308
+
309
+<details open><summary>Config</summary>
310
+
311
+```yaml
312
+jobs:
313
+ - name: databases
314
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
315
+ profiles:
316
+ - sql_database
317
+ - postgres_flexible
318
+ - redis_cache
319
+ auth:
320
+ mode: service_principal
321
+ mode_service_principal:
322
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ client_secret: "your-client-secret"
325
+
326
+```
327
+</details>
328
+
329
+###### Filter by resource group
330
+
331
+Only monitor resources in specific resource groups.
332
+
333
+<details open><summary>Config</summary>
334
+
335
+```yaml
336
+jobs:
337
+ - name: prod-rg
338
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
339
+ resource_groups:
340
+ - production-rg
341
+ - staging-rg
342
+ auth:
343
+ mode: default
344
+
345
+```
346
+</details>
347
+
348
+###### Azure Government cloud
349
+
350
+Connect to Azure Government cloud environment.
351
+
352
+<details open><summary>Config</summary>
353
+
354
+```yaml
355
+jobs:
356
+ - name: gov
357
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ cloud: government
359
+ auth:
360
+ mode: service_principal
361
+ mode_service_principal:
362
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
363
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
364
+ client_secret: "your-client-secret"
365
+
366
+```
367
+</details>
368
+
369
+
370
+
371
+## Troubleshooting
372
+
373
+### Debug Mode
374
+
375
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
376
+
377
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
378
+should give you clues as to why the collector isn't working.
379
+
380
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
381
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
382
+
383
+ ```bash
384
+ cd /usr/libexec/netdata/plugins.d/
385
+ ```
386
+
387
+- Switch to the `netdata` user.
388
+
389
+ ```bash
390
+ sudo -u netdata -s
391
+ ```
392
+
393
+- Run the `go.d.plugin` to debug the collector:
394
+
395
+ ```bash
396
+ ./go.d.plugin -d -m azure_monitor
397
+ ```
398
+
399
+ To debug a specific job:
400
+
401
+ ```bash
402
+ ./go.d.plugin -d -m azure_monitor -j jobName
403
+ ```
404
+
405
+### Getting Logs
406
+
407
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
408
+
409
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
410
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
411
+
412
+#### System with systemd
413
+
414
+Use the following command to view logs generated since the last Netdata service restart:
415
+
416
+```bash
417
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
418
+```
419
+
420
+#### System without systemd
421
+
422
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
423
+
424
+```bash
425
+grep azure_monitor /var/log/netdata/collector.log
426
+```
427
+
428
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
429
+
430
+#### Docker Container
431
+
432
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
433
+
434
+```bash
435
+docker logs netdata 2>&1 | grep azure_monitor
436
+```
437
+
438
+### No metrics are collected
439
+
440
+Verify the following:
441
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
442
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
443
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
444
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
445
+
446
+
447
+### Missing metrics for some resource types
448
+
449
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
450
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
451
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
452
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
453
+
454
+
455
+### Metrics appear delayed
456
+
457
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
458
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
459
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
460
+
461
+
462
+### Authentication errors in sovereign clouds
463
+
464
+For Azure Government or Azure China clouds, set the `cloud` parameter:
465
+- Azure Government: `cloud: government`
466
+- Azure China (21Vianet): `cloud: china`
467
+
468
+Ensure the service principal is registered in the correct cloud tenant.
469
+
470
+
471
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_logic_apps_workflow.md
new
+445
@@ -0,0 +1,445 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_logic_apps_workflow.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Logic Apps Workflow"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'logic', 'apps', 'workflow', 'integration', 'automation']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Logic Apps Workflow
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Logic Apps workflow execution including run completions and failures, action execution counts, trigger firing rates, run and action latency, billable executions, and action-level success and failure breakdowns.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.logic_apps.run_lifecycle | started, completed, succeeded, failed, cancelled | runs/s |
88
+| azure_monitor.logic_apps.run_latency | all, success | seconds |
89
+| azure_monitor.logic_apps.run_failure_rate | failure_rate | percentage |
90
+| azure_monitor.logic_apps.run_throttling | during_run, at_start | events/s |
91
+| azure_monitor.logic_apps.action_lifecycle | started, completed, succeeded, failed, skipped | actions/s |
92
+| azure_monitor.logic_apps.action_latency | all, success | seconds |
93
+| azure_monitor.logic_apps.action_throttling | total | events/s |
94
+| azure_monitor.logic_apps.trigger_lifecycle | started, completed, succeeded, fired, failed, skipped | triggers/s |
95
+| azure_monitor.logic_apps.trigger_latency | all, fire, success | seconds |
96
+| azure_monitor.logic_apps.trigger_throttling | total | events/s |
97
+| azure_monitor.logic_apps.billable_executions | total, actions, triggers | executions/s |
98
+| azure_monitor.logic_apps.billing_by_type | native, standard_connector, storage | operations/s |
99
+| azure_monitor.logic_apps.agent | loop_executions, completion_overflow, prompt_overflow | events/s |
100
+
101
+
102
+
103
+## Alerts
104
+
105
+
106
+The following alerts are available:
107
+
108
+| Alert name | On metric | Description |
109
+|:------------|:----------|:------------|
110
+| [ am_logic_apps_run_failure_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.run_failure_rate | Logic Apps run failure rate on ${label:resource_name} |
111
+| [ am_logic_apps_runs_failed ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.run_lifecycle | Logic Apps failed runs on ${label:resource_name} |
112
+| [ am_logic_apps_run_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.run_latency | Logic Apps run latency on ${label:resource_name} |
113
+| [ am_logic_apps_run_throttled ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.run_throttling | Logic Apps run throttling on ${label:resource_name} |
114
+| [ am_logic_apps_actions_failed ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.action_lifecycle | Logic Apps failed actions on ${label:resource_name} |
115
+| [ am_logic_apps_action_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.action_latency | Logic Apps action latency on ${label:resource_name} |
116
+| [ am_logic_apps_action_throttled ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.action_throttling | Logic Apps action throttling on ${label:resource_name} |
117
+| [ am_logic_apps_triggers_failed ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.trigger_lifecycle | Logic Apps failed triggers on ${label:resource_name} |
118
+| [ am_logic_apps_trigger_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.trigger_latency | Logic Apps trigger latency on ${label:resource_name} |
119
+| [ am_logic_apps_trigger_throttled ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.trigger_throttling | Logic Apps trigger throttling on ${label:resource_name} |
120
+| [ am_logic_apps_completion_token_overflow ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.agent | Logic Apps completion token overflow on ${label:resource_name} |
121
+| [ am_logic_apps_prompt_token_overflow ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_logic_apps.conf) | azure_monitor.logic_apps.agent | Logic Apps prompt token overflow on ${label:resource_name} |
122
+
123
+
124
+## Setup
125
+
126
+
127
+You can configure the **azure_monitor** collector in two ways:
128
+
129
+| Method | Best for | How to |
130
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
131
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
132
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
133
+
134
+:::important
135
+
136
+UI configuration requires paid Netdata Cloud plan.
137
+
138
+:::
139
+
140
+
141
+### Prerequisites
142
+
143
+#### Create an Azure monitoring principal
144
+
145
+Create a service principal or use a managed identity with the following permissions:
146
+
147
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
148
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
149
+
150
+For service principal authentication:
151
+```bash
152
+# Create the service principal
153
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
154
+ --scopes /subscriptions/<subscription-id>
155
+
156
+# Note the appId (client_id), password (client_secret), and tenant
157
+```
158
+
159
+For managed identity (on Azure VMs, VMSS, or AKS):
160
+```bash
161
+# Assign Monitoring Reader role to the VM's managed identity
162
+az role assignment create --assignee <managed-identity-principal-id> \
163
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
164
+```
165
+
166
+
167
+
168
+### Configuration
169
+
170
+#### Options
171
+
172
+The following options can be defined globally: update_every, autodetection_retry.
173
+
174
+Profile files are loaded from:
175
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
176
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
177
+
178
+User profile files with the same filename override stock profiles.
179
+
180
+
181
+<details open><summary>Config options</summary>
182
+
183
+
184
+
185
+| Group | Option | Description | Default | Required |
186
+|:------|:-----|:------------|:--------|:---------:|
187
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
188
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
189
+| **Target** | subscription_id | Azure subscription ID. | | yes |
190
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
191
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
192
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
193
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
194
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
195
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
196
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
197
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
198
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
199
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
200
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
201
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
202
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
203
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
204
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
205
+
206
+
207
+</details>
208
+
209
+
210
+#### via UI
211
+
212
+Configure the **azure_monitor** collector from the Netdata web interface:
213
+
214
+1. Go to **Nodes**.
215
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
216
+3. The **Collectors → Jobs** view opens by default.
217
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
218
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
219
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
220
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
221
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
222
+
223
+
224
+#### via File
225
+
226
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
227
+
228
+The file format is YAML. Generally, the structure is:
229
+
230
+```yaml
231
+update_every: 1
232
+autodetection_retry: 0
233
+jobs:
234
+ - name: some_name1
235
+ - name: some_name2
236
+```
237
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
238
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
239
+
240
+```bash
241
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
242
+sudo ./edit-config go.d/azure_monitor.conf
243
+```
244
+
245
+##### Examples
246
+
247
+###### Service principal (auto-discover all resources)
248
+
249
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
250
+
251
+```yaml
252
+jobs:
253
+ - name: prod
254
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
255
+ auth:
256
+ mode: service_principal
257
+ mode_service_principal:
258
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
259
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
260
+ client_secret: "your-client-secret"
261
+
262
+```
263
+###### Managed identity (Azure VM/VMSS/AKS)
264
+
265
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
266
+
267
+<details open><summary>Config</summary>
268
+
269
+```yaml
270
+jobs:
271
+ - name: prod
272
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
273
+ auth:
274
+ mode: managed_identity
275
+
276
+```
277
+</details>
278
+
279
+###### Specific profiles only
280
+
281
+Monitor only specific Azure services instead of auto-discovering all resource types.
282
+
283
+<details open><summary>Config</summary>
284
+
285
+```yaml
286
+jobs:
287
+ - name: databases
288
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
289
+ profiles:
290
+ - sql_database
291
+ - postgres_flexible
292
+ - redis_cache
293
+ auth:
294
+ mode: service_principal
295
+ mode_service_principal:
296
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
297
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
298
+ client_secret: "your-client-secret"
299
+
300
+```
301
+</details>
302
+
303
+###### Filter by resource group
304
+
305
+Only monitor resources in specific resource groups.
306
+
307
+<details open><summary>Config</summary>
308
+
309
+```yaml
310
+jobs:
311
+ - name: prod-rg
312
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
313
+ resource_groups:
314
+ - production-rg
315
+ - staging-rg
316
+ auth:
317
+ mode: default
318
+
319
+```
320
+</details>
321
+
322
+###### Azure Government cloud
323
+
324
+Connect to Azure Government cloud environment.
325
+
326
+<details open><summary>Config</summary>
327
+
328
+```yaml
329
+jobs:
330
+ - name: gov
331
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
332
+ cloud: government
333
+ auth:
334
+ mode: service_principal
335
+ mode_service_principal:
336
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
337
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
338
+ client_secret: "your-client-secret"
339
+
340
+```
341
+</details>
342
+
343
+
344
+
345
+## Troubleshooting
346
+
347
+### Debug Mode
348
+
349
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
350
+
351
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
352
+should give you clues as to why the collector isn't working.
353
+
354
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
355
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
356
+
357
+ ```bash
358
+ cd /usr/libexec/netdata/plugins.d/
359
+ ```
360
+
361
+- Switch to the `netdata` user.
362
+
363
+ ```bash
364
+ sudo -u netdata -s
365
+ ```
366
+
367
+- Run the `go.d.plugin` to debug the collector:
368
+
369
+ ```bash
370
+ ./go.d.plugin -d -m azure_monitor
371
+ ```
372
+
373
+ To debug a specific job:
374
+
375
+ ```bash
376
+ ./go.d.plugin -d -m azure_monitor -j jobName
377
+ ```
378
+
379
+### Getting Logs
380
+
381
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
382
+
383
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
384
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
385
+
386
+#### System with systemd
387
+
388
+Use the following command to view logs generated since the last Netdata service restart:
389
+
390
+```bash
391
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
392
+```
393
+
394
+#### System without systemd
395
+
396
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
397
+
398
+```bash
399
+grep azure_monitor /var/log/netdata/collector.log
400
+```
401
+
402
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
403
+
404
+#### Docker Container
405
+
406
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
407
+
408
+```bash
409
+docker logs netdata 2>&1 | grep azure_monitor
410
+```
411
+
412
+### No metrics are collected
413
+
414
+Verify the following:
415
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
416
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
417
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
418
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
419
+
420
+
421
+### Missing metrics for some resource types
422
+
423
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
424
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
425
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
426
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
427
+
428
+
429
+### Metrics appear delayed
430
+
431
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
432
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
433
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
434
+
435
+
436
+### Authentication errors in sovereign clouds
437
+
438
+For Azure Government or Azure China clouds, set the `cloud` parameter:
439
+- Azure Government: `cloud: government`
440
+- Azure China (21Vianet): `cloud: china`
441
+
442
+Ensure the service principal is registered in the correct cloud tenant.
443
+
444
+
445
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_machine_learning_workspace.md
new
+465
@@ -0,0 +1,465 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_machine_learning_workspace.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Machine Learning Workspace"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'machine', 'learning', 'ml', 'ai', 'workspace']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Machine Learning Workspace
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Machine Learning workspaces including active model deployments and registered models, pipeline run completions and failures, compute node utilization and preemptions, quota usage, managed endpoint request latency and rates, estimated GPU utilization, and storage utilization.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.machine_learning.agent_events | agent_events, thread_events | events/s |
88
+| azure_monitor.machine_learning.agent_messages | messages | messages/s |
89
+| azure_monitor.machine_learning.agent_runs | runs | runs/s |
90
+| azure_monitor.machine_learning.agent_tokens | tokens | tokens/s |
91
+| azure_monitor.machine_learning.agent_tool_calls | tool_calls | calls/s |
92
+| azure_monitor.machine_learning.agent_indexed_files | indexed_files | files/s |
93
+| azure_monitor.machine_learning.model_deployments | started, succeeded, failed | deployments/s |
94
+| azure_monitor.machine_learning.model_registrations | succeeded, failed | registrations/s |
95
+| azure_monitor.machine_learning.cluster_cores | active, idle, leaving, preempted, unusable | cores |
96
+| azure_monitor.machine_learning.total_cores | total | cores |
97
+| azure_monitor.machine_learning.cluster_nodes | active, idle, leaving, preempted, unusable | nodes |
98
+| azure_monitor.machine_learning.total_nodes | total | nodes |
99
+| azure_monitor.machine_learning.quota_utilization | utilization | percentage |
100
+| azure_monitor.machine_learning.cpu_utilization | cluster_cpu, node_cpu | percentage |
101
+| azure_monitor.machine_learning.cpu_millicores | used, capacity | millicores |
102
+| azure_monitor.machine_learning.cpu_memory_utilization | utilization | percentage |
103
+| azure_monitor.machine_learning.cpu_memory_megabytes | used, capacity | megabytes |
104
+| azure_monitor.machine_learning.gpu_utilization | cluster_gpu, node_gpu | percentage |
105
+| azure_monitor.machine_learning.gpu_milligpus | used, capacity | milliGPUs |
106
+| azure_monitor.machine_learning.gpu_memory_utilization | cluster_gpu_memory, node_gpu_memory | percentage |
107
+| azure_monitor.machine_learning.gpu_memory_megabytes | used, capacity | megabytes |
108
+| azure_monitor.machine_learning.gpu_energy | energy | joules/s |
109
+| azure_monitor.machine_learning.disk_usage | used, available | megabytes |
110
+| azure_monitor.machine_learning.disk_io | read, write | megabytes/s |
111
+| azure_monitor.machine_learning.network_traffic | in, out | megabytes/s |
112
+| azure_monitor.machine_learning.infiniband_traffic | receive, transmit | megabytes/s |
113
+| azure_monitor.machine_learning.storage_api_calls | success, failure | calls/s |
114
+| azure_monitor.machine_learning.run_lifecycle | not_started, starting, preparing, provisioning, queued, started | runs/s |
115
+| azure_monitor.machine_learning.run_completion | completed, failed, finalizing, cancelled, cancel_requested, not_responding | runs/s |
116
+| azure_monitor.machine_learning.run_issues | errors, warnings | events/s |
117
+
118
+
119
+
120
+## Alerts
121
+
122
+
123
+The following alerts are available:
124
+
125
+| Alert name | On metric | Description |
126
+|:------------|:----------|:------------|
127
+| [ am_ml_quota_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.quota_utilization | ML quota utilization on ${label:resource_name} |
128
+| [ am_ml_unusable_cores ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.cluster_cores | ML unusable cores on ${label:resource_name} |
129
+| [ am_ml_preempted_cores ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.cluster_cores | ML preempted cores on ${label:resource_name} |
130
+| [ am_ml_unusable_nodes ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.cluster_nodes | ML unusable nodes on ${label:resource_name} |
131
+| [ am_ml_cpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.cpu_utilization | ML CPU utilization on ${label:resource_name} |
132
+| [ am_ml_cpu_memory_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.cpu_memory_utilization | ML CPU memory utilization on ${label:resource_name} |
133
+| [ am_ml_gpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.gpu_utilization | ML GPU utilization on ${label:resource_name} |
134
+| [ am_ml_gpu_memory_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.gpu_memory_utilization | ML GPU memory utilization on ${label:resource_name} |
135
+| [ am_ml_disk_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.disk_usage | ML disk utilization on ${label:resource_name} |
136
+| [ am_ml_model_deploy_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.model_deployments | ML model deployment failures on ${label:resource_name} |
137
+| [ am_ml_model_register_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.model_registrations | ML model registration failures on ${label:resource_name} |
138
+| [ am_ml_failed_runs ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.run_completion | ML failed runs on ${label:resource_name} |
139
+| [ am_ml_not_responding_runs ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.run_completion | ML not-responding runs on ${label:resource_name} |
140
+| [ am_ml_run_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.run_issues | ML run errors on ${label:resource_name} |
141
+| [ am_ml_storage_api_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_machine_learning.conf) | azure_monitor.machine_learning.storage_api_calls | ML storage API failure rate on ${label:resource_name} |
142
+
143
+
144
+## Setup
145
+
146
+
147
+You can configure the **azure_monitor** collector in two ways:
148
+
149
+| Method | Best for | How to |
150
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
151
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
152
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
153
+
154
+:::important
155
+
156
+UI configuration requires paid Netdata Cloud plan.
157
+
158
+:::
159
+
160
+
161
+### Prerequisites
162
+
163
+#### Create an Azure monitoring principal
164
+
165
+Create a service principal or use a managed identity with the following permissions:
166
+
167
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
168
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
169
+
170
+For service principal authentication:
171
+```bash
172
+# Create the service principal
173
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
174
+ --scopes /subscriptions/<subscription-id>
175
+
176
+# Note the appId (client_id), password (client_secret), and tenant
177
+```
178
+
179
+For managed identity (on Azure VMs, VMSS, or AKS):
180
+```bash
181
+# Assign Monitoring Reader role to the VM's managed identity
182
+az role assignment create --assignee <managed-identity-principal-id> \
183
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
184
+```
185
+
186
+
187
+
188
+### Configuration
189
+
190
+#### Options
191
+
192
+The following options can be defined globally: update_every, autodetection_retry.
193
+
194
+Profile files are loaded from:
195
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
196
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
197
+
198
+User profile files with the same filename override stock profiles.
199
+
200
+
201
+<details open><summary>Config options</summary>
202
+
203
+
204
+
205
+| Group | Option | Description | Default | Required |
206
+|:------|:-----|:------------|:--------|:---------:|
207
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
208
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
209
+| **Target** | subscription_id | Azure subscription ID. | | yes |
210
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
211
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
212
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
213
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
214
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
215
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
216
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
217
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
218
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
219
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
220
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
221
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
222
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
223
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
224
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
225
+
226
+
227
+</details>
228
+
229
+
230
+#### via UI
231
+
232
+Configure the **azure_monitor** collector from the Netdata web interface:
233
+
234
+1. Go to **Nodes**.
235
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
236
+3. The **Collectors → Jobs** view opens by default.
237
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
238
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
239
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
240
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
241
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
242
+
243
+
244
+#### via File
245
+
246
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
247
+
248
+The file format is YAML. Generally, the structure is:
249
+
250
+```yaml
251
+update_every: 1
252
+autodetection_retry: 0
253
+jobs:
254
+ - name: some_name1
255
+ - name: some_name2
256
+```
257
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
258
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
259
+
260
+```bash
261
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
262
+sudo ./edit-config go.d/azure_monitor.conf
263
+```
264
+
265
+##### Examples
266
+
267
+###### Service principal (auto-discover all resources)
268
+
269
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
270
+
271
+```yaml
272
+jobs:
273
+ - name: prod
274
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
275
+ auth:
276
+ mode: service_principal
277
+ mode_service_principal:
278
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
279
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
280
+ client_secret: "your-client-secret"
281
+
282
+```
283
+###### Managed identity (Azure VM/VMSS/AKS)
284
+
285
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
286
+
287
+<details open><summary>Config</summary>
288
+
289
+```yaml
290
+jobs:
291
+ - name: prod
292
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
293
+ auth:
294
+ mode: managed_identity
295
+
296
+```
297
+</details>
298
+
299
+###### Specific profiles only
300
+
301
+Monitor only specific Azure services instead of auto-discovering all resource types.
302
+
303
+<details open><summary>Config</summary>
304
+
305
+```yaml
306
+jobs:
307
+ - name: databases
308
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
309
+ profiles:
310
+ - sql_database
311
+ - postgres_flexible
312
+ - redis_cache
313
+ auth:
314
+ mode: service_principal
315
+ mode_service_principal:
316
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
317
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
318
+ client_secret: "your-client-secret"
319
+
320
+```
321
+</details>
322
+
323
+###### Filter by resource group
324
+
325
+Only monitor resources in specific resource groups.
326
+
327
+<details open><summary>Config</summary>
328
+
329
+```yaml
330
+jobs:
331
+ - name: prod-rg
332
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
333
+ resource_groups:
334
+ - production-rg
335
+ - staging-rg
336
+ auth:
337
+ mode: default
338
+
339
+```
340
+</details>
341
+
342
+###### Azure Government cloud
343
+
344
+Connect to Azure Government cloud environment.
345
+
346
+<details open><summary>Config</summary>
347
+
348
+```yaml
349
+jobs:
350
+ - name: gov
351
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
352
+ cloud: government
353
+ auth:
354
+ mode: service_principal
355
+ mode_service_principal:
356
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ client_secret: "your-client-secret"
359
+
360
+```
361
+</details>
362
+
363
+
364
+
365
+## Troubleshooting
366
+
367
+### Debug Mode
368
+
369
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
370
+
371
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
372
+should give you clues as to why the collector isn't working.
373
+
374
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
375
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
376
+
377
+ ```bash
378
+ cd /usr/libexec/netdata/plugins.d/
379
+ ```
380
+
381
+- Switch to the `netdata` user.
382
+
383
+ ```bash
384
+ sudo -u netdata -s
385
+ ```
386
+
387
+- Run the `go.d.plugin` to debug the collector:
388
+
389
+ ```bash
390
+ ./go.d.plugin -d -m azure_monitor
391
+ ```
392
+
393
+ To debug a specific job:
394
+
395
+ ```bash
396
+ ./go.d.plugin -d -m azure_monitor -j jobName
397
+ ```
398
+
399
+### Getting Logs
400
+
401
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
402
+
403
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
404
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
405
+
406
+#### System with systemd
407
+
408
+Use the following command to view logs generated since the last Netdata service restart:
409
+
410
+```bash
411
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
412
+```
413
+
414
+#### System without systemd
415
+
416
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
417
+
418
+```bash
419
+grep azure_monitor /var/log/netdata/collector.log
420
+```
421
+
422
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
423
+
424
+#### Docker Container
425
+
426
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
427
+
428
+```bash
429
+docker logs netdata 2>&1 | grep azure_monitor
430
+```
431
+
432
+### No metrics are collected
433
+
434
+Verify the following:
435
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
436
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
437
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
438
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
439
+
440
+
441
+### Missing metrics for some resource types
442
+
443
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
444
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
445
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
446
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
447
+
448
+
449
+### Metrics appear delayed
450
+
451
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
452
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
453
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
454
+
455
+
456
+### Authentication errors in sovereign clouds
457
+
458
+For Azure Government or Azure China clouds, set the `cloud` parameter:
459
+- Azure Government: `cloud: government`
460
+- Azure China (21Vianet): `cloud: china`
461
+
462
+Ensure the service principal is registered in the correct cloud tenant.
463
+
464
+
465
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md
+383
-5
@@ -1,20 +1,398 @@
1
<!--startmeta
2
-custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/README.md"
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md"
3
meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
sidebar_label: "Azure Monitor"
5
learn_status: "Published"
6
-learn_rel_path: "Collecting Metrics/Cloud"
7
-keywords: ["azure", "monitor", "cloud"]
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'monitor', 'cloud']
8
message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
endmeta-->
10
11
# Azure Monitor
12
13
+
14
<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
16
+
17
Plugin: go.d.plugin
18
Module: azure_monitor
19
18
-This collector discovers Azure resources with Resource Graph and collects service metrics using the Azure Monitor Metrics batch API.
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+This collector monitors Azure resources through the Azure Monitor Metrics API. It automatically discovers
25
+resources in your subscription and collects platform metrics based on configurable profiles, providing
26
+visibility into the health and performance of over 35 Azure service types.
27
+
28
+
29
+The collector uses Azure SDK clients for:
30
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
31
+- Resource discovery via Azure Resource Graph queries
32
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
33
+
34
+
35
+This collector is supported on all platforms.
36
+
37
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
38
+
39
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
40
+
41
+
42
+### Default Behavior
43
+
44
+#### Auto-Detection
45
+
46
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
47
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
48
+
49
+
50
+#### Limits
51
+
52
+Azure Monitor metrics granularity is typically 1 minute.
53
+The collector enforces a minimum collection interval of 60 seconds.
54
+
55
+
56
+#### Performance Impact
57
+
58
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
59
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
60
+
61
+
62
+## Metrics
63
+
64
+Metrics depend on which Azure Monitor profiles are enabled. Each profile corresponds to an Azure
65
+service type and defines the specific metrics collected. With the default `profiles: [auto]` setting,
66
+profiles are automatically enabled for resource types found in your subscription.
67
+
68
+See the service-specific integrations below for detailed metrics lists.
69
+
70
+
71
+
72
+## Alerts
73
+
74
+There are no alerts configured by default for this integration.
75
+
76
+
77
+## Setup
78
+
79
+
80
+You can configure the **azure_monitor** collector in two ways:
81
+
82
+| Method | Best for | How to |
83
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
84
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
85
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
86
+
87
+:::important
88
+
89
+UI configuration requires paid Netdata Cloud plan.
90
+
91
+:::
92
+
93
+
94
+### Prerequisites
95
+
96
+#### Create an Azure monitoring principal
97
+
98
+Create a service principal or use a managed identity with the following permissions:
99
+
100
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
101
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
102
+
103
+For service principal authentication:
104
+```bash
105
+# Create the service principal
106
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
107
+ --scopes /subscriptions/<subscription-id>
108
+
109
+# Note the appId (client_id), password (client_secret), and tenant
110
+```
111
+
112
+For managed identity (on Azure VMs, VMSS, or AKS):
113
+```bash
114
+# Assign Monitoring Reader role to the VM's managed identity
115
+az role assignment create --assignee <managed-identity-principal-id> \
116
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
117
+```
118
+
119
+
120
+
121
+### Configuration
122
+
123
+#### Options
124
+
125
+The following options can be defined globally: update_every, autodetection_retry.
126
+
127
+Profile files are loaded from:
128
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
129
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
130
+
131
+User profile files with the same filename override stock profiles.
132
+
133
+
134
+<details open><summary>Config options</summary>
135
+
136
+
137
+
138
+| Group | Option | Description | Default | Required |
139
+|:------|:-----|:------------|:--------|:---------:|
140
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
141
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
142
+| **Target** | subscription_id | Azure subscription ID. | | yes |
143
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
144
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
145
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
146
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
147
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
148
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
149
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
150
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
151
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
152
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
153
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
154
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
155
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
156
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
157
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
158
+
159
+
160
+</details>
161
+
162
+
163
+#### via UI
164
+
165
+Configure the **azure_monitor** collector from the Netdata web interface:
166
+
167
+1. Go to **Nodes**.
168
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
169
+3. The **Collectors → Jobs** view opens by default.
170
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
171
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
172
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
173
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
174
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
175
+
176
+
177
+#### via File
178
+
179
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
180
+
181
+The file format is YAML. Generally, the structure is:
182
+
183
+```yaml
184
+update_every: 1
185
+autodetection_retry: 0
186
+jobs:
187
+ - name: some_name1
188
+ - name: some_name2
189
+```
190
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
191
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
192
+
193
+```bash
194
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
195
+sudo ./edit-config go.d/azure_monitor.conf
196
+```
197
+
198
+##### Examples
199
+
200
+###### Service principal (auto-discover all resources)
201
+
202
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
203
+
204
+```yaml
205
+jobs:
206
+ - name: prod
207
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
208
+ auth:
209
+ mode: service_principal
210
+ mode_service_principal:
211
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
212
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
213
+ client_secret: "your-client-secret"
214
+
215
+```
216
+###### Managed identity (Azure VM/VMSS/AKS)
217
+
218
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
219
+
220
+<details open><summary>Config</summary>
221
+
222
+```yaml
223
+jobs:
224
+ - name: prod
225
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
226
+ auth:
227
+ mode: managed_identity
228
+
229
+```
230
+</details>
231
+
232
+###### Specific profiles only
233
+
234
+Monitor only specific Azure services instead of auto-discovering all resource types.
235
+
236
+<details open><summary>Config</summary>
237
+
238
+```yaml
239
+jobs:
240
+ - name: databases
241
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ profiles:
243
+ - sql_database
244
+ - postgres_flexible
245
+ - redis_cache
246
+ auth:
247
+ mode: service_principal
248
+ mode_service_principal:
249
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
250
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
+ client_secret: "your-client-secret"
252
+
253
+```
254
+</details>
255
+
256
+###### Filter by resource group
257
+
258
+Only monitor resources in specific resource groups.
259
+
260
+<details open><summary>Config</summary>
261
+
262
+```yaml
263
+jobs:
264
+ - name: prod-rg
265
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
266
+ resource_groups:
267
+ - production-rg
268
+ - staging-rg
269
+ auth:
270
+ mode: default
271
+
272
+```
273
+</details>
274
+
275
+###### Azure Government cloud
276
+
277
+Connect to Azure Government cloud environment.
278
+
279
+<details open><summary>Config</summary>
280
+
281
+```yaml
282
+jobs:
283
+ - name: gov
284
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
285
+ cloud: government
286
+ auth:
287
+ mode: service_principal
288
+ mode_service_principal:
289
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
290
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
291
+ client_secret: "your-client-secret"
292
+
293
+```
294
+</details>
295
+
296
+
297
+
298
+## Troubleshooting
299
+
300
+### Debug Mode
301
+
302
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
303
+
304
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
305
+should give you clues as to why the collector isn't working.
306
+
307
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
308
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
309
+
310
+ ```bash
311
+ cd /usr/libexec/netdata/plugins.d/
312
+ ```
313
+
314
+- Switch to the `netdata` user.
315
+
316
+ ```bash
317
+ sudo -u netdata -s
318
+ ```
319
+
320
+- Run the `go.d.plugin` to debug the collector:
321
+
322
+ ```bash
323
+ ./go.d.plugin -d -m azure_monitor
324
+ ```
325
+
326
+ To debug a specific job:
327
+
328
+ ```bash
329
+ ./go.d.plugin -d -m azure_monitor -j jobName
330
+ ```
331
+
332
+### Getting Logs
333
+
334
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
335
+
336
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
337
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
338
+
339
+#### System with systemd
340
+
341
+Use the following command to view logs generated since the last Netdata service restart:
342
+
343
+```bash
344
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
345
+```
346
+
347
+#### System without systemd
348
+
349
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
350
+
351
+```bash
352
+grep azure_monitor /var/log/netdata/collector.log
353
+```
354
+
355
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
356
+
357
+#### Docker Container
358
+
359
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
360
+
361
+```bash
362
+docker logs netdata 2>&1 | grep azure_monitor
363
+```
364
+
365
+### No metrics are collected
366
+
367
+Verify the following:
368
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
369
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
370
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
371
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
372
+
373
+
374
+### Missing metrics for some resource types
375
+
376
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
377
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
378
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
379
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
380
+
381
+
382
+### Metrics appear delayed
383
+
384
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
385
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
386
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
387
+
388
+
389
+### Authentication errors in sovereign clouds
390
+
391
+For Azure Government or Azure China clouds, set the `cloud` parameter:
392
+- Azure Government: `cloud: government`
393
+- Azure China (21Vianet): `cloud: china`
394
+
395
+Ensure the service principal is registered in the correct cloud tenant.
396
+
397
+
398
20
-See [collector README](https://github.com/netdata/netdata/tree/master/src/go/plugin/go.d/collector/azure_monitor) for complete setup and configuration details.
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_mysql_flexible_server.md
new
+472
@@ -0,0 +1,472 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_mysql_flexible_server.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure MySQL Flexible Server"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'mysql', 'database', 'flexible']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure MySQL Flexible Server
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor MySQL Flexible Server including active connections, aborted connections, query rates, replication lag, storage utilization, CPU and memory usage, IO operations, InnoDB buffer pool efficiency, network throughput, and HA replication status.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.mysql_flexible.cpu | average | percentage |
88
+| azure_monitor.mysql_flexible.memory | average | percentage |
89
+| azure_monitor.mysql_flexible.io_utilization | average | percentage |
90
+| azure_monitor.mysql_flexible.cpu_credits | consumed, remaining | credits |
91
+| azure_monitor.mysql_flexible.ha_status | io, sql | status |
92
+| azure_monitor.mysql_flexible.replica_status | io, sql | status |
93
+| azure_monitor.mysql_flexible.uptime | uptime | seconds |
94
+| azure_monitor.mysql_flexible.replication_lag | replica, ha | seconds |
95
+| azure_monitor.mysql_flexible.innodb_row_lock_time | average | milliseconds |
96
+| azure_monitor.mysql_flexible.innodb_row_lock_waits | total | waits/s |
97
+| azure_monitor.mysql_flexible.aborted_connections | total | connections/s |
98
+| azure_monitor.mysql_flexible.active_connections | average | connections |
99
+| azure_monitor.mysql_flexible.total_connections | total | connections/s |
100
+| azure_monitor.mysql_flexible.active_transactions | average | transactions |
101
+| azure_monitor.mysql_flexible.threads_running | maximum | threads |
102
+| azure_monitor.mysql_flexible.queries | total, slow | queries/s |
103
+| azure_monitor.mysql_flexible.dml_statements | select, insert, update, delete | statements/s |
104
+| azure_monitor.mysql_flexible.ddl_statements | create_table, alter_table, drop_table, create_db, drop_db | statements/s |
105
+| azure_monitor.mysql_flexible.lock_deadlocks | total | deadlocks/s |
106
+| azure_monitor.mysql_flexible.lock_timeouts | total | timeouts/s |
107
+| azure_monitor.mysql_flexible.innodb_buffer_pool_pages | data, dirty, free, flushed | pages |
108
+| azure_monitor.mysql_flexible.innodb_buffer_pool_io | read_requests, disk_reads | requests/s |
109
+| azure_monitor.mysql_flexible.innodb_data_writes | total | writes/s |
110
+| azure_monitor.mysql_flexible.sort_merge_passes | total | passes/s |
111
+| azure_monitor.mysql_flexible.network | in, out | bytes/s |
112
+| azure_monitor.mysql_flexible.storage_io | total | operations/s |
113
+| azure_monitor.mysql_flexible.storage | used, limit | bytes |
114
+| azure_monitor.mysql_flexible.storage_utilization | average | percentage |
115
+| azure_monitor.mysql_flexible.storage_breakdown | data, ibdata1, binlog, others | bytes |
116
+| azure_monitor.mysql_flexible.backup_storage | used | bytes |
117
+| azure_monitor.mysql_flexible.serverlog_storage | used, limit | bytes |
118
+| azure_monitor.mysql_flexible.serverlog_storage_utilization | average | percentage |
119
+| azure_monitor.mysql_flexible.history_list_length | maximum | entries |
120
+
121
+
122
+
123
+## Alerts
124
+
125
+
126
+The following alerts are available:
127
+
128
+| Alert name | On metric | Description |
129
+|:------------|:----------|:------------|
130
+| [ am_mysql_flexible_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.cpu | MySQL Flexible CPU on ${label:resource_name} |
131
+| [ am_mysql_flexible_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.memory | MySQL Flexible memory on ${label:resource_name} |
132
+| [ am_mysql_flexible_io_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.io_utilization | MySQL Flexible I/O utilization on ${label:resource_name} |
133
+| [ am_mysql_flexible_storage_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.storage_utilization | MySQL Flexible storage on ${label:resource_name} |
134
+| [ am_mysql_flexible_serverlog_storage_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.serverlog_storage_utilization | MySQL Flexible server log storage on ${label:resource_name} |
135
+| [ am_mysql_flexible_cpu_credits_remaining ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.cpu_credits | MySQL Flexible CPU credits low on ${label:resource_name} |
136
+| [ am_mysql_flexible_ha_io_status ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.ha_status | MySQL Flexible HA IO thread down on ${label:resource_name} |
137
+| [ am_mysql_flexible_ha_sql_status ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.ha_status | MySQL Flexible HA SQL thread down on ${label:resource_name} |
138
+| [ am_mysql_flexible_replica_io_status ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.replica_status | MySQL Flexible replica IO thread down on ${label:resource_name} |
139
+| [ am_mysql_flexible_replica_sql_status ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.replica_status | MySQL Flexible replica SQL thread down on ${label:resource_name} |
140
+| [ am_mysql_flexible_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.replication_lag | MySQL Flexible replica lag on ${label:resource_name} |
141
+| [ am_mysql_flexible_ha_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.replication_lag | MySQL Flexible HA replication lag on ${label:resource_name} |
142
+| [ am_mysql_flexible_innodb_row_lock_time ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.innodb_row_lock_time | MySQL Flexible row lock wait time on ${label:resource_name} |
143
+| [ am_mysql_flexible_lock_deadlocks ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.lock_deadlocks | MySQL Flexible deadlocks on ${label:resource_name} |
144
+| [ am_mysql_flexible_lock_timeouts ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.lock_timeouts | MySQL Flexible lock timeouts on ${label:resource_name} |
145
+| [ am_mysql_flexible_aborted_connections ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.aborted_connections | MySQL Flexible aborted connections on ${label:resource_name} |
146
+| [ am_mysql_flexible_slow_queries ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.queries | MySQL Flexible slow queries on ${label:resource_name} |
147
+| [ am_mysql_flexible_innodb_row_lock_waits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.innodb_row_lock_waits | MySQL Flexible row lock waits on ${label:resource_name} |
148
+| [ am_mysql_flexible_history_list_length ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_mysql_flexible.conf) | azure_monitor.mysql_flexible.history_list_length | MySQL Flexible history list length on ${label:resource_name} |
149
+
150
+
151
+## Setup
152
+
153
+
154
+You can configure the **azure_monitor** collector in two ways:
155
+
156
+| Method | Best for | How to |
157
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
158
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
159
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
160
+
161
+:::important
162
+
163
+UI configuration requires paid Netdata Cloud plan.
164
+
165
+:::
166
+
167
+
168
+### Prerequisites
169
+
170
+#### Create an Azure monitoring principal
171
+
172
+Create a service principal or use a managed identity with the following permissions:
173
+
174
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
175
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
176
+
177
+For service principal authentication:
178
+```bash
179
+# Create the service principal
180
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
181
+ --scopes /subscriptions/<subscription-id>
182
+
183
+# Note the appId (client_id), password (client_secret), and tenant
184
+```
185
+
186
+For managed identity (on Azure VMs, VMSS, or AKS):
187
+```bash
188
+# Assign Monitoring Reader role to the VM's managed identity
189
+az role assignment create --assignee <managed-identity-principal-id> \
190
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
191
+```
192
+
193
+
194
+
195
+### Configuration
196
+
197
+#### Options
198
+
199
+The following options can be defined globally: update_every, autodetection_retry.
200
+
201
+Profile files are loaded from:
202
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
203
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
204
+
205
+User profile files with the same filename override stock profiles.
206
+
207
+
208
+<details open><summary>Config options</summary>
209
+
210
+
211
+
212
+| Group | Option | Description | Default | Required |
213
+|:------|:-----|:------------|:--------|:---------:|
214
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
215
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
216
+| **Target** | subscription_id | Azure subscription ID. | | yes |
217
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
218
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
219
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
220
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
221
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
222
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
223
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
224
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
225
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
226
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
227
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
228
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
229
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
230
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
231
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
232
+
233
+
234
+</details>
235
+
236
+
237
+#### via UI
238
+
239
+Configure the **azure_monitor** collector from the Netdata web interface:
240
+
241
+1. Go to **Nodes**.
242
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
243
+3. The **Collectors → Jobs** view opens by default.
244
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
245
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
246
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
247
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
248
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
249
+
250
+
251
+#### via File
252
+
253
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
254
+
255
+The file format is YAML. Generally, the structure is:
256
+
257
+```yaml
258
+update_every: 1
259
+autodetection_retry: 0
260
+jobs:
261
+ - name: some_name1
262
+ - name: some_name2
263
+```
264
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
265
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
266
+
267
+```bash
268
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
269
+sudo ./edit-config go.d/azure_monitor.conf
270
+```
271
+
272
+##### Examples
273
+
274
+###### Service principal (auto-discover all resources)
275
+
276
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
277
+
278
+```yaml
279
+jobs:
280
+ - name: prod
281
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
282
+ auth:
283
+ mode: service_principal
284
+ mode_service_principal:
285
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
286
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
287
+ client_secret: "your-client-secret"
288
+
289
+```
290
+###### Managed identity (Azure VM/VMSS/AKS)
291
+
292
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
293
+
294
+<details open><summary>Config</summary>
295
+
296
+```yaml
297
+jobs:
298
+ - name: prod
299
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
300
+ auth:
301
+ mode: managed_identity
302
+
303
+```
304
+</details>
305
+
306
+###### Specific profiles only
307
+
308
+Monitor only specific Azure services instead of auto-discovering all resource types.
309
+
310
+<details open><summary>Config</summary>
311
+
312
+```yaml
313
+jobs:
314
+ - name: databases
315
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
316
+ profiles:
317
+ - sql_database
318
+ - postgres_flexible
319
+ - redis_cache
320
+ auth:
321
+ mode: service_principal
322
+ mode_service_principal:
323
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ client_secret: "your-client-secret"
326
+
327
+```
328
+</details>
329
+
330
+###### Filter by resource group
331
+
332
+Only monitor resources in specific resource groups.
333
+
334
+<details open><summary>Config</summary>
335
+
336
+```yaml
337
+jobs:
338
+ - name: prod-rg
339
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
340
+ resource_groups:
341
+ - production-rg
342
+ - staging-rg
343
+ auth:
344
+ mode: default
345
+
346
+```
347
+</details>
348
+
349
+###### Azure Government cloud
350
+
351
+Connect to Azure Government cloud environment.
352
+
353
+<details open><summary>Config</summary>
354
+
355
+```yaml
356
+jobs:
357
+ - name: gov
358
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
+ cloud: government
360
+ auth:
361
+ mode: service_principal
362
+ mode_service_principal:
363
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
364
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
365
+ client_secret: "your-client-secret"
366
+
367
+```
368
+</details>
369
+
370
+
371
+
372
+## Troubleshooting
373
+
374
+### Debug Mode
375
+
376
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
377
+
378
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
379
+should give you clues as to why the collector isn't working.
380
+
381
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
382
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
383
+
384
+ ```bash
385
+ cd /usr/libexec/netdata/plugins.d/
386
+ ```
387
+
388
+- Switch to the `netdata` user.
389
+
390
+ ```bash
391
+ sudo -u netdata -s
392
+ ```
393
+
394
+- Run the `go.d.plugin` to debug the collector:
395
+
396
+ ```bash
397
+ ./go.d.plugin -d -m azure_monitor
398
+ ```
399
+
400
+ To debug a specific job:
401
+
402
+ ```bash
403
+ ./go.d.plugin -d -m azure_monitor -j jobName
404
+ ```
405
+
406
+### Getting Logs
407
+
408
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
409
+
410
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
411
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
412
+
413
+#### System with systemd
414
+
415
+Use the following command to view logs generated since the last Netdata service restart:
416
+
417
+```bash
418
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
419
+```
420
+
421
+#### System without systemd
422
+
423
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
424
+
425
+```bash
426
+grep azure_monitor /var/log/netdata/collector.log
427
+```
428
+
429
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
430
+
431
+#### Docker Container
432
+
433
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
434
+
435
+```bash
436
+docker logs netdata 2>&1 | grep azure_monitor
437
+```
438
+
439
+### No metrics are collected
440
+
441
+Verify the following:
442
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
443
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
444
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
445
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
446
+
447
+
448
+### Missing metrics for some resource types
449
+
450
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
451
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
452
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
453
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
454
+
455
+
456
+### Metrics appear delayed
457
+
458
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
459
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
460
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
461
+
462
+
463
+### Authentication errors in sovereign clouds
464
+
465
+For Azure Government or Azure China clouds, set the `cloud` parameter:
466
+- Azure Government: `cloud: government`
467
+- Azure China (21Vianet): `cloud: china`
468
+
469
+Ensure the service principal is registered in the correct cloud tenant.
470
+
471
+
472
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_nat_gateway.md
new
+430
@@ -0,0 +1,430 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_nat_gateway.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure NAT Gateway"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'nat', 'gateway', 'networking']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure NAT Gateway
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor NAT Gateway including byte and packet counts, connection counts, dropped packets, total SNAT connection counts, and datapath availability.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.nat_gateway.datapath_availability | average | percentage |
88
+| azure_monitor.nat_gateway.byte_throughput | total | bytes/s |
89
+| azure_monitor.nat_gateway.packet_throughput | total | packets/s |
90
+| azure_monitor.nat_gateway.dropped_packets | total | packets/s |
91
+| azure_monitor.nat_gateway.snat_connections | total | connections/s |
92
+| azure_monitor.nat_gateway.total_connections | total | connections/s |
93
+
94
+
95
+
96
+## Alerts
97
+
98
+
99
+The following alerts are available:
100
+
101
+| Alert name | On metric | Description |
102
+|:------------|:----------|:------------|
103
+| [ am_nat_gateway_datapath_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_nat_gateway.conf) | azure_monitor.nat_gateway.datapath_availability | NAT Gateway availability on ${label:resource_name} |
104
+| [ am_nat_gateway_dropped_packets ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_nat_gateway.conf) | azure_monitor.nat_gateway.dropped_packets | NAT Gateway packet drops on ${label:resource_name} |
105
+| [ am_nat_gateway_snat_connections ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_nat_gateway.conf) | azure_monitor.nat_gateway.snat_connections | NAT Gateway SNAT connections on ${label:resource_name} |
106
+| [ am_nat_gateway_total_connections ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_nat_gateway.conf) | azure_monitor.nat_gateway.total_connections | NAT Gateway total connections on ${label:resource_name} |
107
+
108
+
109
+## Setup
110
+
111
+
112
+You can configure the **azure_monitor** collector in two ways:
113
+
114
+| Method | Best for | How to |
115
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
116
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
117
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
118
+
119
+:::important
120
+
121
+UI configuration requires paid Netdata Cloud plan.
122
+
123
+:::
124
+
125
+
126
+### Prerequisites
127
+
128
+#### Create an Azure monitoring principal
129
+
130
+Create a service principal or use a managed identity with the following permissions:
131
+
132
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
133
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
134
+
135
+For service principal authentication:
136
+```bash
137
+# Create the service principal
138
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
139
+ --scopes /subscriptions/<subscription-id>
140
+
141
+# Note the appId (client_id), password (client_secret), and tenant
142
+```
143
+
144
+For managed identity (on Azure VMs, VMSS, or AKS):
145
+```bash
146
+# Assign Monitoring Reader role to the VM's managed identity
147
+az role assignment create --assignee <managed-identity-principal-id> \
148
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+```
150
+
151
+
152
+
153
+### Configuration
154
+
155
+#### Options
156
+
157
+The following options can be defined globally: update_every, autodetection_retry.
158
+
159
+Profile files are loaded from:
160
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
161
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+
163
+User profile files with the same filename override stock profiles.
164
+
165
+
166
+<details open><summary>Config options</summary>
167
+
168
+
169
+
170
+| Group | Option | Description | Default | Required |
171
+|:------|:-----|:------------|:--------|:---------:|
172
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
173
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
174
+| **Target** | subscription_id | Azure subscription ID. | | yes |
175
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
176
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
177
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
179
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
180
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
181
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
182
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
183
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
184
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
190
+
191
+
192
+</details>
193
+
194
+
195
+#### via UI
196
+
197
+Configure the **azure_monitor** collector from the Netdata web interface:
198
+
199
+1. Go to **Nodes**.
200
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
201
+3. The **Collectors → Jobs** view opens by default.
202
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
203
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
204
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
205
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
206
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
207
+
208
+
209
+#### via File
210
+
211
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
212
+
213
+The file format is YAML. Generally, the structure is:
214
+
215
+```yaml
216
+update_every: 1
217
+autodetection_retry: 0
218
+jobs:
219
+ - name: some_name1
220
+ - name: some_name2
221
+```
222
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
223
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
224
+
225
+```bash
226
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
227
+sudo ./edit-config go.d/azure_monitor.conf
228
+```
229
+
230
+##### Examples
231
+
232
+###### Service principal (auto-discover all resources)
233
+
234
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
235
+
236
+```yaml
237
+jobs:
238
+ - name: prod
239
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
240
+ auth:
241
+ mode: service_principal
242
+ mode_service_principal:
243
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
244
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
245
+ client_secret: "your-client-secret"
246
+
247
+```
248
+###### Managed identity (Azure VM/VMSS/AKS)
249
+
250
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
251
+
252
+<details open><summary>Config</summary>
253
+
254
+```yaml
255
+jobs:
256
+ - name: prod
257
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
258
+ auth:
259
+ mode: managed_identity
260
+
261
+```
262
+</details>
263
+
264
+###### Specific profiles only
265
+
266
+Monitor only specific Azure services instead of auto-discovering all resource types.
267
+
268
+<details open><summary>Config</summary>
269
+
270
+```yaml
271
+jobs:
272
+ - name: databases
273
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
274
+ profiles:
275
+ - sql_database
276
+ - postgres_flexible
277
+ - redis_cache
278
+ auth:
279
+ mode: service_principal
280
+ mode_service_principal:
281
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
282
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ client_secret: "your-client-secret"
284
+
285
+```
286
+</details>
287
+
288
+###### Filter by resource group
289
+
290
+Only monitor resources in specific resource groups.
291
+
292
+<details open><summary>Config</summary>
293
+
294
+```yaml
295
+jobs:
296
+ - name: prod-rg
297
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
298
+ resource_groups:
299
+ - production-rg
300
+ - staging-rg
301
+ auth:
302
+ mode: default
303
+
304
+```
305
+</details>
306
+
307
+###### Azure Government cloud
308
+
309
+Connect to Azure Government cloud environment.
310
+
311
+<details open><summary>Config</summary>
312
+
313
+```yaml
314
+jobs:
315
+ - name: gov
316
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
317
+ cloud: government
318
+ auth:
319
+ mode: service_principal
320
+ mode_service_principal:
321
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ client_secret: "your-client-secret"
324
+
325
+```
326
+</details>
327
+
328
+
329
+
330
+## Troubleshooting
331
+
332
+### Debug Mode
333
+
334
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
335
+
336
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
337
+should give you clues as to why the collector isn't working.
338
+
339
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
340
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
341
+
342
+ ```bash
343
+ cd /usr/libexec/netdata/plugins.d/
344
+ ```
345
+
346
+- Switch to the `netdata` user.
347
+
348
+ ```bash
349
+ sudo -u netdata -s
350
+ ```
351
+
352
+- Run the `go.d.plugin` to debug the collector:
353
+
354
+ ```bash
355
+ ./go.d.plugin -d -m azure_monitor
356
+ ```
357
+
358
+ To debug a specific job:
359
+
360
+ ```bash
361
+ ./go.d.plugin -d -m azure_monitor -j jobName
362
+ ```
363
+
364
+### Getting Logs
365
+
366
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
367
+
368
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
369
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
370
+
371
+#### System with systemd
372
+
373
+Use the following command to view logs generated since the last Netdata service restart:
374
+
375
+```bash
376
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
377
+```
378
+
379
+#### System without systemd
380
+
381
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
382
+
383
+```bash
384
+grep azure_monitor /var/log/netdata/collector.log
385
+```
386
+
387
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
388
+
389
+#### Docker Container
390
+
391
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
392
+
393
+```bash
394
+docker logs netdata 2>&1 | grep azure_monitor
395
+```
396
+
397
+### No metrics are collected
398
+
399
+Verify the following:
400
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
401
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
402
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
403
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
404
+
405
+
406
+### Missing metrics for some resource types
407
+
408
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
409
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
410
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
411
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
412
+
413
+
414
+### Metrics appear delayed
415
+
416
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
417
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
418
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
419
+
420
+
421
+### Authentication errors in sovereign clouds
422
+
423
+For Azure Government or Azure China clouds, set the `cloud` parameter:
424
+- Azure Government: `cloud: government`
425
+- Azure China (21Vianet): `cloud: china`
426
+
427
+Ensure the service principal is registered in the correct cloud tenant.
428
+
429
+
430
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_postgresql_flexible_server.md
new
+482
@@ -0,0 +1,482 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_postgresql_flexible_server.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure PostgreSQL Flexible Server"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'postgresql', 'postgres', 'database', 'flexible']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure PostgreSQL Flexible Server
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor PostgreSQL Flexible Server including active connections, transaction rates, replication lag, storage and backup utilization, CPU and memory usage, IO throughput, autovacuum activity, PgBouncer connection pooling, database sessions, and burstable instance CPU credits.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.postgres_flexible.cpu | average | percentage |
88
+| azure_monitor.postgres_flexible.memory | average | percentage |
89
+| azure_monitor.postgres_flexible.availability | maximum | state |
90
+| azure_monitor.postgres_flexible.iops | total, read, write | operations/s |
91
+| azure_monitor.postgres_flexible.disk_throughput | read, write | bytes/s |
92
+| azure_monitor.postgres_flexible.disk_saturation | bandwidth, iops | percentage |
93
+| azure_monitor.postgres_flexible.disk_queue_depth | average | operations |
94
+| azure_monitor.postgres_flexible.buffer_cache | hits, reads | blocks/s |
95
+| azure_monitor.postgres_flexible.active_connections | average | connections |
96
+| azure_monitor.postgres_flexible.connection_rate | succeeded, failed | connections/s |
97
+| azure_monitor.postgres_flexible.transactions | committed, rolled_back | transactions/s |
98
+| azure_monitor.postgres_flexible.transaction_rate | total, xact_total | transactions/s |
99
+| azure_monitor.postgres_flexible.deadlocks | total | deadlocks/s |
100
+| azure_monitor.postgres_flexible.tuple_reads | returned, fetched | tuples/s |
101
+| azure_monitor.postgres_flexible.tuple_writes | inserted, updated, deleted | tuples/s |
102
+| azure_monitor.postgres_flexible.temp_files | total | files/s |
103
+| azure_monitor.postgres_flexible.temp_bytes | total | bytes/s |
104
+| azure_monitor.postgres_flexible.long_running | query, transaction | seconds |
105
+| azure_monitor.postgres_flexible.xid_usage | max_used, oldest_xmin | transactions |
106
+| azure_monitor.postgres_flexible.xmin_age | maximum | transactions |
107
+| azure_monitor.postgres_flexible.network | in, out | bytes/s |
108
+| azure_monitor.postgres_flexible.storage | used, free | bytes |
109
+| azure_monitor.postgres_flexible.storage_utilization | average | percentage |
110
+| azure_monitor.postgres_flexible.wal_storage | used | bytes |
111
+| azure_monitor.postgres_flexible.database_size | average | bytes |
112
+| azure_monitor.postgres_flexible.backup_storage | used | bytes |
113
+| azure_monitor.postgres_flexible.replication_lag_time | average | seconds |
114
+| azure_monitor.postgres_flexible.replication_lag_bytes | physical, logical | bytes |
115
+| azure_monitor.postgres_flexible.sessions_by_state | maximum | sessions |
116
+| azure_monitor.postgres_flexible.sessions_by_wait_event_type | maximum | sessions |
117
+| azure_monitor.postgres_flexible.autovacuum_operations | vacuum, autovacuum, analyze, autoanalyze | operations |
118
+| azure_monitor.postgres_flexible.autovacuum_table_coverage | vacuumed, autovacuumed, analyzed, autoanalyzed, total | tables |
119
+| azure_monitor.postgres_flexible.tuple_liveness | live, dead, modified_since_analyze | tuples |
120
+| azure_monitor.postgres_flexible.bloat | maximum | percentage |
121
+| azure_monitor.postgres_flexible.backend_count | maximum | connections |
122
+| azure_monitor.postgres_flexible.pgbouncer_client_connections | active, waiting | connections |
123
+| azure_monitor.postgres_flexible.pgbouncer_server_connections | active, idle | connections |
124
+| azure_monitor.postgres_flexible.pgbouncer_pool_count | pools | pools |
125
+| azure_monitor.postgres_flexible.pgbouncer_pooled_connections | pooled_connections | connections |
126
+| azure_monitor.postgres_flexible.cpu_credits | consumed, remaining | credits |
127
+| azure_monitor.postgres_flexible.postmaster_cpu | average | percentage |
128
+| azure_monitor.postgres_flexible.max_connections | maximum | connections |
129
+| azure_monitor.postgres_flexible.tcp_connection_backlog | maximum | connections |
130
+
131
+
132
+
133
+## Alerts
134
+
135
+
136
+The following alerts are available:
137
+
138
+| Alert name | On metric | Description |
139
+|:------------|:----------|:------------|
140
+| [ am_postgres_flexible_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.availability | PostgreSQL Flexible Server down on ${label:resource_name} |
141
+| [ am_postgres_flexible_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.cpu | PostgreSQL Flexible CPU on ${label:resource_name} |
142
+| [ am_postgres_flexible_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.memory | PostgreSQL Flexible memory on ${label:resource_name} |
143
+| [ am_postgres_flexible_storage_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.storage_utilization | PostgreSQL Flexible storage on ${label:resource_name} |
144
+| [ am_postgres_flexible_disk_bandwidth_saturation ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.disk_saturation | PostgreSQL Flexible disk bandwidth saturation on ${label:resource_name} |
145
+| [ am_postgres_flexible_disk_iops_saturation ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.disk_saturation | PostgreSQL Flexible disk IOPS saturation on ${label:resource_name} |
146
+| [ am_postgres_flexible_disk_queue_depth ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.disk_queue_depth | PostgreSQL Flexible disk queue depth on ${label:resource_name} |
147
+| [ am_postgres_flexible_failed_connections ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.connection_rate | PostgreSQL Flexible failed connections on ${label:resource_name} |
148
+| [ am_postgres_flexible_tcp_connection_backlog ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.tcp_connection_backlog | PostgreSQL Flexible TCP connection backlog on ${label:resource_name} |
149
+| [ am_postgres_flexible_deadlocks ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.deadlocks | PostgreSQL Flexible deadlocks on ${label:resource_name} |
150
+| [ am_postgres_flexible_rollback_ratio ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.transactions | PostgreSQL Flexible rollback ratio on ${label:resource_name} |
151
+| [ am_postgres_flexible_longest_query ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.long_running | PostgreSQL Flexible long running query on ${label:resource_name} |
152
+| [ am_postgres_flexible_longest_transaction ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.long_running | PostgreSQL Flexible long running transaction on ${label:resource_name} |
153
+| [ am_postgres_flexible_xid_usage ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.xid_usage | PostgreSQL Flexible transaction ID usage on ${label:resource_name} |
154
+| [ am_postgres_flexible_xmin_age ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.xmin_age | PostgreSQL Flexible backend xmin age on ${label:resource_name} |
155
+| [ am_postgres_flexible_bloat ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.bloat | PostgreSQL Flexible table bloat on ${label:resource_name} |
156
+| [ am_postgres_flexible_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.replication_lag_time | PostgreSQL Flexible replication lag on ${label:resource_name} |
157
+| [ am_postgres_flexible_cpu_credits_remaining ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.cpu_credits | PostgreSQL Flexible CPU credits low on ${label:resource_name} |
158
+| [ am_postgres_flexible_temp_bytes ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_postgres_flexible.conf) | azure_monitor.postgres_flexible.temp_bytes | PostgreSQL Flexible temp file I/O on ${label:resource_name} |
159
+
160
+
161
+## Setup
162
+
163
+
164
+You can configure the **azure_monitor** collector in two ways:
165
+
166
+| Method | Best for | How to |
167
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
168
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
169
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
170
+
171
+:::important
172
+
173
+UI configuration requires paid Netdata Cloud plan.
174
+
175
+:::
176
+
177
+
178
+### Prerequisites
179
+
180
+#### Create an Azure monitoring principal
181
+
182
+Create a service principal or use a managed identity with the following permissions:
183
+
184
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
185
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
186
+
187
+For service principal authentication:
188
+```bash
189
+# Create the service principal
190
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
191
+ --scopes /subscriptions/<subscription-id>
192
+
193
+# Note the appId (client_id), password (client_secret), and tenant
194
+```
195
+
196
+For managed identity (on Azure VMs, VMSS, or AKS):
197
+```bash
198
+# Assign Monitoring Reader role to the VM's managed identity
199
+az role assignment create --assignee <managed-identity-principal-id> \
200
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
201
+```
202
+
203
+
204
+
205
+### Configuration
206
+
207
+#### Options
208
+
209
+The following options can be defined globally: update_every, autodetection_retry.
210
+
211
+Profile files are loaded from:
212
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
213
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
214
+
215
+User profile files with the same filename override stock profiles.
216
+
217
+
218
+<details open><summary>Config options</summary>
219
+
220
+
221
+
222
+| Group | Option | Description | Default | Required |
223
+|:------|:-----|:------------|:--------|:---------:|
224
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
225
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
226
+| **Target** | subscription_id | Azure subscription ID. | | yes |
227
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
228
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
229
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
230
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
231
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
232
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
233
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
234
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
235
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
236
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
237
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
238
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
239
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
240
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
241
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
242
+
243
+
244
+</details>
245
+
246
+
247
+#### via UI
248
+
249
+Configure the **azure_monitor** collector from the Netdata web interface:
250
+
251
+1. Go to **Nodes**.
252
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
253
+3. The **Collectors → Jobs** view opens by default.
254
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
255
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
256
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
257
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
258
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
259
+
260
+
261
+#### via File
262
+
263
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
264
+
265
+The file format is YAML. Generally, the structure is:
266
+
267
+```yaml
268
+update_every: 1
269
+autodetection_retry: 0
270
+jobs:
271
+ - name: some_name1
272
+ - name: some_name2
273
+```
274
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
275
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
276
+
277
+```bash
278
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
279
+sudo ./edit-config go.d/azure_monitor.conf
280
+```
281
+
282
+##### Examples
283
+
284
+###### Service principal (auto-discover all resources)
285
+
286
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
287
+
288
+```yaml
289
+jobs:
290
+ - name: prod
291
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
292
+ auth:
293
+ mode: service_principal
294
+ mode_service_principal:
295
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
296
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
297
+ client_secret: "your-client-secret"
298
+
299
+```
300
+###### Managed identity (Azure VM/VMSS/AKS)
301
+
302
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
303
+
304
+<details open><summary>Config</summary>
305
+
306
+```yaml
307
+jobs:
308
+ - name: prod
309
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
310
+ auth:
311
+ mode: managed_identity
312
+
313
+```
314
+</details>
315
+
316
+###### Specific profiles only
317
+
318
+Monitor only specific Azure services instead of auto-discovering all resource types.
319
+
320
+<details open><summary>Config</summary>
321
+
322
+```yaml
323
+jobs:
324
+ - name: databases
325
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ profiles:
327
+ - sql_database
328
+ - postgres_flexible
329
+ - redis_cache
330
+ auth:
331
+ mode: service_principal
332
+ mode_service_principal:
333
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
334
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
335
+ client_secret: "your-client-secret"
336
+
337
+```
338
+</details>
339
+
340
+###### Filter by resource group
341
+
342
+Only monitor resources in specific resource groups.
343
+
344
+<details open><summary>Config</summary>
345
+
346
+```yaml
347
+jobs:
348
+ - name: prod-rg
349
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
350
+ resource_groups:
351
+ - production-rg
352
+ - staging-rg
353
+ auth:
354
+ mode: default
355
+
356
+```
357
+</details>
358
+
359
+###### Azure Government cloud
360
+
361
+Connect to Azure Government cloud environment.
362
+
363
+<details open><summary>Config</summary>
364
+
365
+```yaml
366
+jobs:
367
+ - name: gov
368
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
369
+ cloud: government
370
+ auth:
371
+ mode: service_principal
372
+ mode_service_principal:
373
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
374
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
375
+ client_secret: "your-client-secret"
376
+
377
+```
378
+</details>
379
+
380
+
381
+
382
+## Troubleshooting
383
+
384
+### Debug Mode
385
+
386
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
387
+
388
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
389
+should give you clues as to why the collector isn't working.
390
+
391
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
392
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
393
+
394
+ ```bash
395
+ cd /usr/libexec/netdata/plugins.d/
396
+ ```
397
+
398
+- Switch to the `netdata` user.
399
+
400
+ ```bash
401
+ sudo -u netdata -s
402
+ ```
403
+
404
+- Run the `go.d.plugin` to debug the collector:
405
+
406
+ ```bash
407
+ ./go.d.plugin -d -m azure_monitor
408
+ ```
409
+
410
+ To debug a specific job:
411
+
412
+ ```bash
413
+ ./go.d.plugin -d -m azure_monitor -j jobName
414
+ ```
415
+
416
+### Getting Logs
417
+
418
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
419
+
420
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
421
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
422
+
423
+#### System with systemd
424
+
425
+Use the following command to view logs generated since the last Netdata service restart:
426
+
427
+```bash
428
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
429
+```
430
+
431
+#### System without systemd
432
+
433
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
434
+
435
+```bash
436
+grep azure_monitor /var/log/netdata/collector.log
437
+```
438
+
439
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
440
+
441
+#### Docker Container
442
+
443
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
444
+
445
+```bash
446
+docker logs netdata 2>&1 | grep azure_monitor
447
+```
448
+
449
+### No metrics are collected
450
+
451
+Verify the following:
452
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
453
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
454
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
455
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
456
+
457
+
458
+### Missing metrics for some resource types
459
+
460
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
461
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
462
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
463
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
464
+
465
+
466
+### Metrics appear delayed
467
+
468
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
469
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
470
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
471
+
472
+
473
+### Authentication errors in sovereign clouds
474
+
475
+For Azure Government or Azure China clouds, set the `cloud` parameter:
476
+- Azure Government: `cloud: government`
477
+- Azure China (21Vianet): `cloud: china`
478
+
479
+Ensure the service principal is registered in the correct cloud tenant.
480
+
481
+
482
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_service_bus_namespace.md
new
+448
@@ -0,0 +1,448 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_service_bus_namespace.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Service Bus Namespace"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'service', 'bus', 'messaging', 'queue', 'topic']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Service Bus Namespace
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Service Bus namespaces including incoming and outgoing message rates, active connections, active and dead-lettered message counts, scheduled message counts, completed and abandoned requests, server errors, throttled requests, CPU and memory utilization, and pending checkpoint operations.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.service_bus.message_flow | in, out | messages/s |
88
+| azure_monitor.service_bus.message_operations | completed, abandoned | messages/s |
89
+| azure_monitor.service_bus.queue_depth | active | messages |
90
+| azure_monitor.service_bus.problem_messages | dead_lettered, scheduled | messages |
91
+| azure_monitor.service_bus.requests | incoming, successful | requests/s |
92
+| azure_monitor.service_bus.errors | server, user, throttled | errors/s |
93
+| azure_monitor.service_bus.connections | active | connections |
94
+| azure_monitor.service_bus.connection_events | opened, closed | connections |
95
+| azure_monitor.service_bus.namespace_size | average | bytes |
96
+| azure_monitor.service_bus.data_throughput | in, out | bytes/s |
97
+| azure_monitor.service_bus.total_messages | total | messages |
98
+| azure_monitor.service_bus.namespace_resources | cpu, memory | percentage |
99
+| azure_monitor.service_bus.send_latency | average | milliseconds |
100
+| azure_monitor.service_bus.replication_lag | messages | messages |
101
+| azure_monitor.service_bus.replication_lag_duration | duration | seconds |
102
+| azure_monitor.service_bus.checkpoint_operations | pending | operations |
103
+
104
+
105
+
106
+## Alerts
107
+
108
+
109
+The following alerts are available:
110
+
111
+| Alert name | On metric | Description |
112
+|:------------|:----------|:------------|
113
+| [ am_service_bus_server_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.errors | Service Bus server errors on ${label:resource_name} |
114
+| [ am_service_bus_throttled_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.errors | Service Bus throttled requests on ${label:resource_name} |
115
+| [ am_service_bus_user_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.errors | Service Bus user errors on ${label:resource_name} |
116
+| [ am_service_bus_namespace_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.namespace_resources | Service Bus namespace CPU on ${label:resource_name} |
117
+| [ am_service_bus_namespace_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.namespace_resources | Service Bus namespace memory on ${label:resource_name} |
118
+| [ am_service_bus_send_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.send_latency | Service Bus send latency on ${label:resource_name} |
119
+| [ am_service_bus_dead_lettered_messages ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.problem_messages | Service Bus dead-lettered messages on ${label:resource_name} |
120
+| [ am_service_bus_active_messages ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.queue_depth | Service Bus queue depth on ${label:resource_name} |
121
+| [ am_service_bus_request_success_rate ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.requests | Service Bus request success rate on ${label:resource_name} |
122
+| [ am_service_bus_abandoned_messages ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.message_operations | Service Bus abandoned messages on ${label:resource_name} |
123
+| [ am_service_bus_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.replication_lag | Service Bus replication lag on ${label:resource_name} |
124
+| [ am_service_bus_replication_lag_duration ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_service_bus.conf) | azure_monitor.service_bus.replication_lag_duration | Service Bus replication lag duration on ${label:resource_name} |
125
+
126
+
127
+## Setup
128
+
129
+
130
+You can configure the **azure_monitor** collector in two ways:
131
+
132
+| Method | Best for | How to |
133
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
134
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
135
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
136
+
137
+:::important
138
+
139
+UI configuration requires paid Netdata Cloud plan.
140
+
141
+:::
142
+
143
+
144
+### Prerequisites
145
+
146
+#### Create an Azure monitoring principal
147
+
148
+Create a service principal or use a managed identity with the following permissions:
149
+
150
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
151
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
152
+
153
+For service principal authentication:
154
+```bash
155
+# Create the service principal
156
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
157
+ --scopes /subscriptions/<subscription-id>
158
+
159
+# Note the appId (client_id), password (client_secret), and tenant
160
+```
161
+
162
+For managed identity (on Azure VMs, VMSS, or AKS):
163
+```bash
164
+# Assign Monitoring Reader role to the VM's managed identity
165
+az role assignment create --assignee <managed-identity-principal-id> \
166
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
167
+```
168
+
169
+
170
+
171
+### Configuration
172
+
173
+#### Options
174
+
175
+The following options can be defined globally: update_every, autodetection_retry.
176
+
177
+Profile files are loaded from:
178
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
179
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
180
+
181
+User profile files with the same filename override stock profiles.
182
+
183
+
184
+<details open><summary>Config options</summary>
185
+
186
+
187
+
188
+| Group | Option | Description | Default | Required |
189
+|:------|:-----|:------------|:--------|:---------:|
190
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
191
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
192
+| **Target** | subscription_id | Azure subscription ID. | | yes |
193
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
194
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
195
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
196
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
197
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
201
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
202
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
203
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
204
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
205
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
206
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
207
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
208
+
209
+
210
+</details>
211
+
212
+
213
+#### via UI
214
+
215
+Configure the **azure_monitor** collector from the Netdata web interface:
216
+
217
+1. Go to **Nodes**.
218
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
219
+3. The **Collectors → Jobs** view opens by default.
220
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
221
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
222
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
223
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
224
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
225
+
226
+
227
+#### via File
228
+
229
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
230
+
231
+The file format is YAML. Generally, the structure is:
232
+
233
+```yaml
234
+update_every: 1
235
+autodetection_retry: 0
236
+jobs:
237
+ - name: some_name1
238
+ - name: some_name2
239
+```
240
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
241
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
242
+
243
+```bash
244
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
245
+sudo ./edit-config go.d/azure_monitor.conf
246
+```
247
+
248
+##### Examples
249
+
250
+###### Service principal (auto-discover all resources)
251
+
252
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
253
+
254
+```yaml
255
+jobs:
256
+ - name: prod
257
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
258
+ auth:
259
+ mode: service_principal
260
+ mode_service_principal:
261
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
262
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
263
+ client_secret: "your-client-secret"
264
+
265
+```
266
+###### Managed identity (Azure VM/VMSS/AKS)
267
+
268
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
269
+
270
+<details open><summary>Config</summary>
271
+
272
+```yaml
273
+jobs:
274
+ - name: prod
275
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
276
+ auth:
277
+ mode: managed_identity
278
+
279
+```
280
+</details>
281
+
282
+###### Specific profiles only
283
+
284
+Monitor only specific Azure services instead of auto-discovering all resource types.
285
+
286
+<details open><summary>Config</summary>
287
+
288
+```yaml
289
+jobs:
290
+ - name: databases
291
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
292
+ profiles:
293
+ - sql_database
294
+ - postgres_flexible
295
+ - redis_cache
296
+ auth:
297
+ mode: service_principal
298
+ mode_service_principal:
299
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
300
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
301
+ client_secret: "your-client-secret"
302
+
303
+```
304
+</details>
305
+
306
+###### Filter by resource group
307
+
308
+Only monitor resources in specific resource groups.
309
+
310
+<details open><summary>Config</summary>
311
+
312
+```yaml
313
+jobs:
314
+ - name: prod-rg
315
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
316
+ resource_groups:
317
+ - production-rg
318
+ - staging-rg
319
+ auth:
320
+ mode: default
321
+
322
+```
323
+</details>
324
+
325
+###### Azure Government cloud
326
+
327
+Connect to Azure Government cloud environment.
328
+
329
+<details open><summary>Config</summary>
330
+
331
+```yaml
332
+jobs:
333
+ - name: gov
334
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
335
+ cloud: government
336
+ auth:
337
+ mode: service_principal
338
+ mode_service_principal:
339
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
340
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
341
+ client_secret: "your-client-secret"
342
+
343
+```
344
+</details>
345
+
346
+
347
+
348
+## Troubleshooting
349
+
350
+### Debug Mode
351
+
352
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
353
+
354
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
355
+should give you clues as to why the collector isn't working.
356
+
357
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
358
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
359
+
360
+ ```bash
361
+ cd /usr/libexec/netdata/plugins.d/
362
+ ```
363
+
364
+- Switch to the `netdata` user.
365
+
366
+ ```bash
367
+ sudo -u netdata -s
368
+ ```
369
+
370
+- Run the `go.d.plugin` to debug the collector:
371
+
372
+ ```bash
373
+ ./go.d.plugin -d -m azure_monitor
374
+ ```
375
+
376
+ To debug a specific job:
377
+
378
+ ```bash
379
+ ./go.d.plugin -d -m azure_monitor -j jobName
380
+ ```
381
+
382
+### Getting Logs
383
+
384
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
385
+
386
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
387
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
388
+
389
+#### System with systemd
390
+
391
+Use the following command to view logs generated since the last Netdata service restart:
392
+
393
+```bash
394
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
395
+```
396
+
397
+#### System without systemd
398
+
399
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
400
+
401
+```bash
402
+grep azure_monitor /var/log/netdata/collector.log
403
+```
404
+
405
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
406
+
407
+#### Docker Container
408
+
409
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
410
+
411
+```bash
412
+docker logs netdata 2>&1 | grep azure_monitor
413
+```
414
+
415
+### No metrics are collected
416
+
417
+Verify the following:
418
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
419
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
420
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
421
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
422
+
423
+
424
+### Missing metrics for some resource types
425
+
426
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
427
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
428
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
429
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
430
+
431
+
432
+### Metrics appear delayed
433
+
434
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
435
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
436
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
437
+
438
+
439
+### Authentication errors in sovereign clouds
440
+
441
+For Azure Government or Azure China clouds, set the `cloud` parameter:
442
+- Azure Government: `cloud: government`
443
+- Azure China (21Vianet): `cloud: china`
444
+
445
+Ensure the service principal is registered in the correct cloud tenant.
446
+
447
+
448
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_database.md
new
+461
@@ -0,0 +1,461 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_database.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure SQL Database"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'sql', 'database', 'mssql', 'dtu']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure SQL Database
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor SQL Database performance including CPU and DTU utilization, storage consumption, active sessions and workers, deadlocks, IO rates, tempdb usage, in-memory OLTP storage, and serverless auto-pause and billing metrics.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.sql_database.cpu | average, maximum | percentage |
88
+| azure_monitor.sql_database.instance_cpu | average | percentage |
89
+| azure_monitor.sql_database.instance_memory | average | percentage |
90
+| azure_monitor.sql_database.dtu_consumption | average | percentage |
91
+| azure_monitor.sql_database.io_utilization | data_read, log_write | percentage |
92
+| azure_monitor.sql_database.resource_utilization | workers, sessions | percentage |
93
+| azure_monitor.sql_database.availability | average | percentage |
94
+| azure_monitor.sql_database.connections | successful, failed_system, failed_user, firewall_blocked | connections/s |
95
+| azure_monitor.sql_database.deadlocks | total | deadlocks/s |
96
+| azure_monitor.sql_database.storage | used, allocated | bytes |
97
+| azure_monitor.sql_database.storage_utilization | average | percentage |
98
+| azure_monitor.sql_database.tempdb_size | data, log | KiB |
99
+| azure_monitor.sql_database.tempdb_log_utilization | average | percentage |
100
+| azure_monitor.sql_database.vcore_usage | used, limit | vCores |
101
+| azure_monitor.sql_database.dtu_usage | used, limit | DTU |
102
+| azure_monitor.sql_database.sessions_count | average | sessions |
103
+| azure_monitor.sql_database.replication_lag | average | seconds |
104
+| azure_monitor.sql_database.xtp_storage | average | percentage |
105
+| azure_monitor.sql_database.serverless_utilization | cpu, memory | percentage |
106
+| azure_monitor.sql_database.serverless_billing | total, ha_replicas | vCore-seconds/s |
107
+| azure_monitor.sql_database.free_tier_usage | consumed, remaining | vCore-seconds |
108
+| azure_monitor.sql_database.ledger_digest | success, failed | events/s |
109
+
110
+
111
+
112
+## Alerts
113
+
114
+
115
+The following alerts are available:
116
+
117
+| Alert name | On metric | Description |
118
+|:------------|:----------|:------------|
119
+| [ am_sql_database_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.availability | SQL Database availability on ${label:resource_name} |
120
+| [ am_sql_database_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.cpu | SQL Database CPU on ${label:resource_name} |
121
+| [ am_sql_database_instance_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.instance_cpu | SQL Database instance CPU on ${label:resource_name} |
122
+| [ am_sql_database_instance_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.instance_memory | SQL Database instance memory on ${label:resource_name} |
123
+| [ am_sql_database_dtu_consumption ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.dtu_consumption | SQL Database DTU consumption on ${label:resource_name} |
124
+| [ am_sql_database_data_io ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.io_utilization | SQL Database data I/O on ${label:resource_name} |
125
+| [ am_sql_database_log_write ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.io_utilization | SQL Database log write I/O on ${label:resource_name} |
126
+| [ am_sql_database_workers ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.resource_utilization | SQL Database worker utilization on ${label:resource_name} |
127
+| [ am_sql_database_sessions ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.resource_utilization | SQL Database session utilization on ${label:resource_name} |
128
+| [ am_sql_database_connection_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.connections | SQL Database system connection failures on ${label:resource_name} |
129
+| [ am_sql_database_firewall_blocks ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.connections | SQL Database firewall blocks on ${label:resource_name} |
130
+| [ am_sql_database_deadlocks ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.deadlocks | SQL Database deadlocks on ${label:resource_name} |
131
+| [ am_sql_database_storage_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.storage_utilization | SQL Database storage utilization on ${label:resource_name} |
132
+| [ am_sql_database_xtp_storage ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.xtp_storage | SQL Database In-Memory OLTP storage on ${label:resource_name} |
133
+| [ am_sql_database_tempdb_log_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.tempdb_log_utilization | SQL Database tempdb log utilization on ${label:resource_name} |
134
+| [ am_sql_database_replication_lag ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.replication_lag | SQL Database replication lag on ${label:resource_name} |
135
+| [ am_sql_database_serverless_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.serverless_utilization | SQL Database serverless CPU on ${label:resource_name} |
136
+| [ am_sql_database_serverless_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.serverless_utilization | SQL Database serverless memory on ${label:resource_name} |
137
+| [ am_sql_database_ledger_digest_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_database.conf) | azure_monitor.sql_database.ledger_digest | SQL Database ledger digest failures on ${label:resource_name} |
138
+
139
+
140
+## Setup
141
+
142
+
143
+You can configure the **azure_monitor** collector in two ways:
144
+
145
+| Method | Best for | How to |
146
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
147
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
148
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
149
+
150
+:::important
151
+
152
+UI configuration requires paid Netdata Cloud plan.
153
+
154
+:::
155
+
156
+
157
+### Prerequisites
158
+
159
+#### Create an Azure monitoring principal
160
+
161
+Create a service principal or use a managed identity with the following permissions:
162
+
163
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
164
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
165
+
166
+For service principal authentication:
167
+```bash
168
+# Create the service principal
169
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
170
+ --scopes /subscriptions/<subscription-id>
171
+
172
+# Note the appId (client_id), password (client_secret), and tenant
173
+```
174
+
175
+For managed identity (on Azure VMs, VMSS, or AKS):
176
+```bash
177
+# Assign Monitoring Reader role to the VM's managed identity
178
+az role assignment create --assignee <managed-identity-principal-id> \
179
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
180
+```
181
+
182
+
183
+
184
+### Configuration
185
+
186
+#### Options
187
+
188
+The following options can be defined globally: update_every, autodetection_retry.
189
+
190
+Profile files are loaded from:
191
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
192
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
193
+
194
+User profile files with the same filename override stock profiles.
195
+
196
+
197
+<details open><summary>Config options</summary>
198
+
199
+
200
+
201
+| Group | Option | Description | Default | Required |
202
+|:------|:-----|:------------|:--------|:---------:|
203
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
204
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
205
+| **Target** | subscription_id | Azure subscription ID. | | yes |
206
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
207
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
208
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
209
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
210
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
211
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
212
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
213
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
214
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
215
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
216
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
217
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
218
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
219
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
220
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
221
+
222
+
223
+</details>
224
+
225
+
226
+#### via UI
227
+
228
+Configure the **azure_monitor** collector from the Netdata web interface:
229
+
230
+1. Go to **Nodes**.
231
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
232
+3. The **Collectors → Jobs** view opens by default.
233
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
234
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
235
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
236
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
237
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
238
+
239
+
240
+#### via File
241
+
242
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
243
+
244
+The file format is YAML. Generally, the structure is:
245
+
246
+```yaml
247
+update_every: 1
248
+autodetection_retry: 0
249
+jobs:
250
+ - name: some_name1
251
+ - name: some_name2
252
+```
253
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
254
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
255
+
256
+```bash
257
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
258
+sudo ./edit-config go.d/azure_monitor.conf
259
+```
260
+
261
+##### Examples
262
+
263
+###### Service principal (auto-discover all resources)
264
+
265
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
266
+
267
+```yaml
268
+jobs:
269
+ - name: prod
270
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
271
+ auth:
272
+ mode: service_principal
273
+ mode_service_principal:
274
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
275
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
276
+ client_secret: "your-client-secret"
277
+
278
+```
279
+###### Managed identity (Azure VM/VMSS/AKS)
280
+
281
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
282
+
283
+<details open><summary>Config</summary>
284
+
285
+```yaml
286
+jobs:
287
+ - name: prod
288
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
289
+ auth:
290
+ mode: managed_identity
291
+
292
+```
293
+</details>
294
+
295
+###### Specific profiles only
296
+
297
+Monitor only specific Azure services instead of auto-discovering all resource types.
298
+
299
+<details open><summary>Config</summary>
300
+
301
+```yaml
302
+jobs:
303
+ - name: databases
304
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
305
+ profiles:
306
+ - sql_database
307
+ - postgres_flexible
308
+ - redis_cache
309
+ auth:
310
+ mode: service_principal
311
+ mode_service_principal:
312
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
313
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
314
+ client_secret: "your-client-secret"
315
+
316
+```
317
+</details>
318
+
319
+###### Filter by resource group
320
+
321
+Only monitor resources in specific resource groups.
322
+
323
+<details open><summary>Config</summary>
324
+
325
+```yaml
326
+jobs:
327
+ - name: prod-rg
328
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
329
+ resource_groups:
330
+ - production-rg
331
+ - staging-rg
332
+ auth:
333
+ mode: default
334
+
335
+```
336
+</details>
337
+
338
+###### Azure Government cloud
339
+
340
+Connect to Azure Government cloud environment.
341
+
342
+<details open><summary>Config</summary>
343
+
344
+```yaml
345
+jobs:
346
+ - name: gov
347
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
348
+ cloud: government
349
+ auth:
350
+ mode: service_principal
351
+ mode_service_principal:
352
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
+ client_secret: "your-client-secret"
355
+
356
+```
357
+</details>
358
+
359
+
360
+
361
+## Troubleshooting
362
+
363
+### Debug Mode
364
+
365
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
366
+
367
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
368
+should give you clues as to why the collector isn't working.
369
+
370
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
371
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
372
+
373
+ ```bash
374
+ cd /usr/libexec/netdata/plugins.d/
375
+ ```
376
+
377
+- Switch to the `netdata` user.
378
+
379
+ ```bash
380
+ sudo -u netdata -s
381
+ ```
382
+
383
+- Run the `go.d.plugin` to debug the collector:
384
+
385
+ ```bash
386
+ ./go.d.plugin -d -m azure_monitor
387
+ ```
388
+
389
+ To debug a specific job:
390
+
391
+ ```bash
392
+ ./go.d.plugin -d -m azure_monitor -j jobName
393
+ ```
394
+
395
+### Getting Logs
396
+
397
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
398
+
399
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
400
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
401
+
402
+#### System with systemd
403
+
404
+Use the following command to view logs generated since the last Netdata service restart:
405
+
406
+```bash
407
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
408
+```
409
+
410
+#### System without systemd
411
+
412
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
413
+
414
+```bash
415
+grep azure_monitor /var/log/netdata/collector.log
416
+```
417
+
418
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
419
+
420
+#### Docker Container
421
+
422
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
423
+
424
+```bash
425
+docker logs netdata 2>&1 | grep azure_monitor
426
+```
427
+
428
+### No metrics are collected
429
+
430
+Verify the following:
431
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
432
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
433
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
434
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
435
+
436
+
437
+### Missing metrics for some resource types
438
+
439
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
440
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
441
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
442
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
443
+
444
+
445
+### Metrics appear delayed
446
+
447
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
448
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
449
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
450
+
451
+
452
+### Authentication errors in sovereign clouds
453
+
454
+For Azure Government or Azure China clouds, set the `cloud` parameter:
455
+- Azure Government: `cloud: government`
456
+- Azure China (21Vianet): `cloud: china`
457
+
458
+Ensure the service principal is registered in the correct cloud tenant.
459
+
460
+
461
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_elastic_pool.md
new
+450
@@ -0,0 +1,450 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_elastic_pool.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure SQL Elastic Pool"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'sql', 'elastic', 'pool', 'database', 'mssql']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure SQL Elastic Pool
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor SQL Elastic Pool resource consumption including eDTU and CPU utilization, storage usage, active sessions and workers, IO rates, tempdb usage, and in-memory OLTP storage across all databases in the pool.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.sql_elastic_pool.cpu | average, maximum | percentage |
88
+| azure_monitor.sql_elastic_pool.instance_cpu | average | percentage |
89
+| azure_monitor.sql_elastic_pool.instance_memory | average | percentage |
90
+| azure_monitor.sql_elastic_pool.dtu_consumption | average | percentage |
91
+| azure_monitor.sql_elastic_pool.io_utilization | data_read, log_write | percentage |
92
+| azure_monitor.sql_elastic_pool.resource_utilization | workers, sessions | percentage |
93
+| azure_monitor.sql_elastic_pool.sessions_count | average | sessions |
94
+| azure_monitor.sql_elastic_pool.storage | used, allocated, limit | bytes |
95
+| azure_monitor.sql_elastic_pool.storage_utilization | used, allocated | percentage |
96
+| azure_monitor.sql_elastic_pool.xtp_storage | average | percentage |
97
+| azure_monitor.sql_elastic_pool.tempdb_size | data, log | KiB |
98
+| azure_monitor.sql_elastic_pool.tempdb_log_utilization | average | percentage |
99
+| azure_monitor.sql_elastic_pool.vcore_usage | used, limit | vCores |
100
+| azure_monitor.sql_elastic_pool.edtu_usage | used, limit | eDTU |
101
+| azure_monitor.sql_elastic_pool.serverless_utilization | cpu, memory | percentage |
102
+| azure_monitor.sql_elastic_pool.serverless_billing | total | vCore-seconds/s |
103
+
104
+
105
+
106
+## Alerts
107
+
108
+
109
+The following alerts are available:
110
+
111
+| Alert name | On metric | Description |
112
+|:------------|:----------|:------------|
113
+| [ am_sql_elastic_pool_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.cpu | SQL Elastic Pool CPU on ${label:resource_name} |
114
+| [ am_sql_elastic_pool_instance_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.instance_cpu | SQL Elastic Pool instance CPU on ${label:resource_name} |
115
+| [ am_sql_elastic_pool_instance_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.instance_memory | SQL Elastic Pool instance memory on ${label:resource_name} |
116
+| [ am_sql_elastic_pool_dtu_consumption ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.dtu_consumption | SQL Elastic Pool DTU consumption on ${label:resource_name} |
117
+| [ am_sql_elastic_pool_data_io ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.io_utilization | SQL Elastic Pool data I/O on ${label:resource_name} |
118
+| [ am_sql_elastic_pool_log_write ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.io_utilization | SQL Elastic Pool log write on ${label:resource_name} |
119
+| [ am_sql_elastic_pool_workers ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.resource_utilization | SQL Elastic Pool workers on ${label:resource_name} |
120
+| [ am_sql_elastic_pool_sessions ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.resource_utilization | SQL Elastic Pool sessions on ${label:resource_name} |
121
+| [ am_sql_elastic_pool_storage_used ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.storage_utilization | SQL Elastic Pool storage used on ${label:resource_name} |
122
+| [ am_sql_elastic_pool_storage_allocated ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.storage_utilization | SQL Elastic Pool storage allocated on ${label:resource_name} |
123
+| [ am_sql_elastic_pool_xtp_storage ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.xtp_storage | SQL Elastic Pool In-Memory OLTP storage on ${label:resource_name} |
124
+| [ am_sql_elastic_pool_tempdb_log ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.tempdb_log_utilization | SQL Elastic Pool tempdb log on ${label:resource_name} |
125
+| [ am_sql_elastic_pool_serverless_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.serverless_utilization | SQL Elastic Pool serverless CPU on ${label:resource_name} |
126
+| [ am_sql_elastic_pool_serverless_memory ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_elastic_pool.conf) | azure_monitor.sql_elastic_pool.serverless_utilization | SQL Elastic Pool serverless memory on ${label:resource_name} |
127
+
128
+
129
+## Setup
130
+
131
+
132
+You can configure the **azure_monitor** collector in two ways:
133
+
134
+| Method | Best for | How to |
135
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
136
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
137
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
138
+
139
+:::important
140
+
141
+UI configuration requires paid Netdata Cloud plan.
142
+
143
+:::
144
+
145
+
146
+### Prerequisites
147
+
148
+#### Create an Azure monitoring principal
149
+
150
+Create a service principal or use a managed identity with the following permissions:
151
+
152
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
153
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
154
+
155
+For service principal authentication:
156
+```bash
157
+# Create the service principal
158
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
159
+ --scopes /subscriptions/<subscription-id>
160
+
161
+# Note the appId (client_id), password (client_secret), and tenant
162
+```
163
+
164
+For managed identity (on Azure VMs, VMSS, or AKS):
165
+```bash
166
+# Assign Monitoring Reader role to the VM's managed identity
167
+az role assignment create --assignee <managed-identity-principal-id> \
168
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
169
+```
170
+
171
+
172
+
173
+### Configuration
174
+
175
+#### Options
176
+
177
+The following options can be defined globally: update_every, autodetection_retry.
178
+
179
+Profile files are loaded from:
180
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
181
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
182
+
183
+User profile files with the same filename override stock profiles.
184
+
185
+
186
+<details open><summary>Config options</summary>
187
+
188
+
189
+
190
+| Group | Option | Description | Default | Required |
191
+|:------|:-----|:------------|:--------|:---------:|
192
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
193
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
194
+| **Target** | subscription_id | Azure subscription ID. | | yes |
195
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
196
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
197
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
198
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
199
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
203
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
204
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
205
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
206
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
207
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
208
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
209
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
210
+
211
+
212
+</details>
213
+
214
+
215
+#### via UI
216
+
217
+Configure the **azure_monitor** collector from the Netdata web interface:
218
+
219
+1. Go to **Nodes**.
220
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
221
+3. The **Collectors → Jobs** view opens by default.
222
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
223
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
224
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
225
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
226
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
227
+
228
+
229
+#### via File
230
+
231
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
232
+
233
+The file format is YAML. Generally, the structure is:
234
+
235
+```yaml
236
+update_every: 1
237
+autodetection_retry: 0
238
+jobs:
239
+ - name: some_name1
240
+ - name: some_name2
241
+```
242
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
243
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
244
+
245
+```bash
246
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
247
+sudo ./edit-config go.d/azure_monitor.conf
248
+```
249
+
250
+##### Examples
251
+
252
+###### Service principal (auto-discover all resources)
253
+
254
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
255
+
256
+```yaml
257
+jobs:
258
+ - name: prod
259
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
260
+ auth:
261
+ mode: service_principal
262
+ mode_service_principal:
263
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
264
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
265
+ client_secret: "your-client-secret"
266
+
267
+```
268
+###### Managed identity (Azure VM/VMSS/AKS)
269
+
270
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
271
+
272
+<details open><summary>Config</summary>
273
+
274
+```yaml
275
+jobs:
276
+ - name: prod
277
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
278
+ auth:
279
+ mode: managed_identity
280
+
281
+```
282
+</details>
283
+
284
+###### Specific profiles only
285
+
286
+Monitor only specific Azure services instead of auto-discovering all resource types.
287
+
288
+<details open><summary>Config</summary>
289
+
290
+```yaml
291
+jobs:
292
+ - name: databases
293
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
294
+ profiles:
295
+ - sql_database
296
+ - postgres_flexible
297
+ - redis_cache
298
+ auth:
299
+ mode: service_principal
300
+ mode_service_principal:
301
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
302
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
303
+ client_secret: "your-client-secret"
304
+
305
+```
306
+</details>
307
+
308
+###### Filter by resource group
309
+
310
+Only monitor resources in specific resource groups.
311
+
312
+<details open><summary>Config</summary>
313
+
314
+```yaml
315
+jobs:
316
+ - name: prod-rg
317
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
318
+ resource_groups:
319
+ - production-rg
320
+ - staging-rg
321
+ auth:
322
+ mode: default
323
+
324
+```
325
+</details>
326
+
327
+###### Azure Government cloud
328
+
329
+Connect to Azure Government cloud environment.
330
+
331
+<details open><summary>Config</summary>
332
+
333
+```yaml
334
+jobs:
335
+ - name: gov
336
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
337
+ cloud: government
338
+ auth:
339
+ mode: service_principal
340
+ mode_service_principal:
341
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
342
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
343
+ client_secret: "your-client-secret"
344
+
345
+```
346
+</details>
347
+
348
+
349
+
350
+## Troubleshooting
351
+
352
+### Debug Mode
353
+
354
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
355
+
356
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
357
+should give you clues as to why the collector isn't working.
358
+
359
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
360
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
361
+
362
+ ```bash
363
+ cd /usr/libexec/netdata/plugins.d/
364
+ ```
365
+
366
+- Switch to the `netdata` user.
367
+
368
+ ```bash
369
+ sudo -u netdata -s
370
+ ```
371
+
372
+- Run the `go.d.plugin` to debug the collector:
373
+
374
+ ```bash
375
+ ./go.d.plugin -d -m azure_monitor
376
+ ```
377
+
378
+ To debug a specific job:
379
+
380
+ ```bash
381
+ ./go.d.plugin -d -m azure_monitor -j jobName
382
+ ```
383
+
384
+### Getting Logs
385
+
386
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
387
+
388
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
389
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
390
+
391
+#### System with systemd
392
+
393
+Use the following command to view logs generated since the last Netdata service restart:
394
+
395
+```bash
396
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
397
+```
398
+
399
+#### System without systemd
400
+
401
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
402
+
403
+```bash
404
+grep azure_monitor /var/log/netdata/collector.log
405
+```
406
+
407
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
408
+
409
+#### Docker Container
410
+
411
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
412
+
413
+```bash
414
+docker logs netdata 2>&1 | grep azure_monitor
415
+```
416
+
417
+### No metrics are collected
418
+
419
+Verify the following:
420
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
421
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
422
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
423
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
424
+
425
+
426
+### Missing metrics for some resource types
427
+
428
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
429
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
430
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
431
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
432
+
433
+
434
+### Metrics appear delayed
435
+
436
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
437
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
438
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
439
+
440
+
441
+### Authentication errors in sovereign clouds
442
+
443
+For Azure Government or Azure China clouds, set the `cloud` parameter:
444
+- Azure Government: `cloud: government`
445
+- Azure China (21Vianet): `cloud: china`
446
+
447
+Ensure the service principal is registered in the correct cloud tenant.
448
+
449
+
450
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_managed_instance.md
new
+428
@@ -0,0 +1,428 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_managed_instance.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure SQL Managed Instance"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'sql', 'managed', 'instance', 'database', 'mssql']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure SQL Managed Instance
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor SQL Managed Instance performance including virtual core CPU utilization, storage consumption, IO throughput, and average request wait times.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.sql_managed_instance.cpu | average, maximum | percentage |
88
+| azure_monitor.sql_managed_instance.io_throughput | read, written | bytes/s |
89
+| azure_monitor.sql_managed_instance.io_requests | average | requests/s |
90
+| azure_monitor.sql_managed_instance.storage | reserved, used | MiB |
91
+| azure_monitor.sql_managed_instance.virtual_core_count | average | cores |
92
+
93
+
94
+
95
+## Alerts
96
+
97
+
98
+The following alerts are available:
99
+
100
+| Alert name | On metric | Description |
101
+|:------------|:----------|:------------|
102
+| [ am_sql_managed_instance_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_managed_instance.conf) | azure_monitor.sql_managed_instance.cpu | SQL MI CPU utilization on ${label:resource_name} |
103
+| [ am_sql_managed_instance_storage_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_managed_instance.conf) | azure_monitor.sql_managed_instance.storage | SQL MI storage utilization on ${label:resource_name} |
104
+| [ am_sql_managed_instance_io_requests ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_sql_managed_instance.conf) | azure_monitor.sql_managed_instance.io_requests | SQL MI I/O requests on ${label:resource_name} |
105
+
106
+
107
+## Setup
108
+
109
+
110
+You can configure the **azure_monitor** collector in two ways:
111
+
112
+| Method | Best for | How to |
113
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
114
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
115
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
116
+
117
+:::important
118
+
119
+UI configuration requires paid Netdata Cloud plan.
120
+
121
+:::
122
+
123
+
124
+### Prerequisites
125
+
126
+#### Create an Azure monitoring principal
127
+
128
+Create a service principal or use a managed identity with the following permissions:
129
+
130
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
131
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
132
+
133
+For service principal authentication:
134
+```bash
135
+# Create the service principal
136
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
137
+ --scopes /subscriptions/<subscription-id>
138
+
139
+# Note the appId (client_id), password (client_secret), and tenant
140
+```
141
+
142
+For managed identity (on Azure VMs, VMSS, or AKS):
143
+```bash
144
+# Assign Monitoring Reader role to the VM's managed identity
145
+az role assignment create --assignee <managed-identity-principal-id> \
146
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+```
148
+
149
+
150
+
151
+### Configuration
152
+
153
+#### Options
154
+
155
+The following options can be defined globally: update_every, autodetection_retry.
156
+
157
+Profile files are loaded from:
158
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
159
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+
161
+User profile files with the same filename override stock profiles.
162
+
163
+
164
+<details open><summary>Config options</summary>
165
+
166
+
167
+
168
+| Group | Option | Description | Default | Required |
169
+|:------|:-----|:------------|:--------|:---------:|
170
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
171
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
172
+| **Target** | subscription_id | Azure subscription ID. | | yes |
173
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
174
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
175
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
176
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
177
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
178
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
179
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
180
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
181
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
182
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
184
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
185
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
186
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
187
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
188
+
189
+
190
+</details>
191
+
192
+
193
+#### via UI
194
+
195
+Configure the **azure_monitor** collector from the Netdata web interface:
196
+
197
+1. Go to **Nodes**.
198
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
199
+3. The **Collectors → Jobs** view opens by default.
200
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
201
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
202
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
203
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
204
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
205
+
206
+
207
+#### via File
208
+
209
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
210
+
211
+The file format is YAML. Generally, the structure is:
212
+
213
+```yaml
214
+update_every: 1
215
+autodetection_retry: 0
216
+jobs:
217
+ - name: some_name1
218
+ - name: some_name2
219
+```
220
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
221
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
222
+
223
+```bash
224
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
225
+sudo ./edit-config go.d/azure_monitor.conf
226
+```
227
+
228
+##### Examples
229
+
230
+###### Service principal (auto-discover all resources)
231
+
232
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
233
+
234
+```yaml
235
+jobs:
236
+ - name: prod
237
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
238
+ auth:
239
+ mode: service_principal
240
+ mode_service_principal:
241
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
242
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
243
+ client_secret: "your-client-secret"
244
+
245
+```
246
+###### Managed identity (Azure VM/VMSS/AKS)
247
+
248
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
249
+
250
+<details open><summary>Config</summary>
251
+
252
+```yaml
253
+jobs:
254
+ - name: prod
255
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
256
+ auth:
257
+ mode: managed_identity
258
+
259
+```
260
+</details>
261
+
262
+###### Specific profiles only
263
+
264
+Monitor only specific Azure services instead of auto-discovering all resource types.
265
+
266
+<details open><summary>Config</summary>
267
+
268
+```yaml
269
+jobs:
270
+ - name: databases
271
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
272
+ profiles:
273
+ - sql_database
274
+ - postgres_flexible
275
+ - redis_cache
276
+ auth:
277
+ mode: service_principal
278
+ mode_service_principal:
279
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
280
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
281
+ client_secret: "your-client-secret"
282
+
283
+```
284
+</details>
285
+
286
+###### Filter by resource group
287
+
288
+Only monitor resources in specific resource groups.
289
+
290
+<details open><summary>Config</summary>
291
+
292
+```yaml
293
+jobs:
294
+ - name: prod-rg
295
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
296
+ resource_groups:
297
+ - production-rg
298
+ - staging-rg
299
+ auth:
300
+ mode: default
301
+
302
+```
303
+</details>
304
+
305
+###### Azure Government cloud
306
+
307
+Connect to Azure Government cloud environment.
308
+
309
+<details open><summary>Config</summary>
310
+
311
+```yaml
312
+jobs:
313
+ - name: gov
314
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
315
+ cloud: government
316
+ auth:
317
+ mode: service_principal
318
+ mode_service_principal:
319
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ client_secret: "your-client-secret"
322
+
323
+```
324
+</details>
325
+
326
+
327
+
328
+## Troubleshooting
329
+
330
+### Debug Mode
331
+
332
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
333
+
334
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
335
+should give you clues as to why the collector isn't working.
336
+
337
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
338
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
339
+
340
+ ```bash
341
+ cd /usr/libexec/netdata/plugins.d/
342
+ ```
343
+
344
+- Switch to the `netdata` user.
345
+
346
+ ```bash
347
+ sudo -u netdata -s
348
+ ```
349
+
350
+- Run the `go.d.plugin` to debug the collector:
351
+
352
+ ```bash
353
+ ./go.d.plugin -d -m azure_monitor
354
+ ```
355
+
356
+ To debug a specific job:
357
+
358
+ ```bash
359
+ ./go.d.plugin -d -m azure_monitor -j jobName
360
+ ```
361
+
362
+### Getting Logs
363
+
364
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
365
+
366
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
367
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
368
+
369
+#### System with systemd
370
+
371
+Use the following command to view logs generated since the last Netdata service restart:
372
+
373
+```bash
374
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
375
+```
376
+
377
+#### System without systemd
378
+
379
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
380
+
381
+```bash
382
+grep azure_monitor /var/log/netdata/collector.log
383
+```
384
+
385
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
386
+
387
+#### Docker Container
388
+
389
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
390
+
391
+```bash
392
+docker logs netdata 2>&1 | grep azure_monitor
393
+```
394
+
395
+### No metrics are collected
396
+
397
+Verify the following:
398
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
399
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
400
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
401
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
402
+
403
+
404
+### Missing metrics for some resource types
405
+
406
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
407
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
408
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
409
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
410
+
411
+
412
+### Metrics appear delayed
413
+
414
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
415
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
416
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
417
+
418
+
419
+### Authentication errors in sovereign clouds
420
+
421
+For Azure Government or Azure China clouds, set the `cloud` parameter:
422
+- Azure Government: `cloud: government`
423
+- Azure China (21Vianet): `cloud: china`
424
+
425
+Ensure the service principal is registered in the correct cloud tenant.
426
+
427
+
428
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_storage_account.md
new
+431
@@ -0,0 +1,431 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_storage_account.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Storage Account"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'storage', 'blob', 'account', 'files', 'queue', 'table']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Storage Account
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Storage Account operations including transaction counts, availability percentages, success and end-to-end latency, ingress and egress throughput, and used capacity.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.storage_accounts.availability | average | percentage |
88
+| azure_monitor.storage_accounts.transactions | total | transactions/s |
89
+| azure_monitor.storage_accounts.throughput | ingress, egress | bytes/s |
90
+| azure_monitor.storage_accounts.e2e_latency | average, maximum | milliseconds |
91
+| azure_monitor.storage_accounts.server_latency | average, maximum | milliseconds |
92
+| azure_monitor.storage_accounts.used_capacity | average | bytes |
93
+
94
+
95
+
96
+## Alerts
97
+
98
+
99
+The following alerts are available:
100
+
101
+| Alert name | On metric | Description |
102
+|:------------|:----------|:------------|
103
+| [ am_storage_accounts_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_storage_accounts.conf) | azure_monitor.storage_accounts.availability | Storage availability on ${label:resource_name} |
104
+| [ am_storage_accounts_e2e_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_storage_accounts.conf) | azure_monitor.storage_accounts.e2e_latency | Storage E2E latency on ${label:resource_name} |
105
+| [ am_storage_accounts_e2e_latency_peak ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_storage_accounts.conf) | azure_monitor.storage_accounts.e2e_latency | Storage E2E latency peak on ${label:resource_name} |
106
+| [ am_storage_accounts_server_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_storage_accounts.conf) | azure_monitor.storage_accounts.server_latency | Storage server latency on ${label:resource_name} |
107
+| [ am_storage_accounts_server_latency_peak ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_storage_accounts.conf) | azure_monitor.storage_accounts.server_latency | Storage server latency peak on ${label:resource_name} |
108
+
109
+
110
+## Setup
111
+
112
+
113
+You can configure the **azure_monitor** collector in two ways:
114
+
115
+| Method | Best for | How to |
116
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
117
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
118
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
119
+
120
+:::important
121
+
122
+UI configuration requires paid Netdata Cloud plan.
123
+
124
+:::
125
+
126
+
127
+### Prerequisites
128
+
129
+#### Create an Azure monitoring principal
130
+
131
+Create a service principal or use a managed identity with the following permissions:
132
+
133
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
134
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
135
+
136
+For service principal authentication:
137
+```bash
138
+# Create the service principal
139
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
140
+ --scopes /subscriptions/<subscription-id>
141
+
142
+# Note the appId (client_id), password (client_secret), and tenant
143
+```
144
+
145
+For managed identity (on Azure VMs, VMSS, or AKS):
146
+```bash
147
+# Assign Monitoring Reader role to the VM's managed identity
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
150
+```
151
+
152
+
153
+
154
+### Configuration
155
+
156
+#### Options
157
+
158
+The following options can be defined globally: update_every, autodetection_retry.
159
+
160
+Profile files are loaded from:
161
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
162
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
163
+
164
+User profile files with the same filename override stock profiles.
165
+
166
+
167
+<details open><summary>Config options</summary>
168
+
169
+
170
+
171
+| Group | Option | Description | Default | Required |
172
+|:------|:-----|:------------|:--------|:---------:|
173
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
174
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
175
+| **Target** | subscription_id | Azure subscription ID. | | yes |
176
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
177
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
178
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
180
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
181
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
182
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
183
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
184
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
185
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
186
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
191
+
192
+
193
+</details>
194
+
195
+
196
+#### via UI
197
+
198
+Configure the **azure_monitor** collector from the Netdata web interface:
199
+
200
+1. Go to **Nodes**.
201
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
202
+3. The **Collectors → Jobs** view opens by default.
203
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
204
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
205
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
206
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
207
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
208
+
209
+
210
+#### via File
211
+
212
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
213
+
214
+The file format is YAML. Generally, the structure is:
215
+
216
+```yaml
217
+update_every: 1
218
+autodetection_retry: 0
219
+jobs:
220
+ - name: some_name1
221
+ - name: some_name2
222
+```
223
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
224
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
225
+
226
+```bash
227
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
228
+sudo ./edit-config go.d/azure_monitor.conf
229
+```
230
+
231
+##### Examples
232
+
233
+###### Service principal (auto-discover all resources)
234
+
235
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
236
+
237
+```yaml
238
+jobs:
239
+ - name: prod
240
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
241
+ auth:
242
+ mode: service_principal
243
+ mode_service_principal:
244
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
245
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
246
+ client_secret: "your-client-secret"
247
+
248
+```
249
+###### Managed identity (Azure VM/VMSS/AKS)
250
+
251
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
252
+
253
+<details open><summary>Config</summary>
254
+
255
+```yaml
256
+jobs:
257
+ - name: prod
258
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
259
+ auth:
260
+ mode: managed_identity
261
+
262
+```
263
+</details>
264
+
265
+###### Specific profiles only
266
+
267
+Monitor only specific Azure services instead of auto-discovering all resource types.
268
+
269
+<details open><summary>Config</summary>
270
+
271
+```yaml
272
+jobs:
273
+ - name: databases
274
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
275
+ profiles:
276
+ - sql_database
277
+ - postgres_flexible
278
+ - redis_cache
279
+ auth:
280
+ mode: service_principal
281
+ mode_service_principal:
282
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
284
+ client_secret: "your-client-secret"
285
+
286
+```
287
+</details>
288
+
289
+###### Filter by resource group
290
+
291
+Only monitor resources in specific resource groups.
292
+
293
+<details open><summary>Config</summary>
294
+
295
+```yaml
296
+jobs:
297
+ - name: prod-rg
298
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ resource_groups:
300
+ - production-rg
301
+ - staging-rg
302
+ auth:
303
+ mode: default
304
+
305
+```
306
+</details>
307
+
308
+###### Azure Government cloud
309
+
310
+Connect to Azure Government cloud environment.
311
+
312
+<details open><summary>Config</summary>
313
+
314
+```yaml
315
+jobs:
316
+ - name: gov
317
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
318
+ cloud: government
319
+ auth:
320
+ mode: service_principal
321
+ mode_service_principal:
322
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ client_secret: "your-client-secret"
325
+
326
+```
327
+</details>
328
+
329
+
330
+
331
+## Troubleshooting
332
+
333
+### Debug Mode
334
+
335
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
336
+
337
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
338
+should give you clues as to why the collector isn't working.
339
+
340
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
341
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
342
+
343
+ ```bash
344
+ cd /usr/libexec/netdata/plugins.d/
345
+ ```
346
+
347
+- Switch to the `netdata` user.
348
+
349
+ ```bash
350
+ sudo -u netdata -s
351
+ ```
352
+
353
+- Run the `go.d.plugin` to debug the collector:
354
+
355
+ ```bash
356
+ ./go.d.plugin -d -m azure_monitor
357
+ ```
358
+
359
+ To debug a specific job:
360
+
361
+ ```bash
362
+ ./go.d.plugin -d -m azure_monitor -j jobName
363
+ ```
364
+
365
+### Getting Logs
366
+
367
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
368
+
369
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
370
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
371
+
372
+#### System with systemd
373
+
374
+Use the following command to view logs generated since the last Netdata service restart:
375
+
376
+```bash
377
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
378
+```
379
+
380
+#### System without systemd
381
+
382
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
383
+
384
+```bash
385
+grep azure_monitor /var/log/netdata/collector.log
386
+```
387
+
388
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
389
+
390
+#### Docker Container
391
+
392
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
393
+
394
+```bash
395
+docker logs netdata 2>&1 | grep azure_monitor
396
+```
397
+
398
+### No metrics are collected
399
+
400
+Verify the following:
401
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
402
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
403
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
404
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
405
+
406
+
407
+### Missing metrics for some resource types
408
+
409
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
410
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
411
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
412
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
413
+
414
+
415
+### Metrics appear delayed
416
+
417
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
418
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
419
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
420
+
421
+
422
+### Authentication errors in sovereign clouds
423
+
424
+For Azure Government or Azure China clouds, set the `cloud` parameter:
425
+- Azure Government: `cloud: government`
426
+- Azure China (21Vianet): `cloud: china`
427
+
428
+Ensure the service principal is registered in the correct cloud tenant.
429
+
430
+
431
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_stream_analytics_job.md
new
+440
@@ -0,0 +1,440 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_stream_analytics_job.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Stream Analytics Job"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'stream', 'analytics', 'streaming', 'realtime']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Stream Analytics Job
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Stream Analytics jobs including input and output event counts, streaming unit utilization, watermark delay, backlogged input events, runtime and data conversion errors, out-of-order events, and late input events.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.stream_analytics.event_flow | in, out | events/s |
88
+| azure_monitor.stream_analytics.input_data_throughput | received | bytes/s |
89
+| azure_monitor.stream_analytics.input_sources_received | received | sources/s |
90
+| azure_monitor.stream_analytics.event_timing | late, early, out_of_order | events/s |
91
+| azure_monitor.stream_analytics.backlogged_events | backlogged | events |
92
+| azure_monitor.stream_analytics.watermark_delay | delay | seconds |
93
+| azure_monitor.stream_analytics.errors | runtime, data_conversion, deserialization | errors/s |
94
+| azure_monitor.stream_analytics.function_requests | total, failed | requests/s |
95
+| azure_monitor.stream_analytics.function_events | input | events/s |
96
+| azure_monitor.stream_analytics.resource_utilization | cpu, su_memory | percentage |
97
+
98
+
99
+
100
+## Alerts
101
+
102
+
103
+The following alerts are available:
104
+
105
+| Alert name | On metric | Description |
106
+|:------------|:----------|:------------|
107
+| [ am_stream_analytics_su_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.resource_utilization | Stream Analytics SU utilization on ${label:resource_name} |
108
+| [ am_stream_analytics_cpu_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.resource_utilization | Stream Analytics CPU on ${label:resource_name} |
109
+| [ am_stream_analytics_watermark_delay ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.watermark_delay | Stream Analytics watermark delay on ${label:resource_name} |
110
+| [ am_stream_analytics_runtime_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.errors | Stream Analytics runtime errors on ${label:resource_name} |
111
+| [ am_stream_analytics_data_conversion_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.errors | Stream Analytics conversion errors on ${label:resource_name} |
112
+| [ am_stream_analytics_deserialization_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.errors | Stream Analytics deserialization errors on ${label:resource_name} |
113
+| [ am_stream_analytics_out_of_order_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.event_timing | Stream Analytics out-of-order events on ${label:resource_name} |
114
+| [ am_stream_analytics_late_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.event_timing | Stream Analytics late events on ${label:resource_name} |
115
+| [ am_stream_analytics_backlogged_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.backlogged_events | Stream Analytics backlog on ${label:resource_name} |
116
+| [ am_stream_analytics_function_failures ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_stream_analytics.conf) | azure_monitor.stream_analytics.function_requests | Stream Analytics ML function failures on ${label:resource_name} |
117
+
118
+
119
+## Setup
120
+
121
+
122
+You can configure the **azure_monitor** collector in two ways:
123
+
124
+| Method | Best for | How to |
125
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
126
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
127
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
128
+
129
+:::important
130
+
131
+UI configuration requires paid Netdata Cloud plan.
132
+
133
+:::
134
+
135
+
136
+### Prerequisites
137
+
138
+#### Create an Azure monitoring principal
139
+
140
+Create a service principal or use a managed identity with the following permissions:
141
+
142
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
143
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
144
+
145
+For service principal authentication:
146
+```bash
147
+# Create the service principal
148
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
149
+ --scopes /subscriptions/<subscription-id>
150
+
151
+# Note the appId (client_id), password (client_secret), and tenant
152
+```
153
+
154
+For managed identity (on Azure VMs, VMSS, or AKS):
155
+```bash
156
+# Assign Monitoring Reader role to the VM's managed identity
157
+az role assignment create --assignee <managed-identity-principal-id> \
158
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
159
+```
160
+
161
+
162
+
163
+### Configuration
164
+
165
+#### Options
166
+
167
+The following options can be defined globally: update_every, autodetection_retry.
168
+
169
+Profile files are loaded from:
170
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
171
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
172
+
173
+User profile files with the same filename override stock profiles.
174
+
175
+
176
+<details open><summary>Config options</summary>
177
+
178
+
179
+
180
+| Group | Option | Description | Default | Required |
181
+|:------|:-----|:------------|:--------|:---------:|
182
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
183
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
184
+| **Target** | subscription_id | Azure subscription ID. | | yes |
185
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
186
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
187
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
188
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
189
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
190
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
191
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
192
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
193
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
194
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
195
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
196
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
197
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
198
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
199
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
200
+
201
+
202
+</details>
203
+
204
+
205
+#### via UI
206
+
207
+Configure the **azure_monitor** collector from the Netdata web interface:
208
+
209
+1. Go to **Nodes**.
210
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
211
+3. The **Collectors → Jobs** view opens by default.
212
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
213
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
214
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
215
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
216
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
217
+
218
+
219
+#### via File
220
+
221
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
222
+
223
+The file format is YAML. Generally, the structure is:
224
+
225
+```yaml
226
+update_every: 1
227
+autodetection_retry: 0
228
+jobs:
229
+ - name: some_name1
230
+ - name: some_name2
231
+```
232
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
233
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
234
+
235
+```bash
236
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
237
+sudo ./edit-config go.d/azure_monitor.conf
238
+```
239
+
240
+##### Examples
241
+
242
+###### Service principal (auto-discover all resources)
243
+
244
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
245
+
246
+```yaml
247
+jobs:
248
+ - name: prod
249
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
250
+ auth:
251
+ mode: service_principal
252
+ mode_service_principal:
253
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
254
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
255
+ client_secret: "your-client-secret"
256
+
257
+```
258
+###### Managed identity (Azure VM/VMSS/AKS)
259
+
260
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
261
+
262
+<details open><summary>Config</summary>
263
+
264
+```yaml
265
+jobs:
266
+ - name: prod
267
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
268
+ auth:
269
+ mode: managed_identity
270
+
271
+```
272
+</details>
273
+
274
+###### Specific profiles only
275
+
276
+Monitor only specific Azure services instead of auto-discovering all resource types.
277
+
278
+<details open><summary>Config</summary>
279
+
280
+```yaml
281
+jobs:
282
+ - name: databases
283
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
284
+ profiles:
285
+ - sql_database
286
+ - postgres_flexible
287
+ - redis_cache
288
+ auth:
289
+ mode: service_principal
290
+ mode_service_principal:
291
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
292
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
293
+ client_secret: "your-client-secret"
294
+
295
+```
296
+</details>
297
+
298
+###### Filter by resource group
299
+
300
+Only monitor resources in specific resource groups.
301
+
302
+<details open><summary>Config</summary>
303
+
304
+```yaml
305
+jobs:
306
+ - name: prod-rg
307
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
308
+ resource_groups:
309
+ - production-rg
310
+ - staging-rg
311
+ auth:
312
+ mode: default
313
+
314
+```
315
+</details>
316
+
317
+###### Azure Government cloud
318
+
319
+Connect to Azure Government cloud environment.
320
+
321
+<details open><summary>Config</summary>
322
+
323
+```yaml
324
+jobs:
325
+ - name: gov
326
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
327
+ cloud: government
328
+ auth:
329
+ mode: service_principal
330
+ mode_service_principal:
331
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
332
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
333
+ client_secret: "your-client-secret"
334
+
335
+```
336
+</details>
337
+
338
+
339
+
340
+## Troubleshooting
341
+
342
+### Debug Mode
343
+
344
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
345
+
346
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
347
+should give you clues as to why the collector isn't working.
348
+
349
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
350
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
351
+
352
+ ```bash
353
+ cd /usr/libexec/netdata/plugins.d/
354
+ ```
355
+
356
+- Switch to the `netdata` user.
357
+
358
+ ```bash
359
+ sudo -u netdata -s
360
+ ```
361
+
362
+- Run the `go.d.plugin` to debug the collector:
363
+
364
+ ```bash
365
+ ./go.d.plugin -d -m azure_monitor
366
+ ```
367
+
368
+ To debug a specific job:
369
+
370
+ ```bash
371
+ ./go.d.plugin -d -m azure_monitor -j jobName
372
+ ```
373
+
374
+### Getting Logs
375
+
376
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
377
+
378
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
379
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
380
+
381
+#### System with systemd
382
+
383
+Use the following command to view logs generated since the last Netdata service restart:
384
+
385
+```bash
386
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
387
+```
388
+
389
+#### System without systemd
390
+
391
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
392
+
393
+```bash
394
+grep azure_monitor /var/log/netdata/collector.log
395
+```
396
+
397
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
398
+
399
+#### Docker Container
400
+
401
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
402
+
403
+```bash
404
+docker logs netdata 2>&1 | grep azure_monitor
405
+```
406
+
407
+### No metrics are collected
408
+
409
+Verify the following:
410
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
411
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
412
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
413
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
414
+
415
+
416
+### Missing metrics for some resource types
417
+
418
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
419
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
420
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
421
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
422
+
423
+
424
+### Metrics appear delayed
425
+
426
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
427
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
428
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
429
+
430
+
431
+### Authentication errors in sovereign clouds
432
+
433
+For Azure Government or Azure China clouds, set the `cloud` parameter:
434
+- Azure Government: `cloud: government`
435
+- Azure China (21Vianet): `cloud: china`
436
+
437
+Ensure the service principal is registered in the correct cloud tenant.
438
+
439
+
440
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_synapse_analytics_workspace.md
new
+446
@@ -0,0 +1,446 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_synapse_analytics_workspace.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Synapse Analytics Workspace"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'synapse', 'analytics', 'data', 'warehouse', 'sql']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Synapse Analytics Workspace
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Synapse Analytics workspaces including pipeline and activity run metrics, SQL request counts and data processing volumes, data flow activity execution, integration runtime CPU and memory utilization, and link table event processing.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.synapse.builtin_sql_pool_data_processed | processed | bytes/s |
88
+| azure_monitor.synapse.builtin_sql_pool_login_attempts | login_attempts | attempts/s |
89
+| azure_monitor.synapse.builtin_sql_pool_requests | requests | requests/s |
90
+| azure_monitor.synapse.activity_runs | ended | runs/s |
91
+| azure_monitor.synapse.pipeline_runs | ended | runs/s |
92
+| azure_monitor.synapse.trigger_runs | ended | runs/s |
93
+| azure_monitor.synapse.link_connection_events | events | events/s |
94
+| azure_monitor.synapse.link_table_events | events | events/s |
95
+| azure_monitor.synapse.link_processed_rows | changed_rows | rows/s |
96
+| azure_monitor.synapse.link_data_volume | processed | bytes/s |
97
+| azure_monitor.synapse.link_processing_latency | average | seconds |
98
+| azure_monitor.synapse.streaming_event_flow | in, out | events/s |
99
+| azure_monitor.synapse.streaming_input_throughput | received | bytes/s |
100
+| azure_monitor.synapse.streaming_input_sources | received | sources/s |
101
+| azure_monitor.synapse.streaming_event_timing | late, early, out_of_order, backlogged | events/s |
102
+| azure_monitor.synapse.streaming_watermark_delay | delay | seconds |
103
+| azure_monitor.synapse.streaming_errors | runtime, data_conversion, deserialization | errors/s |
104
+| azure_monitor.synapse.streaming_resource_utilization | utilization | percentage |
105
+
106
+
107
+
108
+## Alerts
109
+
110
+
111
+The following alerts are available:
112
+
113
+| Alert name | On metric | Description |
114
+|:------------|:----------|:------------|
115
+| [ am_synapse_streaming_resource_utilization ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_resource_utilization | Synapse streaming SU utilization on ${label:resource_name} |
116
+| [ am_synapse_streaming_runtime_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_errors | Synapse streaming runtime errors on ${label:resource_name} |
117
+| [ am_synapse_streaming_data_errors ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_errors | Synapse streaming data errors on ${label:resource_name} |
118
+| [ am_synapse_streaming_watermark_delay ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_watermark_delay | Synapse streaming watermark delay on ${label:resource_name} |
119
+| [ am_synapse_streaming_late_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_event_timing | Synapse streaming late events on ${label:resource_name} |
120
+| [ am_synapse_streaming_out_of_order_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_event_timing | Synapse streaming out-of-order events on ${label:resource_name} |
121
+| [ am_synapse_streaming_backlogged_events ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.streaming_event_timing | Synapse streaming backlogged events on ${label:resource_name} |
122
+| [ am_synapse_link_processing_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_synapse.conf) | azure_monitor.synapse.link_processing_latency | Synapse Link processing latency on ${label:resource_name} |
123
+
124
+
125
+## Setup
126
+
127
+
128
+You can configure the **azure_monitor** collector in two ways:
129
+
130
+| Method | Best for | How to |
131
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
132
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
133
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
134
+
135
+:::important
136
+
137
+UI configuration requires paid Netdata Cloud plan.
138
+
139
+:::
140
+
141
+
142
+### Prerequisites
143
+
144
+#### Create an Azure monitoring principal
145
+
146
+Create a service principal or use a managed identity with the following permissions:
147
+
148
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
149
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
150
+
151
+For service principal authentication:
152
+```bash
153
+# Create the service principal
154
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
155
+ --scopes /subscriptions/<subscription-id>
156
+
157
+# Note the appId (client_id), password (client_secret), and tenant
158
+```
159
+
160
+For managed identity (on Azure VMs, VMSS, or AKS):
161
+```bash
162
+# Assign Monitoring Reader role to the VM's managed identity
163
+az role assignment create --assignee <managed-identity-principal-id> \
164
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
165
+```
166
+
167
+
168
+
169
+### Configuration
170
+
171
+#### Options
172
+
173
+The following options can be defined globally: update_every, autodetection_retry.
174
+
175
+Profile files are loaded from:
176
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
177
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
178
+
179
+User profile files with the same filename override stock profiles.
180
+
181
+
182
+<details open><summary>Config options</summary>
183
+
184
+
185
+
186
+| Group | Option | Description | Default | Required |
187
+|:------|:-----|:------------|:--------|:---------:|
188
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
189
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
190
+| **Target** | subscription_id | Azure subscription ID. | | yes |
191
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
192
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
193
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
194
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
195
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
196
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
197
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
198
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
199
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
200
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
201
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
202
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
203
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
204
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
205
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
206
+
207
+
208
+</details>
209
+
210
+
211
+#### via UI
212
+
213
+Configure the **azure_monitor** collector from the Netdata web interface:
214
+
215
+1. Go to **Nodes**.
216
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
217
+3. The **Collectors → Jobs** view opens by default.
218
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
219
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
220
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
221
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
222
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
223
+
224
+
225
+#### via File
226
+
227
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
228
+
229
+The file format is YAML. Generally, the structure is:
230
+
231
+```yaml
232
+update_every: 1
233
+autodetection_retry: 0
234
+jobs:
235
+ - name: some_name1
236
+ - name: some_name2
237
+```
238
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
239
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
240
+
241
+```bash
242
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
243
+sudo ./edit-config go.d/azure_monitor.conf
244
+```
245
+
246
+##### Examples
247
+
248
+###### Service principal (auto-discover all resources)
249
+
250
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
251
+
252
+```yaml
253
+jobs:
254
+ - name: prod
255
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
256
+ auth:
257
+ mode: service_principal
258
+ mode_service_principal:
259
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
260
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
261
+ client_secret: "your-client-secret"
262
+
263
+```
264
+###### Managed identity (Azure VM/VMSS/AKS)
265
+
266
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
267
+
268
+<details open><summary>Config</summary>
269
+
270
+```yaml
271
+jobs:
272
+ - name: prod
273
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
274
+ auth:
275
+ mode: managed_identity
276
+
277
+```
278
+</details>
279
+
280
+###### Specific profiles only
281
+
282
+Monitor only specific Azure services instead of auto-discovering all resource types.
283
+
284
+<details open><summary>Config</summary>
285
+
286
+```yaml
287
+jobs:
288
+ - name: databases
289
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
290
+ profiles:
291
+ - sql_database
292
+ - postgres_flexible
293
+ - redis_cache
294
+ auth:
295
+ mode: service_principal
296
+ mode_service_principal:
297
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
298
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
299
+ client_secret: "your-client-secret"
300
+
301
+```
302
+</details>
303
+
304
+###### Filter by resource group
305
+
306
+Only monitor resources in specific resource groups.
307
+
308
+<details open><summary>Config</summary>
309
+
310
+```yaml
311
+jobs:
312
+ - name: prod-rg
313
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
314
+ resource_groups:
315
+ - production-rg
316
+ - staging-rg
317
+ auth:
318
+ mode: default
319
+
320
+```
321
+</details>
322
+
323
+###### Azure Government cloud
324
+
325
+Connect to Azure Government cloud environment.
326
+
327
+<details open><summary>Config</summary>
328
+
329
+```yaml
330
+jobs:
331
+ - name: gov
332
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
333
+ cloud: government
334
+ auth:
335
+ mode: service_principal
336
+ mode_service_principal:
337
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
338
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
339
+ client_secret: "your-client-secret"
340
+
341
+```
342
+</details>
343
+
344
+
345
+
346
+## Troubleshooting
347
+
348
+### Debug Mode
349
+
350
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
351
+
352
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
353
+should give you clues as to why the collector isn't working.
354
+
355
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
356
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
357
+
358
+ ```bash
359
+ cd /usr/libexec/netdata/plugins.d/
360
+ ```
361
+
362
+- Switch to the `netdata` user.
363
+
364
+ ```bash
365
+ sudo -u netdata -s
366
+ ```
367
+
368
+- Run the `go.d.plugin` to debug the collector:
369
+
370
+ ```bash
371
+ ./go.d.plugin -d -m azure_monitor
372
+ ```
373
+
374
+ To debug a specific job:
375
+
376
+ ```bash
377
+ ./go.d.plugin -d -m azure_monitor -j jobName
378
+ ```
379
+
380
+### Getting Logs
381
+
382
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
383
+
384
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
385
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
386
+
387
+#### System with systemd
388
+
389
+Use the following command to view logs generated since the last Netdata service restart:
390
+
391
+```bash
392
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
393
+```
394
+
395
+#### System without systemd
396
+
397
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
398
+
399
+```bash
400
+grep azure_monitor /var/log/netdata/collector.log
401
+```
402
+
403
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
404
+
405
+#### Docker Container
406
+
407
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
408
+
409
+```bash
410
+docker logs netdata 2>&1 | grep azure_monitor
411
+```
412
+
413
+### No metrics are collected
414
+
415
+Verify the following:
416
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
417
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
418
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
419
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
420
+
421
+
422
+### Missing metrics for some resource types
423
+
424
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
425
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
426
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
427
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
428
+
429
+
430
+### Metrics appear delayed
431
+
432
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
433
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
434
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
435
+
436
+
437
+### Authentication errors in sovereign clouds
438
+
439
+For Azure Government or Azure China clouds, set the `cloud` parameter:
440
+- Azure Government: `cloud: government`
441
+- Azure China (21Vianet): `cloud: china`
442
+
443
+Ensure the service principal is registered in the correct cloud tenant.
444
+
445
+
446
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine.md
new
+480
@@ -0,0 +1,480 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Virtual Machine"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'virtual', 'machine', 'vm', 'compute', 'iaas']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Virtual Machine
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Azure Virtual Machines including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.virtual_machines.cpu | average | percentage |
88
+| azure_monitor.virtual_machines.cpu_credits | consumed, remaining | credits |
89
+| azure_monitor.virtual_machines.memory | available | bytes |
90
+| azure_monitor.virtual_machines.memory_percentage | available | percentage |
91
+| azure_monitor.virtual_machines.availability | average | state |
92
+| azure_monitor.virtual_machines.network_traffic | in, out | bytes/s |
93
+| azure_monitor.virtual_machines.network_flows | in, out | flows |
94
+| azure_monitor.virtual_machines.network_flow_creation_rate | in, out | flows/s |
95
+| azure_monitor.virtual_machines.disk_throughput | read, write | bytes/s |
96
+| azure_monitor.virtual_machines.disk_iops | read, write | operations/s |
97
+| azure_monitor.virtual_machines.os_disk_throughput | read, write | bytes/s |
98
+| azure_monitor.virtual_machines.os_disk_iops | read, write | operations/s |
99
+| azure_monitor.virtual_machines.os_disk_latency | average | milliseconds |
100
+| azure_monitor.virtual_machines.os_disk_queue_depth | average | operations |
101
+| azure_monitor.virtual_machines.os_disk_throttling | bandwidth, iops | percentage |
102
+| azure_monitor.virtual_machines.os_disk_burst_capacity | max_burst, target | bytes/s |
103
+| azure_monitor.virtual_machines.os_disk_burst_iops_capacity | max_burst, target | iops |
104
+| azure_monitor.virtual_machines.os_disk_burst_credits | bandwidth, io | percentage |
105
+| azure_monitor.virtual_machines.data_disk_throughput | read, write | bytes/s |
106
+| azure_monitor.virtual_machines.data_disk_iops | read, write | operations/s |
107
+| azure_monitor.virtual_machines.data_disk_latency | average | milliseconds |
108
+| azure_monitor.virtual_machines.data_disk_queue_depth | average | operations |
109
+| azure_monitor.virtual_machines.data_disk_throttling | bandwidth, iops | percentage |
110
+| azure_monitor.virtual_machines.data_disk_burst_capacity | max_burst, target | bytes/s |
111
+| azure_monitor.virtual_machines.data_disk_burst_iops_capacity | max_burst, target | iops |
112
+| azure_monitor.virtual_machines.data_disk_burst_credits | bandwidth, io | percentage |
113
+| azure_monitor.virtual_machines.temp_disk_throughput | read, write | bytes/s |
114
+| azure_monitor.virtual_machines.temp_disk_iops | read, write | operations/s |
115
+| azure_monitor.virtual_machines.temp_disk_latency | average | milliseconds |
116
+| azure_monitor.virtual_machines.temp_disk_queue_depth | average | operations |
117
+| azure_monitor.virtual_machines.premium_data_disk_cache | hit, miss | percentage |
118
+| azure_monitor.virtual_machines.premium_os_disk_cache | hit, miss | percentage |
119
+| azure_monitor.virtual_machines.vm_cached_throttling | bandwidth, iops | percentage |
120
+| azure_monitor.virtual_machines.vm_uncached_throttling | bandwidth, iops | percentage |
121
+| azure_monitor.virtual_machines.vm_cached_burst_credits | bandwidth, io | percentage |
122
+| azure_monitor.virtual_machines.vm_uncached_burst_credits | bandwidth, io | percentage |
123
+
124
+
125
+
126
+## Alerts
127
+
128
+
129
+The following alerts are available:
130
+
131
+| Alert name | On metric | Description |
132
+|:------------|:----------|:------------|
133
+| [ am_vm_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.cpu | VM CPU on ${label:resource_name} |
134
+| [ am_vm_cpu_credits_remaining ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.cpu_credits | VM CPU credits low on ${label:resource_name} |
135
+| [ am_vm_memory_available ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.memory_percentage | VM available memory on ${label:resource_name} |
136
+| [ am_vm_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.availability | VM unavailable ${label:resource_name} |
137
+| [ am_vm_os_disk_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.os_disk_latency | VM OS disk latency on ${label:resource_name} |
138
+| [ am_vm_os_disk_queue_depth ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.os_disk_queue_depth | VM OS disk queue depth on ${label:resource_name} |
139
+| [ am_vm_os_disk_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.os_disk_throttling | VM OS disk bandwidth throttling on ${label:resource_name} |
140
+| [ am_vm_os_disk_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.os_disk_throttling | VM OS disk IOPS throttling on ${label:resource_name} |
141
+| [ am_vm_os_disk_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.os_disk_burst_credits | VM OS disk burst bandwidth credits on ${label:resource_name} |
142
+| [ am_vm_os_disk_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.os_disk_burst_credits | VM OS disk burst IO credits on ${label:resource_name} |
143
+| [ am_vm_data_disk_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.data_disk_latency | VM data disk latency on ${label:resource_name} |
144
+| [ am_vm_data_disk_queue_depth ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.data_disk_queue_depth | VM data disk queue depth on ${label:resource_name} |
145
+| [ am_vm_data_disk_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.data_disk_throttling | VM data disk bandwidth throttling on ${label:resource_name} |
146
+| [ am_vm_data_disk_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.data_disk_throttling | VM data disk IOPS throttling on ${label:resource_name} |
147
+| [ am_vm_data_disk_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.data_disk_burst_credits | VM data disk burst bandwidth credits on ${label:resource_name} |
148
+| [ am_vm_data_disk_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.data_disk_burst_credits | VM data disk burst IO credits on ${label:resource_name} |
149
+| [ am_vm_cached_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_cached_throttling | VM cached bandwidth throttling on ${label:resource_name} |
150
+| [ am_vm_cached_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_cached_throttling | VM cached IOPS throttling on ${label:resource_name} |
151
+| [ am_vm_uncached_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_uncached_throttling | VM uncached bandwidth throttling on ${label:resource_name} |
152
+| [ am_vm_uncached_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_uncached_throttling | VM uncached IOPS throttling on ${label:resource_name} |
153
+| [ am_vm_cached_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_cached_burst_credits | VM cached burst bandwidth credits on ${label:resource_name} |
154
+| [ am_vm_cached_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_cached_burst_credits | VM cached burst IO credits on ${label:resource_name} |
155
+| [ am_vm_uncached_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_uncached_burst_credits | VM uncached burst bandwidth credits on ${label:resource_name} |
156
+| [ am_vm_uncached_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_virtual_machines.conf) | azure_monitor.virtual_machines.vm_uncached_burst_credits | VM uncached burst IO credits on ${label:resource_name} |
157
+
158
+
159
+## Setup
160
+
161
+
162
+You can configure the **azure_monitor** collector in two ways:
163
+
164
+| Method | Best for | How to |
165
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
166
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
167
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
168
+
169
+:::important
170
+
171
+UI configuration requires paid Netdata Cloud plan.
172
+
173
+:::
174
+
175
+
176
+### Prerequisites
177
+
178
+#### Create an Azure monitoring principal
179
+
180
+Create a service principal or use a managed identity with the following permissions:
181
+
182
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
183
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
184
+
185
+For service principal authentication:
186
+```bash
187
+# Create the service principal
188
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
189
+ --scopes /subscriptions/<subscription-id>
190
+
191
+# Note the appId (client_id), password (client_secret), and tenant
192
+```
193
+
194
+For managed identity (on Azure VMs, VMSS, or AKS):
195
+```bash
196
+# Assign Monitoring Reader role to the VM's managed identity
197
+az role assignment create --assignee <managed-identity-principal-id> \
198
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
199
+```
200
+
201
+
202
+
203
+### Configuration
204
+
205
+#### Options
206
+
207
+The following options can be defined globally: update_every, autodetection_retry.
208
+
209
+Profile files are loaded from:
210
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
211
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
212
+
213
+User profile files with the same filename override stock profiles.
214
+
215
+
216
+<details open><summary>Config options</summary>
217
+
218
+
219
+
220
+| Group | Option | Description | Default | Required |
221
+|:------|:-----|:------------|:--------|:---------:|
222
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
223
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
224
+| **Target** | subscription_id | Azure subscription ID. | | yes |
225
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
226
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
227
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
228
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
229
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
230
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
231
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
232
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
233
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
234
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
235
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
236
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
237
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
238
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
239
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
240
+
241
+
242
+</details>
243
+
244
+
245
+#### via UI
246
+
247
+Configure the **azure_monitor** collector from the Netdata web interface:
248
+
249
+1. Go to **Nodes**.
250
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
251
+3. The **Collectors → Jobs** view opens by default.
252
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
253
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
254
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
255
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
256
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
257
+
258
+
259
+#### via File
260
+
261
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
262
+
263
+The file format is YAML. Generally, the structure is:
264
+
265
+```yaml
266
+update_every: 1
267
+autodetection_retry: 0
268
+jobs:
269
+ - name: some_name1
270
+ - name: some_name2
271
+```
272
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
273
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
274
+
275
+```bash
276
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
277
+sudo ./edit-config go.d/azure_monitor.conf
278
+```
279
+
280
+##### Examples
281
+
282
+###### Service principal (auto-discover all resources)
283
+
284
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
285
+
286
+```yaml
287
+jobs:
288
+ - name: prod
289
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
290
+ auth:
291
+ mode: service_principal
292
+ mode_service_principal:
293
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
294
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
295
+ client_secret: "your-client-secret"
296
+
297
+```
298
+###### Managed identity (Azure VM/VMSS/AKS)
299
+
300
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
301
+
302
+<details open><summary>Config</summary>
303
+
304
+```yaml
305
+jobs:
306
+ - name: prod
307
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
308
+ auth:
309
+ mode: managed_identity
310
+
311
+```
312
+</details>
313
+
314
+###### Specific profiles only
315
+
316
+Monitor only specific Azure services instead of auto-discovering all resource types.
317
+
318
+<details open><summary>Config</summary>
319
+
320
+```yaml
321
+jobs:
322
+ - name: databases
323
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ profiles:
325
+ - sql_database
326
+ - postgres_flexible
327
+ - redis_cache
328
+ auth:
329
+ mode: service_principal
330
+ mode_service_principal:
331
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
332
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
333
+ client_secret: "your-client-secret"
334
+
335
+```
336
+</details>
337
+
338
+###### Filter by resource group
339
+
340
+Only monitor resources in specific resource groups.
341
+
342
+<details open><summary>Config</summary>
343
+
344
+```yaml
345
+jobs:
346
+ - name: prod-rg
347
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
348
+ resource_groups:
349
+ - production-rg
350
+ - staging-rg
351
+ auth:
352
+ mode: default
353
+
354
+```
355
+</details>
356
+
357
+###### Azure Government cloud
358
+
359
+Connect to Azure Government cloud environment.
360
+
361
+<details open><summary>Config</summary>
362
+
363
+```yaml
364
+jobs:
365
+ - name: gov
366
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
367
+ cloud: government
368
+ auth:
369
+ mode: service_principal
370
+ mode_service_principal:
371
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
372
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
373
+ client_secret: "your-client-secret"
374
+
375
+```
376
+</details>
377
+
378
+
379
+
380
+## Troubleshooting
381
+
382
+### Debug Mode
383
+
384
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
385
+
386
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
387
+should give you clues as to why the collector isn't working.
388
+
389
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
390
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
391
+
392
+ ```bash
393
+ cd /usr/libexec/netdata/plugins.d/
394
+ ```
395
+
396
+- Switch to the `netdata` user.
397
+
398
+ ```bash
399
+ sudo -u netdata -s
400
+ ```
401
+
402
+- Run the `go.d.plugin` to debug the collector:
403
+
404
+ ```bash
405
+ ./go.d.plugin -d -m azure_monitor
406
+ ```
407
+
408
+ To debug a specific job:
409
+
410
+ ```bash
411
+ ./go.d.plugin -d -m azure_monitor -j jobName
412
+ ```
413
+
414
+### Getting Logs
415
+
416
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
417
+
418
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
419
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
420
+
421
+#### System with systemd
422
+
423
+Use the following command to view logs generated since the last Netdata service restart:
424
+
425
+```bash
426
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
427
+```
428
+
429
+#### System without systemd
430
+
431
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
432
+
433
+```bash
434
+grep azure_monitor /var/log/netdata/collector.log
435
+```
436
+
437
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
438
+
439
+#### Docker Container
440
+
441
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
442
+
443
+```bash
444
+docker logs netdata 2>&1 | grep azure_monitor
445
+```
446
+
447
+### No metrics are collected
448
+
449
+Verify the following:
450
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
451
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
452
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
453
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
454
+
455
+
456
+### Missing metrics for some resource types
457
+
458
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
459
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
460
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
461
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
462
+
463
+
464
+### Metrics appear delayed
465
+
466
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
467
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
468
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
469
+
470
+
471
+### Authentication errors in sovereign clouds
472
+
473
+For Azure Government or Azure China clouds, set the `cloud` parameter:
474
+- Azure Government: `cloud: government`
475
+- Azure China (21Vianet): `cloud: china`
476
+
477
+Ensure the service principal is registered in the correct cloud tenant.
478
+
479
+
480
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine_scale_set.md
new
+482
@@ -0,0 +1,482 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine_scale_set.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure Virtual Machine Scale Set"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'vmss', 'scale', 'set', 'virtual', 'machine', 'compute']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure Virtual Machine Scale Set
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor Virtual Machine Scale Sets including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates across all instances in the scale set.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.vmss.cpu | average | percentage |
88
+| azure_monitor.vmss.cpu_credits | consumed, remaining | credits |
89
+| azure_monitor.vmss.memory | available | bytes |
90
+| azure_monitor.vmss.memory_percentage | available | percentage |
91
+| azure_monitor.vmss.availability | average | state |
92
+| azure_monitor.vmss.network_traffic | in, out | bytes/s |
93
+| azure_monitor.vmss.network_flows | in, out | flows |
94
+| azure_monitor.vmss.network_flow_creation_rate | in, out | flows/s |
95
+| azure_monitor.vmss.disk_throughput | read, write | bytes/s |
96
+| azure_monitor.vmss.disk_iops | read, write | operations/s |
97
+| azure_monitor.vmss.os_disk_throughput | read, write | bytes/s |
98
+| azure_monitor.vmss.os_disk_iops | read, write | operations/s |
99
+| azure_monitor.vmss.os_disk_latency | average | milliseconds |
100
+| azure_monitor.vmss.os_disk_queue_depth | average | operations |
101
+| azure_monitor.vmss.os_disk_throttling | bandwidth, iops | percentage |
102
+| azure_monitor.vmss.os_disk_burst_capacity | max_burst, target | bytes/s |
103
+| azure_monitor.vmss.os_disk_burst_iops_capacity | max_burst, target | iops |
104
+| azure_monitor.vmss.os_disk_burst_credits | bandwidth, io | percentage |
105
+| azure_monitor.vmss.data_disk_throughput | read, write | bytes/s |
106
+| azure_monitor.vmss.data_disk_iops | read, write | operations/s |
107
+| azure_monitor.vmss.data_disk_latency | average | milliseconds |
108
+| azure_monitor.vmss.data_disk_queue_depth | average | operations |
109
+| azure_monitor.vmss.data_disk_throttling | bandwidth, iops | percentage |
110
+| azure_monitor.vmss.data_disk_burst_capacity | max_burst, target | bytes/s |
111
+| azure_monitor.vmss.data_disk_burst_iops_capacity | max_burst, target | iops |
112
+| azure_monitor.vmss.data_disk_burst_credits | bandwidth, io | percentage |
113
+| azure_monitor.vmss.temp_disk_throughput | read, write | bytes/s |
114
+| azure_monitor.vmss.temp_disk_iops | read, write | operations/s |
115
+| azure_monitor.vmss.temp_disk_latency | average | milliseconds |
116
+| azure_monitor.vmss.temp_disk_queue_depth | average | operations |
117
+| azure_monitor.vmss.premium_data_disk_cache | hit, miss | percentage |
118
+| azure_monitor.vmss.premium_os_disk_cache | hit, miss | percentage |
119
+| azure_monitor.vmss.vm_cached_throttling | bandwidth, iops | percentage |
120
+| azure_monitor.vmss.vm_uncached_throttling | bandwidth, iops | percentage |
121
+| azure_monitor.vmss.vm_cached_burst_credits | bandwidth, io | percentage |
122
+| azure_monitor.vmss.vm_uncached_burst_credits | bandwidth, io | percentage |
123
+
124
+
125
+
126
+## Alerts
127
+
128
+
129
+The following alerts are available:
130
+
131
+| Alert name | On metric | Description |
132
+|:------------|:----------|:------------|
133
+| [ am_vmss_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.cpu | VMSS CPU utilization on ${label:resource_name} |
134
+| [ am_vmss_cpu_credits_remaining ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.cpu_credits | VMSS CPU credits low on ${label:resource_name} |
135
+| [ am_vmss_memory_available ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.memory_percentage | VMSS available memory low on ${label:resource_name} |
136
+| [ am_vmss_availability ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.availability | VMSS availability degraded on ${label:resource_name} |
137
+| [ am_vmss_os_disk_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.os_disk_latency | VMSS OS disk latency on ${label:resource_name} |
138
+| [ am_vmss_os_disk_queue_depth ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.os_disk_queue_depth | VMSS OS disk queue depth on ${label:resource_name} |
139
+| [ am_vmss_os_disk_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.os_disk_throttling | VMSS OS disk bandwidth consumed on ${label:resource_name} |
140
+| [ am_vmss_os_disk_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.os_disk_throttling | VMSS OS disk IOPS consumed on ${label:resource_name} |
141
+| [ am_vmss_os_disk_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.os_disk_burst_credits | VMSS OS disk burst BPS credits depleting on ${label:resource_name} |
142
+| [ am_vmss_os_disk_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.os_disk_burst_credits | VMSS OS disk burst IO credits depleting on ${label:resource_name} |
143
+| [ am_vmss_data_disk_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.data_disk_latency | VMSS data disk latency on ${label:resource_name} |
144
+| [ am_vmss_data_disk_queue_depth ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.data_disk_queue_depth | VMSS data disk queue depth on ${label:resource_name} |
145
+| [ am_vmss_data_disk_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.data_disk_throttling | VMSS data disk bandwidth consumed on ${label:resource_name} |
146
+| [ am_vmss_data_disk_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.data_disk_throttling | VMSS data disk IOPS consumed on ${label:resource_name} |
147
+| [ am_vmss_data_disk_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.data_disk_burst_credits | VMSS data disk burst BPS credits depleting on ${label:resource_name} |
148
+| [ am_vmss_data_disk_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.data_disk_burst_credits | VMSS data disk burst IO credits depleting on ${label:resource_name} |
149
+| [ am_vmss_temp_disk_latency ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.temp_disk_latency | VMSS temp disk latency on ${label:resource_name} |
150
+| [ am_vmss_temp_disk_queue_depth ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.temp_disk_queue_depth | VMSS temp disk queue depth on ${label:resource_name} |
151
+| [ am_vmss_vm_cached_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_cached_throttling | VMSS cached bandwidth consumed on ${label:resource_name} |
152
+| [ am_vmss_vm_cached_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_cached_throttling | VMSS cached IOPS consumed on ${label:resource_name} |
153
+| [ am_vmss_vm_uncached_bandwidth_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_uncached_throttling | VMSS uncached bandwidth consumed on ${label:resource_name} |
154
+| [ am_vmss_vm_uncached_iops_throttling ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_uncached_throttling | VMSS uncached IOPS consumed on ${label:resource_name} |
155
+| [ am_vmss_vm_cached_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_cached_burst_credits | VMSS cached burst BPS credits depleting on ${label:resource_name} |
156
+| [ am_vmss_vm_cached_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_cached_burst_credits | VMSS cached burst IO credits depleting on ${label:resource_name} |
157
+| [ am_vmss_vm_uncached_burst_bps_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_uncached_burst_credits | VMSS uncached burst BPS credits depleting on ${label:resource_name} |
158
+| [ am_vmss_vm_uncached_burst_io_credits ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vmss.conf) | azure_monitor.vmss.vm_uncached_burst_credits | VMSS uncached burst IO credits depleting on ${label:resource_name} |
159
+
160
+
161
+## Setup
162
+
163
+
164
+You can configure the **azure_monitor** collector in two ways:
165
+
166
+| Method | Best for | How to |
167
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
168
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
169
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
170
+
171
+:::important
172
+
173
+UI configuration requires paid Netdata Cloud plan.
174
+
175
+:::
176
+
177
+
178
+### Prerequisites
179
+
180
+#### Create an Azure monitoring principal
181
+
182
+Create a service principal or use a managed identity with the following permissions:
183
+
184
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
185
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
186
+
187
+For service principal authentication:
188
+```bash
189
+# Create the service principal
190
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
191
+ --scopes /subscriptions/<subscription-id>
192
+
193
+# Note the appId (client_id), password (client_secret), and tenant
194
+```
195
+
196
+For managed identity (on Azure VMs, VMSS, or AKS):
197
+```bash
198
+# Assign Monitoring Reader role to the VM's managed identity
199
+az role assignment create --assignee <managed-identity-principal-id> \
200
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
201
+```
202
+
203
+
204
+
205
+### Configuration
206
+
207
+#### Options
208
+
209
+The following options can be defined globally: update_every, autodetection_retry.
210
+
211
+Profile files are loaded from:
212
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
213
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
214
+
215
+User profile files with the same filename override stock profiles.
216
+
217
+
218
+<details open><summary>Config options</summary>
219
+
220
+
221
+
222
+| Group | Option | Description | Default | Required |
223
+|:------|:-----|:------------|:--------|:---------:|
224
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
225
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
226
+| **Target** | subscription_id | Azure subscription ID. | | yes |
227
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
228
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
229
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
230
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
231
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
232
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
233
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
234
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
235
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
236
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
237
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
238
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
239
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
240
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
241
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
242
+
243
+
244
+</details>
245
+
246
+
247
+#### via UI
248
+
249
+Configure the **azure_monitor** collector from the Netdata web interface:
250
+
251
+1. Go to **Nodes**.
252
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
253
+3. The **Collectors → Jobs** view opens by default.
254
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
255
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
256
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
257
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
258
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
259
+
260
+
261
+#### via File
262
+
263
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
264
+
265
+The file format is YAML. Generally, the structure is:
266
+
267
+```yaml
268
+update_every: 1
269
+autodetection_retry: 0
270
+jobs:
271
+ - name: some_name1
272
+ - name: some_name2
273
+```
274
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
275
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
276
+
277
+```bash
278
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
279
+sudo ./edit-config go.d/azure_monitor.conf
280
+```
281
+
282
+##### Examples
283
+
284
+###### Service principal (auto-discover all resources)
285
+
286
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
287
+
288
+```yaml
289
+jobs:
290
+ - name: prod
291
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
292
+ auth:
293
+ mode: service_principal
294
+ mode_service_principal:
295
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
296
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
297
+ client_secret: "your-client-secret"
298
+
299
+```
300
+###### Managed identity (Azure VM/VMSS/AKS)
301
+
302
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
303
+
304
+<details open><summary>Config</summary>
305
+
306
+```yaml
307
+jobs:
308
+ - name: prod
309
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
310
+ auth:
311
+ mode: managed_identity
312
+
313
+```
314
+</details>
315
+
316
+###### Specific profiles only
317
+
318
+Monitor only specific Azure services instead of auto-discovering all resource types.
319
+
320
+<details open><summary>Config</summary>
321
+
322
+```yaml
323
+jobs:
324
+ - name: databases
325
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ profiles:
327
+ - sql_database
328
+ - postgres_flexible
329
+ - redis_cache
330
+ auth:
331
+ mode: service_principal
332
+ mode_service_principal:
333
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
334
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
335
+ client_secret: "your-client-secret"
336
+
337
+```
338
+</details>
339
+
340
+###### Filter by resource group
341
+
342
+Only monitor resources in specific resource groups.
343
+
344
+<details open><summary>Config</summary>
345
+
346
+```yaml
347
+jobs:
348
+ - name: prod-rg
349
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
350
+ resource_groups:
351
+ - production-rg
352
+ - staging-rg
353
+ auth:
354
+ mode: default
355
+
356
+```
357
+</details>
358
+
359
+###### Azure Government cloud
360
+
361
+Connect to Azure Government cloud environment.
362
+
363
+<details open><summary>Config</summary>
364
+
365
+```yaml
366
+jobs:
367
+ - name: gov
368
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
369
+ cloud: government
370
+ auth:
371
+ mode: service_principal
372
+ mode_service_principal:
373
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
374
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
375
+ client_secret: "your-client-secret"
376
+
377
+```
378
+</details>
379
+
380
+
381
+
382
+## Troubleshooting
383
+
384
+### Debug Mode
385
+
386
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
387
+
388
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
389
+should give you clues as to why the collector isn't working.
390
+
391
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
392
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
393
+
394
+ ```bash
395
+ cd /usr/libexec/netdata/plugins.d/
396
+ ```
397
+
398
+- Switch to the `netdata` user.
399
+
400
+ ```bash
401
+ sudo -u netdata -s
402
+ ```
403
+
404
+- Run the `go.d.plugin` to debug the collector:
405
+
406
+ ```bash
407
+ ./go.d.plugin -d -m azure_monitor
408
+ ```
409
+
410
+ To debug a specific job:
411
+
412
+ ```bash
413
+ ./go.d.plugin -d -m azure_monitor -j jobName
414
+ ```
415
+
416
+### Getting Logs
417
+
418
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
419
+
420
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
421
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
422
+
423
+#### System with systemd
424
+
425
+Use the following command to view logs generated since the last Netdata service restart:
426
+
427
+```bash
428
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
429
+```
430
+
431
+#### System without systemd
432
+
433
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
434
+
435
+```bash
436
+grep azure_monitor /var/log/netdata/collector.log
437
+```
438
+
439
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
440
+
441
+#### Docker Container
442
+
443
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
444
+
445
+```bash
446
+docker logs netdata 2>&1 | grep azure_monitor
447
+```
448
+
449
+### No metrics are collected
450
+
451
+Verify the following:
452
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
453
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
454
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
455
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
456
+
457
+
458
+### Missing metrics for some resource types
459
+
460
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
461
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
462
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
463
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
464
+
465
+
466
+### Metrics appear delayed
467
+
468
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
469
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
470
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
471
+
472
+
473
+### Authentication errors in sovereign clouds
474
+
475
+For Azure Government or Azure China clouds, set the `cloud` parameter:
476
+- Azure Government: `cloud: government`
477
+- Azure China (21Vianet): `cloud: china`
478
+
479
+Ensure the service principal is registered in the correct cloud tenant.
480
+
481
+
482
+
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_vpn_gateway.md
new
+473
@@ -0,0 +1,473 @@
1
+<!--startmeta
2
+custom_edit_url: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_vpn_gateway.md"
3
+meta_yaml: "https://github.com/netdata/netdata/edit/master/src/go/plugin/go.d/collector/azure_monitor/metadata.yaml"
4
+sidebar_label: "Azure VPN Gateway"
5
+learn_status: "Published"
6
+learn_rel_path: "Collecting Metrics/Cloud and DevOps"
7
+keywords: ['azure', 'vpn', 'gateway', 'networking', 'ipsec', 'tunnel']
8
+message: "DO NOT EDIT THIS FILE DIRECTLY, IT IS GENERATED BY THE COLLECTOR'S metadata.yaml FILE"
9
+endmeta-->
10
+
11
+# Azure VPN Gateway
12
+
13
+
14
+<img src="https://netdata.cloud/img/azure.svg" width="150"/>
15
+
16
+
17
+Plugin: go.d.plugin
18
+Module: azure_monitor
19
+
20
+<img src="https://img.shields.io/badge/maintained%20by-Netdata-%2300ab44" />
21
+
22
+## Overview
23
+
24
+Monitor VPN Gateway including site-to-site bandwidth and BGP peer status, point-to-site connection counts and bandwidth, per-tunnel ingress and egress traffic with packet counts and drops, IPsec security association counts, route table sizes, NAT flow counts and packet translations, and gateway-level bandwidth utilization.
25
+
26
+
27
+The collector uses Azure SDK clients for:
28
+- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
+- Resource discovery via Azure Resource Graph queries
30
+- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
31
+
32
+
33
+This collector is supported on all platforms.
34
+
35
+This collector supports collecting metrics from multiple instances of this integration, including remote instances.
36
+
37
+The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
38
+
39
+
40
+### Default Behavior
41
+
42
+#### Auto-Detection
43
+
44
+When `profiles` includes `auto` (the default), the collector queries Azure Resource Graph
45
+to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
46
+
47
+
48
+#### Limits
49
+
50
+Azure Monitor metrics granularity is typically 1 minute.
51
+The collector enforces a minimum collection interval of 60 seconds.
52
+
53
+
54
+#### Performance Impact
55
+
56
+The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
+Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
58
+
59
+
60
+## Metrics
61
+
62
+Metrics grouped by *scope*.
63
+
64
+The scope defines the instance that the metric belongs to. An instance is uniquely identified by a set of labels.
65
+
66
+
67
+
68
+### Per resource
69
+
70
+These metrics refer to each monitored Azure resource.
71
+
72
+Labels:
73
+
74
+| Label | Description |
75
+|:-----------|:----------------|
76
+| resource_name | The Azure resource name. |
77
+| resource_group | The Azure resource group. |
78
+| region | The Azure region where the resource is deployed. |
79
+| resource_type | The Azure resource type identifier. |
80
+| profile | The Azure Monitor profile id. |
81
+| resource_uid | The unique Azure resource identifier. |
82
+
83
+Metrics:
84
+
85
+| Metric | Dimensions | Unit |
86
+|:------|:----------|:----|
87
+| azure_monitor.vpn_gateway.s2s_bandwidth | average | bytes/s |
88
+| azure_monitor.vpn_gateway.p2s_bandwidth | average | bytes/s |
89
+| azure_monitor.vpn_gateway.p2s_connections | total | connections |
90
+| azure_monitor.vpn_gateway.gateway_flows | inbound, outbound | flows |
91
+| azure_monitor.vpn_gateway.tunnel_bandwidth | average | bytes/s |
92
+| azure_monitor.vpn_gateway.tunnel_bytes | egress, ingress | bytes/s |
93
+| azure_monitor.vpn_gateway.tunnel_packets | egress, ingress | packets/s |
94
+| azure_monitor.vpn_gateway.tunnel_peak_pps | peak | packets/s |
95
+| azure_monitor.vpn_gateway.tunnel_total_flows | total | flows/s |
96
+| azure_monitor.vpn_gateway.tunnel_nat_allocations | total | allocations/s |
97
+| azure_monitor.vpn_gateway.tunnel_nated_bytes | nated, reverse_nated | bytes/s |
98
+| azure_monitor.vpn_gateway.tunnel_nated_packets | nated, reverse_nated | packets/s |
99
+| azure_monitor.vpn_gateway.tunnel_nat_flows | total | flows/s |
100
+| azure_monitor.vpn_gateway.tunnel_packet_drops | egress, ingress | packets/s |
101
+| azure_monitor.vpn_gateway.tunnel_ts_mismatch_drops | egress, ingress | packets/s |
102
+| azure_monitor.vpn_gateway.tunnel_nat_packet_drops | total | packets/s |
103
+| azure_monitor.vpn_gateway.ipsec_sa | mmsa, qmsa | associations |
104
+| azure_monitor.vpn_gateway.bgp_peer_status | average | status |
105
+| azure_monitor.vpn_gateway.bgp_routes | advertised, learned | routes/s |
106
+| azure_monitor.vpn_gateway.route_counts | user_vpn, vnet_prefix | routes |
107
+| azure_monitor.vpn_gateway.er_gateway_bandwidth | average | bits/s |
108
+| azure_monitor.vpn_gateway.er_gateway_cpu | average | percentage |
109
+| azure_monitor.vpn_gateway.er_gateway_packets | average | packets/s |
110
+| azure_monitor.vpn_gateway.er_gateway_active_flows | average | flows |
111
+| azure_monitor.vpn_gateway.er_gateway_routes_advertised | maximum | routes |
112
+| azure_monitor.vpn_gateway.er_gateway_routes_learned | maximum | routes |
113
+| azure_monitor.vpn_gateway.er_gateway_route_changes | total | changes/s |
114
+| azure_monitor.vpn_gateway.er_gateway_max_flows_rate | maximum | flows/s |
115
+| azure_monitor.vpn_gateway.er_gateway_vm_count | maximum | VMs |
116
+| azure_monitor.vpn_gateway.scalable_er_bandwidth | average | bits/s |
117
+| azure_monitor.vpn_gateway.scalable_er_cpu | average | percentage |
118
+| azure_monitor.vpn_gateway.scalable_er_packets | average | packets/s |
119
+| azure_monitor.vpn_gateway.scalable_er_active_flows | average | flows |
120
+| azure_monitor.vpn_gateway.scalable_er_routes_advertised | maximum | routes |
121
+| azure_monitor.vpn_gateway.scalable_er_routes_learned | maximum | routes |
122
+| azure_monitor.vpn_gateway.scalable_er_route_changes | total | changes/s |
123
+| azure_monitor.vpn_gateway.scalable_er_max_flows_rate | maximum | flows/s |
124
+| azure_monitor.vpn_gateway.scalable_er_vm_count | maximum | VMs |
125
+| azure_monitor.vpn_gateway.scalable_er_scale_units | maximum | units |
126
+
127
+
128
+
129
+## Alerts
130
+
131
+
132
+The following alerts are available:
133
+
134
+| Alert name | On metric | Description |
135
+|:------------|:----------|:------------|
136
+| [ am_vpn_gateway_tunnel_packet_drops ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.tunnel_packet_drops | VPN Gateway tunnel packet drops on ${label:resource_name} |
137
+| [ am_vpn_gateway_tunnel_ts_mismatch_drops ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.tunnel_ts_mismatch_drops | VPN Gateway TS mismatch drops on ${label:resource_name} |
138
+| [ am_vpn_gateway_tunnel_nat_packet_drops ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.tunnel_nat_packet_drops | VPN Gateway NAT packet drops on ${label:resource_name} |
139
+| [ am_vpn_gateway_bgp_peer_status ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.bgp_peer_status | VPN Gateway BGP peer down on ${label:resource_name} |
140
+| [ am_vpn_gateway_er_gateway_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.er_gateway_cpu | VPN GW ExpressRoute CPU on ${label:resource_name} |
141
+| [ am_vpn_gateway_er_gateway_active_flows ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.er_gateway_active_flows | VPN GW ExpressRoute active flows on ${label:resource_name} |
142
+| [ am_vpn_gateway_er_gateway_route_changes ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.er_gateway_route_changes | VPN GW ExpressRoute route churn on ${label:resource_name} |
143
+| [ am_vpn_gateway_er_gateway_routes_advertised ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.er_gateway_routes_advertised | VPN GW ExpressRoute routes advertised on ${label:resource_name} |
144
+| [ am_vpn_gateway_er_gateway_routes_learned ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.er_gateway_routes_learned | VPN GW ExpressRoute routes learned on ${label:resource_name} |
145
+| [ am_vpn_gateway_scalable_er_cpu ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.scalable_er_cpu | VPN GW Scalable ER CPU on ${label:resource_name} |
146
+| [ am_vpn_gateway_scalable_er_active_flows ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.scalable_er_active_flows | VPN GW Scalable ER active flows on ${label:resource_name} |
147
+| [ am_vpn_gateway_scalable_er_route_changes ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.scalable_er_route_changes | VPN GW Scalable ER route churn on ${label:resource_name} |
148
+| [ am_vpn_gateway_scalable_er_routes_advertised ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.scalable_er_routes_advertised | VPN GW Scalable ER routes advertised on ${label:resource_name} |
149
+| [ am_vpn_gateway_scalable_er_routes_learned ](https://github.com/netdata/netdata/blob/master/src/health/health.d/azure_monitor_vpn_gateway.conf) | azure_monitor.vpn_gateway.scalable_er_routes_learned | VPN GW Scalable ER routes learned on ${label:resource_name} |
150
+
151
+
152
+## Setup
153
+
154
+
155
+You can configure the **azure_monitor** collector in two ways:
156
+
157
+| Method | Best for | How to |
158
+|-----------------------|------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
159
+| [**UI**](#via-ui) | Fast setup without editing files | Go to **Nodes → Configure this node → Collectors → Jobs**, search for **azure_monitor**, then click **+** to add a job. |
160
+| [**File**](#via-file) | If you prefer configuring via file, or need to automate deployments (e.g., with Ansible) | Edit `go.d/azure_monitor.conf` and add a job. |
161
+
162
+:::important
163
+
164
+UI configuration requires paid Netdata Cloud plan.
165
+
166
+:::
167
+
168
+
169
+### Prerequisites
170
+
171
+#### Create an Azure monitoring principal
172
+
173
+Create a service principal or use a managed identity with the following permissions:
174
+
175
+1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
176
+2. **Reader** role for Azure Resource Graph queries (for resource discovery)
177
+
178
+For service principal authentication:
179
+```bash
180
+# Create the service principal
181
+az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
182
+ --scopes /subscriptions/<subscription-id>
183
+
184
+# Note the appId (client_id), password (client_secret), and tenant
185
+```
186
+
187
+For managed identity (on Azure VMs, VMSS, or AKS):
188
+```bash
189
+# Assign Monitoring Reader role to the VM's managed identity
190
+az role assignment create --assignee <managed-identity-principal-id> \
191
+ --role "Monitoring Reader" --scope /subscriptions/<subscription-id>
192
+```
193
+
194
+
195
+
196
+### Configuration
197
+
198
+#### Options
199
+
200
+The following options can be defined globally: update_every, autodetection_retry.
201
+
202
+Profile files are loaded from:
203
+- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
204
+- User: `/etc/netdata/go.d/azure_monitor.profiles/`
205
+
206
+User profile files with the same filename override stock profiles.
207
+
208
+
209
+<details open><summary>Config options</summary>
210
+
211
+
212
+
213
+| Group | Option | Description | Default | Required |
214
+|:------|:-----|:------------|:--------|:---------:|
215
+| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
216
+| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
217
+| **Target** | subscription_id | Azure subscription ID. | | yes |
218
+| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
219
+| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
220
+| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
221
+| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
222
+| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
223
+| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
224
+| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
225
+| **Profiles** | profiles | Profile ids to enable. Use `auto` to discover resource types via Azure Resource Graph and enable matching profiles. Combine with explicit ids: `[auto, custom_profile]`. | [auto] | no |
226
+| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
227
+| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
228
+| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
229
+| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
230
+| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
231
+| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
232
+| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
233
+
234
+
235
+</details>
236
+
237
+
238
+#### via UI
239
+
240
+Configure the **azure_monitor** collector from the Netdata web interface:
241
+
242
+1. Go to **Nodes**.
243
+2. Select the node **where you want the azure_monitor data-collection job to run** and click the :gear: (**Configure this node**). That node will run the data collection.
244
+3. The **Collectors → Jobs** view opens by default.
245
+4. In the Search box, type _azure_monitor_ (or scroll the list) to locate the **azure_monitor** collector.
246
+5. Click the **+** next to the **azure_monitor** collector to add a new job.
247
+6. Fill in the job fields, then click **Test** to verify the configuration and **Submit** to save.
248
+ - **Test** runs the job with the provided settings and shows whether data can be collected.
249
+ - If it fails, an error message appears with details (for example, connection refused, timeout, or command execution errors), so you can adjust and retest.
250
+
251
+
252
+#### via File
253
+
254
+The configuration file name for this integration is `go.d/azure_monitor.conf`.
255
+
256
+The file format is YAML. Generally, the structure is:
257
+
258
+```yaml
259
+update_every: 1
260
+autodetection_retry: 0
261
+jobs:
262
+ - name: some_name1
263
+ - name: some_name2
264
+```
265
+You can edit the configuration file using the [`edit-config`](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#edit-configuration-files) script from the
266
+Netdata [config directory](https://github.com/netdata/netdata/blob/master/docs/netdata-agent/configuration/README.md#locate-your-config-directory).
267
+
268
+```bash
269
+cd /etc/netdata 2>/dev/null || cd /opt/netdata/etc/netdata
270
+sudo ./edit-config go.d/azure_monitor.conf
271
+```
272
+
273
+##### Examples
274
+
275
+###### Service principal (auto-discover all resources)
276
+
277
+Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
278
+
279
+```yaml
280
+jobs:
281
+ - name: prod
282
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
283
+ auth:
284
+ mode: service_principal
285
+ mode_service_principal:
286
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
287
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
288
+ client_secret: "your-client-secret"
289
+
290
+```
291
+###### Managed identity (Azure VM/VMSS/AKS)
292
+
293
+Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
294
+
295
+<details open><summary>Config</summary>
296
+
297
+```yaml
298
+jobs:
299
+ - name: prod
300
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
301
+ auth:
302
+ mode: managed_identity
303
+
304
+```
305
+</details>
306
+
307
+###### Specific profiles only
308
+
309
+Monitor only specific Azure services instead of auto-discovering all resource types.
310
+
311
+<details open><summary>Config</summary>
312
+
313
+```yaml
314
+jobs:
315
+ - name: databases
316
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
317
+ profiles:
318
+ - sql_database
319
+ - postgres_flexible
320
+ - redis_cache
321
+ auth:
322
+ mode: service_principal
323
+ mode_service_principal:
324
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ client_secret: "your-client-secret"
327
+
328
+```
329
+</details>
330
+
331
+###### Filter by resource group
332
+
333
+Only monitor resources in specific resource groups.
334
+
335
+<details open><summary>Config</summary>
336
+
337
+```yaml
338
+jobs:
339
+ - name: prod-rg
340
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
341
+ resource_groups:
342
+ - production-rg
343
+ - staging-rg
344
+ auth:
345
+ mode: default
346
+
347
+```
348
+</details>
349
+
350
+###### Azure Government cloud
351
+
352
+Connect to Azure Government cloud environment.
353
+
354
+<details open><summary>Config</summary>
355
+
356
+```yaml
357
+jobs:
358
+ - name: gov
359
+ subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
+ cloud: government
361
+ auth:
362
+ mode: service_principal
363
+ mode_service_principal:
364
+ tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
365
+ client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
366
+ client_secret: "your-client-secret"
367
+
368
+```
369
+</details>
370
+
371
+
372
+
373
+## Troubleshooting
374
+
375
+### Debug Mode
376
+
377
+**Important**: Debug mode is not supported for data collection jobs created via the UI using the Dyncfg feature.
378
+
379
+To troubleshoot issues with the `azure_monitor` collector, run the `go.d.plugin` with the debug option enabled. The output
380
+should give you clues as to why the collector isn't working.
381
+
382
+- Navigate to the `plugins.d` directory, usually at `/usr/libexec/netdata/plugins.d/`. If that's not the case on
383
+ your system, open `netdata.conf` and look for the `plugins` setting under `[directories]`.
384
+
385
+ ```bash
386
+ cd /usr/libexec/netdata/plugins.d/
387
+ ```
388
+
389
+- Switch to the `netdata` user.
390
+
391
+ ```bash
392
+ sudo -u netdata -s
393
+ ```
394
+
395
+- Run the `go.d.plugin` to debug the collector:
396
+
397
+ ```bash
398
+ ./go.d.plugin -d -m azure_monitor
399
+ ```
400
+
401
+ To debug a specific job:
402
+
403
+ ```bash
404
+ ./go.d.plugin -d -m azure_monitor -j jobName
405
+ ```
406
+
407
+### Getting Logs
408
+
409
+If you're encountering problems with the `azure_monitor` collector, follow these steps to retrieve logs and identify potential issues:
410
+
411
+- **Run the command** specific to your system (systemd, non-systemd, or Docker container).
412
+- **Examine the output** for any warnings or error messages that might indicate issues. These messages should provide clues about the root cause of the problem.
413
+
414
+#### System with systemd
415
+
416
+Use the following command to view logs generated since the last Netdata service restart:
417
+
418
+```bash
419
+journalctl _SYSTEMD_INVOCATION_ID="$(systemctl show --value --property=InvocationID netdata)" --namespace=netdata --grep azure_monitor
420
+```
421
+
422
+#### System without systemd
423
+
424
+Locate the collector log file, typically at `/var/log/netdata/collector.log`, and use `grep` to filter for collector's name:
425
+
426
+```bash
427
+grep azure_monitor /var/log/netdata/collector.log
428
+```
429
+
430
+**Note**: This method shows logs from all restarts. Focus on the **latest entries** for troubleshooting current issues.
431
+
432
+#### Docker Container
433
+
434
+If your Netdata runs in a Docker container named "netdata" (replace if different), use this command:
435
+
436
+```bash
437
+docker logs netdata 2>&1 | grep azure_monitor
438
+```
439
+
440
+### No metrics are collected
441
+
442
+Verify the following:
443
+1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
444
+2. The `subscription_id` in the configuration matches the subscription containing the target resources.
445
+3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
446
+4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
447
+
448
+
449
+### Missing metrics for some resource types
450
+
451
+Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
452
+1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
453
+2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
454
+3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
455
+
456
+
457
+### Metrics appear delayed
458
+
459
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
460
+If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
461
+Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
462
+
463
+
464
+### Authentication errors in sovereign clouds
465
+
466
+For Azure Government or Azure China clouds, set the `cloud` parameter:
467
+- Azure Government: `cloud: government`
468
+- Azure China (21Vianet): `cloud: china`
469
+
470
+Ensure the service principal is registered in the correct cloud tenant.
471
+
472
+
473
+
src/go/plugin/go.d/collector/mssql/integrations/microsoft_sql_server.md
+3
@@ -1004,3 +1004,6 @@ Ensure SQL Server is configured for mixed mode authentication if using SQL login
1004
1005
The monitoring user needs VIEW SERVER STATE permission.
1006
Grant it with: `GRANT VIEW SERVER STATE TO netdata_user;`
1007
+
1008
+
1009
+
src/go/plugin/go.d/collector/postgres/integrations/postgresql.md
+2
@@ -751,3 +751,5 @@ If your Netdata runs in a Docker container named "netdata" (replace if different
751
```bash
752
docker logs netdata 2>&1 | grep postgres
753
```
754
+
755
+
src/go/plugin/go.d/collector/sql/integrations/sql_databases_generic.md
+2
@@ -897,3 +897,5 @@ If your Netdata runs in a Docker container named "netdata" (replace if different
897
```bash
898
docker logs netdata 2>&1 | grep sql
899
```
900
+
901
+