Regenerate integrations docs (#22096)
Co-authored-by: ilyam8 <22274335+ilyam8@users.noreply.github.com>
Netdata bot committed
Mar 31, 2026 at 13:22 UTC
9384321bee9fcebd5dfbf758c807bcddfd8fb7b8
41 files changed
+9622
-3526
src/collectors/COLLECTORS.md
+40
-40
@@ -315,46 +315,46 @@ Need a dedicated integration? [Submit a feature request](https://github.com/netd
315
|-------------|-------------|
316
| [AWS EC2 Compute instances](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/aws_ec2_compute_instances.md) | Track AWS EC2 instances key metrics for optimized performance and cost management. |
317
| [AWS Quota](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/aws_quota.md) | Monitor AWS service quotas for effective resource usage and cost management. |
318
-| [Azure API Management](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_api_management.md) | Monitor API Management gateway performance including request throughput, response status codes, gateway and backend response times, failed request counts, capacity utilization, event hub events, websocket message counts, and network connection status. |
319
-| [Azure App Service](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_app_service.md) | Monitor App Service web applications including HTTP request rates and response status codes, response times, CPU and memory usage, network throughput, file IO operations, .NET runtime statistics (threads, GC, assemblies), Azure Functions execution counts and units, and Flex Consumption plan metrics. |
320
-| [Azure Application Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_gateway.md) | Monitor Application Gateway performance including throughput and traffic volume, request rates and response status codes, backend health and latency breakdown (connect, first byte, last byte), client latency, current and new connections, WebSocket sessions, capacity and compute units, CPU utilization, TLS connections, and WAF security events including rule matches, challenges, and penalty box activity. |
321
-| [Azure Application Insights](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_insights.md) | Monitor application performance through Application Insights including availability test results and duration, server request rates and response times, dependency call tracking and failures, exception rates by source, browser page load timing breakdown, process CPU and memory usage, IO rates, HTTP request queue depth, page views, and trace volume. |
322
-| [Azure Cache for Redis](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cache_for_redis.md) | Monitor Azure Cache for Redis including cache hit and miss rates, read and write throughput, server load and CPU utilization, memory usage, connected clients, operations per second, command processing rates, latency percentiles, key eviction and expiration, and geo-replication health and sync status. |
323
-| [Azure Cognitive Services](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cognitive_services.md) | Monitor Azure AI and Cognitive Services including API call volume, success and client error rates, response latency, token processing rates for language models, content safety filtering, fine-tuning operations, provisioned throughput utilization, rate-limiting events, active inference connections, and context token cache performance. |
324
-| [Azure Container Apps](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_apps.md) | Monitor Container Apps including CPU and memory usage, network traffic, replica counts, request processing rates, response times, restart frequency, and resource reservation utilization. |
325
-| [Azure Container Instances](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_instances.md) | Monitor Container Instance groups including CPU and memory usage and network bytes transferred in and out. |
326
-| [Azure Container Registry](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_registry.md) | Monitor Container Registry including storage usage, successful and failed pull and push operation counts, and task run duration. |
327
-| [Azure Cosmos DB Account](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cosmos_db_account.md) | Monitor Cosmos DB accounts including request unit consumption and throttling, document counts and storage, data and index sizes, replication latency, availability percentages, provisioned throughput utilization, and normalized RU consumption per partition. |
328
-| [Azure Data Explorer Cluster](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_explorer_cluster.md) | Monitor Azure Data Explorer (Kusto) clusters including ingestion latency, volume, and success rates, query performance and concurrency, cache utilization, CPU and memory usage, export operations, streaming ingest throughput, materialized view health, instance counts, and follower lag. |
329
-| [Azure Data Factory](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_factory.md) | Monitor Data Factory including pipeline, activity, and trigger run success and failure counts, integration runtime CPU and memory utilization, available capacity and queue lengths, SSIS package execution rates, copy operations throughput, data flow processing metrics, and overall factory resource utilization. |
330
-| [Azure Event Grid Topic](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_grid_topic.md) | Monitor Event Grid topics including publish success and failure counts, publish latency, event delivery and routing rates, delivery success and failure counts, dead-lettered events, and matched event routing. |
331
-| [Azure Event Hubs Namespace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_hubs_namespace.md) | Monitor Event Hubs namespaces including incoming and outgoing message rates, byte throughput, captured messages and bytes, throttled and quota-exceeded request counts, active connections, and total connection counts. |
332
-| [Azure ExpressRoute Circuit](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_circuit.md) | Monitor ExpressRoute circuits including bits per second in and out, ARP and BGP availability percentages, packet drops, and QoS bit rate throughput. |
333
-| [Azure ExpressRoute Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_gateway.md) | Monitor ExpressRoute gateways including bits and packets per second for ingress and egress, connection counts, CPU utilization, active flow counts, and gateway scale unit counts. |
334
-| [Azure Firewall](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_firewall.md) | Monitor Azure Firewall including data processed, throughput, application and network rule hit counts, SNAT port utilization, health state percentage, and latency probes. |
335
-| [Azure Front Door](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_front_door.md) | Monitor Azure Front Door including request counts and rates, response sizes, total latency, origin health probe percentages, origin request counts, origin latency, WAF request counts by action and rule, and WebSocket connection metrics. |
336
-| [Azure Functions](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_functions.md) | Monitor Azure Functions execution including function invocation counts, execution units (MB-milliseconds), HTTP request rates and response codes, CPU and memory consumption, and Flex Consumption plan metrics for always-ready and on-demand instances. |
337
-| [Azure IoT Hub](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_iot_hub.md) | Monitor IoT Hub including device telemetry message rates and quota usage, routing delivery and latency, device twin read and write operations, direct method invocations, cloud-to-device messaging and feedback, job completion rates, device connection and authentication events, and event grid publish status. |
338
-| [Azure Key Vault](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_key_vault.md) | Monitor Key Vault including overall vault availability, API saturation approaching service limits, and service API hit and latency metrics. |
339
-| [Azure Kubernetes Service Cluster](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_kubernetes_service_cluster.md) | Monitor AKS cluster health including API server and etcd resource usage, pod scheduling status and readiness, node capacity and conditions, cluster autoscaler behavior, and per-node CPU, memory, disk, and network utilization. |
340
-| [Azure Load Balancer](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_load_balancer.md) | Monitor Azure Load Balancer health and throughput including data path and health probe availability, SYN and SNAT connection counts, byte and packet throughput, allocated and used SNAT ports, and connection attempt rates. |
341
-| [Azure Log Analytics Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_log_analytics_workspace.md) | Monitor Log Analytics workspaces including ingestion volume and latency, query execution counts and volume, available storage capacity, and per-table breakdowns of ingestion rates and billing volume. |
342
-| [Azure Logic Apps Workflow](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_logic_apps_workflow.md) | Monitor Logic Apps workflow execution including run completions and failures, action execution counts, trigger firing rates, run and action latency, billable executions, and action-level success and failure breakdowns. |
343
-| [Azure Machine Learning Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_machine_learning_workspace.md) | Monitor Azure Machine Learning workspaces including active model deployments and registered models, pipeline run completions and failures, compute node utilization and preemptions, quota usage, managed endpoint request latency and rates, estimated GPU utilization, and storage utilization. |
344
-| [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) | This collector monitors Azure resources through the Azure Monitor Metrics API. |
345
-| [Azure MySQL Flexible Server](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_mysql_flexible_server.md) | Monitor MySQL Flexible Server including active connections, aborted connections, query rates, replication lag, storage utilization, CPU and memory usage, IO operations, InnoDB buffer pool efficiency, network throughput, and HA replication status. |
346
-| [Azure NAT Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_nat_gateway.md) | Monitor NAT Gateway including byte and packet counts, connection counts, dropped packets, total SNAT connection counts, and datapath availability. |
347
-| [Azure PostgreSQL Flexible Server](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_postgresql_flexible_server.md) | Monitor PostgreSQL Flexible Server including active connections, transaction rates, replication lag, storage and backup utilization, CPU and memory usage, IO throughput, autovacuum activity, PgBouncer connection pooling, database sessions, and burstable instance CPU credits. |
348
-| [Azure Service Bus Namespace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_service_bus_namespace.md) | Monitor Service Bus namespaces including incoming and outgoing message rates, active connections, active and dead-lettered message counts, scheduled message counts, completed and abandoned requests, server errors, throttled requests, CPU and memory utilization, and pending checkpoint operations. |
349
-| [Azure SQL Database](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_database.md) | Monitor SQL Database performance including CPU and DTU utilization, storage consumption, active sessions and workers, deadlocks, IO rates, tempdb usage, in-memory OLTP storage, and serverless auto-pause and billing metrics. |
350
-| [Azure SQL Elastic Pool](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_elastic_pool.md) | Monitor SQL Elastic Pool resource consumption including eDTU and CPU utilization, storage usage, active sessions and workers, IO rates, tempdb usage, and in-memory OLTP storage across all databases in the pool. |
351
-| [Azure SQL Managed Instance](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_managed_instance.md) | Monitor SQL Managed Instance performance including virtual core CPU utilization, storage consumption, IO throughput, and average request wait times. |
352
-| [Azure Storage Account](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_storage_account.md) | Monitor Azure Storage Account operations including transaction counts, availability percentages, success and end-to-end latency, ingress and egress throughput, and used capacity. |
353
-| [Azure Stream Analytics Job](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_stream_analytics_job.md) | Monitor Stream Analytics jobs including input and output event counts, streaming unit utilization, watermark delay, backlogged input events, runtime and data conversion errors, out-of-order events, and late input events. |
354
-| [Azure Synapse Analytics Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_synapse_analytics_workspace.md) | Monitor Synapse Analytics workspaces including pipeline and activity run metrics, SQL request counts and data processing volumes, data flow activity execution, integration runtime CPU and memory utilization, and link table event processing. |
355
-| [Azure Virtual Machine](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine.md) | Monitor Azure Virtual Machines including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates. |
356
-| [Azure Virtual Machine Scale Set](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine_scale_set.md) | Monitor Virtual Machine Scale Sets including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates across all instances in the scale set. |
357
-| [Azure VPN Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_vpn_gateway.md) | Monitor VPN Gateway including site-to-site bandwidth and BGP peer status, point-to-site connection counts and bandwidth, per-tunnel ingress and egress traffic with packet counts and drops, IPsec security association counts, route table sizes, NAT flow counts and packet translations, and gateway-level bandwidth utilization. |
318
+| [Azure API Management](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_api_management.md) | :::info |
319
+| [Azure App Service](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_app_service.md) | :::info |
320
+| [Azure Application Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_gateway.md) | :::info |
321
+| [Azure Application Insights](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_insights.md) | :::info |
322
+| [Azure Cache for Redis](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cache_for_redis.md) | :::info |
323
+| [Azure Cognitive Services](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cognitive_services.md) | :::info |
324
+| [Azure Container Apps](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_apps.md) | :::info |
325
+| [Azure Container Instances](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_instances.md) | :::info |
326
+| [Azure Container Registry](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_registry.md) | :::info |
327
+| [Azure Cosmos DB Account](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cosmos_db_account.md) | :::info |
328
+| [Azure Data Explorer Cluster](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_explorer_cluster.md) | :::info |
329
+| [Azure Data Factory](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_factory.md) | :::info |
330
+| [Azure Event Grid Topic](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_grid_topic.md) | :::info |
331
+| [Azure Event Hubs Namespace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_hubs_namespace.md) | :::info |
332
+| [Azure ExpressRoute Circuit](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_circuit.md) | :::info |
333
+| [Azure ExpressRoute Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_gateway.md) | :::info |
334
+| [Azure Firewall](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_firewall.md) | :::info |
335
+| [Azure Front Door](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_front_door.md) | :::info |
336
+| [Azure Functions](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_functions.md) | :::info |
337
+| [Azure IoT Hub](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_iot_hub.md) | :::info |
338
+| [Azure Key Vault](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_key_vault.md) | :::info |
339
+| [Azure Kubernetes Service Cluster](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_kubernetes_service_cluster.md) | :::info |
340
+| [Azure Load Balancer](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_load_balancer.md) | :::info |
341
+| [Azure Log Analytics Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_log_analytics_workspace.md) | :::info |
342
+| [Azure Logic Apps Workflow](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_logic_apps_workflow.md) | :::info |
343
+| [Azure Machine Learning Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_machine_learning_workspace.md) | :::info |
344
+| [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) | This collector provides real-time visibility into your Azure infrastructure by collecting platform metrics from the Azure Monitor Metrics API. |
345
+| [Azure MySQL Flexible Server](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_mysql_flexible_server.md) | :::info |
346
+| [Azure NAT Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_nat_gateway.md) | :::info |
347
+| [Azure PostgreSQL Flexible Server](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_postgresql_flexible_server.md) | :::info |
348
+| [Azure Service Bus Namespace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_service_bus_namespace.md) | :::info |
349
+| [Azure SQL Database](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_database.md) | :::info |
350
+| [Azure SQL Elastic Pool](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_elastic_pool.md) | :::info |
351
+| [Azure SQL Managed Instance](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_managed_instance.md) | :::info |
352
+| [Azure Storage Account](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_storage_account.md) | :::info |
353
+| [Azure Stream Analytics Job](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_stream_analytics_job.md) | :::info |
354
+| [Azure Synapse Analytics Workspace](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_synapse_analytics_workspace.md) | :::info |
355
+| [Azure Virtual Machine](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine.md) | :::info |
356
+| [Azure Virtual Machine Scale Set](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine_scale_set.md) | :::info |
357
+| [Azure VPN Gateway](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_vpn_gateway.md) | :::info |
358
| [BOSH](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/bosh.md) | Keep an eye on BOSH deployment metrics for improved cloud orchestration and resource management. |
359
| [Cloud Foundry](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/cloud_foundry.md) | Track Cloud Foundry platform metrics for optimized application deployment and management. |
360
| [Cloud Foundry Firehose](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/prometheus/integrations/cloud_foundry_firehose.md) | Monitor Cloud Foundry Firehose metrics for comprehensive platform diagnostics and management. |
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_api_management.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor API Management gateway performance including request throughput, response status codes, gateway and backend response times, failed request counts, capacity utilization, event hub events, websocket message counts, and network connection status.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure API Management with metrics covering:
31
+
32
+- **Requests** -- gateway request rate
33
+- **Latency** -- request duration (overall and backend response time)
34
+- **Compute** -- gateway CPU and memory utilization
35
+- **Capacity** -- capacity utilization percentage
36
+- **Events** -- EventHub events (successful/failed/dropped/rejected/throttled/timed out), EventHub bytes sent
37
+- **WebSockets** -- WebSocket connection attempts, WebSocket messages
38
+- **Network** -- network connectivity status
39
+
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
+
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -321,6 +458,7 @@ Labels:
458
| region | The Azure region where the resource is deployed. |
459
| resource_type | The Azure resource type identifier. |
460
| profile | The Azure Monitor profile id. |
461
+| subscription_id | The Azure subscription identifier. |
462
| resource_uid | The unique Azure resource identifier. |
463
464
Metrics:
@@ -409,31 +547,46 @@ docker logs netdata 2>&1 | grep azure_monitor
547
548
### No metrics are collected
549
412
-Verify the following:
413
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
414
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
415
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
416
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
550
+Check the following:
551
+
552
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
553
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
554
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
555
+- **Collector logs** -- Check for authentication or API errors:
556
+ ```bash
557
+ # systemd
558
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
559
+ # non-systemd
560
+ grep azure_monitor /var/log/netdata/collector.log
561
+ ```
562
563
564
### Missing metrics for some resource types
565
421
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
422
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
423
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
424
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
566
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
567
+
568
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
569
+- **Verify a built-in profile exists** -- List available profiles:
570
+ ```bash
571
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
572
+ ```
573
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
574
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
575
576
427
-### Metrics appear delayed
577
+### Charts have gaps or incomplete data
578
429
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
430
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
431
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
579
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
580
+
581
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
582
+- Slower time-grain batches automatically use a larger effective offset when needed.
583
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
584
585
586
### Authentication errors in sovereign clouds
587
588
For Azure Government or Azure China clouds, set the `cloud` parameter:
589
+
590
- Azure Government: `cloud: government`
591
- Azure China (21Vianet): `cloud: china`
592
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_app_service.md
+242
-87
@@ -21,40 +21,81 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor App Service web applications including HTTP request rates and response status codes, response times, CPU and memory usage, network throughput, file IO operations, .NET runtime statistics (threads, GC, assemblies), Azure Functions execution counts and units, and Flex Consumption plan metrics.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure App Service with metrics covering:
31
+
32
+- **Requests** -- HTTP request rate, response status codes (2xx/3xx/4xx/5xx), error detail (401/403/404/406)
33
+- **Performance** -- response time, request queue depth
34
+- **Compute** -- CPU utilization, CPU time consumed
35
+- **Memory** -- memory usage (average working set, working set, private bytes)
36
+- **Network** -- network traffic (received/sent), I/O throughput (read/write/other)
37
+- **I/O** -- I/O operations (read/write/other), file handles
38
+- **.NET runtime** -- threads, GC collections (gen0/gen1/gen2), loaded assemblies, app domains
39
+- **Functions** -- function executions and execution units (MB-ms), always-ready and on-demand units
40
+- **Health** -- health check status
41
+
42
+
43
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
44
45
46
This collector is supported on all platforms.
47
48
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
49
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
50
+The service principal or managed identity requires these Azure RBAC roles:
51
+
52
+| Role | Purpose | Scope |
53
+|:-----|:--------|:------|
54
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
55
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
56
57
58
### Default Behavior
59
60
#### Auto-Detection
61
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
62
+The collector has two discovery phases:
63
+
64
+**Bootstrap (first run)**
65
+
66
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
67
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
68
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
69
+- A single job can monitor multiple subscriptions.
70
+
71
+**Runtime (periodic refresh)**
72
+
73
+- Periodically re-discovers resources for **already-active profile types only**.
74
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
75
+
76
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
77
78
79
#### Limits
80
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
81
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
82
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
83
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
84
85
86
#### Performance Impact
87
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
88
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
89
+
90
+**Default concurrency and batching limits:**
91
+
92
+| Setting | Default | Description |
93
+|:--------|:--------|:------------|
94
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
95
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
96
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
97
+
98
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
99
100
101
## Setup
@@ -78,25 +119,36 @@ UI configuration requires paid Netdata Cloud plan.
119
120
#### Create an Azure monitoring principal
121
81
-Create a service principal or use a managed identity with the following permissions:
122
+The collector requires a service principal or managed identity with two Azure RBAC roles:
123
+
124
+| Role | Purpose |
125
+|:-----|:--------|
126
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
127
+| **Reader** | Query Azure Resource Graph for resource discovery |
128
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
129
+**Option A: Service principal**
130
86
-For service principal authentication:
131
```bash
88
-# Create the service principal
132
+# Create service principal with Monitoring Reader role
133
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
134
--scopes /subscriptions/<subscription-id>
135
136
+# Add the Reader role for resource discovery
137
+az role assignment create --assignee <appId-from-above> \
138
+ --role "Reader" --scope /subscriptions/<subscription-id>
139
+
140
# Note the appId (client_id), password (client_secret), and tenant
141
```
142
95
-For managed identity (on Azure VMs, VMSS, or AKS):
143
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
144
+
145
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
146
+# Assign both roles to the VM's managed identity
147
az role assignment create --assignee <managed-identity-principal-id> \
148
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Reader" --scope /subscriptions/<subscription-id>
152
```
153
154
@@ -105,13 +157,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
157
158
#### Options
159
108
-The following options can be defined globally: update_every, autodetection_retry.
160
+The following options can be defined globally: `update_every`, `autodetection_retry`.
161
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+**Profile file locations:**
163
114
-User profile files with the same filename override stock profiles.
164
+| Type | Path |
165
+|:-----|:-----|
166
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
167
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
168
+
169
+User profile files with the same `id` as a stock profile override it.
170
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
171
172
173
<details open><summary>Config options</summary>
@@ -122,25 +178,103 @@ User profile files with the same filename override stock profiles.
178
|:------|:-----|:------------|:--------|:---------:|
179
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
180
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
181
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
182
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
186
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
191
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
192
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
195
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
196
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
197
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
198
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
199
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
204
+<a id="option-collection-query-offset"></a>
205
+##### query_offset
206
+
207
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
208
+
209
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
210
+
211
+- **Default (180s)** works for most services.
212
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
213
+- **Increase to 240-300s** if you still see gaps or missing data points.
214
+- **Do not set below 60s** -- metrics will likely be incomplete.
215
+
216
+
217
+<a id="option-authentication-auth-mode"></a>
218
+##### auth.mode
219
+
220
+Determines how the collector authenticates with Azure.
221
+
222
+| Mode | When to use | Required options |
223
+|:-----|:------------|:-----------------|
224
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
225
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
226
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
227
+
228
+
229
+<a id="option-discovery-discovery-mode"></a>
230
+##### discovery.mode
231
+
232
+Controls how the collector finds candidate Azure resources.
233
+
234
+| Mode | Behavior |
235
+|:-----|:---------|
236
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
237
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
238
+
239
+
240
+<a id="option-discovery-discovery-mode-query-kql"></a>
241
+##### discovery.mode_query.kql
242
+
243
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
244
+
245
+The query **must** project these five columns:
246
+
247
+| Column | Description |
248
+|:-------|:------------|
249
+| `id` | Full Azure resource ID (ARM format) |
250
+| `name` | Resource name |
251
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
252
+| `resourceGroup` | Resource group name |
253
+| `location` | Azure region |
254
+
255
+Example:
256
+
257
+```
258
+resources
259
+| where tags.env =~ "prod"
260
+| project id, name, type, resourceGroup, location
261
+```
262
+
263
+
264
+<a id="option-profiles-profiles-mode"></a>
265
+##### profiles.mode
266
+
267
+Controls how the collector decides which metric profiles to activate.
268
+
269
+| Mode | Behavior |
270
+|:-----|:---------|
271
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
272
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
273
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
274
+
275
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
276
+
277
+
278
279
</details>
280
@@ -182,14 +316,28 @@ sudo ./edit-config go.d/azure_monitor.conf
316
317
##### Examples
318
185
-###### Service principal (auto-discover all resources)
319
+###### Service principal with structured discovery
320
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
321
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
322
323
```yaml
324
jobs:
325
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ subscription_ids:
327
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
328
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
329
+ discovery:
330
+ mode: filters
331
+ mode_filters:
332
+ resource_groups:
333
+ - production-rg
334
+ regions:
335
+ - eastus
336
+ tags:
337
+ env:
338
+ - prod
339
+ profiles:
340
+ mode: auto
341
auth:
342
mode: service_principal
343
mode_service_principal:
@@ -198,59 +346,49 @@ jobs:
346
client_secret: "your-client-secret"
347
348
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
349
+###### Managed identity with exact profiles
350
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
351
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
352
353
<details open><summary>Config</summary>
354
355
```yaml
356
jobs:
357
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ subscription_ids:
359
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
361
+ mode: exact
362
+ mode_exact:
363
+ names:
364
+ - sql_database
365
+ - postgres_flexible
366
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
367
+ mode: managed_identity
368
369
```
370
</details>
371
241
-###### Filter by resource group
372
+###### Custom Azure Resource Graph KQL
373
243
-Only monitor resources in specific resource groups.
374
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
375
376
<details open><summary>Config</summary>
377
378
```yaml
379
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
380
+ - name: prod-query
381
+ subscription_ids:
382
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
383
+ discovery:
384
+ mode: query
385
+ mode_query:
386
+ kql: |
387
+ resources
388
+ | where tags.env =~ "prod"
389
+ | project id, name, type, resourceGroup, location
390
+ profiles:
391
+ mode: auto
392
auth:
393
mode: default
394
@@ -259,14 +397,15 @@ jobs:
397
398
###### Azure Government cloud
399
262
-Connect to Azure Government cloud environment.
400
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
401
402
<details open><summary>Config</summary>
403
404
```yaml
405
jobs:
406
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
+ subscription_ids:
408
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
409
cloud: government
410
auth:
411
mode: service_principal
@@ -318,6 +457,7 @@ Labels:
457
| region | The Azure region where the resource is deployed. |
458
| resource_type | The Azure resource type identifier. |
459
| profile | The Azure Monitor profile id. |
460
+| subscription_id | The Azure subscription identifier. |
461
| resource_uid | The unique Azure resource identifier. |
462
463
Metrics:
@@ -422,31 +562,46 @@ docker logs netdata 2>&1 | grep azure_monitor
562
563
### No metrics are collected
564
425
-Verify the following:
426
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
427
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
428
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
429
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
565
+Check the following:
566
+
567
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
568
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
569
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
570
+- **Collector logs** -- Check for authentication or API errors:
571
+ ```bash
572
+ # systemd
573
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
574
+ # non-systemd
575
+ grep azure_monitor /var/log/netdata/collector.log
576
+ ```
577
578
579
### Missing metrics for some resource types
580
434
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
435
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
436
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
437
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
581
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
582
+
583
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
584
+- **Verify a built-in profile exists** -- List available profiles:
585
+ ```bash
586
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
587
+ ```
588
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
589
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
590
591
440
-### Metrics appear delayed
592
+### Charts have gaps or incomplete data
593
442
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
443
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
444
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
594
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
595
+
596
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
597
+- Slower time-grain batches automatically use a larger effective offset when needed.
598
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
599
600
601
### Authentication errors in sovereign clouds
602
603
For Azure Government or Azure China clouds, set the `cloud` parameter:
604
+
605
- Azure Government: `cloud: government`
606
- Azure China (21Vianet): `cloud: china`
607
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_gateway.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Application Gateway performance including throughput and traffic volume, request rates and response status codes, backend health and latency breakdown (connect, first byte, last byte), client latency, current and new connections, WebSocket sessions, capacity and compute units, CPU utilization, TLS connections, and WAF security events including rule matches, challenges, and penalty box activity.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Application Gateway with metrics covering:
31
+
32
+- **Traffic** -- throughput, traffic volume (received/sent), request rates (total/failed)
33
+- **Response** -- gateway and backend response status codes
34
+- **Backend** -- backend health (healthy/unhealthy hosts), backend latency (connect, first byte, last byte)
35
+- **Client** -- client latency (total time, client RTT)
36
+- **Connections** -- current and new connections, TLS connections, WebSocket connections
37
+- **Capacity** -- capacity units, compute units, billed/fixed billed, CPU utilization
38
+- **WAF** -- WAF requests (total/blocked/matched), rule matches (managed/custom/bot), challenges, penalty box
39
+
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
+
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -317,6 +454,7 @@ Labels:
454
| region | The Azure region where the resource is deployed. |
455
| resource_type | The Azure resource type identifier. |
456
| profile | The Azure Monitor profile id. |
457
+| subscription_id | The Azure subscription identifier. |
458
| resource_uid | The unique Azure resource identifier. |
459
460
Metrics:
@@ -415,31 +553,46 @@ docker logs netdata 2>&1 | grep azure_monitor
553
554
### No metrics are collected
555
418
-Verify the following:
419
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
420
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
421
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
422
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
556
+Check the following:
557
+
558
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
559
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
560
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
561
+- **Collector logs** -- Check for authentication or API errors:
562
+ ```bash
563
+ # systemd
564
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
565
+ # non-systemd
566
+ grep azure_monitor /var/log/netdata/collector.log
567
+ ```
568
569
570
### Missing metrics for some resource types
571
427
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
428
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
429
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
430
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
572
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
573
+
574
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
575
+- **Verify a built-in profile exists** -- List available profiles:
576
+ ```bash
577
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
578
+ ```
579
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
580
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
581
582
433
-### Metrics appear delayed
583
+### Charts have gaps or incomplete data
584
435
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
436
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
437
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
585
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
586
+
587
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
588
+- Slower time-grain batches automatically use a larger effective offset when needed.
589
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
590
591
592
### Authentication errors in sovereign clouds
593
594
For Azure Government or Azure China clouds, set the `cloud` parameter:
595
+
596
- Azure Government: `cloud: government`
597
- Azure China (21Vianet): `cloud: china`
598
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_application_insights.md
+241
-87
@@ -21,40 +21,80 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor application performance through Application Insights including availability test results and duration, server request rates and response times, dependency call tracking and failures, exception rates by source, browser page load timing breakdown, process CPU and memory usage, IO rates, HTTP request queue depth, page views, and trace volume.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Application Insights with metrics covering:
31
+
32
+- **Availability** -- availability test percentage, test duration
33
+- **Requests** -- server request rate, HTTP request rate, HTTP request execution time, request queue depth
34
+- **Responses** -- server response time, server requests (total/failed)
35
+- **Dependencies** -- dependency calls (total/failed), dependency duration
36
+- **Exceptions** -- exception rate, exceptions by source (total/browser/server)
37
+- **Browser** -- page load time, browser timing breakdown (network/send/receive/processing), page views
38
+- **Process** -- CPU utilization (process/processor), memory (available/private), I/O rate
39
+- **Traces** -- trace volume
40
+
41
+
42
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
43
44
45
This collector is supported on all platforms.
46
47
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
48
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
49
+The service principal or managed identity requires these Azure RBAC roles:
50
+
51
+| Role | Purpose | Scope |
52
+|:-----|:--------|:------|
53
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
54
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
55
56
57
### Default Behavior
58
59
#### Auto-Detection
60
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
61
+The collector has two discovery phases:
62
+
63
+**Bootstrap (first run)**
64
+
65
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
66
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
67
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
68
+- A single job can monitor multiple subscriptions.
69
+
70
+**Runtime (periodic refresh)**
71
+
72
+- Periodically re-discovers resources for **already-active profile types only**.
73
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
74
+
75
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
76
77
78
#### Limits
79
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
80
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
81
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
82
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
83
84
85
#### Performance Impact
86
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
87
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
88
+
89
+**Default concurrency and batching limits:**
90
+
91
+| Setting | Default | Description |
92
+|:--------|:--------|:------------|
93
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
94
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
95
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
96
+
97
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
98
99
100
## Setup
@@ -78,25 +118,36 @@ UI configuration requires paid Netdata Cloud plan.
118
119
#### Create an Azure monitoring principal
120
81
-Create a service principal or use a managed identity with the following permissions:
121
+The collector requires a service principal or managed identity with two Azure RBAC roles:
122
+
123
+| Role | Purpose |
124
+|:-----|:--------|
125
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
126
+| **Reader** | Query Azure Resource Graph for resource discovery |
127
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
128
+**Option A: Service principal**
129
86
-For service principal authentication:
130
```bash
88
-# Create the service principal
131
+# Create service principal with Monitoring Reader role
132
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
133
--scopes /subscriptions/<subscription-id>
134
135
+# Add the Reader role for resource discovery
136
+az role assignment create --assignee <appId-from-above> \
137
+ --role "Reader" --scope /subscriptions/<subscription-id>
138
+
139
# Note the appId (client_id), password (client_secret), and tenant
140
```
141
95
-For managed identity (on Azure VMs, VMSS, or AKS):
142
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
143
+
144
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
145
+# Assign both roles to the VM's managed identity
146
az role assignment create --assignee <managed-identity-principal-id> \
147
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
148
+
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Reader" --scope /subscriptions/<subscription-id>
151
```
152
153
@@ -105,13 +156,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
156
157
#### Options
158
108
-The following options can be defined globally: update_every, autodetection_retry.
159
+The following options can be defined globally: `update_every`, `autodetection_retry`.
160
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
161
+**Profile file locations:**
162
114
-User profile files with the same filename override stock profiles.
163
+| Type | Path |
164
+|:-----|:-----|
165
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
166
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
167
+
168
+User profile files with the same `id` as a stock profile override it.
169
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
170
171
172
<details open><summary>Config options</summary>
@@ -122,25 +177,103 @@ User profile files with the same filename override stock profiles.
177
|:------|:-----|:------------|:--------|:---------:|
178
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
179
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
181
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
184
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
185
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
190
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
191
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
194
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
195
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
196
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
197
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
198
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
199
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
200
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
201
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
203
+<a id="option-collection-query-offset"></a>
204
+##### query_offset
205
+
206
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
207
+
208
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
209
+
210
+- **Default (180s)** works for most services.
211
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
212
+- **Increase to 240-300s** if you still see gaps or missing data points.
213
+- **Do not set below 60s** -- metrics will likely be incomplete.
214
+
215
+
216
+<a id="option-authentication-auth-mode"></a>
217
+##### auth.mode
218
+
219
+Determines how the collector authenticates with Azure.
220
+
221
+| Mode | When to use | Required options |
222
+|:-----|:------------|:-----------------|
223
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
224
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
225
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
226
+
227
+
228
+<a id="option-discovery-discovery-mode"></a>
229
+##### discovery.mode
230
+
231
+Controls how the collector finds candidate Azure resources.
232
+
233
+| Mode | Behavior |
234
+|:-----|:---------|
235
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
236
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
237
+
238
+
239
+<a id="option-discovery-discovery-mode-query-kql"></a>
240
+##### discovery.mode_query.kql
241
+
242
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
243
+
244
+The query **must** project these five columns:
245
+
246
+| Column | Description |
247
+|:-------|:------------|
248
+| `id` | Full Azure resource ID (ARM format) |
249
+| `name` | Resource name |
250
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
251
+| `resourceGroup` | Resource group name |
252
+| `location` | Azure region |
253
+
254
+Example:
255
+
256
+```
257
+resources
258
+| where tags.env =~ "prod"
259
+| project id, name, type, resourceGroup, location
260
+```
261
+
262
+
263
+<a id="option-profiles-profiles-mode"></a>
264
+##### profiles.mode
265
+
266
+Controls how the collector decides which metric profiles to activate.
267
+
268
+| Mode | Behavior |
269
+|:-----|:---------|
270
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
271
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
272
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
273
+
274
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
275
+
276
+
277
278
</details>
279
@@ -182,14 +315,28 @@ sudo ./edit-config go.d/azure_monitor.conf
315
316
##### Examples
317
185
-###### Service principal (auto-discover all resources)
318
+###### Service principal with structured discovery
319
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
320
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
321
322
```yaml
323
jobs:
324
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ subscription_ids:
326
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
327
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
328
+ discovery:
329
+ mode: filters
330
+ mode_filters:
331
+ resource_groups:
332
+ - production-rg
333
+ regions:
334
+ - eastus
335
+ tags:
336
+ env:
337
+ - prod
338
+ profiles:
339
+ mode: auto
340
auth:
341
mode: service_principal
342
mode_service_principal:
@@ -198,59 +345,49 @@ jobs:
345
client_secret: "your-client-secret"
346
347
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
348
+###### Managed identity with exact profiles
349
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
350
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
351
352
<details open><summary>Config</summary>
353
354
```yaml
355
jobs:
356
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ subscription_ids:
358
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
360
+ mode: exact
361
+ mode_exact:
362
+ names:
363
+ - sql_database
364
+ - postgres_flexible
365
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
366
+ mode: managed_identity
367
368
```
369
</details>
370
241
-###### Filter by resource group
371
+###### Custom Azure Resource Graph KQL
372
243
-Only monitor resources in specific resource groups.
373
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
374
375
<details open><summary>Config</summary>
376
377
```yaml
378
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
379
+ - name: prod-query
380
+ subscription_ids:
381
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
382
+ discovery:
383
+ mode: query
384
+ mode_query:
385
+ kql: |
386
+ resources
387
+ | where tags.env =~ "prod"
388
+ | project id, name, type, resourceGroup, location
389
+ profiles:
390
+ mode: auto
391
auth:
392
mode: default
393
@@ -259,14 +396,15 @@ jobs:
396
397
###### Azure Government cloud
398
262
-Connect to Azure Government cloud environment.
399
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
400
401
<details open><summary>Config</summary>
402
403
```yaml
404
jobs:
405
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
+ subscription_ids:
407
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
cloud: government
409
auth:
410
mode: service_principal
@@ -324,6 +462,7 @@ Labels:
462
| region | The Azure region where the resource is deployed. |
463
| resource_type | The Azure resource type identifier. |
464
| profile | The Azure Monitor profile id. |
465
+| subscription_id | The Azure subscription identifier. |
466
| resource_uid | The unique Azure resource identifier. |
467
468
Metrics:
@@ -423,31 +562,46 @@ docker logs netdata 2>&1 | grep azure_monitor
562
563
### No metrics are collected
564
426
-Verify the following:
427
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
428
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
429
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
430
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
565
+Check the following:
566
+
567
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
568
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
569
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
570
+- **Collector logs** -- Check for authentication or API errors:
571
+ ```bash
572
+ # systemd
573
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
574
+ # non-systemd
575
+ grep azure_monitor /var/log/netdata/collector.log
576
+ ```
577
578
579
### Missing metrics for some resource types
580
435
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
436
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
437
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
438
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
581
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
582
+
583
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
584
+- **Verify a built-in profile exists** -- List available profiles:
585
+ ```bash
586
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
587
+ ```
588
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
589
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
590
591
441
-### Metrics appear delayed
592
+### Charts have gaps or incomplete data
593
443
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
444
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
445
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
594
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
595
+
596
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
597
+- Slower time-grain batches automatically use a larger effective offset when needed.
598
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
599
600
601
### Authentication errors in sovereign clouds
602
603
For Azure Government or Azure China clouds, set the `cloud` parameter:
604
+
605
- Azure Government: `cloud: government`
606
- Azure China (21Vianet): `cloud: china`
607
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cache_for_redis.md
+242
-87
@@ -21,40 +21,81 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Cache for Redis including cache hit and miss rates, read and write throughput, server load and CPU utilization, memory usage, connected clients, operations per second, command processing rates, latency percentiles, key eviction and expiration, and geo-replication health and sync status. Provides per-shard breakdowns for clustered deployments.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Cache for Redis with metrics covering:
31
+
32
+- **Performance** -- operations/second, command processing rates (get/set), cache hit/miss rates
33
+- **Latency** -- average latency, P99 latency
34
+- **Compute** -- CPU utilization, server load
35
+- **Memory** -- memory usage (used/RSS), memory utilization
36
+- **Connections** -- connected clients, connection rate (created/closed)
37
+- **Keys** -- total keys, evicted keys, expired keys, miss rate
38
+- **Throughput** -- read/write bytes per second
39
+- **Geo-replication** -- replication health, connectivity lag, sync events, data sync offset
40
+- **Per-shard** -- instance-level breakdowns for hit rate, clients, commands, server load, keys, operations, throughput
41
+
42
+
43
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
44
45
46
This collector is supported on all platforms.
47
48
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
49
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
50
+The service principal or managed identity requires these Azure RBAC roles:
51
+
52
+| Role | Purpose | Scope |
53
+|:-----|:--------|:------|
54
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
55
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
56
57
58
### Default Behavior
59
60
#### Auto-Detection
61
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
62
+The collector has two discovery phases:
63
+
64
+**Bootstrap (first run)**
65
+
66
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
67
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
68
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
69
+- A single job can monitor multiple subscriptions.
70
+
71
+**Runtime (periodic refresh)**
72
+
73
+- Periodically re-discovers resources for **already-active profile types only**.
74
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
75
+
76
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
77
78
79
#### Limits
80
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
81
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
82
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
83
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
84
85
86
#### Performance Impact
87
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
88
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
89
+
90
+**Default concurrency and batching limits:**
91
+
92
+| Setting | Default | Description |
93
+|:--------|:--------|:------------|
94
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
95
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
96
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
97
+
98
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
99
100
101
## Setup
@@ -78,25 +119,36 @@ UI configuration requires paid Netdata Cloud plan.
119
120
#### Create an Azure monitoring principal
121
81
-Create a service principal or use a managed identity with the following permissions:
122
+The collector requires a service principal or managed identity with two Azure RBAC roles:
123
+
124
+| Role | Purpose |
125
+|:-----|:--------|
126
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
127
+| **Reader** | Query Azure Resource Graph for resource discovery |
128
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
129
+**Option A: Service principal**
130
86
-For service principal authentication:
131
```bash
88
-# Create the service principal
132
+# Create service principal with Monitoring Reader role
133
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
134
--scopes /subscriptions/<subscription-id>
135
136
+# Add the Reader role for resource discovery
137
+az role assignment create --assignee <appId-from-above> \
138
+ --role "Reader" --scope /subscriptions/<subscription-id>
139
+
140
# Note the appId (client_id), password (client_secret), and tenant
141
```
142
95
-For managed identity (on Azure VMs, VMSS, or AKS):
143
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
144
+
145
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
146
+# Assign both roles to the VM's managed identity
147
az role assignment create --assignee <managed-identity-principal-id> \
148
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Reader" --scope /subscriptions/<subscription-id>
152
```
153
154
@@ -105,13 +157,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
157
158
#### Options
159
108
-The following options can be defined globally: update_every, autodetection_retry.
160
+The following options can be defined globally: `update_every`, `autodetection_retry`.
161
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+**Profile file locations:**
163
114
-User profile files with the same filename override stock profiles.
164
+| Type | Path |
165
+|:-----|:-----|
166
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
167
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
168
+
169
+User profile files with the same `id` as a stock profile override it.
170
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
171
172
173
<details open><summary>Config options</summary>
@@ -122,25 +178,103 @@ User profile files with the same filename override stock profiles.
178
|:------|:-----|:------------|:--------|:---------:|
179
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
180
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
181
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
182
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
186
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
191
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
192
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
195
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
196
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
197
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
198
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
199
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
204
+<a id="option-collection-query-offset"></a>
205
+##### query_offset
206
+
207
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
208
+
209
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
210
+
211
+- **Default (180s)** works for most services.
212
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
213
+- **Increase to 240-300s** if you still see gaps or missing data points.
214
+- **Do not set below 60s** -- metrics will likely be incomplete.
215
+
216
+
217
+<a id="option-authentication-auth-mode"></a>
218
+##### auth.mode
219
+
220
+Determines how the collector authenticates with Azure.
221
+
222
+| Mode | When to use | Required options |
223
+|:-----|:------------|:-----------------|
224
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
225
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
226
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
227
+
228
+
229
+<a id="option-discovery-discovery-mode"></a>
230
+##### discovery.mode
231
+
232
+Controls how the collector finds candidate Azure resources.
233
+
234
+| Mode | Behavior |
235
+|:-----|:---------|
236
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
237
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
238
+
239
+
240
+<a id="option-discovery-discovery-mode-query-kql"></a>
241
+##### discovery.mode_query.kql
242
+
243
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
244
+
245
+The query **must** project these five columns:
246
+
247
+| Column | Description |
248
+|:-------|:------------|
249
+| `id` | Full Azure resource ID (ARM format) |
250
+| `name` | Resource name |
251
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
252
+| `resourceGroup` | Resource group name |
253
+| `location` | Azure region |
254
+
255
+Example:
256
+
257
+```
258
+resources
259
+| where tags.env =~ "prod"
260
+| project id, name, type, resourceGroup, location
261
+```
262
+
263
+
264
+<a id="option-profiles-profiles-mode"></a>
265
+##### profiles.mode
266
+
267
+Controls how the collector decides which metric profiles to activate.
268
+
269
+| Mode | Behavior |
270
+|:-----|:---------|
271
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
272
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
273
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
274
+
275
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
276
+
277
+
278
279
</details>
280
@@ -182,14 +316,28 @@ sudo ./edit-config go.d/azure_monitor.conf
316
317
##### Examples
318
185
-###### Service principal (auto-discover all resources)
319
+###### Service principal with structured discovery
320
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
321
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
322
323
```yaml
324
jobs:
325
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ subscription_ids:
327
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
328
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
329
+ discovery:
330
+ mode: filters
331
+ mode_filters:
332
+ resource_groups:
333
+ - production-rg
334
+ regions:
335
+ - eastus
336
+ tags:
337
+ env:
338
+ - prod
339
+ profiles:
340
+ mode: auto
341
auth:
342
mode: service_principal
343
mode_service_principal:
@@ -198,59 +346,49 @@ jobs:
346
client_secret: "your-client-secret"
347
348
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
349
+###### Managed identity with exact profiles
350
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
351
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
352
353
<details open><summary>Config</summary>
354
355
```yaml
356
jobs:
357
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ subscription_ids:
359
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
361
+ mode: exact
362
+ mode_exact:
363
+ names:
364
+ - sql_database
365
+ - postgres_flexible
366
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
367
+ mode: managed_identity
368
369
```
370
</details>
371
241
-###### Filter by resource group
372
+###### Custom Azure Resource Graph KQL
373
243
-Only monitor resources in specific resource groups.
374
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
375
376
<details open><summary>Config</summary>
377
378
```yaml
379
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
380
+ - name: prod-query
381
+ subscription_ids:
382
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
383
+ discovery:
384
+ mode: query
385
+ mode_query:
386
+ kql: |
387
+ resources
388
+ | where tags.env =~ "prod"
389
+ | project id, name, type, resourceGroup, location
390
+ profiles:
391
+ mode: auto
392
auth:
393
mode: default
394
@@ -259,14 +397,15 @@ jobs:
397
398
###### Azure Government cloud
399
262
-Connect to Azure Government cloud environment.
400
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
401
402
<details open><summary>Config</summary>
403
404
```yaml
405
jobs:
406
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
+ subscription_ids:
408
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
409
cloud: government
410
auth:
411
mode: service_principal
@@ -321,6 +460,7 @@ Labels:
460
| region | The Azure region where the resource is deployed. |
461
| resource_type | The Azure resource type identifier. |
462
| profile | The Azure Monitor profile id. |
463
+| subscription_id | The Azure subscription identifier. |
464
| resource_uid | The unique Azure resource identifier. |
465
466
Metrics:
@@ -431,31 +571,46 @@ docker logs netdata 2>&1 | grep azure_monitor
571
572
### No metrics are collected
573
434
-Verify the following:
435
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
436
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
437
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
438
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
574
+Check the following:
575
+
576
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
577
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
578
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
579
+- **Collector logs** -- Check for authentication or API errors:
580
+ ```bash
581
+ # systemd
582
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
583
+ # non-systemd
584
+ grep azure_monitor /var/log/netdata/collector.log
585
+ ```
586
587
588
### Missing metrics for some resource types
589
443
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
444
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
445
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
446
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
590
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
591
+
592
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
593
+- **Verify a built-in profile exists** -- List available profiles:
594
+ ```bash
595
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
596
+ ```
597
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
598
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
599
600
449
-### Metrics appear delayed
601
+### Charts have gaps or incomplete data
602
451
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
452
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
453
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
603
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
604
+
605
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
606
+- Slower time-grain batches automatically use a larger effective offset when needed.
607
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
608
609
610
### Authentication errors in sovereign clouds
611
612
For Azure Government or Azure China clouds, set the `cloud` parameter:
613
+
614
- Azure Government: `cloud: government`
615
- Azure China (21Vianet): `cloud: china`
616
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cognitive_services.md
+245
-87
@@ -21,40 +21,84 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure AI and Cognitive Services including API call volume, success and client error rates, response latency, token processing rates for language models, content safety filtering, fine-tuning operations, provisioned throughput utilization, rate-limiting events, active inference connections, and context token cache performance.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure AI and Cognitive Services with metrics covering:
31
+
32
+- **API calls** -- total calls (successful/blocked/token), model requests, OpenAI requests
33
+- **Errors** -- total, client, and server errors, rate limiting events
34
+- **Latency** -- service latency, model latency (time to response/first token/between tokens/last byte)
35
+- **Tokens** -- model token usage (input/output/total), OpenAI token usage (prompt/generated), cache tokens (read/write)
36
+- **Availability** -- service availability, model availability, OpenAI availability
37
+- **Content safety** -- content moderation calls (text/image), safety system events, harmful/blocked requests
38
+- **Speech** -- transcription, translation, synthesis, speaker recognition, voice training/hosting
39
+- **Vision** -- computer vision and custom vision transactions, images stored, training time
40
+- **Translator** -- text and document translation (standard/custom)
41
+- **Provisioned** -- model provisioned utilization, OpenAI provisioned-managed utilization
42
+- **Fine-tuning** -- training hours
43
+- **Personalizer** -- events, rewards, actions, feature cardinality
44
+
45
+
46
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
47
48
49
This collector is supported on all platforms.
50
51
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
52
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
53
+The service principal or managed identity requires these Azure RBAC roles:
54
+
55
+| Role | Purpose | Scope |
56
+|:-----|:--------|:------|
57
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
58
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
59
60
61
### Default Behavior
62
63
#### Auto-Detection
64
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
65
+The collector has two discovery phases:
66
+
67
+**Bootstrap (first run)**
68
+
69
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
70
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
71
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
72
+- A single job can monitor multiple subscriptions.
73
+
74
+**Runtime (periodic refresh)**
75
+
76
+- Periodically re-discovers resources for **already-active profile types only**.
77
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
78
+
79
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
80
81
82
#### Limits
83
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
84
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
85
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
86
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
87
88
89
#### Performance Impact
90
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
91
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
92
+
93
+**Default concurrency and batching limits:**
94
+
95
+| Setting | Default | Description |
96
+|:--------|:--------|:------------|
97
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
98
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
99
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
100
+
101
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
102
103
104
## Setup
@@ -78,25 +122,36 @@ UI configuration requires paid Netdata Cloud plan.
122
123
#### Create an Azure monitoring principal
124
81
-Create a service principal or use a managed identity with the following permissions:
125
+The collector requires a service principal or managed identity with two Azure RBAC roles:
126
+
127
+| Role | Purpose |
128
+|:-----|:--------|
129
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
130
+| **Reader** | Query Azure Resource Graph for resource discovery |
131
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
132
+**Option A: Service principal**
133
86
-For service principal authentication:
134
```bash
88
-# Create the service principal
135
+# Create service principal with Monitoring Reader role
136
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
137
--scopes /subscriptions/<subscription-id>
138
139
+# Add the Reader role for resource discovery
140
+az role assignment create --assignee <appId-from-above> \
141
+ --role "Reader" --scope /subscriptions/<subscription-id>
142
+
143
# Note the appId (client_id), password (client_secret), and tenant
144
```
145
95
-For managed identity (on Azure VMs, VMSS, or AKS):
146
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
147
+
148
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
149
+# Assign both roles to the VM's managed identity
150
az role assignment create --assignee <managed-identity-principal-id> \
151
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
152
+
153
+az role assignment create --assignee <managed-identity-principal-id> \
154
+ --role "Reader" --scope /subscriptions/<subscription-id>
155
```
156
157
@@ -105,13 +160,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
160
161
#### Options
162
108
-The following options can be defined globally: update_every, autodetection_retry.
163
+The following options can be defined globally: `update_every`, `autodetection_retry`.
164
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
165
+**Profile file locations:**
166
114
-User profile files with the same filename override stock profiles.
167
+| Type | Path |
168
+|:-----|:-----|
169
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
170
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
171
+
172
+User profile files with the same `id` as a stock profile override it.
173
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
174
175
176
<details open><summary>Config options</summary>
@@ -122,25 +181,103 @@ User profile files with the same filename override stock profiles.
181
|:------|:-----|:------------|:--------|:---------:|
182
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
183
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
184
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
185
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
186
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
187
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
188
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
189
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
190
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
191
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
192
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
193
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
194
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
195
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
196
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
197
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
198
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
199
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
200
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
201
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
202
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
203
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
204
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
205
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
206
207
+<a id="option-collection-query-offset"></a>
208
+##### query_offset
209
+
210
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
211
+
212
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
213
+
214
+- **Default (180s)** works for most services.
215
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
216
+- **Increase to 240-300s** if you still see gaps or missing data points.
217
+- **Do not set below 60s** -- metrics will likely be incomplete.
218
+
219
+
220
+<a id="option-authentication-auth-mode"></a>
221
+##### auth.mode
222
+
223
+Determines how the collector authenticates with Azure.
224
+
225
+| Mode | When to use | Required options |
226
+|:-----|:------------|:-----------------|
227
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
228
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
229
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
230
+
231
+
232
+<a id="option-discovery-discovery-mode"></a>
233
+##### discovery.mode
234
+
235
+Controls how the collector finds candidate Azure resources.
236
+
237
+| Mode | Behavior |
238
+|:-----|:---------|
239
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
240
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
241
+
242
+
243
+<a id="option-discovery-discovery-mode-query-kql"></a>
244
+##### discovery.mode_query.kql
245
+
246
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
247
+
248
+The query **must** project these five columns:
249
+
250
+| Column | Description |
251
+|:-------|:------------|
252
+| `id` | Full Azure resource ID (ARM format) |
253
+| `name` | Resource name |
254
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
255
+| `resourceGroup` | Resource group name |
256
+| `location` | Azure region |
257
+
258
+Example:
259
+
260
+```
261
+resources
262
+| where tags.env =~ "prod"
263
+| project id, name, type, resourceGroup, location
264
+```
265
+
266
+
267
+<a id="option-profiles-profiles-mode"></a>
268
+##### profiles.mode
269
+
270
+Controls how the collector decides which metric profiles to activate.
271
+
272
+| Mode | Behavior |
273
+|:-----|:---------|
274
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
275
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
276
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
277
+
278
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
279
+
280
+
281
282
</details>
283
@@ -182,14 +319,28 @@ sudo ./edit-config go.d/azure_monitor.conf
319
320
##### Examples
321
185
-###### Service principal (auto-discover all resources)
322
+###### Service principal with structured discovery
323
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
324
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
325
326
```yaml
327
jobs:
328
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
329
+ subscription_ids:
330
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
331
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
332
+ discovery:
333
+ mode: filters
334
+ mode_filters:
335
+ resource_groups:
336
+ - production-rg
337
+ regions:
338
+ - eastus
339
+ tags:
340
+ env:
341
+ - prod
342
+ profiles:
343
+ mode: auto
344
auth:
345
mode: service_principal
346
mode_service_principal:
@@ -198,59 +349,49 @@ jobs:
349
client_secret: "your-client-secret"
350
351
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
352
+###### Managed identity with exact profiles
353
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
354
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
355
356
<details open><summary>Config</summary>
357
358
```yaml
359
jobs:
360
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
361
+ subscription_ids:
362
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
363
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
364
+ mode: exact
365
+ mode_exact:
366
+ names:
367
+ - sql_database
368
+ - postgres_flexible
369
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
370
+ mode: managed_identity
371
372
```
373
</details>
374
241
-###### Filter by resource group
375
+###### Custom Azure Resource Graph KQL
376
243
-Only monitor resources in specific resource groups.
377
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
378
379
<details open><summary>Config</summary>
380
381
```yaml
382
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
383
+ - name: prod-query
384
+ subscription_ids:
385
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
386
+ discovery:
387
+ mode: query
388
+ mode_query:
389
+ kql: |
390
+ resources
391
+ | where tags.env =~ "prod"
392
+ | project id, name, type, resourceGroup, location
393
+ profiles:
394
+ mode: auto
395
auth:
396
mode: default
397
@@ -259,14 +400,15 @@ jobs:
400
401
###### Azure Government cloud
402
262
-Connect to Azure Government cloud environment.
403
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
404
405
<details open><summary>Config</summary>
406
407
```yaml
408
jobs:
409
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
410
+ subscription_ids:
411
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
412
cloud: government
413
auth:
414
mode: service_principal
@@ -327,6 +469,7 @@ Labels:
469
| region | The Azure region where the resource is deployed. |
470
| resource_type | The Azure resource type identifier. |
471
| profile | The Azure Monitor profile id. |
472
+| subscription_id | The Azure subscription identifier. |
473
| resource_uid | The unique Azure resource identifier. |
474
475
Metrics:
@@ -478,31 +621,46 @@ docker logs netdata 2>&1 | grep azure_monitor
621
622
### No metrics are collected
623
481
-Verify the following:
482
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
483
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
484
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
485
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
624
+Check the following:
625
+
626
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
627
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
628
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
629
+- **Collector logs** -- Check for authentication or API errors:
630
+ ```bash
631
+ # systemd
632
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
633
+ # non-systemd
634
+ grep azure_monitor /var/log/netdata/collector.log
635
+ ```
636
637
638
### Missing metrics for some resource types
639
490
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
491
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
492
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
493
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
640
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
641
+
642
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
643
+- **Verify a built-in profile exists** -- List available profiles:
644
+ ```bash
645
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
646
+ ```
647
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
648
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
649
650
496
-### Metrics appear delayed
651
+### Charts have gaps or incomplete data
652
498
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
499
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
500
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
653
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
654
+
655
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
656
+- Slower time-grain batches automatically use a larger effective offset when needed.
657
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
658
659
660
### Authentication errors in sovereign clouds
661
662
For Azure Government or Azure China clouds, set the `cloud` parameter:
663
+
664
- Azure Government: `cloud: government`
665
- Azure China (21Vianet): `cloud: china`
666
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_apps.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Container Apps including CPU and memory usage, network traffic, replica counts, request processing rates, response times, restart frequency, and resource reservation utilization.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Container Apps with metrics covering:
31
+
32
+- **Compute** -- CPU usage (nanocores and percentage), GPU utilization
33
+- **Memory** -- memory working set, memory percentage, JVM memory (total/pool/buffer)
34
+- **Requests** -- request rate, response time
35
+- **Network** -- network traffic (received/sent), resiliency pending connections and timeouts
36
+- **Replicas** -- replica count, restart count, reserved cores
37
+- **JVM** -- thread count, GC collections and duration, buffer count
38
+- **Resiliency** -- host ejections, request retries
39
+
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
+
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -321,6 +458,7 @@ Labels:
458
| region | The Azure region where the resource is deployed. |
459
| resource_type | The Azure resource type identifier. |
460
| profile | The Azure Monitor profile id. |
461
+| subscription_id | The Azure subscription identifier. |
462
| resource_uid | The unique Azure resource identifier. |
463
464
Metrics:
@@ -421,31 +559,46 @@ docker logs netdata 2>&1 | grep azure_monitor
559
560
### No metrics are collected
561
424
-Verify the following:
425
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
426
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
427
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
428
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
562
+Check the following:
563
+
564
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
565
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
566
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
567
+- **Collector logs** -- Check for authentication or API errors:
568
+ ```bash
569
+ # systemd
570
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
571
+ # non-systemd
572
+ grep azure_monitor /var/log/netdata/collector.log
573
+ ```
574
575
576
### Missing metrics for some resource types
577
433
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
434
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
435
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
436
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
578
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
579
+
580
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
581
+- **Verify a built-in profile exists** -- List available profiles:
582
+ ```bash
583
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
584
+ ```
585
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
586
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
587
588
439
-### Metrics appear delayed
589
+### Charts have gaps or incomplete data
590
441
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
442
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
443
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
591
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
592
+
593
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
594
+- Slower time-grain batches automatically use a larger effective offset when needed.
595
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
596
597
598
### Authentication errors in sovereign clouds
599
600
For Azure Government or Azure China clouds, set the `cloud` parameter:
601
+
602
- Azure Government: `cloud: government`
603
- Azure China (21Vianet): `cloud: china`
604
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_instances.md
+236
-87
@@ -21,40 +21,75 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Container Instance groups including CPU and memory usage and network bytes transferred in and out.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Container Instances with metrics covering:
31
+
32
+- **Compute** -- CPU usage (average/max)
33
+- **Memory** -- memory usage (average/max)
34
+- **Network** -- network traffic (received/sent)
35
+
36
+
37
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
38
39
40
This collector is supported on all platforms.
41
42
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
43
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
44
+The service principal or managed identity requires these Azure RBAC roles:
45
+
46
+| Role | Purpose | Scope |
47
+|:-----|:--------|:------|
48
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
49
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
50
51
52
### Default Behavior
53
54
#### Auto-Detection
55
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
56
+The collector has two discovery phases:
57
+
58
+**Bootstrap (first run)**
59
+
60
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
61
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
62
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
63
+- A single job can monitor multiple subscriptions.
64
+
65
+**Runtime (periodic refresh)**
66
+
67
+- Periodically re-discovers resources for **already-active profile types only**.
68
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
69
+
70
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
71
72
73
#### Limits
74
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
75
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
76
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
77
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
78
79
80
#### Performance Impact
81
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
82
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
83
+
84
+**Default concurrency and batching limits:**
85
+
86
+| Setting | Default | Description |
87
+|:--------|:--------|:------------|
88
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
89
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
90
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
91
+
92
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
93
94
95
## Setup
@@ -78,25 +113,36 @@ UI configuration requires paid Netdata Cloud plan.
113
114
#### Create an Azure monitoring principal
115
81
-Create a service principal or use a managed identity with the following permissions:
116
+The collector requires a service principal or managed identity with two Azure RBAC roles:
117
+
118
+| Role | Purpose |
119
+|:-----|:--------|
120
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
121
+| **Reader** | Query Azure Resource Graph for resource discovery |
122
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
123
+**Option A: Service principal**
124
86
-For service principal authentication:
125
```bash
88
-# Create the service principal
126
+# Create service principal with Monitoring Reader role
127
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
128
--scopes /subscriptions/<subscription-id>
129
130
+# Add the Reader role for resource discovery
131
+az role assignment create --assignee <appId-from-above> \
132
+ --role "Reader" --scope /subscriptions/<subscription-id>
133
+
134
# Note the appId (client_id), password (client_secret), and tenant
135
```
136
95
-For managed identity (on Azure VMs, VMSS, or AKS):
137
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
138
+
139
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
140
+# Assign both roles to the VM's managed identity
141
az role assignment create --assignee <managed-identity-principal-id> \
142
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
143
+
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Reader" --scope /subscriptions/<subscription-id>
146
```
147
148
@@ -105,13 +151,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
151
152
#### Options
153
108
-The following options can be defined globally: update_every, autodetection_retry.
154
+The following options can be defined globally: `update_every`, `autodetection_retry`.
155
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
156
+**Profile file locations:**
157
114
-User profile files with the same filename override stock profiles.
158
+| Type | Path |
159
+|:-----|:-----|
160
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
161
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
162
+
163
+User profile files with the same `id` as a stock profile override it.
164
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
165
166
167
<details open><summary>Config options</summary>
@@ -122,25 +172,103 @@ User profile files with the same filename override stock profiles.
172
|:------|:-----|:------------|:--------|:---------:|
173
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
174
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
175
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
176
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
177
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
179
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
180
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
181
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
183
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
184
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
185
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
186
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
187
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
189
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
190
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
191
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
192
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
193
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
194
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
195
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
196
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
197
198
+<a id="option-collection-query-offset"></a>
199
+##### query_offset
200
+
201
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
202
+
203
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
204
+
205
+- **Default (180s)** works for most services.
206
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
207
+- **Increase to 240-300s** if you still see gaps or missing data points.
208
+- **Do not set below 60s** -- metrics will likely be incomplete.
209
+
210
+
211
+<a id="option-authentication-auth-mode"></a>
212
+##### auth.mode
213
+
214
+Determines how the collector authenticates with Azure.
215
+
216
+| Mode | When to use | Required options |
217
+|:-----|:------------|:-----------------|
218
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
219
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
220
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
221
+
222
+
223
+<a id="option-discovery-discovery-mode"></a>
224
+##### discovery.mode
225
+
226
+Controls how the collector finds candidate Azure resources.
227
+
228
+| Mode | Behavior |
229
+|:-----|:---------|
230
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
231
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
232
+
233
+
234
+<a id="option-discovery-discovery-mode-query-kql"></a>
235
+##### discovery.mode_query.kql
236
+
237
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
238
+
239
+The query **must** project these five columns:
240
+
241
+| Column | Description |
242
+|:-------|:------------|
243
+| `id` | Full Azure resource ID (ARM format) |
244
+| `name` | Resource name |
245
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
246
+| `resourceGroup` | Resource group name |
247
+| `location` | Azure region |
248
+
249
+Example:
250
+
251
+```
252
+resources
253
+| where tags.env =~ "prod"
254
+| project id, name, type, resourceGroup, location
255
+```
256
+
257
+
258
+<a id="option-profiles-profiles-mode"></a>
259
+##### profiles.mode
260
+
261
+Controls how the collector decides which metric profiles to activate.
262
+
263
+| Mode | Behavior |
264
+|:-----|:---------|
265
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
266
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
267
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
268
+
269
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
270
+
271
+
272
273
</details>
274
@@ -182,14 +310,28 @@ sudo ./edit-config go.d/azure_monitor.conf
310
311
##### Examples
312
185
-###### Service principal (auto-discover all resources)
313
+###### Service principal with structured discovery
314
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
315
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
316
317
```yaml
318
jobs:
319
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ subscription_ids:
321
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
322
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
323
+ discovery:
324
+ mode: filters
325
+ mode_filters:
326
+ resource_groups:
327
+ - production-rg
328
+ regions:
329
+ - eastus
330
+ tags:
331
+ env:
332
+ - prod
333
+ profiles:
334
+ mode: auto
335
auth:
336
mode: service_principal
337
mode_service_principal:
@@ -198,59 +340,49 @@ jobs:
340
client_secret: "your-client-secret"
341
342
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
343
+###### Managed identity with exact profiles
344
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
345
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
346
347
<details open><summary>Config</summary>
348
349
```yaml
350
jobs:
351
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
352
+ subscription_ids:
353
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
355
+ mode: exact
356
+ mode_exact:
357
+ names:
358
+ - sql_database
359
+ - postgres_flexible
360
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
361
+ mode: managed_identity
362
363
```
364
</details>
365
241
-###### Filter by resource group
366
+###### Custom Azure Resource Graph KQL
367
243
-Only monitor resources in specific resource groups.
368
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
369
370
<details open><summary>Config</summary>
371
372
```yaml
373
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
374
+ - name: prod-query
375
+ subscription_ids:
376
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
377
+ discovery:
378
+ mode: query
379
+ mode_query:
380
+ kql: |
381
+ resources
382
+ | where tags.env =~ "prod"
383
+ | project id, name, type, resourceGroup, location
384
+ profiles:
385
+ mode: auto
386
auth:
387
mode: default
388
@@ -259,14 +391,15 @@ jobs:
391
392
###### Azure Government cloud
393
262
-Connect to Azure Government cloud environment.
394
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
395
396
<details open><summary>Config</summary>
397
398
```yaml
399
jobs:
400
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
401
+ subscription_ids:
402
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
cloud: government
404
auth:
405
mode: service_principal
@@ -314,6 +447,7 @@ Labels:
447
| region | The Azure region where the resource is deployed. |
448
| resource_type | The Azure resource type identifier. |
449
| profile | The Azure Monitor profile id. |
450
+| subscription_id | The Azure subscription identifier. |
451
| resource_uid | The unique Azure resource identifier. |
452
453
Metrics:
@@ -395,31 +529,46 @@ docker logs netdata 2>&1 | grep azure_monitor
529
530
### No metrics are collected
531
398
-Verify the following:
399
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
400
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
401
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
402
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
532
+Check the following:
533
+
534
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
535
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
536
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
537
+- **Collector logs** -- Check for authentication or API errors:
538
+ ```bash
539
+ # systemd
540
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
541
+ # non-systemd
542
+ grep azure_monitor /var/log/netdata/collector.log
543
+ ```
544
545
546
### Missing metrics for some resource types
547
407
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
408
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
409
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
410
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
548
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
549
+
550
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
551
+- **Verify a built-in profile exists** -- List available profiles:
552
+ ```bash
553
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
554
+ ```
555
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
556
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
557
558
413
-### Metrics appear delayed
559
+### Charts have gaps or incomplete data
560
415
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
416
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
417
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
561
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
562
+
563
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
564
+- Slower time-grain batches automatically use a larger effective offset when needed.
565
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
566
567
568
### Authentication errors in sovereign clouds
569
570
For Azure Government or Azure China clouds, set the `cloud` parameter:
571
+
572
- Azure Government: `cloud: government`
573
- Azure China (21Vianet): `cloud: china`
574
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_container_registry.md
+236
-87
@@ -21,40 +21,75 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Container Registry including storage usage, successful and failed pull and push operation counts, and task run duration.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Container Registry with metrics covering:
31
+
32
+- **Operations** -- image pulls (successful/total), image pushes (successful/total)
33
+- **Storage** -- storage used
34
+- **Tasks** -- task run duration, agent pool CPU time
35
+
36
+
37
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
38
39
40
This collector is supported on all platforms.
41
42
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
43
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
44
+The service principal or managed identity requires these Azure RBAC roles:
45
+
46
+| Role | Purpose | Scope |
47
+|:-----|:--------|:------|
48
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
49
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
50
51
52
### Default Behavior
53
54
#### Auto-Detection
55
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
56
+The collector has two discovery phases:
57
+
58
+**Bootstrap (first run)**
59
+
60
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
61
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
62
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
63
+- A single job can monitor multiple subscriptions.
64
+
65
+**Runtime (periodic refresh)**
66
+
67
+- Periodically re-discovers resources for **already-active profile types only**.
68
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
69
+
70
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
71
72
73
#### Limits
74
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
75
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
76
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
77
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
78
79
80
#### Performance Impact
81
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
82
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
83
+
84
+**Default concurrency and batching limits:**
85
+
86
+| Setting | Default | Description |
87
+|:--------|:--------|:------------|
88
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
89
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
90
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
91
+
92
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
93
94
95
## Setup
@@ -78,25 +113,36 @@ UI configuration requires paid Netdata Cloud plan.
113
114
#### Create an Azure monitoring principal
115
81
-Create a service principal or use a managed identity with the following permissions:
116
+The collector requires a service principal or managed identity with two Azure RBAC roles:
117
+
118
+| Role | Purpose |
119
+|:-----|:--------|
120
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
121
+| **Reader** | Query Azure Resource Graph for resource discovery |
122
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
123
+**Option A: Service principal**
124
86
-For service principal authentication:
125
```bash
88
-# Create the service principal
126
+# Create service principal with Monitoring Reader role
127
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
128
--scopes /subscriptions/<subscription-id>
129
130
+# Add the Reader role for resource discovery
131
+az role assignment create --assignee <appId-from-above> \
132
+ --role "Reader" --scope /subscriptions/<subscription-id>
133
+
134
# Note the appId (client_id), password (client_secret), and tenant
135
```
136
95
-For managed identity (on Azure VMs, VMSS, or AKS):
137
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
138
+
139
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
140
+# Assign both roles to the VM's managed identity
141
az role assignment create --assignee <managed-identity-principal-id> \
142
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
143
+
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Reader" --scope /subscriptions/<subscription-id>
146
```
147
148
@@ -105,13 +151,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
151
152
#### Options
153
108
-The following options can be defined globally: update_every, autodetection_retry.
154
+The following options can be defined globally: `update_every`, `autodetection_retry`.
155
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
156
+**Profile file locations:**
157
114
-User profile files with the same filename override stock profiles.
158
+| Type | Path |
159
+|:-----|:-----|
160
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
161
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
162
+
163
+User profile files with the same `id` as a stock profile override it.
164
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
165
166
167
<details open><summary>Config options</summary>
@@ -122,25 +172,103 @@ User profile files with the same filename override stock profiles.
172
|:------|:-----|:------------|:--------|:---------:|
173
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
174
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
175
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
176
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
177
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
179
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
180
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
181
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
183
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
184
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
185
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
186
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
187
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
189
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
190
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
191
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
192
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
193
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
194
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
195
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
196
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
197
198
+<a id="option-collection-query-offset"></a>
199
+##### query_offset
200
+
201
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
202
+
203
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
204
+
205
+- **Default (180s)** works for most services.
206
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
207
+- **Increase to 240-300s** if you still see gaps or missing data points.
208
+- **Do not set below 60s** -- metrics will likely be incomplete.
209
+
210
+
211
+<a id="option-authentication-auth-mode"></a>
212
+##### auth.mode
213
+
214
+Determines how the collector authenticates with Azure.
215
+
216
+| Mode | When to use | Required options |
217
+|:-----|:------------|:-----------------|
218
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
219
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
220
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
221
+
222
+
223
+<a id="option-discovery-discovery-mode"></a>
224
+##### discovery.mode
225
+
226
+Controls how the collector finds candidate Azure resources.
227
+
228
+| Mode | Behavior |
229
+|:-----|:---------|
230
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
231
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
232
+
233
+
234
+<a id="option-discovery-discovery-mode-query-kql"></a>
235
+##### discovery.mode_query.kql
236
+
237
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
238
+
239
+The query **must** project these five columns:
240
+
241
+| Column | Description |
242
+|:-------|:------------|
243
+| `id` | Full Azure resource ID (ARM format) |
244
+| `name` | Resource name |
245
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
246
+| `resourceGroup` | Resource group name |
247
+| `location` | Azure region |
248
+
249
+Example:
250
+
251
+```
252
+resources
253
+| where tags.env =~ "prod"
254
+| project id, name, type, resourceGroup, location
255
+```
256
+
257
+
258
+<a id="option-profiles-profiles-mode"></a>
259
+##### profiles.mode
260
+
261
+Controls how the collector decides which metric profiles to activate.
262
+
263
+| Mode | Behavior |
264
+|:-----|:---------|
265
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
266
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
267
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
268
+
269
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
270
+
271
+
272
273
</details>
274
@@ -182,14 +310,28 @@ sudo ./edit-config go.d/azure_monitor.conf
310
311
##### Examples
312
185
-###### Service principal (auto-discover all resources)
313
+###### Service principal with structured discovery
314
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
315
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
316
317
```yaml
318
jobs:
319
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ subscription_ids:
321
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
322
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
323
+ discovery:
324
+ mode: filters
325
+ mode_filters:
326
+ resource_groups:
327
+ - production-rg
328
+ regions:
329
+ - eastus
330
+ tags:
331
+ env:
332
+ - prod
333
+ profiles:
334
+ mode: auto
335
auth:
336
mode: service_principal
337
mode_service_principal:
@@ -198,59 +340,49 @@ jobs:
340
client_secret: "your-client-secret"
341
342
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
343
+###### Managed identity with exact profiles
344
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
345
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
346
347
<details open><summary>Config</summary>
348
349
```yaml
350
jobs:
351
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
352
+ subscription_ids:
353
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
355
+ mode: exact
356
+ mode_exact:
357
+ names:
358
+ - sql_database
359
+ - postgres_flexible
360
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
361
+ mode: managed_identity
362
363
```
364
</details>
365
241
-###### Filter by resource group
366
+###### Custom Azure Resource Graph KQL
367
243
-Only monitor resources in specific resource groups.
368
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
369
370
<details open><summary>Config</summary>
371
372
```yaml
373
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
374
+ - name: prod-query
375
+ subscription_ids:
376
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
377
+ discovery:
378
+ mode: query
379
+ mode_query:
380
+ kql: |
381
+ resources
382
+ | where tags.env =~ "prod"
383
+ | project id, name, type, resourceGroup, location
384
+ profiles:
385
+ mode: auto
386
auth:
387
mode: default
388
@@ -259,14 +391,15 @@ jobs:
391
392
###### Azure Government cloud
393
262
-Connect to Azure Government cloud environment.
394
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
395
396
<details open><summary>Config</summary>
397
398
```yaml
399
jobs:
400
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
401
+ subscription_ids:
402
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
cloud: government
404
auth:
405
mode: service_principal
@@ -312,6 +445,7 @@ Labels:
445
| region | The Azure region where the resource is deployed. |
446
| resource_type | The Azure resource type identifier. |
447
| profile | The Azure Monitor profile id. |
448
+| subscription_id | The Azure subscription identifier. |
449
| resource_uid | The unique Azure resource identifier. |
450
451
Metrics:
@@ -395,31 +529,46 @@ docker logs netdata 2>&1 | grep azure_monitor
529
530
### No metrics are collected
531
398
-Verify the following:
399
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
400
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
401
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
402
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
532
+Check the following:
533
+
534
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
535
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
536
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
537
+- **Collector logs** -- Check for authentication or API errors:
538
+ ```bash
539
+ # systemd
540
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
541
+ # non-systemd
542
+ grep azure_monitor /var/log/netdata/collector.log
543
+ ```
544
545
546
### Missing metrics for some resource types
547
407
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
408
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
409
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
410
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
548
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
549
+
550
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
551
+- **Verify a built-in profile exists** -- List available profiles:
552
+ ```bash
553
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
554
+ ```
555
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
556
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
557
558
413
-### Metrics appear delayed
559
+### Charts have gaps or incomplete data
560
415
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
416
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
417
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
561
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
562
+
563
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
564
+- Slower time-grain batches automatically use a larger effective offset when needed.
565
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
566
567
568
### Authentication errors in sovereign clouds
569
570
For Azure Government or Azure China clouds, set the `cloud` parameter:
571
+
572
- Azure Government: `cloud: government`
573
- Azure China (21Vianet): `cloud: china`
574
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_cosmos_db_account.md
+239
-87
@@ -21,40 +21,78 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Cosmos DB accounts including request unit consumption and throttling, document counts and storage, data and index sizes, replication latency, availability percentages, provisioned throughput utilization, and normalized RU consumption per partition.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Cosmos DB with metrics covering:
31
+
32
+- **Request units** -- RU consumption, provisioned throughput (provisioned/autoscale), normalized RU per partition
33
+- **Storage** -- data, index, and quota storage, physical partition size
34
+- **Latency** -- server-side latency (direct/gateway), replication latency
35
+- **Requests** -- total requests, API requests (Mongo/Cassandra/Gremlin), metadata requests, dedicated gateway requests
36
+- **Availability** -- service availability percentage
37
+- **Advanced** -- document count, partition count, dedicated gateway CPU/memory, integrated cache hit rate
38
+
39
+
40
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
41
42
43
This collector is supported on all platforms.
44
45
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
46
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
47
+The service principal or managed identity requires these Azure RBAC roles:
48
+
49
+| Role | Purpose | Scope |
50
+|:-----|:--------|:------|
51
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
52
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
53
54
55
### Default Behavior
56
57
#### Auto-Detection
58
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
59
+The collector has two discovery phases:
60
+
61
+**Bootstrap (first run)**
62
+
63
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
64
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
65
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
66
+- A single job can monitor multiple subscriptions.
67
+
68
+**Runtime (periodic refresh)**
69
+
70
+- Periodically re-discovers resources for **already-active profile types only**.
71
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
72
+
73
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
74
75
76
#### Limits
77
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
78
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
79
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
80
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
81
82
83
#### Performance Impact
84
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
85
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
86
+
87
+**Default concurrency and batching limits:**
88
+
89
+| Setting | Default | Description |
90
+|:--------|:--------|:------------|
91
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
92
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
93
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
94
+
95
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
96
97
98
## Setup
@@ -78,25 +116,36 @@ UI configuration requires paid Netdata Cloud plan.
116
117
#### Create an Azure monitoring principal
118
81
-Create a service principal or use a managed identity with the following permissions:
119
+The collector requires a service principal or managed identity with two Azure RBAC roles:
120
+
121
+| Role | Purpose |
122
+|:-----|:--------|
123
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
124
+| **Reader** | Query Azure Resource Graph for resource discovery |
125
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
126
+**Option A: Service principal**
127
86
-For service principal authentication:
128
```bash
88
-# Create the service principal
129
+# Create service principal with Monitoring Reader role
130
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
131
--scopes /subscriptions/<subscription-id>
132
133
+# Add the Reader role for resource discovery
134
+az role assignment create --assignee <appId-from-above> \
135
+ --role "Reader" --scope /subscriptions/<subscription-id>
136
+
137
# Note the appId (client_id), password (client_secret), and tenant
138
```
139
95
-For managed identity (on Azure VMs, VMSS, or AKS):
140
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
141
+
142
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
143
+# Assign both roles to the VM's managed identity
144
az role assignment create --assignee <managed-identity-principal-id> \
145
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+
147
+az role assignment create --assignee <managed-identity-principal-id> \
148
+ --role "Reader" --scope /subscriptions/<subscription-id>
149
```
150
151
@@ -105,13 +154,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
154
155
#### Options
156
108
-The following options can be defined globally: update_every, autodetection_retry.
157
+The following options can be defined globally: `update_every`, `autodetection_retry`.
158
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+**Profile file locations:**
160
114
-User profile files with the same filename override stock profiles.
161
+| Type | Path |
162
+|:-----|:-----|
163
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
164
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
165
+
166
+User profile files with the same `id` as a stock profile override it.
167
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
168
169
170
<details open><summary>Config options</summary>
@@ -122,25 +175,103 @@ User profile files with the same filename override stock profiles.
175
|:------|:-----|:------------|:--------|:---------:|
176
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
177
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
178
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
179
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
183
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
186
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
187
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
188
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
189
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
192
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
193
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
194
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
195
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
196
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
197
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
198
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
199
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
200
201
+<a id="option-collection-query-offset"></a>
202
+##### query_offset
203
+
204
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
205
+
206
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
207
+
208
+- **Default (180s)** works for most services.
209
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
210
+- **Increase to 240-300s** if you still see gaps or missing data points.
211
+- **Do not set below 60s** -- metrics will likely be incomplete.
212
+
213
+
214
+<a id="option-authentication-auth-mode"></a>
215
+##### auth.mode
216
+
217
+Determines how the collector authenticates with Azure.
218
+
219
+| Mode | When to use | Required options |
220
+|:-----|:------------|:-----------------|
221
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
222
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
223
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
224
+
225
+
226
+<a id="option-discovery-discovery-mode"></a>
227
+##### discovery.mode
228
+
229
+Controls how the collector finds candidate Azure resources.
230
+
231
+| Mode | Behavior |
232
+|:-----|:---------|
233
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
234
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
235
+
236
+
237
+<a id="option-discovery-discovery-mode-query-kql"></a>
238
+##### discovery.mode_query.kql
239
+
240
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
241
+
242
+The query **must** project these five columns:
243
+
244
+| Column | Description |
245
+|:-------|:------------|
246
+| `id` | Full Azure resource ID (ARM format) |
247
+| `name` | Resource name |
248
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
249
+| `resourceGroup` | Resource group name |
250
+| `location` | Azure region |
251
+
252
+Example:
253
+
254
+```
255
+resources
256
+| where tags.env =~ "prod"
257
+| project id, name, type, resourceGroup, location
258
+```
259
+
260
+
261
+<a id="option-profiles-profiles-mode"></a>
262
+##### profiles.mode
263
+
264
+Controls how the collector decides which metric profiles to activate.
265
+
266
+| Mode | Behavior |
267
+|:-----|:---------|
268
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
269
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
270
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
271
+
272
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
273
+
274
+
275
276
</details>
277
@@ -182,14 +313,28 @@ sudo ./edit-config go.d/azure_monitor.conf
313
314
##### Examples
315
185
-###### Service principal (auto-discover all resources)
316
+###### Service principal with structured discovery
317
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
318
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
319
320
```yaml
321
jobs:
322
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ subscription_ids:
324
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
325
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
326
+ discovery:
327
+ mode: filters
328
+ mode_filters:
329
+ resource_groups:
330
+ - production-rg
331
+ regions:
332
+ - eastus
333
+ tags:
334
+ env:
335
+ - prod
336
+ profiles:
337
+ mode: auto
338
auth:
339
mode: service_principal
340
mode_service_principal:
@@ -198,59 +343,49 @@ jobs:
343
client_secret: "your-client-secret"
344
345
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
346
+###### Managed identity with exact profiles
347
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
348
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
349
350
<details open><summary>Config</summary>
351
352
```yaml
353
jobs:
354
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
+ subscription_ids:
356
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
358
+ mode: exact
359
+ mode_exact:
360
+ names:
361
+ - sql_database
362
+ - postgres_flexible
363
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
364
+ mode: managed_identity
365
366
```
367
</details>
368
241
-###### Filter by resource group
369
+###### Custom Azure Resource Graph KQL
370
243
-Only monitor resources in specific resource groups.
371
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
372
373
<details open><summary>Config</summary>
374
375
```yaml
376
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
377
+ - name: prod-query
378
+ subscription_ids:
379
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
380
+ discovery:
381
+ mode: query
382
+ mode_query:
383
+ kql: |
384
+ resources
385
+ | where tags.env =~ "prod"
386
+ | project id, name, type, resourceGroup, location
387
+ profiles:
388
+ mode: auto
389
auth:
390
mode: default
391
@@ -259,14 +394,15 @@ jobs:
394
395
###### Azure Government cloud
396
262
-Connect to Azure Government cloud environment.
397
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
398
399
<details open><summary>Config</summary>
400
401
```yaml
402
jobs:
403
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
+ subscription_ids:
405
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
cloud: government
407
auth:
408
mode: service_principal
@@ -320,6 +456,7 @@ Labels:
456
| region | The Azure region where the resource is deployed. |
457
| resource_type | The Azure resource type identifier. |
458
| profile | The Azure Monitor profile id. |
459
+| subscription_id | The Azure subscription identifier. |
460
| resource_uid | The unique Azure resource identifier. |
461
462
Metrics:
@@ -418,31 +555,46 @@ docker logs netdata 2>&1 | grep azure_monitor
555
556
### No metrics are collected
557
421
-Verify the following:
422
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
423
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
424
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
425
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
558
+Check the following:
559
+
560
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
561
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
562
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
563
+- **Collector logs** -- Check for authentication or API errors:
564
+ ```bash
565
+ # systemd
566
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
567
+ # non-systemd
568
+ grep azure_monitor /var/log/netdata/collector.log
569
+ ```
570
571
572
### Missing metrics for some resource types
573
430
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
431
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
432
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
433
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
574
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
575
+
576
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
577
+- **Verify a built-in profile exists** -- List available profiles:
578
+ ```bash
579
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
580
+ ```
581
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
582
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
583
584
436
-### Metrics appear delayed
585
+### Charts have gaps or incomplete data
586
438
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
439
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
440
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
587
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
588
+
589
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
590
+- Slower time-grain batches automatically use a larger effective offset when needed.
591
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
592
593
594
### Authentication errors in sovereign clouds
595
596
For Azure Government or Azure China clouds, set the `cloud` parameter:
597
+
598
- Azure Government: `cloud: government`
599
- Azure China (21Vianet): `cloud: china`
600
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_explorer_cluster.md
+243
-87
@@ -21,40 +21,82 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Data Explorer (Kusto) clusters including ingestion latency, volume, and success rates, query performance and concurrency, cache utilization, CPU and memory usage, export operations, streaming ingest throughput, materialized view health, instance counts, and follower lag.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Data Explorer (Kusto) with metrics covering:
31
+
32
+- **Ingestion** -- ingestion latency, volume, result (success/failure), queue length, batch processing
33
+- **Queries** -- query count, query duration, concurrent queries, throttled queries/commands
34
+- **Streaming ingest** -- data rate, duration, result, utilization
35
+- **Cache** -- cache and ingestion utilization
36
+- **Compute** -- CPU utilization
37
+- **Export** -- continuous export records, result, lateness, pending jobs, export utilization
38
+- **Materialized views** -- view health, age, data loss, records in delta, extents rebuild
39
+- **Cluster** -- instance count (average/max/min), keep alive, total extents
40
+- **Events** -- events received/processed/dropped, blobs received/processed/dropped
41
+- **Advanced** -- follower latency, discovery latency, weak consistency latency, partitioning percentage
42
+
43
+
44
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
45
46
47
This collector is supported on all platforms.
48
49
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
50
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
51
+The service principal or managed identity requires these Azure RBAC roles:
52
+
53
+| Role | Purpose | Scope |
54
+|:-----|:--------|:------|
55
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
56
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
57
58
59
### Default Behavior
60
61
#### Auto-Detection
62
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
63
+The collector has two discovery phases:
64
+
65
+**Bootstrap (first run)**
66
+
67
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
68
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
69
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
70
+- A single job can monitor multiple subscriptions.
71
+
72
+**Runtime (periodic refresh)**
73
+
74
+- Periodically re-discovers resources for **already-active profile types only**.
75
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
76
+
77
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
78
79
80
#### Limits
81
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
82
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
83
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
84
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
85
86
87
#### Performance Impact
88
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
89
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
90
+
91
+**Default concurrency and batching limits:**
92
+
93
+| Setting | Default | Description |
94
+|:--------|:--------|:------------|
95
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
96
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
97
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
98
+
99
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
100
101
102
## Setup
@@ -78,25 +120,36 @@ UI configuration requires paid Netdata Cloud plan.
120
121
#### Create an Azure monitoring principal
122
81
-Create a service principal or use a managed identity with the following permissions:
123
+The collector requires a service principal or managed identity with two Azure RBAC roles:
124
+
125
+| Role | Purpose |
126
+|:-----|:--------|
127
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
128
+| **Reader** | Query Azure Resource Graph for resource discovery |
129
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
130
+**Option A: Service principal**
131
86
-For service principal authentication:
132
```bash
88
-# Create the service principal
133
+# Create service principal with Monitoring Reader role
134
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
135
--scopes /subscriptions/<subscription-id>
136
137
+# Add the Reader role for resource discovery
138
+az role assignment create --assignee <appId-from-above> \
139
+ --role "Reader" --scope /subscriptions/<subscription-id>
140
+
141
# Note the appId (client_id), password (client_secret), and tenant
142
```
143
95
-For managed identity (on Azure VMs, VMSS, or AKS):
144
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
145
+
146
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
147
+# Assign both roles to the VM's managed identity
148
az role assignment create --assignee <managed-identity-principal-id> \
149
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
150
+
151
+az role assignment create --assignee <managed-identity-principal-id> \
152
+ --role "Reader" --scope /subscriptions/<subscription-id>
153
```
154
155
@@ -105,13 +158,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
158
159
#### Options
160
108
-The following options can be defined globally: update_every, autodetection_retry.
161
+The following options can be defined globally: `update_every`, `autodetection_retry`.
162
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
163
+**Profile file locations:**
164
114
-User profile files with the same filename override stock profiles.
165
+| Type | Path |
166
+|:-----|:-----|
167
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
168
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
169
+
170
+User profile files with the same `id` as a stock profile override it.
171
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
172
173
174
<details open><summary>Config options</summary>
@@ -122,25 +179,103 @@ User profile files with the same filename override stock profiles.
179
|:------|:-----|:------------|:--------|:---------:|
180
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
181
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
182
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
183
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
185
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
186
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
187
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
189
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
190
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
191
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
192
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
193
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
195
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
196
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
197
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
198
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
199
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
200
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
201
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
202
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
203
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
204
205
+<a id="option-collection-query-offset"></a>
206
+##### query_offset
207
+
208
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
209
+
210
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
211
+
212
+- **Default (180s)** works for most services.
213
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
214
+- **Increase to 240-300s** if you still see gaps or missing data points.
215
+- **Do not set below 60s** -- metrics will likely be incomplete.
216
+
217
+
218
+<a id="option-authentication-auth-mode"></a>
219
+##### auth.mode
220
+
221
+Determines how the collector authenticates with Azure.
222
+
223
+| Mode | When to use | Required options |
224
+|:-----|:------------|:-----------------|
225
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
226
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
227
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
228
+
229
+
230
+<a id="option-discovery-discovery-mode"></a>
231
+##### discovery.mode
232
+
233
+Controls how the collector finds candidate Azure resources.
234
+
235
+| Mode | Behavior |
236
+|:-----|:---------|
237
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
238
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
239
+
240
+
241
+<a id="option-discovery-discovery-mode-query-kql"></a>
242
+##### discovery.mode_query.kql
243
+
244
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
245
+
246
+The query **must** project these five columns:
247
+
248
+| Column | Description |
249
+|:-------|:------------|
250
+| `id` | Full Azure resource ID (ARM format) |
251
+| `name` | Resource name |
252
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
253
+| `resourceGroup` | Resource group name |
254
+| `location` | Azure region |
255
+
256
+Example:
257
+
258
+```
259
+resources
260
+| where tags.env =~ "prod"
261
+| project id, name, type, resourceGroup, location
262
+```
263
+
264
+
265
+<a id="option-profiles-profiles-mode"></a>
266
+##### profiles.mode
267
+
268
+Controls how the collector decides which metric profiles to activate.
269
+
270
+| Mode | Behavior |
271
+|:-----|:---------|
272
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
273
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
274
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
275
+
276
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
277
+
278
+
279
280
</details>
281
@@ -182,14 +317,28 @@ sudo ./edit-config go.d/azure_monitor.conf
317
318
##### Examples
319
185
-###### Service principal (auto-discover all resources)
320
+###### Service principal with structured discovery
321
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
322
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
323
324
```yaml
325
jobs:
326
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
327
+ subscription_ids:
328
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
329
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
330
+ discovery:
331
+ mode: filters
332
+ mode_filters:
333
+ resource_groups:
334
+ - production-rg
335
+ regions:
336
+ - eastus
337
+ tags:
338
+ env:
339
+ - prod
340
+ profiles:
341
+ mode: auto
342
auth:
343
mode: service_principal
344
mode_service_principal:
@@ -198,59 +347,49 @@ jobs:
347
client_secret: "your-client-secret"
348
349
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
350
+###### Managed identity with exact profiles
351
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
352
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
353
354
<details open><summary>Config</summary>
355
356
```yaml
357
jobs:
358
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
+ subscription_ids:
360
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
361
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
362
+ mode: exact
363
+ mode_exact:
364
+ names:
365
+ - sql_database
366
+ - postgres_flexible
367
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
368
+ mode: managed_identity
369
370
```
371
</details>
372
241
-###### Filter by resource group
373
+###### Custom Azure Resource Graph KQL
374
243
-Only monitor resources in specific resource groups.
375
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
376
377
<details open><summary>Config</summary>
378
379
```yaml
380
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
381
+ - name: prod-query
382
+ subscription_ids:
383
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
384
+ discovery:
385
+ mode: query
386
+ mode_query:
387
+ kql: |
388
+ resources
389
+ | where tags.env =~ "prod"
390
+ | project id, name, type, resourceGroup, location
391
+ profiles:
392
+ mode: auto
393
auth:
394
mode: default
395
@@ -259,14 +398,15 @@ jobs:
398
399
###### Azure Government cloud
400
262
-Connect to Azure Government cloud environment.
401
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
402
403
<details open><summary>Config</summary>
404
405
```yaml
406
jobs:
407
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
+ subscription_ids:
409
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
410
cloud: government
411
auth:
412
mode: service_principal
@@ -331,6 +471,7 @@ Labels:
471
| region | The Azure region where the resource is deployed. |
472
| resource_type | The Azure resource type identifier. |
473
| profile | The Azure Monitor profile id. |
474
+| subscription_id | The Azure subscription identifier. |
475
| resource_uid | The unique Azure resource identifier. |
476
477
Metrics:
@@ -453,31 +594,46 @@ docker logs netdata 2>&1 | grep azure_monitor
594
595
### No metrics are collected
596
456
-Verify the following:
457
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
458
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
459
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
460
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
597
+Check the following:
598
+
599
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
600
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
601
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
602
+- **Collector logs** -- Check for authentication or API errors:
603
+ ```bash
604
+ # systemd
605
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
606
+ # non-systemd
607
+ grep azure_monitor /var/log/netdata/collector.log
608
+ ```
609
610
611
### Missing metrics for some resource types
612
465
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
466
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
467
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
468
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
613
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
614
+
615
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
616
+- **Verify a built-in profile exists** -- List available profiles:
617
+ ```bash
618
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
619
+ ```
620
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
621
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
622
623
471
-### Metrics appear delayed
624
+### Charts have gaps or incomplete data
625
473
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
474
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
475
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
626
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
627
+
628
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
629
+- Slower time-grain batches automatically use a larger effective offset when needed.
630
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
631
632
633
### Authentication errors in sovereign clouds
634
635
For Azure Government or Azure China clouds, set the `cloud` parameter:
636
+
637
- Azure Government: `cloud: government`
638
- Azure China (21Vianet): `cloud: china`
639
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_data_factory.md
+241
-87
@@ -21,40 +21,80 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Data Factory including pipeline, activity, and trigger run success and failure counts, integration runtime CPU and memory utilization, available capacity and queue lengths, SSIS package execution rates, copy operations throughput, data flow processing metrics, and overall factory resource utilization.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Data Factory with metrics covering:
31
+
32
+- **Pipeline runs** -- pipeline runs (succeeded/failed/cancelled), elapsed time runs
33
+- **Activity runs** -- activity runs (succeeded/failed/cancelled)
34
+- **Trigger runs** -- trigger runs (succeeded/failed/cancelled)
35
+- **Integration runtime** -- IR CPU and memory utilization, available nodes, queue length, task pickup delay
36
+- **SSIS** -- SSIS package executions (succeeded/failed/cancelled), IR start/stop operations
37
+- **MVNet IR** -- pipeline and copy capacity/utilization, external capacity, queue lengths
38
+- **Airflow IR** -- CPU and memory, DAG processing, task instances, scheduler activity, triggers, pool slots
39
+- **Factory** -- entity count, factory size (current/max allowed)
40
+
41
+
42
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
43
44
45
This collector is supported on all platforms.
46
47
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
48
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
49
+The service principal or managed identity requires these Azure RBAC roles:
50
+
51
+| Role | Purpose | Scope |
52
+|:-----|:--------|:------|
53
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
54
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
55
56
57
### Default Behavior
58
59
#### Auto-Detection
60
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
61
+The collector has two discovery phases:
62
+
63
+**Bootstrap (first run)**
64
+
65
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
66
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
67
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
68
+- A single job can monitor multiple subscriptions.
69
+
70
+**Runtime (periodic refresh)**
71
+
72
+- Periodically re-discovers resources for **already-active profile types only**.
73
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
74
+
75
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
76
77
78
#### Limits
79
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
80
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
81
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
82
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
83
84
85
#### Performance Impact
86
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
87
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
88
+
89
+**Default concurrency and batching limits:**
90
+
91
+| Setting | Default | Description |
92
+|:--------|:--------|:------------|
93
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
94
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
95
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
96
+
97
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
98
99
100
## Setup
@@ -78,25 +118,36 @@ UI configuration requires paid Netdata Cloud plan.
118
119
#### Create an Azure monitoring principal
120
81
-Create a service principal or use a managed identity with the following permissions:
121
+The collector requires a service principal or managed identity with two Azure RBAC roles:
122
+
123
+| Role | Purpose |
124
+|:-----|:--------|
125
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
126
+| **Reader** | Query Azure Resource Graph for resource discovery |
127
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
128
+**Option A: Service principal**
129
86
-For service principal authentication:
130
```bash
88
-# Create the service principal
131
+# Create service principal with Monitoring Reader role
132
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
133
--scopes /subscriptions/<subscription-id>
134
135
+# Add the Reader role for resource discovery
136
+az role assignment create --assignee <appId-from-above> \
137
+ --role "Reader" --scope /subscriptions/<subscription-id>
138
+
139
# Note the appId (client_id), password (client_secret), and tenant
140
```
141
95
-For managed identity (on Azure VMs, VMSS, or AKS):
142
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
143
+
144
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
145
+# Assign both roles to the VM's managed identity
146
az role assignment create --assignee <managed-identity-principal-id> \
147
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
148
+
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Reader" --scope /subscriptions/<subscription-id>
151
```
152
153
@@ -105,13 +156,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
156
157
#### Options
158
108
-The following options can be defined globally: update_every, autodetection_retry.
159
+The following options can be defined globally: `update_every`, `autodetection_retry`.
160
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
161
+**Profile file locations:**
162
114
-User profile files with the same filename override stock profiles.
163
+| Type | Path |
164
+|:-----|:-----|
165
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
166
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
167
+
168
+User profile files with the same `id` as a stock profile override it.
169
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
170
171
172
<details open><summary>Config options</summary>
@@ -122,25 +177,103 @@ User profile files with the same filename override stock profiles.
177
|:------|:-----|:------------|:--------|:---------:|
178
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
179
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
181
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
184
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
185
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
190
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
191
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
194
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
195
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
196
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
197
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
198
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
199
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
200
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
201
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
203
+<a id="option-collection-query-offset"></a>
204
+##### query_offset
205
+
206
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
207
+
208
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
209
+
210
+- **Default (180s)** works for most services.
211
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
212
+- **Increase to 240-300s** if you still see gaps or missing data points.
213
+- **Do not set below 60s** -- metrics will likely be incomplete.
214
+
215
+
216
+<a id="option-authentication-auth-mode"></a>
217
+##### auth.mode
218
+
219
+Determines how the collector authenticates with Azure.
220
+
221
+| Mode | When to use | Required options |
222
+|:-----|:------------|:-----------------|
223
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
224
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
225
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
226
+
227
+
228
+<a id="option-discovery-discovery-mode"></a>
229
+##### discovery.mode
230
+
231
+Controls how the collector finds candidate Azure resources.
232
+
233
+| Mode | Behavior |
234
+|:-----|:---------|
235
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
236
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
237
+
238
+
239
+<a id="option-discovery-discovery-mode-query-kql"></a>
240
+##### discovery.mode_query.kql
241
+
242
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
243
+
244
+The query **must** project these five columns:
245
+
246
+| Column | Description |
247
+|:-------|:------------|
248
+| `id` | Full Azure resource ID (ARM format) |
249
+| `name` | Resource name |
250
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
251
+| `resourceGroup` | Resource group name |
252
+| `location` | Azure region |
253
+
254
+Example:
255
+
256
+```
257
+resources
258
+| where tags.env =~ "prod"
259
+| project id, name, type, resourceGroup, location
260
+```
261
+
262
+
263
+<a id="option-profiles-profiles-mode"></a>
264
+##### profiles.mode
265
+
266
+Controls how the collector decides which metric profiles to activate.
267
+
268
+| Mode | Behavior |
269
+|:-----|:---------|
270
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
271
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
272
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
273
+
274
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
275
+
276
+
277
278
</details>
279
@@ -182,14 +315,28 @@ sudo ./edit-config go.d/azure_monitor.conf
315
316
##### Examples
317
185
-###### Service principal (auto-discover all resources)
318
+###### Service principal with structured discovery
319
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
320
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
321
322
```yaml
323
jobs:
324
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ subscription_ids:
326
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
327
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
328
+ discovery:
329
+ mode: filters
330
+ mode_filters:
331
+ resource_groups:
332
+ - production-rg
333
+ regions:
334
+ - eastus
335
+ tags:
336
+ env:
337
+ - prod
338
+ profiles:
339
+ mode: auto
340
auth:
341
mode: service_principal
342
mode_service_principal:
@@ -198,59 +345,49 @@ jobs:
345
client_secret: "your-client-secret"
346
347
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
348
+###### Managed identity with exact profiles
349
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
350
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
351
352
<details open><summary>Config</summary>
353
354
```yaml
355
jobs:
356
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ subscription_ids:
358
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
360
+ mode: exact
361
+ mode_exact:
362
+ names:
363
+ - sql_database
364
+ - postgres_flexible
365
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
366
+ mode: managed_identity
367
368
```
369
</details>
370
241
-###### Filter by resource group
371
+###### Custom Azure Resource Graph KQL
372
243
-Only monitor resources in specific resource groups.
373
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
374
375
<details open><summary>Config</summary>
376
377
```yaml
378
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
379
+ - name: prod-query
380
+ subscription_ids:
381
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
382
+ discovery:
383
+ mode: query
384
+ mode_query:
385
+ kql: |
386
+ resources
387
+ | where tags.env =~ "prod"
388
+ | project id, name, type, resourceGroup, location
389
+ profiles:
390
+ mode: auto
391
auth:
392
mode: default
393
@@ -259,14 +396,15 @@ jobs:
396
397
###### Azure Government cloud
398
262
-Connect to Azure Government cloud environment.
399
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
400
401
<details open><summary>Config</summary>
402
403
```yaml
404
jobs:
405
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
+ subscription_ids:
407
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
cloud: government
409
auth:
410
mode: service_principal
@@ -336,6 +474,7 @@ Labels:
474
| region | The Azure region where the resource is deployed. |
475
| resource_type | The Azure resource type identifier. |
476
| profile | The Azure Monitor profile id. |
477
+| subscription_id | The Azure subscription identifier. |
478
| resource_uid | The unique Azure resource identifier. |
479
480
Metrics:
@@ -462,31 +601,46 @@ docker logs netdata 2>&1 | grep azure_monitor
601
602
### No metrics are collected
603
465
-Verify the following:
466
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
467
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
468
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
469
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
604
+Check the following:
605
+
606
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
607
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
608
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
609
+- **Collector logs** -- Check for authentication or API errors:
610
+ ```bash
611
+ # systemd
612
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
613
+ # non-systemd
614
+ grep azure_monitor /var/log/netdata/collector.log
615
+ ```
616
617
618
### Missing metrics for some resource types
619
474
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
475
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
476
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
477
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
620
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
621
+
622
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
623
+- **Verify a built-in profile exists** -- List available profiles:
624
+ ```bash
625
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
626
+ ```
627
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
628
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
629
630
480
-### Metrics appear delayed
631
+### Charts have gaps or incomplete data
632
482
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
483
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
484
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
633
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
634
+
635
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
636
+- Slower time-grain batches automatically use a larger effective offset when needed.
637
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
638
639
640
### Authentication errors in sovereign clouds
641
642
For Azure Government or Azure China clouds, set the `cloud` parameter:
643
+
644
- Azure Government: `cloud: government`
645
- Azure China (21Vianet): `cloud: china`
646
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_grid_topic.md
+237
-87
@@ -21,40 +21,76 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Event Grid topics including publish success and failure counts, publish latency, event delivery and routing rates, delivery success and failure counts, dead-lettered events, and matched event routing.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Event Grid with metrics covering:
31
+
32
+- **Publishing** -- publish rate (success/failed), publish latency
33
+- **Delivery** -- events delivered, failed, dropped, and dead-lettered
34
+- **Routing** -- matched and unmatched event routing, destination processing duration
35
+- **Filters** -- advanced filter evaluations
36
+
37
+
38
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
39
40
41
This collector is supported on all platforms.
42
43
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
44
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
45
+The service principal or managed identity requires these Azure RBAC roles:
46
+
47
+| Role | Purpose | Scope |
48
+|:-----|:--------|:------|
49
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
50
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
51
52
53
### Default Behavior
54
55
#### Auto-Detection
56
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
57
+The collector has two discovery phases:
58
+
59
+**Bootstrap (first run)**
60
+
61
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
62
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
63
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
64
+- A single job can monitor multiple subscriptions.
65
+
66
+**Runtime (periodic refresh)**
67
+
68
+- Periodically re-discovers resources for **already-active profile types only**.
69
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
70
+
71
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
72
73
74
#### Limits
75
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
76
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
77
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
78
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
79
80
81
#### Performance Impact
82
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
83
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
84
+
85
+**Default concurrency and batching limits:**
86
+
87
+| Setting | Default | Description |
88
+|:--------|:--------|:------------|
89
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
90
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
91
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
92
+
93
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
94
95
96
## Setup
@@ -78,25 +114,36 @@ UI configuration requires paid Netdata Cloud plan.
114
115
#### Create an Azure monitoring principal
116
81
-Create a service principal or use a managed identity with the following permissions:
117
+The collector requires a service principal or managed identity with two Azure RBAC roles:
118
+
119
+| Role | Purpose |
120
+|:-----|:--------|
121
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
122
+| **Reader** | Query Azure Resource Graph for resource discovery |
123
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
124
+**Option A: Service principal**
125
86
-For service principal authentication:
126
```bash
88
-# Create the service principal
127
+# Create service principal with Monitoring Reader role
128
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
129
--scopes /subscriptions/<subscription-id>
130
131
+# Add the Reader role for resource discovery
132
+az role assignment create --assignee <appId-from-above> \
133
+ --role "Reader" --scope /subscriptions/<subscription-id>
134
+
135
# Note the appId (client_id), password (client_secret), and tenant
136
```
137
95
-For managed identity (on Azure VMs, VMSS, or AKS):
138
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
139
+
140
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
141
+# Assign both roles to the VM's managed identity
142
az role assignment create --assignee <managed-identity-principal-id> \
143
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
144
+
145
+az role assignment create --assignee <managed-identity-principal-id> \
146
+ --role "Reader" --scope /subscriptions/<subscription-id>
147
```
148
149
@@ -105,13 +152,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
152
153
#### Options
154
108
-The following options can be defined globally: update_every, autodetection_retry.
155
+The following options can be defined globally: `update_every`, `autodetection_retry`.
156
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
157
+**Profile file locations:**
158
114
-User profile files with the same filename override stock profiles.
159
+| Type | Path |
160
+|:-----|:-----|
161
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
162
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
163
+
164
+User profile files with the same `id` as a stock profile override it.
165
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
166
167
168
<details open><summary>Config options</summary>
@@ -122,25 +173,103 @@ User profile files with the same filename override stock profiles.
173
|:------|:-----|:------------|:--------|:---------:|
174
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
176
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
177
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
180
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
181
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
184
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
185
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
186
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
187
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
190
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
191
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
192
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
193
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
194
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
195
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
196
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
197
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
198
199
+<a id="option-collection-query-offset"></a>
200
+##### query_offset
201
+
202
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
203
+
204
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
205
+
206
+- **Default (180s)** works for most services.
207
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
208
+- **Increase to 240-300s** if you still see gaps or missing data points.
209
+- **Do not set below 60s** -- metrics will likely be incomplete.
210
+
211
+
212
+<a id="option-authentication-auth-mode"></a>
213
+##### auth.mode
214
+
215
+Determines how the collector authenticates with Azure.
216
+
217
+| Mode | When to use | Required options |
218
+|:-----|:------------|:-----------------|
219
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
220
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
221
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
222
+
223
+
224
+<a id="option-discovery-discovery-mode"></a>
225
+##### discovery.mode
226
+
227
+Controls how the collector finds candidate Azure resources.
228
+
229
+| Mode | Behavior |
230
+|:-----|:---------|
231
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
232
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
233
+
234
+
235
+<a id="option-discovery-discovery-mode-query-kql"></a>
236
+##### discovery.mode_query.kql
237
+
238
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
239
+
240
+The query **must** project these five columns:
241
+
242
+| Column | Description |
243
+|:-------|:------------|
244
+| `id` | Full Azure resource ID (ARM format) |
245
+| `name` | Resource name |
246
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
247
+| `resourceGroup` | Resource group name |
248
+| `location` | Azure region |
249
+
250
+Example:
251
+
252
+```
253
+resources
254
+| where tags.env =~ "prod"
255
+| project id, name, type, resourceGroup, location
256
+```
257
+
258
+
259
+<a id="option-profiles-profiles-mode"></a>
260
+##### profiles.mode
261
+
262
+Controls how the collector decides which metric profiles to activate.
263
+
264
+| Mode | Behavior |
265
+|:-----|:---------|
266
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
267
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
268
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
269
+
270
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
271
+
272
+
273
274
</details>
275
@@ -182,14 +311,28 @@ sudo ./edit-config go.d/azure_monitor.conf
311
312
##### Examples
313
185
-###### Service principal (auto-discover all resources)
314
+###### Service principal with structured discovery
315
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
316
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
317
318
```yaml
319
jobs:
320
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ subscription_ids:
322
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
323
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
324
+ discovery:
325
+ mode: filters
326
+ mode_filters:
327
+ resource_groups:
328
+ - production-rg
329
+ regions:
330
+ - eastus
331
+ tags:
332
+ env:
333
+ - prod
334
+ profiles:
335
+ mode: auto
336
auth:
337
mode: service_principal
338
mode_service_principal:
@@ -198,59 +341,49 @@ jobs:
341
client_secret: "your-client-secret"
342
343
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
344
+###### Managed identity with exact profiles
345
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
346
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
347
348
<details open><summary>Config</summary>
349
350
```yaml
351
jobs:
352
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ subscription_ids:
354
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
356
+ mode: exact
357
+ mode_exact:
358
+ names:
359
+ - sql_database
360
+ - postgres_flexible
361
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
362
+ mode: managed_identity
363
364
```
365
</details>
366
241
-###### Filter by resource group
367
+###### Custom Azure Resource Graph KQL
368
243
-Only monitor resources in specific resource groups.
369
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
370
371
<details open><summary>Config</summary>
372
373
```yaml
374
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
375
+ - name: prod-query
376
+ subscription_ids:
377
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ discovery:
379
+ mode: query
380
+ mode_query:
381
+ kql: |
382
+ resources
383
+ | where tags.env =~ "prod"
384
+ | project id, name, type, resourceGroup, location
385
+ profiles:
386
+ mode: auto
387
auth:
388
mode: default
389
@@ -259,14 +392,15 @@ jobs:
392
393
###### Azure Government cloud
394
262
-Connect to Azure Government cloud environment.
395
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
396
397
<details open><summary>Config</summary>
398
399
```yaml
400
jobs:
401
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
402
+ subscription_ids:
403
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
cloud: government
405
auth:
406
mode: service_principal
@@ -316,6 +450,7 @@ Labels:
450
| region | The Azure region where the resource is deployed. |
451
| resource_type | The Azure resource type identifier. |
452
| profile | The Azure Monitor profile id. |
453
+| subscription_id | The Azure subscription identifier. |
454
| resource_uid | The unique Azure resource identifier. |
455
456
Metrics:
@@ -400,31 +535,46 @@ docker logs netdata 2>&1 | grep azure_monitor
535
536
### No metrics are collected
537
403
-Verify the following:
404
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
405
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
406
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
407
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
538
+Check the following:
539
+
540
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
541
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
542
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
543
+- **Collector logs** -- Check for authentication or API errors:
544
+ ```bash
545
+ # systemd
546
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
547
+ # non-systemd
548
+ grep azure_monitor /var/log/netdata/collector.log
549
+ ```
550
551
552
### Missing metrics for some resource types
553
412
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
413
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
414
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
415
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
554
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
555
+
556
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
557
+- **Verify a built-in profile exists** -- List available profiles:
558
+ ```bash
559
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
560
+ ```
561
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
562
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
563
564
418
-### Metrics appear delayed
565
+### Charts have gaps or incomplete data
566
420
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
421
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
422
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
567
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
568
+
569
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
570
+- Slower time-grain batches automatically use a larger effective offset when needed.
571
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
572
573
574
### Authentication errors in sovereign clouds
575
576
For Azure Government or Azure China clouds, set the `cloud` parameter:
577
+
578
- Azure Government: `cloud: government`
579
- Azure China (21Vianet): `cloud: china`
580
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_event_hubs_namespace.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Event Hubs namespaces including incoming and outgoing message rates, byte throughput, captured messages and bytes, throttled and quota-exceeded request counts, active connections, and total connection counts.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Event Hubs with metrics covering:
31
+
32
+- **Messages** -- message flow (in/out), captured messages and bytes
33
+- **Throughput** -- data throughput (in/out bytes per second)
34
+- **Connections** -- active connections, connection events (opened/closed)
35
+- **Requests** -- incoming and successful request rates
36
+- **Errors** -- server errors, user errors, throttled requests, quota exceeded
37
+- **Replication** -- replication lag (messages and duration)
38
+- **Resources** -- namespace size, CPU and memory utilization
39
+
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
+
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -320,6 +457,7 @@ Labels:
457
| region | The Azure region where the resource is deployed. |
458
| resource_type | The Azure resource type identifier. |
459
| profile | The Azure Monitor profile id. |
460
+| subscription_id | The Azure subscription identifier. |
461
| resource_uid | The unique Azure resource identifier. |
462
463
Metrics:
@@ -411,31 +549,46 @@ docker logs netdata 2>&1 | grep azure_monitor
549
550
### No metrics are collected
551
414
-Verify the following:
415
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
416
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
417
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
418
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
552
+Check the following:
553
+
554
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
555
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
556
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
557
+- **Collector logs** -- Check for authentication or API errors:
558
+ ```bash
559
+ # systemd
560
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
561
+ # non-systemd
562
+ grep azure_monitor /var/log/netdata/collector.log
563
+ ```
564
565
566
### Missing metrics for some resource types
567
423
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
424
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
425
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
426
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
568
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
569
+
570
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
571
+- **Verify a built-in profile exists** -- List available profiles:
572
+ ```bash
573
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
574
+ ```
575
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
576
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
577
578
429
-### Metrics appear delayed
579
+### Charts have gaps or incomplete data
580
431
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
432
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
433
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
581
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
582
+
583
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
584
+- Slower time-grain batches automatically use a larger effective offset when needed.
585
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
586
587
588
### Authentication errors in sovereign clouds
589
590
For Azure Government or Azure China clouds, set the `cloud` parameter:
591
+
592
- Azure Government: `cloud: government`
593
- Azure China (21Vianet): `cloud: china`
594
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_circuit.md
+238
-87
@@ -21,40 +21,77 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor ExpressRoute circuits including bits per second in and out, ARP and BGP availability percentages, packet drops, and QoS bit rate throughput.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure ExpressRoute Circuit with metrics covering:
31
+
32
+- **Throughput** -- circuit throughput (bits/s in/out), GlobalReach throughput
33
+- **Availability** -- ARP availability, BGP availability
34
+- **Bandwidth** -- bandwidth utilization (ingress/egress)
35
+- **QoS** -- QoS dropped bits (in/out)
36
+- **Routes** -- FastPath routes count
37
+
38
+
39
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
40
41
42
This collector is supported on all platforms.
43
44
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
45
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
46
+The service principal or managed identity requires these Azure RBAC roles:
47
+
48
+| Role | Purpose | Scope |
49
+|:-----|:--------|:------|
50
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
51
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
52
53
54
### Default Behavior
55
56
#### Auto-Detection
57
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
58
+The collector has two discovery phases:
59
+
60
+**Bootstrap (first run)**
61
+
62
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
63
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
64
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
65
+- A single job can monitor multiple subscriptions.
66
+
67
+**Runtime (periodic refresh)**
68
+
69
+- Periodically re-discovers resources for **already-active profile types only**.
70
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
71
+
72
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
73
74
75
#### Limits
76
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
77
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
78
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
79
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
80
81
82
#### Performance Impact
83
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
84
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
85
+
86
+**Default concurrency and batching limits:**
87
+
88
+| Setting | Default | Description |
89
+|:--------|:--------|:------------|
90
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
91
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
92
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
93
+
94
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
95
96
97
## Setup
@@ -78,25 +115,36 @@ UI configuration requires paid Netdata Cloud plan.
115
116
#### Create an Azure monitoring principal
117
81
-Create a service principal or use a managed identity with the following permissions:
118
+The collector requires a service principal or managed identity with two Azure RBAC roles:
119
+
120
+| Role | Purpose |
121
+|:-----|:--------|
122
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
123
+| **Reader** | Query Azure Resource Graph for resource discovery |
124
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
125
+**Option A: Service principal**
126
86
-For service principal authentication:
127
```bash
88
-# Create the service principal
128
+# Create service principal with Monitoring Reader role
129
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
130
--scopes /subscriptions/<subscription-id>
131
132
+# Add the Reader role for resource discovery
133
+az role assignment create --assignee <appId-from-above> \
134
+ --role "Reader" --scope /subscriptions/<subscription-id>
135
+
136
# Note the appId (client_id), password (client_secret), and tenant
137
```
138
95
-For managed identity (on Azure VMs, VMSS, or AKS):
139
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
140
+
141
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
142
+# Assign both roles to the VM's managed identity
143
az role assignment create --assignee <managed-identity-principal-id> \
144
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
145
+
146
+az role assignment create --assignee <managed-identity-principal-id> \
147
+ --role "Reader" --scope /subscriptions/<subscription-id>
148
```
149
150
@@ -105,13 +153,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
153
154
#### Options
155
108
-The following options can be defined globally: update_every, autodetection_retry.
156
+The following options can be defined globally: `update_every`, `autodetection_retry`.
157
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
158
+**Profile file locations:**
159
114
-User profile files with the same filename override stock profiles.
160
+| Type | Path |
161
+|:-----|:-----|
162
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
163
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
164
+
165
+User profile files with the same `id` as a stock profile override it.
166
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
167
168
169
<details open><summary>Config options</summary>
@@ -122,25 +174,103 @@ User profile files with the same filename override stock profiles.
174
|:------|:-----|:------------|:--------|:---------:|
175
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
176
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
177
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
178
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
181
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
182
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
187
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
188
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
191
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
192
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
193
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
194
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
195
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
196
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
197
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
198
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
199
200
+<a id="option-collection-query-offset"></a>
201
+##### query_offset
202
+
203
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
204
+
205
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
206
+
207
+- **Default (180s)** works for most services.
208
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
209
+- **Increase to 240-300s** if you still see gaps or missing data points.
210
+- **Do not set below 60s** -- metrics will likely be incomplete.
211
+
212
+
213
+<a id="option-authentication-auth-mode"></a>
214
+##### auth.mode
215
+
216
+Determines how the collector authenticates with Azure.
217
+
218
+| Mode | When to use | Required options |
219
+|:-----|:------------|:-----------------|
220
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
221
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
222
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
223
+
224
+
225
+<a id="option-discovery-discovery-mode"></a>
226
+##### discovery.mode
227
+
228
+Controls how the collector finds candidate Azure resources.
229
+
230
+| Mode | Behavior |
231
+|:-----|:---------|
232
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
233
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
234
+
235
+
236
+<a id="option-discovery-discovery-mode-query-kql"></a>
237
+##### discovery.mode_query.kql
238
+
239
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
240
+
241
+The query **must** project these five columns:
242
+
243
+| Column | Description |
244
+|:-------|:------------|
245
+| `id` | Full Azure resource ID (ARM format) |
246
+| `name` | Resource name |
247
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
248
+| `resourceGroup` | Resource group name |
249
+| `location` | Azure region |
250
+
251
+Example:
252
+
253
+```
254
+resources
255
+| where tags.env =~ "prod"
256
+| project id, name, type, resourceGroup, location
257
+```
258
+
259
+
260
+<a id="option-profiles-profiles-mode"></a>
261
+##### profiles.mode
262
+
263
+Controls how the collector decides which metric profiles to activate.
264
+
265
+| Mode | Behavior |
266
+|:-----|:---------|
267
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
268
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
269
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
270
+
271
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
272
+
273
+
274
275
</details>
276
@@ -182,14 +312,28 @@ sudo ./edit-config go.d/azure_monitor.conf
312
313
##### Examples
314
185
-###### Service principal (auto-discover all resources)
315
+###### Service principal with structured discovery
316
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
317
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
318
319
```yaml
320
jobs:
321
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ subscription_ids:
323
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
324
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
325
+ discovery:
326
+ mode: filters
327
+ mode_filters:
328
+ resource_groups:
329
+ - production-rg
330
+ regions:
331
+ - eastus
332
+ tags:
333
+ env:
334
+ - prod
335
+ profiles:
336
+ mode: auto
337
auth:
338
mode: service_principal
339
mode_service_principal:
@@ -198,59 +342,49 @@ jobs:
342
client_secret: "your-client-secret"
343
344
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
345
+###### Managed identity with exact profiles
346
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
347
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
348
349
<details open><summary>Config</summary>
350
351
```yaml
352
jobs:
353
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
+ subscription_ids:
355
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
357
+ mode: exact
358
+ mode_exact:
359
+ names:
360
+ - sql_database
361
+ - postgres_flexible
362
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
363
+ mode: managed_identity
364
365
```
366
</details>
367
241
-###### Filter by resource group
368
+###### Custom Azure Resource Graph KQL
369
243
-Only monitor resources in specific resource groups.
370
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
371
372
<details open><summary>Config</summary>
373
374
```yaml
375
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
376
+ - name: prod-query
377
+ subscription_ids:
378
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
379
+ discovery:
380
+ mode: query
381
+ mode_query:
382
+ kql: |
383
+ resources
384
+ | where tags.env =~ "prod"
385
+ | project id, name, type, resourceGroup, location
386
+ profiles:
387
+ mode: auto
388
auth:
389
mode: default
390
@@ -259,14 +393,15 @@ jobs:
393
394
###### Azure Government cloud
395
262
-Connect to Azure Government cloud environment.
396
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
397
398
<details open><summary>Config</summary>
399
400
```yaml
401
jobs:
402
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
+ subscription_ids:
404
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
cloud: government
406
auth:
407
mode: service_principal
@@ -316,6 +451,7 @@ Labels:
451
| region | The Azure region where the resource is deployed. |
452
| resource_type | The Azure resource type identifier. |
453
| profile | The Azure Monitor profile id. |
454
+| subscription_id | The Azure subscription identifier. |
455
| resource_uid | The unique Azure resource identifier. |
456
457
Metrics:
@@ -401,31 +537,46 @@ docker logs netdata 2>&1 | grep azure_monitor
537
538
### No metrics are collected
539
404
-Verify the following:
405
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
406
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
407
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
408
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
540
+Check the following:
541
+
542
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
543
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
544
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
545
+- **Collector logs** -- Check for authentication or API errors:
546
+ ```bash
547
+ # systemd
548
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
549
+ # non-systemd
550
+ grep azure_monitor /var/log/netdata/collector.log
551
+ ```
552
553
554
### Missing metrics for some resource types
555
413
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
414
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
415
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
416
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
556
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
557
+
558
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
559
+- **Verify a built-in profile exists** -- List available profiles:
560
+ ```bash
561
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
562
+ ```
563
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
564
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
565
566
419
-### Metrics appear delayed
567
+### Charts have gaps or incomplete data
568
421
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
422
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
423
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
569
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
570
+
571
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
572
+- Slower time-grain batches automatically use a larger effective offset when needed.
573
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
574
575
576
### Authentication errors in sovereign clouds
577
578
For Azure Government or Azure China clouds, set the `cloud` parameter:
579
+
580
- Azure Government: `cloud: government`
581
- Azure China (21Vianet): `cloud: china`
582
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_expressroute_gateway.md
+238
-87
@@ -21,40 +21,77 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor ExpressRoute gateways including bits and packets per second for ingress and egress, connection counts, CPU utilization, active flow counts, and gateway scale unit counts.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure ExpressRoute Gateway with metrics covering:
31
+
32
+- **Throughput** -- gateway throughput, connection throughput (bits/s in/out), packets per second
33
+- **Compute** -- CPU utilization
34
+- **Flows** -- active flows, max flow creation rate
35
+- **Routes** -- routes advertised to peer, routes learned from peer, route changes
36
+- **Scale** -- VMs in VNet
37
+
38
+
39
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
40
41
42
This collector is supported on all platforms.
43
44
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
45
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
46
+The service principal or managed identity requires these Azure RBAC roles:
47
+
48
+| Role | Purpose | Scope |
49
+|:-----|:--------|:------|
50
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
51
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
52
53
54
### Default Behavior
55
56
#### Auto-Detection
57
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
58
+The collector has two discovery phases:
59
+
60
+**Bootstrap (first run)**
61
+
62
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
63
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
64
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
65
+- A single job can monitor multiple subscriptions.
66
+
67
+**Runtime (periodic refresh)**
68
+
69
+- Periodically re-discovers resources for **already-active profile types only**.
70
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
71
+
72
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
73
74
75
#### Limits
76
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
77
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
78
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
79
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
80
81
82
#### Performance Impact
83
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
84
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
85
+
86
+**Default concurrency and batching limits:**
87
+
88
+| Setting | Default | Description |
89
+|:--------|:--------|:------------|
90
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
91
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
92
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
93
+
94
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
95
96
97
## Setup
@@ -78,25 +115,36 @@ UI configuration requires paid Netdata Cloud plan.
115
116
#### Create an Azure monitoring principal
117
81
-Create a service principal or use a managed identity with the following permissions:
118
+The collector requires a service principal or managed identity with two Azure RBAC roles:
119
+
120
+| Role | Purpose |
121
+|:-----|:--------|
122
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
123
+| **Reader** | Query Azure Resource Graph for resource discovery |
124
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
125
+**Option A: Service principal**
126
86
-For service principal authentication:
127
```bash
88
-# Create the service principal
128
+# Create service principal with Monitoring Reader role
129
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
130
--scopes /subscriptions/<subscription-id>
131
132
+# Add the Reader role for resource discovery
133
+az role assignment create --assignee <appId-from-above> \
134
+ --role "Reader" --scope /subscriptions/<subscription-id>
135
+
136
# Note the appId (client_id), password (client_secret), and tenant
137
```
138
95
-For managed identity (on Azure VMs, VMSS, or AKS):
139
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
140
+
141
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
142
+# Assign both roles to the VM's managed identity
143
az role assignment create --assignee <managed-identity-principal-id> \
144
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
145
+
146
+az role assignment create --assignee <managed-identity-principal-id> \
147
+ --role "Reader" --scope /subscriptions/<subscription-id>
148
```
149
150
@@ -105,13 +153,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
153
154
#### Options
155
108
-The following options can be defined globally: update_every, autodetection_retry.
156
+The following options can be defined globally: `update_every`, `autodetection_retry`.
157
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
158
+**Profile file locations:**
159
114
-User profile files with the same filename override stock profiles.
160
+| Type | Path |
161
+|:-----|:-----|
162
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
163
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
164
+
165
+User profile files with the same `id` as a stock profile override it.
166
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
167
168
169
<details open><summary>Config options</summary>
@@ -122,25 +174,103 @@ User profile files with the same filename override stock profiles.
174
|:------|:-----|:------------|:--------|:---------:|
175
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
176
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
177
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
178
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
181
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
182
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
187
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
188
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
191
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
192
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
193
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
194
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
195
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
196
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
197
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
198
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
199
200
+<a id="option-collection-query-offset"></a>
201
+##### query_offset
202
+
203
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
204
+
205
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
206
+
207
+- **Default (180s)** works for most services.
208
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
209
+- **Increase to 240-300s** if you still see gaps or missing data points.
210
+- **Do not set below 60s** -- metrics will likely be incomplete.
211
+
212
+
213
+<a id="option-authentication-auth-mode"></a>
214
+##### auth.mode
215
+
216
+Determines how the collector authenticates with Azure.
217
+
218
+| Mode | When to use | Required options |
219
+|:-----|:------------|:-----------------|
220
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
221
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
222
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
223
+
224
+
225
+<a id="option-discovery-discovery-mode"></a>
226
+##### discovery.mode
227
+
228
+Controls how the collector finds candidate Azure resources.
229
+
230
+| Mode | Behavior |
231
+|:-----|:---------|
232
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
233
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
234
+
235
+
236
+<a id="option-discovery-discovery-mode-query-kql"></a>
237
+##### discovery.mode_query.kql
238
+
239
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
240
+
241
+The query **must** project these five columns:
242
+
243
+| Column | Description |
244
+|:-------|:------------|
245
+| `id` | Full Azure resource ID (ARM format) |
246
+| `name` | Resource name |
247
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
248
+| `resourceGroup` | Resource group name |
249
+| `location` | Azure region |
250
+
251
+Example:
252
+
253
+```
254
+resources
255
+| where tags.env =~ "prod"
256
+| project id, name, type, resourceGroup, location
257
+```
258
+
259
+
260
+<a id="option-profiles-profiles-mode"></a>
261
+##### profiles.mode
262
+
263
+Controls how the collector decides which metric profiles to activate.
264
+
265
+| Mode | Behavior |
266
+|:-----|:---------|
267
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
268
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
269
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
270
+
271
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
272
+
273
+
274
275
</details>
276
@@ -182,14 +312,28 @@ sudo ./edit-config go.d/azure_monitor.conf
312
313
##### Examples
314
185
-###### Service principal (auto-discover all resources)
315
+###### Service principal with structured discovery
316
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
317
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
318
319
```yaml
320
jobs:
321
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ subscription_ids:
323
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
324
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
325
+ discovery:
326
+ mode: filters
327
+ mode_filters:
328
+ resource_groups:
329
+ - production-rg
330
+ regions:
331
+ - eastus
332
+ tags:
333
+ env:
334
+ - prod
335
+ profiles:
336
+ mode: auto
337
auth:
338
mode: service_principal
339
mode_service_principal:
@@ -198,59 +342,49 @@ jobs:
342
client_secret: "your-client-secret"
343
344
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
345
+###### Managed identity with exact profiles
346
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
347
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
348
349
<details open><summary>Config</summary>
350
351
```yaml
352
jobs:
353
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
+ subscription_ids:
355
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
357
+ mode: exact
358
+ mode_exact:
359
+ names:
360
+ - sql_database
361
+ - postgres_flexible
362
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
363
+ mode: managed_identity
364
365
```
366
</details>
367
241
-###### Filter by resource group
368
+###### Custom Azure Resource Graph KQL
369
243
-Only monitor resources in specific resource groups.
370
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
371
372
<details open><summary>Config</summary>
373
374
```yaml
375
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
376
+ - name: prod-query
377
+ subscription_ids:
378
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
379
+ discovery:
380
+ mode: query
381
+ mode_query:
382
+ kql: |
383
+ resources
384
+ | where tags.env =~ "prod"
385
+ | project id, name, type, resourceGroup, location
386
+ profiles:
387
+ mode: auto
388
auth:
389
mode: default
390
@@ -259,14 +393,15 @@ jobs:
393
394
###### Azure Government cloud
395
262
-Connect to Azure Government cloud environment.
396
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
397
398
<details open><summary>Config</summary>
399
400
```yaml
401
jobs:
402
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
+ subscription_ids:
404
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
cloud: government
406
auth:
407
mode: service_principal
@@ -315,6 +450,7 @@ Labels:
450
| region | The Azure region where the resource is deployed. |
451
| resource_type | The Azure resource type identifier. |
452
| profile | The Azure Monitor profile id. |
453
+| subscription_id | The Azure subscription identifier. |
454
| resource_uid | The unique Azure resource identifier. |
455
456
Metrics:
@@ -403,31 +539,46 @@ docker logs netdata 2>&1 | grep azure_monitor
539
540
### No metrics are collected
541
406
-Verify the following:
407
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
408
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
409
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
410
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
542
+Check the following:
543
+
544
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
545
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
546
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
547
+- **Collector logs** -- Check for authentication or API errors:
548
+ ```bash
549
+ # systemd
550
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
551
+ # non-systemd
552
+ grep azure_monitor /var/log/netdata/collector.log
553
+ ```
554
555
556
### Missing metrics for some resource types
557
415
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
416
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
417
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
418
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
558
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
559
+
560
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
561
+- **Verify a built-in profile exists** -- List available profiles:
562
+ ```bash
563
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
564
+ ```
565
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
566
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
567
568
421
-### Metrics appear delayed
569
+### Charts have gaps or incomplete data
570
423
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
424
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
425
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
571
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
572
+
573
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
574
+- Slower time-grain batches automatically use a larger effective offset when needed.
575
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
576
577
578
### Authentication errors in sovereign clouds
579
580
For Azure Government or Azure China clouds, set the `cloud` parameter:
581
+
582
- Azure Government: `cloud: government`
583
- Azure China (21Vianet): `cloud: china`
584
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_firewall.md
+239
-87
@@ -21,40 +21,78 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Firewall including data processed, throughput, application and network rule hit counts, SNAT port utilization, health state percentage, and latency probes.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Firewall with metrics covering:
31
+
32
+- **Traffic** -- data processed, throughput (bits/s)
33
+- **Rules** -- application and network rule hit counts
34
+- **SNAT** -- SNAT port utilization
35
+- **Health** -- firewall health state percentage
36
+- **Latency** -- latency probe
37
+- **Capacity** -- observed capacity units
38
+
39
+
40
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
41
42
43
This collector is supported on all platforms.
44
45
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
46
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
47
+The service principal or managed identity requires these Azure RBAC roles:
48
+
49
+| Role | Purpose | Scope |
50
+|:-----|:--------|:------|
51
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
52
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
53
54
55
### Default Behavior
56
57
#### Auto-Detection
58
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
59
+The collector has two discovery phases:
60
+
61
+**Bootstrap (first run)**
62
+
63
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
64
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
65
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
66
+- A single job can monitor multiple subscriptions.
67
+
68
+**Runtime (periodic refresh)**
69
+
70
+- Periodically re-discovers resources for **already-active profile types only**.
71
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
72
+
73
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
74
75
76
#### Limits
77
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
78
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
79
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
80
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
81
82
83
#### Performance Impact
84
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
85
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
86
+
87
+**Default concurrency and batching limits:**
88
+
89
+| Setting | Default | Description |
90
+|:--------|:--------|:------------|
91
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
92
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
93
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
94
+
95
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
96
97
98
## Setup
@@ -78,25 +116,36 @@ UI configuration requires paid Netdata Cloud plan.
116
117
#### Create an Azure monitoring principal
118
81
-Create a service principal or use a managed identity with the following permissions:
119
+The collector requires a service principal or managed identity with two Azure RBAC roles:
120
+
121
+| Role | Purpose |
122
+|:-----|:--------|
123
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
124
+| **Reader** | Query Azure Resource Graph for resource discovery |
125
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
126
+**Option A: Service principal**
127
86
-For service principal authentication:
128
```bash
88
-# Create the service principal
129
+# Create service principal with Monitoring Reader role
130
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
131
--scopes /subscriptions/<subscription-id>
132
133
+# Add the Reader role for resource discovery
134
+az role assignment create --assignee <appId-from-above> \
135
+ --role "Reader" --scope /subscriptions/<subscription-id>
136
+
137
# Note the appId (client_id), password (client_secret), and tenant
138
```
139
95
-For managed identity (on Azure VMs, VMSS, or AKS):
140
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
141
+
142
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
143
+# Assign both roles to the VM's managed identity
144
az role assignment create --assignee <managed-identity-principal-id> \
145
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+
147
+az role assignment create --assignee <managed-identity-principal-id> \
148
+ --role "Reader" --scope /subscriptions/<subscription-id>
149
```
150
151
@@ -105,13 +154,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
154
155
#### Options
156
108
-The following options can be defined globally: update_every, autodetection_retry.
157
+The following options can be defined globally: `update_every`, `autodetection_retry`.
158
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+**Profile file locations:**
160
114
-User profile files with the same filename override stock profiles.
161
+| Type | Path |
162
+|:-----|:-----|
163
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
164
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
165
+
166
+User profile files with the same `id` as a stock profile override it.
167
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
168
169
170
<details open><summary>Config options</summary>
@@ -122,25 +175,103 @@ User profile files with the same filename override stock profiles.
175
|:------|:-----|:------------|:--------|:---------:|
176
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
177
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
178
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
179
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
183
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
186
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
187
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
188
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
189
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
192
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
193
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
194
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
195
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
196
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
197
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
198
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
199
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
200
201
+<a id="option-collection-query-offset"></a>
202
+##### query_offset
203
+
204
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
205
+
206
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
207
+
208
+- **Default (180s)** works for most services.
209
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
210
+- **Increase to 240-300s** if you still see gaps or missing data points.
211
+- **Do not set below 60s** -- metrics will likely be incomplete.
212
+
213
+
214
+<a id="option-authentication-auth-mode"></a>
215
+##### auth.mode
216
+
217
+Determines how the collector authenticates with Azure.
218
+
219
+| Mode | When to use | Required options |
220
+|:-----|:------------|:-----------------|
221
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
222
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
223
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
224
+
225
+
226
+<a id="option-discovery-discovery-mode"></a>
227
+##### discovery.mode
228
+
229
+Controls how the collector finds candidate Azure resources.
230
+
231
+| Mode | Behavior |
232
+|:-----|:---------|
233
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
234
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
235
+
236
+
237
+<a id="option-discovery-discovery-mode-query-kql"></a>
238
+##### discovery.mode_query.kql
239
+
240
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
241
+
242
+The query **must** project these five columns:
243
+
244
+| Column | Description |
245
+|:-------|:------------|
246
+| `id` | Full Azure resource ID (ARM format) |
247
+| `name` | Resource name |
248
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
249
+| `resourceGroup` | Resource group name |
250
+| `location` | Azure region |
251
+
252
+Example:
253
+
254
+```
255
+resources
256
+| where tags.env =~ "prod"
257
+| project id, name, type, resourceGroup, location
258
+```
259
+
260
+
261
+<a id="option-profiles-profiles-mode"></a>
262
+##### profiles.mode
263
+
264
+Controls how the collector decides which metric profiles to activate.
265
+
266
+| Mode | Behavior |
267
+|:-----|:---------|
268
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
269
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
270
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
271
+
272
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
273
+
274
+
275
276
</details>
277
@@ -182,14 +313,28 @@ sudo ./edit-config go.d/azure_monitor.conf
313
314
##### Examples
315
185
-###### Service principal (auto-discover all resources)
316
+###### Service principal with structured discovery
317
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
318
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
319
320
```yaml
321
jobs:
322
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ subscription_ids:
324
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
325
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
326
+ discovery:
327
+ mode: filters
328
+ mode_filters:
329
+ resource_groups:
330
+ - production-rg
331
+ regions:
332
+ - eastus
333
+ tags:
334
+ env:
335
+ - prod
336
+ profiles:
337
+ mode: auto
338
auth:
339
mode: service_principal
340
mode_service_principal:
@@ -198,59 +343,49 @@ jobs:
343
client_secret: "your-client-secret"
344
345
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
346
+###### Managed identity with exact profiles
347
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
348
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
349
350
<details open><summary>Config</summary>
351
352
```yaml
353
jobs:
354
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
+ subscription_ids:
356
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
358
+ mode: exact
359
+ mode_exact:
360
+ names:
361
+ - sql_database
362
+ - postgres_flexible
363
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
364
+ mode: managed_identity
365
366
```
367
</details>
368
241
-###### Filter by resource group
369
+###### Custom Azure Resource Graph KQL
370
243
-Only monitor resources in specific resource groups.
371
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
372
373
<details open><summary>Config</summary>
374
375
```yaml
376
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
377
+ - name: prod-query
378
+ subscription_ids:
379
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
380
+ discovery:
381
+ mode: query
382
+ mode_query:
383
+ kql: |
384
+ resources
385
+ | where tags.env =~ "prod"
386
+ | project id, name, type, resourceGroup, location
387
+ profiles:
388
+ mode: auto
389
auth:
390
mode: default
391
@@ -259,14 +394,15 @@ jobs:
394
395
###### Azure Government cloud
396
262
-Connect to Azure Government cloud environment.
397
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
398
399
<details open><summary>Config</summary>
400
401
```yaml
402
jobs:
403
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
+ subscription_ids:
405
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
cloud: government
407
auth:
408
mode: service_principal
@@ -313,6 +449,7 @@ Labels:
449
| region | The Azure region where the resource is deployed. |
450
| resource_type | The Azure resource type identifier. |
451
| profile | The Azure Monitor profile id. |
452
+| subscription_id | The Azure subscription identifier. |
453
| resource_uid | The unique Azure resource identifier. |
454
455
Metrics:
@@ -398,31 +535,46 @@ docker logs netdata 2>&1 | grep azure_monitor
535
536
### No metrics are collected
537
401
-Verify the following:
402
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
403
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
404
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
405
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
538
+Check the following:
539
+
540
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
541
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
542
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
543
+- **Collector logs** -- Check for authentication or API errors:
544
+ ```bash
545
+ # systemd
546
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
547
+ # non-systemd
548
+ grep azure_monitor /var/log/netdata/collector.log
549
+ ```
550
551
552
### Missing metrics for some resource types
553
410
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
411
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
412
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
413
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
554
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
555
+
556
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
557
+- **Verify a built-in profile exists** -- List available profiles:
558
+ ```bash
559
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
560
+ ```
561
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
562
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
563
564
416
-### Metrics appear delayed
565
+### Charts have gaps or incomplete data
566
418
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
419
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
420
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
567
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
568
+
569
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
570
+- Slower time-grain batches automatically use a larger effective offset when needed.
571
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
572
573
574
### Authentication errors in sovereign clouds
575
576
For Azure Government or Azure China clouds, set the `cloud` parameter:
577
+
578
- Azure Government: `cloud: government`
579
- Azure China (21Vianet): `cloud: china`
580
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_front_door.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Front Door including request counts and rates, response sizes, total latency, origin health probe percentages, origin request counts, origin latency, WAF request counts by action and rule, and WebSocket connection metrics.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Front Door with metrics covering:
31
+
32
+- **Requests** -- client and origin request rates, origin shield requests (to shield/to origin/rate limited)
33
+- **Latency** -- total latency, origin latency
34
+- **Data transfer** -- request and response data transfer, origin shield data transfer, byte hit ratio
35
+- **Errors** -- 4xx and 5xx error rates
36
+- **Origin** -- origin health probe percentage
37
+- **WAF** -- WAF requests, challenges (CAPTCHA/JS challenge)
38
+- **WebSocket** -- WebSocket connections (requested/active), connection duration
39
+
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
+
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -317,6 +454,7 @@ Labels:
454
| region | The Azure region where the resource is deployed. |
455
| resource_type | The Azure resource type identifier. |
456
| profile | The Azure Monitor profile id. |
457
+| subscription_id | The Azure subscription identifier. |
458
| resource_uid | The unique Azure resource identifier. |
459
460
Metrics:
@@ -407,31 +545,46 @@ docker logs netdata 2>&1 | grep azure_monitor
545
546
### No metrics are collected
547
410
-Verify the following:
411
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
412
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
413
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
414
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
548
+Check the following:
549
+
550
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
551
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
552
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
553
+- **Collector logs** -- Check for authentication or API errors:
554
+ ```bash
555
+ # systemd
556
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
557
+ # non-systemd
558
+ grep azure_monitor /var/log/netdata/collector.log
559
+ ```
560
561
562
### Missing metrics for some resource types
563
419
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
420
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
421
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
422
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
564
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
565
+
566
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
567
+- **Verify a built-in profile exists** -- List available profiles:
568
+ ```bash
569
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
570
+ ```
571
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
572
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
573
574
425
-### Metrics appear delayed
575
+### Charts have gaps or incomplete data
576
427
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
428
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
429
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
577
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
578
+
579
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
580
+- Slower time-grain batches automatically use a larger effective offset when needed.
581
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
582
583
584
### Authentication errors in sovereign clouds
585
586
For Azure Government or Azure China clouds, set the `cloud` parameter:
587
+
588
- Azure Government: `cloud: government`
589
- Azure China (21Vianet): `cloud: china`
590
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_functions.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Functions execution including function invocation counts, execution units (MB-milliseconds), HTTP request rates and response codes, CPU and memory consumption, and Flex Consumption plan metrics for always-ready and on-demand instances. Uses the same underlying metrics as App Service since Azure Functions runs on the App Service platform.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
+
28
+:::
29
+
30
+Monitor Azure Functions with metrics covering:
31
+
32
+- **Executions** -- function invocation counts, execution units (MB-milliseconds)
33
+- **HTTP** -- request rates, response status codes
34
+- **Compute** -- CPU utilization, CPU time consumed
35
+- **Memory** -- memory usage (working set, private bytes)
36
+- **Flex Consumption** -- always-ready and on-demand function executions and units
37
+
38
+Azure Functions runs on the App Service platform and shares the same underlying metrics.
39
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
+
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -306,6 +443,7 @@ Labels:
443
| region | The Azure region where the resource is deployed. |
444
| resource_type | The Azure resource type identifier. |
445
| profile | The Azure Monitor profile id. |
446
+| subscription_id | The Azure subscription identifier. |
447
| resource_uid | The unique Azure resource identifier. |
448
449
Metrics:
@@ -410,31 +548,46 @@ docker logs netdata 2>&1 | grep azure_monitor
548
549
### No metrics are collected
550
413
-Verify the following:
414
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
415
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
416
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
417
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
551
+Check the following:
552
+
553
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
554
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
555
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
556
+- **Collector logs** -- Check for authentication or API errors:
557
+ ```bash
558
+ # systemd
559
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
560
+ # non-systemd
561
+ grep azure_monitor /var/log/netdata/collector.log
562
+ ```
563
564
565
### Missing metrics for some resource types
566
422
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
423
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
424
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
425
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
567
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
568
+
569
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
570
+- **Verify a built-in profile exists** -- List available profiles:
571
+ ```bash
572
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
573
+ ```
574
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
575
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
576
577
428
-### Metrics appear delayed
578
+### Charts have gaps or incomplete data
579
430
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
431
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
432
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
580
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
581
+
582
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
583
+- Slower time-grain batches automatically use a larger effective offset when needed.
584
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
585
586
587
### Authentication errors in sovereign clouds
588
589
For Azure Government or Azure China clouds, set the `cloud` parameter:
590
+
591
- Azure Government: `cloud: government`
592
- Azure China (21Vianet): `cloud: china`
593
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_iot_hub.md
+242
-87
@@ -21,40 +21,81 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor IoT Hub including device telemetry message rates and quota usage, routing delivery and latency, device twin read and write operations, direct method invocations, cloud-to-device messaging and feedback, job completion rates, device connection and authentication events, and event grid publish status.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure IoT Hub with metrics covering:
31
+
32
+- **Telemetry** -- device telemetry messages (attempted/sent), throttling errors, daily message quota usage
33
+- **Routing** -- message deliveries by endpoint (Event Hubs, Service Bus, storage), routing latency, delivery status
34
+- **Device twins** -- backend and device twin reads/writes (successful/failed), query results
35
+- **Direct methods** -- method invocations (successful/failed), request/response sizes
36
+- **Cloud-to-device** -- C2D commands (completed/abandoned/rejected), expired messages
37
+- **Jobs** -- job completions, cancellations, list calls, twin update and method job creations
38
+- **Connections** -- successful connections, connected devices, total devices
39
+- **Event Grid** -- Event Grid deliveries, Event Grid latency
40
+- **Data** -- device data usage
41
+
42
+
43
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
44
45
46
This collector is supported on all platforms.
47
48
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
49
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
50
+The service principal or managed identity requires these Azure RBAC roles:
51
+
52
+| Role | Purpose | Scope |
53
+|:-----|:--------|:------|
54
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
55
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
56
57
58
### Default Behavior
59
60
#### Auto-Detection
61
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
62
+The collector has two discovery phases:
63
+
64
+**Bootstrap (first run)**
65
+
66
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
67
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
68
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
69
+- A single job can monitor multiple subscriptions.
70
+
71
+**Runtime (periodic refresh)**
72
+
73
+- Periodically re-discovers resources for **already-active profile types only**.
74
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
75
+
76
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
77
78
79
#### Limits
80
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
81
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
82
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
83
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
84
85
86
#### Performance Impact
87
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
88
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
89
+
90
+**Default concurrency and batching limits:**
91
+
92
+| Setting | Default | Description |
93
+|:--------|:--------|:------------|
94
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
95
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
96
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
97
+
98
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
99
100
101
## Setup
@@ -78,25 +119,36 @@ UI configuration requires paid Netdata Cloud plan.
119
120
#### Create an Azure monitoring principal
121
81
-Create a service principal or use a managed identity with the following permissions:
122
+The collector requires a service principal or managed identity with two Azure RBAC roles:
123
+
124
+| Role | Purpose |
125
+|:-----|:--------|
126
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
127
+| **Reader** | Query Azure Resource Graph for resource discovery |
128
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
129
+**Option A: Service principal**
130
86
-For service principal authentication:
131
```bash
88
-# Create the service principal
132
+# Create service principal with Monitoring Reader role
133
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
134
--scopes /subscriptions/<subscription-id>
135
136
+# Add the Reader role for resource discovery
137
+az role assignment create --assignee <appId-from-above> \
138
+ --role "Reader" --scope /subscriptions/<subscription-id>
139
+
140
# Note the appId (client_id), password (client_secret), and tenant
141
```
142
95
-For managed identity (on Azure VMs, VMSS, or AKS):
143
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
144
+
145
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
146
+# Assign both roles to the VM's managed identity
147
az role assignment create --assignee <managed-identity-principal-id> \
148
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Reader" --scope /subscriptions/<subscription-id>
152
```
153
154
@@ -105,13 +157,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
157
158
#### Options
159
108
-The following options can be defined globally: update_every, autodetection_retry.
160
+The following options can be defined globally: `update_every`, `autodetection_retry`.
161
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+**Profile file locations:**
163
114
-User profile files with the same filename override stock profiles.
164
+| Type | Path |
165
+|:-----|:-----|
166
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
167
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
168
+
169
+User profile files with the same `id` as a stock profile override it.
170
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
171
172
173
<details open><summary>Config options</summary>
@@ -122,25 +178,103 @@ User profile files with the same filename override stock profiles.
178
|:------|:-----|:------------|:--------|:---------:|
179
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
180
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
181
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
182
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
186
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
191
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
192
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
195
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
196
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
197
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
198
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
199
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
204
+<a id="option-collection-query-offset"></a>
205
+##### query_offset
206
+
207
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
208
+
209
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
210
+
211
+- **Default (180s)** works for most services.
212
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
213
+- **Increase to 240-300s** if you still see gaps or missing data points.
214
+- **Do not set below 60s** -- metrics will likely be incomplete.
215
+
216
+
217
+<a id="option-authentication-auth-mode"></a>
218
+##### auth.mode
219
+
220
+Determines how the collector authenticates with Azure.
221
+
222
+| Mode | When to use | Required options |
223
+|:-----|:------------|:-----------------|
224
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
225
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
226
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
227
+
228
+
229
+<a id="option-discovery-discovery-mode"></a>
230
+##### discovery.mode
231
+
232
+Controls how the collector finds candidate Azure resources.
233
+
234
+| Mode | Behavior |
235
+|:-----|:---------|
236
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
237
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
238
+
239
+
240
+<a id="option-discovery-discovery-mode-query-kql"></a>
241
+##### discovery.mode_query.kql
242
+
243
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
244
+
245
+The query **must** project these five columns:
246
+
247
+| Column | Description |
248
+|:-------|:------------|
249
+| `id` | Full Azure resource ID (ARM format) |
250
+| `name` | Resource name |
251
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
252
+| `resourceGroup` | Resource group name |
253
+| `location` | Azure region |
254
+
255
+Example:
256
+
257
+```
258
+resources
259
+| where tags.env =~ "prod"
260
+| project id, name, type, resourceGroup, location
261
+```
262
+
263
+
264
+<a id="option-profiles-profiles-mode"></a>
265
+##### profiles.mode
266
+
267
+Controls how the collector decides which metric profiles to activate.
268
+
269
+| Mode | Behavior |
270
+|:-----|:---------|
271
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
272
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
273
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
274
+
275
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
276
+
277
+
278
279
</details>
280
@@ -182,14 +316,28 @@ sudo ./edit-config go.d/azure_monitor.conf
316
317
##### Examples
318
185
-###### Service principal (auto-discover all resources)
319
+###### Service principal with structured discovery
320
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
321
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
322
323
```yaml
324
jobs:
325
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ subscription_ids:
327
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
328
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
329
+ discovery:
330
+ mode: filters
331
+ mode_filters:
332
+ resource_groups:
333
+ - production-rg
334
+ regions:
335
+ - eastus
336
+ tags:
337
+ env:
338
+ - prod
339
+ profiles:
340
+ mode: auto
341
auth:
342
mode: service_principal
343
mode_service_principal:
@@ -198,59 +346,49 @@ jobs:
346
client_secret: "your-client-secret"
347
348
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
349
+###### Managed identity with exact profiles
350
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
351
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
352
353
<details open><summary>Config</summary>
354
355
```yaml
356
jobs:
357
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ subscription_ids:
359
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
361
+ mode: exact
362
+ mode_exact:
363
+ names:
364
+ - sql_database
365
+ - postgres_flexible
366
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
367
+ mode: managed_identity
368
369
```
370
</details>
371
241
-###### Filter by resource group
372
+###### Custom Azure Resource Graph KQL
373
243
-Only monitor resources in specific resource groups.
374
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
375
376
<details open><summary>Config</summary>
377
378
```yaml
379
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
380
+ - name: prod-query
381
+ subscription_ids:
382
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
383
+ discovery:
384
+ mode: query
385
+ mode_query:
386
+ kql: |
387
+ resources
388
+ | where tags.env =~ "prod"
389
+ | project id, name, type, resourceGroup, location
390
+ profiles:
391
+ mode: auto
392
auth:
393
mode: default
394
@@ -259,14 +397,15 @@ jobs:
397
398
###### Azure Government cloud
399
262
-Connect to Azure Government cloud environment.
400
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
401
402
<details open><summary>Config</summary>
403
404
```yaml
405
jobs:
406
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
+ subscription_ids:
408
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
409
cloud: government
410
auth:
411
mode: service_principal
@@ -327,6 +466,7 @@ Labels:
466
| region | The Azure region where the resource is deployed. |
467
| resource_type | The Azure resource type identifier. |
468
| profile | The Azure Monitor profile id. |
469
+| subscription_id | The Azure subscription identifier. |
470
| resource_uid | The unique Azure resource identifier. |
471
472
Metrics:
@@ -444,31 +584,46 @@ docker logs netdata 2>&1 | grep azure_monitor
584
585
### No metrics are collected
586
447
-Verify the following:
448
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
449
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
450
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
451
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
587
+Check the following:
588
+
589
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
590
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
591
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
592
+- **Collector logs** -- Check for authentication or API errors:
593
+ ```bash
594
+ # systemd
595
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
596
+ # non-systemd
597
+ grep azure_monitor /var/log/netdata/collector.log
598
+ ```
599
600
601
### Missing metrics for some resource types
602
456
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
457
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
458
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
459
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
603
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
604
+
605
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
606
+- **Verify a built-in profile exists** -- List available profiles:
607
+ ```bash
608
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
609
+ ```
610
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
611
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
612
613
462
-### Metrics appear delayed
614
+### Charts have gaps or incomplete data
615
464
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
465
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
466
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
616
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
617
+
618
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
619
+- Slower time-grain batches automatically use a larger effective offset when needed.
620
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
621
622
623
### Authentication errors in sovereign clouds
624
625
For Azure Government or Azure China clouds, set the `cloud` parameter:
626
+
627
- Azure Government: `cloud: government`
628
- Azure China (21Vianet): `cloud: china`
629
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_key_vault.md
+236
-87
@@ -21,40 +21,75 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Key Vault including overall vault availability, API saturation approaching service limits, and service API hit and latency metrics.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Key Vault with metrics covering:
31
+
32
+- **Availability** -- overall vault availability percentage
33
+- **API** -- API activity (hits/results), API latency
34
+- **Saturation** -- API saturation approaching service limits
35
+
36
+
37
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
38
39
40
This collector is supported on all platforms.
41
42
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
43
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
44
+The service principal or managed identity requires these Azure RBAC roles:
45
+
46
+| Role | Purpose | Scope |
47
+|:-----|:--------|:------|
48
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
49
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
50
51
52
### Default Behavior
53
54
#### Auto-Detection
55
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
56
+The collector has two discovery phases:
57
+
58
+**Bootstrap (first run)**
59
+
60
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
61
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
62
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
63
+- A single job can monitor multiple subscriptions.
64
+
65
+**Runtime (periodic refresh)**
66
+
67
+- Periodically re-discovers resources for **already-active profile types only**.
68
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
69
+
70
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
71
72
73
#### Limits
74
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
75
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
76
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
77
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
78
79
80
#### Performance Impact
81
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
82
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
83
+
84
+**Default concurrency and batching limits:**
85
+
86
+| Setting | Default | Description |
87
+|:--------|:--------|:------------|
88
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
89
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
90
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
91
+
92
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
93
94
95
## Setup
@@ -78,25 +113,36 @@ UI configuration requires paid Netdata Cloud plan.
113
114
#### Create an Azure monitoring principal
115
81
-Create a service principal or use a managed identity with the following permissions:
116
+The collector requires a service principal or managed identity with two Azure RBAC roles:
117
+
118
+| Role | Purpose |
119
+|:-----|:--------|
120
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
121
+| **Reader** | Query Azure Resource Graph for resource discovery |
122
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
123
+**Option A: Service principal**
124
86
-For service principal authentication:
125
```bash
88
-# Create the service principal
126
+# Create service principal with Monitoring Reader role
127
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
128
--scopes /subscriptions/<subscription-id>
129
130
+# Add the Reader role for resource discovery
131
+az role assignment create --assignee <appId-from-above> \
132
+ --role "Reader" --scope /subscriptions/<subscription-id>
133
+
134
# Note the appId (client_id), password (client_secret), and tenant
135
```
136
95
-For managed identity (on Azure VMs, VMSS, or AKS):
137
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
138
+
139
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
140
+# Assign both roles to the VM's managed identity
141
az role assignment create --assignee <managed-identity-principal-id> \
142
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
143
+
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Reader" --scope /subscriptions/<subscription-id>
146
```
147
148
@@ -105,13 +151,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
151
152
#### Options
153
108
-The following options can be defined globally: update_every, autodetection_retry.
154
+The following options can be defined globally: `update_every`, `autodetection_retry`.
155
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
156
+**Profile file locations:**
157
114
-User profile files with the same filename override stock profiles.
158
+| Type | Path |
159
+|:-----|:-----|
160
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
161
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
162
+
163
+User profile files with the same `id` as a stock profile override it.
164
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
165
166
167
<details open><summary>Config options</summary>
@@ -122,25 +172,103 @@ User profile files with the same filename override stock profiles.
172
|:------|:-----|:------------|:--------|:---------:|
173
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
174
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
175
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
176
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
177
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
179
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
180
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
181
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
183
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
184
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
185
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
186
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
187
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
189
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
190
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
191
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
192
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
193
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
194
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
195
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
196
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
197
198
+<a id="option-collection-query-offset"></a>
199
+##### query_offset
200
+
201
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
202
+
203
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
204
+
205
+- **Default (180s)** works for most services.
206
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
207
+- **Increase to 240-300s** if you still see gaps or missing data points.
208
+- **Do not set below 60s** -- metrics will likely be incomplete.
209
+
210
+
211
+<a id="option-authentication-auth-mode"></a>
212
+##### auth.mode
213
+
214
+Determines how the collector authenticates with Azure.
215
+
216
+| Mode | When to use | Required options |
217
+|:-----|:------------|:-----------------|
218
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
219
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
220
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
221
+
222
+
223
+<a id="option-discovery-discovery-mode"></a>
224
+##### discovery.mode
225
+
226
+Controls how the collector finds candidate Azure resources.
227
+
228
+| Mode | Behavior |
229
+|:-----|:---------|
230
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
231
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
232
+
233
+
234
+<a id="option-discovery-discovery-mode-query-kql"></a>
235
+##### discovery.mode_query.kql
236
+
237
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
238
+
239
+The query **must** project these five columns:
240
+
241
+| Column | Description |
242
+|:-------|:------------|
243
+| `id` | Full Azure resource ID (ARM format) |
244
+| `name` | Resource name |
245
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
246
+| `resourceGroup` | Resource group name |
247
+| `location` | Azure region |
248
+
249
+Example:
250
+
251
+```
252
+resources
253
+| where tags.env =~ "prod"
254
+| project id, name, type, resourceGroup, location
255
+```
256
+
257
+
258
+<a id="option-profiles-profiles-mode"></a>
259
+##### profiles.mode
260
+
261
+Controls how the collector decides which metric profiles to activate.
262
+
263
+| Mode | Behavior |
264
+|:-----|:---------|
265
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
266
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
267
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
268
+
269
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
270
+
271
+
272
273
</details>
274
@@ -182,14 +310,28 @@ sudo ./edit-config go.d/azure_monitor.conf
310
311
##### Examples
312
185
-###### Service principal (auto-discover all resources)
313
+###### Service principal with structured discovery
314
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
315
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
316
317
```yaml
318
jobs:
319
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ subscription_ids:
321
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
322
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
323
+ discovery:
324
+ mode: filters
325
+ mode_filters:
326
+ resource_groups:
327
+ - production-rg
328
+ regions:
329
+ - eastus
330
+ tags:
331
+ env:
332
+ - prod
333
+ profiles:
334
+ mode: auto
335
auth:
336
mode: service_principal
337
mode_service_principal:
@@ -198,59 +340,49 @@ jobs:
340
client_secret: "your-client-secret"
341
342
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
343
+###### Managed identity with exact profiles
344
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
345
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
346
347
<details open><summary>Config</summary>
348
349
```yaml
350
jobs:
351
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
352
+ subscription_ids:
353
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
355
+ mode: exact
356
+ mode_exact:
357
+ names:
358
+ - sql_database
359
+ - postgres_flexible
360
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
361
+ mode: managed_identity
362
363
```
364
</details>
365
241
-###### Filter by resource group
366
+###### Custom Azure Resource Graph KQL
367
243
-Only monitor resources in specific resource groups.
368
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
369
370
<details open><summary>Config</summary>
371
372
```yaml
373
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
374
+ - name: prod-query
375
+ subscription_ids:
376
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
377
+ discovery:
378
+ mode: query
379
+ mode_query:
380
+ kql: |
381
+ resources
382
+ | where tags.env =~ "prod"
383
+ | project id, name, type, resourceGroup, location
384
+ profiles:
385
+ mode: auto
386
auth:
387
mode: default
388
@@ -259,14 +391,15 @@ jobs:
391
392
###### Azure Government cloud
393
262
-Connect to Azure Government cloud environment.
394
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
395
396
<details open><summary>Config</summary>
397
398
```yaml
399
jobs:
400
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
401
+ subscription_ids:
402
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
cloud: government
404
auth:
405
mode: service_principal
@@ -313,6 +446,7 @@ Labels:
446
| region | The Azure region where the resource is deployed. |
447
| resource_type | The Azure resource type identifier. |
448
| profile | The Azure Monitor profile id. |
449
+| subscription_id | The Azure subscription identifier. |
450
| resource_uid | The unique Azure resource identifier. |
451
452
Metrics:
@@ -395,31 +529,46 @@ docker logs netdata 2>&1 | grep azure_monitor
529
530
### No metrics are collected
531
398
-Verify the following:
399
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
400
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
401
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
402
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
532
+Check the following:
533
+
534
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
535
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
536
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
537
+- **Collector logs** -- Check for authentication or API errors:
538
+ ```bash
539
+ # systemd
540
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
541
+ # non-systemd
542
+ grep azure_monitor /var/log/netdata/collector.log
543
+ ```
544
545
546
### Missing metrics for some resource types
547
407
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
408
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
409
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
410
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
548
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
549
+
550
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
551
+- **Verify a built-in profile exists** -- List available profiles:
552
+ ```bash
553
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
554
+ ```
555
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
556
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
557
558
413
-### Metrics appear delayed
559
+### Charts have gaps or incomplete data
560
415
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
416
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
417
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
561
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
562
+
563
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
564
+- Slower time-grain batches automatically use a larger effective offset when needed.
565
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
566
567
568
### Authentication errors in sovereign clouds
569
570
For Azure Government or Azure China clouds, set the `cloud` parameter:
571
+
572
- Azure Government: `cloud: government`
573
- Azure China (21Vianet): `cloud: china`
574
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_kubernetes_service_cluster.md
+237
-87
@@ -21,40 +21,76 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor AKS cluster health including API server and etcd resource usage, pod scheduling status and readiness, node capacity and conditions, cluster autoscaler behavior, and per-node CPU, memory, disk, and network utilization.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Kubernetes Service (AKS) with metrics covering:
31
+
32
+- **Control plane** -- API server CPU/memory, etcd CPU/memory/database utilization, inflight requests
33
+- **Nodes** -- allocatable CPU/memory, per-node CPU (millicores and %), memory (RSS, working set), disk usage, network traffic
34
+- **Pods** -- pods by phase, pods in ready state, node conditions
35
+- **Autoscaler** -- autoscaler health, unneeded nodes, unschedulable pods
36
+
37
+
38
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
39
40
41
This collector is supported on all platforms.
42
43
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
44
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
45
+The service principal or managed identity requires these Azure RBAC roles:
46
+
47
+| Role | Purpose | Scope |
48
+|:-----|:--------|:------|
49
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
50
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
51
52
53
### Default Behavior
54
55
#### Auto-Detection
56
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
57
+The collector has two discovery phases:
58
+
59
+**Bootstrap (first run)**
60
+
61
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
62
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
63
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
64
+- A single job can monitor multiple subscriptions.
65
+
66
+**Runtime (periodic refresh)**
67
+
68
+- Periodically re-discovers resources for **already-active profile types only**.
69
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
70
+
71
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
72
73
74
#### Limits
75
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
76
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
77
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
78
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
79
80
81
#### Performance Impact
82
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
83
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
84
+
85
+**Default concurrency and batching limits:**
86
+
87
+| Setting | Default | Description |
88
+|:--------|:--------|:------------|
89
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
90
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
91
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
92
+
93
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
94
95
96
## Setup
@@ -78,25 +114,36 @@ UI configuration requires paid Netdata Cloud plan.
114
115
#### Create an Azure monitoring principal
116
81
-Create a service principal or use a managed identity with the following permissions:
117
+The collector requires a service principal or managed identity with two Azure RBAC roles:
118
+
119
+| Role | Purpose |
120
+|:-----|:--------|
121
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
122
+| **Reader** | Query Azure Resource Graph for resource discovery |
123
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
124
+**Option A: Service principal**
125
86
-For service principal authentication:
126
```bash
88
-# Create the service principal
127
+# Create service principal with Monitoring Reader role
128
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
129
--scopes /subscriptions/<subscription-id>
130
131
+# Add the Reader role for resource discovery
132
+az role assignment create --assignee <appId-from-above> \
133
+ --role "Reader" --scope /subscriptions/<subscription-id>
134
+
135
# Note the appId (client_id), password (client_secret), and tenant
136
```
137
95
-For managed identity (on Azure VMs, VMSS, or AKS):
138
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
139
+
140
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
141
+# Assign both roles to the VM's managed identity
142
az role assignment create --assignee <managed-identity-principal-id> \
143
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
144
+
145
+az role assignment create --assignee <managed-identity-principal-id> \
146
+ --role "Reader" --scope /subscriptions/<subscription-id>
147
```
148
149
@@ -105,13 +152,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
152
153
#### Options
154
108
-The following options can be defined globally: update_every, autodetection_retry.
155
+The following options can be defined globally: `update_every`, `autodetection_retry`.
156
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
157
+**Profile file locations:**
158
114
-User profile files with the same filename override stock profiles.
159
+| Type | Path |
160
+|:-----|:-----|
161
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
162
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
163
+
164
+User profile files with the same `id` as a stock profile override it.
165
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
166
167
168
<details open><summary>Config options</summary>
@@ -122,25 +173,103 @@ User profile files with the same filename override stock profiles.
173
|:------|:-----|:------------|:--------|:---------:|
174
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
176
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
177
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
180
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
181
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
184
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
185
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
186
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
187
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
190
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
191
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
192
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
193
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
194
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
195
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
196
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
197
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
198
199
+<a id="option-collection-query-offset"></a>
200
+##### query_offset
201
+
202
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
203
+
204
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
205
+
206
+- **Default (180s)** works for most services.
207
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
208
+- **Increase to 240-300s** if you still see gaps or missing data points.
209
+- **Do not set below 60s** -- metrics will likely be incomplete.
210
+
211
+
212
+<a id="option-authentication-auth-mode"></a>
213
+##### auth.mode
214
+
215
+Determines how the collector authenticates with Azure.
216
+
217
+| Mode | When to use | Required options |
218
+|:-----|:------------|:-----------------|
219
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
220
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
221
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
222
+
223
+
224
+<a id="option-discovery-discovery-mode"></a>
225
+##### discovery.mode
226
+
227
+Controls how the collector finds candidate Azure resources.
228
+
229
+| Mode | Behavior |
230
+|:-----|:---------|
231
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
232
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
233
+
234
+
235
+<a id="option-discovery-discovery-mode-query-kql"></a>
236
+##### discovery.mode_query.kql
237
+
238
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
239
+
240
+The query **must** project these five columns:
241
+
242
+| Column | Description |
243
+|:-------|:------------|
244
+| `id` | Full Azure resource ID (ARM format) |
245
+| `name` | Resource name |
246
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
247
+| `resourceGroup` | Resource group name |
248
+| `location` | Azure region |
249
+
250
+Example:
251
+
252
+```
253
+resources
254
+| where tags.env =~ "prod"
255
+| project id, name, type, resourceGroup, location
256
+```
257
+
258
+
259
+<a id="option-profiles-profiles-mode"></a>
260
+##### profiles.mode
261
+
262
+Controls how the collector decides which metric profiles to activate.
263
+
264
+| Mode | Behavior |
265
+|:-----|:---------|
266
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
267
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
268
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
269
+
270
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
271
+
272
+
273
274
</details>
275
@@ -182,14 +311,28 @@ sudo ./edit-config go.d/azure_monitor.conf
311
312
##### Examples
313
185
-###### Service principal (auto-discover all resources)
314
+###### Service principal with structured discovery
315
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
316
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
317
318
```yaml
319
jobs:
320
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ subscription_ids:
322
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
323
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
324
+ discovery:
325
+ mode: filters
326
+ mode_filters:
327
+ resource_groups:
328
+ - production-rg
329
+ regions:
330
+ - eastus
331
+ tags:
332
+ env:
333
+ - prod
334
+ profiles:
335
+ mode: auto
336
auth:
337
mode: service_principal
338
mode_service_principal:
@@ -198,59 +341,49 @@ jobs:
341
client_secret: "your-client-secret"
342
343
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
344
+###### Managed identity with exact profiles
345
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
346
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
347
348
<details open><summary>Config</summary>
349
350
```yaml
351
jobs:
352
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ subscription_ids:
354
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
356
+ mode: exact
357
+ mode_exact:
358
+ names:
359
+ - sql_database
360
+ - postgres_flexible
361
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
362
+ mode: managed_identity
363
364
```
365
</details>
366
241
-###### Filter by resource group
367
+###### Custom Azure Resource Graph KQL
368
243
-Only monitor resources in specific resource groups.
369
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
370
371
<details open><summary>Config</summary>
372
373
```yaml
374
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
375
+ - name: prod-query
376
+ subscription_ids:
377
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ discovery:
379
+ mode: query
380
+ mode_query:
381
+ kql: |
382
+ resources
383
+ | where tags.env =~ "prod"
384
+ | project id, name, type, resourceGroup, location
385
+ profiles:
386
+ mode: auto
387
auth:
388
mode: default
389
@@ -259,14 +392,15 @@ jobs:
392
393
###### Azure Government cloud
394
262
-Connect to Azure Government cloud environment.
395
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
396
397
<details open><summary>Config</summary>
398
399
```yaml
400
jobs:
401
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
402
+ subscription_ids:
403
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
cloud: government
405
auth:
406
mode: service_principal
@@ -322,6 +456,7 @@ Labels:
456
| region | The Azure region where the resource is deployed. |
457
| resource_type | The Azure resource type identifier. |
458
| profile | The Azure Monitor profile id. |
459
+| subscription_id | The Azure subscription identifier. |
460
| resource_uid | The unique Azure resource identifier. |
461
462
Metrics:
@@ -423,31 +558,46 @@ docker logs netdata 2>&1 | grep azure_monitor
558
559
### No metrics are collected
560
426
-Verify the following:
427
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
428
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
429
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
430
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
561
+Check the following:
562
+
563
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
564
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
565
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
566
+- **Collector logs** -- Check for authentication or API errors:
567
+ ```bash
568
+ # systemd
569
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
570
+ # non-systemd
571
+ grep azure_monitor /var/log/netdata/collector.log
572
+ ```
573
574
575
### Missing metrics for some resource types
576
435
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
436
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
437
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
438
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
577
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
578
+
579
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
580
+- **Verify a built-in profile exists** -- List available profiles:
581
+ ```bash
582
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
583
+ ```
584
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
585
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
586
587
441
-### Metrics appear delayed
588
+### Charts have gaps or incomplete data
589
443
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
444
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
445
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
590
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
591
+
592
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
593
+- Slower time-grain batches automatically use a larger effective offset when needed.
594
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
595
596
597
### Authentication errors in sovereign clouds
598
599
For Azure Government or Azure China clouds, set the `cloud` parameter:
600
+
601
- Azure Government: `cloud: government`
602
- Azure China (21Vianet): `cloud: china`
603
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_load_balancer.md
+237
-87
@@ -21,40 +21,76 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Load Balancer health and throughput including data path and health probe availability, SYN and SNAT connection counts, byte and packet throughput, allocated and used SNAT ports, and connection attempt rates.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Load Balancer with metrics covering:
31
+
32
+- **Availability** -- data path availability, health probe status, global backend availability
33
+- **Throughput** -- byte and packet throughput
34
+- **Connections** -- SNAT connections, SYN packet count
35
+- **SNAT ports** -- allocated and used SNAT ports
36
+
37
+
38
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
39
40
41
This collector is supported on all platforms.
42
43
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
44
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
45
+The service principal or managed identity requires these Azure RBAC roles:
46
+
47
+| Role | Purpose | Scope |
48
+|:-----|:--------|:------|
49
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
50
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
51
52
53
### Default Behavior
54
55
#### Auto-Detection
56
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
57
+The collector has two discovery phases:
58
+
59
+**Bootstrap (first run)**
60
+
61
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
62
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
63
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
64
+- A single job can monitor multiple subscriptions.
65
+
66
+**Runtime (periodic refresh)**
67
+
68
+- Periodically re-discovers resources for **already-active profile types only**.
69
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
70
+
71
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
72
73
74
#### Limits
75
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
76
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
77
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
78
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
79
80
81
#### Performance Impact
82
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
83
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
84
+
85
+**Default concurrency and batching limits:**
86
+
87
+| Setting | Default | Description |
88
+|:--------|:--------|:------------|
89
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
90
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
91
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
92
+
93
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
94
95
96
## Setup
@@ -78,25 +114,36 @@ UI configuration requires paid Netdata Cloud plan.
114
115
#### Create an Azure monitoring principal
116
81
-Create a service principal or use a managed identity with the following permissions:
117
+The collector requires a service principal or managed identity with two Azure RBAC roles:
118
+
119
+| Role | Purpose |
120
+|:-----|:--------|
121
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
122
+| **Reader** | Query Azure Resource Graph for resource discovery |
123
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
124
+**Option A: Service principal**
125
86
-For service principal authentication:
126
```bash
88
-# Create the service principal
127
+# Create service principal with Monitoring Reader role
128
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
129
--scopes /subscriptions/<subscription-id>
130
131
+# Add the Reader role for resource discovery
132
+az role assignment create --assignee <appId-from-above> \
133
+ --role "Reader" --scope /subscriptions/<subscription-id>
134
+
135
# Note the appId (client_id), password (client_secret), and tenant
136
```
137
95
-For managed identity (on Azure VMs, VMSS, or AKS):
138
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
139
+
140
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
141
+# Assign both roles to the VM's managed identity
142
az role assignment create --assignee <managed-identity-principal-id> \
143
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
144
+
145
+az role assignment create --assignee <managed-identity-principal-id> \
146
+ --role "Reader" --scope /subscriptions/<subscription-id>
147
```
148
149
@@ -105,13 +152,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
152
153
#### Options
154
108
-The following options can be defined globally: update_every, autodetection_retry.
155
+The following options can be defined globally: `update_every`, `autodetection_retry`.
156
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
157
+**Profile file locations:**
158
114
-User profile files with the same filename override stock profiles.
159
+| Type | Path |
160
+|:-----|:-----|
161
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
162
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
163
+
164
+User profile files with the same `id` as a stock profile override it.
165
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
166
167
168
<details open><summary>Config options</summary>
@@ -122,25 +173,103 @@ User profile files with the same filename override stock profiles.
173
|:------|:-----|:------------|:--------|:---------:|
174
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
176
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
177
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
180
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
181
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
184
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
185
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
186
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
187
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
190
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
191
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
192
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
193
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
194
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
195
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
196
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
197
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
198
199
+<a id="option-collection-query-offset"></a>
200
+##### query_offset
201
+
202
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
203
+
204
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
205
+
206
+- **Default (180s)** works for most services.
207
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
208
+- **Increase to 240-300s** if you still see gaps or missing data points.
209
+- **Do not set below 60s** -- metrics will likely be incomplete.
210
+
211
+
212
+<a id="option-authentication-auth-mode"></a>
213
+##### auth.mode
214
+
215
+Determines how the collector authenticates with Azure.
216
+
217
+| Mode | When to use | Required options |
218
+|:-----|:------------|:-----------------|
219
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
220
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
221
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
222
+
223
+
224
+<a id="option-discovery-discovery-mode"></a>
225
+##### discovery.mode
226
+
227
+Controls how the collector finds candidate Azure resources.
228
+
229
+| Mode | Behavior |
230
+|:-----|:---------|
231
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
232
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
233
+
234
+
235
+<a id="option-discovery-discovery-mode-query-kql"></a>
236
+##### discovery.mode_query.kql
237
+
238
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
239
+
240
+The query **must** project these five columns:
241
+
242
+| Column | Description |
243
+|:-------|:------------|
244
+| `id` | Full Azure resource ID (ARM format) |
245
+| `name` | Resource name |
246
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
247
+| `resourceGroup` | Resource group name |
248
+| `location` | Azure region |
249
+
250
+Example:
251
+
252
+```
253
+resources
254
+| where tags.env =~ "prod"
255
+| project id, name, type, resourceGroup, location
256
+```
257
+
258
+
259
+<a id="option-profiles-profiles-mode"></a>
260
+##### profiles.mode
261
+
262
+Controls how the collector decides which metric profiles to activate.
263
+
264
+| Mode | Behavior |
265
+|:-----|:---------|
266
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
267
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
268
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
269
+
270
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
271
+
272
+
273
274
</details>
275
@@ -182,14 +311,28 @@ sudo ./edit-config go.d/azure_monitor.conf
311
312
##### Examples
313
185
-###### Service principal (auto-discover all resources)
314
+###### Service principal with structured discovery
315
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
316
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
317
318
```yaml
319
jobs:
320
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ subscription_ids:
322
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
323
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
324
+ discovery:
325
+ mode: filters
326
+ mode_filters:
327
+ resource_groups:
328
+ - production-rg
329
+ regions:
330
+ - eastus
331
+ tags:
332
+ env:
333
+ - prod
334
+ profiles:
335
+ mode: auto
336
auth:
337
mode: service_principal
338
mode_service_principal:
@@ -198,59 +341,49 @@ jobs:
341
client_secret: "your-client-secret"
342
343
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
344
+###### Managed identity with exact profiles
345
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
346
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
347
348
<details open><summary>Config</summary>
349
350
```yaml
351
jobs:
352
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ subscription_ids:
354
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
356
+ mode: exact
357
+ mode_exact:
358
+ names:
359
+ - sql_database
360
+ - postgres_flexible
361
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
362
+ mode: managed_identity
363
364
```
365
</details>
366
241
-###### Filter by resource group
367
+###### Custom Azure Resource Graph KQL
368
243
-Only monitor resources in specific resource groups.
369
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
370
371
<details open><summary>Config</summary>
372
373
```yaml
374
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
375
+ - name: prod-query
376
+ subscription_ids:
377
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ discovery:
379
+ mode: query
380
+ mode_query:
381
+ kql: |
382
+ resources
383
+ | where tags.env =~ "prod"
384
+ | project id, name, type, resourceGroup, location
385
+ profiles:
386
+ mode: auto
387
auth:
388
mode: default
389
@@ -259,14 +392,15 @@ jobs:
392
393
###### Azure Government cloud
394
262
-Connect to Azure Government cloud environment.
395
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
396
397
<details open><summary>Config</summary>
398
399
```yaml
400
jobs:
401
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
402
+ subscription_ids:
403
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
cloud: government
405
auth:
406
mode: service_principal
@@ -314,6 +448,7 @@ Labels:
448
| region | The Azure region where the resource is deployed. |
449
| resource_type | The Azure resource type identifier. |
450
| profile | The Azure Monitor profile id. |
451
+| subscription_id | The Azure subscription identifier. |
452
| resource_uid | The unique Azure resource identifier. |
453
454
Metrics:
@@ -400,31 +535,46 @@ docker logs netdata 2>&1 | grep azure_monitor
535
536
### No metrics are collected
537
403
-Verify the following:
404
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
405
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
406
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
407
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
538
+Check the following:
539
+
540
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
541
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
542
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
543
+- **Collector logs** -- Check for authentication or API errors:
544
+ ```bash
545
+ # systemd
546
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
547
+ # non-systemd
548
+ grep azure_monitor /var/log/netdata/collector.log
549
+ ```
550
551
552
### Missing metrics for some resource types
553
412
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
413
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
414
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
415
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
554
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
555
+
556
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
557
+- **Verify a built-in profile exists** -- List available profiles:
558
+ ```bash
559
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
560
+ ```
561
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
562
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
563
564
418
-### Metrics appear delayed
565
+### Charts have gaps or incomplete data
566
420
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
421
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
422
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
567
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
568
+
569
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
570
+- Slower time-grain batches automatically use a larger effective offset when needed.
571
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
572
573
574
### Authentication errors in sovereign clouds
575
576
For Azure Government or Azure China clouds, set the `cloud` parameter:
577
+
578
- Azure Government: `cloud: government`
579
- Azure China (21Vianet): `cloud: china`
580
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_log_analytics_workspace.md
+237
-87
@@ -21,40 +21,76 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Log Analytics workspaces including ingestion volume and latency, query execution counts and volume, available storage capacity, and per-table breakdowns of ingestion rates and billing volume.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Log Analytics Workspace with metrics covering:
31
+
32
+- **Ingestion** -- ingestion volume (records/s), ingestion latency (average/max/min)
33
+- **Queries** -- query count (total/failed), query availability
34
+- **Export** -- exported data (bytes/s), exported records
35
+- **Legacy agent** -- CPU utilization (processor/privileged/user/idle), memory (available/used/free), disk (free space/utilization/I/O/queue/latency), network (traffic/packets/errors/throughput), swap, paging, system processes, uptime, events, heartbeats, users
36
+
37
+
38
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
39
40
41
This collector is supported on all platforms.
42
43
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
44
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
45
+The service principal or managed identity requires these Azure RBAC roles:
46
+
47
+| Role | Purpose | Scope |
48
+|:-----|:--------|:------|
49
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
50
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
51
52
53
### Default Behavior
54
55
#### Auto-Detection
56
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
57
+The collector has two discovery phases:
58
+
59
+**Bootstrap (first run)**
60
+
61
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
62
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
63
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
64
+- A single job can monitor multiple subscriptions.
65
+
66
+**Runtime (periodic refresh)**
67
+
68
+- Periodically re-discovers resources for **already-active profile types only**.
69
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
70
+
71
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
72
73
74
#### Limits
75
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
76
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
77
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
78
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
79
80
81
#### Performance Impact
82
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
83
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
84
+
85
+**Default concurrency and batching limits:**
86
+
87
+| Setting | Default | Description |
88
+|:--------|:--------|:------------|
89
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
90
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
91
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
92
+
93
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
94
95
96
## Setup
@@ -78,25 +114,36 @@ UI configuration requires paid Netdata Cloud plan.
114
115
#### Create an Azure monitoring principal
116
81
-Create a service principal or use a managed identity with the following permissions:
117
+The collector requires a service principal or managed identity with two Azure RBAC roles:
118
+
119
+| Role | Purpose |
120
+|:-----|:--------|
121
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
122
+| **Reader** | Query Azure Resource Graph for resource discovery |
123
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
124
+**Option A: Service principal**
125
86
-For service principal authentication:
126
```bash
88
-# Create the service principal
127
+# Create service principal with Monitoring Reader role
128
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
129
--scopes /subscriptions/<subscription-id>
130
131
+# Add the Reader role for resource discovery
132
+az role assignment create --assignee <appId-from-above> \
133
+ --role "Reader" --scope /subscriptions/<subscription-id>
134
+
135
# Note the appId (client_id), password (client_secret), and tenant
136
```
137
95
-For managed identity (on Azure VMs, VMSS, or AKS):
138
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
139
+
140
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
141
+# Assign both roles to the VM's managed identity
142
az role assignment create --assignee <managed-identity-principal-id> \
143
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
144
+
145
+az role assignment create --assignee <managed-identity-principal-id> \
146
+ --role "Reader" --scope /subscriptions/<subscription-id>
147
```
148
149
@@ -105,13 +152,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
152
153
#### Options
154
108
-The following options can be defined globally: update_every, autodetection_retry.
155
+The following options can be defined globally: `update_every`, `autodetection_retry`.
156
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
157
+**Profile file locations:**
158
114
-User profile files with the same filename override stock profiles.
159
+| Type | Path |
160
+|:-----|:-----|
161
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
162
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
163
+
164
+User profile files with the same `id` as a stock profile override it.
165
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
166
167
168
<details open><summary>Config options</summary>
@@ -122,25 +173,103 @@ User profile files with the same filename override stock profiles.
173
|:------|:-----|:------------|:--------|:---------:|
174
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
176
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
177
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
180
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
181
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
184
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
185
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
186
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
187
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
190
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
191
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
192
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
193
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
194
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
195
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
196
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
197
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
198
199
+<a id="option-collection-query-offset"></a>
200
+##### query_offset
201
+
202
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
203
+
204
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
205
+
206
+- **Default (180s)** works for most services.
207
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
208
+- **Increase to 240-300s** if you still see gaps or missing data points.
209
+- **Do not set below 60s** -- metrics will likely be incomplete.
210
+
211
+
212
+<a id="option-authentication-auth-mode"></a>
213
+##### auth.mode
214
+
215
+Determines how the collector authenticates with Azure.
216
+
217
+| Mode | When to use | Required options |
218
+|:-----|:------------|:-----------------|
219
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
220
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
221
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
222
+
223
+
224
+<a id="option-discovery-discovery-mode"></a>
225
+##### discovery.mode
226
+
227
+Controls how the collector finds candidate Azure resources.
228
+
229
+| Mode | Behavior |
230
+|:-----|:---------|
231
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
232
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
233
+
234
+
235
+<a id="option-discovery-discovery-mode-query-kql"></a>
236
+##### discovery.mode_query.kql
237
+
238
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
239
+
240
+The query **must** project these five columns:
241
+
242
+| Column | Description |
243
+|:-------|:------------|
244
+| `id` | Full Azure resource ID (ARM format) |
245
+| `name` | Resource name |
246
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
247
+| `resourceGroup` | Resource group name |
248
+| `location` | Azure region |
249
+
250
+Example:
251
+
252
+```
253
+resources
254
+| where tags.env =~ "prod"
255
+| project id, name, type, resourceGroup, location
256
+```
257
+
258
+
259
+<a id="option-profiles-profiles-mode"></a>
260
+##### profiles.mode
261
+
262
+Controls how the collector decides which metric profiles to activate.
263
+
264
+| Mode | Behavior |
265
+|:-----|:---------|
266
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
267
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
268
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
269
+
270
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
271
+
272
+
273
274
</details>
275
@@ -182,14 +311,28 @@ sudo ./edit-config go.d/azure_monitor.conf
311
312
##### Examples
313
185
-###### Service principal (auto-discover all resources)
314
+###### Service principal with structured discovery
315
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
316
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
317
318
```yaml
319
jobs:
320
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ subscription_ids:
322
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
323
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
324
+ discovery:
325
+ mode: filters
326
+ mode_filters:
327
+ resource_groups:
328
+ - production-rg
329
+ regions:
330
+ - eastus
331
+ tags:
332
+ env:
333
+ - prod
334
+ profiles:
335
+ mode: auto
336
auth:
337
mode: service_principal
338
mode_service_principal:
@@ -198,59 +341,49 @@ jobs:
341
client_secret: "your-client-secret"
342
343
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
344
+###### Managed identity with exact profiles
345
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
346
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
347
348
<details open><summary>Config</summary>
349
350
```yaml
351
jobs:
352
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ subscription_ids:
354
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
356
+ mode: exact
357
+ mode_exact:
358
+ names:
359
+ - sql_database
360
+ - postgres_flexible
361
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
362
+ mode: managed_identity
363
364
```
365
</details>
366
241
-###### Filter by resource group
367
+###### Custom Azure Resource Graph KQL
368
243
-Only monitor resources in specific resource groups.
369
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
370
371
<details open><summary>Config</summary>
372
373
```yaml
374
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
375
+ - name: prod-query
376
+ subscription_ids:
377
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ discovery:
379
+ mode: query
380
+ mode_query:
381
+ kql: |
382
+ resources
383
+ | where tags.env =~ "prod"
384
+ | project id, name, type, resourceGroup, location
385
+ profiles:
386
+ mode: auto
387
auth:
388
mode: default
389
@@ -259,14 +392,15 @@ jobs:
392
393
###### Azure Government cloud
394
262
-Connect to Azure Government cloud environment.
395
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
396
397
<details open><summary>Config</summary>
398
399
```yaml
400
jobs:
401
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
402
+ subscription_ids:
403
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
cloud: government
405
auth:
406
mode: service_principal
@@ -324,6 +458,7 @@ Labels:
458
| region | The Azure region where the resource is deployed. |
459
| resource_type | The Azure resource type identifier. |
460
| profile | The Azure Monitor profile id. |
461
+| subscription_id | The Azure subscription identifier. |
462
| resource_uid | The unique Azure resource identifier. |
463
464
Metrics:
@@ -439,31 +574,46 @@ docker logs netdata 2>&1 | grep azure_monitor
574
575
### No metrics are collected
576
442
-Verify the following:
443
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
444
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
445
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
446
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
577
+Check the following:
578
+
579
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
580
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
581
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
582
+- **Collector logs** -- Check for authentication or API errors:
583
+ ```bash
584
+ # systemd
585
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
586
+ # non-systemd
587
+ grep azure_monitor /var/log/netdata/collector.log
588
+ ```
589
590
591
### Missing metrics for some resource types
592
451
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
452
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
453
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
454
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
593
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
594
+
595
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
596
+- **Verify a built-in profile exists** -- List available profiles:
597
+ ```bash
598
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
599
+ ```
600
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
601
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
602
603
457
-### Metrics appear delayed
604
+### Charts have gaps or incomplete data
605
459
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
460
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
461
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
606
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
607
+
608
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
609
+- Slower time-grain batches automatically use a larger effective offset when needed.
610
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
611
612
613
### Authentication errors in sovereign clouds
614
615
For Azure Government or Azure China clouds, set the `cloud` parameter:
616
+
617
- Azure Government: `cloud: government`
618
- Azure China (21Vianet): `cloud: china`
619
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_logic_apps_workflow.md
+239
-87
@@ -21,40 +21,78 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Logic Apps workflow execution including run completions and failures, action execution counts, trigger firing rates, run and action latency, billable executions, and action-level success and failure breakdowns.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Logic Apps with metrics covering:
31
+
32
+- **Runs** -- run lifecycle (started/completed/succeeded/failed/cancelled), run failure rate
33
+- **Actions** -- action lifecycle (started/completed/succeeded/failed/skipped)
34
+- **Triggers** -- trigger lifecycle (started/completed/succeeded/fired/failed/skipped)
35
+- **Latency** -- run, action, and trigger latency
36
+- **Billing** -- billable executions (total/actions/triggers), billing by type (native/connector/storage)
37
+- **Throttling** -- run, action, and trigger throttling events
38
+
39
+
40
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
41
42
43
This collector is supported on all platforms.
44
45
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
46
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
47
+The service principal or managed identity requires these Azure RBAC roles:
48
+
49
+| Role | Purpose | Scope |
50
+|:-----|:--------|:------|
51
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
52
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
53
54
55
### Default Behavior
56
57
#### Auto-Detection
58
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
59
+The collector has two discovery phases:
60
+
61
+**Bootstrap (first run)**
62
+
63
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
64
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
65
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
66
+- A single job can monitor multiple subscriptions.
67
+
68
+**Runtime (periodic refresh)**
69
+
70
+- Periodically re-discovers resources for **already-active profile types only**.
71
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
72
+
73
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
74
75
76
#### Limits
77
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
78
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
79
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
80
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
81
82
83
#### Performance Impact
84
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
85
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
86
+
87
+**Default concurrency and batching limits:**
88
+
89
+| Setting | Default | Description |
90
+|:--------|:--------|:------------|
91
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
92
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
93
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
94
+
95
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
96
97
98
## Setup
@@ -78,25 +116,36 @@ UI configuration requires paid Netdata Cloud plan.
116
117
#### Create an Azure monitoring principal
118
81
-Create a service principal or use a managed identity with the following permissions:
119
+The collector requires a service principal or managed identity with two Azure RBAC roles:
120
+
121
+| Role | Purpose |
122
+|:-----|:--------|
123
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
124
+| **Reader** | Query Azure Resource Graph for resource discovery |
125
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
126
+**Option A: Service principal**
127
86
-For service principal authentication:
128
```bash
88
-# Create the service principal
129
+# Create service principal with Monitoring Reader role
130
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
131
--scopes /subscriptions/<subscription-id>
132
133
+# Add the Reader role for resource discovery
134
+az role assignment create --assignee <appId-from-above> \
135
+ --role "Reader" --scope /subscriptions/<subscription-id>
136
+
137
# Note the appId (client_id), password (client_secret), and tenant
138
```
139
95
-For managed identity (on Azure VMs, VMSS, or AKS):
140
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
141
+
142
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
143
+# Assign both roles to the VM's managed identity
144
az role assignment create --assignee <managed-identity-principal-id> \
145
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+
147
+az role assignment create --assignee <managed-identity-principal-id> \
148
+ --role "Reader" --scope /subscriptions/<subscription-id>
149
```
150
151
@@ -105,13 +154,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
154
155
#### Options
156
108
-The following options can be defined globally: update_every, autodetection_retry.
157
+The following options can be defined globally: `update_every`, `autodetection_retry`.
158
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+**Profile file locations:**
160
114
-User profile files with the same filename override stock profiles.
161
+| Type | Path |
162
+|:-----|:-----|
163
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
164
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
165
+
166
+User profile files with the same `id` as a stock profile override it.
167
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
168
169
170
<details open><summary>Config options</summary>
@@ -122,25 +175,103 @@ User profile files with the same filename override stock profiles.
175
|:------|:-----|:------------|:--------|:---------:|
176
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
177
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
178
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
179
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
183
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
186
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
187
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
188
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
189
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
192
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
193
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
194
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
195
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
196
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
197
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
198
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
199
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
200
201
+<a id="option-collection-query-offset"></a>
202
+##### query_offset
203
+
204
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
205
+
206
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
207
+
208
+- **Default (180s)** works for most services.
209
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
210
+- **Increase to 240-300s** if you still see gaps or missing data points.
211
+- **Do not set below 60s** -- metrics will likely be incomplete.
212
+
213
+
214
+<a id="option-authentication-auth-mode"></a>
215
+##### auth.mode
216
+
217
+Determines how the collector authenticates with Azure.
218
+
219
+| Mode | When to use | Required options |
220
+|:-----|:------------|:-----------------|
221
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
222
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
223
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
224
+
225
+
226
+<a id="option-discovery-discovery-mode"></a>
227
+##### discovery.mode
228
+
229
+Controls how the collector finds candidate Azure resources.
230
+
231
+| Mode | Behavior |
232
+|:-----|:---------|
233
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
234
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
235
+
236
+
237
+<a id="option-discovery-discovery-mode-query-kql"></a>
238
+##### discovery.mode_query.kql
239
+
240
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
241
+
242
+The query **must** project these five columns:
243
+
244
+| Column | Description |
245
+|:-------|:------------|
246
+| `id` | Full Azure resource ID (ARM format) |
247
+| `name` | Resource name |
248
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
249
+| `resourceGroup` | Resource group name |
250
+| `location` | Azure region |
251
+
252
+Example:
253
+
254
+```
255
+resources
256
+| where tags.env =~ "prod"
257
+| project id, name, type, resourceGroup, location
258
+```
259
+
260
+
261
+<a id="option-profiles-profiles-mode"></a>
262
+##### profiles.mode
263
+
264
+Controls how the collector decides which metric profiles to activate.
265
+
266
+| Mode | Behavior |
267
+|:-----|:---------|
268
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
269
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
270
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
271
+
272
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
273
+
274
+
275
276
</details>
277
@@ -182,14 +313,28 @@ sudo ./edit-config go.d/azure_monitor.conf
313
314
##### Examples
315
185
-###### Service principal (auto-discover all resources)
316
+###### Service principal with structured discovery
317
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
318
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
319
320
```yaml
321
jobs:
322
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ subscription_ids:
324
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
325
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
326
+ discovery:
327
+ mode: filters
328
+ mode_filters:
329
+ resource_groups:
330
+ - production-rg
331
+ regions:
332
+ - eastus
333
+ tags:
334
+ env:
335
+ - prod
336
+ profiles:
337
+ mode: auto
338
auth:
339
mode: service_principal
340
mode_service_principal:
@@ -198,59 +343,49 @@ jobs:
343
client_secret: "your-client-secret"
344
345
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
346
+###### Managed identity with exact profiles
347
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
348
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
349
350
<details open><summary>Config</summary>
351
352
```yaml
353
jobs:
354
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
+ subscription_ids:
356
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
358
+ mode: exact
359
+ mode_exact:
360
+ names:
361
+ - sql_database
362
+ - postgres_flexible
363
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
364
+ mode: managed_identity
365
366
```
367
</details>
368
241
-###### Filter by resource group
369
+###### Custom Azure Resource Graph KQL
370
243
-Only monitor resources in specific resource groups.
371
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
372
373
<details open><summary>Config</summary>
374
375
```yaml
376
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
377
+ - name: prod-query
378
+ subscription_ids:
379
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
380
+ discovery:
381
+ mode: query
382
+ mode_query:
383
+ kql: |
384
+ resources
385
+ | where tags.env =~ "prod"
386
+ | project id, name, type, resourceGroup, location
387
+ profiles:
388
+ mode: auto
389
auth:
390
mode: default
391
@@ -259,14 +394,15 @@ jobs:
394
395
###### Azure Government cloud
396
262
-Connect to Azure Government cloud environment.
397
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
398
399
<details open><summary>Config</summary>
400
401
```yaml
402
jobs:
403
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
+ subscription_ids:
405
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
cloud: government
407
auth:
408
mode: service_principal
@@ -322,6 +458,7 @@ Labels:
458
| region | The Azure region where the resource is deployed. |
459
| resource_type | The Azure resource type identifier. |
460
| profile | The Azure Monitor profile id. |
461
+| subscription_id | The Azure subscription identifier. |
462
| resource_uid | The unique Azure resource identifier. |
463
464
Metrics:
@@ -413,31 +550,46 @@ docker logs netdata 2>&1 | grep azure_monitor
550
551
### No metrics are collected
552
416
-Verify the following:
417
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
418
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
419
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
420
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
553
+Check the following:
554
+
555
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
556
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
557
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
558
+- **Collector logs** -- Check for authentication or API errors:
559
+ ```bash
560
+ # systemd
561
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
562
+ # non-systemd
563
+ grep azure_monitor /var/log/netdata/collector.log
564
+ ```
565
566
567
### Missing metrics for some resource types
568
425
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
426
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
427
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
428
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
569
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
570
+
571
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
572
+- **Verify a built-in profile exists** -- List available profiles:
573
+ ```bash
574
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
575
+ ```
576
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
577
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
578
579
431
-### Metrics appear delayed
580
+### Charts have gaps or incomplete data
581
433
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
434
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
435
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
582
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
583
+
584
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
585
+- Slower time-grain batches automatically use a larger effective offset when needed.
586
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
587
588
589
### Authentication errors in sovereign clouds
590
591
For Azure Government or Azure China clouds, set the `cloud` parameter:
592
+
593
- Azure Government: `cloud: government`
594
- Azure China (21Vianet): `cloud: china`
595
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_machine_learning_workspace.md
+242
-87
@@ -21,40 +21,81 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Machine Learning workspaces including active model deployments and registered models, pipeline run completions and failures, compute node utilization and preemptions, quota usage, managed endpoint request latency and rates, estimated GPU utilization, and storage utilization.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Machine Learning with metrics covering:
31
+
32
+- **Compute** -- CPU utilization and millicores (used/capacity), CPU memory (used/capacity)
33
+- **GPU** -- GPU utilization (cluster/node), GPU memory (used/capacity), GPU energy
34
+- **Cluster** -- total cores and nodes, cluster cores and nodes by state (active/idle/leaving/preempted/unusable)
35
+- **Runs** -- run completion (completed/failed/cancelled), run lifecycle, run issues (errors/warnings)
36
+- **Models** -- model registrations (succeeded/failed), model deployments (started/succeeded/failed)
37
+- **Quota** -- quota utilization
38
+- **Storage** -- disk I/O (read/write), disk usage (used/available), storage API calls
39
+- **Network** -- network traffic (in/out), InfiniBand traffic
40
+- **AI agents** -- agent runs, messages, tokens, tool calls, events, indexed files
41
+
42
+
43
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
44
45
46
This collector is supported on all platforms.
47
48
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
49
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
50
+The service principal or managed identity requires these Azure RBAC roles:
51
+
52
+| Role | Purpose | Scope |
53
+|:-----|:--------|:------|
54
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
55
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
56
57
58
### Default Behavior
59
60
#### Auto-Detection
61
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
62
+The collector has two discovery phases:
63
+
64
+**Bootstrap (first run)**
65
+
66
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
67
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
68
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
69
+- A single job can monitor multiple subscriptions.
70
+
71
+**Runtime (periodic refresh)**
72
+
73
+- Periodically re-discovers resources for **already-active profile types only**.
74
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
75
+
76
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
77
78
79
#### Limits
80
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
81
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
82
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
83
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
84
85
86
#### Performance Impact
87
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
88
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
89
+
90
+**Default concurrency and batching limits:**
91
+
92
+| Setting | Default | Description |
93
+|:--------|:--------|:------------|
94
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
95
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
96
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
97
+
98
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
99
100
101
## Setup
@@ -78,25 +119,36 @@ UI configuration requires paid Netdata Cloud plan.
119
120
#### Create an Azure monitoring principal
121
81
-Create a service principal or use a managed identity with the following permissions:
122
+The collector requires a service principal or managed identity with two Azure RBAC roles:
123
+
124
+| Role | Purpose |
125
+|:-----|:--------|
126
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
127
+| **Reader** | Query Azure Resource Graph for resource discovery |
128
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
129
+**Option A: Service principal**
130
86
-For service principal authentication:
131
```bash
88
-# Create the service principal
132
+# Create service principal with Monitoring Reader role
133
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
134
--scopes /subscriptions/<subscription-id>
135
136
+# Add the Reader role for resource discovery
137
+az role assignment create --assignee <appId-from-above> \
138
+ --role "Reader" --scope /subscriptions/<subscription-id>
139
+
140
# Note the appId (client_id), password (client_secret), and tenant
141
```
142
95
-For managed identity (on Azure VMs, VMSS, or AKS):
143
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
144
+
145
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
146
+# Assign both roles to the VM's managed identity
147
az role assignment create --assignee <managed-identity-principal-id> \
148
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Reader" --scope /subscriptions/<subscription-id>
152
```
153
154
@@ -105,13 +157,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
157
158
#### Options
159
108
-The following options can be defined globally: update_every, autodetection_retry.
160
+The following options can be defined globally: `update_every`, `autodetection_retry`.
161
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+**Profile file locations:**
163
114
-User profile files with the same filename override stock profiles.
164
+| Type | Path |
165
+|:-----|:-----|
166
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
167
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
168
+
169
+User profile files with the same `id` as a stock profile override it.
170
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
171
172
173
<details open><summary>Config options</summary>
@@ -122,25 +178,103 @@ User profile files with the same filename override stock profiles.
178
|:------|:-----|:------------|:--------|:---------:|
179
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
180
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
181
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
182
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
186
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
191
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
192
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
195
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
196
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
197
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
198
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
199
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
204
+<a id="option-collection-query-offset"></a>
205
+##### query_offset
206
+
207
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
208
+
209
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
210
+
211
+- **Default (180s)** works for most services.
212
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
213
+- **Increase to 240-300s** if you still see gaps or missing data points.
214
+- **Do not set below 60s** -- metrics will likely be incomplete.
215
+
216
+
217
+<a id="option-authentication-auth-mode"></a>
218
+##### auth.mode
219
+
220
+Determines how the collector authenticates with Azure.
221
+
222
+| Mode | When to use | Required options |
223
+|:-----|:------------|:-----------------|
224
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
225
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
226
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
227
+
228
+
229
+<a id="option-discovery-discovery-mode"></a>
230
+##### discovery.mode
231
+
232
+Controls how the collector finds candidate Azure resources.
233
+
234
+| Mode | Behavior |
235
+|:-----|:---------|
236
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
237
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
238
+
239
+
240
+<a id="option-discovery-discovery-mode-query-kql"></a>
241
+##### discovery.mode_query.kql
242
+
243
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
244
+
245
+The query **must** project these five columns:
246
+
247
+| Column | Description |
248
+|:-------|:------------|
249
+| `id` | Full Azure resource ID (ARM format) |
250
+| `name` | Resource name |
251
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
252
+| `resourceGroup` | Resource group name |
253
+| `location` | Azure region |
254
+
255
+Example:
256
+
257
+```
258
+resources
259
+| where tags.env =~ "prod"
260
+| project id, name, type, resourceGroup, location
261
+```
262
+
263
+
264
+<a id="option-profiles-profiles-mode"></a>
265
+##### profiles.mode
266
+
267
+Controls how the collector decides which metric profiles to activate.
268
+
269
+| Mode | Behavior |
270
+|:-----|:---------|
271
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
272
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
273
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
274
+
275
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
276
+
277
+
278
279
</details>
280
@@ -182,14 +316,28 @@ sudo ./edit-config go.d/azure_monitor.conf
316
317
##### Examples
318
185
-###### Service principal (auto-discover all resources)
319
+###### Service principal with structured discovery
320
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
321
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
322
323
```yaml
324
jobs:
325
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ subscription_ids:
327
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
328
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
329
+ discovery:
330
+ mode: filters
331
+ mode_filters:
332
+ resource_groups:
333
+ - production-rg
334
+ regions:
335
+ - eastus
336
+ tags:
337
+ env:
338
+ - prod
339
+ profiles:
340
+ mode: auto
341
auth:
342
mode: service_principal
343
mode_service_principal:
@@ -198,59 +346,49 @@ jobs:
346
client_secret: "your-client-secret"
347
348
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
349
+###### Managed identity with exact profiles
350
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
351
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
352
353
<details open><summary>Config</summary>
354
355
```yaml
356
jobs:
357
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ subscription_ids:
359
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
361
+ mode: exact
362
+ mode_exact:
363
+ names:
364
+ - sql_database
365
+ - postgres_flexible
366
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
367
+ mode: managed_identity
368
369
```
370
</details>
371
241
-###### Filter by resource group
372
+###### Custom Azure Resource Graph KQL
373
243
-Only monitor resources in specific resource groups.
374
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
375
376
<details open><summary>Config</summary>
377
378
```yaml
379
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
380
+ - name: prod-query
381
+ subscription_ids:
382
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
383
+ discovery:
384
+ mode: query
385
+ mode_query:
386
+ kql: |
387
+ resources
388
+ | where tags.env =~ "prod"
389
+ | project id, name, type, resourceGroup, location
390
+ profiles:
391
+ mode: auto
392
auth:
393
mode: default
394
@@ -259,14 +397,15 @@ jobs:
397
398
###### Azure Government cloud
399
262
-Connect to Azure Government cloud environment.
400
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
401
402
<details open><summary>Config</summary>
403
404
```yaml
405
jobs:
406
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
+ subscription_ids:
408
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
409
cloud: government
410
auth:
411
mode: service_principal
@@ -325,6 +464,7 @@ Labels:
464
| region | The Azure region where the resource is deployed. |
465
| resource_type | The Azure resource type identifier. |
466
| profile | The Azure Monitor profile id. |
467
+| subscription_id | The Azure subscription identifier. |
468
| resource_uid | The unique Azure resource identifier. |
469
470
Metrics:
@@ -433,31 +573,46 @@ docker logs netdata 2>&1 | grep azure_monitor
573
574
### No metrics are collected
575
436
-Verify the following:
437
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
438
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
439
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
440
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
576
+Check the following:
577
+
578
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
579
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
580
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
581
+- **Collector logs** -- Check for authentication or API errors:
582
+ ```bash
583
+ # systemd
584
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
585
+ # non-systemd
586
+ grep azure_monitor /var/log/netdata/collector.log
587
+ ```
588
589
590
### Missing metrics for some resource types
591
445
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
446
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
447
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
448
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
592
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
593
+
594
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
595
+- **Verify a built-in profile exists** -- List available profiles:
596
+ ```bash
597
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
598
+ ```
599
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
600
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
601
602
451
-### Metrics appear delayed
603
+### Charts have gaps or incomplete data
604
453
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
454
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
455
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
605
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
606
+
607
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
608
+- Slower time-grain batches automatically use a larger effective offset when needed.
609
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
610
611
612
### Authentication errors in sovereign clouds
613
614
For Azure Government or Azure China clouds, set the `cloud` parameter:
615
+
616
- Azure Government: `cloud: government`
617
- Azure China (21Vianet): `cloud: china`
618
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md
+236
-93
@@ -21,42 +21,72 @@ Module: azure_monitor
21
22
## Overview
23
24
-This collector monitors Azure resources through the Azure Monitor Metrics API. It automatically discovers
25
-resources in your subscription and collects platform metrics based on configurable profiles, providing
26
-visibility into the health and performance of over 35 Azure service types.
24
+This collector provides real-time visibility into your Azure infrastructure by collecting platform metrics from the Azure Monitor Metrics API.
25
26
+**Key capabilities:**
27
29
-The collector uses Azure SDK clients for:
30
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
31
-- Resource discovery via Azure Resource Graph queries
32
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+- **Multi-subscription** -- monitor resources across one or more Azure subscriptions in a single job
29
+- **Automatic service detection** -- discovers resources and enables matching metric profiles without manual configuration
30
+- **38 built-in service profiles** -- covers databases, compute, networking, storage, AI, analytics, and more
31
+- **Flexible discovery** -- use structured filters (resource groups, regions, tags) or a custom Azure Resource Graph KQL query
32
+
33
+
34
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
35
36
37
This collector is supported on all platforms.
38
39
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
40
39
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
41
+The service principal or managed identity requires these Azure RBAC roles:
42
+
43
+| Role | Purpose | Scope |
44
+|:-----|:--------|:------|
45
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
46
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
47
48
49
### Default Behavior
50
51
#### Auto-Detection
52
46
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
47
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
53
+The collector has two discovery phases:
54
+
55
+**Bootstrap (first run)**
56
+
57
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
58
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
59
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
60
+- A single job can monitor multiple subscriptions.
61
+
62
+**Runtime (periodic refresh)**
63
+
64
+- Periodically re-discovers resources for **already-active profile types only**.
65
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
66
+
67
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
68
69
70
#### Limits
71
52
-Azure Monitor metrics granularity is typically 1 minute.
53
-The collector enforces a minimum collection interval of 60 seconds.
72
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
73
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
74
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
75
76
77
#### Performance Impact
78
58
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
59
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
79
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
80
+
81
+**Default concurrency and batching limits:**
82
+
83
+| Setting | Default | Description |
84
+|:--------|:--------|:------------|
85
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
86
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
87
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
88
+
89
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
90
91
92
## Setup
@@ -80,25 +110,36 @@ UI configuration requires paid Netdata Cloud plan.
110
111
#### Create an Azure monitoring principal
112
83
-Create a service principal or use a managed identity with the following permissions:
113
+The collector requires a service principal or managed identity with two Azure RBAC roles:
114
+
115
+| Role | Purpose |
116
+|:-----|:--------|
117
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
118
+| **Reader** | Query Azure Resource Graph for resource discovery |
119
85
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
86
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
120
+**Option A: Service principal**
121
88
-For service principal authentication:
122
```bash
90
-# Create the service principal
123
+# Create service principal with Monitoring Reader role
124
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
125
--scopes /subscriptions/<subscription-id>
126
127
+# Add the Reader role for resource discovery
128
+az role assignment create --assignee <appId-from-above> \
129
+ --role "Reader" --scope /subscriptions/<subscription-id>
130
+
131
# Note the appId (client_id), password (client_secret), and tenant
132
```
133
97
-For managed identity (on Azure VMs, VMSS, or AKS):
134
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
135
+
136
```bash
99
-# Assign Monitoring Reader role to the VM's managed identity
137
+# Assign both roles to the VM's managed identity
138
az role assignment create --assignee <managed-identity-principal-id> \
139
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
140
+
141
+az role assignment create --assignee <managed-identity-principal-id> \
142
+ --role "Reader" --scope /subscriptions/<subscription-id>
143
```
144
145
@@ -107,13 +148,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
148
149
#### Options
150
110
-The following options can be defined globally: update_every, autodetection_retry.
151
+The following options can be defined globally: `update_every`, `autodetection_retry`.
152
+
153
+**Profile file locations:**
154
112
-Profile files are loaded from:
113
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
114
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
155
+| Type | Path |
156
+|:-----|:-----|
157
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
158
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
159
116
-User profile files with the same filename override stock profiles.
160
+User profile files with the same `id` as a stock profile override it.
161
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
162
163
164
<details open><summary>Config options</summary>
@@ -124,25 +169,103 @@ User profile files with the same filename override stock profiles.
169
|:------|:-----|:------------|:--------|:---------:|
170
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
171
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
127
-| **Target** | subscription_id | Azure subscription ID. | | yes |
172
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
173
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
129
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
130
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
174
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
175
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
132
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
133
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
134
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
135
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
136
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
137
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
138
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
139
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
176
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
177
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
178
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
179
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
180
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
181
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
182
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
183
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
184
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
185
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
186
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
187
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
188
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
189
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
190
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
191
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
192
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
193
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
194
195
+<a id="option-collection-query-offset"></a>
196
+##### query_offset
197
+
198
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
199
+
200
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
201
+
202
+- **Default (180s)** works for most services.
203
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
204
+- **Increase to 240-300s** if you still see gaps or missing data points.
205
+- **Do not set below 60s** -- metrics will likely be incomplete.
206
+
207
+
208
+<a id="option-authentication-auth-mode"></a>
209
+##### auth.mode
210
+
211
+Determines how the collector authenticates with Azure.
212
+
213
+| Mode | When to use | Required options |
214
+|:-----|:------------|:-----------------|
215
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
216
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
217
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
218
+
219
+
220
+<a id="option-discovery-discovery-mode"></a>
221
+##### discovery.mode
222
+
223
+Controls how the collector finds candidate Azure resources.
224
+
225
+| Mode | Behavior |
226
+|:-----|:---------|
227
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
228
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
229
+
230
+
231
+<a id="option-discovery-discovery-mode-query-kql"></a>
232
+##### discovery.mode_query.kql
233
+
234
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
235
+
236
+The query **must** project these five columns:
237
+
238
+| Column | Description |
239
+|:-------|:------------|
240
+| `id` | Full Azure resource ID (ARM format) |
241
+| `name` | Resource name |
242
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
243
+| `resourceGroup` | Resource group name |
244
+| `location` | Azure region |
245
+
246
+Example:
247
+
248
+```
249
+resources
250
+| where tags.env =~ "prod"
251
+| project id, name, type, resourceGroup, location
252
+```
253
+
254
+
255
+<a id="option-profiles-profiles-mode"></a>
256
+##### profiles.mode
257
+
258
+Controls how the collector decides which metric profiles to activate.
259
+
260
+| Mode | Behavior |
261
+|:-----|:---------|
262
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
263
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
264
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
265
+
266
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
267
+
268
+
269
270
</details>
271
@@ -184,14 +307,28 @@ sudo ./edit-config go.d/azure_monitor.conf
307
308
##### Examples
309
187
-###### Service principal (auto-discover all resources)
310
+###### Service principal with structured discovery
311
189
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
312
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
313
314
```yaml
315
jobs:
316
- name: prod
194
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
317
+ subscription_ids:
318
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
319
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
320
+ discovery:
321
+ mode: filters
322
+ mode_filters:
323
+ resource_groups:
324
+ - production-rg
325
+ regions:
326
+ - eastus
327
+ tags:
328
+ env:
329
+ - prod
330
+ profiles:
331
+ mode: auto
332
auth:
333
mode: service_principal
334
mode_service_principal:
@@ -200,59 +337,49 @@ jobs:
337
client_secret: "your-client-secret"
338
339
```
203
-###### Managed identity (Azure VM/VMSS/AKS)
340
+###### Managed identity with exact profiles
341
205
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
206
-
207
-<details open><summary>Config</summary>
208
-
209
-```yaml
210
-jobs:
211
- - name: prod
212
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
213
- auth:
214
- mode: managed_identity
215
-
216
-```
217
-</details>
218
-
219
-###### Specific profiles only
220
-
221
-Monitor only specific Azure services instead of auto-discovering all resource types.
342
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
343
344
<details open><summary>Config</summary>
345
346
```yaml
347
jobs:
348
- name: databases
228
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
349
+ subscription_ids:
350
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
351
profiles:
230
- - sql_database
231
- - postgres_flexible
232
- - redis_cache
352
+ mode: exact
353
+ mode_exact:
354
+ names:
355
+ - sql_database
356
+ - postgres_flexible
357
auth:
234
- mode: service_principal
235
- mode_service_principal:
236
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
237
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
238
- client_secret: "your-client-secret"
358
+ mode: managed_identity
359
360
```
361
</details>
362
243
-###### Filter by resource group
363
+###### Custom Azure Resource Graph KQL
364
245
-Only monitor resources in specific resource groups.
365
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
366
367
<details open><summary>Config</summary>
368
369
```yaml
370
jobs:
251
- - name: prod-rg
252
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
253
- resource_groups:
254
- - production-rg
255
- - staging-rg
371
+ - name: prod-query
372
+ subscription_ids:
373
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
374
+ discovery:
375
+ mode: query
376
+ mode_query:
377
+ kql: |
378
+ resources
379
+ | where tags.env =~ "prod"
380
+ | project id, name, type, resourceGroup, location
381
+ profiles:
382
+ mode: auto
383
auth:
384
mode: default
385
@@ -261,14 +388,15 @@ jobs:
388
389
###### Azure Government cloud
390
264
-Connect to Azure Government cloud environment.
391
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
392
393
<details open><summary>Config</summary>
394
395
```yaml
396
jobs:
397
- name: gov
271
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
398
+ subscription_ids:
399
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
400
cloud: government
401
auth:
402
mode: service_principal
@@ -289,11 +417,11 @@ There are no alerts configured by default for this integration.
417
418
## Metrics
419
292
-Metrics depend on which Azure Monitor profiles are enabled. Each profile corresponds to an Azure
293
-service type and defines the specific metrics collected. With the default `profiles: [auto]` setting,
294
-profiles are automatically enabled for resource types found in your subscription.
420
+The metrics collected depend on which Azure Monitor profiles are active. Each profile corresponds to an Azure service (e.g., SQL Database, Virtual Machines) and defines the specific charts and metrics for that service.
421
296
-See the service-specific integrations below for detailed metrics lists.
422
+With the default `profiles.mode: auto`, profiles are activated automatically based on the resource types found in your subscriptions.
423
+
424
+**See the service-specific integrations below for detailed metric lists per Azure service.**
425
426
427
@@ -366,31 +494,46 @@ docker logs netdata 2>&1 | grep azure_monitor
494
495
### No metrics are collected
496
369
-Verify the following:
370
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
371
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
372
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
373
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
497
+Check the following:
498
+
499
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
500
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
501
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
502
+- **Collector logs** -- Check for authentication or API errors:
503
+ ```bash
504
+ # systemd
505
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
506
+ # non-systemd
507
+ grep azure_monitor /var/log/netdata/collector.log
508
+ ```
509
510
511
### Missing metrics for some resource types
512
378
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
379
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
380
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
381
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
513
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
514
+
515
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
516
+- **Verify a built-in profile exists** -- List available profiles:
517
+ ```bash
518
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
519
+ ```
520
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
521
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
522
523
384
-### Metrics appear delayed
524
+### Charts have gaps or incomplete data
525
386
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
387
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
388
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
526
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
527
+
528
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
529
+- Slower time-grain batches automatically use a larger effective offset when needed.
530
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
531
532
533
### Authentication errors in sovereign clouds
534
535
For Azure Government or Azure China clouds, set the `cloud` parameter:
536
+
537
- Azure Government: `cloud: government`
538
- Azure China (21Vianet): `cloud: china`
539
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_mysql_flexible_server.md
+242
-87
@@ -21,40 +21,81 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor MySQL Flexible Server including active connections, aborted connections, query rates, replication lag, storage utilization, CPU and memory usage, IO operations, InnoDB buffer pool efficiency, network throughput, and HA replication status.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure MySQL Flexible Server with metrics covering:
31
+
32
+- **Compute** -- CPU utilization, CPU credits (consumed/remaining)
33
+- **Memory** -- memory utilization
34
+- **Storage** -- storage used/limit, storage breakdown (data/ibdata1/binlog), backup storage, server log storage, I/O utilization
35
+- **Connections** -- active connections, aborted connections, total connections, threads running
36
+- **Queries** -- queries (total/slow), DML statements (select/insert/update/delete), DDL statements
37
+- **InnoDB** -- buffer pool I/O (read requests/disk reads), buffer pool pages, data writes, row lock time/waits
38
+- **Replication** -- replication lag (replica/HA), HA status (I/O/SQL), replica status
39
+- **Network** -- network traffic (in/out)
40
+- **Health** -- uptime, deadlocks, lock timeouts
41
+
42
+
43
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
44
45
46
This collector is supported on all platforms.
47
48
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
49
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
50
+The service principal or managed identity requires these Azure RBAC roles:
51
+
52
+| Role | Purpose | Scope |
53
+|:-----|:--------|:------|
54
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
55
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
56
57
58
### Default Behavior
59
60
#### Auto-Detection
61
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
62
+The collector has two discovery phases:
63
+
64
+**Bootstrap (first run)**
65
+
66
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
67
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
68
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
69
+- A single job can monitor multiple subscriptions.
70
+
71
+**Runtime (periodic refresh)**
72
+
73
+- Periodically re-discovers resources for **already-active profile types only**.
74
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
75
+
76
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
77
78
79
#### Limits
80
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
81
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
82
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
83
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
84
85
86
#### Performance Impact
87
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
88
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
89
+
90
+**Default concurrency and batching limits:**
91
+
92
+| Setting | Default | Description |
93
+|:--------|:--------|:------------|
94
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
95
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
96
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
97
+
98
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
99
100
101
## Setup
@@ -78,25 +119,36 @@ UI configuration requires paid Netdata Cloud plan.
119
120
#### Create an Azure monitoring principal
121
81
-Create a service principal or use a managed identity with the following permissions:
122
+The collector requires a service principal or managed identity with two Azure RBAC roles:
123
+
124
+| Role | Purpose |
125
+|:-----|:--------|
126
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
127
+| **Reader** | Query Azure Resource Graph for resource discovery |
128
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
129
+**Option A: Service principal**
130
86
-For service principal authentication:
131
```bash
88
-# Create the service principal
132
+# Create service principal with Monitoring Reader role
133
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
134
--scopes /subscriptions/<subscription-id>
135
136
+# Add the Reader role for resource discovery
137
+az role assignment create --assignee <appId-from-above> \
138
+ --role "Reader" --scope /subscriptions/<subscription-id>
139
+
140
# Note the appId (client_id), password (client_secret), and tenant
141
```
142
95
-For managed identity (on Azure VMs, VMSS, or AKS):
143
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
144
+
145
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
146
+# Assign both roles to the VM's managed identity
147
az role assignment create --assignee <managed-identity-principal-id> \
148
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Reader" --scope /subscriptions/<subscription-id>
152
```
153
154
@@ -105,13 +157,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
157
158
#### Options
159
108
-The following options can be defined globally: update_every, autodetection_retry.
160
+The following options can be defined globally: `update_every`, `autodetection_retry`.
161
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+**Profile file locations:**
163
114
-User profile files with the same filename override stock profiles.
164
+| Type | Path |
165
+|:-----|:-----|
166
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
167
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
168
+
169
+User profile files with the same `id` as a stock profile override it.
170
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
171
172
173
<details open><summary>Config options</summary>
@@ -122,25 +178,103 @@ User profile files with the same filename override stock profiles.
178
|:------|:-----|:------------|:--------|:---------:|
179
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
180
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
181
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
182
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
186
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
191
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
192
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
195
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
196
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
197
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
198
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
199
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
204
+<a id="option-collection-query-offset"></a>
205
+##### query_offset
206
+
207
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
208
+
209
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
210
+
211
+- **Default (180s)** works for most services.
212
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
213
+- **Increase to 240-300s** if you still see gaps or missing data points.
214
+- **Do not set below 60s** -- metrics will likely be incomplete.
215
+
216
+
217
+<a id="option-authentication-auth-mode"></a>
218
+##### auth.mode
219
+
220
+Determines how the collector authenticates with Azure.
221
+
222
+| Mode | When to use | Required options |
223
+|:-----|:------------|:-----------------|
224
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
225
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
226
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
227
+
228
+
229
+<a id="option-discovery-discovery-mode"></a>
230
+##### discovery.mode
231
+
232
+Controls how the collector finds candidate Azure resources.
233
+
234
+| Mode | Behavior |
235
+|:-----|:---------|
236
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
237
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
238
+
239
+
240
+<a id="option-discovery-discovery-mode-query-kql"></a>
241
+##### discovery.mode_query.kql
242
+
243
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
244
+
245
+The query **must** project these five columns:
246
+
247
+| Column | Description |
248
+|:-------|:------------|
249
+| `id` | Full Azure resource ID (ARM format) |
250
+| `name` | Resource name |
251
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
252
+| `resourceGroup` | Resource group name |
253
+| `location` | Azure region |
254
+
255
+Example:
256
+
257
+```
258
+resources
259
+| where tags.env =~ "prod"
260
+| project id, name, type, resourceGroup, location
261
+```
262
+
263
+
264
+<a id="option-profiles-profiles-mode"></a>
265
+##### profiles.mode
266
+
267
+Controls how the collector decides which metric profiles to activate.
268
+
269
+| Mode | Behavior |
270
+|:-----|:---------|
271
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
272
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
273
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
274
+
275
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
276
+
277
+
278
279
</details>
280
@@ -182,14 +316,28 @@ sudo ./edit-config go.d/azure_monitor.conf
316
317
##### Examples
318
185
-###### Service principal (auto-discover all resources)
319
+###### Service principal with structured discovery
320
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
321
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
322
323
```yaml
324
jobs:
325
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ subscription_ids:
327
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
328
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
329
+ discovery:
330
+ mode: filters
331
+ mode_filters:
332
+ resource_groups:
333
+ - production-rg
334
+ regions:
335
+ - eastus
336
+ tags:
337
+ env:
338
+ - prod
339
+ profiles:
340
+ mode: auto
341
auth:
342
mode: service_principal
343
mode_service_principal:
@@ -198,59 +346,49 @@ jobs:
346
client_secret: "your-client-secret"
347
348
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
349
+###### Managed identity with exact profiles
350
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
351
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
352
353
<details open><summary>Config</summary>
354
355
```yaml
356
jobs:
357
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ subscription_ids:
359
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
361
+ mode: exact
362
+ mode_exact:
363
+ names:
364
+ - sql_database
365
+ - postgres_flexible
366
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
367
+ mode: managed_identity
368
369
```
370
</details>
371
241
-###### Filter by resource group
372
+###### Custom Azure Resource Graph KQL
373
243
-Only monitor resources in specific resource groups.
374
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
375
376
<details open><summary>Config</summary>
377
378
```yaml
379
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
380
+ - name: prod-query
381
+ subscription_ids:
382
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
383
+ discovery:
384
+ mode: query
385
+ mode_query:
386
+ kql: |
387
+ resources
388
+ | where tags.env =~ "prod"
389
+ | project id, name, type, resourceGroup, location
390
+ profiles:
391
+ mode: auto
392
auth:
393
mode: default
394
@@ -259,14 +397,15 @@ jobs:
397
398
###### Azure Government cloud
399
262
-Connect to Azure Government cloud environment.
400
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
401
402
<details open><summary>Config</summary>
403
404
```yaml
405
jobs:
406
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
+ subscription_ids:
408
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
409
cloud: government
410
auth:
411
mode: service_principal
@@ -329,6 +468,7 @@ Labels:
468
| region | The Azure region where the resource is deployed. |
469
| resource_type | The Azure resource type identifier. |
470
| profile | The Azure Monitor profile id. |
471
+| subscription_id | The Azure subscription identifier. |
472
| resource_uid | The unique Azure resource identifier. |
473
474
Metrics:
@@ -440,31 +580,46 @@ docker logs netdata 2>&1 | grep azure_monitor
580
581
### No metrics are collected
582
443
-Verify the following:
444
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
445
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
446
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
447
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
583
+Check the following:
584
+
585
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
586
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
587
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
588
+- **Collector logs** -- Check for authentication or API errors:
589
+ ```bash
590
+ # systemd
591
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
592
+ # non-systemd
593
+ grep azure_monitor /var/log/netdata/collector.log
594
+ ```
595
596
597
### Missing metrics for some resource types
598
452
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
453
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
454
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
455
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
599
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
600
+
601
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
602
+- **Verify a built-in profile exists** -- List available profiles:
603
+ ```bash
604
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
605
+ ```
606
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
607
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
608
609
458
-### Metrics appear delayed
610
+### Charts have gaps or incomplete data
611
460
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
461
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
462
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
612
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
613
+
614
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
615
+- Slower time-grain batches automatically use a larger effective offset when needed.
616
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
617
618
619
### Authentication errors in sovereign clouds
620
621
For Azure Government or Azure China clouds, set the `cloud` parameter:
622
+
623
- Azure Government: `cloud: government`
624
- Azure China (21Vianet): `cloud: china`
625
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_nat_gateway.md
+237
-87
@@ -21,40 +21,76 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor NAT Gateway including byte and packet counts, connection counts, dropped packets, total SNAT connection counts, and datapath availability.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure NAT Gateway with metrics covering:
31
+
32
+- **Throughput** -- byte and packet throughput
33
+- **Connections** -- SNAT connections, total SNAT connections
34
+- **Drops** -- dropped packets
35
+- **Availability** -- datapath availability percentage
36
+
37
+
38
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
39
40
41
This collector is supported on all platforms.
42
43
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
44
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
45
+The service principal or managed identity requires these Azure RBAC roles:
46
+
47
+| Role | Purpose | Scope |
48
+|:-----|:--------|:------|
49
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
50
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
51
52
53
### Default Behavior
54
55
#### Auto-Detection
56
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
57
+The collector has two discovery phases:
58
+
59
+**Bootstrap (first run)**
60
+
61
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
62
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
63
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
64
+- A single job can monitor multiple subscriptions.
65
+
66
+**Runtime (periodic refresh)**
67
+
68
+- Periodically re-discovers resources for **already-active profile types only**.
69
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
70
+
71
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
72
73
74
#### Limits
75
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
76
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
77
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
78
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
79
80
81
#### Performance Impact
82
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
83
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
84
+
85
+**Default concurrency and batching limits:**
86
+
87
+| Setting | Default | Description |
88
+|:--------|:--------|:------------|
89
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
90
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
91
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
92
+
93
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
94
95
96
## Setup
@@ -78,25 +114,36 @@ UI configuration requires paid Netdata Cloud plan.
114
115
#### Create an Azure monitoring principal
116
81
-Create a service principal or use a managed identity with the following permissions:
117
+The collector requires a service principal or managed identity with two Azure RBAC roles:
118
+
119
+| Role | Purpose |
120
+|:-----|:--------|
121
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
122
+| **Reader** | Query Azure Resource Graph for resource discovery |
123
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
124
+**Option A: Service principal**
125
86
-For service principal authentication:
126
```bash
88
-# Create the service principal
127
+# Create service principal with Monitoring Reader role
128
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
129
--scopes /subscriptions/<subscription-id>
130
131
+# Add the Reader role for resource discovery
132
+az role assignment create --assignee <appId-from-above> \
133
+ --role "Reader" --scope /subscriptions/<subscription-id>
134
+
135
# Note the appId (client_id), password (client_secret), and tenant
136
```
137
95
-For managed identity (on Azure VMs, VMSS, or AKS):
138
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
139
+
140
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
141
+# Assign both roles to the VM's managed identity
142
az role assignment create --assignee <managed-identity-principal-id> \
143
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
144
+
145
+az role assignment create --assignee <managed-identity-principal-id> \
146
+ --role "Reader" --scope /subscriptions/<subscription-id>
147
```
148
149
@@ -105,13 +152,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
152
153
#### Options
154
108
-The following options can be defined globally: update_every, autodetection_retry.
155
+The following options can be defined globally: `update_every`, `autodetection_retry`.
156
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
157
+**Profile file locations:**
158
114
-User profile files with the same filename override stock profiles.
159
+| Type | Path |
160
+|:-----|:-----|
161
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
162
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
163
+
164
+User profile files with the same `id` as a stock profile override it.
165
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
166
167
168
<details open><summary>Config options</summary>
@@ -122,25 +173,103 @@ User profile files with the same filename override stock profiles.
173
|:------|:-----|:------------|:--------|:---------:|
174
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
175
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
176
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
177
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
180
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
181
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
184
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
185
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
186
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
187
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
190
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
191
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
192
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
193
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
194
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
195
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
196
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
197
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
198
199
+<a id="option-collection-query-offset"></a>
200
+##### query_offset
201
+
202
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
203
+
204
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
205
+
206
+- **Default (180s)** works for most services.
207
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
208
+- **Increase to 240-300s** if you still see gaps or missing data points.
209
+- **Do not set below 60s** -- metrics will likely be incomplete.
210
+
211
+
212
+<a id="option-authentication-auth-mode"></a>
213
+##### auth.mode
214
+
215
+Determines how the collector authenticates with Azure.
216
+
217
+| Mode | When to use | Required options |
218
+|:-----|:------------|:-----------------|
219
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
220
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
221
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
222
+
223
+
224
+<a id="option-discovery-discovery-mode"></a>
225
+##### discovery.mode
226
+
227
+Controls how the collector finds candidate Azure resources.
228
+
229
+| Mode | Behavior |
230
+|:-----|:---------|
231
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
232
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
233
+
234
+
235
+<a id="option-discovery-discovery-mode-query-kql"></a>
236
+##### discovery.mode_query.kql
237
+
238
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
239
+
240
+The query **must** project these five columns:
241
+
242
+| Column | Description |
243
+|:-------|:------------|
244
+| `id` | Full Azure resource ID (ARM format) |
245
+| `name` | Resource name |
246
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
247
+| `resourceGroup` | Resource group name |
248
+| `location` | Azure region |
249
+
250
+Example:
251
+
252
+```
253
+resources
254
+| where tags.env =~ "prod"
255
+| project id, name, type, resourceGroup, location
256
+```
257
+
258
+
259
+<a id="option-profiles-profiles-mode"></a>
260
+##### profiles.mode
261
+
262
+Controls how the collector decides which metric profiles to activate.
263
+
264
+| Mode | Behavior |
265
+|:-----|:---------|
266
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
267
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
268
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
269
+
270
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
271
+
272
+
273
274
</details>
275
@@ -182,14 +311,28 @@ sudo ./edit-config go.d/azure_monitor.conf
311
312
##### Examples
313
185
-###### Service principal (auto-discover all resources)
314
+###### Service principal with structured discovery
315
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
316
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
317
318
```yaml
319
jobs:
320
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
321
+ subscription_ids:
322
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
323
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
324
+ discovery:
325
+ mode: filters
326
+ mode_filters:
327
+ resource_groups:
328
+ - production-rg
329
+ regions:
330
+ - eastus
331
+ tags:
332
+ env:
333
+ - prod
334
+ profiles:
335
+ mode: auto
336
auth:
337
mode: service_principal
338
mode_service_principal:
@@ -198,59 +341,49 @@ jobs:
341
client_secret: "your-client-secret"
342
343
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
344
+###### Managed identity with exact profiles
345
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
346
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
347
348
<details open><summary>Config</summary>
349
350
```yaml
351
jobs:
352
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
353
+ subscription_ids:
354
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
356
+ mode: exact
357
+ mode_exact:
358
+ names:
359
+ - sql_database
360
+ - postgres_flexible
361
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
362
+ mode: managed_identity
363
364
```
365
</details>
366
241
-###### Filter by resource group
367
+###### Custom Azure Resource Graph KQL
368
243
-Only monitor resources in specific resource groups.
369
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
370
371
<details open><summary>Config</summary>
372
373
```yaml
374
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
375
+ - name: prod-query
376
+ subscription_ids:
377
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
378
+ discovery:
379
+ mode: query
380
+ mode_query:
381
+ kql: |
382
+ resources
383
+ | where tags.env =~ "prod"
384
+ | project id, name, type, resourceGroup, location
385
+ profiles:
386
+ mode: auto
387
auth:
388
mode: default
389
@@ -259,14 +392,15 @@ jobs:
392
393
###### Azure Government cloud
394
262
-Connect to Azure Government cloud environment.
395
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
396
397
<details open><summary>Config</summary>
398
399
```yaml
400
jobs:
401
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
402
+ subscription_ids:
403
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
cloud: government
405
auth:
406
mode: service_principal
@@ -314,6 +448,7 @@ Labels:
448
| region | The Azure region where the resource is deployed. |
449
| resource_type | The Azure resource type identifier. |
450
| profile | The Azure Monitor profile id. |
451
+| subscription_id | The Azure subscription identifier. |
452
| resource_uid | The unique Azure resource identifier. |
453
454
Metrics:
@@ -398,31 +533,46 @@ docker logs netdata 2>&1 | grep azure_monitor
533
534
### No metrics are collected
535
401
-Verify the following:
402
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
403
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
404
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
405
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
536
+Check the following:
537
+
538
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
539
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
540
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
541
+- **Collector logs** -- Check for authentication or API errors:
542
+ ```bash
543
+ # systemd
544
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
545
+ # non-systemd
546
+ grep azure_monitor /var/log/netdata/collector.log
547
+ ```
548
549
550
### Missing metrics for some resource types
551
410
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
411
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
412
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
413
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
552
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
553
+
554
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
555
+- **Verify a built-in profile exists** -- List available profiles:
556
+ ```bash
557
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
558
+ ```
559
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
560
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
561
562
416
-### Metrics appear delayed
563
+### Charts have gaps or incomplete data
564
418
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
419
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
420
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
565
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
566
+
567
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
568
+- Slower time-grain batches automatically use a larger effective offset when needed.
569
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
570
571
572
### Authentication errors in sovereign clouds
573
574
For Azure Government or Azure China clouds, set the `cloud` parameter:
575
+
576
- Azure Government: `cloud: government`
577
- Azure China (21Vianet): `cloud: china`
578
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_postgresql_flexible_server.md
+242
-87
@@ -21,40 +21,81 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor PostgreSQL Flexible Server including active connections, transaction rates, replication lag, storage and backup utilization, CPU and memory usage, IO throughput, autovacuum activity, PgBouncer connection pooling, database sessions, and burstable instance CPU credits.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure PostgreSQL Flexible Server with metrics covering:
31
+
32
+- **Compute** -- CPU utilization, burstable CPU credits (consumed/remaining)
33
+- **Memory** -- memory utilization
34
+- **Storage** -- storage used/free, backup storage, WAL storage, database size, disk queue depth and saturation
35
+- **I/O** -- IOPS (read/write), disk throughput, temp bytes/files
36
+- **Connections** -- active connections, connection rate, max connections, PgBouncer client/server/pooled connections
37
+- **Database** -- transaction rate, commits/rollbacks, tuple reads/writes, replication lag (time/bytes), deadlocks
38
+- **Maintenance** -- autovacuum operations, table coverage, bloat percentage, buffer cache hit rate
39
+- **Sessions** -- sessions by state and wait event type, backend count
40
+- **Availability** -- database alive state
41
+
42
+
43
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
44
45
46
This collector is supported on all platforms.
47
48
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
49
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
50
+The service principal or managed identity requires these Azure RBAC roles:
51
+
52
+| Role | Purpose | Scope |
53
+|:-----|:--------|:------|
54
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
55
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
56
57
58
### Default Behavior
59
60
#### Auto-Detection
61
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
62
+The collector has two discovery phases:
63
+
64
+**Bootstrap (first run)**
65
+
66
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
67
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
68
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
69
+- A single job can monitor multiple subscriptions.
70
+
71
+**Runtime (periodic refresh)**
72
+
73
+- Periodically re-discovers resources for **already-active profile types only**.
74
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
75
+
76
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
77
78
79
#### Limits
80
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
81
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
82
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
83
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
84
85
86
#### Performance Impact
87
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
88
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
89
+
90
+**Default concurrency and batching limits:**
91
+
92
+| Setting | Default | Description |
93
+|:--------|:--------|:------------|
94
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
95
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
96
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
97
+
98
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
99
100
101
## Setup
@@ -78,25 +119,36 @@ UI configuration requires paid Netdata Cloud plan.
119
120
#### Create an Azure monitoring principal
121
81
-Create a service principal or use a managed identity with the following permissions:
122
+The collector requires a service principal or managed identity with two Azure RBAC roles:
123
+
124
+| Role | Purpose |
125
+|:-----|:--------|
126
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
127
+| **Reader** | Query Azure Resource Graph for resource discovery |
128
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
129
+**Option A: Service principal**
130
86
-For service principal authentication:
131
```bash
88
-# Create the service principal
132
+# Create service principal with Monitoring Reader role
133
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
134
--scopes /subscriptions/<subscription-id>
135
136
+# Add the Reader role for resource discovery
137
+az role assignment create --assignee <appId-from-above> \
138
+ --role "Reader" --scope /subscriptions/<subscription-id>
139
+
140
# Note the appId (client_id), password (client_secret), and tenant
141
```
142
95
-For managed identity (on Azure VMs, VMSS, or AKS):
143
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
144
+
145
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
146
+# Assign both roles to the VM's managed identity
147
az role assignment create --assignee <managed-identity-principal-id> \
148
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
149
+
150
+az role assignment create --assignee <managed-identity-principal-id> \
151
+ --role "Reader" --scope /subscriptions/<subscription-id>
152
```
153
154
@@ -105,13 +157,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
157
158
#### Options
159
108
-The following options can be defined globally: update_every, autodetection_retry.
160
+The following options can be defined globally: `update_every`, `autodetection_retry`.
161
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
162
+**Profile file locations:**
163
114
-User profile files with the same filename override stock profiles.
164
+| Type | Path |
165
+|:-----|:-----|
166
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
167
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
168
+
169
+User profile files with the same `id` as a stock profile override it.
170
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
171
172
173
<details open><summary>Config options</summary>
@@ -122,25 +178,103 @@ User profile files with the same filename override stock profiles.
178
|:------|:-----|:------------|:--------|:---------:|
179
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
180
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
181
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
182
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
185
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
186
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
189
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
190
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
191
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
192
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
195
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
196
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
197
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
198
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
199
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
200
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
201
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
202
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
203
204
+<a id="option-collection-query-offset"></a>
205
+##### query_offset
206
+
207
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
208
+
209
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
210
+
211
+- **Default (180s)** works for most services.
212
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
213
+- **Increase to 240-300s** if you still see gaps or missing data points.
214
+- **Do not set below 60s** -- metrics will likely be incomplete.
215
+
216
+
217
+<a id="option-authentication-auth-mode"></a>
218
+##### auth.mode
219
+
220
+Determines how the collector authenticates with Azure.
221
+
222
+| Mode | When to use | Required options |
223
+|:-----|:------------|:-----------------|
224
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
225
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
226
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
227
+
228
+
229
+<a id="option-discovery-discovery-mode"></a>
230
+##### discovery.mode
231
+
232
+Controls how the collector finds candidate Azure resources.
233
+
234
+| Mode | Behavior |
235
+|:-----|:---------|
236
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
237
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
238
+
239
+
240
+<a id="option-discovery-discovery-mode-query-kql"></a>
241
+##### discovery.mode_query.kql
242
+
243
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
244
+
245
+The query **must** project these five columns:
246
+
247
+| Column | Description |
248
+|:-------|:------------|
249
+| `id` | Full Azure resource ID (ARM format) |
250
+| `name` | Resource name |
251
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
252
+| `resourceGroup` | Resource group name |
253
+| `location` | Azure region |
254
+
255
+Example:
256
+
257
+```
258
+resources
259
+| where tags.env =~ "prod"
260
+| project id, name, type, resourceGroup, location
261
+```
262
+
263
+
264
+<a id="option-profiles-profiles-mode"></a>
265
+##### profiles.mode
266
+
267
+Controls how the collector decides which metric profiles to activate.
268
+
269
+| Mode | Behavior |
270
+|:-----|:---------|
271
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
272
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
273
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
274
+
275
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
276
+
277
+
278
279
</details>
280
@@ -182,14 +316,28 @@ sudo ./edit-config go.d/azure_monitor.conf
316
317
##### Examples
318
185
-###### Service principal (auto-discover all resources)
319
+###### Service principal with structured discovery
320
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
321
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
322
323
```yaml
324
jobs:
325
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
326
+ subscription_ids:
327
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
328
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
329
+ discovery:
330
+ mode: filters
331
+ mode_filters:
332
+ resource_groups:
333
+ - production-rg
334
+ regions:
335
+ - eastus
336
+ tags:
337
+ env:
338
+ - prod
339
+ profiles:
340
+ mode: auto
341
auth:
342
mode: service_principal
343
mode_service_principal:
@@ -198,59 +346,49 @@ jobs:
346
client_secret: "your-client-secret"
347
348
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
349
+###### Managed identity with exact profiles
350
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
351
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
352
353
<details open><summary>Config</summary>
354
355
```yaml
356
jobs:
357
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
+ subscription_ids:
359
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
360
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
361
+ mode: exact
362
+ mode_exact:
363
+ names:
364
+ - sql_database
365
+ - postgres_flexible
366
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
367
+ mode: managed_identity
368
369
```
370
</details>
371
241
-###### Filter by resource group
372
+###### Custom Azure Resource Graph KQL
373
243
-Only monitor resources in specific resource groups.
374
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
375
376
<details open><summary>Config</summary>
377
378
```yaml
379
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
380
+ - name: prod-query
381
+ subscription_ids:
382
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
383
+ discovery:
384
+ mode: query
385
+ mode_query:
386
+ kql: |
387
+ resources
388
+ | where tags.env =~ "prod"
389
+ | project id, name, type, resourceGroup, location
390
+ profiles:
391
+ mode: auto
392
auth:
393
mode: default
394
@@ -259,14 +397,15 @@ jobs:
397
398
###### Azure Government cloud
399
262
-Connect to Azure Government cloud environment.
400
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
401
402
<details open><summary>Config</summary>
403
404
```yaml
405
jobs:
406
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
+ subscription_ids:
408
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
409
cloud: government
410
auth:
411
mode: service_principal
@@ -329,6 +468,7 @@ Labels:
468
| region | The Azure region where the resource is deployed. |
469
| resource_type | The Azure resource type identifier. |
470
| profile | The Azure Monitor profile id. |
471
+| subscription_id | The Azure subscription identifier. |
472
| resource_uid | The unique Azure resource identifier. |
473
474
Metrics:
@@ -450,31 +590,46 @@ docker logs netdata 2>&1 | grep azure_monitor
590
591
### No metrics are collected
592
453
-Verify the following:
454
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
455
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
456
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
457
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
593
+Check the following:
594
+
595
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
596
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
597
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
598
+- **Collector logs** -- Check for authentication or API errors:
599
+ ```bash
600
+ # systemd
601
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
602
+ # non-systemd
603
+ grep azure_monitor /var/log/netdata/collector.log
604
+ ```
605
606
607
### Missing metrics for some resource types
608
462
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
463
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
464
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
465
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
609
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
610
+
611
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
612
+- **Verify a built-in profile exists** -- List available profiles:
613
+ ```bash
614
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
615
+ ```
616
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
617
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
618
619
468
-### Metrics appear delayed
620
+### Charts have gaps or incomplete data
621
470
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
471
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
472
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
622
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
623
+
624
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
625
+- Slower time-grain batches automatically use a larger effective offset when needed.
626
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
627
628
629
### Authentication errors in sovereign clouds
630
631
For Azure Government or Azure China clouds, set the `cloud` parameter:
632
+
633
- Azure Government: `cloud: government`
634
- Azure China (21Vianet): `cloud: china`
635
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_service_bus_namespace.md
+241
-87
@@ -21,40 +21,80 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Service Bus namespaces including incoming and outgoing message rates, active connections, active and dead-lettered message counts, scheduled message counts, completed and abandoned requests, server errors, throttled requests, CPU and memory utilization, and pending checkpoint operations.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Service Bus with metrics covering:
31
+
32
+- **Messages** -- message flow (in/out), active messages, dead-lettered messages, scheduled messages, queue depth
33
+- **Throughput** -- data throughput (in/out bytes per second)
34
+- **Operations** -- completed and abandoned message operations, send latency
35
+- **Connections** -- active connections, connection events (opened/closed)
36
+- **Requests** -- incoming and successful request rates
37
+- **Errors** -- server errors, user errors, throttled requests
38
+- **Replication** -- replication lag (messages and duration)
39
+- **Resources** -- namespace size, CPU and memory utilization, pending checkpoint operations
40
+
41
+
42
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
43
44
45
This collector is supported on all platforms.
46
47
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
48
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
49
+The service principal or managed identity requires these Azure RBAC roles:
50
+
51
+| Role | Purpose | Scope |
52
+|:-----|:--------|:------|
53
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
54
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
55
56
57
### Default Behavior
58
59
#### Auto-Detection
60
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
61
+The collector has two discovery phases:
62
+
63
+**Bootstrap (first run)**
64
+
65
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
66
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
67
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
68
+- A single job can monitor multiple subscriptions.
69
+
70
+**Runtime (periodic refresh)**
71
+
72
+- Periodically re-discovers resources for **already-active profile types only**.
73
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
74
+
75
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
76
77
78
#### Limits
79
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
80
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
81
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
82
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
83
84
85
#### Performance Impact
86
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
87
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
88
+
89
+**Default concurrency and batching limits:**
90
+
91
+| Setting | Default | Description |
92
+|:--------|:--------|:------------|
93
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
94
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
95
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
96
+
97
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
98
99
100
## Setup
@@ -78,25 +118,36 @@ UI configuration requires paid Netdata Cloud plan.
118
119
#### Create an Azure monitoring principal
120
81
-Create a service principal or use a managed identity with the following permissions:
121
+The collector requires a service principal or managed identity with two Azure RBAC roles:
122
+
123
+| Role | Purpose |
124
+|:-----|:--------|
125
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
126
+| **Reader** | Query Azure Resource Graph for resource discovery |
127
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
128
+**Option A: Service principal**
129
86
-For service principal authentication:
130
```bash
88
-# Create the service principal
131
+# Create service principal with Monitoring Reader role
132
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
133
--scopes /subscriptions/<subscription-id>
134
135
+# Add the Reader role for resource discovery
136
+az role assignment create --assignee <appId-from-above> \
137
+ --role "Reader" --scope /subscriptions/<subscription-id>
138
+
139
# Note the appId (client_id), password (client_secret), and tenant
140
```
141
95
-For managed identity (on Azure VMs, VMSS, or AKS):
142
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
143
+
144
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
145
+# Assign both roles to the VM's managed identity
146
az role assignment create --assignee <managed-identity-principal-id> \
147
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
148
+
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Reader" --scope /subscriptions/<subscription-id>
151
```
152
153
@@ -105,13 +156,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
156
157
#### Options
158
108
-The following options can be defined globally: update_every, autodetection_retry.
159
+The following options can be defined globally: `update_every`, `autodetection_retry`.
160
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
161
+**Profile file locations:**
162
114
-User profile files with the same filename override stock profiles.
163
+| Type | Path |
164
+|:-----|:-----|
165
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
166
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
167
+
168
+User profile files with the same `id` as a stock profile override it.
169
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
170
171
172
<details open><summary>Config options</summary>
@@ -122,25 +177,103 @@ User profile files with the same filename override stock profiles.
177
|:------|:-----|:------------|:--------|:---------:|
178
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
179
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
181
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
184
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
185
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
190
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
191
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
194
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
195
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
196
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
197
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
198
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
199
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
200
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
201
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
203
+<a id="option-collection-query-offset"></a>
204
+##### query_offset
205
+
206
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
207
+
208
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
209
+
210
+- **Default (180s)** works for most services.
211
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
212
+- **Increase to 240-300s** if you still see gaps or missing data points.
213
+- **Do not set below 60s** -- metrics will likely be incomplete.
214
+
215
+
216
+<a id="option-authentication-auth-mode"></a>
217
+##### auth.mode
218
+
219
+Determines how the collector authenticates with Azure.
220
+
221
+| Mode | When to use | Required options |
222
+|:-----|:------------|:-----------------|
223
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
224
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
225
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
226
+
227
+
228
+<a id="option-discovery-discovery-mode"></a>
229
+##### discovery.mode
230
+
231
+Controls how the collector finds candidate Azure resources.
232
+
233
+| Mode | Behavior |
234
+|:-----|:---------|
235
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
236
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
237
+
238
+
239
+<a id="option-discovery-discovery-mode-query-kql"></a>
240
+##### discovery.mode_query.kql
241
+
242
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
243
+
244
+The query **must** project these five columns:
245
+
246
+| Column | Description |
247
+|:-------|:------------|
248
+| `id` | Full Azure resource ID (ARM format) |
249
+| `name` | Resource name |
250
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
251
+| `resourceGroup` | Resource group name |
252
+| `location` | Azure region |
253
+
254
+Example:
255
+
256
+```
257
+resources
258
+| where tags.env =~ "prod"
259
+| project id, name, type, resourceGroup, location
260
+```
261
+
262
+
263
+<a id="option-profiles-profiles-mode"></a>
264
+##### profiles.mode
265
+
266
+Controls how the collector decides which metric profiles to activate.
267
+
268
+| Mode | Behavior |
269
+|:-----|:---------|
270
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
271
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
272
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
273
+
274
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
275
+
276
+
277
278
</details>
279
@@ -182,14 +315,28 @@ sudo ./edit-config go.d/azure_monitor.conf
315
316
##### Examples
317
185
-###### Service principal (auto-discover all resources)
318
+###### Service principal with structured discovery
319
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
320
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
321
322
```yaml
323
jobs:
324
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ subscription_ids:
326
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
327
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
328
+ discovery:
329
+ mode: filters
330
+ mode_filters:
331
+ resource_groups:
332
+ - production-rg
333
+ regions:
334
+ - eastus
335
+ tags:
336
+ env:
337
+ - prod
338
+ profiles:
339
+ mode: auto
340
auth:
341
mode: service_principal
342
mode_service_principal:
@@ -198,59 +345,49 @@ jobs:
345
client_secret: "your-client-secret"
346
347
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
348
+###### Managed identity with exact profiles
349
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
350
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
351
352
<details open><summary>Config</summary>
353
354
```yaml
355
jobs:
356
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ subscription_ids:
358
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
360
+ mode: exact
361
+ mode_exact:
362
+ names:
363
+ - sql_database
364
+ - postgres_flexible
365
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
366
+ mode: managed_identity
367
368
```
369
</details>
370
241
-###### Filter by resource group
371
+###### Custom Azure Resource Graph KQL
372
243
-Only monitor resources in specific resource groups.
373
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
374
375
<details open><summary>Config</summary>
376
377
```yaml
378
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
379
+ - name: prod-query
380
+ subscription_ids:
381
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
382
+ discovery:
383
+ mode: query
384
+ mode_query:
385
+ kql: |
386
+ resources
387
+ | where tags.env =~ "prod"
388
+ | project id, name, type, resourceGroup, location
389
+ profiles:
390
+ mode: auto
391
auth:
392
mode: default
393
@@ -259,14 +396,15 @@ jobs:
396
397
###### Azure Government cloud
398
262
-Connect to Azure Government cloud environment.
399
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
400
401
<details open><summary>Config</summary>
402
403
```yaml
404
jobs:
405
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
+ subscription_ids:
407
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
cloud: government
409
auth:
410
mode: service_principal
@@ -322,6 +460,7 @@ Labels:
460
| region | The Azure region where the resource is deployed. |
461
| resource_type | The Azure resource type identifier. |
462
| profile | The Azure Monitor profile id. |
463
+| subscription_id | The Azure subscription identifier. |
464
| resource_uid | The unique Azure resource identifier. |
465
466
Metrics:
@@ -416,31 +555,46 @@ docker logs netdata 2>&1 | grep azure_monitor
555
556
### No metrics are collected
557
419
-Verify the following:
420
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
421
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
422
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
423
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
558
+Check the following:
559
+
560
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
561
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
562
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
563
+- **Collector logs** -- Check for authentication or API errors:
564
+ ```bash
565
+ # systemd
566
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
567
+ # non-systemd
568
+ grep azure_monitor /var/log/netdata/collector.log
569
+ ```
570
571
572
### Missing metrics for some resource types
573
428
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
429
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
430
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
431
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
574
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
575
+
576
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
577
+- **Verify a built-in profile exists** -- List available profiles:
578
+ ```bash
579
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
580
+ ```
581
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
582
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
583
584
434
-### Metrics appear delayed
585
+### Charts have gaps or incomplete data
586
436
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
437
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
438
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
587
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
588
+
589
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
590
+- Slower time-grain batches automatically use a larger effective offset when needed.
591
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
592
593
594
### Authentication errors in sovereign clouds
595
596
For Azure Government or Azure China clouds, set the `cloud` parameter:
597
+
598
- Azure Government: `cloud: government`
599
- Azure China (21Vianet): `cloud: china`
600
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_database.md
+240
-87
@@ -21,40 +21,79 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor SQL Database performance including CPU and DTU utilization, storage consumption, active sessions and workers, deadlocks, IO rates, tempdb usage, in-memory OLTP storage, and serverless auto-pause and billing metrics.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure SQL Database with metrics covering:
31
+
32
+- **CPU and DTU** -- CPU utilization (average/max), instance CPU, DTU consumption, vCore usage
33
+- **Memory** -- instance memory utilization
34
+- **Storage** -- data and allocated storage, storage utilization, tempdb size, in-memory OLTP storage
35
+- **I/O** -- data read and log write utilization, tempdb log utilization
36
+- **Connections** -- successful, failed, and firewall-blocked connections, active sessions and workers
37
+- **Availability** -- database availability percentage
38
+- **Advanced** -- deadlocks, replication lag, serverless CPU/memory/billing, ledger digest, free tier usage
39
+
40
+
41
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
42
43
44
This collector is supported on all platforms.
45
46
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
47
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
48
+The service principal or managed identity requires these Azure RBAC roles:
49
+
50
+| Role | Purpose | Scope |
51
+|:-----|:--------|:------|
52
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
53
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
54
55
56
### Default Behavior
57
58
#### Auto-Detection
59
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
60
+The collector has two discovery phases:
61
+
62
+**Bootstrap (first run)**
63
+
64
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
65
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
66
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
67
+- A single job can monitor multiple subscriptions.
68
+
69
+**Runtime (periodic refresh)**
70
+
71
+- Periodically re-discovers resources for **already-active profile types only**.
72
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
73
+
74
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
75
76
77
#### Limits
78
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
79
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
80
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
81
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
82
83
84
#### Performance Impact
85
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
86
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
87
+
88
+**Default concurrency and batching limits:**
89
+
90
+| Setting | Default | Description |
91
+|:--------|:--------|:------------|
92
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
93
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
94
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
95
+
96
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
97
98
99
## Setup
@@ -78,25 +117,36 @@ UI configuration requires paid Netdata Cloud plan.
117
118
#### Create an Azure monitoring principal
119
81
-Create a service principal or use a managed identity with the following permissions:
120
+The collector requires a service principal or managed identity with two Azure RBAC roles:
121
+
122
+| Role | Purpose |
123
+|:-----|:--------|
124
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
125
+| **Reader** | Query Azure Resource Graph for resource discovery |
126
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
127
+**Option A: Service principal**
128
86
-For service principal authentication:
129
```bash
88
-# Create the service principal
130
+# Create service principal with Monitoring Reader role
131
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
132
--scopes /subscriptions/<subscription-id>
133
134
+# Add the Reader role for resource discovery
135
+az role assignment create --assignee <appId-from-above> \
136
+ --role "Reader" --scope /subscriptions/<subscription-id>
137
+
138
# Note the appId (client_id), password (client_secret), and tenant
139
```
140
95
-For managed identity (on Azure VMs, VMSS, or AKS):
141
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
142
+
143
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
144
+# Assign both roles to the VM's managed identity
145
az role assignment create --assignee <managed-identity-principal-id> \
146
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
147
+
148
+az role assignment create --assignee <managed-identity-principal-id> \
149
+ --role "Reader" --scope /subscriptions/<subscription-id>
150
```
151
152
@@ -105,13 +155,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
155
156
#### Options
157
108
-The following options can be defined globally: update_every, autodetection_retry.
158
+The following options can be defined globally: `update_every`, `autodetection_retry`.
159
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
160
+**Profile file locations:**
161
114
-User profile files with the same filename override stock profiles.
162
+| Type | Path |
163
+|:-----|:-----|
164
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
165
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
166
+
167
+User profile files with the same `id` as a stock profile override it.
168
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
169
170
171
<details open><summary>Config options</summary>
@@ -122,25 +176,103 @@ User profile files with the same filename override stock profiles.
176
|:------|:-----|:------------|:--------|:---------:|
177
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
178
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
179
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
180
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
183
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
184
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
187
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
188
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
189
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
190
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
193
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
194
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
195
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
196
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
197
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
198
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
199
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
200
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
201
202
+<a id="option-collection-query-offset"></a>
203
+##### query_offset
204
+
205
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
206
+
207
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
208
+
209
+- **Default (180s)** works for most services.
210
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
211
+- **Increase to 240-300s** if you still see gaps or missing data points.
212
+- **Do not set below 60s** -- metrics will likely be incomplete.
213
+
214
+
215
+<a id="option-authentication-auth-mode"></a>
216
+##### auth.mode
217
+
218
+Determines how the collector authenticates with Azure.
219
+
220
+| Mode | When to use | Required options |
221
+|:-----|:------------|:-----------------|
222
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
223
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
224
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
225
+
226
+
227
+<a id="option-discovery-discovery-mode"></a>
228
+##### discovery.mode
229
+
230
+Controls how the collector finds candidate Azure resources.
231
+
232
+| Mode | Behavior |
233
+|:-----|:---------|
234
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
235
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
236
+
237
+
238
+<a id="option-discovery-discovery-mode-query-kql"></a>
239
+##### discovery.mode_query.kql
240
+
241
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
242
+
243
+The query **must** project these five columns:
244
+
245
+| Column | Description |
246
+|:-------|:------------|
247
+| `id` | Full Azure resource ID (ARM format) |
248
+| `name` | Resource name |
249
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
250
+| `resourceGroup` | Resource group name |
251
+| `location` | Azure region |
252
+
253
+Example:
254
+
255
+```
256
+resources
257
+| where tags.env =~ "prod"
258
+| project id, name, type, resourceGroup, location
259
+```
260
+
261
+
262
+<a id="option-profiles-profiles-mode"></a>
263
+##### profiles.mode
264
+
265
+Controls how the collector decides which metric profiles to activate.
266
+
267
+| Mode | Behavior |
268
+|:-----|:---------|
269
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
270
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
271
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
272
+
273
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
274
+
275
+
276
277
</details>
278
@@ -182,14 +314,28 @@ sudo ./edit-config go.d/azure_monitor.conf
314
315
##### Examples
316
185
-###### Service principal (auto-discover all resources)
317
+###### Service principal with structured discovery
318
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
319
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
320
321
```yaml
322
jobs:
323
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
324
+ subscription_ids:
325
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
326
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
327
+ discovery:
328
+ mode: filters
329
+ mode_filters:
330
+ resource_groups:
331
+ - production-rg
332
+ regions:
333
+ - eastus
334
+ tags:
335
+ env:
336
+ - prod
337
+ profiles:
338
+ mode: auto
339
auth:
340
mode: service_principal
341
mode_service_principal:
@@ -198,59 +344,49 @@ jobs:
344
client_secret: "your-client-secret"
345
346
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
347
+###### Managed identity with exact profiles
348
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
349
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
350
351
<details open><summary>Config</summary>
352
353
```yaml
354
jobs:
355
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
+ subscription_ids:
357
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
358
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
359
+ mode: exact
360
+ mode_exact:
361
+ names:
362
+ - sql_database
363
+ - postgres_flexible
364
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
365
+ mode: managed_identity
366
367
```
368
</details>
369
241
-###### Filter by resource group
370
+###### Custom Azure Resource Graph KQL
371
243
-Only monitor resources in specific resource groups.
372
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
373
374
<details open><summary>Config</summary>
375
376
```yaml
377
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
378
+ - name: prod-query
379
+ subscription_ids:
380
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
381
+ discovery:
382
+ mode: query
383
+ mode_query:
384
+ kql: |
385
+ resources
386
+ | where tags.env =~ "prod"
387
+ | project id, name, type, resourceGroup, location
388
+ profiles:
389
+ mode: auto
390
auth:
391
mode: default
392
@@ -259,14 +395,15 @@ jobs:
395
396
###### Azure Government cloud
397
262
-Connect to Azure Government cloud environment.
398
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
399
400
<details open><summary>Config</summary>
401
402
```yaml
403
jobs:
404
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
+ subscription_ids:
406
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
407
cloud: government
408
auth:
409
mode: service_principal
@@ -329,6 +466,7 @@ Labels:
466
| region | The Azure region where the resource is deployed. |
467
| resource_type | The Azure resource type identifier. |
468
| profile | The Azure Monitor profile id. |
469
+| subscription_id | The Azure subscription identifier. |
470
| resource_uid | The unique Azure resource identifier. |
471
472
Metrics:
@@ -429,31 +567,46 @@ docker logs netdata 2>&1 | grep azure_monitor
567
568
### No metrics are collected
569
432
-Verify the following:
433
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
434
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
435
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
436
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
570
+Check the following:
571
+
572
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
573
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
574
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
575
+- **Collector logs** -- Check for authentication or API errors:
576
+ ```bash
577
+ # systemd
578
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
579
+ # non-systemd
580
+ grep azure_monitor /var/log/netdata/collector.log
581
+ ```
582
583
584
### Missing metrics for some resource types
585
441
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
442
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
443
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
444
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
586
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
587
+
588
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
589
+- **Verify a built-in profile exists** -- List available profiles:
590
+ ```bash
591
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
592
+ ```
593
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
594
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
595
596
447
-### Metrics appear delayed
597
+### Charts have gaps or incomplete data
598
449
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
450
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
451
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
599
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
600
+
601
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
602
+- Slower time-grain batches automatically use a larger effective offset when needed.
603
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
604
605
606
### Authentication errors in sovereign clouds
607
608
For Azure Government or Azure China clouds, set the `cloud` parameter:
609
+
610
- Azure Government: `cloud: government`
611
- Azure China (21Vianet): `cloud: china`
612
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_elastic_pool.md
+241
-87
@@ -21,40 +21,80 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor SQL Elastic Pool resource consumption including eDTU and CPU utilization, storage usage, active sessions and workers, IO rates, tempdb usage, and in-memory OLTP storage across all databases in the pool.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
+
28
+:::
29
+
30
+Monitor Azure SQL Elastic Pool with metrics covering:
31
+
32
+- **CPU and DTU** -- CPU utilization (average/max), instance CPU, DTU consumption, eDTU and vCore usage
33
+- **Memory** -- instance memory utilization
34
+- **Storage** -- data and allocated storage, storage utilization, tempdb size, in-memory OLTP storage
35
+- **I/O** -- data read and log write utilization, tempdb log utilization
36
+- **Sessions** -- active sessions and workers count, serverless CPU/memory utilization
37
+- **Billing** -- serverless billing (vCore-seconds)
38
+
39
+Metrics are aggregated across all databases in the pool.
40
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
41
+
42
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
43
44
45
This collector is supported on all platforms.
46
47
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
48
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
49
+The service principal or managed identity requires these Azure RBAC roles:
50
+
51
+| Role | Purpose | Scope |
52
+|:-----|:--------|:------|
53
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
54
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
55
56
57
### Default Behavior
58
59
#### Auto-Detection
60
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
61
+The collector has two discovery phases:
62
+
63
+**Bootstrap (first run)**
64
+
65
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
66
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
67
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
68
+- A single job can monitor multiple subscriptions.
69
+
70
+**Runtime (periodic refresh)**
71
+
72
+- Periodically re-discovers resources for **already-active profile types only**.
73
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
74
+
75
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
76
77
78
#### Limits
79
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
80
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
81
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
82
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
83
84
85
#### Performance Impact
86
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
87
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
88
+
89
+**Default concurrency and batching limits:**
90
+
91
+| Setting | Default | Description |
92
+|:--------|:--------|:------------|
93
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
94
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
95
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
96
+
97
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
98
99
100
## Setup
@@ -78,25 +118,36 @@ UI configuration requires paid Netdata Cloud plan.
118
119
#### Create an Azure monitoring principal
120
81
-Create a service principal or use a managed identity with the following permissions:
121
+The collector requires a service principal or managed identity with two Azure RBAC roles:
122
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
123
+| Role | Purpose |
124
+|:-----|:--------|
125
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
126
+| **Reader** | Query Azure Resource Graph for resource discovery |
127
+
128
+**Option A: Service principal**
129
86
-For service principal authentication:
130
```bash
88
-# Create the service principal
131
+# Create service principal with Monitoring Reader role
132
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
133
--scopes /subscriptions/<subscription-id>
134
135
+# Add the Reader role for resource discovery
136
+az role assignment create --assignee <appId-from-above> \
137
+ --role "Reader" --scope /subscriptions/<subscription-id>
138
+
139
# Note the appId (client_id), password (client_secret), and tenant
140
```
141
95
-For managed identity (on Azure VMs, VMSS, or AKS):
142
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
143
+
144
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
145
+# Assign both roles to the VM's managed identity
146
az role assignment create --assignee <managed-identity-principal-id> \
147
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
148
+
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Reader" --scope /subscriptions/<subscription-id>
151
```
152
153
@@ -105,13 +156,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
156
157
#### Options
158
108
-The following options can be defined globally: update_every, autodetection_retry.
159
+The following options can be defined globally: `update_every`, `autodetection_retry`.
160
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
161
+**Profile file locations:**
162
114
-User profile files with the same filename override stock profiles.
163
+| Type | Path |
164
+|:-----|:-----|
165
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
166
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
167
+
168
+User profile files with the same `id` as a stock profile override it.
169
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
170
171
172
<details open><summary>Config options</summary>
@@ -122,25 +177,103 @@ User profile files with the same filename override stock profiles.
177
|:------|:-----|:------------|:--------|:---------:|
178
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
179
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
181
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
184
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
185
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
190
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
191
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
194
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
195
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
196
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
197
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
198
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
199
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
200
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
201
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
203
+<a id="option-collection-query-offset"></a>
204
+##### query_offset
205
+
206
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
207
+
208
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
209
+
210
+- **Default (180s)** works for most services.
211
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
212
+- **Increase to 240-300s** if you still see gaps or missing data points.
213
+- **Do not set below 60s** -- metrics will likely be incomplete.
214
+
215
+
216
+<a id="option-authentication-auth-mode"></a>
217
+##### auth.mode
218
+
219
+Determines how the collector authenticates with Azure.
220
+
221
+| Mode | When to use | Required options |
222
+|:-----|:------------|:-----------------|
223
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
224
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
225
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
226
+
227
+
228
+<a id="option-discovery-discovery-mode"></a>
229
+##### discovery.mode
230
+
231
+Controls how the collector finds candidate Azure resources.
232
+
233
+| Mode | Behavior |
234
+|:-----|:---------|
235
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
236
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
237
+
238
+
239
+<a id="option-discovery-discovery-mode-query-kql"></a>
240
+##### discovery.mode_query.kql
241
+
242
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
243
+
244
+The query **must** project these five columns:
245
+
246
+| Column | Description |
247
+|:-------|:------------|
248
+| `id` | Full Azure resource ID (ARM format) |
249
+| `name` | Resource name |
250
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
251
+| `resourceGroup` | Resource group name |
252
+| `location` | Azure region |
253
+
254
+Example:
255
+
256
+```
257
+resources
258
+| where tags.env =~ "prod"
259
+| project id, name, type, resourceGroup, location
260
+```
261
+
262
+
263
+<a id="option-profiles-profiles-mode"></a>
264
+##### profiles.mode
265
+
266
+Controls how the collector decides which metric profiles to activate.
267
+
268
+| Mode | Behavior |
269
+|:-----|:---------|
270
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
271
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
272
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
273
+
274
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
275
+
276
+
277
278
</details>
279
@@ -182,14 +315,28 @@ sudo ./edit-config go.d/azure_monitor.conf
315
316
##### Examples
317
185
-###### Service principal (auto-discover all resources)
318
+###### Service principal with structured discovery
319
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
320
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
321
322
```yaml
323
jobs:
324
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ subscription_ids:
326
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
327
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
328
+ discovery:
329
+ mode: filters
330
+ mode_filters:
331
+ resource_groups:
332
+ - production-rg
333
+ regions:
334
+ - eastus
335
+ tags:
336
+ env:
337
+ - prod
338
+ profiles:
339
+ mode: auto
340
auth:
341
mode: service_principal
342
mode_service_principal:
@@ -198,59 +345,49 @@ jobs:
345
client_secret: "your-client-secret"
346
347
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
348
+###### Managed identity with exact profiles
349
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
350
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
351
352
<details open><summary>Config</summary>
353
354
```yaml
355
jobs:
356
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ subscription_ids:
358
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
360
+ mode: exact
361
+ mode_exact:
362
+ names:
363
+ - sql_database
364
+ - postgres_flexible
365
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
366
+ mode: managed_identity
367
368
```
369
</details>
370
241
-###### Filter by resource group
371
+###### Custom Azure Resource Graph KQL
372
243
-Only monitor resources in specific resource groups.
373
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
374
375
<details open><summary>Config</summary>
376
377
```yaml
378
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
379
+ - name: prod-query
380
+ subscription_ids:
381
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
382
+ discovery:
383
+ mode: query
384
+ mode_query:
385
+ kql: |
386
+ resources
387
+ | where tags.env =~ "prod"
388
+ | project id, name, type, resourceGroup, location
389
+ profiles:
390
+ mode: auto
391
auth:
392
mode: default
393
@@ -259,14 +396,15 @@ jobs:
396
397
###### Azure Government cloud
398
262
-Connect to Azure Government cloud environment.
399
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
400
401
<details open><summary>Config</summary>
402
403
```yaml
404
jobs:
405
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
+ subscription_ids:
407
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
cloud: government
409
auth:
410
mode: service_principal
@@ -324,6 +462,7 @@ Labels:
462
| region | The Azure region where the resource is deployed. |
463
| resource_type | The Azure resource type identifier. |
464
| profile | The Azure Monitor profile id. |
465
+| subscription_id | The Azure subscription identifier. |
466
| resource_uid | The unique Azure resource identifier. |
467
468
Metrics:
@@ -418,31 +557,46 @@ docker logs netdata 2>&1 | grep azure_monitor
557
558
### No metrics are collected
559
421
-Verify the following:
422
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
423
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
424
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
425
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
560
+Check the following:
561
+
562
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
563
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
564
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
565
+- **Collector logs** -- Check for authentication or API errors:
566
+ ```bash
567
+ # systemd
568
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
569
+ # non-systemd
570
+ grep azure_monitor /var/log/netdata/collector.log
571
+ ```
572
573
574
### Missing metrics for some resource types
575
430
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
431
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
432
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
433
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
576
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
577
+
578
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
579
+- **Verify a built-in profile exists** -- List available profiles:
580
+ ```bash
581
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
582
+ ```
583
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
584
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
585
586
436
-### Metrics appear delayed
587
+### Charts have gaps or incomplete data
588
438
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
439
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
440
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
589
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
590
+
591
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
592
+- Slower time-grain batches automatically use a larger effective offset when needed.
593
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
594
595
596
### Authentication errors in sovereign clouds
597
598
For Azure Government or Azure China clouds, set the `cloud` parameter:
599
+
600
- Azure Government: `cloud: government`
601
- Azure China (21Vianet): `cloud: china`
602
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_sql_managed_instance.md
+236
-87
@@ -21,40 +21,75 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor SQL Managed Instance performance including virtual core CPU utilization, storage consumption, IO throughput, and average request wait times.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure SQL Managed Instance with metrics covering:
31
+
32
+- **Compute** -- CPU utilization (average/max), virtual core count
33
+- **Storage** -- reserved and used storage
34
+- **I/O** -- read/write throughput, I/O request rate
35
+
36
+
37
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
38
39
40
This collector is supported on all platforms.
41
42
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
43
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
44
+The service principal or managed identity requires these Azure RBAC roles:
45
+
46
+| Role | Purpose | Scope |
47
+|:-----|:--------|:------|
48
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
49
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
50
51
52
### Default Behavior
53
54
#### Auto-Detection
55
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
56
+The collector has two discovery phases:
57
+
58
+**Bootstrap (first run)**
59
+
60
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
61
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
62
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
63
+- A single job can monitor multiple subscriptions.
64
+
65
+**Runtime (periodic refresh)**
66
+
67
+- Periodically re-discovers resources for **already-active profile types only**.
68
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
69
+
70
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
71
72
73
#### Limits
74
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
75
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
76
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
77
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
78
79
80
#### Performance Impact
81
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
82
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
83
+
84
+**Default concurrency and batching limits:**
85
+
86
+| Setting | Default | Description |
87
+|:--------|:--------|:------------|
88
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
89
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
90
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
91
+
92
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
93
94
95
## Setup
@@ -78,25 +113,36 @@ UI configuration requires paid Netdata Cloud plan.
113
114
#### Create an Azure monitoring principal
115
81
-Create a service principal or use a managed identity with the following permissions:
116
+The collector requires a service principal or managed identity with two Azure RBAC roles:
117
+
118
+| Role | Purpose |
119
+|:-----|:--------|
120
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
121
+| **Reader** | Query Azure Resource Graph for resource discovery |
122
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
123
+**Option A: Service principal**
124
86
-For service principal authentication:
125
```bash
88
-# Create the service principal
126
+# Create service principal with Monitoring Reader role
127
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
128
--scopes /subscriptions/<subscription-id>
129
130
+# Add the Reader role for resource discovery
131
+az role assignment create --assignee <appId-from-above> \
132
+ --role "Reader" --scope /subscriptions/<subscription-id>
133
+
134
# Note the appId (client_id), password (client_secret), and tenant
135
```
136
95
-For managed identity (on Azure VMs, VMSS, or AKS):
137
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
138
+
139
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
140
+# Assign both roles to the VM's managed identity
141
az role assignment create --assignee <managed-identity-principal-id> \
142
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
143
+
144
+az role assignment create --assignee <managed-identity-principal-id> \
145
+ --role "Reader" --scope /subscriptions/<subscription-id>
146
```
147
148
@@ -105,13 +151,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
151
152
#### Options
153
108
-The following options can be defined globally: update_every, autodetection_retry.
154
+The following options can be defined globally: `update_every`, `autodetection_retry`.
155
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
156
+**Profile file locations:**
157
114
-User profile files with the same filename override stock profiles.
158
+| Type | Path |
159
+|:-----|:-----|
160
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
161
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
162
+
163
+User profile files with the same `id` as a stock profile override it.
164
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
165
166
167
<details open><summary>Config options</summary>
@@ -122,25 +172,103 @@ User profile files with the same filename override stock profiles.
172
|:------|:-----|:------------|:--------|:---------:|
173
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
174
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
175
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
176
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
177
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
178
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
179
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
180
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
181
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
182
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
183
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
184
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
185
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
186
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
187
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
188
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
189
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
190
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
191
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
192
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
193
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
194
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
195
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
196
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
197
198
+<a id="option-collection-query-offset"></a>
199
+##### query_offset
200
+
201
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
202
+
203
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
204
+
205
+- **Default (180s)** works for most services.
206
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
207
+- **Increase to 240-300s** if you still see gaps or missing data points.
208
+- **Do not set below 60s** -- metrics will likely be incomplete.
209
+
210
+
211
+<a id="option-authentication-auth-mode"></a>
212
+##### auth.mode
213
+
214
+Determines how the collector authenticates with Azure.
215
+
216
+| Mode | When to use | Required options |
217
+|:-----|:------------|:-----------------|
218
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
219
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
220
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
221
+
222
+
223
+<a id="option-discovery-discovery-mode"></a>
224
+##### discovery.mode
225
+
226
+Controls how the collector finds candidate Azure resources.
227
+
228
+| Mode | Behavior |
229
+|:-----|:---------|
230
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
231
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
232
+
233
+
234
+<a id="option-discovery-discovery-mode-query-kql"></a>
235
+##### discovery.mode_query.kql
236
+
237
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
238
+
239
+The query **must** project these five columns:
240
+
241
+| Column | Description |
242
+|:-------|:------------|
243
+| `id` | Full Azure resource ID (ARM format) |
244
+| `name` | Resource name |
245
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
246
+| `resourceGroup` | Resource group name |
247
+| `location` | Azure region |
248
+
249
+Example:
250
+
251
+```
252
+resources
253
+| where tags.env =~ "prod"
254
+| project id, name, type, resourceGroup, location
255
+```
256
+
257
+
258
+<a id="option-profiles-profiles-mode"></a>
259
+##### profiles.mode
260
+
261
+Controls how the collector decides which metric profiles to activate.
262
+
263
+| Mode | Behavior |
264
+|:-----|:---------|
265
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
266
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
267
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
268
+
269
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
270
+
271
+
272
273
</details>
274
@@ -182,14 +310,28 @@ sudo ./edit-config go.d/azure_monitor.conf
310
311
##### Examples
312
185
-###### Service principal (auto-discover all resources)
313
+###### Service principal with structured discovery
314
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
315
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
316
317
```yaml
318
jobs:
319
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
320
+ subscription_ids:
321
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
322
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
323
+ discovery:
324
+ mode: filters
325
+ mode_filters:
326
+ resource_groups:
327
+ - production-rg
328
+ regions:
329
+ - eastus
330
+ tags:
331
+ env:
332
+ - prod
333
+ profiles:
334
+ mode: auto
335
auth:
336
mode: service_principal
337
mode_service_principal:
@@ -198,59 +340,49 @@ jobs:
340
client_secret: "your-client-secret"
341
342
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
343
+###### Managed identity with exact profiles
344
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
345
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
346
347
<details open><summary>Config</summary>
348
349
```yaml
350
jobs:
351
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
352
+ subscription_ids:
353
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
355
+ mode: exact
356
+ mode_exact:
357
+ names:
358
+ - sql_database
359
+ - postgres_flexible
360
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
361
+ mode: managed_identity
362
363
```
364
</details>
365
241
-###### Filter by resource group
366
+###### Custom Azure Resource Graph KQL
367
243
-Only monitor resources in specific resource groups.
368
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
369
370
<details open><summary>Config</summary>
371
372
```yaml
373
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
374
+ - name: prod-query
375
+ subscription_ids:
376
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
377
+ discovery:
378
+ mode: query
379
+ mode_query:
380
+ kql: |
381
+ resources
382
+ | where tags.env =~ "prod"
383
+ | project id, name, type, resourceGroup, location
384
+ profiles:
385
+ mode: auto
386
auth:
387
mode: default
388
@@ -259,14 +391,15 @@ jobs:
391
392
###### Azure Government cloud
393
262
-Connect to Azure Government cloud environment.
394
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
395
396
<details open><summary>Config</summary>
397
398
```yaml
399
jobs:
400
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
401
+ subscription_ids:
402
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
cloud: government
404
auth:
405
mode: service_principal
@@ -313,6 +446,7 @@ Labels:
446
| region | The Azure region where the resource is deployed. |
447
| resource_type | The Azure resource type identifier. |
448
| profile | The Azure Monitor profile id. |
449
+| subscription_id | The Azure subscription identifier. |
450
| resource_uid | The unique Azure resource identifier. |
451
452
Metrics:
@@ -396,31 +530,46 @@ docker logs netdata 2>&1 | grep azure_monitor
530
531
### No metrics are collected
532
399
-Verify the following:
400
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
401
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
402
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
403
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
533
+Check the following:
534
+
535
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
536
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
537
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
538
+- **Collector logs** -- Check for authentication or API errors:
539
+ ```bash
540
+ # systemd
541
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
542
+ # non-systemd
543
+ grep azure_monitor /var/log/netdata/collector.log
544
+ ```
545
546
547
### Missing metrics for some resource types
548
408
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
409
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
410
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
411
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
549
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
550
+
551
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
552
+- **Verify a built-in profile exists** -- List available profiles:
553
+ ```bash
554
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
555
+ ```
556
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
557
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
558
559
414
-### Metrics appear delayed
560
+### Charts have gaps or incomplete data
561
416
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
417
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
418
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
562
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
563
+
564
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
565
+- Slower time-grain batches automatically use a larger effective offset when needed.
566
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
567
568
569
### Authentication errors in sovereign clouds
570
571
For Azure Government or Azure China clouds, set the `cloud` parameter:
572
+
573
- Azure Government: `cloud: government`
574
- Azure China (21Vianet): `cloud: china`
575
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_storage_account.md
+238
-87
@@ -21,40 +21,77 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Storage Account operations including transaction counts, availability percentages, success and end-to-end latency, ingress and egress throughput, and used capacity.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Storage Account with metrics covering:
31
+
32
+- **Transactions** -- transaction count
33
+- **Latency** -- end-to-end latency and server latency (average/max)
34
+- **Throughput** -- ingress and egress bytes per second
35
+- **Availability** -- service availability percentage
36
+- **Capacity** -- used storage capacity
37
+
38
+
39
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
40
41
42
This collector is supported on all platforms.
43
44
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
45
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
46
+The service principal or managed identity requires these Azure RBAC roles:
47
+
48
+| Role | Purpose | Scope |
49
+|:-----|:--------|:------|
50
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
51
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
52
53
54
### Default Behavior
55
56
#### Auto-Detection
57
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
58
+The collector has two discovery phases:
59
+
60
+**Bootstrap (first run)**
61
+
62
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
63
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
64
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
65
+- A single job can monitor multiple subscriptions.
66
+
67
+**Runtime (periodic refresh)**
68
+
69
+- Periodically re-discovers resources for **already-active profile types only**.
70
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
71
+
72
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
73
74
75
#### Limits
76
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
77
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
78
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
79
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
80
81
82
#### Performance Impact
83
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
84
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
85
+
86
+**Default concurrency and batching limits:**
87
+
88
+| Setting | Default | Description |
89
+|:--------|:--------|:------------|
90
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
91
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
92
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
93
+
94
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
95
96
97
## Setup
@@ -78,25 +115,36 @@ UI configuration requires paid Netdata Cloud plan.
115
116
#### Create an Azure monitoring principal
117
81
-Create a service principal or use a managed identity with the following permissions:
118
+The collector requires a service principal or managed identity with two Azure RBAC roles:
119
+
120
+| Role | Purpose |
121
+|:-----|:--------|
122
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
123
+| **Reader** | Query Azure Resource Graph for resource discovery |
124
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
125
+**Option A: Service principal**
126
86
-For service principal authentication:
127
```bash
88
-# Create the service principal
128
+# Create service principal with Monitoring Reader role
129
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
130
--scopes /subscriptions/<subscription-id>
131
132
+# Add the Reader role for resource discovery
133
+az role assignment create --assignee <appId-from-above> \
134
+ --role "Reader" --scope /subscriptions/<subscription-id>
135
+
136
# Note the appId (client_id), password (client_secret), and tenant
137
```
138
95
-For managed identity (on Azure VMs, VMSS, or AKS):
139
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
140
+
141
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
142
+# Assign both roles to the VM's managed identity
143
az role assignment create --assignee <managed-identity-principal-id> \
144
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
145
+
146
+az role assignment create --assignee <managed-identity-principal-id> \
147
+ --role "Reader" --scope /subscriptions/<subscription-id>
148
```
149
150
@@ -105,13 +153,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
153
154
#### Options
155
108
-The following options can be defined globally: update_every, autodetection_retry.
156
+The following options can be defined globally: `update_every`, `autodetection_retry`.
157
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
158
+**Profile file locations:**
159
114
-User profile files with the same filename override stock profiles.
160
+| Type | Path |
161
+|:-----|:-----|
162
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
163
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
164
+
165
+User profile files with the same `id` as a stock profile override it.
166
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
167
168
169
<details open><summary>Config options</summary>
@@ -122,25 +174,103 @@ User profile files with the same filename override stock profiles.
174
|:------|:-----|:------------|:--------|:---------:|
175
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
176
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
177
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
178
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
181
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
182
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
187
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
188
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
191
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
192
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
193
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
194
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
195
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
196
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
197
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
198
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
199
200
+<a id="option-collection-query-offset"></a>
201
+##### query_offset
202
+
203
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
204
+
205
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
206
+
207
+- **Default (180s)** works for most services.
208
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
209
+- **Increase to 240-300s** if you still see gaps or missing data points.
210
+- **Do not set below 60s** -- metrics will likely be incomplete.
211
+
212
+
213
+<a id="option-authentication-auth-mode"></a>
214
+##### auth.mode
215
+
216
+Determines how the collector authenticates with Azure.
217
+
218
+| Mode | When to use | Required options |
219
+|:-----|:------------|:-----------------|
220
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
221
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
222
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
223
+
224
+
225
+<a id="option-discovery-discovery-mode"></a>
226
+##### discovery.mode
227
+
228
+Controls how the collector finds candidate Azure resources.
229
+
230
+| Mode | Behavior |
231
+|:-----|:---------|
232
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
233
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
234
+
235
+
236
+<a id="option-discovery-discovery-mode-query-kql"></a>
237
+##### discovery.mode_query.kql
238
+
239
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
240
+
241
+The query **must** project these five columns:
242
+
243
+| Column | Description |
244
+|:-------|:------------|
245
+| `id` | Full Azure resource ID (ARM format) |
246
+| `name` | Resource name |
247
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
248
+| `resourceGroup` | Resource group name |
249
+| `location` | Azure region |
250
+
251
+Example:
252
+
253
+```
254
+resources
255
+| where tags.env =~ "prod"
256
+| project id, name, type, resourceGroup, location
257
+```
258
+
259
+
260
+<a id="option-profiles-profiles-mode"></a>
261
+##### profiles.mode
262
+
263
+Controls how the collector decides which metric profiles to activate.
264
+
265
+| Mode | Behavior |
266
+|:-----|:---------|
267
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
268
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
269
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
270
+
271
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
272
+
273
+
274
275
</details>
276
@@ -182,14 +312,28 @@ sudo ./edit-config go.d/azure_monitor.conf
312
313
##### Examples
314
185
-###### Service principal (auto-discover all resources)
315
+###### Service principal with structured discovery
316
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
317
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
318
319
```yaml
320
jobs:
321
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ subscription_ids:
323
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
324
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
325
+ discovery:
326
+ mode: filters
327
+ mode_filters:
328
+ resource_groups:
329
+ - production-rg
330
+ regions:
331
+ - eastus
332
+ tags:
333
+ env:
334
+ - prod
335
+ profiles:
336
+ mode: auto
337
auth:
338
mode: service_principal
339
mode_service_principal:
@@ -198,59 +342,49 @@ jobs:
342
client_secret: "your-client-secret"
343
344
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
345
+###### Managed identity with exact profiles
346
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
347
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
348
349
<details open><summary>Config</summary>
350
351
```yaml
352
jobs:
353
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
+ subscription_ids:
355
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
357
+ mode: exact
358
+ mode_exact:
359
+ names:
360
+ - sql_database
361
+ - postgres_flexible
362
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
363
+ mode: managed_identity
364
365
```
366
</details>
367
241
-###### Filter by resource group
368
+###### Custom Azure Resource Graph KQL
369
243
-Only monitor resources in specific resource groups.
370
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
371
372
<details open><summary>Config</summary>
373
374
```yaml
375
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
376
+ - name: prod-query
377
+ subscription_ids:
378
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
379
+ discovery:
380
+ mode: query
381
+ mode_query:
382
+ kql: |
383
+ resources
384
+ | where tags.env =~ "prod"
385
+ | project id, name, type, resourceGroup, location
386
+ profiles:
387
+ mode: auto
388
auth:
389
mode: default
390
@@ -259,14 +393,15 @@ jobs:
393
394
###### Azure Government cloud
395
262
-Connect to Azure Government cloud environment.
396
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
397
398
<details open><summary>Config</summary>
399
400
```yaml
401
jobs:
402
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
+ subscription_ids:
404
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
cloud: government
406
auth:
407
mode: service_principal
@@ -315,6 +450,7 @@ Labels:
450
| region | The Azure region where the resource is deployed. |
451
| resource_type | The Azure resource type identifier. |
452
| profile | The Azure Monitor profile id. |
453
+| subscription_id | The Azure subscription identifier. |
454
| resource_uid | The unique Azure resource identifier. |
455
456
Metrics:
@@ -399,31 +535,46 @@ docker logs netdata 2>&1 | grep azure_monitor
535
536
### No metrics are collected
537
402
-Verify the following:
403
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
404
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
405
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
406
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
538
+Check the following:
539
+
540
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
541
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
542
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
543
+- **Collector logs** -- Check for authentication or API errors:
544
+ ```bash
545
+ # systemd
546
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
547
+ # non-systemd
548
+ grep azure_monitor /var/log/netdata/collector.log
549
+ ```
550
551
552
### Missing metrics for some resource types
553
411
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
412
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
413
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
414
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
554
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
555
+
556
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
557
+- **Verify a built-in profile exists** -- List available profiles:
558
+ ```bash
559
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
560
+ ```
561
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
562
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
563
564
417
-### Metrics appear delayed
565
+### Charts have gaps or incomplete data
566
419
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
420
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
421
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
567
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
568
+
569
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
570
+- Slower time-grain batches automatically use a larger effective offset when needed.
571
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
572
573
574
### Authentication errors in sovereign clouds
575
576
For Azure Government or Azure China clouds, set the `cloud` parameter:
577
+
578
- Azure Government: `cloud: government`
579
- Azure China (21Vianet): `cloud: china`
580
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_stream_analytics_job.md
+239
-87
@@ -21,40 +21,78 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Stream Analytics jobs including input and output event counts, streaming unit utilization, watermark delay, backlogged input events, runtime and data conversion errors, out-of-order events, and late input events.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Stream Analytics with metrics covering:
31
+
32
+- **Events** -- event flow (in/out), backlogged input events
33
+- **Errors** -- runtime errors, data conversion errors, deserialization errors
34
+- **Timing** -- late, early, and out-of-order events, watermark delay
35
+- **Resources** -- CPU and streaming unit memory utilization
36
+- **Input** -- input data throughput, input sources received
37
+- **Functions** -- function events and requests (total/failed)
38
+
39
+
40
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
41
42
43
This collector is supported on all platforms.
44
45
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
46
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
47
+The service principal or managed identity requires these Azure RBAC roles:
48
+
49
+| Role | Purpose | Scope |
50
+|:-----|:--------|:------|
51
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
52
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
53
54
55
### Default Behavior
56
57
#### Auto-Detection
58
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
59
+The collector has two discovery phases:
60
+
61
+**Bootstrap (first run)**
62
+
63
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
64
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
65
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
66
+- A single job can monitor multiple subscriptions.
67
+
68
+**Runtime (periodic refresh)**
69
+
70
+- Periodically re-discovers resources for **already-active profile types only**.
71
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
72
+
73
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
74
75
76
#### Limits
77
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
78
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
79
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
80
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
81
82
83
#### Performance Impact
84
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
85
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
86
+
87
+**Default concurrency and batching limits:**
88
+
89
+| Setting | Default | Description |
90
+|:--------|:--------|:------------|
91
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
92
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
93
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
94
+
95
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
96
97
98
## Setup
@@ -78,25 +116,36 @@ UI configuration requires paid Netdata Cloud plan.
116
117
#### Create an Azure monitoring principal
118
81
-Create a service principal or use a managed identity with the following permissions:
119
+The collector requires a service principal or managed identity with two Azure RBAC roles:
120
+
121
+| Role | Purpose |
122
+|:-----|:--------|
123
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
124
+| **Reader** | Query Azure Resource Graph for resource discovery |
125
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
126
+**Option A: Service principal**
127
86
-For service principal authentication:
128
```bash
88
-# Create the service principal
129
+# Create service principal with Monitoring Reader role
130
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
131
--scopes /subscriptions/<subscription-id>
132
133
+# Add the Reader role for resource discovery
134
+az role assignment create --assignee <appId-from-above> \
135
+ --role "Reader" --scope /subscriptions/<subscription-id>
136
+
137
# Note the appId (client_id), password (client_secret), and tenant
138
```
139
95
-For managed identity (on Azure VMs, VMSS, or AKS):
140
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
141
+
142
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
143
+# Assign both roles to the VM's managed identity
144
az role assignment create --assignee <managed-identity-principal-id> \
145
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
146
+
147
+az role assignment create --assignee <managed-identity-principal-id> \
148
+ --role "Reader" --scope /subscriptions/<subscription-id>
149
```
150
151
@@ -105,13 +154,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
154
155
#### Options
156
108
-The following options can be defined globally: update_every, autodetection_retry.
157
+The following options can be defined globally: `update_every`, `autodetection_retry`.
158
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
159
+**Profile file locations:**
160
114
-User profile files with the same filename override stock profiles.
161
+| Type | Path |
162
+|:-----|:-----|
163
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
164
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
165
+
166
+User profile files with the same `id` as a stock profile override it.
167
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
168
169
170
<details open><summary>Config options</summary>
@@ -122,25 +175,103 @@ User profile files with the same filename override stock profiles.
175
|:------|:-----|:------------|:--------|:---------:|
176
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
177
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
178
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
179
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
181
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
182
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
183
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
185
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
186
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
187
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
188
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
189
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
191
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
192
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
193
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
194
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
195
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
196
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
197
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
198
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
199
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
200
201
+<a id="option-collection-query-offset"></a>
202
+##### query_offset
203
+
204
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
205
+
206
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
207
+
208
+- **Default (180s)** works for most services.
209
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
210
+- **Increase to 240-300s** if you still see gaps or missing data points.
211
+- **Do not set below 60s** -- metrics will likely be incomplete.
212
+
213
+
214
+<a id="option-authentication-auth-mode"></a>
215
+##### auth.mode
216
+
217
+Determines how the collector authenticates with Azure.
218
+
219
+| Mode | When to use | Required options |
220
+|:-----|:------------|:-----------------|
221
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
222
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
223
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
224
+
225
+
226
+<a id="option-discovery-discovery-mode"></a>
227
+##### discovery.mode
228
+
229
+Controls how the collector finds candidate Azure resources.
230
+
231
+| Mode | Behavior |
232
+|:-----|:---------|
233
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
234
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
235
+
236
+
237
+<a id="option-discovery-discovery-mode-query-kql"></a>
238
+##### discovery.mode_query.kql
239
+
240
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
241
+
242
+The query **must** project these five columns:
243
+
244
+| Column | Description |
245
+|:-------|:------------|
246
+| `id` | Full Azure resource ID (ARM format) |
247
+| `name` | Resource name |
248
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
249
+| `resourceGroup` | Resource group name |
250
+| `location` | Azure region |
251
+
252
+Example:
253
+
254
+```
255
+resources
256
+| where tags.env =~ "prod"
257
+| project id, name, type, resourceGroup, location
258
+```
259
+
260
+
261
+<a id="option-profiles-profiles-mode"></a>
262
+##### profiles.mode
263
+
264
+Controls how the collector decides which metric profiles to activate.
265
+
266
+| Mode | Behavior |
267
+|:-----|:---------|
268
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
269
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
270
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
271
+
272
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
273
+
274
+
275
276
</details>
277
@@ -182,14 +313,28 @@ sudo ./edit-config go.d/azure_monitor.conf
313
314
##### Examples
315
185
-###### Service principal (auto-discover all resources)
316
+###### Service principal with structured discovery
317
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
318
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
319
320
```yaml
321
jobs:
322
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
323
+ subscription_ids:
324
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
325
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
326
+ discovery:
327
+ mode: filters
328
+ mode_filters:
329
+ resource_groups:
330
+ - production-rg
331
+ regions:
332
+ - eastus
333
+ tags:
334
+ env:
335
+ - prod
336
+ profiles:
337
+ mode: auto
338
auth:
339
mode: service_principal
340
mode_service_principal:
@@ -198,59 +343,49 @@ jobs:
343
client_secret: "your-client-secret"
344
345
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
346
+###### Managed identity with exact profiles
347
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
348
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
349
350
<details open><summary>Config</summary>
351
352
```yaml
353
jobs:
354
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
355
+ subscription_ids:
356
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
358
+ mode: exact
359
+ mode_exact:
360
+ names:
361
+ - sql_database
362
+ - postgres_flexible
363
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
364
+ mode: managed_identity
365
366
```
367
</details>
368
241
-###### Filter by resource group
369
+###### Custom Azure Resource Graph KQL
370
243
-Only monitor resources in specific resource groups.
371
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
372
373
<details open><summary>Config</summary>
374
375
```yaml
376
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
377
+ - name: prod-query
378
+ subscription_ids:
379
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
380
+ discovery:
381
+ mode: query
382
+ mode_query:
383
+ kql: |
384
+ resources
385
+ | where tags.env =~ "prod"
386
+ | project id, name, type, resourceGroup, location
387
+ profiles:
388
+ mode: auto
389
auth:
390
mode: default
391
@@ -259,14 +394,15 @@ jobs:
394
395
###### Azure Government cloud
396
262
-Connect to Azure Government cloud environment.
397
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
398
399
<details open><summary>Config</summary>
400
401
```yaml
402
jobs:
403
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
404
+ subscription_ids:
405
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
cloud: government
407
auth:
408
mode: service_principal
@@ -320,6 +456,7 @@ Labels:
456
| region | The Azure region where the resource is deployed. |
457
| resource_type | The Azure resource type identifier. |
458
| profile | The Azure Monitor profile id. |
459
+| subscription_id | The Azure subscription identifier. |
460
| resource_uid | The unique Azure resource identifier. |
461
462
Metrics:
@@ -408,31 +545,46 @@ docker logs netdata 2>&1 | grep azure_monitor
545
546
### No metrics are collected
547
411
-Verify the following:
412
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
413
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
414
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
415
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
548
+Check the following:
549
+
550
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
551
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
552
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
553
+- **Collector logs** -- Check for authentication or API errors:
554
+ ```bash
555
+ # systemd
556
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
557
+ # non-systemd
558
+ grep azure_monitor /var/log/netdata/collector.log
559
+ ```
560
561
562
### Missing metrics for some resource types
563
420
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
421
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
422
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
423
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
564
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
565
+
566
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
567
+- **Verify a built-in profile exists** -- List available profiles:
568
+ ```bash
569
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
570
+ ```
571
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
572
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
573
574
426
-### Metrics appear delayed
575
+### Charts have gaps or incomplete data
576
428
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
429
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
430
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
577
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
578
+
579
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
580
+- Slower time-grain batches automatically use a larger effective offset when needed.
581
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
582
583
584
### Authentication errors in sovereign clouds
585
586
For Azure Government or Azure China clouds, set the `cloud` parameter:
587
+
588
- Azure Government: `cloud: government`
589
- Azure China (21Vianet): `cloud: china`
590
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_synapse_analytics_workspace.md
+238
-87
@@ -21,40 +21,77 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Synapse Analytics workspaces including pipeline and activity run metrics, SQL request counts and data processing volumes, data flow activity execution, integration runtime CPU and memory utilization, and link table event processing.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Synapse Analytics with metrics covering:
31
+
32
+- **Pipeline** -- pipeline runs, activity runs, trigger runs
33
+- **SQL pool** -- built-in SQL pool requests, login attempts, data processed
34
+- **Streaming** -- event flow (in/out), event timing (late/early/out-of-order/backlogged), watermark delay, resource utilization, errors
35
+- **Streaming I/O** -- input data throughput, input sources received
36
+- **Link** -- connection events, processed data volume, changed rows, processing latency, table events
37
+
38
+
39
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
40
41
42
This collector is supported on all platforms.
43
44
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
45
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
46
+The service principal or managed identity requires these Azure RBAC roles:
47
+
48
+| Role | Purpose | Scope |
49
+|:-----|:--------|:------|
50
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
51
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
52
53
54
### Default Behavior
55
56
#### Auto-Detection
57
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
58
+The collector has two discovery phases:
59
+
60
+**Bootstrap (first run)**
61
+
62
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
63
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
64
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
65
+- A single job can monitor multiple subscriptions.
66
+
67
+**Runtime (periodic refresh)**
68
+
69
+- Periodically re-discovers resources for **already-active profile types only**.
70
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
71
+
72
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
73
74
75
#### Limits
76
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
77
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
78
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
79
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
80
81
82
#### Performance Impact
83
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
84
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
85
+
86
+**Default concurrency and batching limits:**
87
+
88
+| Setting | Default | Description |
89
+|:--------|:--------|:------------|
90
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
91
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
92
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
93
+
94
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
95
96
97
## Setup
@@ -78,25 +115,36 @@ UI configuration requires paid Netdata Cloud plan.
115
116
#### Create an Azure monitoring principal
117
81
-Create a service principal or use a managed identity with the following permissions:
118
+The collector requires a service principal or managed identity with two Azure RBAC roles:
119
+
120
+| Role | Purpose |
121
+|:-----|:--------|
122
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
123
+| **Reader** | Query Azure Resource Graph for resource discovery |
124
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
125
+**Option A: Service principal**
126
86
-For service principal authentication:
127
```bash
88
-# Create the service principal
128
+# Create service principal with Monitoring Reader role
129
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
130
--scopes /subscriptions/<subscription-id>
131
132
+# Add the Reader role for resource discovery
133
+az role assignment create --assignee <appId-from-above> \
134
+ --role "Reader" --scope /subscriptions/<subscription-id>
135
+
136
# Note the appId (client_id), password (client_secret), and tenant
137
```
138
95
-For managed identity (on Azure VMs, VMSS, or AKS):
139
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
140
+
141
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
142
+# Assign both roles to the VM's managed identity
143
az role assignment create --assignee <managed-identity-principal-id> \
144
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
145
+
146
+az role assignment create --assignee <managed-identity-principal-id> \
147
+ --role "Reader" --scope /subscriptions/<subscription-id>
148
```
149
150
@@ -105,13 +153,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
153
154
#### Options
155
108
-The following options can be defined globally: update_every, autodetection_retry.
156
+The following options can be defined globally: `update_every`, `autodetection_retry`.
157
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
158
+**Profile file locations:**
159
114
-User profile files with the same filename override stock profiles.
160
+| Type | Path |
161
+|:-----|:-----|
162
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
163
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
164
+
165
+User profile files with the same `id` as a stock profile override it.
166
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
167
168
169
<details open><summary>Config options</summary>
@@ -122,25 +174,103 @@ User profile files with the same filename override stock profiles.
174
|:------|:-----|:------------|:--------|:---------:|
175
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
176
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
177
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
178
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
179
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
180
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
181
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
182
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
183
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
184
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
185
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
186
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
187
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
188
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
189
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
190
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
191
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
192
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
193
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
194
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
195
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
196
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
197
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
198
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
199
200
+<a id="option-collection-query-offset"></a>
201
+##### query_offset
202
+
203
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
204
+
205
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
206
+
207
+- **Default (180s)** works for most services.
208
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
209
+- **Increase to 240-300s** if you still see gaps or missing data points.
210
+- **Do not set below 60s** -- metrics will likely be incomplete.
211
+
212
+
213
+<a id="option-authentication-auth-mode"></a>
214
+##### auth.mode
215
+
216
+Determines how the collector authenticates with Azure.
217
+
218
+| Mode | When to use | Required options |
219
+|:-----|:------------|:-----------------|
220
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
221
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
222
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
223
+
224
+
225
+<a id="option-discovery-discovery-mode"></a>
226
+##### discovery.mode
227
+
228
+Controls how the collector finds candidate Azure resources.
229
+
230
+| Mode | Behavior |
231
+|:-----|:---------|
232
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
233
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
234
+
235
+
236
+<a id="option-discovery-discovery-mode-query-kql"></a>
237
+##### discovery.mode_query.kql
238
+
239
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
240
+
241
+The query **must** project these five columns:
242
+
243
+| Column | Description |
244
+|:-------|:------------|
245
+| `id` | Full Azure resource ID (ARM format) |
246
+| `name` | Resource name |
247
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
248
+| `resourceGroup` | Resource group name |
249
+| `location` | Azure region |
250
+
251
+Example:
252
+
253
+```
254
+resources
255
+| where tags.env =~ "prod"
256
+| project id, name, type, resourceGroup, location
257
+```
258
+
259
+
260
+<a id="option-profiles-profiles-mode"></a>
261
+##### profiles.mode
262
+
263
+Controls how the collector decides which metric profiles to activate.
264
+
265
+| Mode | Behavior |
266
+|:-----|:---------|
267
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
268
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
269
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
270
+
271
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
272
+
273
+
274
275
</details>
276
@@ -182,14 +312,28 @@ sudo ./edit-config go.d/azure_monitor.conf
312
313
##### Examples
314
185
-###### Service principal (auto-discover all resources)
315
+###### Service principal with structured discovery
316
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
317
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
318
319
```yaml
320
jobs:
321
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
322
+ subscription_ids:
323
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
324
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
325
+ discovery:
326
+ mode: filters
327
+ mode_filters:
328
+ resource_groups:
329
+ - production-rg
330
+ regions:
331
+ - eastus
332
+ tags:
333
+ env:
334
+ - prod
335
+ profiles:
336
+ mode: auto
337
auth:
338
mode: service_principal
339
mode_service_principal:
@@ -198,59 +342,49 @@ jobs:
342
client_secret: "your-client-secret"
343
344
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
345
+###### Managed identity with exact profiles
346
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
347
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
348
349
<details open><summary>Config</summary>
350
351
```yaml
352
jobs:
353
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
354
+ subscription_ids:
355
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
356
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
357
+ mode: exact
358
+ mode_exact:
359
+ names:
360
+ - sql_database
361
+ - postgres_flexible
362
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
363
+ mode: managed_identity
364
365
```
366
</details>
367
241
-###### Filter by resource group
368
+###### Custom Azure Resource Graph KQL
369
243
-Only monitor resources in specific resource groups.
370
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
371
372
<details open><summary>Config</summary>
373
374
```yaml
375
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
376
+ - name: prod-query
377
+ subscription_ids:
378
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
379
+ discovery:
380
+ mode: query
381
+ mode_query:
382
+ kql: |
383
+ resources
384
+ | where tags.env =~ "prod"
385
+ | project id, name, type, resourceGroup, location
386
+ profiles:
387
+ mode: auto
388
auth:
389
mode: default
390
@@ -259,14 +393,15 @@ jobs:
393
394
###### Azure Government cloud
395
262
-Connect to Azure Government cloud environment.
396
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
397
398
<details open><summary>Config</summary>
399
400
```yaml
401
jobs:
402
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
403
+ subscription_ids:
404
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
405
cloud: government
406
auth:
407
mode: service_principal
@@ -318,6 +453,7 @@ Labels:
453
| region | The Azure region where the resource is deployed. |
454
| resource_type | The Azure resource type identifier. |
455
| profile | The Azure Monitor profile id. |
456
+| subscription_id | The Azure subscription identifier. |
457
| resource_uid | The unique Azure resource identifier. |
458
459
Metrics:
@@ -414,31 +550,46 @@ docker logs netdata 2>&1 | grep azure_monitor
550
551
### No metrics are collected
552
417
-Verify the following:
418
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
419
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
420
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
421
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
553
+Check the following:
554
+
555
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
556
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
557
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
558
+- **Collector logs** -- Check for authentication or API errors:
559
+ ```bash
560
+ # systemd
561
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
562
+ # non-systemd
563
+ grep azure_monitor /var/log/netdata/collector.log
564
+ ```
565
566
567
### Missing metrics for some resource types
568
426
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
427
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
428
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
429
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
569
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
570
+
571
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
572
+- **Verify a built-in profile exists** -- List available profiles:
573
+ ```bash
574
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
575
+ ```
576
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
577
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
578
579
432
-### Metrics appear delayed
580
+### Charts have gaps or incomplete data
581
434
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
435
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
436
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
582
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
583
+
584
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
585
+- Slower time-grain batches automatically use a larger effective offset when needed.
586
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
587
588
589
### Authentication errors in sovereign clouds
590
591
For Azure Government or Azure China clouds, set the `cloud` parameter:
592
+
593
- Azure Government: `cloud: government`
594
- Azure China (21Vianet): `cloud: china`
595
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine.md
+241
-87
@@ -21,40 +21,80 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Azure Virtual Machines including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure Virtual Machines with metrics covering:
31
+
32
+- **Compute** -- CPU utilization, CPU credits (consumed/remaining)
33
+- **Memory** -- available memory (bytes and percentage)
34
+- **Disk** -- IOPS, throughput, latency, and queue depth for OS, data, and temp disks
35
+- **Disk burst** -- burst credits and capacity for OS and data disks, VM-level cached/uncached burst credits
36
+- **Disk cache** -- premium OS and data disk cache hit/miss rates
37
+- **Network** -- traffic in/out, network flows, flow creation rate
38
+- **Availability** -- VM availability state
39
+- **Throttling** -- cached and uncached I/O bandwidth/IOPS throttling
40
+
41
+
42
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
43
44
45
This collector is supported on all platforms.
46
47
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
48
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
49
+The service principal or managed identity requires these Azure RBAC roles:
50
+
51
+| Role | Purpose | Scope |
52
+|:-----|:--------|:------|
53
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
54
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
55
56
57
### Default Behavior
58
59
#### Auto-Detection
60
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
61
+The collector has two discovery phases:
62
+
63
+**Bootstrap (first run)**
64
+
65
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
66
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
67
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
68
+- A single job can monitor multiple subscriptions.
69
+
70
+**Runtime (periodic refresh)**
71
+
72
+- Periodically re-discovers resources for **already-active profile types only**.
73
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
74
+
75
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
76
77
78
#### Limits
79
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
80
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
81
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
82
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
83
84
85
#### Performance Impact
86
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
87
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
88
+
89
+**Default concurrency and batching limits:**
90
+
91
+| Setting | Default | Description |
92
+|:--------|:--------|:------------|
93
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
94
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
95
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
96
+
97
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
98
99
100
## Setup
@@ -78,25 +118,36 @@ UI configuration requires paid Netdata Cloud plan.
118
119
#### Create an Azure monitoring principal
120
81
-Create a service principal or use a managed identity with the following permissions:
121
+The collector requires a service principal or managed identity with two Azure RBAC roles:
122
+
123
+| Role | Purpose |
124
+|:-----|:--------|
125
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
126
+| **Reader** | Query Azure Resource Graph for resource discovery |
127
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
128
+**Option A: Service principal**
129
86
-For service principal authentication:
130
```bash
88
-# Create the service principal
131
+# Create service principal with Monitoring Reader role
132
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
133
--scopes /subscriptions/<subscription-id>
134
135
+# Add the Reader role for resource discovery
136
+az role assignment create --assignee <appId-from-above> \
137
+ --role "Reader" --scope /subscriptions/<subscription-id>
138
+
139
# Note the appId (client_id), password (client_secret), and tenant
140
```
141
95
-For managed identity (on Azure VMs, VMSS, or AKS):
142
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
143
+
144
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
145
+# Assign both roles to the VM's managed identity
146
az role assignment create --assignee <managed-identity-principal-id> \
147
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
148
+
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Reader" --scope /subscriptions/<subscription-id>
151
```
152
153
@@ -105,13 +156,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
156
157
#### Options
158
108
-The following options can be defined globally: update_every, autodetection_retry.
159
+The following options can be defined globally: `update_every`, `autodetection_retry`.
160
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
161
+**Profile file locations:**
162
114
-User profile files with the same filename override stock profiles.
163
+| Type | Path |
164
+|:-----|:-----|
165
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
166
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
167
+
168
+User profile files with the same `id` as a stock profile override it.
169
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
170
171
172
<details open><summary>Config options</summary>
@@ -122,25 +177,103 @@ User profile files with the same filename override stock profiles.
177
|:------|:-----|:------------|:--------|:---------:|
178
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
179
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
181
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
184
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
185
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
190
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
191
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
194
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
195
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
196
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
197
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
198
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
199
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
200
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
201
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
203
+<a id="option-collection-query-offset"></a>
204
+##### query_offset
205
+
206
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
207
+
208
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
209
+
210
+- **Default (180s)** works for most services.
211
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
212
+- **Increase to 240-300s** if you still see gaps or missing data points.
213
+- **Do not set below 60s** -- metrics will likely be incomplete.
214
+
215
+
216
+<a id="option-authentication-auth-mode"></a>
217
+##### auth.mode
218
+
219
+Determines how the collector authenticates with Azure.
220
+
221
+| Mode | When to use | Required options |
222
+|:-----|:------------|:-----------------|
223
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
224
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
225
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
226
+
227
+
228
+<a id="option-discovery-discovery-mode"></a>
229
+##### discovery.mode
230
+
231
+Controls how the collector finds candidate Azure resources.
232
+
233
+| Mode | Behavior |
234
+|:-----|:---------|
235
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
236
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
237
+
238
+
239
+<a id="option-discovery-discovery-mode-query-kql"></a>
240
+##### discovery.mode_query.kql
241
+
242
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
243
+
244
+The query **must** project these five columns:
245
+
246
+| Column | Description |
247
+|:-------|:------------|
248
+| `id` | Full Azure resource ID (ARM format) |
249
+| `name` | Resource name |
250
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
251
+| `resourceGroup` | Resource group name |
252
+| `location` | Azure region |
253
+
254
+Example:
255
+
256
+```
257
+resources
258
+| where tags.env =~ "prod"
259
+| project id, name, type, resourceGroup, location
260
+```
261
+
262
+
263
+<a id="option-profiles-profiles-mode"></a>
264
+##### profiles.mode
265
+
266
+Controls how the collector decides which metric profiles to activate.
267
+
268
+| Mode | Behavior |
269
+|:-----|:---------|
270
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
271
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
272
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
273
+
274
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
275
+
276
+
277
278
</details>
279
@@ -182,14 +315,28 @@ sudo ./edit-config go.d/azure_monitor.conf
315
316
##### Examples
317
185
-###### Service principal (auto-discover all resources)
318
+###### Service principal with structured discovery
319
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
320
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
321
322
```yaml
323
jobs:
324
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ subscription_ids:
326
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
327
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
328
+ discovery:
329
+ mode: filters
330
+ mode_filters:
331
+ resource_groups:
332
+ - production-rg
333
+ regions:
334
+ - eastus
335
+ tags:
336
+ env:
337
+ - prod
338
+ profiles:
339
+ mode: auto
340
auth:
341
mode: service_principal
342
mode_service_principal:
@@ -198,59 +345,49 @@ jobs:
345
client_secret: "your-client-secret"
346
347
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
348
+###### Managed identity with exact profiles
349
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
350
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
351
352
<details open><summary>Config</summary>
353
354
```yaml
355
jobs:
356
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ subscription_ids:
358
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
360
+ mode: exact
361
+ mode_exact:
362
+ names:
363
+ - sql_database
364
+ - postgres_flexible
365
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
366
+ mode: managed_identity
367
368
```
369
</details>
370
241
-###### Filter by resource group
371
+###### Custom Azure Resource Graph KQL
372
243
-Only monitor resources in specific resource groups.
373
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
374
375
<details open><summary>Config</summary>
376
377
```yaml
378
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
379
+ - name: prod-query
380
+ subscription_ids:
381
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
382
+ discovery:
383
+ mode: query
384
+ mode_query:
385
+ kql: |
386
+ resources
387
+ | where tags.env =~ "prod"
388
+ | project id, name, type, resourceGroup, location
389
+ profiles:
390
+ mode: auto
391
auth:
392
mode: default
393
@@ -259,14 +396,15 @@ jobs:
396
397
###### Azure Government cloud
398
262
-Connect to Azure Government cloud environment.
399
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
400
401
<details open><summary>Config</summary>
402
403
```yaml
404
jobs:
405
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
+ subscription_ids:
407
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
cloud: government
409
auth:
410
mode: service_principal
@@ -334,6 +472,7 @@ Labels:
472
| region | The Azure region where the resource is deployed. |
473
| resource_type | The Azure resource type identifier. |
474
| profile | The Azure Monitor profile id. |
475
+| subscription_id | The Azure subscription identifier. |
476
| resource_uid | The unique Azure resource identifier. |
477
478
Metrics:
@@ -448,31 +587,46 @@ docker logs netdata 2>&1 | grep azure_monitor
587
588
### No metrics are collected
589
451
-Verify the following:
452
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
453
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
454
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
455
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
590
+Check the following:
591
+
592
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
593
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
594
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
595
+- **Collector logs** -- Check for authentication or API errors:
596
+ ```bash
597
+ # systemd
598
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
599
+ # non-systemd
600
+ grep azure_monitor /var/log/netdata/collector.log
601
+ ```
602
603
604
### Missing metrics for some resource types
605
460
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
461
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
462
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
463
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
606
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
607
+
608
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
609
+- **Verify a built-in profile exists** -- List available profiles:
610
+ ```bash
611
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
612
+ ```
613
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
614
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
615
616
466
-### Metrics appear delayed
617
+### Charts have gaps or incomplete data
618
468
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
469
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
470
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
619
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
620
+
621
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
622
+- Slower time-grain batches automatically use a larger effective offset when needed.
623
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
624
625
626
### Authentication errors in sovereign clouds
627
628
For Azure Government or Azure China clouds, set the `cloud` parameter:
629
+
630
- Azure Government: `cloud: government`
631
- Azure China (21Vianet): `cloud: china`
632
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_virtual_machine_scale_set.md
+243
-87
@@ -21,40 +21,82 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor Virtual Machine Scale Sets including CPU utilization, available memory percentage, disk IOPS and throughput for OS, data, temp, and premium cache disks, disk burst and VM-level burst credit balances, network traffic, and inbound/outbound flow creation rates across all instances in the scale set.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
+
28
+:::
29
+
30
+Monitor Azure Virtual Machine Scale Sets with metrics covering:
31
+
32
+- **Compute** -- CPU utilization, CPU credits (consumed/remaining)
33
+- **Memory** -- available memory (bytes and percentage)
34
+- **Disk** -- IOPS, throughput, latency, and queue depth for OS, data, and temp disks
35
+- **Disk burst** -- burst credits and capacity for OS and data disks, VM-level cached/uncached burst credits
36
+- **Disk cache** -- premium OS and data disk cache hit/miss rates
37
+- **Network** -- traffic in/out, network flows, flow creation rate
38
+- **Availability** -- VMSS availability state
39
+- **Throttling** -- cached and uncached I/O bandwidth/IOPS throttling
40
+
41
+Metrics are aggregated across all instances in the scale set.
42
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
43
+
44
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
45
46
47
This collector is supported on all platforms.
48
49
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
50
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
51
+The service principal or managed identity requires these Azure RBAC roles:
52
+
53
+| Role | Purpose | Scope |
54
+|:-----|:--------|:------|
55
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
56
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
57
58
59
### Default Behavior
60
61
#### Auto-Detection
62
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
63
+The collector has two discovery phases:
64
+
65
+**Bootstrap (first run)**
66
+
67
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
68
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
69
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
70
+- A single job can monitor multiple subscriptions.
71
+
72
+**Runtime (periodic refresh)**
73
+
74
+- Periodically re-discovers resources for **already-active profile types only**.
75
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
76
+
77
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
78
79
80
#### Limits
81
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
82
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
83
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
84
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
85
86
87
#### Performance Impact
88
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
89
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
90
+
91
+**Default concurrency and batching limits:**
92
+
93
+| Setting | Default | Description |
94
+|:--------|:--------|:------------|
95
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
96
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
97
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
98
+
99
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
100
101
102
## Setup
@@ -78,25 +120,36 @@ UI configuration requires paid Netdata Cloud plan.
120
121
#### Create an Azure monitoring principal
122
81
-Create a service principal or use a managed identity with the following permissions:
123
+The collector requires a service principal or managed identity with two Azure RBAC roles:
124
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
125
+| Role | Purpose |
126
+|:-----|:--------|
127
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
128
+| **Reader** | Query Azure Resource Graph for resource discovery |
129
+
130
+**Option A: Service principal**
131
86
-For service principal authentication:
132
```bash
88
-# Create the service principal
133
+# Create service principal with Monitoring Reader role
134
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
135
--scopes /subscriptions/<subscription-id>
136
137
+# Add the Reader role for resource discovery
138
+az role assignment create --assignee <appId-from-above> \
139
+ --role "Reader" --scope /subscriptions/<subscription-id>
140
+
141
# Note the appId (client_id), password (client_secret), and tenant
142
```
143
95
-For managed identity (on Azure VMs, VMSS, or AKS):
144
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
145
+
146
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
147
+# Assign both roles to the VM's managed identity
148
az role assignment create --assignee <managed-identity-principal-id> \
149
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
150
+
151
+az role assignment create --assignee <managed-identity-principal-id> \
152
+ --role "Reader" --scope /subscriptions/<subscription-id>
153
```
154
155
@@ -105,13 +158,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
158
159
#### Options
160
108
-The following options can be defined globally: update_every, autodetection_retry.
161
+The following options can be defined globally: `update_every`, `autodetection_retry`.
162
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
163
+**Profile file locations:**
164
114
-User profile files with the same filename override stock profiles.
165
+| Type | Path |
166
+|:-----|:-----|
167
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
168
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
169
+
170
+User profile files with the same `id` as a stock profile override it.
171
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
172
173
174
<details open><summary>Config options</summary>
@@ -122,25 +179,103 @@ User profile files with the same filename override stock profiles.
179
|:------|:-----|:------------|:--------|:---------:|
180
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
181
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
182
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
183
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
184
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
185
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
186
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
187
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
188
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
189
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
190
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
191
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
192
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
193
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
194
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
195
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
196
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
197
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
198
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
199
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
200
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
201
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
202
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
203
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
204
205
+<a id="option-collection-query-offset"></a>
206
+##### query_offset
207
+
208
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
209
+
210
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
211
+
212
+- **Default (180s)** works for most services.
213
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
214
+- **Increase to 240-300s** if you still see gaps or missing data points.
215
+- **Do not set below 60s** -- metrics will likely be incomplete.
216
+
217
+
218
+<a id="option-authentication-auth-mode"></a>
219
+##### auth.mode
220
+
221
+Determines how the collector authenticates with Azure.
222
+
223
+| Mode | When to use | Required options |
224
+|:-----|:------------|:-----------------|
225
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
226
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
227
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
228
+
229
+
230
+<a id="option-discovery-discovery-mode"></a>
231
+##### discovery.mode
232
+
233
+Controls how the collector finds candidate Azure resources.
234
+
235
+| Mode | Behavior |
236
+|:-----|:---------|
237
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
238
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
239
+
240
+
241
+<a id="option-discovery-discovery-mode-query-kql"></a>
242
+##### discovery.mode_query.kql
243
+
244
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
245
+
246
+The query **must** project these five columns:
247
+
248
+| Column | Description |
249
+|:-------|:------------|
250
+| `id` | Full Azure resource ID (ARM format) |
251
+| `name` | Resource name |
252
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
253
+| `resourceGroup` | Resource group name |
254
+| `location` | Azure region |
255
+
256
+Example:
257
+
258
+```
259
+resources
260
+| where tags.env =~ "prod"
261
+| project id, name, type, resourceGroup, location
262
+```
263
+
264
+
265
+<a id="option-profiles-profiles-mode"></a>
266
+##### profiles.mode
267
+
268
+Controls how the collector decides which metric profiles to activate.
269
+
270
+| Mode | Behavior |
271
+|:-----|:---------|
272
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
273
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
274
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
275
+
276
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
277
+
278
+
279
280
</details>
281
@@ -182,14 +317,28 @@ sudo ./edit-config go.d/azure_monitor.conf
317
318
##### Examples
319
185
-###### Service principal (auto-discover all resources)
320
+###### Service principal with structured discovery
321
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
322
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
323
324
```yaml
325
jobs:
326
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
327
+ subscription_ids:
328
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
329
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
330
+ discovery:
331
+ mode: filters
332
+ mode_filters:
333
+ resource_groups:
334
+ - production-rg
335
+ regions:
336
+ - eastus
337
+ tags:
338
+ env:
339
+ - prod
340
+ profiles:
341
+ mode: auto
342
auth:
343
mode: service_principal
344
mode_service_principal:
@@ -198,59 +347,49 @@ jobs:
347
client_secret: "your-client-secret"
348
349
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
350
+###### Managed identity with exact profiles
351
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
352
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
353
354
<details open><summary>Config</summary>
355
356
```yaml
357
jobs:
358
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
+ subscription_ids:
360
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
361
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
362
+ mode: exact
363
+ mode_exact:
364
+ names:
365
+ - sql_database
366
+ - postgres_flexible
367
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
368
+ mode: managed_identity
369
370
```
371
</details>
372
241
-###### Filter by resource group
373
+###### Custom Azure Resource Graph KQL
374
243
-Only monitor resources in specific resource groups.
375
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
376
377
<details open><summary>Config</summary>
378
379
```yaml
380
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
381
+ - name: prod-query
382
+ subscription_ids:
383
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
384
+ discovery:
385
+ mode: query
386
+ mode_query:
387
+ kql: |
388
+ resources
389
+ | where tags.env =~ "prod"
390
+ | project id, name, type, resourceGroup, location
391
+ profiles:
392
+ mode: auto
393
auth:
394
mode: default
395
@@ -259,14 +398,15 @@ jobs:
398
399
###### Azure Government cloud
400
262
-Connect to Azure Government cloud environment.
401
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
402
403
<details open><summary>Config</summary>
404
405
```yaml
406
jobs:
407
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
+ subscription_ids:
409
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
410
cloud: government
411
auth:
412
mode: service_principal
@@ -336,6 +476,7 @@ Labels:
476
| region | The Azure region where the resource is deployed. |
477
| resource_type | The Azure resource type identifier. |
478
| profile | The Azure Monitor profile id. |
479
+| subscription_id | The Azure subscription identifier. |
480
| resource_uid | The unique Azure resource identifier. |
481
482
Metrics:
@@ -450,31 +591,46 @@ docker logs netdata 2>&1 | grep azure_monitor
591
592
### No metrics are collected
593
453
-Verify the following:
454
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
455
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
456
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
457
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
594
+Check the following:
595
+
596
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
597
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
598
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
599
+- **Collector logs** -- Check for authentication or API errors:
600
+ ```bash
601
+ # systemd
602
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
603
+ # non-systemd
604
+ grep azure_monitor /var/log/netdata/collector.log
605
+ ```
606
607
608
### Missing metrics for some resource types
609
462
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
463
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
464
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
465
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
610
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
611
+
612
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
613
+- **Verify a built-in profile exists** -- List available profiles:
614
+ ```bash
615
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
616
+ ```
617
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
618
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
619
620
468
-### Metrics appear delayed
621
+### Charts have gaps or incomplete data
622
470
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
471
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
472
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
623
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
624
+
625
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
626
+- Slower time-grain batches automatically use a larger effective offset when needed.
627
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
628
629
630
### Authentication errors in sovereign clouds
631
632
For Azure Government or Azure China clouds, set the `cloud` parameter:
633
+
634
- Azure Government: `cloud: government`
635
- Azure China (21Vianet): `cloud: china`
636
src/go/plugin/go.d/collector/azure_monitor/integrations/azure_vpn_gateway.md
+241
-87
@@ -21,40 +21,80 @@ Module: azure_monitor
21
22
## Overview
23
24
-Monitor VPN Gateway including site-to-site bandwidth and BGP peer status, point-to-site connection counts and bandwidth, per-tunnel ingress and egress traffic with packet counts and drops, IPsec security association counts, route table sizes, NAT flow counts and packet translations, and gateway-level bandwidth utilization.
24
+:::info
25
26
+This is part of the [Azure Monitor](https://github.com/netdata/netdata/blob/master/src/go/plugin/go.d/collector/azure_monitor/integrations/azure_monitor.md) collector. No separate setup is needed -- a single Azure Monitor job discovers and monitors all supported resource types automatically.
27
27
-The collector uses Azure SDK clients for:
28
-- Authentication via Entra ID (service principal, managed identity, or default credentials)
29
-- Resource discovery via Azure Resource Graph queries
30
-- Metrics collection via Azure Monitor Metrics batch API, grouped by region and time grain
28
+:::
29
+
30
+Monitor Azure VPN Gateway with metrics covering:
31
+
32
+- **Site-to-site** -- S2S bandwidth, tunnel bandwidth, tunnel bytes (ingress/egress), tunnel packets and drops
33
+- **Point-to-site** -- P2S connection count, P2S bandwidth
34
+- **BGP** -- BGP peer status, routes advertised and learned
35
+- **ExpressRoute** -- ExpressRoute gateway bandwidth, CPU, packets, active flows, route changes, VMs in VNet
36
+- **IPsec** -- MMSA and QMSA security association counts
37
+- **NAT** -- NAT flows, NAT allocations, NATed bytes and packets, NAT packet drops
38
+- **Routes** -- user VPN and VNet prefix route counts
39
+- **Flows** -- gateway inbound/outbound flows, tunnel total flows, peak PPS, TS mismatch drops
40
+
41
+
42
+It uses the [Azure Monitor Metrics batch API](https://learn.microsoft.com/en-us/azure/azure-monitor/essentials/migrate-to-batch-api) to collect metrics, grouping requests by subscription, region, and time grain. Resources are discovered via [Azure Resource Graph](https://learn.microsoft.com/en-us/azure/governance/resource-graph/overview) queries at startup and refreshed periodically. Authentication is handled through [Microsoft Entra ID](https://learn.microsoft.com/en-us/entra/identity/) (service principal, managed identity, or default credentials).
43
44
45
This collector is supported on all platforms.
46
47
This collector supports collecting metrics from multiple instances of this integration, including remote instances.
48
37
-The monitoring principal needs read access to Azure Resource Graph and Azure Monitor metrics for target resources.
49
+The service principal or managed identity requires these Azure RBAC roles:
50
+
51
+| Role | Purpose | Scope |
52
+|:-----|:--------|:------|
53
+| **Monitoring Reader** | Read Azure Monitor metrics | Subscription or resource group |
54
+| **Reader** | Query Azure Resource Graph for resource discovery | Subscription or resource group |
55
56
57
### Default Behavior
58
59
#### Auto-Detection
60
44
-When `profile_selection_mode` is `auto` (the default), the collector queries Azure Resource Graph
45
-to discover which resource types exist in the subscription and enables matching built-in profiles automatically.
61
+The collector has two discovery phases:
62
+
63
+**Bootstrap (first run)**
64
+
65
+- With the default `profiles.mode: auto`, the collector queries Azure Resource Graph within the configured `subscription_ids` to find candidate resources.
66
+- It matches discovered resource types against built-in profiles and automatically enables the relevant ones.
67
+- Discovery scope can be narrowed using `discovery.mode: filters` (resource groups, regions, tags) or replaced entirely with `discovery.mode: query` for a custom KQL query.
68
+- A single job can monitor multiple subscriptions.
69
+
70
+**Runtime (periodic refresh)**
71
+
72
+- Periodically re-discovers resources for **already-active profile types only**.
73
+- Controlled by `discovery.refresh_every` (default: 300 seconds, set to 0 to disable).
74
+
75
+> **Important:** Runtime refresh does not activate new profiles. If a new resource type appears after bootstrap, restart the collector to pick it up.
76
77
78
#### Limits
79
50
-Azure Monitor metrics granularity is typically 1 minute.
51
-The collector enforces a minimum collection interval of 60 seconds.
80
+- **Minimum collection interval:** 60 seconds (enforced). Azure Monitor metrics granularity is typically 1 minute.
81
+- **Metrics reporting delay:** Azure Monitor metrics have a 1-3 minute reporting delay. The collector uses `query_offset` (default: 180s) as a minimum offset and automatically uses a larger effective offset for slower time-grain batches when needed.
82
+- **API throttling:** Azure Monitor applies per-subscription rate limits. The collector uses bounded concurrency and batching to stay within limits, but monitoring many resources in a single subscription may require tuning `limits.*` options.
83
84
85
#### Performance Impact
86
56
-The collector uses bounded request concurrency and batches resources and metrics to minimize API calls.
57
-Default limits: 4 concurrent queries, 50 resources per batch, 20 metrics per query.
87
+The collector batches resources and metrics to minimize Azure API calls and uses bounded concurrency to avoid overwhelming the API.
88
+
89
+**Default concurrency and batching limits:**
90
+
91
+| Setting | Default | Description |
92
+|:--------|:--------|:------------|
93
+| `limits.max_concurrency` | 4 | Maximum concurrent batch queries |
94
+| `limits.max_batch_resources` | 50 | Maximum resources per batch request |
95
+| `limits.max_metrics_per_query` | 20 | Maximum metrics per batch request |
96
+
97
+For large deployments, consider splitting resources across multiple jobs. If you hit Azure API rate limits, reduce `max_concurrency`.
98
99
100
## Setup
@@ -78,25 +118,36 @@ UI configuration requires paid Netdata Cloud plan.
118
119
#### Create an Azure monitoring principal
120
81
-Create a service principal or use a managed identity with the following permissions:
121
+The collector requires a service principal or managed identity with two Azure RBAC roles:
122
+
123
+| Role | Purpose |
124
+|:-----|:--------|
125
+| **Monitoring Reader** | Access Azure Monitor metrics for target resources |
126
+| **Reader** | Query Azure Resource Graph for resource discovery |
127
83
-1. **Monitoring Reader** role on the target subscription or resource groups (for Azure Monitor metrics access)
84
-2. **Reader** role for Azure Resource Graph queries (for resource discovery)
128
+**Option A: Service principal**
129
86
-For service principal authentication:
130
```bash
88
-# Create the service principal
131
+# Create service principal with Monitoring Reader role
132
az ad sp create-for-rbac --name "netdata-monitor" --role "Monitoring Reader" \
133
--scopes /subscriptions/<subscription-id>
134
135
+# Add the Reader role for resource discovery
136
+az role assignment create --assignee <appId-from-above> \
137
+ --role "Reader" --scope /subscriptions/<subscription-id>
138
+
139
# Note the appId (client_id), password (client_secret), and tenant
140
```
141
95
-For managed identity (on Azure VMs, VMSS, or AKS):
142
+**Option B: Managed identity** (Azure VMs, VMSS, or AKS)
143
+
144
```bash
97
-# Assign Monitoring Reader role to the VM's managed identity
145
+# Assign both roles to the VM's managed identity
146
az role assignment create --assignee <managed-identity-principal-id> \
147
--role "Monitoring Reader" --scope /subscriptions/<subscription-id>
148
+
149
+az role assignment create --assignee <managed-identity-principal-id> \
150
+ --role "Reader" --scope /subscriptions/<subscription-id>
151
```
152
153
@@ -105,13 +156,17 @@ az role assignment create --assignee <managed-identity-principal-id> \
156
157
#### Options
158
108
-The following options can be defined globally: update_every, autodetection_retry.
159
+The following options can be defined globally: `update_every`, `autodetection_retry`.
160
110
-Profile files are loaded from:
111
-- Stock: `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/`
112
-- User: `/etc/netdata/go.d/azure_monitor.profiles/`
161
+**Profile file locations:**
162
114
-User profile files with the same filename override stock profiles.
163
+| Type | Path |
164
+|:-----|:-----|
165
+| Stock profiles | `/usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` |
166
+| User overrides | `/etc/netdata/go.d/azure_monitor.profiles/` |
167
+
168
+User profile files with the same `id` as a stock profile override it.
169
+Custom profiles extend the collector's catalog -- they do not replace the discovery mechanism.
170
171
172
<details open><summary>Config options</summary>
@@ -122,25 +177,103 @@ User profile files with the same filename override stock profiles.
177
|:------|:-----|:------------|:--------|:---------:|
178
| **Collection** | update_every | Data collection interval (seconds). Must be at least 60. | 60 | no |
179
| | autodetection_retry | Autodetection retry interval (seconds). Set 0 to disable. | 0 | no |
125
-| **Target** | subscription_id | Azure subscription ID. | | yes |
180
+| | subscription_ids | List of Azure subscription IDs to monitor. Used as the scope for resource discovery. | | yes |
181
| | cloud | Azure cloud environment: `public`, `government`, or `china`. | public | no |
127
-| **Collection** | discovery_every | Resource discovery interval in seconds. | 300 | no |
128
-| | query_offset | Offset in seconds for metric query windows. Increase if metrics appear incomplete. | 180 | no |
182
+| | [query_offset](#option-collection-query-offset) | Minimum offset (seconds) subtracted from metric query windows. Increase if metrics appear incomplete. | 180 | no |
183
| | timeout | Timeout for Azure Resource Graph and Azure Monitor API requests, in seconds. | 30 | no |
130
-| **Limits** | max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
131
-| | max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
132
-| | max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
133
-| **Profiles** | profile_selection_mode | Profile selection mode: `auto` discovers matching profiles via Azure Resource Graph, `exact` uses only listed profile ids, `combined` merges listed ids with auto-discovered profiles. | auto | no |
134
-| | profile_selection_mode_exact.profiles | Profile ids to enable (used when `profile_selection_mode` is `exact`). | [] | no |
135
-| | profile_selection_mode_combined.profiles | Profile ids to merge with auto-discovered profiles (used when `profile_selection_mode` is `combined`). | [] | no |
136
-| **Filters** | resource_groups | Optional list of resource group names to restrict monitoring scope. | [] | no |
137
-| **Authentication** | auth.mode | Authentication mode: `service_principal`, `managed_identity`, or `default`. | | yes |
184
+| **Authentication** | [auth.mode](#option-authentication-auth-mode) | Authentication method: `service_principal`, `managed_identity`, or `default`. | | yes |
185
| | auth.mode_service_principal.tenant_id | Entra ID tenant ID (required for `service_principal` mode). | | no |
186
| | auth.mode_service_principal.client_id | Entra ID application (client) ID (required for `service_principal` mode). | | no |
187
| | auth.mode_service_principal.client_secret | Entra ID client secret (required for `service_principal` mode). | | no |
188
| | auth.mode_managed_identity.client_id | Client ID for user-assigned managed identity. Leave empty for system-assigned. | | no |
189
+| **Discovery** | discovery.refresh_every | Interval (seconds) for refreshing discovered resources. Set `0` to disable runtime re-discovery after bootstrap. | 300 | no |
190
+| | [discovery.mode](#option-discovery-discovery-mode) | Resource discovery method: `filters` (structured filters) or `query` (custom KQL). | filters | no |
191
+| | discovery.mode_filters.resource_groups | Optional list of Azure resource groups to include in `filters` mode. | [] | no |
192
+| | discovery.mode_filters.regions | Optional list of Azure regions to include in `filters` mode. | [] | no |
193
+| | discovery.mode_filters.tags | Optional exact-match tag filters for `filters` mode. Keys are matched case-insensitively and values case-sensitively. | {} | no |
194
+| | [discovery.mode_query.kql](#option-discovery-discovery-mode-query-kql) | Custom Azure Resource Graph KQL for `query` mode. Must project `id`, `name`, `type`, `resourceGroup`, `location`. | | no |
195
+| **Profiles** | [profiles.mode](#option-profiles-profiles-mode) | How profiles are selected: `auto` (discover from resources), `exact` (explicit list), or `combined` (both). | auto | no |
196
+| | profiles.mode_exact.names | Explicit profile file basenames used by `exact` mode. Matching is case-insensitive. | [] | no |
197
+| | profiles.mode_combined.names | Explicit profile file basenames merged with auto-discovered profiles in `combined` mode. Matching is case-insensitive. | [] | no |
198
+| **Limits** | limits.max_concurrency | Maximum concurrent batch queries to Azure Monitor. | 4 | no |
199
+| | limits.max_batch_resources | Maximum resources per Azure Monitor batch request. | 50 | no |
200
+| | limits.max_metrics_per_query | Maximum metrics per Azure Monitor batch request. | 20 | no |
201
| **Virtual Node** | vnode | Associates this data collection job with a [Virtual Node](https://learn.netdata.cloud/docs/netdata-agent/configuration/organize-systems-metrics-and-alerts#virtual-nodes). | | no |
202
203
+<a id="option-collection-query-offset"></a>
204
+##### query_offset
205
+
206
+Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector subtracts this offset from the current time when building metric query windows to avoid fetching incomplete data points.
207
+
208
+The configured `query_offset` acts as a minimum floor. For slower metric batches, the collector automatically uses a larger effective offset when the batch time grain is longer than the configured value.
209
+
210
+- **Default (180s)** works for most services.
211
+- **Longer time grains** (for example `PT5M`) automatically use at least one full time grain as the effective offset.
212
+- **Increase to 240-300s** if you still see gaps or missing data points.
213
+- **Do not set below 60s** -- metrics will likely be incomplete.
214
+
215
+
216
+<a id="option-authentication-auth-mode"></a>
217
+##### auth.mode
218
+
219
+Determines how the collector authenticates with Azure.
220
+
221
+| Mode | When to use | Required options |
222
+|:-----|:------------|:-----------------|
223
+| `service_principal` | Running outside Azure, or when you need explicit credentials | `tenant_id`, `client_id`, `client_secret` |
224
+| `managed_identity` | Running on Azure VMs, VMSS, or AKS with a managed identity | Optionally `client_id` for user-assigned identity |
225
+| `default` | Uses the Azure SDK default credential chain (environment variables, managed identity, Azure CLI, etc.) | None |
226
+
227
+
228
+<a id="option-discovery-discovery-mode"></a>
229
+##### discovery.mode
230
+
231
+Controls how the collector finds candidate Azure resources.
232
+
233
+| Mode | Behavior |
234
+|:-----|:---------|
235
+| `filters` | Builds an Azure Resource Graph query from the structured `mode_filters.*` options (resource groups, regions, tags). This is the default. |
236
+| `query` | Uses the raw KQL you provide in `discovery.mode_query.kql`. The query must project `id`, `name`, `type`, `resourceGroup`, and `location`. |
237
+
238
+
239
+<a id="option-discovery-discovery-mode-query-kql"></a>
240
+##### discovery.mode_query.kql
241
+
242
+A raw Azure Resource Graph KQL query used when `discovery.mode` is `query`.
243
+
244
+The query **must** project these five columns:
245
+
246
+| Column | Description |
247
+|:-------|:------------|
248
+| `id` | Full Azure resource ID (ARM format) |
249
+| `name` | Resource name |
250
+| `type` | Resource type (e.g., `microsoft.sql/servers/databases`) |
251
+| `resourceGroup` | Resource group name |
252
+| `location` | Azure region |
253
+
254
+Example:
255
+
256
+```
257
+resources
258
+| where tags.env =~ "prod"
259
+| project id, name, type, resourceGroup, location
260
+```
261
+
262
+
263
+<a id="option-profiles-profiles-mode"></a>
264
+##### profiles.mode
265
+
266
+Controls how the collector decides which metric profiles to activate.
267
+
268
+| Mode | Behavior |
269
+|:-----|:---------|
270
+| `auto` | Discovers resource types in your subscriptions and enables matching built-in profiles automatically. This is the default. |
271
+| `exact` | Uses only the profile basenames listed under `profiles.mode_exact.names`. No auto-discovery. |
272
+| `combined` | Merges auto-discovered profiles with the basenames listed under `profiles.mode_combined.names`. |
273
+
274
+Profile basename matching is case-insensitive. A basename is the profile filename without the `.yaml` / `.yml` suffix.
275
+
276
+
277
278
</details>
279
@@ -182,14 +315,28 @@ sudo ./edit-config go.d/azure_monitor.conf
315
316
##### Examples
317
185
-###### Service principal (auto-discover all resources)
318
+###### Service principal with structured discovery
319
187
-Authenticate with a service principal and auto-discover all supported Azure resource types in the subscription.
320
+Authenticate with a service principal and auto-discover resources across two subscriptions, filtered to the `production-rg` resource group in `eastus` with the tag `env=prod`.
321
322
```yaml
323
jobs:
324
- name: prod
192
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
325
+ subscription_ids:
326
+ - "aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa"
327
+ - "bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb"
328
+ discovery:
329
+ mode: filters
330
+ mode_filters:
331
+ resource_groups:
332
+ - production-rg
333
+ regions:
334
+ - eastus
335
+ tags:
336
+ env:
337
+ - prod
338
+ profiles:
339
+ mode: auto
340
auth:
341
mode: service_principal
342
mode_service_principal:
@@ -198,59 +345,49 @@ jobs:
345
client_secret: "your-client-secret"
346
347
```
201
-###### Managed identity (Azure VM/VMSS/AKS)
202
-
203
-Use the managed identity of the Azure VM, VMSS, or AKS node where Netdata is running.
204
-
205
-<details open><summary>Config</summary>
206
-
207
-```yaml
208
-jobs:
209
- - name: prod
210
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
211
- auth:
212
- mode: managed_identity
213
-
214
-```
215
-</details>
348
+###### Managed identity with exact profiles
349
217
-###### Specific profiles only
218
-
219
-Monitor only specific Azure services instead of auto-discovering all resource types.
350
+Use a managed identity (on an Azure VM, VMSS, or AKS) and monitor only SQL Database and PostgreSQL Flexible Server resources -- skip auto-discovery of other services.
351
352
<details open><summary>Config</summary>
353
354
```yaml
355
jobs:
356
- name: databases
226
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
357
+ subscription_ids:
358
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
359
profiles:
228
- - sql_database
229
- - postgres_flexible
230
- - redis_cache
360
+ mode: exact
361
+ mode_exact:
362
+ names:
363
+ - sql_database
364
+ - postgres_flexible
365
auth:
232
- mode: service_principal
233
- mode_service_principal:
234
- tenant_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
235
- client_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
236
- client_secret: "your-client-secret"
366
+ mode: managed_identity
367
368
```
369
</details>
370
241
-###### Filter by resource group
371
+###### Custom Azure Resource Graph KQL
372
243
-Only monitor resources in specific resource groups.
373
+Replace the built-in discovery filters with your own KQL query. Useful when you need joins, computed columns, or filtering logic that structured filters cannot express.
374
375
<details open><summary>Config</summary>
376
377
```yaml
378
jobs:
249
- - name: prod-rg
250
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
251
- resource_groups:
252
- - production-rg
253
- - staging-rg
379
+ - name: prod-query
380
+ subscription_ids:
381
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
382
+ discovery:
383
+ mode: query
384
+ mode_query:
385
+ kql: |
386
+ resources
387
+ | where tags.env =~ "prod"
388
+ | project id, name, type, resourceGroup, location
389
+ profiles:
390
+ mode: auto
391
auth:
392
mode: default
393
@@ -259,14 +396,15 @@ jobs:
396
397
###### Azure Government cloud
398
262
-Connect to Azure Government cloud environment.
399
+Connect to an Azure Government environment. Set `cloud: government` to use the correct authentication and API endpoints.
400
401
<details open><summary>Config</summary>
402
403
```yaml
404
jobs:
405
- name: gov
269
- subscription_id: "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
406
+ subscription_ids:
407
+ - "xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx"
408
cloud: government
409
auth:
410
mode: service_principal
@@ -324,6 +462,7 @@ Labels:
462
| region | The Azure region where the resource is deployed. |
463
| resource_type | The Azure resource type identifier. |
464
| profile | The Azure Monitor profile id. |
465
+| subscription_id | The Azure subscription identifier. |
466
| resource_uid | The unique Azure resource identifier. |
467
468
Metrics:
@@ -441,31 +580,46 @@ docker logs netdata 2>&1 | grep azure_monitor
580
581
### No metrics are collected
582
444
-Verify the following:
445
-1. The service principal or managed identity has **Monitoring Reader** role on the subscription or resource group.
446
-2. The `subscription_id` in the configuration matches the subscription containing the target resources.
447
-3. Target resources are running and producing metrics (check Azure Portal > Metrics for the resource).
448
-4. Check the Netdata error log for authentication or API errors: `grep azure_monitor /var/log/netdata/error.log`.
583
+Check the following:
584
+
585
+- **Permissions** -- The principal has both **Monitoring Reader** and **Reader** roles on the target subscription.
586
+- **Subscription IDs** -- The `subscription_ids` list includes the correct subscription(s).
587
+- **Resources are active** -- Verify in Azure Portal > Metrics that the resources are producing metrics.
588
+- **Collector logs** -- Check for authentication or API errors:
589
+ ```bash
590
+ # systemd
591
+ journalctl -u netdata --namespace=netdata --grep azure_monitor --since "5 minutes ago"
592
+ # non-systemd
593
+ grep azure_monitor /var/log/netdata/collector.log
594
+ ```
595
596
597
### Missing metrics for some resource types
598
453
-Azure Monitor profiles are matched by resource type. If a resource type exists but no metrics appear:
454
-1. Ensure `profiles: [auto]` (default) is set, or the specific profile id is listed.
455
-2. Verify the resource type matches a built-in profile. Run `ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/` to see available profiles.
456
-3. Some metrics require the resource to be actively processing data (e.g., IoT Hub telemetry metrics only appear when devices send messages).
599
+Profiles are matched by Azure resource type. If a resource type exists but metrics are missing:
600
+
601
+- **Check profile mode** -- Ensure `profiles.mode: auto` (default), or explicitly list the profile basename under `profiles.mode_exact.names` or `profiles.mode_combined.names`.
602
+- **Verify a built-in profile exists** -- List available profiles:
603
+ ```bash
604
+ ls /usr/lib/netdata/conf.d/go.d/azure_monitor.profiles/default/
605
+ ```
606
+- **Check resource activity** -- Some metrics only appear when the resource is actively processing data (e.g., IoT Hub telemetry metrics require devices to be sending messages).
607
+- **New resource types after startup** -- Runtime discovery does not activate new profiles. Restart the collector if new resource types were added after bootstrap.
608
609
459
-### Metrics appear delayed
610
+### Charts have gaps or incomplete data
611
461
-Azure Monitor metrics have a built-in reporting delay of 1-3 minutes. The collector uses a `query_offset` (default: 180 seconds) to account for this.
462
-If metrics are missing or incomplete, try increasing `query_offset` to 240 or 300 seconds.
463
-Some metrics with longer time grains (e.g., PT5M) may take up to 5 minutes to appear.
612
+Azure Monitor metrics have a built-in reporting delay of **1-3 minutes**.
613
+
614
+- The collector uses `query_offset` (default: **180 seconds**) as the minimum offset for metric query windows.
615
+- Slower time-grain batches automatically use a larger effective offset when needed.
616
+- If metrics are still missing or incomplete, increase `query_offset` to **240** or **300** seconds.
617
618
619
### Authentication errors in sovereign clouds
620
621
For Azure Government or Azure China clouds, set the `cloud` parameter:
622
+
623
- Azure Government: `cloud: government`
624
- Azure China (21Vianet): `cloud: china`
625