@cryptotaxi247 / netdata / commits / cc5b8f138

ai-docs (#21043)

* ai-docs * ai-docs * updating docs and images * mdx to md

Shyam Sreevalsan committed Sep 25, 2025 at 21:59 UTC cc5b8f13869151dfd8f4f90e23bb2847e4ac8d36
16 files changed +685 -447
docs/.map/map.csv
+35 -27
@@ -110,13 +110,13 @@ https://github.com/netdata/netdata/edit/master/src/collectors/README.md,Collecti
110 https://github.com/netdata/netdata/edit/master/src/collectors/REFERENCE.md,Collectors configuration,Published,Collecting Metrics,
111 https://github.com/netdata/agent-service-discovery/edit/master/README.md,Service discovery,Published,Collecting Metrics,
112 https://github.com/netdata/netdata/edit/master/src/collectors/statsd.plugin/README.md,StatsD,Published,Collecting Metrics,
113 -https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/README.md,Metrics Centralization Points,Published,Collecting Metrics/Metrics Centralization Points,
114 -https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/configuration.md,Configuring Metrics Centralization Points,Published,Collecting Metrics/Metrics Centralization Points,
115 -https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/sizing-netdata-parents.md,Sizing Netdata Parents,Published,Collecting Metrics/Metrics Centralization Points,
116 -,Optimizing Netdata Children,Unpublished,Collecting Metrics/Metrics Centralization Points,
117 -https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/clustering-and-high-availability-of-netdata-parents.md,Clustering and High Availability of Netdata Parents,Published,Collecting Metrics/Metrics Centralization Points,
118 -https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/replication-of-past-samples.md,Replication of Past Samples,Published,Collecting Metrics/Metrics Centralization Points,
119 -https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/faq.md,FAQ on Metrics Centralization Points,Published,Collecting Metrics/Metrics Centralization Points,
113 +https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/README.md,Metrics Centralization Points,Published,Netdata Parents/Metrics Centralization Points,
114 +https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/configuration.md,Configuring Metrics Centralization Points,Published,Netdata Parents/Metrics Centralization Points,
115 +https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/sizing-netdata-parents.md,Sizing Netdata Parents,Published,Netdata Parents/Metrics Centralization Points,
116 +,Optimizing Netdata Children,Unpublished,Netdata Parents/Metrics Centralization Points,
117 +https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/clustering-and-high-availability-of-netdata-parents.md,Clustering and High Availability of Netdata Parents,Published,Netdata Parents/Metrics Centralization Points,
118 +https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/replication-of-past-samples.md,Replication of Past Samples,Published,Netdata Parents/Metrics Centralization Points,
119 +https://github.com/netdata/netdata/edit/master/docs/observability-centralization-points/metrics-centralization-points/faq.md,FAQ on Metrics Centralization Points,Published,Netdata Parents/Metrics Centralization Points,
120 https://github.com/netdata/netdata/edit/master/docs/collecting-metrics/system-metrics.md,System metrics,Unpublished,Collecting Metrics,"Netdata collects thousands of metrics from physical and virtual systems, IoT/edge devices, and containers with zero configuration."
121 https://github.com/netdata/netdata/edit/master/docs/collecting-metrics/application-metrics.md,Application metrics,Unpublished,Collecting Metrics,"Monitor and troubleshoot every application on your infrastructure with per-second metrics, zero configuration, and meaningful charts."
122 https://github.com/netdata/netdata/edit/master/docs/collecting-metrics/container-metrics.md,Container metrics,Unpublished,Collecting Metrics,Use Netdata to collect per-second utilization and application-level metrics from Linux/Docker containers and Kubernetes clusters.
@@ -162,26 +162,34 @@ cloud_notifications_integrations,,,,
162 https://github.com/netdata/netdata/edit/master/src/health/REFERENCE.md,Alert Configuration Reference,Published,Alerts & Notifications,
163 https://github.com/netdata/netdata/edit/master/src/web/api/health/README.md,Health API Calls,Published,Alerts & Notifications,
164 ,,,,
165 -https://github.com/netdata/netdata/edit/master/docs/category-overview-pages/machine-learning-and-assisted-troubleshooting.md,AI & ML,Published,AI & ML,
166 -https://github.com/netdata/netdata/edit/master/docs/learn/mcp.md,Model Context Protocol (MCP),Published,AI & ML,
167 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/ai-chat-netdata.md,Chat with Netdata,Published,AI & ML/Chat with Netdata,
168 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/claude-desktop.md,Claude Desktop,Published,AI & ML/Chat with Netdata,
169 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/cursor.md,Cursor,Published,AI & ML/Chat with Netdata,
170 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/jetbrains-ides.md,JetBrains IDEs,Published,AI & ML/Chat with Netdata,
171 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/netdata-web-client.md,Netdata Web Client,Published,AI & ML/Chat with Netdata,
172 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/vs-code.md,Visual Studio Code,Published,AI & ML/Chat with Netdata,
173 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md,DevOps Copilots,Published,AI & ML/DevOps Copilots,
174 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-devops-copilot/claude-code.md,Claude Code,Published,AI & ML/DevOps Copilots,
175 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-devops-copilot/gemini-cli.md,Gemini CLI,Published,AI & ML/DevOps Copilots,
176 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-insights.md,AI Insights,Published,AI & ML,
177 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/anomaly-advisor.md,Anomaly Advisor,Published,AI & ML,
178 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ml-anomaly-detection/ml-anomaly-detection.md,ML Anomaly Detection,Published,AI & ML/ML Anomaly Detection,
179 -https://github.com/netdata/netdata/edit/master/docs/ml-ai/ml-anomaly-detection/ml-accuracy.md,ML Accuracy,Published,AI & ML/ML Anomaly Detection,"Analysis of Netdata's ML anomaly detection accuracy, false positive rates, and comparison with other approaches"
180 -https://github.com/netdata/netdata/edit/master/src/ml/ml-configuration.md,ML Configuration,Published,AI & ML/ML Anomaly Detection,
181 -https://github.com/netdata/netdata/edit/master/docs/metric-correlations.md,Metric Correlations,Published,AI & ML/ML Anomaly Detection,Quickly find metrics and charts closely related to a particular timeframe of interest anywhere in your infrastructure to discover the root cause faster.
182 -https://github.com/netdata/netdata/edit/master/docs/troubleshooting/troubleshoot.md,AI-Powered Alert Troubleshooting,Published,AI & ML,
183 -https://github.com/netdata/netdata/edit/master/docs/troubleshooting/custom-investigations.md,Custom Investigations,Published,AI & ML,
184 -,,,,
165 +https://github.com/netdata/netdata/edit/master/docs/category-overview-pages/machine-learning-and-assisted-troubleshooting.md,Netdata AI,Published,Netdata AI,
166 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-insights.md,Insights,Published,Netdata AI/Insights,
167 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/insights/infrastructure-summary.md,Infrastructure Summary,Published,Netdata AI/Insights,
168 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/insights/performance-optimization.md,Performance Optimization,Published,Netdata AI/Insights,
169 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/insights/capacity-planning.md,Capacity Planning,Published,Netdata AI/Insights,
170 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/insights/anomaly-analysis.md,Anomaly Analysis,Published,Netdata AI/Insights,
171 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/insights/scheduled-reports.md,Scheduled Reports,Published,Netdata AI/Insights,
172 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/investigations/index.md,Investigations,Published,Netdata AI/Investigations,
173 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/investigations/custom-investigations.md,Custom Investigations,Published,Netdata AI/Investigations,
174 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/investigations/scheduled-investigations.md,Scheduled Investigations,Published,Netdata AI/Investigations,
175 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/troubleshooting/index.md,Troubleshooting,Published,Netdata AI/Troubleshooting,
176 +https://github.com/netdata/netdata/edit/master/docs/troubleshooting/troubleshoot.md,Alert Troubleshooting,Published,Netdata AI/Troubleshooting,
177 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/anomaly-advisor.md,Anomaly Advisor,Published,Netdata AI/Troubleshooting,
178 +https://github.com/netdata/netdata/edit/master/docs/metric-correlations.md,Metric Correlations,Published,Netdata AI/Troubleshooting,Quickly find metrics and charts closely related to a particular timeframe of interest anywhere in your infrastructure to discover the root cause faster.
179 +https://github.com/netdata/netdata/edit/master/docs/netdata-ai/troubleshooting/troubleshoot-button.md,Troubleshoot Button,Published,Netdata AI/Troubleshooting,
180 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ml-anomaly-detection/ml-anomaly-detection.md,Anomaly Detection,Published,Netdata AI/Anomaly Detection,
181 +https://github.com/netdata/netdata/edit/master/src/ml/ml-configuration.md,ML Configuration,Published,Netdata AI/Anomaly Detection,
182 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ml-anomaly-detection/ml-accuracy.md,ML Accuracy,Published,Netdata AI/Anomaly Detection,"Analysis of Netdata's ML anomaly detection accuracy, false positive rates, and comparison with other approaches"
183 +https://github.com/netdata/netdata/edit/master/docs/learn/mcp.md,MCP,Published,Netdata AI/MCP,
184 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/ai-chat-netdata.md,Chat with Netdata,Published,Netdata AI/MCP,
185 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md,MCP Clients,Published,Netdata AI/MCP/MCP Clients,
186 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/claude-desktop.md,Claude Desktop,Published,Netdata AI/MCP/MCP Clients,
187 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/cursor.md,Cursor,Published,Netdata AI/MCP/MCP Clients,
188 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/vs-code.md,Visual Studio Code,Published,Netdata AI/MCP/MCP Clients,
189 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/jetbrains-ides.md,JetBrains IDEs,Published,Netdata AI/MCP/MCP Clients,
190 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-chat-netdata/netdata-web-client.md,Netdata Web Client,Published,Netdata AI/MCP/MCP Clients,
191 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-devops-copilot/claude-code.md,Claude Code,Published,Netdata AI/MCP/MCP Clients,
192 +https://github.com/netdata/netdata/edit/master/docs/ml-ai/ai-devops-copilot/gemini-cli.md,Gemini CLI,Published,Netdata AI/MCP/MCP Clients,
193 https://github.com/netdata/netdata/edit/master/docs/netdata-assistant.md,AI powered troubleshooting assistant,Unpublished,AI and Machine Learning,
194 https://github.com/netdata/netdata/edit/master/src/ml/README.md,ML models and anomaly detection,Unpublished,AI and Machine Learning,This is an in-depth look at how Netdata uses ML to detect anomalies.
195 ,,,,
docs/category-overview-pages/machine-learning-and-assisted-troubleshooting.md
+33 -176
@@ -1,198 +1,55 @@
1 -# AI and Machine Learning
1 +# Netdata AI
2
3 -Netdata provides powerful AI-driven capabilities to transform how you monitor and troubleshoot your infrastructure, with more innovations coming soon.
3 +Netdata AI is a set of analysis and troubleshooting capabilities built into Netdata Cloud. It turns high‑fidelity telemetry into explanations, timelines, and recommendations so teams resolve issues faster and document decisions with confidence.
4
5 -## What's Available Today
5 +![Netdata AI overview](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/netdata-ai.png)
6
7 -### 1. AI Chat with Netdata
7 +## Why it’s accurate and powerful
8
9 -**Available Now** - Chat with your infrastructure using natural language
9 +- Per‑second granularity: Every Netdata Agent collects metrics at 1‑second resolution, preserving short‑lived spikes and transient behavior.
10 +- On‑device ML: Unsupervised models run on every agent, continuously scoring anomalies for every metric with zero configuration.
11 +- Evidence‑based correlation: Netdata’s correlation engine relates metrics, anomalies, and events across nodes to form defendable root‑cause hypotheses.
12 +- Full context: Reports and investigations combine statistical summaries, anomaly timelines, alert history, and dependency information.
13
11 -Ask questions about your infrastructure like you're talking to a colleague. Get instant answers about performance, find specific logs, identify top resource consumers, or investigate issues - all through simple conversation. No more complex queries or dashboard hunting.
14 +## Capabilities
15
13 -**Key capabilities**:
16 +### 1) Insights
17
15 -- **Natural language queries** - "Which servers have high CPU usage?" or "Show database errors from last hour" or "What is wrong with my infrastructure now", or "Do a post-mortem analysis of the outage we had yesteday", or "Show me all network dependencies of process X"
16 -- **Multi-node visibility** - Analyzes your entire infrastructure through Netdata Parents
17 -- **Flexible AI options** - Use your existing AI tools or our standalone web chat
18 +Generates on‑demand, professional reports (see [AI Insights](/docs/ml-ai/ai-insights.md)):
19
19 -<details>
20 -<summary><strong>How it works</strong></summary>
20 +- [Infrastructure Summary](/docs/netdata-ai/insights/infrastructure-summary.md) – incident timelines, health, and prioritized actions
21 +- [Performance Optimization](/docs/netdata-ai/insights/performance-optimization.md) – bottlenecks, contention, and concrete tuning steps
22 +- [Capacity Planning](/docs/netdata-ai/insights/capacity-planning.md) – growth projections and exhaustion dates
23 +- [Anomaly Analysis](/docs/netdata-ai/insights/anomaly-analysis.md) – forensics on unusual behavior and likely causes
24
22 -- **MCP integration** - You chat with an LLM, that has access to your observability data, via Model Context Protocol (MCP)
23 -- **Choice of AI providers** - Claude, GPT-4, Gemini, and others
24 -- **Two deployment options** - Use an existing AI client that supports MCP, or use a web page chat we created for it (LLM is pay-per-use with API keys)
25 -- **Real-time data access** - Query live metrics, logs, processes, network connections, and system state
26 -- **Secure connection** - LLM has access to your data via the LLM client
25 +Each report includes an executive summary, evidence, and actionable recommendations. Reports are downloadable as PDFs and shareable with your team. You can also [schedule reports](/docs/netdata-ai/insights/scheduled-reports.md).
26
28 -</details>
27 +### 2) Investigations
28
30 -**Access**: Available now for all Netdata Agent deployments (Standalone and Parents)
29 +Ask open‑ended questions (“what changed here?”, “why did X regress?”) and get a researched answer using your telemetry — see the [Investigations overview](/docs/netdata-ai/investigations/index.md). Launch from the “Troubleshoot with AI” button (captures current scope) or from Insights → New Investigation. Create [Custom Investigations](/docs/netdata-ai/investigations/custom-investigations.md) and set up [Scheduled Investigations](/docs/netdata-ai/investigations/scheduled-investigations.md).
30
32 -[Explore AI Chat →](./chat-with-netdata-mcp)
31 +### 3) Troubleshooting
32
34 -### 2. AI DevOps Copilot
33 +- [Alert Troubleshooting](/docs/troubleshooting/troubleshoot.md) – one‑click analysis for any alert with a root‑cause hypothesis and supporting signals
34 +- [Anomaly Advisor](/docs/ml-ai/anomaly-advisor.md) – interactive exploration of how anomalies propagate across systems
35 +- [Metric Correlations](/docs/metric-correlations.md) – focus on the most relevant charts for any time window
36
36 -**Available Now** - Transform observability into action with CLI AI assistants
37 +See the [Troubleshooting overview](/docs/netdata-ai/troubleshooting/index.md). From any view, use the [Troubleshoot with AI button](/docs/netdata-ai/troubleshooting/troubleshoot-button.md).
38
38 -Combine the power of AI with system automation. CLI-based AI assistants like Claude Code and Gemini CLI can access your Netdata metrics and execute commands, enabling intelligent infrastructure optimization, automated troubleshooting, and configuration management - all driven by real observability data.
39 +### 4) Anomaly Detection
40
40 -**Key capabilities**:
41 +Local, unsupervised ML runs on every agent, learning normal behavior and scoring anomalies for all metrics in real time. Anomaly ribbons appear on charts, and historical scores are stored alongside metrics for analysis. See [ML Anomaly Detection](/docs/ml-ai/ml-anomaly-detection/ml-anomaly-detection.md), configure via [ML Configuration](/src/ml/ml-configuration.md), and review methodology in [ML Accuracy](/docs/ml-ai/ml-anomaly-detection/ml-accuracy.md).
42
42 -- **Observability-driven automation** - AI analyzes metrics and executes fixes
43 -- **Infrastructure optimization** - Automatic tuning based on performance data
44 -- **Intelligent troubleshooting** - From problem detection to resolution
45 -- **Configuration management** - AI-generated configs based on actual usage
43 +### 5) MCP (Model Context Protocol)
44
47 -<details>
48 -<summary><strong>How it works</strong></summary>
45 +Connect AI clients to Netdata’s MCP server to bring live observability into natural‑language workflows and optional automation. Options include [MCP](/docs/learn/mcp.md), [Chat with Netdata](/docs/ml-ai/ai-chat-netdata/ai-chat-netdata.md), and [MCP Clients](/docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md) like Claude Desktop, Cursor, VS Code, JetBrains IDEs, Claude Code, Gemini CLI, and the Netdata Web Client.
46
50 -- **MCP-enabled CLI tools** - Claude Code, Gemini CLI, and others
51 -- **Bidirectional integration** - Read metrics, execute commands
52 -- **Context-aware decisions** - AI understands your infrastructure state
53 -- **Safe execution** - Review AI suggestions before implementation
54 -- **Team collaboration** - Share configurations via version control
47 +## Usage and credits
48
56 -</details>
49 +- Eligible Spaces receive 10 free AI credits; each Insights report, investigation, or alert troubleshooting run consumes 1 AI credit.
50 +- Additional usage is available via AI Credits. Track usage from Settings → Usage & Billing → AI Credits.
51
58 -**Access**: Available now with MCP-supported CLI AI tools
52 +## Note
53
60 -[Explore AI DevOps Copilot →](./ai-devops-copilot/ai-devops-copilot)
61 -
62 -### 3. AI Insights
63 -
64 -**Preview (Netdata Cloud Feature)** - Strategic infrastructure analysis in minutes
65 -
66 -Transform past data into actionable insights with AI-generated reports. Perfect for capacity planning, performance reviews, and executive briefings. Get comprehensive analysis of your infrastructure trends, optimization opportunities, and future requirements - all in professionally formatted PDFs.
67 -
68 -**Four report types**:
69 -
70 -- **Infrastructure Summary** - Complete system health and incident analysis
71 -- **Capacity Planning** - Growth projections and resource recommendations
72 -- **Performance Optimization** - Bottleneck identification and tuning suggestions
73 -- **Anomaly Analysis** - Deep dive into unusual patterns and their impacts
74 -
75 -<details>
76 -<summary><strong>How it works</strong></summary>
77 -
78 -- **2-3 minute generation** - Analyzes historical data comprehensively
79 -- **PDF downloads** - Professional reports ready for sharing
80 -- **Embedded visualizations** - Charts and graphs from your actual data
81 -- **Executive-ready** - Clear summaries with technical details included
82 -- **Secure processing** - Data analyzed then immediately discarded
83 -
84 -</details>
85 -
86 -**Access**:
87 -
88 -- Business subscriptions: Unlimited reports
89 -- Free trial users: Full access during trial
90 -- Community users: 10 free reports ([request early access](https://discord.gg/mPZ6WZKKG2))
91 -
92 -[Explore AI Reports →](./ai-insights)
93 -
94 -
95 -### 4. Anomaly Advisor
96 -
97 -**Available to All** - Revolutionary troubleshooting that finds root causes in minutes
98 -
99 -Stop guessing what went wrong. The Anomaly Advisor instantly shows you how problems cascade across your infrastructure and ranks every metric by anomaly severity. Root causes typically appear in the top 20-30 results, turning hours of investigation into minutes of discovery.
100 -
101 -**Revolutionary approach**:
102 -
103 -- **See cascading effects** - Watch anomalies propagate across systems
104 -- **Automatic ranking** - Every metric scored and sorted by anomaly severity
105 -- **No expertise required** - Works even on unfamiliar systems
106 -
107 -<details>
108 -<summary><strong>How it works</strong></summary>
109 -
110 -- **Data-driven analysis** - No hypotheses needed, the data reveals the story
111 -- **Influence tracking** - Shows what influenced and what was influenced
112 -- **Time window analysis** - Highlight any incident period for investigation
113 -- **Scale-agnostic** - Works identically from 10 to 10,000 nodes
114 -- **Visual propagation** - See anomaly clusters and cascades instantly
115 -
116 -</details>
117 -
118 -**Find it**: Anomalies tab in any Netdata dashboard
119 -
120 -[Learn more about Anomaly Advisor →](./anomaly-advisor)
121 -
122 -### 5. Machine Learning Anomaly Detection
123 -
124 -**Available to All** - Continuous anomaly detection on every metric
125 -
126 -The foundation of Netdata's AI capabilities. Machine learning models run locally on every agent, continuously learning normal patterns and detecting anomalies in real-time. Zero configuration required - it just works, protecting your infrastructure 24/7.
127 -
128 -**Automatic protection**:
129 -
130 -- **Every metric monitored** - ML analyzes all metrics continuously
131 -- **Visual anomaly indicators** - Purple ribbons on every chart show anomaly rates
132 -- **Historical anomaly data** - ML scores saved with metrics for past analysis
133 -- **Zero configuration** - Starts working immediately after installation
134 -
135 -<details>
136 -<summary><strong>How it works</strong></summary>
137 -
138 -- **Local ML engine** - Runs on every Netdata Agent, no cloud dependency
139 -- **Multiple models** - Consensus approach reduces noise and false positives by 99%
140 -- **Integrated storage** - Anomaly scores saved in the database with metrics
141 -- **Historical queries** - Query past anomaly rates just like any other metric
142 -- **Visual integration** - Purple anomaly ribbons appear on all charts automatically
143 -- **Minimal overhead** - Designed for production environments
144 -- **Privacy by design** - Your data never leaves your infrastructure
145 -
146 -</details>
147 -
148 -**Access**: Free for everyone - enabled by default
149 -
150 -[Explore Machine Learning →](./machine-learning-anomaly-detection)
151 -
152 -### 6. AI-Powered Alert Troubleshooting
153 -
154 -When an alert fires, you can now use AI to get a detailed troubleshooting report that determines whether the alert requires immediate action or is just noise. The AI examines your alert's history, correlates it with thousands of other metrics across your infrastructure, and provides actionable insights—all within minutes.
155 -
156 -**Key capabilities**:
157 -- **Automated Analysis:** Click "Ask AI" on any alert to generate a comprehensive troubleshooting report
158 -- **Correlation Discovery:** AI scans thousands of metrics to find what else was behaving abnormally
159 -- **Root Cause Hypothesis:** Get likely root causes with specific metrics and dimensions that matter most
160 -- **Noise Reduction:** Quickly identify false positives versus legitimate issues
161 -
162 -**How to access**:
163 -- From the Alerts tab: Click the "Ask AI" button on any alert
164 -- From the Insights tab: Select "Alert Troubleshooting" and choose an alert
165 -- From email notifications: Click "Troubleshoot with AI" link
166 -
167 -Reports are generated in 1-2 minutes and saved in your Insights tab. All Business plan users get 10 AI troubleshooting sessions per month during trial.
168 -
169 -**Access**: Netdata Cloud Business Feature
170 -
171 -## Coming Soon
172 -
173 -### AI Chat with Netdata (Netdata Cloud version)
174 -
175 -**In Development** - Chat with your entire infrastructure through Netdata Cloud
176 -
177 -Soon, Netdata Cloud will become an MCP server itself. This means you'll be able to chat with your entire infrastructure without setting up local MCP bridges. Get the same natural language capabilities with the added benefits of Cloud's global view, team collaboration, and seamless access from anywhere.
178 -
179 -**What to expect**:
180 -
181 -- Direct MCP integration with Netdata Cloud
182 -- Chat with all your infrastructure from one place
183 -- No local bridge setup required
184 -- Team collaboration on AI conversations
185 -- Access from any device, anywhere
186 -
187 -### AI Weekly Digest
188 -
189 -**In Development (Netdata Cloud)** - Your infrastructure insights delivered weekly
190 -
191 -Stay informed without information overload. The AI Weekly Digest will analyze your infrastructure's performance over the past week and deliver a concise summary of what matters most - trends, issues resolved, optimization opportunities, and what to watch next week.
192 -
193 -**What to expect**:
194 -
195 -- Weekly email summaries customized for your role
196 -- Key metrics and trend analysis
197 -- Proactive recommendations for the week ahead
198 -- Highlights of resolved and ongoing issues
54 +- No model training on your data: information is used only to generate your outputs.
55 +- Despite our best efforts to eliminate inaccuracies, AI responses may sometimes be incorrect, please think carefully before making important changes or decisions.
docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md
+28 -16
@@ -1,32 +1,32 @@
1 -# AI DevOps Copilot
1 +# MCP Clients
2
3 -Command-line AI assistants like **Claude Code** and **Gemini CLI** represent a revolutionary shift in how infrastructure professionals work. These tools combine the power of large language models with access to observability data and the ability to execute system commands, creating unprecedented automation opportunities.
3 +Model Context Protocol (MCP) clients like **Claude Desktop**, **Cursor**, **Visual Studio Code**, **JetBrains IDEs**, **Netdata Web Client**, **Claude Code**, and **Gemini CLI** can connect to Netdata’s MCP server to bring real observability data into your AI workflows. This enables natural‑language analysis with context from your infrastructure and, for CLI tools, optional automation.
4
5 -## The Power of CLI-based AI Assistants
5 +## The power of MCP clients
6
7 ### Key Capabilities
8
9 -**Observability-Driven Operations:**
9 +**Observability‑driven operations**
10
11 - Access real-time metrics and logs from monitoring systems
12 - Analyze performance trends and identify bottlenecks
13 - Correlate issues across multiple systems and services
14
15 -**System Configuration Management:**
15 +**System configuration management**
16
17 - Generate and modify configuration files based on observed conditions
18 - Implement best practices automatically
19 - Adapt configurations to changing requirements
20
21 -**Automated Troubleshooting:**
21 +**Automated troubleshooting**
22
23 - Diagnose issues using multiple data sources
24 - Execute diagnostic commands and interpret results
25 - Implement fixes based on root cause analysis
26
27 -## Observability + Automation Use Cases
27 +## Observability + automation use cases
28
29 -When AI assistants have access to observability data (like Netdata through MCP), they can make informed decisions about system changes:
29 +When MCP clients have access to Netdata, they can make informed decisions about system changes:
30
31 ### Infrastructure Optimization Examples
32
@@ -106,9 +106,9 @@ Keep in mind however, that usually this prompt should be split into multiple sma
106
107 This showcases how AI can combine application expertise, infrastructure knowledge, and observability best practices to create sophisticated testing environments that would typically require weeks of manual setup and deep domain expertise.
108
109 -## ⚠️ Critical Security and Safety Considerations
109 +## ⚠️ Critical security and safety considerations
110
111 -### Command Execution Risks
111 +### Command execution risks
112
113 **LLMs Are Not Infallible:**
114
@@ -122,7 +122,7 @@ This showcases how AI can combine application expertise, infrastructure knowledg
122 - Changes may have cascading effects across interconnected services
123 - Recovery from AI-generated misconfigurations can be time-consuming
124
125 -### Data Privacy and Security Concerns
125 +### Data privacy and security concerns
126
127 **External LLM Provider Exposure:**
128
@@ -138,7 +138,7 @@ This showcases how AI can combine application expertise, infrastructure knowledg
138 - Application secrets and encryption keys
139 - User data and personally identifiable information
140
141 -### Recommended Safe Usage Practices
141 +### Recommended safe usage practices
142
143 **1. Analysis-First Approach:**
144
@@ -176,9 +176,9 @@ high usage and what solutions you recommend
176 - Implement change management processes for AI-suggested modifications
177 - Maintain air-gapped environments for highly sensitive systems
178
179 -## Best Practices for Implementation
179 +## Best practices for implementation
180
181 -### Safe Integration Workflow
181 +### Safe integration workflow
182
183 1. **Discovery Phase:** Let AI analyze your current setup and identify opportunities
184 2. **Planning Phase:** Have AI generate detailed implementation plans with explanations
@@ -187,14 +187,26 @@ high usage and what solutions you recommend
187 5. **Validation Phase:** Verify results match expectations before production deployment
188 6. **Documentation Phase:** Have AI help document the changes and their rationale
189
190 -### Building Trust Over Time
190 +### Building trust over time
191
192 - Start with simple, low-risk tasks to build confidence
193 - Gradually increase complexity as you validate AI accuracy
194 - Develop institutional knowledge about AI strengths and limitations
195 - Create feedback loops to improve AI prompts and instructions
196
197 -### Team Education and Guidelines
197 +### Team education and guidelines
198 +
199 +## Client guides
200 +
201 +See dedicated configuration guides for each client:
202 +
203 +- Claude Desktop
204 +- Cursor
205 +- Visual Studio Code
206 +- JetBrains IDEs
207 +- Netdata Web Client
208 +- Claude Code
209 +- Gemini CLI
210
211 - Train team members on safe AI usage practices
212 - Establish clear guidelines for when AI assistance is appropriate
docs/ml-ai/ai-insights.md
+29 -211
@@ -1,227 +1,45 @@
1 # AI Insights
2
3 -**From hours of debugging to minutes of clarity** - AI Insights transforms your infrastructure monitoring data into professional reports that explain what happened, why it happened, and what to do about it.
3 +AI Insights generates on‑demand reports from your Netdata telemetry to explain what happened, why it happened, and recommended next steps. Reports use per‑second metrics, local anomaly scores, and correlation across nodes, then present evidence and actions in a concise, shareable format.
4
5 -## The Challenge AI Insights Solves
5 +![Insights overview](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/insights.png)
6
7 -Traditional monitoring requires you to manually query metrics, correlate data, and build dashboards during incidents - all while the clock is ticking. Even experienced engineers struggle with:
7 +## Report types
8
9 -- Learning complex query languages (PromQL, SQL) just to ask basic questions
10 -- Building custom dashboards during incidents instead of fixing problems
11 -- Correlating metrics across multiple systems to find root causes
12 -- Translating technical metrics into business impact for stakeholders
13 -- Spending hours on post-incident analysis and reporting
9 +- [Infrastructure Summary](/docs/netdata-ai/insights/infrastructure-summary.md)
10 +- [Performance Optimization](/docs/netdata-ai/insights/performance-optimization.md)
11 +- [Capacity Planning](/docs/netdata-ai/insights/capacity-planning.md)
12 +- [Anomaly Analysis](/docs/netdata-ai/insights/anomaly-analysis.md)
13
15 -**AI Insights eliminates these barriers** by automatically analyzing your infrastructure and delivering comprehensive reports that provide both executive summaries and technical deep-dives.
14 +Schedule recurring runs: [Scheduled Reports](/docs/netdata-ai/insights/scheduled-reports.md)
15
17 -## Why AI Insights Transforms Operations
16 +## Generate a report
17
19 -- **No query languages needed** - Skip the learning curve of PromQL, SQL, or custom dashboards
20 -- **AI with SRE expertise** - Get analysis from an AI trained to think like a senior engineer
21 -- **Root cause, not symptoms** - Understand the cascade of issues, not just surface metrics
22 -- **Business context included** - Reports explain technical issues in terms of business impact
23 -- **Collaborative by design** - Share professional PDFs with stakeholders who need answers, not dashboards
24 -- **Powered by Netdata's ML** - Leverages anomaly scores from ML models trained on every metric
25 -- **Zero configuration needed** - Works immediately with your existing Netdata deployment
18 +1. Open Netdata Cloud → Insights
19 +2. Select a report type
20 +3. Configure time range and scope (rooms/nodes)
21 +4. Optional: adjust sensitivity or focus (varies by report)
22 +5. Click Generate (reports complete in ~2–3 minutes)
23
27 -## Four Specialized Report Types
24 +Reports appear in the Insights tab and are downloadable as PDFs. An email notification is sent when a report is ready.
25
29 -![AI Insights Report Example](https://github.com/user-attachments/assets/c6997afb-94cb-41cc-a038-b384cb92e751)
26 +## Parameters and scope
27
31 -### Infrastructure Summary
28 +- Time range: 6h–30d typical windows; longer ranges supported by some reports
29 +- Scope: entire Space, selected rooms, or specific nodes
30 +- Sensitivity/focus: report‑specific options (see the individual report pages)
31
33 -**Your automated health check and incident analyst**
32 +## Output
33
35 -Perfect for Monday morning reviews, post-incident analysis, or executive updates. This report provides:
34 +- Executive summary with key findings
35 +- Evidence: charts, anomaly timelines, alert/event context
36 +- Recommendations with rationale
37 +- PDF download and shareable view in Netdata Cloud
38
37 -- Complete system health assessment with prioritized issues
38 -- Timeline of incidents and their business impact
39 -- Critical alerts analysis with resolution recommendations
40 -- Top 3 actionable items to improve infrastructure health
41 -- Performance trends across all key metrics
39 +## How it works (high level)
40
43 -**Use cases**: Weekend incident recovery, executive briefings, team handoffs, regular health checks
41 +- Collects the relevant metrics, anomaly scores, and alerts from your agents
42 +- Compresses them into a structured context (summaries, correlations, timelines)
43 +- Uses a model to synthesize explanations and recommended actions from that context
44
45 -### Capacity Planning
46 -
47 -**Stop guessing future needs - get data-driven projections**
48 -
49 -Make informed decisions about infrastructure investments with reports that include:
50 -
51 -- Resource utilization trends and growth patterns
52 -- Predicted capacity exhaustion dates for critical resources
53 -- Specific hardware recommendations based on usage patterns
54 -- Cost optimization opportunities
55 -- Projections for 3 months to 2 years ahead
56 -
57 -**Use cases**: Quarterly planning, budget justification, infrastructure roadmaps, vendor negotiations
58 -
59 -### Performance Optimization
60 -
61 -**Find and fix bottlenecks before users complain**
62 -
63 -Identify inefficiencies and optimization opportunities with:
64 -
65 -- Bottleneck analysis across application, database, network, and storage
66 -- Resource contention patterns and their impact
67 -- Specific tuning recommendations with expected improvements
68 -- Prioritized list of optimizations by potential impact
69 -- Before/after projections for recommended changes
70 -
71 -**Use cases**: Performance audits, system tuning, SRE optimization projects, efficiency improvements
72 -
73 -### Anomaly Analysis
74 -
75 -**Post-incident forensics made simple**
76 -
77 -Understand unusual patterns and prevent future issues with:
78 -
79 -- ML-detected anomalies with severity scoring
80 -- Root cause analysis showing how issues cascaded
81 -- Timeline reconstruction of anomaly propagation
82 -- Correlation between different system anomalies
83 -- Recommendations to prevent recurrence
84 -
85 -**Use cases**: Post-mortems, proactive issue detection, system behavior analysis, troubleshooting
86 -
87 -## Customize Reports to Your Needs
88 -
89 -Each report type offers flexible customization options for content and analysis scope (note: report structure and visual style are standardized for consistency):
90 -
91 -### Time Period Selection
92 -
93 -- **Infrastructure Summary**: Last 24 hours, 48 hours, 7 days, or month
94 -- **Capacity Planning**: Forecast for 3 months, 6 months, 1 year, or 2 years
95 -- **Performance Optimization**: Last 24 hours, 7 days, month, or quarter
96 -- **Anomaly Analysis**: Last 6 hours, 12 hours, 24 hours, or 7 days
97 -
98 -### Scope and Filtering
99 -
100 -- **Node Selection**: Analyze specific servers or your entire infrastructure
101 -- **Metric Categories**: Focus on CPU, Memory, Disk, Network, or Applications
102 -- **Resource Types**: Target Compute, Storage, Network, or Database resources
103 -- **Focus Areas**: Drill into specific performance domains
104 -- **Anomaly Thresholds**: Set sensitivity levels (10%, 20%, or 30%)
105 -
106 -## How AI Insights Works
107 -
108 -### 1. Intelligent Data Collection
109 -
110 -When you request a report, AI Insights:
111 -
112 -- Gathers relevant metrics from your selected time period and nodes
113 -- Collects active alerts and their severity levels
114 -- Retrieves ML-detected anomalies and their scores
115 -- Maps system relationships and dependencies
116 -- Compiles process and application performance data
117 -
118 -### 2. AI-Powered Analysis
119 -
120 -The collected data is analyzed by Anthropic's Claude 3.7 Sonnet model, optimized for infrastructure telemetry analysis using SRE methodologies. This AI model:
121 -
122 -- Applies SRE-level expertise to identify patterns
123 -- Correlates issues across different systems
124 -- Determines root causes vs symptoms
125 -- Prioritizes findings by business impact
126 -- Generates actionable recommendations
127 -
128 -### 3. Professional Report Generation
129 -
130 -Within 2-3 minutes, you receive:
131 -
132 -- **Structured content**: Headers, insights, charts, and tables in logical flow
133 -- **Embedded visualizations**: Charts generated from your actual metrics
134 -- **Executive summary**: High-level findings for stakeholders
135 -- **Technical details**: Deep-dive analysis for engineers
136 -- **Action items**: Prioritized recommendations with clear next steps
137 -- **PDF format**: Professional reports ready for sharing
138 -
139 -### 4. Security and Privacy
140 -
141 -- **In-memory processing**: Data analyzed then immediately discarded
142 -- **No training data**: Your infrastructure data is never used for model training
143 -- **Secure API**: All communications encrypted end-to-end
144 -- **Access controlled**: Respects your existing Netdata permissions
145 -
146 -## Real-World Impact
147 -
148 -From the Inrento fintech case study:
149 -> "AI Insights provided **significant time savings** in identifying and resolving issues. It **drastically reduced the time spent** identifying problems and implementing solutions, leading to **enhanced productivity and performance** with **minimized downtime**."
150 -Teams report that incident analysis that previously took hours of manual investigation now completes in minutes with AI Insights.
151 -
152 -## Perfect For
153 -
154 -- **Incident post-mortems**: Generate comprehensive analysis in minutes, not hours
155 -- **Executive briefings**: Professional PDFs with clear summaries and visualizations
156 -- **Capacity reviews**: Data-driven planning for budget and resource allocation
157 -- **Performance audits**: Regular health checks without manual analysis
158 -- **Team handoffs**: Share context-rich reports instead of dashboard links
159 -- **Compliance reporting**: Document infrastructure state and changes
160 -- **Vendor discussions**: Data-backed evidence for infrastructure decisions
161 -
162 -## Unlike Traditional Monitoring
163 -
164 -AI Insights represents a paradigm shift in infrastructure monitoring:
165 -
166 -| Traditional Monitoring | AI Insights |
167 -|------------------------|-------------|
168 -| Build dashboards during incidents | Get instant analysis |
169 -| Learn query languages | Use natural language selection |
170 -| Manual correlation across metrics | Automatic relationship detection |
171 -| Raw metrics without context | Narrative explanations with context |
172 -| Technical data only | Business impact included |
173 -| Hours of manual analysis | 2-3 minute automated reports |
174 -
175 -## What Sets AI Insights Apart
176 -
177 -Unlike traditional AI monitoring assistants that require extensive configuration or operate as black-box cloud services, AI Insights:
178 -
179 -- **Runs entirely on your infrastructure** - No external dependencies or mysterious cloud processing
180 -- **Uses your actual data** - Not generic patterns or industry averages
181 -- **Provides transparent analysis** - Clear reasoning, not black-box decisions
182 -- **Respects your security** - Data never leaves your control
183 -- **Works instantly** - No training period or configuration required
184 -
185 -## Getting Started
186 -
187 -1. **Access AI Insights** from the Netdata Cloud navigation menu
188 -2. **Select a report type** based on your current need
189 -3. **Customize parameters** like time period and node selection
190 -4. **Generate report** and receive it within 2-3 minutes
191 -5. **Share or download** the PDF for stakeholders
192 -
193 -## Technical Requirements
194 -
195 -- Active Netdata Cloud account
196 -- At least one connected Netdata Agent
197 -- Historical data (minimum 24 hours recommended)
198 -- No additional configuration needed
199 -
200 -## Frequently Asked Questions
201 -
202 -**Q: How far back can AI Insights analyze data?**
203 -A: AI Insights can analyze any data retained by your Netdata agents, from 6 hours to 2 years depending on the report type and your retention settings.
204 -
205 -**Q: Can I schedule regular reports?**
206 -A: Currently reports are generated on-demand. Scheduled reports are on the roadmap.
207 -
208 -**Q: What metrics are included in the analysis?**
209 -A: AI Insights analyzes all metrics collected by your Netdata agents, including system metrics, application metrics, and custom collectors.
210 -
211 -**Q: How does it handle sensitive data?**
212 -A: All data is processed securely and discarded after report generation. No data is stored or used for training.
213 -
214 -**Q: Can I customize the report format?**
215 -A: Report structure and visual style are standardized for consistency and professional presentation. However, you have extensive control over the analysis scope, time periods, metrics, and focus areas through customization parameters.
216 -
217 -## What's Next
218 -
219 -AI Insights continues to evolve with new capabilities planned:
220 -
221 -- Scheduled report generation
222 -- Custom report templates
223 -- API access for automation
224 -- Integration with ticketing systems
225 -- Comparative analysis between time periods
226 -
227 -Experience the future of infrastructure monitoring - transform your data into intelligence with AI Insights.
45 +
docs/netdata-ai/insights/anomaly-analysis.md new
+50
@@ -0,0 +1,50 @@
1 +# Anomaly Analysis
2 +
3 +Get a forensics‑grade explanation of unusual behavior. The Anomaly Analysis report correlates ML‑detected anomalies across nodes and metrics, reconstructs the timeline, and proposes likely root causes with supporting evidence.
4 +
5 +![Anomaly Analysis tab](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/anomaly-analysis.png)
6 +
7 +## When to use it
8 +
9 +- Post‑incident analysis and RCA preparation
10 +- Investigating “what changed here?” on a chart or service
11 +- Validating whether anomalies were symptoms or causes
12 +
13 +## How to generate
14 +
15 +1. In Netdata Cloud, open `Insights`
16 +2. Select `Anomaly Analysis`
17 +3. Choose the time window around the event of interest
18 +4. Scope to affected services/nodes if known
19 +5. Click `Generate`
20 +
21 +## What’s analyzed
22 +
23 +- Agent‑side ML anomaly scores (every metric, every second)
24 +- Temporal propagation of anomalies across metrics/services
25 +- Correlations with alerts, deployments, and configuration changes
26 +- Cross‑node relationships and influence chains
27 +
28 +## What you get
29 +
30 +- Narrative of how the incident unfolded
31 +- Ranked list of likely root causes vs. downstream effects
32 +- Key correlated signals and “why this matters” notes
33 +- Recommendations to prevent recurrence
34 +
35 +![Anomaly Analysis report example](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/anomaly-analysis-report.png)
36 +
37 +## Example: “What changed here?”
38 +
39 +Point the report at a suspicious time window and let it reconstruct the change: which metrics shifted first, where anomalies clustered, and which changes correlate strongly with the observed behavior.
40 +
41 +## Related tools
42 +
43 +- Use the `Anomaly Advisor` tab for interactive exploration
44 +- Combine with `Metric Correlations` to focus the search space
45 +
46 +## Availability and usage
47 +
48 +- Available on Business and Free Trial plans
49 +- Each report consumes 1 AI credit (10 free per month on eligible plans)
50 +
docs/netdata-ai/insights/capacity-planning.md new
+52
@@ -0,0 +1,52 @@
1 +# Capacity Planning
2 +
3 +Stop guessing and plan with confidence. The Capacity Planning report projects growth, highlights inflection points, and recommends concrete hardware or configuration changes backed by your actual utilization trends.
4 +
5 +![Capacity Planning tab](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/capacity-planning.png)
6 +
7 +## When to use it
8 +
9 +- Quarterly/annual planning and budgeting cycles
10 +- Preparing procurement requests and vendor discussions
11 +- Evaluating consolidation and right‑sizing opportunities
12 +
13 +## How to generate
14 +
15 +1. Open `Insights` in Netdata Cloud
16 +2. Select `Capacity Planning`
17 +3. Pick a historical window and forecast horizon (3–24 months)
18 +4. Scope to nodes, rooms, or services
19 +5. Click `Generate`
20 +
21 +## What’s analyzed
22 +
23 +- Historical utilization and growth trends (CPU, memory, storage, network)
24 +- Variability, seasonality, and workload patterns
25 +- Anomaly‑adjusted baselines for accurate projections
26 +- Cross‑node comparisons and consolidation candidates
27 +
28 +## What you get
29 +
30 +- Exhaustion date estimates for key resources
31 +- Headroom analysis and risk categorization
32 +- Concrete recommendations (e.g., instance types, disk tiers, scaling)
33 +- Opportunity map for consolidation and cost savings
34 +
35 +![Capacity Planning report example](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/capacity-planning-report.png)
36 +
37 +## Example: Quarterly planning
38 +
39 +Produce a report that justifies next‑quarter spend: show utilization trends, where headroom is tight, when you’ll breach capacity, and specific remediation options with trade‑offs.
40 +
41 +## Best practices
42 +
43 +- Run monthly; compare sequential reports for trend confidence
44 +- Pair with `Performance Optimization` to validate trade‑offs
45 +- Use room‑level scoping to build service‑oriented plans
46 +
47 +## Availability and usage
48 +
49 +- Available on Business and Free Trial plans
50 +- Each report consumes 1 AI credit (10 free per month on eligible plans)
51 +- Reports are saved in Insights and downloadable as PDFs
52 +
docs/netdata-ai/insights/infrastructure-summary.md new
+58
@@ -0,0 +1,58 @@
1 +# Infrastructure Summary
2 +
3 +The Infrastructure Summary report synthesizes the last hours, days, or weeks of your infrastructure into a concise, shareable narrative. It combines critical timelines, anomaly context, alert analysis, and actionable recommendations so your team can quickly align on what happened and what to do next.
4 +
5 +![Infrastructure Summary tab](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/infrastructure-summary.png)
6 +
7 +## When to use it
8 +
9 +- Monday morning recap of weekend incidents and health trends
10 +- Post-incident executive summary for leadership and stakeholders
11 +- Weekly team handoff and situational awareness
12 +- Baseline health before planned infrastructure changes
13 +
14 +## How to generate
15 +
16 +1. Open Netdata Cloud and go to the `Insights` tab
17 +2. Select `Infrastructure Summary`
18 +3. Choose the time range (last 24h, 48h, 7d, or custom)
19 +4. Scope the analysis to all nodes or a subset (rooms/spaces)
20 +5. Click `Generate`
21 +
22 +Reports typically complete in 2–3 minutes. You’ll see them in Insights and receive an email when ready.
23 +
24 +## What’s included in the report
25 +
26 +- Executive summary of the period with key findings
27 +- Incident timeline with affected services and impact
28 +- Alerts overview: frequency, severity, and patterns
29 +- Detected anomalies with confidence and correlations
30 +- Cross-node correlations and dependency highlights
31 +- Notable configuration changes and deploy events (when available)
32 +- Top recommendations with expected impact and rationale
33 +
34 +![Infrastructure Summary report example](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/infrastructure-summary-report.png)
35 +
36 +## Example: Weekend incident recovery
37 +
38 +Generate a 7‑day summary Monday morning to reconstruct what happened while the team was off: which alerts fired, which services were impacted, and where to focus remediation. Use the recommendations section to triage follow-ups.
39 +
40 +## Tips for best results
41 +
42 +- Scope to the most relevant rooms/services when investigating a targeted issue
43 +- Pair with a dedicated `Anomaly Analysis` report for deep dives
44 +- Save summaries as PDFs for sharing with management or compliance
45 +
46 +## Availability and usage
47 +
48 +- Available in Netdata Cloud for Business and Free Trial
49 +- Each generated report consumes 1 AI credit (10 free per month on eligible plans)
50 +- Data privacy: metrics are summarized into structured context; your data is not used to train foundation models
51 +
52 +## See also
53 +
54 +- Performance Optimization
55 +- Capacity Planning
56 +- Anomaly Analysis
57 +- Scheduled Reports
58 +
docs/netdata-ai/insights/performance-optimization.md new
+54
@@ -0,0 +1,54 @@
1 +# Performance Optimization
2 +
3 +Find bottlenecks before users notice. The Performance Optimization report analyzes contention patterns, throttling risks, and systemic inefficiencies, then produces prioritized, concrete remediation steps tied to your observed workload.
4 +
5 +![Performance Optimization tab](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/performance-optimization.png)
6 +
7 +## When to use it
8 +
9 +- Ongoing SRE/ops optimization workstreams
10 +- After key deploys, major configuration changes, or scaling events
11 +- To prepare proposals for performance investments or capacity changes
12 +
13 +## How to generate
14 +
15 +1. Open the `Insights` tab in Netdata Cloud
16 +2. Select `Performance Optimization`
17 +3. Choose a window (e.g., last 24h, 7d, 30d, or custom)
18 +4. Scope to infrastructure segments (rooms/spaces) or services of interest
19 +5. Click `Generate`
20 +
21 +## What’s analyzed
22 +
23 +- CPU and memory saturation, noisy neighbors, and throttling signals
24 +- Disk IO, queue depths, saturation ratios, filesystem pressure
25 +- Network throughput, packet loss, retransmits, egress hot spots
26 +- Container and pod throttling, OOM risks, scheduling pressure
27 +- Database/service bottlenecks and backpressure evidence
28 +
29 +## What you get
30 +
31 +- Ranked list of bottlenecks with severity and confidence
32 +- Correlated signals to distinguish cause vs. symptom
33 +- Specific tuning and right‑sizing recommendations
34 +- Expected impact estimates where feasible (latency/throughput)
35 +- Before/after projections for planned changes (when applicable)
36 +
37 +![Performance Optimization report example](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/performance-optimization-report.png)
38 +
39 +## Example: Debugging Kubernetes performance
40 +
41 +An SRE investigating cluster slowness sees synthesized findings about container throttling, resource contention on specific nodes, and recommended limit/request adjustments—with nodes and workloads called out explicitly.
42 +
43 +## Best practices
44 +
45 +- Run monthly for baselining; run ad‑hoc after notable changes
46 +- Use findings to drive tickets with clear owners and measurable goals
47 +- Combine with `Capacity Planning` for a balanced performance/cost view
48 +
49 +## Availability and usage
50 +
51 +- Available on Business and Free Trial plans
52 +- Each report consumes 1 AI credit (10 free per month on eligible plans)
53 +- Results are saved in Insights and downloadable as PDFs
54 +
docs/netdata-ai/insights/scheduled-reports.md new
+56
@@ -0,0 +1,56 @@
1 +# Scheduled Reports
2 +
3 +Automate your reporting workflow. Scheduled AI reports let you run Insights and Investigations on a recurring cadence and deliver the results automatically—turning manual, repetitive work into a hands‑off process.
4 +
5 +![Schedule dialog 1](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/schedule1.png)
6 +
7 +## What you can schedule
8 +
9 +- Any pre‑built Insight: Infrastructure Summary, Performance Optimization, Capacity Planning, Anomaly Analysis
10 +- Custom Investigations (your own prompts and scope)
11 +
12 +## How to schedule a report
13 +
14 +1. Go to the `Insights` tab in Netdata Cloud
15 +2. Pick an Insight type or click `New Investigation`
16 +3. Configure the time range and scope
17 +4. Click `Schedule` (next to `Generate`)
18 +5. Choose cadence (daily/weekly/monthly) and time
19 +
20 +At the scheduled time, Netdata AI runs the report and delivers it to your email and the Insights tab.
21 +
22 +![Schedule dialog 2](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/schedule2.png)
23 +
24 +## Example setups
25 +
26 +### Weekly infrastructure health
27 +- Type: Infrastructure Summary
28 +- Time range: Last 7 days
29 +- Schedule: Mondays 09:00
30 +
31 +### Monthly performance optimization
32 +- Type: Performance Optimization
33 +- Time range: Last month
34 +- Schedule: 1st of each month 10:00
35 +
36 +### Automated SLO conformance
37 +- Type: New Investigation
38 +- Prompt: Generate SLO conformance for services X and Y with targets …
39 +- Schedule: Mondays 10:00
40 +
41 +## Managing schedules
42 +
43 +- View, pause, or edit schedules from the Insights tab
44 +- Scheduled runs consume AI credits when they execute
45 +
46 +## Availability and usage
47 +
48 +- Available to Business and Free Trial plans
49 +- Each scheduled run consumes 1 AI credit (10 free/month on eligible plans)
50 +
51 +## Tips
52 +
53 +- Start with weekly summaries to establish a baseline
54 +- Schedule targeted reports for critical services or high‑cost areas
55 +- Use schedules to feed regular Slack/email updates and leadership briefs
56 +
docs/netdata-ai/investigations/custom-investigations.md new
+87
@@ -0,0 +1,87 @@
1 +# Custom Investigations
2 +
3 +Create deeply researched, context‑aware analyses by asking Netdata open‑ended questions about your infrastructure. Custom Investigations correlate metrics, anomalies, and events to answer the questions dashboards can’t—typically in about two minutes.
4 +
5 +![Custom Investigation creation](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/custom-investigation.png)
6 +
7 +## When to use Custom Investigations
8 +
9 +- Troubleshoot complex incidents by delegating parallel investigations
10 +- Analyze deployment or configuration change impact (before/after)
11 +- Optimize performance and cost (identify underutilization and hotspots)
12 +- Explore longer‑term behavioral changes and trends
13 +
14 +## Start an investigation
15 +
16 +Two ways to launch:
17 +
18 +- From anywhere: Click `Troubleshoot with AI` (top‑right). The current view’s scope (chart, dashboard, room, service) is captured automatically; add your question and context.
19 +- From Insights: Go to `Insights` → `New Investigation` for a blank canvas and full control.
20 +
21 +Reports are saved in Insights and you’ll receive an email when ready.
22 +
23 +## Provide good context (get great results)
24 +
25 +Think of this as briefing a teammate. Include time ranges, environments, related services, symptoms, and recent changes.
26 +
27 +### Example 1: Troubleshooting a problem
28 +Request: Why are my checkout‑service pods crashing repeatedly?
29 +
30 +Context:
31 +```
32 +- Started after: deployment at 14:00 UTC of version 2.3.1
33 +- Impact: Customer checkout failures, lost revenue ~$X/hour
34 +- Recent changes: payment gateway integration update; workers 10→20
35 +- Logs: "connection refused to payment-service:8080", "Java heap space"
36 +- Environment: production / eks-prod-us-east-1
37 +- Related: payment-service, inventory-service, redis-session-store
38 +```
39 +
40 +### Example 2: Analyze a change
41 +Request: Compare system metrics before and after the user‑authentication‑service deployment.
42 +
43 +Context:
44 +```
45 +- Service: user-authentication-service v2.2.0
46 +- Deployed: 2025‑01‑24 09:00 UTC
47 +- Changes: JWT→Redis sessions; Argon2 hashing
48 +- Concern: intermittent logouts; rising redis_connected_clients
49 +- Windows: 24h before vs 24h after
50 +```
51 +
52 +### Example 3: Cost optimization
53 +Request: Identify underutilized nodes for cost optimization.
54 +
55 +Context:
56 +```
57 +- Monthly compute: ~$12K
58 +- Mixed workloads (prod + staging)
59 +- Dev envs run 24/7; batch nodes idle 20h/day
60 +- Goal: save $2–3K/month without reliability impact
61 +```
62 +
63 +## Best practices
64 +
65 +1. Be specific: timeframe, environment, services
66 +2. Add helpful context from tickets/Slack/deploy logs
67 +3. Set clear goals (reduce costs, find root cause, etc.)
68 +4. Run multiple investigations in parallel during incidents
69 +
70 +![Custom Investigation report example](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/custom-investigation-report.png)
71 +
72 +## Scheduling
73 +
74 +Automate recurring investigations (weekly health, monthly optimization, SLO conformance) from the `Insights` tab. See `Scheduled Investigations` for examples and setup.
75 +
76 +## Availability and credits
77 +
78 +- Generally available in Netdata Cloud (Business and Free Trial)
79 +- Eligible Spaces receive 10 free AI runs per month; additional usage via AI Credits
80 +- Track usage in `Settings → Usage & Billing → AI Credits`
81 +
82 +## Related
83 +
84 +- `Investigations` overview
85 +- `Scheduled Investigations`
86 +- `Alert Troubleshooting`
87 +
docs/netdata-ai/investigations/index.md new
+70
@@ -0,0 +1,70 @@
1 +# Investigations
2 +
3 +Ask Netdata anything about your infrastructure and get a deeply researched answer in minutes. Investigations turn your question and context into an analysis that correlates metrics, anomalies, and events across your systems.
4 +
5 +## What Investigations are good for
6 +
7 +- Troubleshooting live incidents without manual data wrangling
8 +- Analyzing the impact of deployments or config changes
9 +- Cost and efficiency reviews (identify underutilized resources)
10 +- Exploring longer‑term behavioral changes and trends
11 +
12 +## Starting an investigation
13 +
14 +Two easy entry points:
15 +
16 +- `Troubleshoot with AI` button (top‑right): Captures the current chart, dashboard, or service context automatically, then you add your question
17 +- `Insights` → `New Investigation`: Blank canvas for any custom prompt
18 +
19 +Reports complete in ~2 minutes and are saved in Insights; you’ll get an email when ready.
20 +
21 +## Provide good context (get great results)
22 +
23 +Think of it like briefing a teammate. Include timeframes, environments, related services, symptoms, and recent changes. Example formats:
24 +
25 +### Example: Troubleshoot a problem
26 +Request: Why are my checkout‑service pods crashing repeatedly?
27 +
28 +Context:
29 +```
30 +- Started after: deployment at 14:00 UTC of version 2.3.1
31 +- Impact: Customer checkout failures, lost revenue ~$X/hour
32 +- Recent changes: payment gateway integration update; workers 10→20
33 +- Logs: "connection refused to payment-service:8080", "Java heap space"
34 +- Environment: production / eks-prod-us-east-1
35 +- Related: payment-service, inventory-service, redis-session-store
36 +```
37 +
38 +### Example: Analyze a change
39 +Request: Compare metrics before/after the user‑authentication‑service deploy.
40 +
41 +Context:
42 +```
43 +- Service: user-authentication-service v2.2.0
44 +- Deployed: 2025‑01‑24 09:00 UTC
45 +- Changes: JWT→Redis sessions; Argon2 hashing added
46 +- Concern: intermittent logouts; rising redis_connected_clients
47 +- Windows: 24h before vs 24h after
48 +```
49 +
50 +### Example: Cost optimization
51 +Request: Identify underutilized nodes for cost savings.
52 +
53 +Context:
54 +```
55 +- Monthly compute: ~$12K
56 +- Mixed workloads (prod + staging)
57 +- Dev envs run 24/7; batch nodes idle 20h/day
58 +- Goal: save $2–3K/month without reliability impact
59 +```
60 +
61 +## Availability and credits
62 +
63 +- Available to Business and Free Trial plans
64 +- Each run consumes 1 AI credit (10 free per month on eligible plans)
65 +
66 +## Related documentation
67 +
68 +- [Custom Investigations](/docs/netdata-ai/investigations/custom-investigations.md)
69 +- [Scheduled Investigations](/docs/netdata-ai/investigations/scheduled-investigations.md)
70 +- [Alert Troubleshooting](/docs/troubleshooting/troubleshoot.md)
docs/netdata-ai/investigations/scheduled-investigations.md new
+51
@@ -0,0 +1,51 @@
1 +# Scheduled Investigations
2 +
3 +Automate recurring custom analyses by scheduling your own investigation prompts. Great for weekly health checks, monthly cost reviews, and SLO conformance reporting.
4 +
5 +![Schedule dialog 1](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/schedule1.png)
6 +
7 +## How to schedule
8 +
9 +1. Go to the `Insights` tab → `New Investigation`
10 +2. Enter your prompt and set scope/time window
11 +3. Click `Schedule` and choose cadence (daily/weekly/monthly)
12 +4. Confirm recipients (email) and save
13 +
14 +At the scheduled time, Netdata AI runs the investigation and delivers the report to your email and the Insights tab.
15 +
16 +![Schedule dialog 2](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/schedule2.png)
17 +
18 +## Examples
19 +
20 +### Weekly health check
21 +Prompt:
22 +```
23 +Generate a weekly infrastructure summary for services A, B, C. Include major incidents,
24 +anomalies, capacity risks, and recommended follow‑ups.
25 +```
26 +
27 +### Monthly optimization review
28 +Prompt:
29 +```
30 +Analyze performance regressions and right‑sizing opportunities over the past month for
31 +our Kubernetes workloads in room X. Prioritize actions by potential impact.
32 +```
33 +
34 +### SLO conformance
35 +Prompt:
36 +```
37 +Generate an SLO conformance report for 'user-auth' (99.9% uptime, p95 latency <200ms)
38 +and 'payment-processing' (99.99% uptime, p95 <500ms) for the last 7 days. Include
39 +breaches, contributing factors, and remediation recommendations.
40 +```
41 +
42 +## Manage schedules
43 +
44 +- Edit, pause, or delete schedules from the Insights tab
45 +- Scheduled runs consume AI credits when they execute
46 +
47 +## Availability and credits
48 +
49 +- Available on Business and Free Trial plans
50 +- 10 free AI runs/month on eligible Spaces; additional usage via AI Credits
51 +
docs/netdata-ai/troubleshooting/index.md new
+27
@@ -0,0 +1,27 @@
1 +# Troubleshooting
2 +
3 +Netdata AI accelerates troubleshooting with three complementary tools:
4 +
5 +- Alert Troubleshooting: one‑click analysis from any alert
6 +- Anomaly Advisor: interactive, ML‑driven incident investigation
7 +- Metric Correlations: quickly focus on relevant charts for a time window
8 +
9 +Use Alert Troubleshooting to start from an alert with an automated baseline. Pivot to Anomaly Advisor for propagation analysis and to Metric Correlations to narrow the search space across charts.
10 +
11 +## Alert Troubleshooting
12 +
13 +Generate a report that assesses alert validity, uncovers correlated signals, and proposes a root‑cause hypothesis with supporting evidence. Start from the Alerts tab (`Ask AI`), Insights (`Alert Troubleshooting`), or the link in alert emails.
14 +
15 +## Anomaly Advisor
16 +
17 +Explore incident timelines visually and see how anomalies cascade across your infrastructure. Start from the Anomalies tab in Netdata Cloud.
18 +
19 +## Metric Correlations
20 +
21 +From any dashboard or time window, surface the charts most related to your selection to speed root cause analysis.
22 +
23 +## See also
24 +
25 +- Troubleshoot Button (how to trigger analysis from anywhere)
26 +- Investigations (ask open‑ended questions with rich context)
27 +
docs/netdata-ai/troubleshooting/troubleshoot-button.md new
+41
@@ -0,0 +1,41 @@
1 +# Troubleshoot with AI Button
2 +
3 +Trigger an AI‑powered investigation from anywhere in Netdata Cloud. The `Troubleshoot with AI` button captures your current context (chart, dashboard, room, or service) and launches an investigation with that scope pre‑selected.
4 +
5 +![Troubleshoot with AI button](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/troubleshoot-button.png)
6 +
7 +## Where to find it
8 +
9 +- Alerts tab: `Ask AI` next to any alert
10 +- Insights tab: `Alert Troubleshooting` and `New Investigation`
11 +- Top‑right of most views: `Troubleshoot with AI`
12 +- Alert emails: `Troubleshoot with AI` link
13 +
14 +## How it works
15 +
16 +1. Click `Troubleshoot with AI`
17 +2. Review the captured scope and time window
18 +3. Add your question and any extra context (symptoms, recent changes)
19 +4. Start the investigation
20 +
21 +Within ~2 minutes, you’ll receive a report with:
22 +
23 +- Summary of findings and likely root cause
24 +- Correlated metrics/logs across affected systems
25 +- Suggested next steps with rationale
26 +
27 +## Tips for better results
28 +
29 +- Be explicit about timeframe, environment, and related services
30 +- Paste relevant notes from tickets/Slack/deploy logs
31 +- Run multiple investigations in parallel during incidents
32 +
33 +## Availability and credits
34 +
35 +- Available on Business and Free Trial plans
36 +- Each run consumes 1 AI credit (10 free per month on eligible plans)
37 +
38 +## Privacy
39 +
40 +Your infrastructure data is summarized to a compact context for analysis and is not used to train foundation models.
41 +
docs/troubleshooting/custom-investigations.md
+7 -10
@@ -104,16 +104,13 @@ Click the **"Troubleshoot with AI"** button in the top right corner from any scr
104
105 ### Access and Availability
106
107 -This feature is available in preview mode for:
107 +- Generally available in Netdata Cloud (Business and Free Trial)
108 +- Eligible Spaces receive 10 free AI runs per month; additional usage via AI Credits
109
109 -- All Business and Homelab plan users
110 -- New users get 10 AI investigation sessions per month during their Business plan trial
111 -- Community users can request access by contacting product@netdata.cloud
110 +:::note
111 +Track AI credit usage from `Settings → Usage & Billing → AI Credits`.
112 +:::
113
113 -### Coming Soon
114 +### Scheduling
115
115 -We're actively developing:
116 -
117 -- Scheduled recurring investigations for regular reports
118 -- Custom SLO report templates
119 -- Weekly cost-optimization analyses
116 +You can schedule recurring investigations from the `Insights` tab (daily/weekly/monthly). Use this to automate weekly health checks, monthly optimization reviews, or SLO conformance reports.
docs/troubleshooting/troubleshoot.md
+7 -7
@@ -4,6 +4,8 @@
4
5 When an alert fires, you can use AI to generate a detailed troubleshooting report that analyzes whether the alert requires immediate action or is just noise. The AI examines your alert's history, correlates it with thousands of other metrics across your infrastructure, and provides actionable insights—all within minutes.
6
7 +![Ask AI from Alerts](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/alert-troubleshoot-1.png)
8 +
9 ### Key Benefits
10
11 - **Save hours of manual investigation** - Skip the initial data collection and correlation work
@@ -63,15 +65,13 @@ Reports typically generate in 1-2 minutes. Once complete:
65 - A copy is saved in the **Insights** tab under "Investigations"
66 - You receive an email notification with the analysis summary
67
66 -### Access and Availability
68 +![Alert Troubleshooting report example](https://raw.githubusercontent.com/netdata/docs-images/refs/heads/master/netdata-cloud/netdata-ai/alert-troubleshoot-report.png)
69
68 -This feature is available in preview mode for:
70 +### Access and Availability
71
70 -- All Business and Homelab plan users
71 -- New users get 10 AI troubleshooting sessions per month during their Business plan trial
72 +- Generally available in Netdata Cloud (Business and Free Trial)
73 +- Eligible Spaces receive 10 free AI runs per month; additional usage via AI Credits
74
75 :::note
74 -
75 -Community users can request access by contacting product@netdata.cloud
76 -
76 +Track AI credit usage from `Settings → Usage & Billing → AI Credits`.
77 :::