@cryptotaxi247 / netdata / commits / 1df7ff60c

transfer Learn PR 2473 (#20600)

transfer https://github.com/netdata/learn/pull/2473 with necessary fixes to work

Fotis Voutsas committed Jul 1, 2025 at 13:10 UTC 1df7ff60c6705534235b0cfb8f0cf907f781a0fd
15 files changed +2392 -324
docs/category-overview-pages/machine-learning-and-assisted-troubleshooting.md
+185 -36
@@ -1,54 +1,203 @@
1 -# Netdata AI and Machine Learning
1 +# AI and Machine Learning
2
3 -Boost your monitoring and troubleshooting capabilities with Netdata's AI-powered features.
3 +Netdata provides powerful AI-driven capabilities to transform how you monitor and troubleshoot your infrastructure, with more innovations coming soon.
4
5 -Netdata AI helps you **detect anomalies, understand metric relationships, and resolve issues quickly** with intelligent assistance all designed to make your infrastructure management smarter, faster, and bulletproof.
5 +## What's Available Today
6
7 -## What Can Netdata AI Do For You?
7 +### 1. AI Chat with Netdata
8
9 -Netdata AI combines powerful machine learning capabilities with intuitive interfaces to help you:
9 +**Available Now** - Chat with your infrastructure using natural language
10
11 -1. **Detect anomalies automatically** before they escalate into critical issues
12 -2. **Understand relationships** between metrics during troubleshooting
13 -3. **Get expert guidance** when resolving alerts and performance problems
11 +Ask questions about your infrastructure like you're talking to a colleague. Get instant answers about performance, find specific logs, identify top resource consumers, or investigate issues - all through simple conversation. No more complex queries or dashboard hunting.
12
15 -## Machine Learning and Anomaly Detection
13 +**Key capabilities**:
14
17 -Our ML-powered anomaly detection works silently in the background, monitoring your metrics and identifying unusual patterns.
15 +- **Natural language queries** - "Which servers have high CPU usage?" or "Show database errors from last hour" or "What is wrong with my infrastructure now", or "Do a post-mortem analysis of the outage we had yesteday", or "Show me all network dependencies of process X"
16 +- **Multi-node visibility** - Analyzes your entire infrastructure through Netdata Parents
17 +- **Flexible AI options** - Use your existing AI tools or our standalone web chat
18
19 -| Feature | What It Does For You |
20 -|----------------------------------|------------------------------------------------------------------------------|
21 -| **Unsupervised Learning** | Works automatically without requiring manual training or labeling of data |
22 -| **Multiple Model Consensus** | Reduces false positives by 99% by requiring agreement across multiple models |
23 -| **Real-time Anomaly Bits** | Flags unusual metrics instantly, with zero storage overhead |
24 -| **Anomaly Rate Visualization** | Highlights anomalous time periods in your dashboard for quick investigation |
25 -| **Node-Level Anomaly Detection** | Identifies when your entire system is behaving unusually |
26 -| **Metric Correlations** | Helps you find relationships between metrics to pinpoint root causes |
19 +<details>
20 +<summary><strong>How it works</strong></summary>
21
28 -Learn more in the [Machine Learning and Anomaly Detection](/src/ml/README.md) documentation.
22 +- **MCP integration** - You chat with an LLM, that has access to your observability data, via Model Context Protocol (MCP)
23 +- **Choice of AI providers** - Claude, GPT-4, Gemini, and others
24 +- **Two deployment options** - Use an existing AI client that supports MCP, or use a web page chat we created for it (LLM is pay-per-use with API keys)
25 +- **Real-time data access** - Query live metrics, logs, processes, network connections, and system state
26 +- **Secure connection** - LLM has access to your data via the LLM client
27
30 -## Netdata Assistant
28 +</details>
29
32 -When alerts trigger or anomalies emerge, Netdata Assistant serves as your AI-powered troubleshooting companion.
30 +**Access**: Available now for all Netdata Agent deployments (Standalone and Parents)
31
34 -| Feature | What It Does For You |
35 -|----------------------------|-----------------------------------------------------------------------|
36 -| **Alert Context** | Explains what each alert means and why you should care about it |
37 -| **Guided Troubleshooting** | Offers step-by-step instructions tailored to your specific situation |
38 -| **Persistent Window** | Follows you throughout your dashboards as you investigate issues |
39 -| **Curated Resources** | Provides links to relevant documentation to deepen your understanding |
40 -| **Time-Saving** | Eliminates the need for searching documentation or online forums |
32 +[Explore AI Chat →](./chat-with-netdata-mcp)
33
42 -Learn more about [Netdata Assistant](/docs/netdata-assistant.md) and how it helps streamline your troubleshooting workflow.
34 +---
35
44 -## Getting Started
36 +### 2. AI DevOps Copilot
37
46 -Netdata AI features are enabled by default with the standard installation. The machine learning capabilities require the `dbengine` database mode, which is the default setting.
38 +**Available Now** - Transform observability into action with CLI AI assistants
39
48 -To start exploring:
40 +Combine the power of AI with system automation. CLI-based AI assistants like Claude Code and Gemini CLI can access your Netdata metrics and execute commands, enabling intelligent infrastructure optimization, automated troubleshooting, and configuration management - all driven by real observability data.
41
50 -1. **Anomaly Detection**: Check the [Anomaly Advisor tab](/docs/dashboards-and-charts/anomaly-advisor-tab.md) to see detected anomalies
51 -2. **Metric Correlations**: Use the Metric Correlations button in the dashboard to analyze relationships between metrics
52 -3. **Netdata Assistant**: Click the Assistant button in the Alerts tab when troubleshooting alerts
42 +**Key capabilities**:
43
54 -These AI features work seamlessly with Netdata's other capabilities, enhancing your overall monitoring and troubleshooting experience without requiring any AI expertise.
44 +- **Observability-driven automation** - AI analyzes metrics and executes fixes
45 +- **Infrastructure optimization** - Automatic tuning based on performance data
46 +- **Intelligent troubleshooting** - From problem detection to resolution
47 +- **Configuration management** - AI-generated configs based on actual usage
48 +
49 +<details>
50 +<summary><strong>How it works</strong></summary>
51 +
52 +- **MCP-enabled CLI tools** - Claude Code, Gemini CLI, and others
53 +- **Bidirectional integration** - Read metrics, execute commands
54 +- **Context-aware decisions** - AI understands your infrastructure state
55 +- **Safe execution** - Review AI suggestions before implementation
56 +- **Team collaboration** - Share configurations via version control
57 +
58 +</details>
59 +
60 +**Access**: Available now with MCP-supported CLI AI tools
61 +
62 +[Explore AI DevOps Copilot →](./ai-devops-copilot/ai-devops-copilot)
63 +
64 +---
65 +
66 +### 3. AI Insights
67 +
68 +**Preview (Netdata Cloud Feature)** - Strategic infrastructure analysis in minutes
69 +
70 +Transform past data into actionable insights with AI-generated reports. Perfect for capacity planning, performance reviews, and executive briefings. Get comprehensive analysis of your infrastructure trends, optimization opportunities, and future requirements - all in professionally formatted PDFs.
71 +
72 +**Four report types**:
73 +
74 +- **Infrastructure Summary** - Complete system health and incident analysis
75 +- **Capacity Planning** - Growth projections and resource recommendations
76 +- **Performance Optimization** - Bottleneck identification and tuning suggestions
77 +- **Anomaly Analysis** - Deep dive into unusual patterns and their impacts
78 +
79 +<details>
80 +<summary><strong>How it works</strong></summary>
81 +
82 +- **2-3 minute generation** - Analyzes historical data comprehensively
83 +- **PDF downloads** - Professional reports ready for sharing
84 +- **Embedded visualizations** - Charts and graphs from your actual data
85 +- **Executive-ready** - Clear summaries with technical details included
86 +- **Secure processing** - Data analyzed then immediately discarded
87 +
88 +</details>
89 +
90 +**Access**:
91 +
92 +- Business subscriptions: Unlimited reports
93 +- Free trial users: Full access during trial
94 +- Community users: 10 free reports ([request early access](https://discord.gg/mPZ6WZKKG2))
95 +
96 +[Explore AI Reports →](./ai-insights)
97 +
98 +---
99 +
100 +### 4. Anomaly Advisor
101 +
102 +**Available to All** - Revolutionary troubleshooting that finds root causes in minutes
103 +
104 +Stop guessing what went wrong. The Anomaly Advisor instantly shows you how problems cascade across your infrastructure and ranks every metric by anomaly severity. Root causes typically appear in the top 20-30 results, turning hours of investigation into minutes of discovery.
105 +
106 +**Revolutionary approach**:
107 +
108 +- **See cascading effects** - Watch anomalies propagate across systems
109 +- **Automatic ranking** - Every metric scored and sorted by anomaly severity
110 +- **No expertise required** - Works even on unfamiliar systems
111 +
112 +<details>
113 +<summary><strong>How it works</strong></summary>
114 +
115 +- **Data-driven analysis** - No hypotheses needed, the data reveals the story
116 +- **Influence tracking** - Shows what influenced and what was influenced
117 +- **Time window analysis** - Highlight any incident period for investigation
118 +- **Scale-agnostic** - Works identically from 10 to 10,000 nodes
119 +- **Visual propagation** - See anomaly clusters and cascades instantly
120 +
121 +</details>
122 +
123 +**Find it**: Anomalies tab in any Netdata dashboard
124 +
125 +[Learn more about Anomaly Advisor →](./anomaly-advisor)
126 +
127 +---
128 +
129 +### 5. Machine Learning Anomaly Detection
130 +
131 +**Available to All** - Continuous anomaly detection on every metric
132 +
133 +The foundation of Netdata's AI capabilities. Machine learning models run locally on every agent, continuously learning normal patterns and detecting anomalies in real-time. Zero configuration required - it just works, protecting your infrastructure 24/7.
134 +
135 +**Automatic protection**:
136 +
137 +- **Every metric monitored** - ML analyzes all metrics continuously
138 +- **Visual anomaly indicators** - Purple ribbons on every chart show anomaly rates
139 +- **Historical anomaly data** - ML scores saved with metrics for past analysis
140 +- **Zero configuration** - Starts working immediately after installation
141 +
142 +<details>
143 +<summary><strong>How it works</strong></summary>
144 +
145 +- **Local ML engine** - Runs on every Netdata Agent, no cloud dependency
146 +- **Multiple models** - Consensus approach reduces noise and false positives by 99%
147 +- **Integrated storage** - Anomaly scores saved in the database with metrics
148 +- **Historical queries** - Query past anomaly rates just like any other metric
149 +- **Visual integration** - Purple anomaly ribbons appear on all charts automatically
150 +- **Minimal overhead** - Designed for production environments
151 +- **Privacy by design** - Your data never leaves your infrastructure
152 +
153 +</details>
154 +
155 +**Access**: Free for everyone - enabled by default
156 +
157 +[Explore Machine Learning →](./machine-learning-anomaly-detection)
158 +
159 +## Coming Soon
160 +
161 +### AI Chat with Netdata (Netdata Cloud version)
162 +
163 +**In Development** - Chat with your entire infrastructure through Netdata Cloud
164 +
165 +Soon, Netdata Cloud will become an MCP server itself. This means you'll be able to chat with your entire infrastructure without setting up local MCP bridges. Get the same natural language capabilities with the added benefits of Cloud's global view, team collaboration, and seamless access from anywhere.
166 +
167 +**What to expect**:
168 +
169 +- Direct MCP integration with Netdata Cloud
170 +- Chat with all your infrastructure from one place
171 +- No local bridge setup required
172 +- Team collaboration on AI conversations
173 +- Access from any device, anywhere
174 +
175 +---
176 +
177 +### AI Alert Assistant
178 +
179 +**In Development (Netdata Cloud)** - Real-time AI help when alerts fire
180 +
181 +When critical alerts trigger, you'll get instant AI analysis of the situation. The Alert Assistant will identify likely root causes, provide specific troubleshooting steps, and guide you through resolution - all tailored to your specific infrastructure and alert context.
182 +
183 +**What to expect**:
184 +
185 +- Automatic activation when viewing alerts
186 +- Root cause analysis with confidence scores
187 +- Step-by-step remediation guidance
188 +- Integration with your runbooks and procedures
189 +
190 +---
191 +
192 +### AI Weekly Digest
193 +
194 +**In Development (Netdata Cloud)** - Your infrastructure insights delivered weekly
195 +
196 +Stay informed without information overload. The AI Weekly Digest will analyze your infrastructure's performance over the past week and deliver a concise summary of what matters most - trends, issues resolved, optimization opportunities, and what to watch next week.
197 +
198 +**What to expect**:
199 +
200 +- Weekly email summaries customized for your role
201 +- Key metrics and trend analysis
202 +- Proactive recommendations for the week ahead
203 +- Highlights of resolved and ongoing issues
docs/learn/mcp.md
+196 -92
@@ -1,140 +1,244 @@
1 -# Netdata MCP Integration: AI-Powered Infrastructure Analysis
1 +# Netdata MCP
2
3 -:::info
3 +All Netdata Agents and Parents are Model Context Protocol (MCP) servers, enabling AI assistants to interact with your infrastructure monitoring data.
4
5 -The Netdata MCP Server preview is live. [Get early access](https://b6yi53u6qjm.typeform.com/to/DQi5ibhE?typeform-source=www.netdata.cloud) or visit our GitHub repository for the latest nd-mcp tools and setup instructions.
5 +Every Netdata Agent and Parent includes an MCP server that:
6
7 -:::
7 +- Implements the protocol as WebSocket for transport
8 +- Provides read-only access to metrics, logs, alerts, and live system information
9 +- Requires no additional installation - it's part of Netdata
10
9 -:::note
11 +## Visibility Scope
12
11 -This integration leverages new and evolving AI technologies. While Netdata provides comprehensive infrastructure monitoring capabilities, the AI analysis features **depend on external AI services** and their inherent limitations. The quality and accuracy of AI-generated insights are **subject to the capabilities and constraints of the underlying AI models, not Netdata's monitoring functionality**.
13 +Netdata provides comprehensive access to all available observability data through MCP, including complete metadata:
14
13 -:::
15 +- **Node Discovery** - Hardware specifications, operating system details, version information, streaming topology, and associated metadata
16 +- **Metrics Discovery** - Full-text search capabilities across contexts, instances, dimensions, and labels
17 +- **Function Discovery** - Access to system functions including `processes`, `network-connections`, `streaming`, `systemd-journal`, `windows-events`, etc.
18 +- **Alert Discovery** - Real-time visibility into active and raised alerts
19 +- **Metrics Queries** - Complex aggregations and groupings with ML-powered anomaly detection
20 +- **Metrics Scoring** - Root cause analysis leveraging anomaly detection and metric correlations
21 +- **Alert History** - Complete alert transition logs and state changes
22 +- **Function Execution** - Execute Netdata functions on any connected node (requires Netdata Parent)
23 +- **Log Exploration** - Access logs from any connected node (requires Netdata Parent)
24
15 -## What is MCP?
25 +For sensitive features currently protected by Netdata Cloud SSO, a temporary MCP API key is generated on each Netdata instance. When included in the MCP connection string, this key unlocks access to sensitive data and protected functions (like `systemd-journal`, `windows-events` and `processes`). This temporary API key mechanism will eventually be replaced with a new authentication system integrated with Netdata Cloud.
26
17 -**Model Context Protocol (MCP)** is a new open standard that allows AI assistants to connect directly to your data sources and tools. Think of it as a bridge that lets AI systems access and analyze your real-time infrastructure data instead of just providing generic advice.
27 +AI assistants have different visibility depending on where they connect:
28
19 -:::note
29 +- **Netdata Cloud**: (coming soon) Full visibility across all nodes in your infrastructure
30 +- **Netdata Parent Node**: Visibility across all child nodes connected to that parent
31 +- **Netdata Child/Standalone Node**: Visibility only into that specific node
32
21 -**Learn More**: MCP is an open standard created by Anthropic. For technical details about the protocol specification, visit [Anthropic's MCP documentation](https://modelcontextprotocol.io/).
33 +## Finding the nd-mcp Bridge
34
23 -:::
35 +AI clients like Claude Desktop run locally on your computer and use `stdio` communication. Since your Netdata runs remotely on a server, you need a bridge to convert `stdio` to WebSocket communication.
36
25 -## How Netdata is Pioneering AI-Powered Monitoring
37 +The `nd-mcp` bridge needs to be available on your desktop or laptop where your AI client runs. Since most users run Netdata on remote servers rather than their local machines, you have two options:
38
27 -Netdata is **one of the first monitoring platforms** to integrate MCP, placing us among the pioneers in AI-powered infrastructure analysis. We've built a direct connection between advanced AI models (like Claude) and your Netdata monitoring data, creating an intelligent troubleshooting partner that understands your specific infrastructure.
39 +1. **If you have Netdata installed locally** - Use the existing nd-mcp
40 +2. **If Netdata is only on remote servers** - Build nd-mcp on your desktop/laptop
41
29 -### What Makes This Revolutionary
42 +### Option 1: Using Existing nd-mcp
43
31 -**Traditional monitoring**: You look at charts and graphs, manually correlating data across different services to find problems.
44 +If you have Netdata installed on your desktop/laptop, find the existing bridge:
45
33 -**Netdata with MCP**: AI analyzes your actual monitoring data in real-time, automatically correlates anomalies across your entire infrastructure, and provides expert-level post-mortem analysis.
46 +#### Linux
47
35 -## What This Changes for You
48 +```bash
49 +# Try these locations in order:
50 +which nd-mcp
51 +ls -la /usr/sbin/nd-mcp
52 +ls -la /usr/bin/nd-mcp
53 +ls -la /opt/netdata/usr/bin/nd-mcp
54 +ls -la /usr/local/bin/nd-mcp
55 +ls -la /usr/local/netdata/usr/bin/nd-mcp
56
37 -### Instant Expert Analysis
57 +# Or search for it:
58 +find / -name "nd-mcp" 2>/dev/null
59 +```
60 +
61 +Common locations:
62 +
63 +- **Native packages (apt, yum, etc.)**: `/usr/sbin/nd-mcp` or `/usr/bin/nd-mcp`
64 +- **Static installations**: `/opt/netdata/usr/bin/nd-mcp`
65 +- **Built from source**: `/usr/local/netdata/usr/bin/nd-mcp`
66 +
67 +#### macOS
68 +
69 +```bash
70 +# Try these locations:
71 +which nd-mcp
72 +ls -la /usr/local/bin/nd-mcp
73 +ls -la /usr/local/netdata/usr/bin/nd-mcp
74 +ls -la /opt/homebrew/bin/nd-mcp
75 +
76 +# Or search for it:
77 +find / -name "nd-mcp" 2>/dev/null
78 +```
79 +
80 +#### Windows
81 +
82 +```powershell
83 +# Check common locations:
84 +dir "C:\Program Files\Netdata\usr\bin\nd-mcp.exe"
85 +dir "C:\Netdata\usr\bin\nd-mcp.exe"
86 +# Or search for it:
87 +where nd-mcp.exe
88 +```
89 +
90 +### Option 2: Building nd-mcp for Your Desktop
91 +
92 +If you don't have Netdata installed loca you can build just the nd-mcp bridge. Netdata provides three implementations - choose the one that best fits your environment:
93 +
94 +1. **Go bridge** (recommended) - [Go bridge source code](https://github.com/netdata/netdata/tree/master/src/web/mcp/bridges/stdio-golang)
95 + - Produces a single binary with no dependencies
96 + - Creates executable named `nd-mcp` (`nd-mcp.exe` on windows)
97 + - Includes both `build.sh` and `build.bat` (for Windows)
98
39 -Instead of spending hours analyzing charts during an incident, you can now:
99 +2. **Node.js bridge** - [Node.js bridge source code](https://github.com/netdata/netdata/tree/master/src/web/mcp/bridges/stdio-nodejs)
100 + - Good if you already have Node.js installed
101 + - Creates script named `nd-mcp.js`
102 + - Includes `build.sh`
103
41 -- **Ask natural questions**: "What caused the outage between 13:00-13:30 UTC?"
42 -- **Get comprehensive post-mortems**: AI analyzes your metrics, identifies root causes, and maps dependency chains
43 -- **Understand complex correlations**: AI detects which services were affected and why
104 +3. **Python bridge** - [Python bridge source code](https://github.com/netdata/netdata/tree/master/src/web/mcp/bridges/stdio-python)
105 + - Good if you already have Python installed
106 + - Creates script named `nd-mcp.py`
107 + - Includes `build.sh`
108
45 -### Real-World Example
109 +To build:
110
47 -**The Problem**: Database server goes down, affecting multiple applications across your infrastructure.
111 +```bash
112 +# Clone the Netdata repository
113 +git clone https://github.com/netdata/netdata.git
114 +cd netdata
115
49 -```mermaid
50 -flowchart TD
51 - A("🚨 Database Outage Occurs")
52 -
53 - B("📊 Traditional Approach")
54 - C("🤖 Netdata + MCP")
55 -
56 - D("Manual dashboard review<br/>⏱️ Hours of analysis")
57 - E("AI analyzes all metrics<br/>⚡ Instant correlation")
58 -
59 - F("❌ Slow resolution<br/>Missed dependencies")
60 - G("✅ Fast resolution<br/>Complete root cause")
61 -
62 - A --> B
63 - A --> C
64 - B --> D
65 - C --> E
66 - D --> F
67 - E --> G
68 -
69 - %% Styling
70 - classDef problem fill:#ffeb3b,stroke:#333,stroke-width:2px
71 - classDef traditional fill:#ff7043,stroke:#333,stroke-width:2px
72 - classDef ai fill:#42a5f5,stroke:#333,stroke-width:2px
73 - classDef outcome fill:#f5f5f5,stroke:#333,stroke-width:2px
74 -
75 - class A problem
76 - class B,D traditional
77 - class C,E ai
78 - class F,G outcome
116 +# Choose your preferred implementation
117 +cd src/web/mcp/bridges/stdio-golang/ # or stdio-nodejs/ or stdio-python/
118 +
119 +# Build the bridge
120 +./build.sh # On Windows with the Go version, use build.bat
121 +
122 +# The executable will be created with different names:
123 +# - Go: nd-mcp
124 +# - Node.js: nd-mcp.js
125 +# - Python: nd-mcp.py
126 +
127 +# Test the bridge with your Netdata instance (replace localhost with your Netdata IP)
128 +./nd-mcp ws://localhost:19999/mcp # Go bridge
129 +./nd-mcp.js ws://localhost:19999/mcp # Node.js bridge
130 +./nd-mcp.py ws://localhost:19999/mcp # Python bridge
131 +
132 +# You should see:
133 +# nd-mcp: Connecting to ws://localhost:19999/mcp...
134 +# nd-mcp: Connected
135 +# Press Ctrl+C to stop the test
136 +
137 +# Get the absolute path for your AI client configuration
138 +pwd # Shows current directory
139 +# Example output: /home/user/netdata/src/web/mcp/bridges/stdio-golang
140 +# Your nd-mcp path would be: /home/user/netdata/src/web/mcp/bridges/stdio-golang/nd-mcp
141 ```
142
81 -### Advanced Anomaly Detection
143 +**Important**: When configuring your AI client, use the full absolute path to the executable:
144
83 -Your AI assistant can:
145 +- Go bridge: `/path/to/bridges/stdio-golang/nd-mcp`
146 +- Node.js bridge: `/path/to/bridges/stdio-nodejs/nd-mcp.js`
147 +- Python bridge: `/path/to/bridges/stdio-python/nd-mcp.py`
148
85 -- **Detect anomaly patterns** across your entire infrastructure
86 -- **Identify cascade failures** before they become critical
87 -- **Explain complex service dependencies** in plain language
88 -- **Predict which services will be affected** by specific component failures
149 +### Verify the Bridge Works
150
90 -## Current Capabilities
151 +Once you have nd-mcp (either from existing installation or built), test it:
152
92 -### What You Can Do Today
153 +```bash
154 +# Test connection to your Netdata instance (replace YOUR_NETDATA_IP with actual IP)
155 +/path/to/nd-mcp ws://YOUR_NETDATA_IP:19999/mcp
156
94 -| Capability | What You Get | Key Benefits |
95 -|-----------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------|
96 -| **Infrastructure Analysis** | **Post-mortem analysis** of incidents with **root cause identification**, **real-time anomaly correlation** across all your services, **service dependency mapping** and impact analysis, **performance bottleneck identification** | Understand complex incidents in minutes instead of hours, identify cascading failures before they spread |
97 -| **Intelligent Querying** | Ask questions about your infrastructure in **natural language**, get **explanations of complex metrics** and their relationships, understand **streaming configurations** and network topologies, analyze **resource usage patterns** and trends | No more manual chart correlation - just ask and get expert-level insights |
98 -| **Expert Troubleshooting** | **AI-powered investigation** of performance issues, **automated correlation** of events across your infrastructure, **context-aware recommendations** based on your specific setup | Get the analysis that would normally require a senior engineer, available 24/7 |
157 +# You should see:
158 +# nd-mcp: Connecting to ws://YOUR_NETDATA_IP:19999/mcp...
159 +# nd-mcp: Connected
160 +# Press Ctrl+C to stop the test
161 +```
162
100 -### How to Get Started
163 +## Finding Your API Key
164
102 -1. **Install Netdata nightly build** (required for MCP support)
103 -2. **Set up the MCP bridge** using our nd-mcp tool
104 -3. **Connect your AI assistant** (Claude Desktop recommended)
105 -4. **Start asking questions** about your infrastructure
165 +To access sensitive functions like logs and live system information, you need an API key. Netdata automatically generates an API key on startup. The key is stored in a file on the Netdata server you want to connect to.
166
107 -:::note
167 +You need the API key of the Netdata you will connect to (usually a Netdata Parent).
168
109 -**Current Version**: MCP integration works with both individual Netdata agents and parent agents that aggregate data from multiple nodes. When connected to a parent, AI can analyze data across all nodes that stream to that parent.
169 +**Note**: This temporary API key mechanism will eventually be replaced by integration with Netdata Cloud.
170
111 -:::
171 +### Find the API Key File
172
113 -## What's Coming Next
173 +```bash
174 +# Try the default location first:
175 +sudo cat /var/lib/netdata/mcp_dev_preview_api_key
176
115 -### Netdata Cloud Integration
177 +# For static installations:
178 +sudo cat /opt/netdata/var/lib/netdata/mcp_dev_preview_api_key
179
117 -**The next major step**: MCP will integrate directly with Netdata Cloud, giving you access to your complete infrastructure data:
180 +# If not found, search for it:
181 +sudo find / -name "mcp_dev_preview_api_key" 2>/dev/null
182 +```
183
119 -- **Organization-wide analysis** across all your monitored infrastructure
120 -- **Cross-room correlation** for complex distributed systems
121 -- **Complete infrastructure visibility** through AI analysis
122 -- **Team-wide insights** from your centralized monitoring data
184 +### Copy the API Key
185
124 -### Beyond Single Nodes
186 +The file contains a UUID that looks like:
187
126 -**Current capability**: MCP integration works with individual Netdata agents and parent agents that aggregate multiple nodes.
188 +```
189 +a1b2c3d4-e5f6-7890-abcd-ef1234567890
190 +```
191 +
192 +Copy this entire string - you'll need it for your AI client configuration.
193
128 -**Coming soon**: Direct access to your Netdata Cloud data, enabling AI analysis across your entire infrastructure ecosystem with full organizational context.
194 +### No API Key File?
195
130 -## The Future of Infrastructure Monitoring
196 +If the file doesn't exist:
197
132 -With Netdata MCP, we're moving toward a future where **AI understands your infrastructure** as well as your best engineers. **Troubleshooting becomes conversational** rather than manual chart analysis, **post-mortems write themselves** with complete root cause analysis, and **your monitoring system becomes a true team member** that helps solve problems.
198 +1. Ensure you have a recent version of Netdata
199 +2. Restart Netdata: `sudo systemctl restart netdata`
200 +3. Check the file again after restart
201
134 -:::note
202 +## AI Client Configuration
203
136 -This represents our vision for the future of infrastructure monitoring. While we're making significant progress with MCP integration, these capabilities are aspirational goals we're working toward, not current product features. For now, we are building the foundations for the reality we want to create.
204 +Most AI clients use a similar configuration format:
205
138 -:::
206 +```json
207 +{
208 + "mcpServers": {
209 + "netdata": {
210 + "command": "/usr/sbin/nd-mcp",
211 + "args": [
212 + "ws://IP_OF_YOUR_NETDATA:19999/mcp?api_key=YOUR_API_KEY"
213 + ]
214 + }
215 + }
216 +}
217 +```
218 +
219 +Replace:
220 +
221 +- `/usr/sbin/nd-mcp` - With your actual nd-mcp path
222 +- `IP_OF_YOUR_NETDATA`: Your Netdata instance IP/hostname
223 +- `YOUR_API_KEY`: The API key from the file mentioned above
224 +
225 +### Multiple MCP Servers
226 +
227 +You can configure multiple Netdata instances:
228 +
229 +```json
230 +{
231 + "mcpServers": {
232 + "netdata-production": {
233 + "command": "/usr/sbin/nd-mcp",
234 + "args": ["ws://prod-parent:19999/mcp?api_key=PROD_KEY"]
235 + },
236 + "netdata-testing": {
237 + "command": "/usr/sbin/nd-mcp",
238 + "args": ["ws://test-parent:19999/mcp?api_key=TEST_KEY"]
239 + }
240 + }
241 +}
242 +```
243
140 -**Join us in pioneering the next generation of intelligent infrastructure monitoring.**
244 +Note: Most AI clients have difficulty choosing between multiple MCP servers. You may need to enable/disable them manually.
docs/ml-ai/ai-chat-netdata/ai-chat-netdata.md new
+128
@@ -0,0 +1,128 @@
1 +# AI Chat with Netdata
2 +
3 +Chat with your infrastructure using natural language through two distinct integration architectures.
4 +
5 +## Integration Architecture
6 +
7 +### Method 1: Client-Controlled Communication (Available Now)
8 +
9 +```mermaid
10 +flowchart TB
11 + LLM("🤖 LLM Provider<br/>OpenAI, Anthropic, etc.")
12 +
13 + subgraph infra["Your Infrastructure"]
14 + direction TB
15 + subgraph userLayer[" "]
16 + direction LR
17 + User("👤 User")
18 + Client("💻 AI Client<br/>Claude Desktop, Cursor, etc.")
19 +
20 + User -->|"(1) Ask question"| Client
21 + Client -->|"(8) Display response"| User
22 + end
23 +
24 + Agent("📊 Netdata Agent or Parent<br/>with MCP Server")
25 +
26 + Client -->|"(4) Execute tools"| Agent
27 + Agent -->|"(5) Return data"| Client
28 + end
29 +
30 + Client -->|"(2) Send query"| LLM
31 + LLM -->|"(3) Tool commands"| Client
32 + Client -->|"(6) Send results"| LLM
33 + LLM -->|"(7) Final answer"| Client
34 +
35 + classDef user fill:#e1f5fe,stroke:#01579b,stroke-width:2px
36 + classDef client fill:#f3e5f5,stroke:#4a148c,stroke-width:2px
37 + classDef llm fill:#e8f5e8,stroke:#1b5e20,stroke-width:2px
38 + classDef agent fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
39 + classDef invisible fill:transparent,stroke:transparent
40 +
41 + class User user
42 + class Client client
43 + class LLM llm
44 + class Agent agent
45 + class userLayer invisible
46 +```
47 +
48 +**How it works:**
49 +
50 +1. You ask a question to your AI client
51 +2. LLM responds with tool execution commands
52 +3. Your AI client executes tools against Netdata Agent MCP (locally)
53 +4. Your AI client sends tool responses back to LLM
54 +5. LLM provides the final answer
55 +
56 +**Key characteristics:**
57 +
58 +- Your AI client orchestrates all communication
59 +- Netdata Agent MCP runs locally on your infrastructure
60 +- No internet access required for Netdata Agent
61 +- Full control over data flow and privacy
62 +
63 +### Method 2: LLM-Direct Communication (Coming Soon)
64 +
65 +```mermaid
66 +flowchart TB
67 + LLM("🤖 LLM Provider<br/>OpenAI, Anthropic, etc.")
68 + CloudMCP("☁️ Netdata Cloud<br/>with MCP Server")
69 +
70 + subgraph infra["Your Infrastructure"]
71 + direction TB
72 + subgraph userLayer[" "]
73 + direction LR
74 + User("👤 User")
75 + Client("💻 AI Client")
76 +
77 + User -->|"(1) Ask question"| Client
78 + Client -->|"(6) Display response"| User
79 + end
80 +
81 + Agents("📊 Netdata Agents<br/>and Parents")
82 + end
83 +
84 + Client -->|"(2) Send query"| LLM
85 + LLM -.->|"(3) Access tools"| CloudMCP
86 + CloudMCP -.->|"(4) Return data"| LLM
87 + LLM -->|"(5) Final answer"| Client
88 +
89 + Agents -.-> CloudMCP
90 +
91 + classDef user fill:#e1f5fe,stroke:#01579b,stroke-width:2px
92 + classDef client fill:#f3e5f5,stroke:#4a148c,stroke-width:2px
93 + classDef llm fill:#e8f5e8,stroke:#1b5e20,stroke-width:2px
94 + classDef cloud fill:#e3f2fd,stroke:#0d47a1,stroke-width:2px
95 + classDef agent fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
96 + classDef invisible fill:transparent,stroke:transparent
97 +
98 + class User user
99 + class Client client
100 + class LLM llm
101 + class CloudMCP cloud
102 + class Agents agent
103 + class userLayer invisible
104 +```
105 +
106 +**How it works:**
107 +
108 +1. You ask a question to your AI client
109 +2. LLM directly accesses Netdata Cloud MCP tools
110 +3. LLM provides the final answer with integrated data
111 +
112 +**Key characteristics:**
113 +
114 +- LLM provider manages MCP integration
115 +- Direct connection between LLM and MCP tools
116 +- Netdata Cloud MCP accessible via internet
117 +- Simplified setup, no local MCP configuration needed
118 +
119 +## Quick Comparison
120 +
121 +| Aspect | Method 1: Client-Controlled | Method 2: LLM-Direct |
122 +|--------|---------------------------|---------------------|
123 +| **Availability** | ✅ Available now | 🚧 Coming soon |
124 +| **Setup Complexity** | Moderate (configure AI client + MCP) | Simple (just AI client) |
125 +| **Data Privacy** | Depends on LLM provider | Depends on LLM provider |
126 +| **Internet Requirements** | AI client needs internet, MCP is local | Both AI client and MCP need internet |
127 +| **Supported AI Clients** | Any MCP-aware client (including those using LLM APIs) | Only clients from providers that support MCP on LLM side |
128 +| **Infrastructure Access** | Limited to one Parent's scope | Complete visibility across all infrastructure |
docs/ml-ai/ai-chat-netdata/claude-desktop.md new
+127
@@ -0,0 +1,127 @@
1 +# Claude Desktop
2 +
3 +Configure Claude Desktop to access your Netdata infrastructure through MCP.
4 +
5 +## Prerequisites
6 +
7 +1. **Claude Desktop installed** - Download from [claude.ai/download](https://claude.ai/download)
8 +2. **The IP and port (usually 19999) of a running Netdata Agent** - Prefer a Netdata Parent to get infrastructure level visibility. Currently the latest nightly version of Netdata has MCP support (not released to the stable channel yet). Your AI Client (running on your desktop or laptop) needs to have direct network access to this IP and port.
9 +3. **`nd-mcp` program available on your desktop or laptop** - This is the bridge that translates `stdio` to `websocket`, connecting your AI Client to your Netdata Agent or Parent. [Find its absolute path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
10 +4. **Optionally, the Netdata MCP API key** that unlocks full access to sensitive observability data (protected functions, full access to logs) on your Netdata. Each Netdata Agent or Parent has its own unique API key for MCP - [Find your Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
11 +
12 +## Platform-Specific Installation
13 +
14 +### Windows & macOS
15 +
16 +Download directly from [claude.ai/download](https://claude.ai/download)
17 +
18 +### Linux
19 +
20 +Use the community AppImage project:
21 +
22 +1. Download from [github.com/fsoft72/claude-desktop-to-appimage](https://github.com/fsoft72/claude-desktop-to-appimage)
23 +2. For best experience, install [AppImageLauncher](https://github.com/TheAssassin/AppImageLauncher)
24 +
25 +## Configuration
26 +
27 +1. Open Claude Desktop
28 +2. Navigate to Settings:
29 + - **Windows/Linux**: File → Settings → Developer (or `Ctrl+,`)
30 + - **macOS**: Claude → Settings → Developer (or `Cmd+,`)
31 +3. Click "Edit Config" button
32 +4. Add the Netdata configuration:
33 +
34 +```json
35 +{
36 + "mcpServers": {
37 + "netdata": {
38 + "command": "/usr/sbin/nd-mcp",
39 + "args": [
40 + "ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY"
41 + ]
42 + }
43 + }
44 +}
45 +```
46 +
47 +Replace:
48 +
49 +- `/usr/sbin/nd-mcp` - With your [actual nd-mcp path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
50 +- `YOUR_NETDATA_IP` - IP address or hostname of your Netdata Agent/Parent
51 +- `NETDATA_MCP_API_KEY` - Your [Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
52 +
53 +5. Save the configuration
54 +6. **Restart Claude Desktop** (required for changes to take effect)
55 +
56 +## Verify Connection
57 +
58 +1. Click the "Search and tools" button (below the prompt)
59 +2. You should see "netdata" listed among available tools
60 +3. If not visible, check your configuration and restart
61 +
62 +## Usage Examples
63 +
64 +Simply ask Claude about your infrastructure:
65 +
66 +```
67 +What's the current CPU usage across all my servers?
68 +Show me any anomalies in the last 4 hours
69 +Which processes are consuming the most memory?
70 +Are there any critical alerts active?
71 +Search the logs for authentication failures
72 +```
73 +
74 +## Multiple Environments
75 +
76 +Claude Desktop has limitations with multiple MCP servers. Options:
77 +
78 +### Option 1: Toggle Servers
79 +
80 +Add multiple configurations and enable/disable as needed:
81 +
82 +```json
83 +{
84 + "mcpServers": {
85 + "netdata-production": {
86 + "command": "/usr/sbin/nd-mcp",
87 + "args": ["ws://prod-parent:19999/mcp?api_key=PROD_KEY"]
88 + },
89 + "netdata-staging": {
90 + "command": "/usr/sbin/nd-mcp",
91 + "args": ["ws://stage-parent:19999/mcp?api_key=STAGE_KEY"]
92 + }
93 + }
94 +}
95 +```
96 +
97 +Use the toggle switch in settings to enable only one at a time.
98 +
99 +### Option 2: Single Parent
100 +
101 +Connect to your main Netdata Parent that has visibility across all environments.
102 +
103 +## Troubleshooting
104 +
105 +### Netdata Not Appearing in Tools
106 +
107 +- Ensure configuration file is valid JSON
108 +- Restart Claude Desktop after configuration changes
109 +- Check the bridge path exists and is executable
110 +
111 +### Connection Errors
112 +
113 +- Verify Netdata is accessible from your machine
114 +- Test: `curl http://YOUR_NETDATA_IP:19999/api/v3/info`
115 +- Check firewall rules allow connection to port 19999
116 +
117 +### "Bridge Not Found" Error
118 +
119 +- Verify the nd-mcp path is correct
120 +- Windows users: Include the `.exe` extension
121 +- Ensure Netdata is installed on your local machine (for the bridge)
122 +
123 +### Limited Access to Data
124 +
125 +- Verify API key is included in the connection string
126 +- Ensure the API key file exists on the Netdata server
127 +- Check that functions and logs collectors are enabled
docs/ml-ai/ai-chat-netdata/cursor.md new
+145
@@ -0,0 +1,145 @@
1 +# Cursor
2 +
3 +Configure Cursor IDE to access your Netdata infrastructure through MCP.
4 +
5 +## Prerequisites
6 +
7 +1. **Cursor installed** - Download from [cursor.com](https://www.cursor.com)
8 +2. **The IP and port (usually 19999) of a running Netdata Agent** - Prefer a Netdata Parent to get infrastructure level visibility. Currently the latest nightly version of Netdata has MCP support (not released to the stable channel yet). Your AI Client (running on your desktop or laptop) needs to have direct network access to this IP and port.
9 +3. **`nd-mcp` program available on your desktop or laptop** - This is the bridge that translates `stdio` to `websocket`, connecting your AI Client to your Netdata Agent or Parent. [Find its absolute path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
10 +4. **Optionally, the Netdata MCP API key** that unlocks full access to sensitive observability data (protected functions, full access to logs) on your Netdata. Each Netdata Agent or Parent has its own unique API key for MCP - [Find your Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
11 +
12 +## Configuration
13 +
14 +1. Open Cursor
15 +2. Navigate to Settings:
16 + - **Windows/Linux**: File → Preferences → Settings (or `Ctrl+,`)
17 + - **macOS**: Cursor → Preferences → Settings (or `Cmd+,`)
18 +3. Search for "MCP" in settings
19 +4. Add your Netdata configuration to MCP Servers
20 +
21 +The configuration format:
22 +
23 +```json
24 +{
25 + "mcpServers": {
26 + "netdata": {
27 + "command": "/usr/sbin/nd-mcp",
28 + "args": [
29 + "ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY"
30 + ]
31 + }
32 + }
33 +}
34 +```
35 +
36 +Replace:
37 +
38 +- `/usr/sbin/nd-mcp` - With your [actual nd-mcp path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
39 +- `YOUR_NETDATA_IP` - IP address or hostname of your Netdata Agent/Parent
40 +- `NETDATA_MCP_API_KEY` - Your [Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
41 +
42 +## Using Netdata in Cursor
43 +
44 +### In Chat (Cmd+K)
45 +
46 +Reference Netdata directly in your queries:
47 +
48 +```
49 +@netdata what's the current CPU usage?
50 +@netdata show me database query performance
51 +@netdata are there any anomalies in the web servers?
52 +```
53 +
54 +### In Code Comments
55 +
56 +Get infrastructure context while coding:
57 +
58 +```python
59 +# @netdata what's the typical memory usage of this service?
60 +def process_large_dataset():
61 + # Implementation
62 +```
63 +
64 +### Multi-Model Support
65 +
66 +Cursor's strength is using multiple AI models. You can:
67 +
68 +- Use Claude for complex analysis
69 +- Switch to GPT-4 for different perspectives
70 +- Use smaller models for quick queries
71 +
72 +All models can access your Netdata data through MCP.
73 +
74 +## Multiple Environments
75 +
76 +Cursor allows multiple MCP servers but requires manual toggling:
77 +
78 +```json
79 +{
80 + "mcpServers": {
81 + "netdata-prod": {
82 + "command": "/usr/sbin/nd-mcp",
83 + "args": ["ws://prod-parent:19999/mcp?api_key=PROD_KEY"]
84 + },
85 + "netdata-dev": {
86 + "command": "/usr/sbin/nd-mcp",
87 + "args": ["ws://dev-parent:19999/mcp?api_key=DEV_KEY"]
88 + }
89 + }
90 +}
91 +```
92 +
93 +Use the toggle in settings to enable only the environment you need.
94 +
95 +## Best Practices
96 +
97 +### Infrastructure-Aware Development
98 +
99 +While coding, ask about:
100 +
101 +- Current resource usage of services you're modifying
102 +- Historical performance patterns
103 +- Impact of deployments on system metrics
104 +
105 +### Debugging with Context
106 +
107 +```
108 +@netdata show me the logs when this error last occurred
109 +@netdata what was the system state during the last deployment?
110 +@netdata find correlated metrics during the performance regression
111 +```
112 +
113 +### Performance Optimization
114 +
115 +```
116 +@netdata analyze database query latency patterns
117 +@netdata which endpoints have the highest response times?
118 +@netdata show me resource usage trends for this service
119 +```
120 +
121 +## Troubleshooting
122 +
123 +### MCP Server Not Available
124 +
125 +- Restart Cursor after adding configuration
126 +- Verify JSON syntax in settings
127 +- Check MCP is enabled in Cursor settings
128 +
129 +### Connection Issues
130 +
131 +- Test Netdata accessibility: `curl http://YOUR_NETDATA_IP:19999/api/v3/info`
132 +- Verify bridge path is correct and executable
133 +- Check firewall allows connection to Netdata
134 +
135 +### Multiple Servers Confusion
136 +
137 +- Cursor may query the wrong server if multiple are enabled
138 +- Always disable unused servers
139 +- Name servers clearly (prod, dev, staging)
140 +
141 +### Limited Functionality
142 +
143 +- Ensure API key is included for full access
144 +- Verify Netdata agent is claimed
145 +- Check that required collectors are enabled
docs/ml-ai/ai-chat-netdata/jetbrains-ides.md new
+214
@@ -0,0 +1,214 @@
1 +# JetBrains IDEs
2 +
3 +Configure JetBrains IDEs to access your Netdata infrastructure through MCP.
4 +
5 +## Supported IDEs
6 +
7 +- IntelliJ IDEA
8 +- PyCharm
9 +- WebStorm
10 +- PhpStorm
11 +- GoLand
12 +- DataGrip
13 +- Rider
14 +- CLion
15 +- RubyMine
16 +
17 +## Prerequisites
18 +
19 +1. **JetBrains IDE installed** - Any IDE from the list above
20 +2. **AI Assistant plugin** - Install from IDE marketplace
21 +3. **The IP and port (usually 19999) of a running Netdata Agent** - Prefer a Netdata Parent to get infrastructure level visibility. Currently the latest nightly version of Netdata has MCP support (not released to the stable channel yet). Your AI Client (running on your desktop or laptop) needs to have direct network access to this IP and port.
22 +4. **`nd-mcp` program available on your desktop or laptop** - This is the bridge that translates `stdio` to `websocket`, connecting your AI Client to your Netdata Agent or Parent. [Find its absolute path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
23 +5. **Optionally, the Netdata MCP API key** that unlocks full access to sensitive observability data (protected functions, full access to logs) on your Netdata. Each Netdata Agent or Parent has its own unique API key for MCP - [Find your Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
24 +
25 +## Installing AI Assistant
26 +
27 +1. Open your JetBrains IDE
28 +2. Go to Settings/Preferences:
29 + - **Windows/Linux**: File → Settings → Plugins
30 + - **macOS**: IntelliJ IDEA → Preferences → Plugins
31 +3. Search for "AI Assistant" in Marketplace
32 +4. Install and restart IDE
33 +
34 +## MCP Configuration
35 +
36 +:::note
37 +MCP support in JetBrains IDEs may require additional plugins or configuration. Check the plugin documentation for the latest setup instructions.
38 +:::
39 +
40 +### Method 1: AI Assistant Settings
41 +
42 +1. Go to Settings → Tools → AI Assistant
43 +2. Look for MCP or External Tools configuration
44 +3. Add Netdata MCP server:
45 +
46 +```json
47 +{
48 + "name": "netdata",
49 + "command": "/usr/sbin/nd-mcp",
50 + "args": [
51 + "ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY"
52 + ]
53 +}
54 +```
55 +
56 +### Method 2: External Tools
57 +
58 +If direct MCP support is not available, configure as an External Tool:
59 +
60 +1. Go to Settings → Tools → External Tools
61 +2. Click "+" to add new tool
62 +3. Configure:
63 + - **Name**: Netdata MCP
64 + - **Program**: `/usr/sbin/nd-mcp`
65 + - **Arguments**: `ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY`
66 +
67 +Replace:
68 +
69 +- `/usr/sbin/nd-mcp` - With your [actual nd-mcp path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
70 +- `YOUR_NETDATA_IP` - IP address or hostname of your Netdata Agent/Parent
71 +- `NETDATA_MCP_API_KEY` - Your [Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
72 +
73 +## Usage in Different IDEs
74 +
75 +### IntelliJ IDEA (Java/Kotlin)
76 +
77 +Monitor JVM applications:
78 +
79 +```
80 +// Ask AI Assistant about production performance
81 +"What's the memory usage of our Java services?"
82 +"Show me GC patterns in the last hour"
83 +"Are there any thread pool issues?"
84 +```
85 +
86 +### PyCharm (Python)
87 +
88 +Debug Python applications:
89 +
90 +```python
91 +# Ask: What's the CPU usage when this function runs in production?
92 +def process_data():
93 + pass
94 +
95 +# Ask: Show me memory patterns for the Python workers
96 +```
97 +
98 +### WebStorm (JavaScript/TypeScript)
99 +
100 +Monitor Node.js applications:
101 +
102 +```javascript
103 +// Ask: What's the event loop latency?
104 +// Ask: Show me API endpoint response times
105 +// Ask: Any memory leaks in the Node processes?
106 +```
107 +
108 +### DataGrip (Databases)
109 +
110 +Analyze database performance:
111 +
112 +```sql
113 +-- Ask: Show me database query latency
114 +-- Ask: What's the connection pool usage?
115 +-- Ask: Any slow queries in the last hour?
116 +```
117 +
118 +## IDE-Specific Features
119 +
120 +### Code Annotations
121 +
122 +Add infrastructure context to your code:
123 +
124 +```java
125 +@NetdataMonitor("cpu.usage > 80%")
126 +public void resourceIntensiveMethod() {
127 + // AI Assistant can show real-time metrics
128 +}
129 +```
130 +
131 +### Debugging with Metrics
132 +
133 +While debugging:
134 +
135 +1. Set breakpoint
136 +2. Ask AI Assistant: "What were the system metrics when this code last ran in production?"
137 +3. Get historical context for better debugging
138 +
139 +### Performance Profiling
140 +
141 +Combine IDE profiler with Netdata metrics:
142 +
143 +- Run profiler in IDE
144 +- Ask: "Show me system metrics during the profiling period"
145 +- Correlate application and system performance
146 +
147 +## Best Practices
148 +
149 +### Development Workflow
150 +
151 +1. Before deploying: "What's the current production load?"
152 +2. During testing: "Compare metrics between dev and prod"
153 +3. After deployment: "Show me metrics changes after deployment"
154 +
155 +### Troubleshooting Production Issues
156 +
157 +```
158 +"Show me what happened at 14:32 when the error occurred"
159 +"What were the system resources during the last OutOfMemory error?"
160 +"Find correlated metrics during the last service degradation"
161 +```
162 +
163 +### Capacity Planning
164 +
165 +```
166 +"What's the resource usage trend for this service?"
167 +"Project memory needs based on current growth"
168 +"When will we need to scale based on current patterns?"
169 +```
170 +
171 +## Plugin Alternatives
172 +
173 +If official MCP support is limited, consider:
174 +
175 +### MCP Bridge Plugin
176 +
177 +Search marketplace for:
178 +
179 +- "MCP Client"
180 +- "Model Context Protocol"
181 +- "External AI Tools"
182 +
183 +### Custom Plugin Development
184 +
185 +Create a simple plugin that bridges JetBrains with Netdata:
186 +
187 +1. Use IDE Plugin SDK
188 +2. Implement MCP client
189 +3. Add tool window for Netdata metrics
190 +
191 +## Troubleshooting
192 +
193 +### AI Assistant Not Connecting
194 +
195 +- Check MCP configuration in settings
196 +- Restart IDE after configuration changes
197 +
198 +### No Netdata Option
199 +
200 +- Ensure latest AI Assistant version
201 +- Check for additional MCP plugins
202 +- Try External Tools approach
203 +
204 +### Connection Errors
205 +
206 +- Test Netdata access: `curl http://YOUR_NETDATA_IP:19999/api/v3/info`
207 +- Verify bridge path and permissions
208 +- Check IDE logs for detailed errors
209 +
210 +### Limited Functionality
211 +
212 +- Some IDEs may have restricted AI Assistant features
213 +- Try different JetBrains IDEs for better support
214 +- Consider using Cursor or VS Code for full MCP support
docs/ml-ai/ai-chat-netdata/netdata-web-client.md new
+94
@@ -0,0 +1,94 @@
1 +# Netdata Web Client
2 +
3 +A self-hosted AI chat interface purpose-built for infrastructure observability, featuring advanced cost optimization and multi-provider LLM support.
4 +
5 +![Netdata Web Client Interface](https://github.com/user-attachments/assets/f2facc59-66c1-4ea5-9404-d335e8f67ff2)
6 +
7 +## Purpose-Built for Observability
8 +
9 +The Netdata Web Client is an open-source browser-based AI assistant that connects directly to your Netdata infrastructure via MCP (Model Context Protocol). Unlike generic AI chat interfaces, it's specifically optimized for DevOps and SRE workflows with specialized system prompts and infrastructure-aware features.
10 +
11 +## Key Features
12 +
13 +### Multi-Provider LLM Support
14 +
15 +- **Unified Interface** - Switch between OpenAI (GPT-4), Anthropic (Claude), and Google (Gemini) models within the same conversation
16 +- **Model Discovery** - Automatically detects available models from each provider
17 +- **Provider Arbitrage** - Use cheaper models for simple queries, premium models for complex analysis
18 +
19 +### Advanced Cost Optimization
20 +
21 +- **Real-time Cost Tracking** - See exact costs per message and cumulative session costs
22 +- **Token Accounting** - Detailed breakdown of input, output, cache read, and cache write tokens
23 +- **Smart Context Management**:
24 + - Automatic tool memory pruning after N conversation turns
25 + - Large response summarization using cheaper models
26 + - Configurable cache control strategies
27 + - Auto-summarization when approaching context limits
28 +- **Safety Limits** - Prevents runaway costs with iteration and request size limits
29 +
30 +### Infrastructure-Specific Features
31 +
32 +- **DevOps System Prompts** - Pre-configured for infrastructure monitoring and analysis
33 +- **Time-Aware Analysis** - Understands relative time references ("last night", "this morning")
34 +- **Multi-Chat Architecture** - Run parallel investigations in separate contexts
35 +- **MCP WebSocket Integration** - Real-time connection to multiple Netdata instances
36 +
37 +### Professional Observability Workflow
38 +
39 +- **Persistent Conversations** - All chats auto-saved with full history
40 +- **Accounting Logs** - JSONL export for cost analysis and auditing
41 +- **Error Recovery** - Maintains state through API failures
42 +- **Rich Formatting** - Markdown, code blocks, tables, and ASCII diagrams
43 +
44 +## Cost-Effective Infrastructure Analysis
45 +
46 +The web client's cost optimization features make AI-assisted troubleshooting affordable at scale:
47 +
48 +| Feature | Cost Impact |
49 +|---------|------------|
50 +| Tool Memory Window | -40% context size |
51 +| Response Summarization | -60% for large outputs |
52 +| Smart Caching | -30% for repetitive queries |
53 +| Model Selection | -80% using appropriate models |
54 +
55 +## Get Started
56 +
57 +The complete source code, installation instructions, and documentation are available on GitHub:
58 +
59 +🔗 **[github.com/netdata/netdata/tree/master/src/web/mcp/mcp-web-client](https://github.com/netdata/netdata/tree/master/src/web/mcp/mcp-web-client)**
60 +
61 +### Quick Start
62 +
63 +```bash
64 +# Clone and run locally
65 +git clone https://github.com/netdata/netdata.git
66 +cd netdata/src/web/mcp/mcp-web-client
67 +node llm-proxy.js
68 +# Open http://localhost:3456 in your browser
69 +```
70 +
71 +### Requirements
72 +
73 +- Node.js 18+
74 +- API keys from OpenAI, Anthropic, or Google
75 +- Access to a Netdata instance with MCP enabled
76 +- Modern web browser
77 +
78 +## Why Self-Host?
79 +
80 +- **Data Privacy** - Your infrastructure data never leaves your control
81 +- **Cost Control** - Use your own API keys with transparent pay-per-use pricing
82 +- **Customization** - Modify prompts and behavior for your specific needs
83 +- **No Vendor Lock-in** - Switch LLM providers anytime
84 +
85 +## Ideal For
86 +
87 +- **Cost-Conscious Teams** - Pay only for what you use
88 +- **Security-Focused Organizations** - Keep all data within your infrastructure
89 +- **Advanced Users** - Full control over prompts and model selection
90 +- **Multi-Cloud Environments** - Connect to multiple Netdata instances
91 +
92 +---
93 +
94 +For detailed setup instructions, configuration options, and the complete feature list, visit the [GitHub repository](https://github.com/netdata/netdata/tree/master/src/web/mcp/mcp-web-client).
docs/ml-ai/ai-chat-netdata/vs-code.md new
+245
@@ -0,0 +1,245 @@
1 +# VS Code
2 +
3 +Configure Visual Studio Code extensions to access your Netdata infrastructure through MCP.
4 +
5 +## Available Extensions
6 +
7 +### Continue (Recommended)
8 +
9 +The most popular open-source AI code assistant with MCP support.
10 +
11 +### Cline
12 +
13 +Autonomous coding agent that can use MCP tools.
14 +
15 +## Prerequisites
16 +
17 +1. **VS Code installed** - [Download VS Code](https://code.visualstudio.com)
18 +2. **MCP-compatible extension** - Install from VS Code Marketplace
19 +3. **The IP and port (usually 19999) of a running Netdata Agent** - Prefer a Netdata Parent to get infrastructure level visibility. Currently the latest nightly version of Netdata has MCP support (not released to the stable channel yet). Your AI Client (running on your desktop or laptop) needs to have direct network access to this IP and port.
20 +4. **`nd-mcp` program available on your desktop or laptop** - This is the bridge that translates `stdio` to `websocket`, connecting your AI Client to your Netdata Agent or Parent. [Find its absolute path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
21 +5. **Optionally, the Netdata MCP API key** that unlocks full access to sensitive observability data (protected functions, full access to logs) on your Netdata. Each Netdata Agent or Parent has its own unique API key for MCP - [Find your Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
22 +
23 +## Continue Extension Setup
24 +
25 +### Installation
26 +
27 +1. Open VS Code
28 +2. Go to Extensions (Ctrl+Shift+X)
29 +3. Search for "Continue"
30 +4. Install the Continue extension
31 +5. Reload VS Code
32 +
33 +### Configuration
34 +
35 +1. Open Command Palette (Ctrl+Shift+P)
36 +2. Type "Continue: Open config.json"
37 +3. Add Netdata MCP configuration:
38 +
39 +```json
40 +{
41 + "models": [
42 + {
43 + "title": "Claude 3.5 Sonnet",
44 + "provider": "anthropic",
45 + "model": "claude-3-5-sonnet-20241022",
46 + "apiKey": "YOUR_ANTHROPIC_KEY"
47 + }
48 + ],
49 + "mcpServers": {
50 + "netdata": {
51 + "command": "/usr/sbin/nd-mcp",
52 + "args": [
53 + "ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY"
54 + ]
55 + }
56 + }
57 +}
58 +```
59 +
60 +Replace:
61 +
62 +- `/usr/sbin/nd-mcp` - With your [actual nd-mcp path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
63 +- `YOUR_NETDATA_IP` - IP address or hostname of your Netdata Agent/Parent
64 +- `NETDATA_MCP_API_KEY` - Your [Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
65 +- `YOUR_ANTHROPIC_KEY` - Your Anthropic API key
66 +
67 +### Usage
68 +
69 +Press `Ctrl+L` to open Continue chat, then:
70 +
71 +```
72 +@netdata what's the current CPU usage?
73 +@netdata show me memory trends for the last hour
74 +@netdata are there any anomalies in the database servers?
75 +```
76 +
77 +## Cline Extension Setup
78 +
79 +### Installation
80 +
81 +1. Search for "Cline" in Extensions
82 +2. Install and reload VS Code
83 +
84 +### Configuration
85 +
86 +1. Open Settings (Ctrl+,)
87 +2. Search for "Cline MCP"
88 +3. Add configuration:
89 +
90 +```json
91 +{
92 + "cline.mcpServers": [
93 + {
94 + "name": "netdata",
95 + "command": "/usr/sbin/nd-mcp",
96 + "args": [
97 + "ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY"
98 + ]
99 + }
100 + ]
101 +}
102 +```
103 +
104 +### Usage
105 +
106 +1. Open Cline (Ctrl+Shift+P → "Cline: Open Chat")
107 +2. Cline can autonomously:
108 + - Analyze performance issues
109 + - Create monitoring scripts
110 + - Debug based on metrics
111 +
112 +Example:
113 +
114 +```
115 +Create a Python script that checks Netdata for high CPU usage and sends an alert
116 +```
117 +
118 +## Multiple Environments
119 +
120 +### Workspace-Specific Configuration
121 +
122 +Create `.vscode/settings.json` in your project:
123 +
124 +```json
125 +{
126 + "continue.mcpServers": {
127 + "netdata-prod": {
128 + "command": "/usr/sbin/nd-mcp",
129 + "args": ["ws://prod-parent:19999/mcp?api_key=PROD_NETDATA_MCP_API_KEY"]
130 + }
131 + }
132 +}
133 +```
134 +
135 +### Environment Switching
136 +
137 +Different projects can have different Netdata connections:
138 +
139 +- `~/projects/frontend/.vscode/settings.json` → Frontend servers
140 +- `~/projects/backend/.vscode/settings.json` → Backend servers
141 +- `~/projects/infrastructure/.vscode/settings.json` → All servers
142 +
143 +## Advanced Usage
144 +
145 +### Custom Commands
146 +
147 +Create custom VS Code commands that query Netdata:
148 +
149 +```json
150 +{
151 + "commands": [
152 + {
153 + "command": "netdata.checkHealth",
154 + "title": "Netdata: Check System Health"
155 + }
156 + ]
157 +}
158 +```
159 +
160 +### Task Integration
161 +
162 +Add Netdata checks to tasks.json:
163 +
164 +```json
165 +{
166 + "version": "2.0.0",
167 + "tasks": [
168 + {
169 + "label": "Check Production Metrics",
170 + "type": "shell",
171 + "command": "continue",
172 + "args": ["--ask", "@netdata show current system status"]
173 + }
174 + ]
175 +}
176 +```
177 +
178 +### Snippets with Metrics
179 +
180 +Create snippets that include metric checks:
181 +
182 +```json
183 +{
184 + "Check Performance": {
185 + "prefix": "perf",
186 + "body": [
187 + "// @netdata: Current ${1:CPU} usage?",
188 + "$0"
189 + ]
190 + }
191 +}
192 +```
193 +
194 +## Extension Comparison
195 +
196 +| Feature | Continue | Cline | Codeium | Copilot Chat |
197 +|---------|----------|-------|---------|--------------|
198 +| MCP Support | ✅ Full | ✅ Full | ❓ Check | ❓ Future |
199 +| Autonomous Actions | ❌ | ✅ | ❌ | ❌ |
200 +| Multiple Models | ✅ | ✅ | ❌ | ❌ |
201 +| Free Tier | ❌ | ❌ | ✅ | ❌ |
202 +| Open Source | ✅ | ✅ | ❌ | ❌ |
203 +
204 +## Troubleshooting
205 +
206 +### Extension Not Finding MCP
207 +
208 +- Restart VS Code after configuration
209 +- Check extension logs (Output → Continue/Cline)
210 +- Verify JSON syntax in settings
211 +
212 +### Connection Issues
213 +
214 +- Test Netdata: `curl http://YOUR_NETDATA_IP:19999/api/v3/info`
215 +- Check bridge is executable
216 +- Verify network access from VS Code
217 +
218 +### No Netdata Option
219 +
220 +- Ensure `@netdata` is typed correctly
221 +- Check MCP server is configured
222 +- Try reloading the window (Ctrl+R)
223 +
224 +### Performance Problems
225 +
226 +- Use local Netdata Parent for faster response
227 +- Check extension memory usage
228 +- Disable unused extensions
229 +
230 +## Best Practices
231 +
232 +### Development Workflow
233 +
234 +1. Start coding with infrastructure context
235 +2. Check metrics before optimization
236 +3. Validate changes against production data
237 +4. Monitor impact of deployments
238 +
239 +### Team Collaboration
240 +
241 +Share Netdata configurations:
242 +
243 +- Commit `.vscode/settings.json` for project-specific configs
244 +- Document which Netdata Parent to use
245 +- Create team snippets for common queries
docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md new
+213
@@ -0,0 +1,213 @@
1 +# AI DevOps Copilot
2 +
3 +Command-line AI assistants like **Claude Code** and **Gemini CLI** represent a revolutionary shift in how infrastructure professionals work. These tools combine the power of large language models with access to observability data and the ability to execute system commands, creating unprecedented automation opportunities.
4 +
5 +## The Power of CLI-based AI Assistants
6 +
7 +### Key Capabilities
8 +
9 +**Observability-Driven Operations:**
10 +
11 +- Access real-time metrics and logs from monitoring systems
12 +- Analyze performance trends and identify bottlenecks
13 +- Correlate issues across multiple systems and services
14 +
15 +**System Configuration Management:**
16 +
17 +- Generate and modify configuration files based on observed conditions
18 +- Implement best practices automatically
19 +- Adapt configurations to changing requirements
20 +
21 +**Automated Troubleshooting:**
22 +
23 +- Diagnose issues using multiple data sources
24 +- Execute diagnostic commands and interpret results
25 +- Implement fixes based on root cause analysis
26 +
27 +## Observability + Automation Use Cases
28 +
29 +When AI assistants have access to observability data (like Netdata through MCP), they can make informed decisions about system changes:
30 +
31 +### Infrastructure Optimization Examples
32 +
33 +**Database Performance Tuning:**
34 +
35 +```
36 +PostgreSQL is showing high query response times. Check the metrics and optimize
37 +the configuration.
38 +```
39 +
40 +The AI analyzes connection counts, query performance, and resource usage to adjust connection pools, memory settings, and query optimization parameters.
41 +
42 +**Resource Management:**
43 +
44 +```
45 +This Kubernetes cluster is experiencing frequent pod restarts. Investigate and
46 +fix the resource allocation.
47 +```
48 +
49 +The AI examines CPU, memory, and network metrics to identify resource constraints and adjust limits, requests, and HPA configurations.
50 +
51 +**Storage Optimization:**
52 +
53 +```
54 +Disk usage is growing rapidly on our log servers. Implement appropriate
55 +retention policies.
56 +```
57 +
58 +The AI analyzes disk growth patterns, identifies log volume trends, and configures rotation, compression, and cleanup policies.
59 +
60 +**Network Performance:**
61 +
62 +```
63 +API response times are inconsistent. Check network metrics and optimize the
64 +load balancer configuration.
65 +```
66 +
67 +The AI examines network latency, connection distribution, and backend health to adjust load balancing algorithms and connection settings.
68 +
69 +**Monitoring Setup:**
70 +
71 +```
72 +This server runs Redis but we're not monitoring it properly. Please configure
73 +comprehensive monitoring.
74 +```
75 +
76 +The AI detects the Redis installation, configures appropriate collectors, sets up alerting thresholds, and verifies metric collection.
77 +
78 +**Auto-scaling Configuration:**
79 +
80 +```
81 +Set up intelligent auto-scaling based on current usage patterns I'm seeing.
82 +```
83 +
84 +The AI analyzes historical resource utilization to configure scaling policies, thresholds, and cooldown periods that match actual workload patterns.
85 +
86 +**Complex Test Environment Setup:**
87 +
88 +```
89 +I need a complete test environment that mirrors our production setup: a
90 +multi-tier application with PostgreSQL primary/replica, Redis cluster, message
91 +queues, and load balancers. Set up everything with a Netdata monitoring
92 +everything and realistic test data.
93 +```
94 +
95 +The AI leverages its deep knowledge of application architectures and Netdata's monitoring capabilities to:
96 +
97 +- Deploy and configure all required services with production-like settings
98 +- Set up database replication, clustering, and connection pooling
99 +- Configure realistic test datasets and user simulation
100 +- Implement comprehensive monitoring for all components with appropriate alerts
101 +- Create load testing scenarios that match production traffic patterns
102 +- Establish proper network segmentation and security configurations
103 +- Generate documentation for the test environment and runbooks for common scenarios
104 +
105 +Keep in mind however, that usually this prompt should be split into multiple smaller prompts, so that the LLM can focus on completing a smaller task at a time.
106 +
107 +This showcases how AI can combine application expertise, infrastructure knowledge, and observability best practices to create sophisticated testing environments that would typically require weeks of manual setup and deep domain expertise.
108 +
109 +## ⚠️ Critical Security and Safety Considerations
110 +
111 +### Command Execution Risks
112 +
113 +**LLMs Are Not Infallible:**
114 +
115 +- AI assistants can misinterpret requirements or generate incorrect commands
116 +- Complex system interactions may not be fully understood by the model
117 +- Edge cases and system-specific configurations can lead to unexpected results
118 +
119 +**System Impact Awareness:**
120 +
121 +- Commands can affect system stability, performance, and security
122 +- Changes may have cascading effects across interconnected services
123 +- Recovery from AI-generated misconfigurations can be time-consuming
124 +
125 +### Data Privacy and Security Concerns
126 +
127 +**External LLM Provider Exposure:**
128 +
129 +- All data accessed by the AI (files, configurations, command outputs) is transmitted to external providers
130 +- Sensitive information like passwords, API keys, certificates, and secrets may be inadvertently exposed
131 +- Infrastructure topology, performance metrics, and operational details become visible to third parties
132 +- Compliance requirements (GDPR, HIPAA, SOX) may be violated by external data transmission
133 +
134 +**Network and System Information:**
135 +
136 +- Database connection strings and credentials
137 +- Network topology and security configurations
138 +- Application secrets and encryption keys
139 +- User data and personally identifiable information
140 +
141 +### Recommended Safe Usage Practices
142 +
143 +**1. Analysis-First Approach:**
144 +
145 +```
146 +Instead of: Fix the high CPU usage on server X
147 +Try: Analyze the CPU metrics on server X and explain what might be causing
148 +high usage and what solutions you recommend
149 +```
150 +
151 +**2. Review and Validation:**
152 +
153 +- Always review AI-generated commands before execution
154 +- Test suggestions in development environments first
155 +- Understand the impact and side effects of proposed changes
156 +- Have rollback procedures ready
157 +
158 +**3. Data Sanitization:**
159 +
160 +- Remove or mask sensitive information before sharing with AI
161 +- Use environment variables or placeholder values for secrets
162 +- Avoid sharing production credentials or keys
163 +- Consider using development/staging data for analysis
164 +
165 +**4. Graduated Permissions:**
166 +
167 +- Start with read-only access for analysis
168 +- Grant execution permissions gradually based on trust and validation
169 +- Use separate accounts with limited privileges for AI operations
170 +- Implement audit logging for all AI-initiated changes
171 +
172 +**5. Environment Separation:**
173 +
174 +- Use AI assistance primarily in development and testing environments
175 +- Require manual approval for production changes
176 +- Implement change management processes for AI-suggested modifications
177 +- Maintain air-gapped environments for highly sensitive systems
178 +
179 +## Best Practices for Implementation
180 +
181 +### Safe Integration Workflow
182 +
183 +1. **Discovery Phase:** Let AI analyze your current setup and identify opportunities
184 +2. **Planning Phase:** Have AI generate detailed implementation plans with explanations
185 +3. **Review Phase:** Manually review all suggested changes and commands
186 +4. **Testing Phase:** Implement changes in non-production environments
187 +5. **Validation Phase:** Verify results match expectations before production deployment
188 +6. **Documentation Phase:** Have AI help document the changes and their rationale
189 +
190 +### Building Trust Over Time
191 +
192 +- Start with simple, low-risk tasks to build confidence
193 +- Gradually increase complexity as you validate AI accuracy
194 +- Develop institutional knowledge about AI strengths and limitations
195 +- Create feedback loops to improve AI prompts and instructions
196 +
197 +### Team Education and Guidelines
198 +
199 +- Train team members on safe AI usage practices
200 +- Establish clear guidelines for when AI assistance is appropriate
201 +- Create approval processes for AI-suggested changes
202 +- Share lessons learned and best practices across teams
203 +
204 +## The Future of AI-Driven Operations
205 +
206 +CLI-based AI assistants represent the beginning of a transformation in infrastructure management. As these tools mature, they will likely become central to:
207 +
208 +- **Predictive Operations:** Proactively identifying and preventing issues before they occur
209 +- **Adaptive Infrastructure:** Systems that automatically optimize themselves based on changing conditions
210 +- **Intelligent Automation:** Context-aware automation that understands business impact
211 +- **Enhanced Collaboration:** AI as a knowledgeable team member that augments human expertise
212 +
213 +However, the human element remains crucial for oversight, validation, and strategic decision-making. The most successful implementations will be those that thoughtfully balance AI capabilities with human judgment and appropriate safety measures.
docs/ml-ai/ai-devops-copilot/claude-code.md new
+142
@@ -0,0 +1,142 @@
1 +# Claude Code
2 +
3 +Configure Claude Code to access your Netdata infrastructure through MCP.
4 +
5 +## Prerequisites
6 +
7 +1. **Claude Code installed** - Available at [anthropic.com/claude-code](https://www.anthropic.com/claude-code)
8 +2. **The IP and port (usually 19999) of a running Netdata Agent** - Prefer a Netdata Parent to get infrastructure level visibility. Currently the latest nightly version of Netdata has MCP support (not released to the stable channel yet). Your AI Client (running on your desktop or laptop) needs to have direct network access to this IP and port.
9 +3. **`nd-mcp` program available on your desktop or laptop** - This is the bridge that translates `stdio` to `websocket`, connecting your AI Client to your Netdata Agent or Parent. [Find its absolute path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
10 +4. **Optionally, the Netdata MCP API key** that unlocks full access to sensitive observability data (protected functions, full access to logs) on your Netdata. Each Netdata Agent or Parent has its own unique API key for MCP - [Find your Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
11 +
12 +## Configuration
13 +
14 +Claude Code has comprehensive MCP server management capabilities. For detailed documentation on all configuration options and commands, see the [official Claude Code MCP documentation](https://docs.anthropic.com/en/docs/claude-code/mcp).
15 +
16 +### Adding Netdata MCP Server
17 +
18 +Use Claude Code's built-in MCP commands to add your Netdata server:
19 +
20 +```bash
21 +# Add Netdata MCP server (project-scoped for team sharing)
22 +claude mcp add --scope project netdata /usr/sbin/nd-mcp ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY
23 +
24 +# Or add locally for personal use only
25 +claude mcp add netdata /usr/sbin/nd-mcp ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY
26 +
27 +# List configured servers to verify
28 +claude mcp list
29 +
30 +# Get server details
31 +claude mcp get netdata
32 +```
33 +
34 +Replace:
35 +
36 +- `/usr/sbin/nd-mcp` - With your [actual nd-mcp path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
37 +- `YOUR_NETDATA_IP` - IP address or hostname of your Netdata Agent/Parent
38 +- `NETDATA_MCP_API_KEY` - Your [Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
39 +
40 +**Project-scoped configuration** creates a `.mcp.json` file that can be shared with your team via version control.
41 +
42 +## How to Use
43 +
44 +Claude Code can automatically use Netdata MCP when you ask infrastructure-related questions. If Netdata is your only observability solution configured via MCP, simply ask your question naturally:
45 +
46 +```
47 +What's the current CPU usage across all servers?
48 +Show me any anomalies in the last hour
49 +Which processes are consuming the most memory?
50 +```
51 +
52 +### Explicit MCP Server Selection
53 +
54 +Claude Code also allows you to explicitly specify which MCP server to use with the `/mcp` command:
55 +
56 +1. Open Claude Code in the directory containing `.mcp.json`
57 +2. Type `/mcp` to verify Netdata is available
58 +3. Use `/mcp netdata` followed by your query:
59 +
60 +```
61 +/mcp netdata describe my infrastructure
62 +/mcp netdata what alerts are currently active?
63 +/mcp netdata show me database performance metrics
64 +```
65 +
66 +This is particularly useful when you have multiple MCP servers configured and want to ensure Claude uses the correct one.
67 +
68 +> **💡 Advanced Usage:** Claude Code can combine observability data with system automation for powerful DevOps workflows. Learn about the opportunities and security considerations in [AI DevOps Copilot](/docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md).
69 +
70 +## Project-Based Configuration
71 +
72 +Claude Code's strength is project-specific configurations. So you can have different project directories with different MCP servers on each of them, allowing you to control the MCP servers that will be used, based on the directory from which you started it.
73 +
74 +### Production Environment
75 +
76 +Create `~/projects/production/.mcp.json`:
77 +
78 +```json
79 +{
80 + "mcpServers": {
81 + "netdata": {
82 + "command": "/usr/sbin/nd-mcp",
83 + "args": ["ws://prod-parent.company.com:19999/mcp?api_key=PROD_KEY"]
84 + }
85 + }
86 +}
87 +```
88 +
89 +### Development Environment
90 +
91 +Create `~/projects/development/.mcp.json`:
92 +
93 +```json
94 +{
95 + "mcpServers": {
96 + "netdata": {
97 + "command": "/usr/sbin/nd-mcp",
98 + "args": ["ws://dev-parent.company.com:19999/mcp?api_key=DEV_KEY"]
99 + }
100 + }
101 +}
102 +```
103 +
104 +## Claude Instructions
105 +
106 +Create a `Claude.md` file in your project root with default instructions:
107 +
108 +```markdown
109 +# Claude Instructions
110 +
111 +You have access to Netdata monitoring for our production infrastructure.
112 +
113 +When I ask about performance or issues:
114 +1. Always check current metrics first
115 +2. Look for anomalies in the relevant time period
116 +3. Check logs if investigating errors
117 +4. Provide specific metric values and timestamps
118 +
119 +Our key services to monitor:
120 +- Web servers (nginx)
121 +- Databases (PostgreSQL, Redis)
122 +- Message queues (RabbitMQ)
123 +```
124 +
125 +## Troubleshooting
126 +
127 +### MCP Not Available
128 +
129 +- Ensure `.mcp.json` is in the current directory
130 +- Restart Claude Code after creating the configuration
131 +- Verify the JSON syntax is correct
132 +
133 +### Connection Failed
134 +
135 +- Check Netdata is accessible: `curl http://YOUR_NETDATA_IP:19999/api/v3/info`
136 +- Verify the bridge path exists and is executable
137 +- Ensure API key is correct
138 +
139 +### Limited Data Access
140 +
141 +- Verify API key is included in the connection string
142 +- Check that the Netdata agent is claimed
docs/ml-ai/ai-devops-copilot/gemini-cli.md new
+130
@@ -0,0 +1,130 @@
1 +# Gemini CLI
2 +
3 +Configure Google's Gemini CLI to access your Netdata infrastructure through MCP for powerful AI-driven operations.
4 +
5 +## Prerequisites
6 +
7 +1. **Gemini CLI installed** - Available from [GitHub](https://github.com/google-gemini/gemini-cli)
8 +2. **The IP and port (usually 19999) of a running Netdata Agent** - Prefer a Netdata Parent to get infrastructure level visibility. Currently the latest nightly version of Netdata has MCP support (not released to the stable channel yet). Your AI Client (running on your desktop or laptop) needs to have direct network access to this IP and port.
9 +3. **`nd-mcp` program available on your desktop or laptop** - This is the bridge that translates `stdio` to `websocket`, connecting your AI Client to your Netdata Agent or Parent. [Find its absolute path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
10 +4. **Optionally, the Netdata MCP API key** that unlocks full access to sensitive observability data (protected functions, full access to logs) on your Netdata. Each Netdata Agent or Parent has its own unique API key for MCP - [Find your Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
11 +
12 +## Installation
13 +
14 +```bash
15 +# Run Gemini CLI directly from GitHub
16 +npx https://github.com/google-gemini/gemini-cli
17 +
18 +# Or clone and install locally
19 +git clone https://github.com/google-gemini/gemini-cli.git
20 +cd gemini-cli
21 +npm install
22 +npm run build
23 +```
24 +
25 +## Configuration
26 +
27 +Gemini CLI has built-in MCP server support. For detailed MCP configuration, see the [official MCP documentation](https://github.com/google-gemini/gemini-cli/blob/main/docs/tools/mcp-server.md).
28 +
29 +### Adding Netdata MCP Server
30 +
31 +Configure your Gemini settings to include the Netdata MCP server:
32 +
33 +```bash
34 +# Edit Gemini settings file
35 +~/.gemini/settings.json
36 +```
37 +
38 +Add your Netdata MCP server configuration:
39 +
40 +```json
41 +{
42 + "mcpServers": {
43 + "netdata": {
44 + "command": "/usr/sbin/nd-mcp",
45 + "args": ["ws://YOUR_NETDATA_IP:19999/mcp?api_key=NETDATA_MCP_API_KEY"]
46 + }
47 + }
48 +}
49 +```
50 +
51 +### Verify MCP Configuration
52 +
53 +Use the `/mcp` command to verify your setup:
54 +
55 +```bash
56 +# List configured MCP servers
57 +/mcp
58 +
59 +# Show detailed descriptions of MCP servers and tools
60 +/mcp desc
61 +
62 +# Show MCP server schema details
63 +/mcp schema
64 +```
65 +
66 +Replace:
67 +
68 +- `/usr/sbin/nd-mcp` - With your [actual nd-mcp path](/docs/learn/mcp.md#finding-the-nd-mcp-bridge)
69 +- `YOUR_NETDATA_IP` - IP address or hostname of your Netdata Agent/Parent
70 +- `NETDATA_MCP_API_KEY` - Your [Netdata MCP API key](/docs/learn/mcp.md#finding-your-api-key)
71 +
72 +## How to Use
73 +
74 +Gemini CLI can leverage Netdata's observability data for infrastructure analysis and automation:
75 +
76 +```
77 +What's the current system performance across all monitored servers?
78 +Show me any performance anomalies in the last 2 hours
79 +Which services are consuming the most resources right now?
80 +Analyze the database performance trends over the past week
81 +```
82 +
83 +## Example Workflows
84 +
85 +**Performance Investigation:**
86 +
87 +```
88 +Investigate why our application response times increased this afternoon
89 +```
90 +
91 +**Resource Optimization:**
92 +
93 +```
94 +Check memory usage patterns and suggest optimization strategies
95 +```
96 +
97 +**Alert Analysis:**
98 +
99 +```
100 +Explain the current active alerts and their potential impact
101 +```
102 +
103 +> **💡 Advanced Usage:** Gemini CLI can combine observability data with system automation for powerful DevOps workflows. Learn about the opportunities and security considerations in [AI DevOps Copilot](/docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md).
104 +
105 +## Troubleshooting
106 +
107 +### MCP Connection Issues
108 +
109 +- Verify Netdata is accessible: `curl http://YOUR_NETDATA_IP:19999/api/v3/info`
110 +- Check that the bridge path exists and is executable
111 +- Ensure API key is correct and properly formatted
112 +
113 +### Limited Data Access
114 +
115 +- Verify API key is included in the connection string
116 +- Check that the Netdata agent is properly configured for MCP
117 +- Ensure network connectivity between Gemini CLI and Netdata
118 +
119 +### Command Execution Problems
120 +
121 +- Review command syntax for your specific Gemini CLI version
122 +- Check MCP server configuration parameters
123 +- Verify that MCP protocol is supported in your Gemini CLI installation
124 +
125 +## Documentation Links
126 +
127 +- [Gemini CLI GitHub Repository](https://github.com/google-gemini/gemini-cli)
128 +- [Gemini CLI Official Documentation](https://developers.google.com/gemini-code-assist/docs/gemini-cli)
129 +- [Netdata MCP Setup](/docs/learn/mcp.md)
130 +- [AI DevOps Best Practices](/docs/ml-ai/ai-devops-copilot/ai-devops-copilot.md)
docs/ml-ai/ai-insights.md new
+227
@@ -0,0 +1,227 @@
1 +# AI Insights
2 +
3 +**From hours of debugging to minutes of clarity** - AI Insights transforms your infrastructure monitoring data into professional reports that explain what happened, why it happened, and what to do about it.
4 +
5 +## The Challenge AI Insights Solves
6 +
7 +Traditional monitoring requires you to manually query metrics, correlate data, and build dashboards during incidents - all while the clock is ticking. Even experienced engineers struggle with:
8 +
9 +- Learning complex query languages (PromQL, SQL) just to ask basic questions
10 +- Building custom dashboards during incidents instead of fixing problems
11 +- Correlating metrics across multiple systems to find root causes
12 +- Translating technical metrics into business impact for stakeholders
13 +- Spending hours on post-incident analysis and reporting
14 +
15 +**AI Insights eliminates these barriers** by automatically analyzing your infrastructure and delivering comprehensive reports that provide both executive summaries and technical deep-dives.
16 +
17 +## Why AI Insights Transforms Operations
18 +
19 +- **No query languages needed** - Skip the learning curve of PromQL, SQL, or custom dashboards
20 +- **AI with SRE expertise** - Get analysis from an AI trained to think like a senior engineer
21 +- **Root cause, not symptoms** - Understand the cascade of issues, not just surface metrics
22 +- **Business context included** - Reports explain technical issues in terms of business impact
23 +- **Collaborative by design** - Share professional PDFs with stakeholders who need answers, not dashboards
24 +- **Powered by Netdata's ML** - Leverages anomaly scores from ML models trained on every metric
25 +- **Zero configuration needed** - Works immediately with your existing Netdata deployment
26 +
27 +## Four Specialized Report Types
28 +
29 +![AI Insights Report Example](https://github.com/user-attachments/assets/c6997afb-94cb-41cc-a038-b384cb92e751)
30 +
31 +### Infrastructure Summary
32 +
33 +**Your automated health check and incident analyst**
34 +
35 +Perfect for Monday morning reviews, post-incident analysis, or executive updates. This report provides:
36 +
37 +- Complete system health assessment with prioritized issues
38 +- Timeline of incidents and their business impact
39 +- Critical alerts analysis with resolution recommendations
40 +- Top 3 actionable items to improve infrastructure health
41 +- Performance trends across all key metrics
42 +
43 +**Use cases**: Weekend incident recovery, executive briefings, team handoffs, regular health checks
44 +
45 +### Capacity Planning
46 +
47 +**Stop guessing future needs - get data-driven projections**
48 +
49 +Make informed decisions about infrastructure investments with reports that include:
50 +
51 +- Resource utilization trends and growth patterns
52 +- Predicted capacity exhaustion dates for critical resources
53 +- Specific hardware recommendations based on usage patterns
54 +- Cost optimization opportunities
55 +- Projections for 3 months to 2 years ahead
56 +
57 +**Use cases**: Quarterly planning, budget justification, infrastructure roadmaps, vendor negotiations
58 +
59 +### Performance Optimization
60 +
61 +**Find and fix bottlenecks before users complain**
62 +
63 +Identify inefficiencies and optimization opportunities with:
64 +
65 +- Bottleneck analysis across application, database, network, and storage
66 +- Resource contention patterns and their impact
67 +- Specific tuning recommendations with expected improvements
68 +- Prioritized list of optimizations by potential impact
69 +- Before/after projections for recommended changes
70 +
71 +**Use cases**: Performance audits, system tuning, SRE optimization projects, efficiency improvements
72 +
73 +### Anomaly Analysis
74 +
75 +**Post-incident forensics made simple**
76 +
77 +Understand unusual patterns and prevent future issues with:
78 +
79 +- ML-detected anomalies with severity scoring
80 +- Root cause analysis showing how issues cascaded
81 +- Timeline reconstruction of anomaly propagation
82 +- Correlation between different system anomalies
83 +- Recommendations to prevent recurrence
84 +
85 +**Use cases**: Post-mortems, proactive issue detection, system behavior analysis, troubleshooting
86 +
87 +## Customize Reports to Your Needs
88 +
89 +Each report type offers flexible customization options for content and analysis scope (note: report structure and visual style are standardized for consistency):
90 +
91 +### Time Period Selection
92 +
93 +- **Infrastructure Summary**: Last 24 hours, 48 hours, 7 days, or month
94 +- **Capacity Planning**: Forecast for 3 months, 6 months, 1 year, or 2 years
95 +- **Performance Optimization**: Last 24 hours, 7 days, month, or quarter
96 +- **Anomaly Analysis**: Last 6 hours, 12 hours, 24 hours, or 7 days
97 +
98 +### Scope and Filtering
99 +
100 +- **Node Selection**: Analyze specific servers or your entire infrastructure
101 +- **Metric Categories**: Focus on CPU, Memory, Disk, Network, or Applications
102 +- **Resource Types**: Target Compute, Storage, Network, or Database resources
103 +- **Focus Areas**: Drill into specific performance domains
104 +- **Anomaly Thresholds**: Set sensitivity levels (10%, 20%, or 30%)
105 +
106 +## How AI Insights Works
107 +
108 +### 1. Intelligent Data Collection
109 +
110 +When you request a report, AI Insights:
111 +
112 +- Gathers relevant metrics from your selected time period and nodes
113 +- Collects active alerts and their severity levels
114 +- Retrieves ML-detected anomalies and their scores
115 +- Maps system relationships and dependencies
116 +- Compiles process and application performance data
117 +
118 +### 2. AI-Powered Analysis
119 +
120 +The collected data is analyzed by Anthropic's Claude 3.7 Sonnet model, optimized for infrastructure telemetry analysis using SRE methodologies. This AI model:
121 +
122 +- Applies SRE-level expertise to identify patterns
123 +- Correlates issues across different systems
124 +- Determines root causes vs symptoms
125 +- Prioritizes findings by business impact
126 +- Generates actionable recommendations
127 +
128 +### 3. Professional Report Generation
129 +
130 +Within 2-3 minutes, you receive:
131 +
132 +- **Structured content**: Headers, insights, charts, and tables in logical flow
133 +- **Embedded visualizations**: Charts generated from your actual metrics
134 +- **Executive summary**: High-level findings for stakeholders
135 +- **Technical details**: Deep-dive analysis for engineers
136 +- **Action items**: Prioritized recommendations with clear next steps
137 +- **PDF format**: Professional reports ready for sharing
138 +
139 +### 4. Security and Privacy
140 +
141 +- **In-memory processing**: Data analyzed then immediately discarded
142 +- **No training data**: Your infrastructure data is never used for model training
143 +- **Secure API**: All communications encrypted end-to-end
144 +- **Access controlled**: Respects your existing Netdata permissions
145 +
146 +## Real-World Impact
147 +
148 +From the Inrento fintech case study:
149 +> "AI Insights provided **significant time savings** in identifying and resolving issues. It **drastically reduced the time spent** identifying problems and implementing solutions, leading to **enhanced productivity and performance** with **minimized downtime**."
150 +Teams report that incident analysis that previously took hours of manual investigation now completes in minutes with AI Insights.
151 +
152 +## Perfect For
153 +
154 +- **Incident post-mortems**: Generate comprehensive analysis in minutes, not hours
155 +- **Executive briefings**: Professional PDFs with clear summaries and visualizations
156 +- **Capacity reviews**: Data-driven planning for budget and resource allocation
157 +- **Performance audits**: Regular health checks without manual analysis
158 +- **Team handoffs**: Share context-rich reports instead of dashboard links
159 +- **Compliance reporting**: Document infrastructure state and changes
160 +- **Vendor discussions**: Data-backed evidence for infrastructure decisions
161 +
162 +## Unlike Traditional Monitoring
163 +
164 +AI Insights represents a paradigm shift in infrastructure monitoring:
165 +
166 +| Traditional Monitoring | AI Insights |
167 +|------------------------|-------------|
168 +| Build dashboards during incidents | Get instant analysis |
169 +| Learn query languages | Use natural language selection |
170 +| Manual correlation across metrics | Automatic relationship detection |
171 +| Raw metrics without context | Narrative explanations with context |
172 +| Technical data only | Business impact included |
173 +| Hours of manual analysis | 2-3 minute automated reports |
174 +
175 +## What Sets AI Insights Apart
176 +
177 +Unlike traditional AI monitoring assistants that require extensive configuration or operate as black-box cloud services, AI Insights:
178 +
179 +- **Runs entirely on your infrastructure** - No external dependencies or mysterious cloud processing
180 +- **Uses your actual data** - Not generic patterns or industry averages
181 +- **Provides transparent analysis** - Clear reasoning, not black-box decisions
182 +- **Respects your security** - Data never leaves your control
183 +- **Works instantly** - No training period or configuration required
184 +
185 +## Getting Started
186 +
187 +1. **Access AI Insights** from the Netdata Cloud navigation menu
188 +2. **Select a report type** based on your current need
189 +3. **Customize parameters** like time period and node selection
190 +4. **Generate report** and receive it within 2-3 minutes
191 +5. **Share or download** the PDF for stakeholders
192 +
193 +## Technical Requirements
194 +
195 +- Active Netdata Cloud account
196 +- At least one connected Netdata Agent
197 +- Historical data (minimum 24 hours recommended)
198 +- No additional configuration needed
199 +
200 +## Frequently Asked Questions
201 +
202 +**Q: How far back can AI Insights analyze data?**
203 +A: AI Insights can analyze any data retained by your Netdata agents, from 6 hours to 2 years depending on the report type and your retention settings.
204 +
205 +**Q: Can I schedule regular reports?**
206 +A: Currently reports are generated on-demand. Scheduled reports are on the roadmap.
207 +
208 +**Q: What metrics are included in the analysis?**
209 +A: AI Insights analyzes all metrics collected by your Netdata agents, including system metrics, application metrics, and custom collectors.
210 +
211 +**Q: How does it handle sensitive data?**
212 +A: All data is processed securely and discarded after report generation. No data is stored or used for training.
213 +
214 +**Q: Can I customize the report format?**
215 +A: Report structure and visual style are standardized for consistency and professional presentation. However, you have extensive control over the analysis scope, time periods, metrics, and focus areas through customization parameters.
216 +
217 +## What's Next
218 +
219 +AI Insights continues to evolve with new capabilities planned:
220 +
221 +- Scheduled report generation
222 +- Custom report templates
223 +- API access for automation
224 +- Integration with ticketing systems
225 +- Comparative analysis between time periods
226 +
227 +Experience the future of infrastructure monitoring - transform your data into intelligence with AI Insights.
docs/ml-ai/anomaly-advisor.md new
+86
@@ -0,0 +1,86 @@
1 +# Anomaly Advisor
2 +
3 +The Anomaly Advisor (the "Anomalies" tab on Netdata dashboards) is a troubleshooting assistant that correlates anomalies across your entire infrastructure and presents them as a ranked list of metrics, sorted by anomaly severity.
4 +
5 +Built on three components: per-metric anomaly detection using k-means clustering (18 models per metric), pre-computed Node Anomaly Rate (NAR) correlation charts, and a specialized scoring engine that evaluates thousands of metrics simultaneously. When you highlight an incident timeframe, the scoring engine analyzes all metrics and returns an ordered list - typically placing root causes within the top 30-50 results.
6 +
7 +The system works by detecting anomaly clusters. A single event (like an SSH login, a backup job, or a service restart) typically triggers anomalies across dozens of related metrics - CPU, memory, network, disk I/O, and application-specific counters. The Anomaly Advisor captures these correlations and ranks them by severity, effectively showing you what changed and how those changes cascaded through your infrastructure.
8 +
9 +This approach inverts traditional troubleshooting. Instead of forming hypotheses and validating them one by one, you start with data - a ranked list of what actually deviated from normal patterns. The tool works without requiring system-specific knowledge, though interpreting results still requires engineering expertise.
10 +
11 +Limitations: Works best for sudden changes and patterns not seen in the last 54 hours. Cannot detect anomalies in stopped services (no data = no anomalies). May miss gradually evolving issues where each increment appears normal.
12 +
13 +## System Characteristics
14 +
15 +| Aspect | Implementation | Operational Benefit |
16 +|--------|----------------|-------------------|
17 +| **Data Source** | Per-metric anomaly detection using k-means (k=2) with 18-model consensus | Comprehensive coverage, no blind spots |
18 +| **Correlation Engine** | Pre-computed Node Anomaly Rate (NAR) charts updated in real-time | Instant blast radius visualization |
19 +| **Query Engine** | Specialized scoring engine evaluating thousands of metrics simultaneously | Returns ranked list, not time-series data |
20 +| **Ranking Algorithm** | Anomaly severity scoring across selected time window | Root cause typically in top 30-50 results |
21 +| **Infrastructure View** | Dual charts: % anomalous and absolute count per node | Distinguishes small node spikes from large node issues |
22 +| **Time to Insight** | Highlight timeframe → ranked results in seconds |Minutes to root cause vs hours of hypothesis testing |
23 +| **Expertise Required** | No system-specific knowledge needed to identify anomalies | Minimal expertise to interpret results |
24 +| **Dependency Discovery** | Correlated anomalies reveal component relationships | Exposes hidden infrastructure dependencies |
25 +| **Best Use Cases** | Sudden changes, cascading failures, multi-node incidents | Excellent for "what just happened?" scenarios |
26 +| **Limitations** | Cannot detect stopped services, gradual degradation | Not a replacement for all monitoring |
27 +
28 +:::warning Limitations
29 +
30 +- **Stopped services**: No data = no anomalies detected
31 +- **Gradual degradation**: Changes within 54-hour training window may appear normal
32 +- **Pattern fragments**: If anomaly patterns existed separately in training data, consensus may not trigger
33 +:::
34 +
35 +## Visualizing Cascading Infrastructure Level Effects
36 +
37 +The Anomaly Advisor provides two views of node-level anomalies, revealing how issues propagate across infrastructure. Each node-level chart aggregates the underlying anomaly bits calculated per metric (as described in [Machine Learning Anomaly Detection](/docs/ml-ai/ml-anomaly-detection/ml-anomaly-detection.md)):
38 +
39 +![Anomaly cascading effects visualization](https://github.com/user-attachments/assets/a69b1461-c559-4b22-bb02-045670d84168)
40 +
41 +This visualization shows two distinct anomaly clusters:
42 +
43 +**First cluster (left):**
44 +
45 +- Shows clear propagation: one node spikes first, followed by three more nodes in sequence
46 +- Each subsequent node shows anomalies shortly after the previous one
47 +- Classic cascading pattern where an issue on one node impacts dependent nodes
48 +
49 +**Second cluster (right):**
50 +
51 +- Multiple nodes become anomalous simultaneously
52 +- The final node shows the largest spike (200+ anomalous metrics in absolute count)
53 +- Indicates either a shared resource issue or the final node being the aggregation point
54 +
55 +The two charts provide different perspectives:
56 +
57 +1. **Top chart** - Percentage of anomalous metrics per node (spikes up to 10%)
58 +2. **Bottom chart** - Absolute count of anomalous metrics per node (spikes up to 200+)
59 +
60 +This dual view helps distinguish between:
61 +
62 +- **Small nodes with high anomaly rates** (high percentage, low count)
63 +- **Large nodes with many anomalies** (lower percentage, high count)
64 +
65 +**This visualization provides the infrastructure-level blast radius of any incident.** At a glance, you can see:
66 +
67 +- Which nodes were affected
68 +- When each node was impacted
69 +- The severity of impact on each node
70 +- Whether the issue propagated sequentially or hit multiple nodes simultaneously
71 +
72 +By highlighting any spike or cluster (click and drag on the timeline), the scoring engine analyzes all metrics from all affected nodes during that period. The engine ranks metrics by their anomaly severity within the selected timeframe, returning an ordered list that typically reveals the root cause within the top 30-50 results.
73 +
74 +## How to Use It
75 +
76 +1. **Click the Anomalies tab** in any dashboard
77 +2. **Highlight the incident time window** (click and drag on any chart)
78 +3. **Review the ranked list** of anomalous metrics
79 +4. **Root cause usually surfaces** in top 30-50 metrics
80 +
81 +## Learn More
82 +
83 +For detailed information about using the Anomalies tab, see:
84 +
85 +- [Anomalies Tab Documentation](/docs/dashboards-and-charts/anomaly-advisor-tab.md)
86 +- [Machine Learning Anomaly Detection](/docs/ml-ai/ml-anomaly-detection/ml-anomaly-detection.md) - The foundation powering the Anomaly Advisor
docs/ml-ai/ml-anomaly-detection/ml-anomaly-detection.md new
+260
@@ -0,0 +1,260 @@
1 +# Machine Learning Anomaly Detection
2 +
3 +Netdata uses k-means clustering to detect anomalies for each collected metric automatically.
4 +
5 +The system maintains 18 models per metric, each trained on 6-hour windows at 3-hour intervals, providing approximately 54 hours of rolling behavioral patterns. Anomaly detection occurs in real-time during data collection - a data point is flagged as anomalous only when all 18 models reach consensus, effectively eliminating noise while maintaining sensitivity to genuine issues.
6 +
7 +Anomaly bits are stored alongside metric data in the time-series database, with the same retention period. The query engine calculates anomaly rates dynamically during data aggregation, exposing anomaly information on every chart without additional overhead.
8 +
9 +A dedicated process correlates anomalies across all metrics within each node, generating real-time node-level anomaly charts. This correlation data feeds into Netdata's scoring engine - a specialized query system that can evaluate thousands of metrics simultaneously and return an ordered list ranked by anomaly severity, powering the Anomaly Advisor for rapid root cause analysis.
10 +
11 +## System Characteristics
12 +
13 +| Aspect | Implementation | Benefit |
14 +|--------|----------------|---------|
15 +| **Algorithm** | Unsupervised k-means clustering (k=2) via [dlib](https://github.com/davisking/dlib) | No manual training or labeled data required |
16 +| **Model Architecture** | rolling 18 models per metric, 3-hour staggered training | Eliminates 99% of false positives through consensus |
17 +| **Processing Location** | Edge computation on each Netdata agent | No cloud dependency, no data egress |
18 +| **Resource Usage** | ~18KB RAM per metric, 2-4% of a single CPU for 10k metrics | Predictable linear scaling |
19 +| **Configuration** | Zero-configuration with automatic adaptation | Works instantly on any metric type |
20 +| **Detection Latency** | Real-time during data collection | Anomalies flagged within 1 second |
21 +| **Historical Storage** | Anomaly bit embedded in metric storage | No additional storage overhead |
22 +| **Query Performance** | On-the-fly anomaly rate calculation | No pre-aggregation needed |
23 +| **Time-series Integrity** | Immutable anomaly history | No hindsight bias - shows what was detectable THEN |
24 +| **Coverage** | Every metric, every dimension | No sampling, no blind spots |
25 +| **Correlation Engine** | Real-time anomaly correlation across metrics | Powers Anomaly Advisor for root cause analysis |
26 +| **Alert Philosophy** | Investigation aid, not alert source | Reduces alert fatigue |
27 +
28 +:::note
29 +Netdata avoids deep learning models to maintain lightweight operation on any Linux system. The entire ML system is designed to run efficiently without specialized hardware or dependencies.
30 +:::
31 +
32 +## Types of Anomalies Detected
33 +
34 +| Anomaly Type | Description | Business Impact |
35 +|--------------------------|-------------------------------------------------------------------|------------------------------------------|
36 +| **Point Anomalies** | Unusually high or low values compared to historical data | Early warning of service degradation |
37 +| **Contextual Anomalies** | Sequences of values that deviate from expected patterns | Identification of unusual usage patterns |
38 +| **Collective Anomalies** | Multivariate anomalies where a combination of metrics appears off | Detection of complex system issues |
39 +| **Concept Drifts** | Gradual shifts leading to a new baseline | Recognition of evolving system behavior |
40 +| **Change Points** | Sudden shifts resulting in a new normal state | Identification of system changes |
41 +
42 +## Technical Deep Dive: How Netdata ML Works
43 +
44 +```mermaid
45 +flowchart TD
46 + Raw("**Raw Metrics**<br/><br/>Last 6 Hours")
47 + Preprocess("**Preprocess**<br/><br/>Feature Vectors")
48 + Train("**Train k-means<br/><br/>k=2**")
49 + Model("**Trained Model**")
50 +
51 + M1("Model 1<br/><br/>**Recent Data**")
52 + M2("Model 2<br/><br/>**Older Data**")
53 + M3("Model 3<br/><br/>**Even Older Data**")
54 + MN("Model N<br/><br/>**Up to 54 Hours Old**")
55 +
56 + NewData("**New Metrics**")
57 + DistCalc("**Calculate**<br/><br/>Euclidean Distance<br/>to Cluster Centers")
58 + Threshold("**Distance > 99th**<br/><br/>Percentile?")
59 + FlagA("**Flag as Anomalous**<br/><br/>in This Model")
60 + FlagN("**Flag as Normal**<br/><br/>in This Model")
61 +
62 + AllResults("**Results from All Models**")
63 + AllAgree("**All Models<br/><br/>Agree it's<br/><br/>Anomalous?**")
64 + SetBit("**Set Anomaly Bit = 1**<br/><br/>True")
65 + ClearBit("**Set Anomaly Bit = 0**<br/><br/>False")
66 +
67 + Raw --> Preprocess
68 + Preprocess --> Train
69 + Train --> Model
70 + Model --> M1
71 + Model --> M2
72 + Model --> M3
73 + Model --> MN
74 +
75 + M1 --> NewData
76 + M2 --> NewData
77 + M3 --> NewData
78 + MN --> NewData
79 +
80 + NewData --> DistCalc
81 + DistCalc --> Threshold
82 + Threshold -->|Yes| FlagA
83 + Threshold -->|No| FlagN
84 +
85 + FlagA --> AllResults
86 + FlagN --> ClearBit
87 + AllResults --> AllAgree
88 + AllAgree -->|Yes| SetBit
89 + AllAgree -->|No| ClearBit
90 +
91 + %% Style definitions
92 + classDef neutral fill:#f9f9f9,stroke:#000000,stroke-width:3px,color:#000000,font-size:16px
93 + classDef process fill:#ffeb3b,stroke:#000000,stroke-width:3px,color:#000000,font-size:16px
94 + classDef complete fill:#4caf50,stroke:#000000,stroke-width:3px,color:#000000,font-size:16px
95 + classDef anomaly fill:#f44336,stroke:#000000,stroke-width:3px,color:#000000,font-size:16px
96 +
97 + %% Apply styles
98 + class Raw,Preprocess,NewData,AllResults neutral
99 + class Train,Model,M1,M2,M3,MN,DistCalc,Threshold,AllAgree process
100 + class FlagN,ClearBit complete
101 + class FlagA,SetBit anomaly
102 +```
103 +
104 +### Training & Detection Process
105 +
106 +When you enable ML, Netdata trains an unsupervised model for each of your metrics. By default, this model is a [k-means clustering](https://en.wikipedia.org/wiki/K-means_clustering) algorithm (with k=2) trained on the last 6 hours of your data. Instead of just analyzing raw values, the model works with preprocessed feature vectors to improve your detection accuracy.
107 +
108 +:::important
109 +To reduce false positives in your environment, Netdata trains multiple models per time-series, covering over two days of data. **An anomaly is flagged only if all models agree on it, eliminating 99% of false positives**. This approach of requiring consensus across models trained on different time scales makes the system highly resistant to spurious anomalies while still being sensitive to real issues.
110 +:::
111 +
112 +The anomaly detection algorithm uses the [Euclidean distance](https://en.wikipedia.org/wiki/Euclidean_distance) between recent metric patterns and the learned cluster centers. If this distance exceeds a threshold based on the 99th percentile of training data, that model considers the metric anomalous.
113 +
114 +### The Anomaly Bit
115 +
116 +Each trained model assigns an **anomaly score** at every time step based on how far your data deviates from learned clusters. If the score exceeds the 99th percentile of training data, the **anomaly bit** is set to `true` (100); otherwise, it remains `false` (0).
117 +
118 +**Key benefits you'll experience:**
119 +
120 +- No additional storage overhead since the anomaly bit is embedded in Netdata's floating point number format
121 +- The query engine automatically computes anomaly rates without requiring extra queries
122 +
123 +:::note
124 +The anomaly bit is quite literally a bit in Netdata's [internal storage representation](https://github.com/netdata/netdata/blob/89f22f056ca2aae5d143da9a4e94fcab1f7ee1b8/libnetdata/storage_number/storage_number.c#L83). This ingenious design means that for every metric collected, Netdata can also track whether it's anomalous without increasing storage requirements.
125 +:::
126 +
127 +You can access the anomaly bits through Netdata's API by adding the `options=anomaly-bit` parameter to your query. For example:
128 +
129 +```
130 +https://your-node/api/v3/data?chart=system.cpu&dimensions=user&after=-10&options=anomaly-bit
131 +```
132 +
133 +This would return anomaly bits for the last 10 seconds of CPU user data, with values of either 0 (normal) or 100 (anomalous).
134 +
135 +### Anomaly Rate Calculations
136 +
137 +You can see **Node Anomaly Rate (NAR)** and **Dimension Anomaly Rate (DAR)** calculated based on anomaly bits. Here's an example matrix:
138 +
139 +| Time | d1 | d2 | d3 | d4 | d5 | **NAR** |
140 +|---------|---------|---------|---------|---------|---------|-----------------------|
141 +| t1 | 0 | 0 | 0 | 0 | 0 | **0%** |
142 +| t2 | 0 | 0 | 0 | 0 | 100 | **20%** |
143 +| t3 | 0 | 0 | 0 | 0 | 0 | **0%** |
144 +| t4 | 0 | 100 | 0 | 0 | 0 | **20%** |
145 +| t5 | 100 | 0 | 0 | 0 | 0 | **20%** |
146 +| t6 | 0 | 100 | 100 | 0 | 100 | **60%** |
147 +| t7 | 0 | 100 | 0 | 100 | 0 | **40%** |
148 +| t8 | 0 | 0 | 0 | 0 | 100 | **20%** |
149 +| t9 | 0 | 0 | 100 | 100 | 0 | **40%** |
150 +| t10 | 0 | 0 | 0 | 0 | 0 | **0%** |
151 +| **DAR** | **10%** | **30%** | **20%** | **20%** | **30%** | **_NAR_t1-10 = 22%_** |
152 +
153 +- **DAR (Dimension Anomaly Rate):** Average anomalies for a specific metric over time
154 +- **NAR (Node Anomaly Rate):** Average anomalies across all metrics at a given time
155 +- **Overall anomaly rate:** Computed across your entire dataset for deeper insights
156 +
157 +### Node-Level Anomaly Detection
158 +
159 +Netdata tracks the percentage of anomaly bits over time for you. When the **Node Anomaly Rate (NAR)** exceeds a set threshold and remains high for a period, a **node anomaly event** is triggered. These events are recorded in the `new_anomaly_event` dimension on the `anomaly_detection.anomaly_detection` chart.
160 +
161 +## Available Documentation
162 +
163 +- **[ML Configuration](/src/ml/ml-configuration.md)** - Configuration and tuning guide
164 +- **[Metric Correlations](/docs/metric-correlations.md)** - Finding related metrics during incidents
165 +
166 +## Viewing Anomaly Data in Your Netdata Dashboard
167 +
168 +Once you enable ML, you'll have access to an **Anomaly Detection** menu with key charts:
169 +
170 +- **`anomaly_detection.dimensions`**: Number of dimensions flagged as anomalous
171 +- **`anomaly_detection.anomaly_rate`**: Percentage of anomalous dimensions
172 +- **`anomaly_detection.anomaly_detection`**: Flags (0 or 1) indicating when an anomaly event occurs
173 +
174 +These insights help you quickly assess potential issues and take action before they escalate.
175 +
176 +## Operational Details
177 +
178 +### Why 18 Models?
179 +
180 +The number 18 balances three competing requirements:
181 +
182 +1. **Incremental learning efficiency** - Training 48 hours of data every 3 hours would waste computational resources. Instead, each model trains on just 6 hours of data, with only one new model created every 3 hours.
183 +
184 +2. **Adaptive memory duration** - When an anomaly occurs, the newest model will learn it as "normal" within 3 hours. The system gradually "forgets" this pattern as older models are replaced. With 18 models at 3-hour intervals, complete forgetting takes 54 hours (2.25 days).
185 +
186 +3. **Consensus noise reduction** - Multiple models voting together eliminate random fluctuations. 18 models provide strong consensus without excessive memory use.
187 +
188 +This creates a sliding window memory: recent anomalies become "normal" quickly (within 3 hours for the newest model), while the full consensus takes 54 hours to completely forget an anomalous pattern. This balance prevents both alert fatigue from repeated anomalies and blindness to recurring issues.
189 +
190 +### How Netdata Minimizes Training CPU Impact
191 +
192 +ML typically doubles the agent's CPU usage - from ~2% to ~4% of a single core. This efficiency comes from several optimizations:
193 +
194 +1. **Smart metric filtering** - Metrics with constant or fixed values are automatically excluded from training, eliminating wasted computation on unchanging data.
195 +
196 +2. **Incremental training windows** - Each model trains on only 6 hours of data instead of the full 54-hour history, reducing computational requirements by ~90%.
197 +
198 +3. **Even training distribution** - The agent dynamically throttles model training to spread the work evenly across each 3-hour window, preventing CPU spikes. With 10,000 metrics, this means training ~1 model per second instead of training 10,000 models in a burst.
199 +
200 +4. **Distributed intelligence** - Child agents stream both trained models and anomaly bits to parent agents along with metric data. Parents receive pre-computed ML results, requiring zero additional ML computation for aggregated views.
201 +
202 +This design ensures ML remains lightweight enough to run on production systems without impacting primary workloads.
203 +
204 +**Dynamic prioritization**: ML automatically throttles or even pauses training during:
205 +
206 +- Heavy query load - ensuring dashboards remain responsive
207 +- Parent-child reconnections - prioritizing metric replication
208 +- Any resource contention - backing off to protect core monitoring
209 +
210 +Under these conditions, ML will completely stop training new models to ensure:
211 +
212 +- User queries remain fast and responsive
213 +- Metric streaming completes quickly after network interruptions
214 +- Overall CPU and I/O consumption stays within bounds
215 +
216 +This means ML is truly a background process - it uses spare cycles but immediately yields resources when needed for operational tasks.
217 +
218 +### Storage Impact
219 +
220 +ML has **zero storage overhead** in the time-series database. The anomaly bit uses a previously unused bit in the existing sample storage format - no schema changes or storage expansion required.
221 +
222 +The only storage impact comes from persisting trained models to disk for survival across restarts:
223 +
224 +- Model files are small compared to the time-series data
225 +- Negligible impact on overall storage requirements
226 +- Models are retained only for active metrics
227 +
228 +This means you can enable ML without provisioning additional storage capacity. Anomaly history is retained for the same period as your metrics, with no extra space required.
229 +
230 +**Query performance impact: None**. The anomaly bit is loaded together with metric data in a single disk read - no additional I/O operations required. Querying metrics with anomaly data has the same disk I/O pattern as querying metrics without ML.
231 +
232 +### Cold Start Behavior
233 +
234 +On a freshly installed agent, ML begins detecting anomalies within 10 minutes. However, early detection has important characteristics:
235 +
236 +**Timeline:**
237 +
238 +- **0-10 minutes**: Collecting initial data, no anomaly detection
239 +- **10+ minutes**: First models trained, anomaly detection begins with high sensitivity
240 +- **3 hours**: First model rotation, improved accuracy
241 +- **54 hours**: Full model set established, optimal detection accuracy
242 +
243 +**What to expect:**
244 +
245 +- Initial hours show more anomalies due to limited training data
246 +- False positive rate decreases as models accumulate more behavioral patterns
247 +- Each 3-hour cycle improves detection quality
248 +- After 2-3 days, the system reaches steady-state accuracy
249 +
250 +**Operational tip**: During the first 48 hours after deployment, expect elevated anomaly rates. This is normal as the system learns your infrastructure's patterns. Use this period to observe ML behavior but avoid making critical decisions based solely on early anomaly detection.
251 +
252 +## Getting Started
253 +
254 +ML is enabled by default in recent Netdata versions. To use anomaly detection:
255 +
256 +1. **View anomaly ribbons** - Purple overlays on all charts show anomaly rates
257 +2. **Access Anomaly Advisor** - Click the Anomalies tab for guided troubleshooting
258 +3. **Query historical anomalies** - Use the query engine to analyze past incidents
259 +
260 +[Learn more about the Anomaly Advisor →](/docs/ml-ai/anomaly-advisor.md)
docs/netdata-assistant.md deleted
-196
@@ -1,196 +0,0 @@
1 -# Netdata AI: Alert Assistant & Infrastructure Insights
2 -
3 -**Netdata AI provides intelligent assistance for both immediate alert response and strategic infrastructure analysis** using advanced AI to help you understand incidents quickly and synthesize high-resolution metrics into actionable intelligence.
4 -
5 -This comprehensive system combines **real-time alert assistance** for emergency troubleshooting with **strategic infrastructure insights** for long-term planning, helping you both respond to immediate incidents and make informed decisions about your infrastructure's future.
6 -
7 -## Two Complementary Approaches
8 -
9 -Netdata Insights serves different moments in your engineering workflow through two distinct but complementary capabilities:
10 -
11 -| Aspect | **Real-time Alert Assistant** | **Strategic Insights Reports** |
12 -|----------------------|-----------------------------------|--------------------------------------------------------------------------------------------------|
13 -| **Primary Use** | **Immediate incident response** | Strategic planning & analysis |
14 -| **When You Use It** | During active alerts | Post-incident, planning sessions |
15 -| **Mindset** | "The building is on fire" | "Let's understand and plan better" |
16 -| **Response Time** | **Instant contextual help** | **2-3 minutes for comprehensive analysis** |
17 -| **Scope** | Single alert or immediate issue | Infrastructure-wide trends and patterns |
18 -| **Output** | Quick explanations and next steps | Detailed reports with embedded visualizations, **downloadable as PDFs**, **shareable via email** |
19 -| **Typical Scenario** | 3 AM emergency response | Monday morning incident review |
20 -
21 -:::tip
22 -
23 -These serve fundamentally different moments in an engineer's workflow. The **Assistant** is for high-stress situations when you need immediate context, while **Insights Reports** are for when you have time to think strategically about your infrastructure's health and future needs.
24 -
25 -:::
26 -
27 -:::note
28 -
29 -**Netdata Insights is currently in beta as a research preview:**
30 -
31 -- Available in Netdata Cloud for Business users and Free Trial participants
32 -- Works with any infrastructure where you've deployed Netdata agents
33 -- No additional configuration or new pipelines required
34 -- Everyone gets 10 reports to generate for free
35 -- Community users can get early access via Discord or email to product@netdata.cloud
36 -
37 -:::
38 -
39 -## Real-time Alert Assistant
40 -
41 -**Get immediate context and guidance when alerts fire** - exactly when you need it most, especially during critical situations.
42 -
43 -The Assistant provides instant explanations and troubleshooting steps directly within your alert workflow, helping you understand what's happening without leaving the Netdata interface.
44 -
45 -| Feature | Benefit |
46 -|---------------------------|-----------------------------------------------------------------------------------------------------------------------|
47 -| **Follows Your Workflow** | The Assistant window stays with you as you navigate through Netdata dashboards during your troubleshooting process. |
48 -| **Works at Any Hour** | Especially valuable during after-hours emergencies when you might not have team support available. |
49 -| **Contextual Knowledge** | Combines Netdata's community expertise with the power of large language models to provide relevant advice. |
50 -| **Time-Saving** | Eliminates the need for searches across multiple documentation sources or community forums. |
51 -| **Non-Intrusive** | Provides helpful guidance without taking control away from you - you remain in charge of the troubleshooting process. |
52 -
53 -### Using the Alert Assistant
54 -
55 -<details>
56 -<summary><strong>Accessing the Assistant</strong></summary><br/>
57 -
58 -1. Navigate to the **Alerts** tab.
59 -2. If there are active alerts, the **Actions** column will have an **Assistant** button.
60 -
61 - ![actions column](https://github-production-user-asset-6210df.s3.amazonaws.com/24860547/253559075-815ca123-e2b6-4d44-a780-eeee64cca420.png)
62 -
63 -3. Click the **Assistant** button to open a floating window with tailored troubleshooting insights.
64 -
65 -4. If there are no active alerts, you can still access the Assistant from the **Alert Configuration** view.
66 -
67 -</details>
68 -
69 -<details>
70 -<summary><strong>Understanding Assistant Information</strong></summary><br/>
71 -
72 -When you open the Assistant, you'll see:
73 -
74 -1. **Alert Context**: Explanation of what the alert means and why it's occurring
75 -
76 - ![Netdata Assistant popup](https://github-production-user-asset-6210df.s3.amazonaws.com/24860547/253559645-62850c7b-cd1d-45f2-b2dd-474ecbf2b713.png)
77 -
78 -2. **Troubleshooting Steps**: Recommended actions to address the issue
79 -
80 -3. **Importance Level**: Context on how critical this alert is for your system
81 -
82 -4. **Resource Links**: Curated documentation and external resources for further investigation
83 -
84 - ![useful resources](https://github-production-user-asset-6210df.s3.amazonaws.com/24860547/253560071-e768fa6d-6c9a-4504-bb1f-17d5f4707627.png)
85 -
86 -</details>
87 -
88 -### Real-World Alert Response
89 -
90 -Here's how the Alert Assistant helps in a critical situation:
91 -
92 -```mermaid
93 -flowchart LR
94 - A("**3 AM Alert**<br/><br/>load average 15<br/><br/>Emergency Response")
95 -
96 - B("**Without Assistant**<br/><br/>Google searches<br/><br/>Manual Research")
97 -
98 - C("**With Assistant**<br/><br/>Click Assistant button<br/><br/>Instant Help")
99 -
100 - D("**Time wasted<br/><br/>Stress increased<br/><br/>Delayed Resolution**")
101 -
102 - E("**Explanation of<br/><br/>system load cause<br/><br/>Immediate Context**")
103 -
104 - F("**Specific troubleshooting<br/><br/>steps provided<br/><br/>Guided Actions**")
105 -
106 - G("**Assistant follows as<br/><br/>you check metrics<br/><br/>Continuous Support**")
107 -
108 - H("**Quick access to<br/><br/>additional resources<br/><br/>Extended Learning**")
109 -
110 - I("**Issue resolved faster<br/><br/>with confidence<br/><br/>Successful Resolution**")
111 -
112 - A --> B
113 - A --> C
114 - B --> D
115 - C --> E
116 - E --> F
117 - F --> G
118 - G --> H
119 - H --> I
120 -
121 - %% Style definitions
122 - classDef alert fill:#ffeb3b,stroke:#000000,stroke-width:3px,color:#000000,font-size:18px
123 - classDef neutral fill:#f9f9f9,stroke:#000000,stroke-width:3px,color:#000000,font-size:18px
124 - classDef complete fill:#4caf50,stroke:#000000,stroke-width:3px,color:#000000,font-size:18px
125 - classDef problem fill:#f44336,stroke:#000000,stroke-width:3px,color:#000000,font-size:18px
126 -
127 - %% Apply styles
128 - class A alert
129 - class B,C neutral
130 - class D problem
131 - class E,F,G,H,I complete
132 -```
133 -
134 -## Strategic Insights Reports
135 -
136 -**Generate comprehensive infrastructure analysis** that synthesizes days, weeks, or months of high-resolution data into actionable intelligence for strategic decision-making.
137 -
138 -**Insights Reports transform your raw telemetry data into structured narratives** that help you understand trends, plan capacity, optimize performance, and conduct thorough post-incident analysis.
139 -
140 -### Four Types of Strategic Analysis
141 -
142 -| Report Type | What It Provides | Key Capabilities | Best Used For |
143 -|------------------------------|-----------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------|
144 -| **Infrastructure Summary** | Complete timeline of incidents, performance changes, and system behavior | **What happened**: Timeline reconstruction **Impact assessment**: Affected services **Current status**: Action items | Weekend incident recovery, executive updates, team handoffs |
145 -| **Capacity Planning** | Data-driven projections with concrete recommendations | **Trend analysis**: Resource utilization patterns **Bottleneck prediction**: Inflection-point dates **Scaling recommendations**: Hardware suggestions | Quarterly planning, budget justification, infrastructure roadmaps |
146 -| **Performance Optimization** | Synthesized analysis of system inefficiencies and improvement opportunities | **Contention patterns**: Resource conflicts **Optimization opportunities**: Tuning recommendations **Impact prioritization**: Biggest improvement areas | Performance debugging, system tuning, SRE optimization projects |
147 -| **Anomaly Analysis** | Context-aware detection and explanation of unusual infrastructure behavior | **Pattern recognition**: Abnormal behavior **Root cause analysis**: Why anomalies occurred **Trend correlation**: Cross-infrastructure connections | Post-incident analysis, proactive issue detection, system health assessment |
148 -
149 -:::tip
150 -
151 -**Simply request the analysis you need and get comprehensive reports** that turn months of raw telemetry data into clear, actionable intelligence that helps you make better infrastructure decisions faster.
152 -
153 -:::
154 -
155 -## How Netdata Insights Works
156 -
157 -The system combines three key components to deliver infrastructure intelligence at scale:
158 -
159 -<details>
160 -<summary><strong>1. Data Pipeline</strong></summary><br/>
161 -
162 -Your Netdata agents continue collecting metrics every second, storing them locally as they always have. **When you request analysis, Insights queries relevant time ranges across your infrastructure**, pulling raw metrics, events, and anomaly detection results.
163 -
164 -</details>
165 -
166 -<details>
167 -<summary><strong>2. Context Compression</strong></summary><br/>
168 -
169 -Raw telemetry data is compressed into structured context bundles that include:
170 -
171 -- **Statistical summaries** (percentiles, trends, correlation coefficients)
172 -- **Detected anomalies** with confidence scores and affected metrics
173 -- **Event timelines** (alerts, deployments, configuration changes)
174 -- **Cross-node correlations** and dependency mappings
175 -- **Historical baselines** for comparison
176 -
177 -</details>
178 -
179 -<details>
180 -<summary><strong>3. AI Analysis</strong></summary><br/>
181 -
182 -**Advanced language models process the compressed context** to generate structured reports with natural-language explanations, relevant visualizations, and actionable recommendations.
183 -
184 -:::important
185 -
186 -Your infrastructure data is processed for your reports and then discarded. **We never use them for training or model improvement**.
187 -
188 -:::
189 -
190 -</details>
191 -
192 -## What's Coming Next
193 -
194 -The goal is building **an autonomous debugging partner, not just another chatbot**. A system that scales human decision-making using all the info that Netdata already collects about your infrastructure.
195 -
196 -For details on upcoming features and our product roadmap, [read our full announcement on the Netdata blog](https://www.netdata.cloud/blog/netdata-insights/).
\ No newline at end of file