docs: preserve superseded TechCrunch PRD draft (#377)
Operator-approved; all doc-reference comments resolved. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Juan Manuel Servera committed
Jun 11, 2026 at 14:37 UTC
ee90b383342407dbd2a892ee8d2985a069becacf
1 file changed
+307
docs/processed/PRD-techcrunch-integration-2026-05-cross-source-correlation.md
new
+307
@@ -0,0 +1,307 @@
1
+# PRD: TechCrunch RSS Integration for Cross-Source Trend Correlation
2
+
3
+> **Archival note (2026-06-11):** This is the **earlier, superseded** draft of the TechCrunch
4
+> RSS PRD, preserved for history. It was replaced by the revised, shipped PRD
5
+> [`PRD-techcrunch-integration.md`](./PRD-techcrunch-integration.md)
6
+> ("TechCrunch RSS Integration for Cross-Signal Enrichment"), which is the canonical record of
7
+> the implemented feature. See issue #377 for the reconciliation of the duplicate filename.
8
+
9
+**Author:** Farnsworth (Analyst/Content Curator)
10
+**Date:** 2026-05-19
11
+**Status:** Superseded — see canonical `PRD-techcrunch-integration.md`
12
+**Type:** Feature PRD
13
+**Depends on:** .squad/decisions-archive.md (Decision #7: Crawler Plugin Architecture), docs/processed/PRD-topic-channels.md
14
+
15
+---
16
+
17
+## Executive Summary
18
+
19
+SquadScope currently derives all insights from a single signal source: GitHub activity. While GitHub reveals *what developers are building*, it cannot tell us *why* activity is spiking — whether it's organic community interest, a VC-backed launch, or a viral TechCrunch article driving attention. This PRD proposes integrating TechCrunch's RSS feed as SquadScope's first non-GitHub data source, enabling **cross-source trend correlation** that distinguishes organic momentum from press-driven hype.
20
+
21
+**Key insight:** GitHub star surges often lag TechCrunch coverage by 24–72 hours. Detecting this pattern lets SquadScope editorially distinguish "genuinely important" (organic growth) from "temporarily hyped" (press-driven spike that fades within a week).
22
+
23
+---
24
+
25
+## Problem Statement
26
+
27
+### What GitHub data alone cannot tell us
28
+
29
+1. **Causality is invisible.** A repo gaining 2,000 stars in a week is interesting, but *why* matters editorially. Is it because the project shipped a breakthrough feature, or because TechCrunch wrote about it and HN amplified?
30
+
31
+2. **Funding and launch context is missing.** When a startup raises a $50M Series B and open-sources their core library, GitHub shows a star spike — but without the funding context, the analysis misattributes organic community excitement.
32
+
33
+3. **Industry narrative gaps.** SquadScope's "Gaps" section (what's missing from the conversation) is currently limited to what's absent from GitHub. But sometimes the gap is between what the industry *claims* to care about (per press coverage) and what's actually *being built* (per GitHub).
34
+
35
+4. **Hype detection requires a baseline.** To identify noise, you need to know what the press machine is amplifying. Without press data, everything on GitHub looks equally "organic."
36
+
37
+5. **Prediction accuracy suffers.** The topic-channels PRD envisions a prediction ledger. Cross-referencing press coverage with subsequent GitHub activity dramatically improves prediction calibration.
38
+
39
+---
40
+
41
+## Value Proposition
42
+
43
+### For SquadScope readers
44
+
45
+| Current State (GitHub-only) | With TechCrunch Correlation |
46
+|---|---|
47
+| "Repo X gained 3,000 stars this week" | "Repo X gained 3,000 stars after TechCrunch covered their $30M raise — watch if stars sustain past week 2" |
48
+| "These 5 AI repos are trending" | "3 of 5 trending AI repos correlate with press coverage; 2 show organic growth (stronger signal)" |
49
+| "Gap: No new observability tools" | "Gap: TechCrunch covered 4 observability startups this month, but none have meaningful GitHub traction yet — vaporware risk" |
50
+
51
+### For SquadScope's editorial stance
52
+
53
+- **Critical thinking becomes measurable:** "Press-amplified vs. organically growing" is a concrete, data-backed editorial judgment
54
+- **Signal vs. noise gets sharper:** Hype detection moves from vibes-based to correlation-based
55
+- **The Gaps section gains depth:** Disconnects between press narrative and actual developer activity become visible
56
+
57
+---
58
+
59
+## Correlation Model
60
+
61
+### How TechCrunch articles map to GitHub signals
62
+
63
+```
64
+┌─────────────────┐ ┌──────────────────────┐
65
+│ TechCrunch RSS │ │ GitHub Weekly Crawl │
66
+│ (article feed) │ │ (repo activity) │
67
+└────────┬────────┘ └──────────┬───────────┘
68
+ │ │
69
+ ▼ ▼
70
+┌─────────────────┐ ┌──────────────────────┐
71
+│ Extract: │ │ Extract: │
72
+│ - Company/proj │ │ - Repo name/org │
73
+│ - Category │ │ - Star delta │
74
+│ - Funding amt │ │ - Fork delta │
75
+│ - GitHub links │ │ - Contributor growth │
76
+└────────┬────────┘ └──────────┬───────────┘
77
+ │ │
78
+ └──────────┬───────────────────┘
79
+ ▼
80
+ ┌─────────────────────┐
81
+ │ Correlation Engine │
82
+ │ (fuzzy matching) │
83
+ └──────────┬──────────┘
84
+ ▼
85
+ ┌─────────────────────┐
86
+ │ Annotated Analysis │
87
+ │ - press_correlated │
88
+ │ - organic_growth │
89
+ │ - hype_risk_score │
90
+ └─────────────────────┘
91
+```
92
+
93
+### Correlation heuristics
94
+
95
+1. **Direct link match:** TechCrunch article contains a GitHub URL → exact match to crawled repo
96
+2. **Organization match:** Article mentions company X → match to `github.com/X/*` repos gaining stars
97
+3. **Project name match:** Article title/body contains project name → fuzzy match against repo names in weekly crawl
98
+4. **Category correlation:** Article tagged "AI" published Monday → AI-category repos spiking by Thursday
99
+5. **Temporal lag analysis:** Stars gained within 72 hours of article publication → likely press-correlated
100
+
101
+### Hype risk scoring
102
+
103
+| Pattern | Hype Risk | Editorial Label |
104
+|---------|-----------|-----------------|
105
+| Stars spike post-article, sustain 2+ weeks | Low | "Press-validated, community-sustained" |
106
+| Stars spike post-article, decay within 7 days | High | "Press-driven hype, fading interest" |
107
+| Stars growing before any press coverage | Very Low | "Organic growth — genuinely interesting" |
108
+| Press coverage but no GitHub activity | Medium | "Announced but unbuilt / closed-source" |
109
+
110
+---
111
+
112
+## Technical Approach
113
+
114
+### Data Source: TechCrunch RSS
115
+
116
+- **Feed URL:** `https://techcrunch.com/feed/`
117
+- **Format:** RSS 2.0 / XML
118
+- **Update frequency:** ~20-40 articles/day
119
+- **Relevant categories:** Startups, Apps, AI, Funding, Open Source
120
+- **Rate limits:** None (public RSS)
121
+- **Content available in feed:** Title, excerpt/summary, author, publish date, categories, link
122
+
123
+### Architecture: Fits Decision #7 (Crawler Plugin)
124
+
125
+The existing `DataSource` protocol interface applies directly:
126
+
127
+```python
128
+class TechCrunchSource:
129
+ """Crawler plugin for TechCrunch RSS feed."""
130
+
131
+ def get_name(self) -> str:
132
+ return "techcrunch"
133
+
134
+ def get_rate_limits(self) -> RateLimits:
135
+ return RateLimits(requests_per_hour=10, burst=5)
136
+
137
+ async def crawl(self, config: CrawlConfig) -> CrawlResult:
138
+ """Fetch and parse TechCrunch RSS, extract structured articles."""
139
+ ...
140
+```
141
+
142
+### Data flow integration
143
+
144
+```
145
+Existing: data/raw/YYYY-WNN.json (GitHub crawl)
146
+New: data/raw/YYYY-WNN-techcrunch.json (TechCrunch crawl)
147
+Merged: data/analyzed/YYYY-WNN-summary.md (cross-referenced analysis)
148
+```
149
+
150
+### RSS parsing requirements
151
+
152
+| Requirement | Approach |
153
+|-------------|----------|
154
+| XML parsing | `feedparser` (Python) — battle-tested RSS library |
155
+| Category extraction | Map TC categories to SquadScope topic taxonomy |
156
+| GitHub link extraction | Regex scan article content for `github.com` URLs |
157
+| Entity extraction | Match company/project names against crawled repos |
158
+| Deduplication | Hash on article URL; skip already-processed items |
159
+| Storage | JSON array, same weekly naming as GitHub crawl |
160
+
161
+### Output schema (per article)
162
+
163
+```json
164
+{
165
+ "source": "techcrunch",
166
+ "title": "Anthropic open-sources Claude's tool-use framework",
167
+ "url": "https://techcrunch.com/2026/05/15/...",
168
+ "published_at": "2026-05-15T14:30:00Z",
169
+ "categories": ["ai", "open-source", "funding"],
170
+ "github_links": ["https://github.com/anthropics/tool-use-sdk"],
171
+ "entities": ["Anthropic", "Claude"],
172
+ "funding_amount": null,
173
+ "relevance_score": 0.85
174
+}
175
+```
176
+
177
+### Analyzer changes
178
+
179
+The analyzer prompt gains a new context block:
180
+
181
+```
182
+## Press Context (TechCrunch, week of {date})
183
+{N} articles published relevant to tech/open-source.
184
+Notable coverage:
185
+- {title} ({category}) — mentions {github_links}
186
+- ...
187
+
188
+Cross-reference: For each trending repo, note if press coverage
189
+preceded the star surge. Label as "press-correlated" or "organic."
190
+```
191
+
192
+---
193
+
194
+## Phases
195
+
196
+### Phase 1: RSS Crawl Plugin (1–2 weeks)
197
+
198
+- Implement `TechCrunchSource` crawler plugin
199
+- Parse RSS feed, extract structured article data
200
+- Store as `data/raw/YYYY-WNN-techcrunch.json`
201
+- Filter to tech/open-source relevant articles only
202
+- Basic deduplication
203
+- **Output:** Weekly TechCrunch article JSON alongside GitHub JSON
204
+
205
+### Phase 2: Correlation Engine (2–3 weeks)
206
+
207
+- Implement GitHub URL extraction from articles
208
+- Fuzzy entity matching (company name → GitHub org)
209
+- Temporal correlation (article date vs. star surge timing)
210
+- Add `press_correlated: bool` and `hype_risk: low|medium|high` to repo analysis
211
+- **Output:** Enriched analysis with cross-source annotations
212
+
213
+### Phase 3: Editorial Integration (1–2 weeks)
214
+
215
+- Update analyzer prompt to consume TechCrunch context
216
+- Add "Press vs. Reality" subsection to weekly summary
217
+- Surface disconnects in Gaps section
218
+- Update Hugo templates to render correlation badges
219
+- **Output:** Reader-facing cross-source insights on the published site
220
+
221
+### Phase 4: Prediction Enhancement (future)
222
+
223
+- Track whether press-correlated repos sustain momentum
224
+- Feed correlation accuracy back into prediction ledger
225
+- Calibrate hype risk scoring over time
226
+- **Output:** Improved prediction accuracy in topic channels
227
+
228
+---
229
+
230
+## Cost & Resource Impact
231
+
232
+| Resource | Impact |
233
+|----------|--------|
234
+| RSS fetch | Negligible (1 HTTP request/week, public feed, no auth) |
235
+| Storage | ~50-100 KB/week JSON (40 articles × metadata) |
236
+| Analyzer tokens | +500-800 tokens input context per run (~$0.002/week) |
237
+| API rate limits | Zero impact (RSS is not GitHub API) |
238
+| CI minutes | +5-10 seconds per run (RSS fetch + parse) |
239
+| Dependencies | `feedparser` (Python, MIT license, mature) |
240
+
241
+**Total incremental cost: <$0.01/week.** Trivial relative to base pipeline costs documented in PRD-cost-estimation.md.
242
+
243
+---
244
+
245
+## Risks & Mitigations
246
+
247
+| Risk | Probability | Impact | Mitigation |
248
+|------|------------|--------|------------|
249
+| TechCrunch changes RSS format | Low | Medium | feedparser handles format variations; alert on parse failures |
250
+| RSS feed discontinued | Very Low | Low | Graceful degradation — analysis runs without press context |
251
+| False correlations (noise) | Medium | Medium | Require temporal proximity (72h) + name match confidence >0.7 |
252
+| Over-weighting press signal | Medium | High | Editorial rule: press correlation is annotation, not ranking factor |
253
+| Content extraction blocked | Low | Low | Use RSS summary only, don't scrape full articles |
254
+
255
+---
256
+
257
+## Open Questions
258
+
259
+1. **OQ1: Should we also extract from TechCrunch's category-specific feeds?**
260
+ - `techcrunch.com/category/artificial-intelligence/feed/` for topic-channel alignment
261
+ - Pro: Better relevance filtering. Con: More feeds to manage.
262
+
263
+2. **OQ2: Full article fetch vs. RSS excerpt only?**
264
+ - RSS includes ~200 word excerpt. Full article requires HTTP fetch + HTML parsing.
265
+ - Recommendation: Start with RSS excerpt only. Avoids scraping concerns and ToS issues.
266
+
267
+3. **OQ3: Should correlation annotations be visible to readers or analyst-only?**
268
+ - Option A: Show "📰 Press-correlated" badge on repo entries
269
+ - Option B: Keep as internal signal that shapes editorial tone only
270
+ - Recommendation: Option A for transparency (readers deserve to know *why* something is trending)
271
+
272
+4. **OQ4: Add HackerNews as a second correlation source simultaneously?**
273
+ - HN has an API, overlaps with TechCrunch coverage, and better represents developer sentiment
274
+ - Recommendation: TechCrunch first (simpler, RSS), HN second (API, different signal)
275
+
276
+5. **OQ5: How to handle TechCrunch articles about closed-source products?**
277
+ - Many TC articles cover proprietary SaaS with no GitHub presence
278
+ - Recommendation: Filter to articles containing GitHub links OR open-source keywords only
279
+
280
+---
281
+
282
+## Success Criteria
283
+
284
+| Metric | Target | Measurement |
285
+|--------|--------|-------------|
286
+| Articles crawled per week | 15-40 relevant | Count in weekly JSON |
287
+| Correlation hit rate | >30% of trending repos have press match | Cross-reference accuracy |
288
+| Hype detection accuracy | >70% of "high hype risk" repos show star decay at week +2 | Retrospective validation |
289
+| Reader value signal | Qualitative improvement in Gaps section depth | Editorial review |
290
+| Zero pipeline failures from RSS source | 100% graceful degradation | CI logs |
291
+
292
+---
293
+
294
+## Relationship to Existing PRDs
295
+
296
+- **PRD-topic-channels.md:** TechCrunch correlation enriches per-topic analysis. AI-focused TC articles correlate with `ai-ml` topic channel repos.
297
+- **PRD-cost-estimation.md:** Incremental cost is negligible (<$0.01/week). No tier change needed.
298
+- **.squad/decisions-archive.md Decision #7:** This is the first concrete implementation of the crawler plugin architecture.
299
+- **.squad/decisions-archive.md MCP Tools:** TechCrunch RSS fetch can be an MCP tool, registered in allowlist per Decision 5.
300
+
301
+---
302
+
303
+## Editorial Philosophy Note
304
+
305
+TechCrunch integration does NOT mean SquadScope becomes a TechCrunch aggregator. The feed is a **correlation signal**, not content to republish. SquadScope's voice remains: "Here's what's actually happening on GitHub this week, and here's what the press says is happening. Notice the gap? That's where the real story is."
306
+
307
+The editorial value is in the *delta* between press narrative and developer activity — not in summarizing TechCrunch articles.