docs: preserve superseded TechCrunch PRD draft (#377)

Operator-approved; all doc-reference comments resolved. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

Juan Manuel Servera committed Jun 11, 2026 at 14:37 UTC ee90b383342407dbd2a892ee8d2985a069becacf
1 file changed +307
docs/processed/PRD-techcrunch-integration-2026-05-cross-source-correlation.md new
+307
@@ -0,0 +1,307 @@
1 +# PRD: TechCrunch RSS Integration for Cross-Source Trend Correlation
2 +
3 +> **Archival note (2026-06-11):** This is the **earlier, superseded** draft of the TechCrunch
4 +> RSS PRD, preserved for history. It was replaced by the revised, shipped PRD
5 +> [`PRD-techcrunch-integration.md`](./PRD-techcrunch-integration.md)
6 +> ("TechCrunch RSS Integration for Cross-Signal Enrichment"), which is the canonical record of
7 +> the implemented feature. See issue #377 for the reconciliation of the duplicate filename.
8 +
9 +**Author:** Farnsworth (Analyst/Content Curator)
10 +**Date:** 2026-05-19
11 +**Status:** Superseded — see canonical `PRD-techcrunch-integration.md`
12 +**Type:** Feature PRD
13 +**Depends on:** .squad/decisions-archive.md (Decision #7: Crawler Plugin Architecture), docs/processed/PRD-topic-channels.md
14 +
15 +---
16 +
17 +## Executive Summary
18 +
19 +SquadScope currently derives all insights from a single signal source: GitHub activity. While GitHub reveals *what developers are building*, it cannot tell us *why* activity is spiking — whether it's organic community interest, a VC-backed launch, or a viral TechCrunch article driving attention. This PRD proposes integrating TechCrunch's RSS feed as SquadScope's first non-GitHub data source, enabling **cross-source trend correlation** that distinguishes organic momentum from press-driven hype.
20 +
21 +**Key insight:** GitHub star surges often lag TechCrunch coverage by 24–72 hours. Detecting this pattern lets SquadScope editorially distinguish "genuinely important" (organic growth) from "temporarily hyped" (press-driven spike that fades within a week).
22 +
23 +---
24 +
25 +## Problem Statement
26 +
27 +### What GitHub data alone cannot tell us
28 +
29 +1. **Causality is invisible.** A repo gaining 2,000 stars in a week is interesting, but *why* matters editorially. Is it because the project shipped a breakthrough feature, or because TechCrunch wrote about it and HN amplified?
30 +
31 +2. **Funding and launch context is missing.** When a startup raises a $50M Series B and open-sources their core library, GitHub shows a star spike — but without the funding context, the analysis misattributes organic community excitement.
32 +
33 +3. **Industry narrative gaps.** SquadScope's "Gaps" section (what's missing from the conversation) is currently limited to what's absent from GitHub. But sometimes the gap is between what the industry *claims* to care about (per press coverage) and what's actually *being built* (per GitHub).
34 +
35 +4. **Hype detection requires a baseline.** To identify noise, you need to know what the press machine is amplifying. Without press data, everything on GitHub looks equally "organic."
36 +
37 +5. **Prediction accuracy suffers.** The topic-channels PRD envisions a prediction ledger. Cross-referencing press coverage with subsequent GitHub activity dramatically improves prediction calibration.
38 +
39 +---
40 +
41 +## Value Proposition
42 +
43 +### For SquadScope readers
44 +
45 +| Current State (GitHub-only) | With TechCrunch Correlation |
46 +|---|---|
47 +| "Repo X gained 3,000 stars this week" | "Repo X gained 3,000 stars after TechCrunch covered their $30M raise — watch if stars sustain past week 2" |
48 +| "These 5 AI repos are trending" | "3 of 5 trending AI repos correlate with press coverage; 2 show organic growth (stronger signal)" |
49 +| "Gap: No new observability tools" | "Gap: TechCrunch covered 4 observability startups this month, but none have meaningful GitHub traction yet — vaporware risk" |
50 +
51 +### For SquadScope's editorial stance
52 +
53 +- **Critical thinking becomes measurable:** "Press-amplified vs. organically growing" is a concrete, data-backed editorial judgment
54 +- **Signal vs. noise gets sharper:** Hype detection moves from vibes-based to correlation-based
55 +- **The Gaps section gains depth:** Disconnects between press narrative and actual developer activity become visible
56 +
57 +---
58 +
59 +## Correlation Model
60 +
61 +### How TechCrunch articles map to GitHub signals
62 +
63 +```
64 +┌─────────────────┐ ┌──────────────────────┐
65 +│ TechCrunch RSS │ │ GitHub Weekly Crawl │
66 +│ (article feed) │ │ (repo activity) │
67 +└────────┬────────┘ └──────────┬───────────┘
68 + │ │
69 + ▼ ▼
70 +┌─────────────────┐ ┌──────────────────────┐
71 +│ Extract: │ │ Extract: │
72 +│ - Company/proj │ │ - Repo name/org │
73 +│ - Category │ │ - Star delta │
74 +│ - Funding amt │ │ - Fork delta │
75 +│ - GitHub links │ │ - Contributor growth │
76 +└────────┬────────┘ └──────────┬───────────┘
77 + │ │
78 + └──────────┬───────────────────┘
79 + ▼
80 + ┌─────────────────────┐
81 + │ Correlation Engine │
82 + │ (fuzzy matching) │
83 + └──────────┬──────────┘
84 + ▼
85 + ┌─────────────────────┐
86 + │ Annotated Analysis │
87 + │ - press_correlated │
88 + │ - organic_growth │
89 + │ - hype_risk_score │
90 + └─────────────────────┘
91 +```
92 +
93 +### Correlation heuristics
94 +
95 +1. **Direct link match:** TechCrunch article contains a GitHub URL → exact match to crawled repo
96 +2. **Organization match:** Article mentions company X → match to `github.com/X/*` repos gaining stars
97 +3. **Project name match:** Article title/body contains project name → fuzzy match against repo names in weekly crawl
98 +4. **Category correlation:** Article tagged "AI" published Monday → AI-category repos spiking by Thursday
99 +5. **Temporal lag analysis:** Stars gained within 72 hours of article publication → likely press-correlated
100 +
101 +### Hype risk scoring
102 +
103 +| Pattern | Hype Risk | Editorial Label |
104 +|---------|-----------|-----------------|
105 +| Stars spike post-article, sustain 2+ weeks | Low | "Press-validated, community-sustained" |
106 +| Stars spike post-article, decay within 7 days | High | "Press-driven hype, fading interest" |
107 +| Stars growing before any press coverage | Very Low | "Organic growth — genuinely interesting" |
108 +| Press coverage but no GitHub activity | Medium | "Announced but unbuilt / closed-source" |
109 +
110 +---
111 +
112 +## Technical Approach
113 +
114 +### Data Source: TechCrunch RSS
115 +
116 +- **Feed URL:** `https://techcrunch.com/feed/`
117 +- **Format:** RSS 2.0 / XML
118 +- **Update frequency:** ~20-40 articles/day
119 +- **Relevant categories:** Startups, Apps, AI, Funding, Open Source
120 +- **Rate limits:** None (public RSS)
121 +- **Content available in feed:** Title, excerpt/summary, author, publish date, categories, link
122 +
123 +### Architecture: Fits Decision #7 (Crawler Plugin)
124 +
125 +The existing `DataSource` protocol interface applies directly:
126 +
127 +```python
128 +class TechCrunchSource:
129 + """Crawler plugin for TechCrunch RSS feed."""
130 +
131 + def get_name(self) -> str:
132 + return "techcrunch"
133 +
134 + def get_rate_limits(self) -> RateLimits:
135 + return RateLimits(requests_per_hour=10, burst=5)
136 +
137 + async def crawl(self, config: CrawlConfig) -> CrawlResult:
138 + """Fetch and parse TechCrunch RSS, extract structured articles."""
139 + ...
140 +```
141 +
142 +### Data flow integration
143 +
144 +```
145 +Existing: data/raw/YYYY-WNN.json (GitHub crawl)
146 +New: data/raw/YYYY-WNN-techcrunch.json (TechCrunch crawl)
147 +Merged: data/analyzed/YYYY-WNN-summary.md (cross-referenced analysis)
148 +```
149 +
150 +### RSS parsing requirements
151 +
152 +| Requirement | Approach |
153 +|-------------|----------|
154 +| XML parsing | `feedparser` (Python) — battle-tested RSS library |
155 +| Category extraction | Map TC categories to SquadScope topic taxonomy |
156 +| GitHub link extraction | Regex scan article content for `github.com` URLs |
157 +| Entity extraction | Match company/project names against crawled repos |
158 +| Deduplication | Hash on article URL; skip already-processed items |
159 +| Storage | JSON array, same weekly naming as GitHub crawl |
160 +
161 +### Output schema (per article)
162 +
163 +```json
164 +{
165 + "source": "techcrunch",
166 + "title": "Anthropic open-sources Claude's tool-use framework",
167 + "url": "https://techcrunch.com/2026/05/15/...",
168 + "published_at": "2026-05-15T14:30:00Z",
169 + "categories": ["ai", "open-source", "funding"],
170 + "github_links": ["https://github.com/anthropics/tool-use-sdk"],
171 + "entities": ["Anthropic", "Claude"],
172 + "funding_amount": null,
173 + "relevance_score": 0.85
174 +}
175 +```
176 +
177 +### Analyzer changes
178 +
179 +The analyzer prompt gains a new context block:
180 +
181 +```
182 +## Press Context (TechCrunch, week of {date})
183 +{N} articles published relevant to tech/open-source.
184 +Notable coverage:
185 +- {title} ({category}) — mentions {github_links}
186 +- ...
187 +
188 +Cross-reference: For each trending repo, note if press coverage
189 +preceded the star surge. Label as "press-correlated" or "organic."
190 +```
191 +
192 +---
193 +
194 +## Phases
195 +
196 +### Phase 1: RSS Crawl Plugin (1–2 weeks)
197 +
198 +- Implement `TechCrunchSource` crawler plugin
199 +- Parse RSS feed, extract structured article data
200 +- Store as `data/raw/YYYY-WNN-techcrunch.json`
201 +- Filter to tech/open-source relevant articles only
202 +- Basic deduplication
203 +- **Output:** Weekly TechCrunch article JSON alongside GitHub JSON
204 +
205 +### Phase 2: Correlation Engine (2–3 weeks)
206 +
207 +- Implement GitHub URL extraction from articles
208 +- Fuzzy entity matching (company name → GitHub org)
209 +- Temporal correlation (article date vs. star surge timing)
210 +- Add `press_correlated: bool` and `hype_risk: low|medium|high` to repo analysis
211 +- **Output:** Enriched analysis with cross-source annotations
212 +
213 +### Phase 3: Editorial Integration (1–2 weeks)
214 +
215 +- Update analyzer prompt to consume TechCrunch context
216 +- Add "Press vs. Reality" subsection to weekly summary
217 +- Surface disconnects in Gaps section
218 +- Update Hugo templates to render correlation badges
219 +- **Output:** Reader-facing cross-source insights on the published site
220 +
221 +### Phase 4: Prediction Enhancement (future)
222 +
223 +- Track whether press-correlated repos sustain momentum
224 +- Feed correlation accuracy back into prediction ledger
225 +- Calibrate hype risk scoring over time
226 +- **Output:** Improved prediction accuracy in topic channels
227 +
228 +---
229 +
230 +## Cost & Resource Impact
231 +
232 +| Resource | Impact |
233 +|----------|--------|
234 +| RSS fetch | Negligible (1 HTTP request/week, public feed, no auth) |
235 +| Storage | ~50-100 KB/week JSON (40 articles × metadata) |
236 +| Analyzer tokens | +500-800 tokens input context per run (~$0.002/week) |
237 +| API rate limits | Zero impact (RSS is not GitHub API) |
238 +| CI minutes | +5-10 seconds per run (RSS fetch + parse) |
239 +| Dependencies | `feedparser` (Python, MIT license, mature) |
240 +
241 +**Total incremental cost: <$0.01/week.** Trivial relative to base pipeline costs documented in PRD-cost-estimation.md.
242 +
243 +---
244 +
245 +## Risks & Mitigations
246 +
247 +| Risk | Probability | Impact | Mitigation |
248 +|------|------------|--------|------------|
249 +| TechCrunch changes RSS format | Low | Medium | feedparser handles format variations; alert on parse failures |
250 +| RSS feed discontinued | Very Low | Low | Graceful degradation — analysis runs without press context |
251 +| False correlations (noise) | Medium | Medium | Require temporal proximity (72h) + name match confidence >0.7 |
252 +| Over-weighting press signal | Medium | High | Editorial rule: press correlation is annotation, not ranking factor |
253 +| Content extraction blocked | Low | Low | Use RSS summary only, don't scrape full articles |
254 +
255 +---
256 +
257 +## Open Questions
258 +
259 +1. **OQ1: Should we also extract from TechCrunch's category-specific feeds?**
260 + - `techcrunch.com/category/artificial-intelligence/feed/` for topic-channel alignment
261 + - Pro: Better relevance filtering. Con: More feeds to manage.
262 +
263 +2. **OQ2: Full article fetch vs. RSS excerpt only?**
264 + - RSS includes ~200 word excerpt. Full article requires HTTP fetch + HTML parsing.
265 + - Recommendation: Start with RSS excerpt only. Avoids scraping concerns and ToS issues.
266 +
267 +3. **OQ3: Should correlation annotations be visible to readers or analyst-only?**
268 + - Option A: Show "📰 Press-correlated" badge on repo entries
269 + - Option B: Keep as internal signal that shapes editorial tone only
270 + - Recommendation: Option A for transparency (readers deserve to know *why* something is trending)
271 +
272 +4. **OQ4: Add HackerNews as a second correlation source simultaneously?**
273 + - HN has an API, overlaps with TechCrunch coverage, and better represents developer sentiment
274 + - Recommendation: TechCrunch first (simpler, RSS), HN second (API, different signal)
275 +
276 +5. **OQ5: How to handle TechCrunch articles about closed-source products?**
277 + - Many TC articles cover proprietary SaaS with no GitHub presence
278 + - Recommendation: Filter to articles containing GitHub links OR open-source keywords only
279 +
280 +---
281 +
282 +## Success Criteria
283 +
284 +| Metric | Target | Measurement |
285 +|--------|--------|-------------|
286 +| Articles crawled per week | 15-40 relevant | Count in weekly JSON |
287 +| Correlation hit rate | >30% of trending repos have press match | Cross-reference accuracy |
288 +| Hype detection accuracy | >70% of "high hype risk" repos show star decay at week +2 | Retrospective validation |
289 +| Reader value signal | Qualitative improvement in Gaps section depth | Editorial review |
290 +| Zero pipeline failures from RSS source | 100% graceful degradation | CI logs |
291 +
292 +---
293 +
294 +## Relationship to Existing PRDs
295 +
296 +- **PRD-topic-channels.md:** TechCrunch correlation enriches per-topic analysis. AI-focused TC articles correlate with `ai-ml` topic channel repos.
297 +- **PRD-cost-estimation.md:** Incremental cost is negligible (<$0.01/week). No tier change needed.
298 +- **.squad/decisions-archive.md Decision #7:** This is the first concrete implementation of the crawler plugin architecture.
299 +- **.squad/decisions-archive.md MCP Tools:** TechCrunch RSS fetch can be an MCP tool, registered in allowlist per Decision 5.
300 +
301 +---
302 +
303 +## Editorial Philosophy Note
304 +
305 +TechCrunch integration does NOT mean SquadScope becomes a TechCrunch aggregator. The feed is a **correlation signal**, not content to republish. SquadScope's voice remains: "Here's what's actually happening on GitHub this week, and here's what the press says is happening. Notice the gap? That's where the real story is."
306 +
307 +The editorial value is in the *delta* between press narrative and developer activity — not in summarizing TechCrunch articles.