main
md 128 lines 9.9 KB
Rendered Raw
1 # Research Intelligence Engine
2
3 Phase 3 builds the research evidence foundation. It does not produce final BUY/SELL recommendations.
4
5 ## Pipeline
6
7 Public source -> document discovery -> fetch -> extraction -> normalization -> deduplication -> entity resolution -> structured event extraction -> validation -> research fact store -> deterministic catalyst scoring -> future AI explanation.
8
9 LLMs are optional and may only produce schema-constrained extraction candidates. Invalid enum values, dates, numbers, or unsupported claims are rejected before persistence.
10
11 ## Source Providers
12
13 The source abstraction records source type, supported countries and markets, fetch strategy, reliability level, rate-limit policy, JavaScript requirement, and automatic access status.
14
15 Default providers cover company websites, investor relations, exchange announcements, regulatory filings, government procurement, permitted RSS, permitted news, and a provider-neutral search discovery boundary. Google search result scraping is not implemented.
16
17 Search discovery is discovery only:
18
19 ```text
20 search result/snippet -> candidate URL -> URL/security validation -> fetch original publisher page -> extract document -> validate evidence -> store event
21 ```
22
23 Search snippets are never persisted as research evidence. Implementations must use a supported search API, such as a Google-compatible custom search endpoint, Brave-compatible API, or another approved provider. Provider credentials come only from runtime environment or secret configuration.
24
25 ## Source Safety
26
27 The fetcher must not bypass CAPTCHA, paywalls, authentication, robots restrictions, anti-bot controls, or source terms. Sources that are not permitted for automation must be marked `UNAVAILABLE`, `RESTRICTED`, or `MANUAL_ONLY`.
28
29 Default live fetching is disabled with `AIP_RESEARCH_LIVE_ENABLED=false`. Demo fixtures are labelled `DEMO`.
30
31 ## Phase 3 Controlled Live Source
32
33 The first controlled live ingestion path is limited to AIXTRON SE:
34
35 - Instrument ID: `11111111-1111-1111-1111-111111111111`
36 - Stable identity: `AIXA` / `XETR` / `DE000A0WMPJ6`
37 - Registered source: AIXTRON official press release page on `aixtron.com`
38 - Source reliability: `LEVEL_B`
39 - Source mode: `REAL`
40
41 The frontend does not submit arbitrary URLs. Refresh uses registered source definitions first, then optional bounded search discovery only for categories that remain `NO_EVIDENCE`.
42
43 ## API Contract
44
45 Provider-free readiness and analysis, with an explicit targeted acquisition step:
46
47 ```http
48 GET /api/v1/research/readiness/{globalInstrumentId}
49 POST /api/v1/research/readiness/{globalInstrumentId}/ensure
50 POST /api/v1/research/analysis/{globalInstrumentId}
51 X-Correlation-Id: <caller-generated-id>
52 ```
53
54 The readiness GET and analysis POST never call providers. Only `ensure` may
55 execute planner-selected stale, partial, or missing requirements under the
56 per-instrument single-flight boundary.
57
58 Summary:
59
60 ```http
61 GET /api/v1/research/companies/{instrumentId}/summary
62 ```
63
64 The response contains:
65
66 - `profile`: company and instrument identity
67 - `catalystScore`: deterministic score, bucket scores, confidence, generated timestamp
68 - `recentEvents`: classified evidence with `sourceMode`, reliability, confidence, impact, dates, values, evidence excerpt, source URL, source classification, and supporting source URLs
69 - `documents`: normalized source documents with `sourceMode`, `freshness`, content hash, publication/retrieval/discovery dates, reliability, source classification, independence key, and canonical URL
70 - `dataFreshness`: `REAL`, `DEMO`, `DEMO_FALLBACK`, `SOURCE_UNAVAILABLE`, or `UNAVAILABLE`
71 - `demo`: true only when the returned summary is demo/fallback data
72 - `sourceMix`: document counts by source type
73
74 ## Fetching
75
76 Normal HTTP fetching is the first strategy. The implementation supports configurable user agent, connect/request timeouts, bounded retries, exponential backoff, `Retry-After`, redirect limits, redirect target validation, content-type validation, maximum content size, canonical URL normalization, and SSRF protection.
77
78 Playwright fallback is represented but disabled by default and not bundled in the default image. It may only be enabled for sources that permit JavaScript automation.
79
80 ## Documents
81
82 `ResearchDocument` supports canonical URL, original URL, title, source type/name/classification, publisher, publication/retrieval/discovery dates, language, content type, document type, content hash, instrument/company IDs, status, reliability, discovery provider, and source independence key. Raw and normalized bodies are excluded from API responses.
83
84 PDFs are supported as metadata/reference documents in Phase 3. OCR is not implemented.
85
86 ## Entity Resolution
87
88 Resolution does not use ticker alone. It combines ISIN, company name, aliases, ticker plus exchange, country, and known domains, and returns a confidence score.
89
90 ## Events And Scoring
91
92 `ResearchEvent` captures event type, date, source document, source URL, reliability, confidence, impact, time horizon, monetary values, percentages, customer/counterparty, location, capacity, and a short evidence reference.
93
94 Category scoring distinguishes validated evidence from missing evidence. Each category is reported as `POSITIVE_EVIDENCE`, `NEUTRAL_EVIDENCE`, `NEGATIVE_EVIDENCE`, `MIXED_EVIDENCE`, or `NO_EVIDENCE`; `NO_EVIDENCE` categories have a `null` score and are omitted from the overall catalyst calculation rather than being treated as score 50. Score 50 is reserved for validated neutral evidence.
95
96 Targeted research is an allowlisted discovery pass that runs only for categories with `NO_EVIDENCE` after primary source extraction. Discovery returns registered public sources with approved domains and source metadata; fetching still uses the SSRF, timeout, retry, content-type, redirect, size-limit, and user-agent controls. The system does not expose arbitrary URL fetching.
97
98 Optional search discovery runs after registered-source discovery when enabled with `AIP_RESEARCH_SEARCH_ENABLED=true` and a configured provider. The Google-compatible adapter requires `AIP_RESEARCH_SEARCH_ENDPOINT`, `AIP_RESEARCH_SEARCH_ENGINE_ID`, and a secret-backed `AIP_RESEARCH_SEARCH_API_KEY`; the adapter maps these to provider-specific `q`, `key`, `cx`, and `num` parameters. The Brave-compatible adapter requires `AIP_RESEARCH_SEARCH_ENDPOINT` and a secret-backed `AIP_RESEARCH_SEARCH_API_KEY`; it maps these to provider-specific `q`, `count`, and `X-Subscription-Token` values. It is bounded by max queries per category, max results per query, max documents per refresh, refresh cooldown, and cache TTL settings. Discovered URLs are classified as `OFFICIAL_COMPANY`, `REGULATORY`, `EXCHANGE`, `CUSTOMER`, `PARTNER`, `SUPPLIER`, `REPUTABLE_NEWS`, or `OTHER`; `OTHER` is rejected from scoring.
99
100 Source independence is tracked with canonical URLs and normalized content hashes. Syndicated or copied pages are deduplicated and do not increase independent-source confidence. Conflicting positive and negative evidence is preserved as `MIXED_EVIDENCE`.
101
102 The deterministic `CatalystScore` is 0-100 and uses event weights, source reliability, confidence, impact, and temporal decay. It is not a recommendation engine.
103
104 ## Persistence And Events
105
106 Versioned event contracts exist for `research.document.discovered`, `research.document.fetched`, `research.document.processed`, `research.event.extracted`, `research.event.rejected`, and `research.company.updated`. Events contain IDs and metadata, not raw page bodies or secrets.
107
108 Phase 4 persists normalized research data in PostgreSQL when `AIP_RESEARCH_PERSISTENCE_ENABLED=true`. The research engine uses the `research_documents`, `research_events`, `research_event_sources`, and `research_refresh_runs` tables. Runtime summaries are still built from the same provider-neutral `ResearchDocument` and `ResearchEvent` models, but REAL documents and events are reloaded from durable storage at startup.
109
110 Research schema migrations are owned by the Spring `research-service`, following the existing repository Flyway pattern used by portfolio and broker services. The migration is `services/research-service/src/main/resources/db/migration/V1__research_intelligence.sql`, applied by Flyway as version `1 - research intelligence` into the `research` schema with history table `flyway_schema_history_research`. In Kubernetes DEV, `research-engine` waits for `research-service` readiness when persistence is enabled, so Flyway has completed before the Python engine starts using PostgreSQL. The Python PostgreSQL repository does not create production tables at runtime; it fails fast with `RESEARCH_SCHEMA_NOT_MIGRATED` if the Flyway-owned schema is missing.
111
112 Idempotency strategy:
113
114 - Documents are unique by canonical URL and content hash.
115 - Events are unique by a stable event fingerprint composed from instrument ID, event type, normalized evidence reference, monetary value, customer, and counterparty.
116 - Event-source links are unique by event, document, and evidence excerpt, which allows multiple supporting documents without duplicate scoring contribution.
117 - Refresh runs are append-only audit records with safe status, counters, correlation ID, and sanitized error fields.
118
119 Portfolio-wide research refresh and its durable job polling model were retired
120 in Iteration 4. Portfolio and watchlist summaries remain read-only projections;
121 public company acquisition is keyed by `globalInstrumentId` and initiated only
122 through the requirement-targeted readiness `ensure` endpoint.
123
124 Kafka is not introduced in Phase 4 because no downstream consumer requires decoupled research updates yet. The existing in-process platform event list remains safe metadata only. If a future `ResearchUpdated` Kafka event is added, it must contain only event ID, version, correlation ID, instrument ID, company ID, refresh run ID, research status, and generated time.
125
126 ## Security
127
128 The fetcher rejects localhost, loopback, private ranges, link-local addresses, cloud metadata endpoints, non-http(s) schemes, and unsafe redirect targets. Search API keys are optional runtime secrets and must not be committed.