Add cloud-agent product plan and roadmap (internal)

Strategy, architecture, and phased roadmap for a Copilot-style cloud coding agent on AWS, plus the pricing model. Private-repo planning doc. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

Seto Elkahfi committed Jun 26, 2026 at 16:43 UTC c4eafd9389344575aa9ca6b2422de51302aefd63
1 file changed +419
docs/product/cloud-agent-plan.md new
+419
@@ -0,0 +1,419 @@
1 +# siGit Cloud Agent — product plan and roadmap
2 +
3 +Status: draft v1 (2026-06-26). Owner: product/eng. Audience: internal (private repo).
4 +
5 +A GitHub-Copilot-coding-agent-style product built into sigit.si and siGit Code
6 +Cloud: you give it a task against a repo you host on sigit.si, an isolated AWS
7 +sandbox runs the siGit Code agent loop against that repo, and it returns a branch
8 +plus a proposed pull request. This document is the strategy, the architecture, and
9 +the phased roadmap.
10 +
11 +---
12 +
13 +## 1. What we are building (and what it is not)
14 +
15 +**siGit Cloud Agent** is an asynchronous, server-side coding agent. The user
16 +describes a task ("add pagination to the repos list", "fix the failing
17 +auth test", "upgrade Rails to 8.1"), points it at one of their sigit.si repos and
18 +a base branch, and walks away. The platform provisions an ephemeral AWS sandbox,
19 +clones the repo into it, runs siGit Code headless against the task with cloud
20 +inference, lets it edit files and run build/test commands, then pushes a head
21 +branch back to sigit.si and opens a pull request for human review.
22 +
23 +It is the cloud, autonomous sibling of the two things siGit already ships:
24 +
25 +- **siGit Code** (public Rust CLI / ACP agent) runs *locally / on the desktop*,
26 + interactively, driven by the developer at their keyboard.
27 +- **siGit Cloud Agent** runs *in our cloud*, asynchronously, driven by a task and
28 + reviewed afterward through a PR.
29 +
30 +Reference points in the market: GitHub Copilot coding agent (assign an issue, it
31 +opens a PR), OpenAI Codex cloud, Cursor background agents, Devin. The
32 +differentiator for us is that we already own the whole vertical: the git host, the
33 +agent, the inference, the account, and the billing all live inside siGit. We are
34 +not bolting an agent onto someone else's platform; we are completing a platform we
35 +already run.
36 +
37 +What it is **not**, in v1: it is not an autonomous merger (a human reviews and
38 +merges), not a long-lived persistent dev environment (sandboxes are ephemeral per
39 +run), and not a chat product (the chat tier already exists; this is the
40 +task-to-PR product on top of it).
41 +
42 +---
43 +
44 +## 2. Why siGit is unusually well positioned
45 +
46 +The expensive parts of a cloud-agent product already exist in this repo and the
47 +sibling repos. The new work is mostly orchestration and isolation, not net-new
48 +agent or inference plumbing.
49 +
50 +| Capability the product needs | Already exists | Where |
51 +|---|---|---|
52 +| A git host the agent can clone from and push to | yes | `git_http_controller` (smart-HTTP, bare repos on disk, `Repository#disk_path`) |
53 +| The coding agent itself (edit/run/test loop, tool calling) | yes | `sigit` (public Rust CLI/ACP agent), `onde-cloud` tool-call mapping |
54 +| Hosted inference with auth, identity masking, metering | yes | `Api::V1::ChatCompletionsController` → `OndeCloudService` → Onde Cloud |
55 +| Accounts and a trust boundary that mints scoped tokens | yes | smbCloud auth, `Api::V1` sessions/me, `git_token` exchange |
56 +| Subscription gating + monthly metering | yes | `Subscription`, `CloudUsage`, `User#entitled_to_cloud?` |
57 +
58 +What is **missing** and must be built (see roadmap):
59 +
60 +1. **An execution plane on AWS** (the sandbox machine) and the orchestration to
61 + drive it. This is the heart of the project.
62 +2. **A control plane in Rails**: an `AgentRun` resource, its lifecycle, log
63 + streaming, and web UI.
64 +3. **Pull requests** (and ideally issues). sigit.si has repos, blobs, commits, and
65 + stars, but no `PullRequest` model today. The agent's output is a PR, so a
66 + minimal PR/diff/review surface is a hard dependency. This can be scoped down to
67 + a "compare and open PR" view for v1.
68 +4. **Per-run scoped credentials**: short-lived git-push and inference tokens minted
69 + server-side, budget-bounded, never long-lived in the sandbox.
70 +5. **A new metering/pricing dimension**: agent runs consume sandbox compute *and*
71 + inference tokens, so COGS has two drivers, not one.
72 +
73 +---
74 +
75 +## 3. Architecture
76 +
77 +Two planes plus the existing inference path. Control plane is Rails (sigit-si);
78 +execution plane is AWS; inference reuses the existing `/api/v1/chat/completions`
79 +proxy unchanged.
80 +
81 +```
82 +┌─ Control plane (Rails, sigit-si) ───────────────────────────────────────────┐
83 +│ Web UI: repo "Agent" tab → run form, live transcript, diff, "Open PR" │
84 +│ AgentRunsController + Api::V1::AgentRunsController (create/show/cancel/log) │
85 +│ AgentRun model (lifecycle state machine) │
86 +│ AgentRunJob → provisions sandbox, monitors, collects result │
87 +│ Mints per-run scoped tokens (git push + inference), enforces caps/metering │
88 +└──────────────────────────────────────────────────────────────────────────────┘
89 + │ RunTask (aws-sdk) ▲ SSE/webhook: logs, status, diff
90 + ▼ │
91 +┌─ Execution plane (AWS) ─────────────────────────────────────────────────────┐
92 +│ Ephemeral sandbox (Fargate task v1 → Firecracker microVM at scale) │
93 +│ ├─ clones repo from sigit.si over smart-HTTP (scoped git token) │
94 +│ ├─ runs siGit Code headless against the task │
95 +│ │ OPENAI_BASE_URL=https://sigit.si/api/v1 (per-run inference token) │
96 +│ ├─ edits files, runs build/test in restricted shell │
97 +│ └─ pushes head branch back to sigit.si (scoped git token) │
98 +│ Private subnet, egress allowlist (sigit.si + package registries only) │
99 +│ Hard caps: wall-clock, CPU/mem, token budget, max tool calls │
100 +└──────────────────────────────────────────────────────────────────────────────┘
101 + │ /v1/chat/completions (per-run token)
102 + ▼
103 +┌─ Inference (unchanged) ─────────────────────────────────────────────────────┐
104 +│ Api::V1::ChatCompletionsController → OndeCloudService → Onde Cloud → upstream│
105 +│ Existing entitlement gate, allowance metering, and identity masking apply. │
106 +└──────────────────────────────────────────────────────────────────────────────┘
107 +```
108 +
109 +### 3.1 Control plane (Rails)
110 +
111 +**`AgentRun` model** (new table). Belongs to `user` and `repository`. Fields:
112 +
113 +- `status`: `queued`, `provisioning`, `running`, `pushing`, `needs_input`,
114 + `completed`, `failed`, `canceled` (state machine; one-way transitions logged).
115 +- `task_prompt` (text), `base_branch`, `head_branch` (generated, e.g.
116 + `agent/<run-id>-<slug>`), `pull_request_id` (nullable until pushed).
117 +- `sandbox_ref` (ECS task ARN / microVM id), `region`.
118 +- Budgets and accounting: `token_budget`, `tokens_used`, `wall_clock_limit_s`,
119 + `started_at`, `finished_at`, `exit_reason`.
120 +- `transcript_url` (S3 pointer for the full log), plus a tail kept in Postgres for
121 + the live view.
122 +
123 +**Controllers**: a web `AgentRunsController` (HTML, Turbo) under the repo, and an
124 +`Api::V1::AgentRunsController` so the CLI and desktop can trigger and follow runs.
125 +Actions: `create`, `index`, `show`, `cancel`, `messages#create` (steer a running
126 +agent), and a `logs` SSE endpoint that relays the live transcript (reuse the
127 +`ActionController::Live` pattern already used for chat streaming).
128 +
129 +**`AgentRunJob`** (Active Job, on the existing DB-backed queue): transitions the
130 +run to `provisioning`, calls AWS to start the sandbox, persists the sandbox ref,
131 +then hands off to monitoring. Cancellation and timeout both tear the sandbox down.
132 +
133 +**Web UI**: a new "Agent" (or "Tasks") tab on the repository page. A run form
134 +(task prompt, base branch, optional model tier). A live transcript panel (Turbo
135 +Streams fed by the SSE relay). On completion, a diff view and an "Open pull
136 +request" action.
137 +
138 +### 3.2 Execution plane (AWS) — the sandbox machine
139 +
140 +This is the core AWS decision and the riskiest surface, because the sandbox runs
141 +build and test commands over user code.
142 +
143 +**Runtime choice.**
144 +
145 +- **v1: AWS Fargate (ECS) ephemeral tasks.** One task per run. Scales to zero, pay
146 + per second, decent container isolation, no servers to manage, and `RunTask` is a
147 + single SDK call from a Rails job. Fast to ship. This is the recommendation for
148 + v1.
149 +- **At scale: Firecracker microVMs.** For stronger isolation of untrusted code and
150 + for snapshot/restore warm pools (sub-second starts), move the sandbox to
151 + Firecracker microVMs on bare-metal EC2 (the model E2B / Modal / Codex-style
152 + sandboxes use). More operational weight; defer until run volume and the threat
153 + model justify it. gVisor or Kata on EC2 is a middle option if Fargate isolation
154 + proves insufficient before we are ready for Firecracker.
155 +
156 +**The sandbox image.** A container that bundles headless siGit Code plus a base
157 +toolchain (git, common language runtimes). The agent boots, reads the run spec
158 +from an injected env/file, clones, works, and pushes. Per-language base images (or
159 +a `.sigit/agent.yml` setup step, see Phase 3) keep cold builds fast.
160 +
161 +**Orchestration.** v1 keeps it simple: the Rails `AgentRunJob` calls ECS `RunTask`
162 +directly via `aws-sdk-ecs`, passing the run spec as container overrides, and polls
163 +task status (or receives EventBridge task-state-change events into a webhook). If
164 +the lifecycle grows (retries, multi-step, fan-out), promote to Step Functions.
165 +Avoid Step Functions on day one; it is premature.
166 +
167 +**Networking and isolation (load-bearing for safety).**
168 +
169 +- Sandbox runs in a **private subnet**. Egress through a NAT restricted by an
170 + **allowlist**: sigit.si (git + inference) and an explicit set of package
171 + registries (rubygems, npm, pypi, crates, etc.). Everything else is denied. This
172 + is the primary control against data exfiltration, SSRF against internal
173 + services, and crypto-mining abuse.
174 +- **No inbound.** The sandbox is not reachable from the internet.
175 +- Per-run IAM role scoped to only what the task needs; no broad account access
176 + inside the sandbox.
177 +
178 +**Credentials (mint short-lived, never long-lived).** The sandbox receives:
179 +
180 +- a **git token scoped to the single repo and ideally the single head branch**,
181 + valid for the run only (extends the existing `git_token` exchange);
182 +- an **inference token** minted per run, carrying the user's entitlement and a
183 + hard token budget, pointed at `https://sigit.si/api/v1`.
184 +
185 +Both expire when the run ends. The sandbox never holds `app_secret` or any
186 +long-lived credential, mirroring the existing public-client rule.
187 +
188 +**Logs and artifacts.** The agent streams transcript chunks back to Rails (the SSE
189 +relay) for the live view and writes the full transcript and build logs to S3
190 +(pointer stored on `AgentRun`). CloudWatch captures infra-level logs.
191 +
192 +### 3.3 Inference path (reused as-is)
193 +
194 +The sandboxed agent sets `OPENAI_BASE_URL=https://sigit.si/api/v1` and uses the
195 +per-run inference token as its bearer. That flows through the **existing**
196 +`Api::V1::ChatCompletionsController`: entitlement gate, allowance metering, and the
197 +`SIGIT_IDENTITY_PROMPT` identity masking all apply with no change. This is a major
198 +reason the project is tractable: the autonomous agent is just another client of an
199 +inference endpoint we already operate and protect.
200 +
201 +This also gives a clean answer to the existing "inference token ↔ Onde Cloud auth"
202 +open item: for cloud-agent runs the token is minted server-side with a budget, so
203 +there is no public client holding credentials at all.
204 +
205 +### 3.4 Git and PR flow
206 +
207 +The agent pushes the head branch via the existing `git-receive-pack` endpoint. On
208 +push, the platform creates (or links) the PR record. Because **no PR model exists
209 +yet**, scope for v1:
210 +
211 +- minimal `PullRequest` model (base/head branch, repo, author, status, title,
212 + body), a diff/compare view (we already render blobs and commits, so the diff
213 + renderer is incremental), and "open / close / merge" actions for the reviewer;
214 +- the agent fills in title and body from its summary. PR prose must stay neutral
215 + and must never name the upstream model or provider (same rule as chat).
216 +
217 +Issues (assign-an-issue-to-the-agent) are a Phase 3 surface and depend on an issue
218 +model that also does not exist yet.
219 +
220 +---
221 +
222 +## 4. Safety, guardrails, and abuse
223 +
224 +Principal-engineer non-negotiables, because this executes code on our infra on
225 +behalf of users:
226 +
227 +- **Hard caps per run**: wall-clock timeout, CPU/memory limits, inference token
228 + budget (tied to `CloudUsage`/a new `AgentUsage`), max tool calls, max sandbox
229 + lifetime. A run that blows any cap is killed and marked `failed` with a reason.
230 +- **Network egress allowlist** (section 3.2). The single most important control.
231 +- **Ephemeral, scoped credentials** only. Nothing long-lived in the sandbox.
232 +- **Human-in-the-loop by default**: the agent proposes a PR; it does not merge. No
233 + auto-merge in v1.
234 +- **Concurrency limits per plan**: caps simultaneous runs per user to bound spend
235 + and abuse.
236 +- **Identity hygiene**: transcripts, PR titles/bodies, and error messages stay
237 + neutral; never disclose the upstream model or provider (existing rule extends to
238 + agent output).
239 +- **Cancellation is real**: cancel tears down the sandbox and revokes the run's
240 + tokens.
241 +- **Idempotency**: run creation and sandbox start are idempotent so retries cannot
242 + double-spend.
243 +
244 +A short threat-model doc is a Phase 0 deliverable (exfiltration, SSRF to internal
245 +metadata endpoints, resource abuse / mining, secret leakage from the user's own
246 +repo, prompt injection from repo contents steering the agent).
247 +
248 +---
249 +
250 +## 5. Pricing and packaging
251 +
252 +Agent runs cost us **sandbox compute (Fargate per-second) + inference tokens**, so
253 +metering must capture both and price above blended COGS. The numbers below are a
254 +working model with the inputs we cannot yet read from this repo marked
255 +`[FILL FROM PHASE 0]`. The metering schema and billing copy should be designed
256 +against this formula now; the dollar figures get populated once the Phase 0
257 +dogfood produces measured per-run token counts.
258 +
259 +### 5.1 Unit economics: cost per run (COGS)
260 +
261 +```
262 +cost_per_run = sandbox_cost + inference_cost
263 +
264 +sandbox_cost = vcpu_count * $0.04048/vcpu-hr * hours
265 + + mem_gb * $0.004445/gb-hr * hours (Fargate, us-east-1)
266 +
267 +inference_cost = tokens_per_run / 1_000_000 * blended_upstream_cost_per_mtok
268 +```
269 +
270 +Worked example, sandbox side (known, fixed): 1 vCPU + 2 GB for a 20-minute run:
271 +
272 +```
273 +(1 * 0.04048 + 2 * 0.004445) * (20/60) = $0.0165 per run
274 +```
275 +
276 +Sandbox compute is rounding error. **Inference dominates**, and it is the unknown:
277 +
278 +```
279 +tokens_per_run = [FILL FROM PHASE 0] (expect 200K – 2M+)
280 +blended_upstream_cost_per_mtok = [FILL: Onde/upstream blended $/Mtok]
281 +
282 +inference_cost_per_run = (tokens_per_run / 1e6) * blended_cost_per_mtok
283 +cost_per_run = 0.0165 + inference_cost_per_run
284 +```
285 +
286 +Conclusion that holds regardless of the blanks: **price on metered inference per
287 +run; treat sandbox-minutes as a kill-switch guardrail, not a billing axis.**
288 +
289 +### 5.2 Packaging (allowance model, mirrors `Subscription::CLOUD_ALLOWANCE`)
290 +
291 +Agent runs are an included monthly bundle per plan, with overage or an add-on for
292 +heavy use. Suggested `Subscription` constant shape:
293 +
294 +```ruby
295 +# Included siGit Cloud Agent runs per billing period. Tunable pricing knob.
296 +AGENT_RUNS = { "pro" => [FILL], "team" => [FILL] }.freeze # e.g. pro 20, team 75
297 +TRIAL_AGENT_RUNS = [FILL] # e.g. 2
298 +AGENT_OVERAGE_PER_RUN = [FILL] # $/run past the bundle
299 +```
300 +
301 +| Plan | Price (today) | Included agent runs/mo | Overage | Notes |
302 +|---|---|---|---|---|
303 +| Free | $0 | 0 | n/a | local siGit Code only, no cloud agent |
304 +| Trial | $0 (14 days) | `[FILL: ~2]` | n/a | felt-value, hard run cap |
305 +| Pro | $20/mo (existing) | `[FILL: ~20]` | `[FILL: $/run]` | bundle sized so COGS < ~X% of $20 |
306 +| Team | (existing) | `[FILL: ~75]` | `[FILL: $/run]` | higher bundle + concurrency |
307 +
308 +Sizing rule for the bundle: pick included-runs so that **bundle COGS stays under a
309 +target fraction of plan price** (e.g. included-runs * cost_per_run ≤ 40% of MRR),
310 +leaving margin for the chat allowance those plans already include. With
311 +`cost_per_run` from 5.1 unknown, the bundle count is the lever set last.
312 +
313 +### 5.3 Metering
314 +
315 +Add `AgentUsage` keyed by `(user, billing_period_key)` (parallel to `CloudUsage`),
316 +recording per period: `runs_count`, `agent_seconds`, `tokens_used`. Runs decrement
317 +the plan's run bundle; tokens already flow through the existing cloud allowance and
318 +its 429 cap, so a single run can never escape the token budget. `GET
319 +/api/v1/billing` gains `agent_runs_used` + `agent_runs_allowance` alongside the
320 +existing cloud fields.
321 +
322 +### 5.4 Guardrails (worst-case cost containment)
323 +
324 +- **Per-run token budget** minted into the run token, independent of the monthly
325 + allowance, so one pathological run is bounded.
326 +- **Per-run wall-clock + sandbox-lifetime cap** (kills runaway compute).
327 +- **Per-plan concurrency cap** bounds simultaneous spend.
328 +- **Trial run cap** (`TRIAL_AGENT_RUNS`) prevents trial-driven upstream bills,
329 + mirroring `TRIAL_ALLOWANCE`.
330 +
331 +### 5.5 What Phase 0 must measure to finalize this
332 +
333 +1. **`tokens_per_run`** distribution (p50/p90/p99) over 20–30 real dogfood runs.
334 +2. **`blended_upstream_cost_per_mtok`** (from the Onde/upstream cost, internal).
335 +3. Resulting **`cost_per_run`** distribution, then back-solve included-run bundles
336 + per tier against the target-margin rule in 5.2.
337 +
338 +Until 1 and 2 are real, every dollar above is a placeholder by design.
339 +
340 +---
341 +
342 +## 6. Roadmap (phased)
343 +
344 +Estimates are calendar weeks for a small team; adjust to staffing. Each phase ends
345 +with a demoable, dogfoodable increment.
346 +
347 +### Phase 0 — Foundations and spikes (2–3 weeks)
348 +- Decide sandbox runtime: **Fargate for v1** (documented path to Firecracker).
349 +- Build the headless siGit Code container image; prove end to end **manually**: a
350 + container clones a sigit.si repo, runs the agent against a task with inference via
351 + `/api/v1`, edits files, runs a test, and pushes a branch.
352 +- Per-run scoped token minting (git + inference) in Rails.
353 +- Threat model + egress allowlist design.
354 +- Add `aws-sdk-ecs`/`aws-sdk-core` to the Gemfile; stand up the VPC/subnet/NAT and
355 + the task definition in IaC.
356 +- **Exit:** one task goes from prompt to pushed branch by hand.
357 +
358 +### Phase 1 — Private alpha, single happy path (3–4 weeks)
359 +- `AgentRun` model + state machine; web `AgentRunsController`; `AgentRunJob` calling
360 + ECS `RunTask`.
361 +- Repo "Agent" tab: run form, live transcript via SSE/Turbo, final diff view.
362 +- Branch push + a minimal compare view (full PR model can lag one phase).
363 +- Caps: wall-clock timeout, concurrency = 1, token budget tied to metering.
364 +- Internal-only feature flag; dogfood on our own repos.
365 +- **Exit:** team members trigger runs from the web and review the diff.
366 +
367 +### Phase 2 — Beta, productized (4–6 weeks)
368 +- Minimal `PullRequest` model + diff/review UI so agent output is a real PR.
369 +- Steerability: follow-up messages to a running agent, cancel, re-run.
370 +- Metering + pricing surface: `AgentUsage`, plan caps, billing page copy.
371 +- Hardening: enforce egress allowlist, isolation review, token scoping, per-plan
372 + concurrency, abuse controls.
373 +- Trigger surfaces: `sigit cloud run` in the CLI and a desktop entry point, both
374 + hitting `Api::V1::AgentRunsController`.
375 +- Cold-start work (prebaked base images; warm pool if needed).
376 +- **Exit:** invite-only beta with real external users and real billing.
377 +
378 +### Phase 3 — GA and advanced (ongoing)
379 +- Issues + assign-an-issue-to-the-agent (needs an issue model).
380 +- Custom environments: repo-level `.sigit/agent.yml` (setup steps, allowed
381 + commands, language matrix), the analog of Copilot's environment customization.
382 +- Firecracker microVM sandboxes + snapshot warm pools for isolation and sub-second
383 + starts at volume.
384 +- Dependency/repo caching (S3/EFS layers) for faster, cheaper runs.
385 +- Parallel runs, plan-then-execute, multi-file refactors at scale.
386 +- Observability and an eval harness (task success rate, PR acceptance rate, cost
387 + per accepted PR) to drive model-tier and prompt tuning.
388 +
389 +---
390 +
391 +## 7. Key decisions and open dependencies
392 +
393 +- **PRs/issues do not exist on sigit.si yet.** A minimal PR surface is on the
394 + critical path (Phase 2); issues gate the issue-to-PR flow (Phase 3). Confirm
395 + scope early.
396 +- **Sandbox isolation level.** Fargate for v1; reassess against the threat model
397 + before opening to untrusted external users at volume; Firecracker is the scale
398 + answer.
399 +- **Inference token model.** Mint per-run, server-side, budget-bounded. This also
400 + closes the standing "inference token ↔ Onde Cloud auth" gap for this path.
401 +- **COGS visibility.** Must meter sandbox time *and* tokens from day one, or
402 + pricing flies blind.
403 +- **No AWS SDK in the app yet.** Adds a new infra dependency and IAM/VPC footprint
404 + to own and secure.
405 +- **Naming.** "siGit Cloud Agent" is a working name; align with the existing
406 + brand rules (`siGit Code`, `siGit Code Cloud`, `Onde Cloud`) before anything
407 + ships to users.
408 +
409 +---
410 +
411 +## 8. North-star and success metrics
412 +
413 +- **North star:** accepted agent PRs per active paying user per month.
414 +- **Quality:** task success rate (run produces a mergeable PR without human
415 + fixes), PR acceptance rate, median time-to-PR.
416 +- **Economics:** cost per accepted PR (sandbox + tokens), gross margin per plan.
417 +- **Reliability/safety:** zero egress-allowlist escapes, zero credential leaks,
418 + sandbox p95 cold start, run failure rate by cause.
419 +```