Add cloud-agent product plan and roadmap (internal)
Strategy, architecture, and phased roadmap for a Copilot-style cloud coding agent on AWS, plus the pricing model. Private-repo planning doc. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Seto Elkahfi committed
Jun 26, 2026 at 16:43 UTC
c4eafd9389344575aa9ca6b2422de51302aefd63
1 file changed
+419
docs/product/cloud-agent-plan.md
new
+419
@@ -0,0 +1,419 @@
1
+# siGit Cloud Agent — product plan and roadmap
2
+
3
+Status: draft v1 (2026-06-26). Owner: product/eng. Audience: internal (private repo).
4
+
5
+A GitHub-Copilot-coding-agent-style product built into sigit.si and siGit Code
6
+Cloud: you give it a task against a repo you host on sigit.si, an isolated AWS
7
+sandbox runs the siGit Code agent loop against that repo, and it returns a branch
8
+plus a proposed pull request. This document is the strategy, the architecture, and
9
+the phased roadmap.
10
+
11
+---
12
+
13
+## 1. What we are building (and what it is not)
14
+
15
+**siGit Cloud Agent** is an asynchronous, server-side coding agent. The user
16
+describes a task ("add pagination to the repos list", "fix the failing
17
+auth test", "upgrade Rails to 8.1"), points it at one of their sigit.si repos and
18
+a base branch, and walks away. The platform provisions an ephemeral AWS sandbox,
19
+clones the repo into it, runs siGit Code headless against the task with cloud
20
+inference, lets it edit files and run build/test commands, then pushes a head
21
+branch back to sigit.si and opens a pull request for human review.
22
+
23
+It is the cloud, autonomous sibling of the two things siGit already ships:
24
+
25
+- **siGit Code** (public Rust CLI / ACP agent) runs *locally / on the desktop*,
26
+ interactively, driven by the developer at their keyboard.
27
+- **siGit Cloud Agent** runs *in our cloud*, asynchronously, driven by a task and
28
+ reviewed afterward through a PR.
29
+
30
+Reference points in the market: GitHub Copilot coding agent (assign an issue, it
31
+opens a PR), OpenAI Codex cloud, Cursor background agents, Devin. The
32
+differentiator for us is that we already own the whole vertical: the git host, the
33
+agent, the inference, the account, and the billing all live inside siGit. We are
34
+not bolting an agent onto someone else's platform; we are completing a platform we
35
+already run.
36
+
37
+What it is **not**, in v1: it is not an autonomous merger (a human reviews and
38
+merges), not a long-lived persistent dev environment (sandboxes are ephemeral per
39
+run), and not a chat product (the chat tier already exists; this is the
40
+task-to-PR product on top of it).
41
+
42
+---
43
+
44
+## 2. Why siGit is unusually well positioned
45
+
46
+The expensive parts of a cloud-agent product already exist in this repo and the
47
+sibling repos. The new work is mostly orchestration and isolation, not net-new
48
+agent or inference plumbing.
49
+
50
+| Capability the product needs | Already exists | Where |
51
+|---|---|---|
52
+| A git host the agent can clone from and push to | yes | `git_http_controller` (smart-HTTP, bare repos on disk, `Repository#disk_path`) |
53
+| The coding agent itself (edit/run/test loop, tool calling) | yes | `sigit` (public Rust CLI/ACP agent), `onde-cloud` tool-call mapping |
54
+| Hosted inference with auth, identity masking, metering | yes | `Api::V1::ChatCompletionsController` → `OndeCloudService` → Onde Cloud |
55
+| Accounts and a trust boundary that mints scoped tokens | yes | smbCloud auth, `Api::V1` sessions/me, `git_token` exchange |
56
+| Subscription gating + monthly metering | yes | `Subscription`, `CloudUsage`, `User#entitled_to_cloud?` |
57
+
58
+What is **missing** and must be built (see roadmap):
59
+
60
+1. **An execution plane on AWS** (the sandbox machine) and the orchestration to
61
+ drive it. This is the heart of the project.
62
+2. **A control plane in Rails**: an `AgentRun` resource, its lifecycle, log
63
+ streaming, and web UI.
64
+3. **Pull requests** (and ideally issues). sigit.si has repos, blobs, commits, and
65
+ stars, but no `PullRequest` model today. The agent's output is a PR, so a
66
+ minimal PR/diff/review surface is a hard dependency. This can be scoped down to
67
+ a "compare and open PR" view for v1.
68
+4. **Per-run scoped credentials**: short-lived git-push and inference tokens minted
69
+ server-side, budget-bounded, never long-lived in the sandbox.
70
+5. **A new metering/pricing dimension**: agent runs consume sandbox compute *and*
71
+ inference tokens, so COGS has two drivers, not one.
72
+
73
+---
74
+
75
+## 3. Architecture
76
+
77
+Two planes plus the existing inference path. Control plane is Rails (sigit-si);
78
+execution plane is AWS; inference reuses the existing `/api/v1/chat/completions`
79
+proxy unchanged.
80
+
81
+```
82
+┌─ Control plane (Rails, sigit-si) ───────────────────────────────────────────┐
83
+│ Web UI: repo "Agent" tab → run form, live transcript, diff, "Open PR" │
84
+│ AgentRunsController + Api::V1::AgentRunsController (create/show/cancel/log) │
85
+│ AgentRun model (lifecycle state machine) │
86
+│ AgentRunJob → provisions sandbox, monitors, collects result │
87
+│ Mints per-run scoped tokens (git push + inference), enforces caps/metering │
88
+└──────────────────────────────────────────────────────────────────────────────┘
89
+ │ RunTask (aws-sdk) ▲ SSE/webhook: logs, status, diff
90
+ ▼ │
91
+┌─ Execution plane (AWS) ─────────────────────────────────────────────────────┐
92
+│ Ephemeral sandbox (Fargate task v1 → Firecracker microVM at scale) │
93
+│ ├─ clones repo from sigit.si over smart-HTTP (scoped git token) │
94
+│ ├─ runs siGit Code headless against the task │
95
+│ │ OPENAI_BASE_URL=https://sigit.si/api/v1 (per-run inference token) │
96
+│ ├─ edits files, runs build/test in restricted shell │
97
+│ └─ pushes head branch back to sigit.si (scoped git token) │
98
+│ Private subnet, egress allowlist (sigit.si + package registries only) │
99
+│ Hard caps: wall-clock, CPU/mem, token budget, max tool calls │
100
+└──────────────────────────────────────────────────────────────────────────────┘
101
+ │ /v1/chat/completions (per-run token)
102
+ ▼
103
+┌─ Inference (unchanged) ─────────────────────────────────────────────────────┐
104
+│ Api::V1::ChatCompletionsController → OndeCloudService → Onde Cloud → upstream│
105
+│ Existing entitlement gate, allowance metering, and identity masking apply. │
106
+└──────────────────────────────────────────────────────────────────────────────┘
107
+```
108
+
109
+### 3.1 Control plane (Rails)
110
+
111
+**`AgentRun` model** (new table). Belongs to `user` and `repository`. Fields:
112
+
113
+- `status`: `queued`, `provisioning`, `running`, `pushing`, `needs_input`,
114
+ `completed`, `failed`, `canceled` (state machine; one-way transitions logged).
115
+- `task_prompt` (text), `base_branch`, `head_branch` (generated, e.g.
116
+ `agent/<run-id>-<slug>`), `pull_request_id` (nullable until pushed).
117
+- `sandbox_ref` (ECS task ARN / microVM id), `region`.
118
+- Budgets and accounting: `token_budget`, `tokens_used`, `wall_clock_limit_s`,
119
+ `started_at`, `finished_at`, `exit_reason`.
120
+- `transcript_url` (S3 pointer for the full log), plus a tail kept in Postgres for
121
+ the live view.
122
+
123
+**Controllers**: a web `AgentRunsController` (HTML, Turbo) under the repo, and an
124
+`Api::V1::AgentRunsController` so the CLI and desktop can trigger and follow runs.
125
+Actions: `create`, `index`, `show`, `cancel`, `messages#create` (steer a running
126
+agent), and a `logs` SSE endpoint that relays the live transcript (reuse the
127
+`ActionController::Live` pattern already used for chat streaming).
128
+
129
+**`AgentRunJob`** (Active Job, on the existing DB-backed queue): transitions the
130
+run to `provisioning`, calls AWS to start the sandbox, persists the sandbox ref,
131
+then hands off to monitoring. Cancellation and timeout both tear the sandbox down.
132
+
133
+**Web UI**: a new "Agent" (or "Tasks") tab on the repository page. A run form
134
+(task prompt, base branch, optional model tier). A live transcript panel (Turbo
135
+Streams fed by the SSE relay). On completion, a diff view and an "Open pull
136
+request" action.
137
+
138
+### 3.2 Execution plane (AWS) — the sandbox machine
139
+
140
+This is the core AWS decision and the riskiest surface, because the sandbox runs
141
+build and test commands over user code.
142
+
143
+**Runtime choice.**
144
+
145
+- **v1: AWS Fargate (ECS) ephemeral tasks.** One task per run. Scales to zero, pay
146
+ per second, decent container isolation, no servers to manage, and `RunTask` is a
147
+ single SDK call from a Rails job. Fast to ship. This is the recommendation for
148
+ v1.
149
+- **At scale: Firecracker microVMs.** For stronger isolation of untrusted code and
150
+ for snapshot/restore warm pools (sub-second starts), move the sandbox to
151
+ Firecracker microVMs on bare-metal EC2 (the model E2B / Modal / Codex-style
152
+ sandboxes use). More operational weight; defer until run volume and the threat
153
+ model justify it. gVisor or Kata on EC2 is a middle option if Fargate isolation
154
+ proves insufficient before we are ready for Firecracker.
155
+
156
+**The sandbox image.** A container that bundles headless siGit Code plus a base
157
+toolchain (git, common language runtimes). The agent boots, reads the run spec
158
+from an injected env/file, clones, works, and pushes. Per-language base images (or
159
+a `.sigit/agent.yml` setup step, see Phase 3) keep cold builds fast.
160
+
161
+**Orchestration.** v1 keeps it simple: the Rails `AgentRunJob` calls ECS `RunTask`
162
+directly via `aws-sdk-ecs`, passing the run spec as container overrides, and polls
163
+task status (or receives EventBridge task-state-change events into a webhook). If
164
+the lifecycle grows (retries, multi-step, fan-out), promote to Step Functions.
165
+Avoid Step Functions on day one; it is premature.
166
+
167
+**Networking and isolation (load-bearing for safety).**
168
+
169
+- Sandbox runs in a **private subnet**. Egress through a NAT restricted by an
170
+ **allowlist**: sigit.si (git + inference) and an explicit set of package
171
+ registries (rubygems, npm, pypi, crates, etc.). Everything else is denied. This
172
+ is the primary control against data exfiltration, SSRF against internal
173
+ services, and crypto-mining abuse.
174
+- **No inbound.** The sandbox is not reachable from the internet.
175
+- Per-run IAM role scoped to only what the task needs; no broad account access
176
+ inside the sandbox.
177
+
178
+**Credentials (mint short-lived, never long-lived).** The sandbox receives:
179
+
180
+- a **git token scoped to the single repo and ideally the single head branch**,
181
+ valid for the run only (extends the existing `git_token` exchange);
182
+- an **inference token** minted per run, carrying the user's entitlement and a
183
+ hard token budget, pointed at `https://sigit.si/api/v1`.
184
+
185
+Both expire when the run ends. The sandbox never holds `app_secret` or any
186
+long-lived credential, mirroring the existing public-client rule.
187
+
188
+**Logs and artifacts.** The agent streams transcript chunks back to Rails (the SSE
189
+relay) for the live view and writes the full transcript and build logs to S3
190
+(pointer stored on `AgentRun`). CloudWatch captures infra-level logs.
191
+
192
+### 3.3 Inference path (reused as-is)
193
+
194
+The sandboxed agent sets `OPENAI_BASE_URL=https://sigit.si/api/v1` and uses the
195
+per-run inference token as its bearer. That flows through the **existing**
196
+`Api::V1::ChatCompletionsController`: entitlement gate, allowance metering, and the
197
+`SIGIT_IDENTITY_PROMPT` identity masking all apply with no change. This is a major
198
+reason the project is tractable: the autonomous agent is just another client of an
199
+inference endpoint we already operate and protect.
200
+
201
+This also gives a clean answer to the existing "inference token ↔ Onde Cloud auth"
202
+open item: for cloud-agent runs the token is minted server-side with a budget, so
203
+there is no public client holding credentials at all.
204
+
205
+### 3.4 Git and PR flow
206
+
207
+The agent pushes the head branch via the existing `git-receive-pack` endpoint. On
208
+push, the platform creates (or links) the PR record. Because **no PR model exists
209
+yet**, scope for v1:
210
+
211
+- minimal `PullRequest` model (base/head branch, repo, author, status, title,
212
+ body), a diff/compare view (we already render blobs and commits, so the diff
213
+ renderer is incremental), and "open / close / merge" actions for the reviewer;
214
+- the agent fills in title and body from its summary. PR prose must stay neutral
215
+ and must never name the upstream model or provider (same rule as chat).
216
+
217
+Issues (assign-an-issue-to-the-agent) are a Phase 3 surface and depend on an issue
218
+model that also does not exist yet.
219
+
220
+---
221
+
222
+## 4. Safety, guardrails, and abuse
223
+
224
+Principal-engineer non-negotiables, because this executes code on our infra on
225
+behalf of users:
226
+
227
+- **Hard caps per run**: wall-clock timeout, CPU/memory limits, inference token
228
+ budget (tied to `CloudUsage`/a new `AgentUsage`), max tool calls, max sandbox
229
+ lifetime. A run that blows any cap is killed and marked `failed` with a reason.
230
+- **Network egress allowlist** (section 3.2). The single most important control.
231
+- **Ephemeral, scoped credentials** only. Nothing long-lived in the sandbox.
232
+- **Human-in-the-loop by default**: the agent proposes a PR; it does not merge. No
233
+ auto-merge in v1.
234
+- **Concurrency limits per plan**: caps simultaneous runs per user to bound spend
235
+ and abuse.
236
+- **Identity hygiene**: transcripts, PR titles/bodies, and error messages stay
237
+ neutral; never disclose the upstream model or provider (existing rule extends to
238
+ agent output).
239
+- **Cancellation is real**: cancel tears down the sandbox and revokes the run's
240
+ tokens.
241
+- **Idempotency**: run creation and sandbox start are idempotent so retries cannot
242
+ double-spend.
243
+
244
+A short threat-model doc is a Phase 0 deliverable (exfiltration, SSRF to internal
245
+metadata endpoints, resource abuse / mining, secret leakage from the user's own
246
+repo, prompt injection from repo contents steering the agent).
247
+
248
+---
249
+
250
+## 5. Pricing and packaging
251
+
252
+Agent runs cost us **sandbox compute (Fargate per-second) + inference tokens**, so
253
+metering must capture both and price above blended COGS. The numbers below are a
254
+working model with the inputs we cannot yet read from this repo marked
255
+`[FILL FROM PHASE 0]`. The metering schema and billing copy should be designed
256
+against this formula now; the dollar figures get populated once the Phase 0
257
+dogfood produces measured per-run token counts.
258
+
259
+### 5.1 Unit economics: cost per run (COGS)
260
+
261
+```
262
+cost_per_run = sandbox_cost + inference_cost
263
+
264
+sandbox_cost = vcpu_count * $0.04048/vcpu-hr * hours
265
+ + mem_gb * $0.004445/gb-hr * hours (Fargate, us-east-1)
266
+
267
+inference_cost = tokens_per_run / 1_000_000 * blended_upstream_cost_per_mtok
268
+```
269
+
270
+Worked example, sandbox side (known, fixed): 1 vCPU + 2 GB for a 20-minute run:
271
+
272
+```
273
+(1 * 0.04048 + 2 * 0.004445) * (20/60) = $0.0165 per run
274
+```
275
+
276
+Sandbox compute is rounding error. **Inference dominates**, and it is the unknown:
277
+
278
+```
279
+tokens_per_run = [FILL FROM PHASE 0] (expect 200K – 2M+)
280
+blended_upstream_cost_per_mtok = [FILL: Onde/upstream blended $/Mtok]
281
+
282
+inference_cost_per_run = (tokens_per_run / 1e6) * blended_cost_per_mtok
283
+cost_per_run = 0.0165 + inference_cost_per_run
284
+```
285
+
286
+Conclusion that holds regardless of the blanks: **price on metered inference per
287
+run; treat sandbox-minutes as a kill-switch guardrail, not a billing axis.**
288
+
289
+### 5.2 Packaging (allowance model, mirrors `Subscription::CLOUD_ALLOWANCE`)
290
+
291
+Agent runs are an included monthly bundle per plan, with overage or an add-on for
292
+heavy use. Suggested `Subscription` constant shape:
293
+
294
+```ruby
295
+# Included siGit Cloud Agent runs per billing period. Tunable pricing knob.
296
+AGENT_RUNS = { "pro" => [FILL], "team" => [FILL] }.freeze # e.g. pro 20, team 75
297
+TRIAL_AGENT_RUNS = [FILL] # e.g. 2
298
+AGENT_OVERAGE_PER_RUN = [FILL] # $/run past the bundle
299
+```
300
+
301
+| Plan | Price (today) | Included agent runs/mo | Overage | Notes |
302
+|---|---|---|---|---|
303
+| Free | $0 | 0 | n/a | local siGit Code only, no cloud agent |
304
+| Trial | $0 (14 days) | `[FILL: ~2]` | n/a | felt-value, hard run cap |
305
+| Pro | $20/mo (existing) | `[FILL: ~20]` | `[FILL: $/run]` | bundle sized so COGS < ~X% of $20 |
306
+| Team | (existing) | `[FILL: ~75]` | `[FILL: $/run]` | higher bundle + concurrency |
307
+
308
+Sizing rule for the bundle: pick included-runs so that **bundle COGS stays under a
309
+target fraction of plan price** (e.g. included-runs * cost_per_run ≤ 40% of MRR),
310
+leaving margin for the chat allowance those plans already include. With
311
+`cost_per_run` from 5.1 unknown, the bundle count is the lever set last.
312
+
313
+### 5.3 Metering
314
+
315
+Add `AgentUsage` keyed by `(user, billing_period_key)` (parallel to `CloudUsage`),
316
+recording per period: `runs_count`, `agent_seconds`, `tokens_used`. Runs decrement
317
+the plan's run bundle; tokens already flow through the existing cloud allowance and
318
+its 429 cap, so a single run can never escape the token budget. `GET
319
+/api/v1/billing` gains `agent_runs_used` + `agent_runs_allowance` alongside the
320
+existing cloud fields.
321
+
322
+### 5.4 Guardrails (worst-case cost containment)
323
+
324
+- **Per-run token budget** minted into the run token, independent of the monthly
325
+ allowance, so one pathological run is bounded.
326
+- **Per-run wall-clock + sandbox-lifetime cap** (kills runaway compute).
327
+- **Per-plan concurrency cap** bounds simultaneous spend.
328
+- **Trial run cap** (`TRIAL_AGENT_RUNS`) prevents trial-driven upstream bills,
329
+ mirroring `TRIAL_ALLOWANCE`.
330
+
331
+### 5.5 What Phase 0 must measure to finalize this
332
+
333
+1. **`tokens_per_run`** distribution (p50/p90/p99) over 20–30 real dogfood runs.
334
+2. **`blended_upstream_cost_per_mtok`** (from the Onde/upstream cost, internal).
335
+3. Resulting **`cost_per_run`** distribution, then back-solve included-run bundles
336
+ per tier against the target-margin rule in 5.2.
337
+
338
+Until 1 and 2 are real, every dollar above is a placeholder by design.
339
+
340
+---
341
+
342
+## 6. Roadmap (phased)
343
+
344
+Estimates are calendar weeks for a small team; adjust to staffing. Each phase ends
345
+with a demoable, dogfoodable increment.
346
+
347
+### Phase 0 — Foundations and spikes (2–3 weeks)
348
+- Decide sandbox runtime: **Fargate for v1** (documented path to Firecracker).
349
+- Build the headless siGit Code container image; prove end to end **manually**: a
350
+ container clones a sigit.si repo, runs the agent against a task with inference via
351
+ `/api/v1`, edits files, runs a test, and pushes a branch.
352
+- Per-run scoped token minting (git + inference) in Rails.
353
+- Threat model + egress allowlist design.
354
+- Add `aws-sdk-ecs`/`aws-sdk-core` to the Gemfile; stand up the VPC/subnet/NAT and
355
+ the task definition in IaC.
356
+- **Exit:** one task goes from prompt to pushed branch by hand.
357
+
358
+### Phase 1 — Private alpha, single happy path (3–4 weeks)
359
+- `AgentRun` model + state machine; web `AgentRunsController`; `AgentRunJob` calling
360
+ ECS `RunTask`.
361
+- Repo "Agent" tab: run form, live transcript via SSE/Turbo, final diff view.
362
+- Branch push + a minimal compare view (full PR model can lag one phase).
363
+- Caps: wall-clock timeout, concurrency = 1, token budget tied to metering.
364
+- Internal-only feature flag; dogfood on our own repos.
365
+- **Exit:** team members trigger runs from the web and review the diff.
366
+
367
+### Phase 2 — Beta, productized (4–6 weeks)
368
+- Minimal `PullRequest` model + diff/review UI so agent output is a real PR.
369
+- Steerability: follow-up messages to a running agent, cancel, re-run.
370
+- Metering + pricing surface: `AgentUsage`, plan caps, billing page copy.
371
+- Hardening: enforce egress allowlist, isolation review, token scoping, per-plan
372
+ concurrency, abuse controls.
373
+- Trigger surfaces: `sigit cloud run` in the CLI and a desktop entry point, both
374
+ hitting `Api::V1::AgentRunsController`.
375
+- Cold-start work (prebaked base images; warm pool if needed).
376
+- **Exit:** invite-only beta with real external users and real billing.
377
+
378
+### Phase 3 — GA and advanced (ongoing)
379
+- Issues + assign-an-issue-to-the-agent (needs an issue model).
380
+- Custom environments: repo-level `.sigit/agent.yml` (setup steps, allowed
381
+ commands, language matrix), the analog of Copilot's environment customization.
382
+- Firecracker microVM sandboxes + snapshot warm pools for isolation and sub-second
383
+ starts at volume.
384
+- Dependency/repo caching (S3/EFS layers) for faster, cheaper runs.
385
+- Parallel runs, plan-then-execute, multi-file refactors at scale.
386
+- Observability and an eval harness (task success rate, PR acceptance rate, cost
387
+ per accepted PR) to drive model-tier and prompt tuning.
388
+
389
+---
390
+
391
+## 7. Key decisions and open dependencies
392
+
393
+- **PRs/issues do not exist on sigit.si yet.** A minimal PR surface is on the
394
+ critical path (Phase 2); issues gate the issue-to-PR flow (Phase 3). Confirm
395
+ scope early.
396
+- **Sandbox isolation level.** Fargate for v1; reassess against the threat model
397
+ before opening to untrusted external users at volume; Firecracker is the scale
398
+ answer.
399
+- **Inference token model.** Mint per-run, server-side, budget-bounded. This also
400
+ closes the standing "inference token ↔ Onde Cloud auth" gap for this path.
401
+- **COGS visibility.** Must meter sandbox time *and* tokens from day one, or
402
+ pricing flies blind.
403
+- **No AWS SDK in the app yet.** Adds a new infra dependency and IAM/VPC footprint
404
+ to own and secure.
405
+- **Naming.** "siGit Cloud Agent" is a working name; align with the existing
406
+ brand rules (`siGit Code`, `siGit Code Cloud`, `Onde Cloud`) before anything
407
+ ships to users.
408
+
409
+---
410
+
411
+## 8. North-star and success metrics
412
+
413
+- **North star:** accepted agent PRs per active paying user per month.
414
+- **Quality:** task success rate (run produces a mergeable PR without human
415
+ fixes), PR acceptance rate, median time-to-PR.
416
+- **Economics:** cost per accepted PR (sandbox + tokens), gross margin per plan.
417
+- **Reliability/safety:** zero egress-allowlist escapes, zero credential leaks,
418
+ sandbox p95 cold start, run failure rate by cause.
419
+```