@setoelkahfi / sigit / commits / 240f46c

Add headless one-shot mode: sigit run

A new subcommand executes a single task non-interactively and exits, built for cloud runners (siGit Code Cloud Agent) and scripting: sigit run [--prompt <text> | --prompt-file <path>] [--cwd <dir>] [--max-rounds <n>] [--output jsonl|text] - Progress streams as JSONL events on stdout (run_started, turn_text, tool_call, tool_result, compaction, result); logs stay on stderr. - Exit codes: 0 completed, 1 run failed, 2 usage/config error. - Provider resolves like other modes (OPENAI_BASE_URL/OPENAI_API_KEY override first, signed-in cloud next) and never falls back to on-device inference, so a headless host cannot trigger a multi-GB model download. - The tool loop mirrors the ACP prompt handler: permission gate, auto-compaction between rounds, forced text reply on the last round. A tool that would prompt (Decision::Ask) is declined with a pointer to SIGIT_PERMISSIONS=allow instead of hanging. - Integration tests drive the real binary against a scripted OpenAI-compatible endpoint, covering the tool-free path, the deny-without-override path, and the allow-override path. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A1xu3s7ey2EqDNjgtt2XFA Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

siGit Code Session committed Jul 6, 2026 at 17:46 UTC 240f46c9e96c716da6ba41dbf185218375b11056
5 files changed +821 -1
CHANGELOG.md
+6
@@ -1,5 +1,11 @@
1 # Changelog
2
3 +## Unreleased
4 +
5 +### What changed
6 +
7 +- New headless one-shot mode: `sigit run [--prompt <text> | --prompt-file <path>] [--cwd <dir>] [--max-rounds <n>] [--output jsonl|text]` runs a single task non-interactively and exits. Progress is emitted as JSONL events on stdout (`run_started`, `turn_text`, `tool_call`, `tool_result`, `compaction`, and a final `result`), logs stay on stderr, and exit codes are 0 (completed), 1 (run failed), 2 (usage/configuration error). The provider resolves like other modes (`OPENAI_BASE_URL`/`OPENAI_API_KEY` override first), never falling back to on-device inference, and a tool that would normally prompt for permission is declined with a pointer to `SIGIT_PERMISSIONS=allow` instead of hanging. This is the execution engine for siGit Code Cloud Agent runners
8 +
9 ## 1.3.2
10
11 Adds a tool permission system with plan mode, durable sessions with context
CLAUDE.md
+6 -1
@@ -15,7 +15,12 @@ is a TTY:
15 it relies on fd redirection to keep logs out of the TUI, so Windows gets ACP mode only.
16
17 Before the TTY/ACP split, `main` also dispatches the account subcommands `sigit login`,
18 -`sigit logout`, `sigit whoami` (see `src/main.rs` `main()`).
18 +`sigit logout`, `sigit whoami`, and the headless one-shot mode `sigit run` (see `src/main.rs`
19 +`main()`). `sigit run` (`src/headless.rs`) executes a single task non-interactively — JSONL
20 +progress events on stdout, logs on stderr, exit 0/1/2 — and is what cloud runners (siGit Code
21 +Cloud Agent) drive; it never falls back to on-device inference and declines permission prompts
22 +unless `SIGIT_PERMISSIONS=allow` is set. Its tool loop mirrors `handle_prompt`; keep the two in
23 +sync.
24
25 ## Working in this repo
26
src/headless.rs new
+497
@@ -0,0 +1,497 @@
1 +//! Headless one-shot mode: `sigit run`.
2 +//!
3 +//! Runs a single task non-interactively and exits: the prompt arrives as a
4 +//! CLI flag, progress is emitted as JSONL events on stdout (one object per
5 +//! line), and logs stay on stderr. Built for cloud runners (siGit Code Cloud
6 +//! Agent) and scripting, where nobody is present to answer a permission
7 +//! prompt — on `Decision::Ask` the tool is declined with a pointer to
8 +//! `SIGIT_PERMISSIONS=allow` instead of blocking.
9 +//!
10 +//! The tool loop mirrors the ACP prompt handler (`handle_prompt` in
11 +//! `main.rs`): permission gate → execute → feed results back, with
12 +//! auto-compaction between rounds and a forced text reply on the final
13 +//! round. Keep the two in sync when changing loop semantics.
14 +
15 +use std::io::Write as _;
16 +use std::path::PathBuf;
17 +use std::sync::Arc;
18 +
19 +use serde_json::json;
20 +
21 +use crate::backend::{self, InferenceBackend, OpenAiBackend, ToolResult, ToolSpec};
22 +use crate::{permissions, provider, tools};
23 +
24 +/// Headless runs default to a higher round cap than interactive prompts: an
25 +/// autonomous task routinely needs long edit/build/test chains and there is
26 +/// no user present to re-prompt a stopped run.
27 +const DEFAULT_MAX_ROUNDS: usize = 40;
28 +
29 +/// Cap on `arguments`/`output` strings embedded in JSONL events. Full outputs
30 +/// still reach the model; events only need enough for a live transcript.
31 +const EVENT_FIELD_MAX_CHARS: usize = 4_000;
32 +
33 +const USAGE: &str = "usage: sigit run [--prompt <text> | --prompt-file <path>] \
34 + [--cwd <dir>] [--max-rounds <n>] [--output jsonl|text]";
35 +
36 +#[derive(Debug, Clone, Copy, PartialEq, Eq)]
37 +pub enum OutputMode {
38 + Jsonl,
39 + Text,
40 +}
41 +
42 +#[derive(Debug)]
43 +pub struct HeadlessOptions {
44 + pub prompt: String,
45 + pub cwd: PathBuf,
46 + pub max_rounds: usize,
47 + pub output: OutputMode,
48 +}
49 +
50 +/// The value following a flag, or a usage error naming the flag.
51 +fn next_value(args: &mut impl Iterator<Item = String>, flag: &str) -> Result<String, String> {
52 + args.next()
53 + .ok_or_else(|| format!("{flag} needs a value\n{USAGE}"))
54 +}
55 +
56 +/// Parse `sigit run` arguments (everything after the subcommand).
57 +pub fn parse_args(mut args: impl Iterator<Item = String>) -> Result<HeadlessOptions, String> {
58 + let mut prompt: Option<String> = None;
59 + let mut cwd: Option<PathBuf> = None;
60 + let mut max_rounds = DEFAULT_MAX_ROUNDS;
61 + let mut output = OutputMode::Jsonl;
62 +
63 + while let Some(flag) = args.next() {
64 + match flag.as_str() {
65 + "--prompt" => {
66 + let value = next_value(&mut args, "--prompt")?;
67 + if prompt.is_some() {
68 + return Err(format!("give --prompt or --prompt-file once\n{USAGE}"));
69 + }
70 + prompt = Some(value);
71 + }
72 + "--prompt-file" => {
73 + let path = next_value(&mut args, "--prompt-file")?;
74 + if prompt.is_some() {
75 + return Err(format!("give --prompt or --prompt-file once\n{USAGE}"));
76 + }
77 + let text = std::fs::read_to_string(&path)
78 + .map_err(|error| format!("cannot read --prompt-file {path}: {error}"))?;
79 + prompt = Some(text);
80 + }
81 + "--cwd" => {
82 + cwd = Some(PathBuf::from(next_value(&mut args, "--cwd")?));
83 + }
84 + "--max-rounds" => {
85 + max_rounds = next_value(&mut args, "--max-rounds")?
86 + .parse::<usize>()
87 + .ok()
88 + .filter(|n| *n > 0)
89 + .ok_or_else(|| format!("--max-rounds needs a positive integer\n{USAGE}"))?;
90 + }
91 + "--output" => {
92 + output = match next_value(&mut args, "--output")?.as_str() {
93 + "jsonl" => OutputMode::Jsonl,
94 + "text" => OutputMode::Text,
95 + other => return Err(format!("unknown --output {other}\n{USAGE}")),
96 + };
97 + }
98 + other => return Err(format!("unknown argument {other}\n{USAGE}")),
99 + }
100 + }
101 +
102 + let prompt = prompt
103 + .map(|text| text.trim().to_string())
104 + .filter(|text| !text.is_empty())
105 + .ok_or_else(|| format!("a non-empty --prompt or --prompt-file is required\n{USAGE}"))?;
106 +
107 + let cwd = cwd.unwrap_or_else(|| PathBuf::from("."));
108 + let cwd = cwd
109 + .canonicalize()
110 + .map_err(|error| format!("--cwd {}: {error}", cwd.display()))?;
111 +
112 + Ok(HeadlessOptions {
113 + prompt,
114 + cwd,
115 + max_rounds,
116 + output,
117 + })
118 +}
119 +
120 +/// Entry point for `sigit run`. Never returns on failure paths — exits the
121 +/// process with 0 (run completed), 1 (run failed), or 2 (usage/config error).
122 +pub async fn run(args: impl Iterator<Item = String>) -> anyhow::Result<()> {
123 + let options = match parse_args(args) {
124 + Ok(options) => options,
125 + Err(message) => {
126 + eprintln!("sigit run: {message}");
127 + std::process::exit(2);
128 + }
129 + };
130 +
131 + // Provider: the explicit override (env / providers.toml) first — this is
132 + // how a cloud runner injects a per-run endpoint and token — then the
133 + // signed-in cloud as a convenience. Never fall back to on-device: a
134 + // headless host should not silently download a multi-GB model.
135 + let Some(config) =
136 + provider::active_provider().or_else(|| provider::cloud_tier_provider("large"))
137 + else {
138 + eprintln!(
139 + "sigit run: no inference provider configured. Set OPENAI_BASE_URL and \
140 + OPENAI_API_KEY (and optionally SIGIT_MODEL), configure providers.toml, \
141 + or sign in with `sigit login`."
142 + );
143 + std::process::exit(2);
144 + };
145 +
146 + if let Err(error) = std::env::set_current_dir(&options.cwd) {
147 + eprintln!("sigit run: cannot enter {}: {error}", options.cwd.display());
148 + std::process::exit(2);
149 + }
150 +
151 + let system_prompt = format!(
152 + "{}\n\n{}",
153 + crate::system_prompt_for_model(true),
154 + crate::session_context_message(&options.cwd)
155 + );
156 + let backend: Arc<dyn InferenceBackend> = Arc::new(OpenAiBackend::new(
157 + config.base_url.clone(),
158 + config.api_key.clone(),
159 + config.model.clone(),
160 + Some(system_prompt),
161 + ));
162 + crate::register_subagent_factory_for(&config);
163 + let tools = crate::agent_tools_as_specs();
164 +
165 + let emitter = Emitter {
166 + mode: options.output,
167 + };
168 + emitter.event(json!({
169 + "type": "run_started",
170 + "cwd": options.cwd.display().to_string(),
171 + "model": config.model,
172 + "max_rounds": options.max_rounds,
173 + }));
174 +
175 + let rounds = match drive_loop(&backend, &tools, &options, &emitter).await {
176 + Ok((summary, rounds)) => {
177 + emitter.event(json!({
178 + "type": "result",
179 + "status": "completed",
180 + "summary": summary,
181 + "rounds": rounds,
182 + }));
183 + rounds
184 + }
185 + Err((error, rounds)) => {
186 + emitter.event(json!({
187 + "type": "result",
188 + "status": "failed",
189 + "error": error,
190 + "rounds": rounds,
191 + }));
192 + std::process::exit(1);
193 + }
194 + };
195 + log::info!("headless run completed after {rounds} tool round(s)");
196 + Ok(())
197 +}
198 +
199 +/// The tool loop. Returns `(summary, rounds)` or `(error, rounds)`.
200 +///
201 +/// Keep in sync with `handle_prompt` in `main.rs`: same permission gate, same
202 +/// auto-compaction trigger, same force-text final round.
203 +async fn drive_loop(
204 + backend: &Arc<dyn InferenceBackend>,
205 + tools: &[ToolSpec],
206 + options: &HeadlessOptions,
207 + emitter: &Emitter,
208 +) -> Result<(String, usize), (String, usize)> {
209 + // Permission decisions are per-session state; a headless process is one
210 + // session. There are no grants to accumulate (nobody can answer "always
211 + // allow"), the id only namespaces the lookup.
212 + let session = format!("headless-{}", std::process::id());
213 +
214 + let mut result = backend
215 + .send_message_with_tools(&options.prompt, tools, None)
216 + .await
217 + .map_err(|error| (format!("inference failed: {error}"), 0))?;
218 + emitter.turn_text(&result.text);
219 +
220 + let mut round = 0usize;
221 +
222 + while !result.tool_calls.is_empty() && round < options.max_rounds {
223 + round += 1;
224 +
225 + // Auto-compaction: long tool runs grow history fast; fold it into a
226 + // summary before the next round rather than blowing the window.
227 + let estimate = backend::estimate_tokens(&backend.history_snapshot().await);
228 + if estimate > backend::DEFAULT_CONTEXT_TOKEN_BUDGET {
229 + match backend.compact_history(backend::COMPACT_KEEP_LAST).await {
230 + Ok(()) => {
231 + let after = backend::estimate_tokens(&backend.history_snapshot().await);
232 + emitter.event(json!({
233 + "type": "compaction",
234 + "approx_tokens_before": estimate,
235 + "approx_tokens_after": after,
236 + }));
237 + }
238 + Err(error) => log::warn!("headless compaction failed: {error}"),
239 + }
240 + }
241 +
242 + let mut tool_results = Vec::new();
243 + for call in &result.tool_calls {
244 + emitter.tool_call(call);
245 + let (output, denied) = match permissions::decision_for(&session, &call.name) {
246 + permissions::Decision::Allow => (
247 + tools::execute_tool(&call.name, &call.arguments).await,
248 + false,
249 + ),
250 + permissions::Decision::Deny(reason) => {
251 + log::info!("headless: {} denied by policy", call.name);
252 + (reason, true)
253 + }
254 + permissions::Decision::Ask => {
255 + log::info!("headless: {} needs approval, declining", call.name);
256 + (
257 + format!(
258 + "`{}` was not executed: headless mode cannot prompt for \
259 + permission. Run with SIGIT_PERMISSIONS=allow to auto-approve \
260 + mutating tools, or grant this tool in settings.toml.",
261 + call.name
262 + ),
263 + true,
264 + )
265 + }
266 + };
267 + emitter.tool_result(call, &output, denied);
268 + tool_results.push(ToolResult {
269 + tool_call_id: call.id.clone(),
270 + content: output,
271 + });
272 + }
273 +
274 + let next_tools = if round < options.max_rounds {
275 + Some(tools)
276 + } else {
277 + None // last round: force a text reply
278 + };
279 + result = backend
280 + .send_tool_results(tool_results, next_tools, None)
281 + .await
282 + .map_err(|error| (format!("inference failed: {error}"), round))?;
283 + emitter.turn_text(&result.text);
284 + }
285 +
286 + let (_think, visible) = crate::chat::strip_think_blocks(&result.text);
287 + let summary = if visible.trim().is_empty() {
288 + "The run finished without a final summary.".to_string()
289 + } else {
290 + visible.trim().to_string()
291 + };
292 + Ok((summary, round))
293 +}
294 +
295 +// ── Event output ─────────────────────────────────────────────────────────────
296 +
297 +struct Emitter {
298 + mode: OutputMode,
299 +}
300 +
301 +impl Emitter {
302 + /// Write one event. JSONL mode prints the object as-is; text mode renders
303 + /// a human-oriented line per event kind.
304 + fn event(&self, event: serde_json::Value) {
305 + match self.mode {
306 + OutputMode::Jsonl => {
307 + let mut stdout = std::io::stdout().lock();
308 + let _ = writeln!(stdout, "{event}");
309 + let _ = stdout.flush();
310 + }
311 + OutputMode::Text => {
312 + let line = match event["type"].as_str() {
313 + Some("run_started") => format!(
314 + "▶ run started in {} (model {})",
315 + event["cwd"].as_str().unwrap_or("?"),
316 + event["model"].as_str().unwrap_or("?"),
317 + ),
318 + Some("turn_text") => event["text"].as_str().unwrap_or_default().to_string(),
319 + Some("tool_call") => format!(
320 + "→ {}({})",
321 + event["name"].as_str().unwrap_or("?"),
322 + event["arguments"].as_str().unwrap_or_default(),
323 + ),
324 + Some("tool_result") => format!(
325 + "← {} ({} chars{})",
326 + event["name"].as_str().unwrap_or("?"),
327 + event["output_chars"].as_u64().unwrap_or(0),
328 + if event["denied"].as_bool().unwrap_or(false) {
329 + ", denied"
330 + } else {
331 + ""
332 + },
333 + ),
334 + Some("compaction") => "… compacted conversation history".to_string(),
335 + Some("result") => match event["status"].as_str() {
336 + Some("completed") => format!(
337 + "✔ completed\n{}",
338 + event["summary"].as_str().unwrap_or_default()
339 + ),
340 + _ => format!("✘ failed: {}", event["error"].as_str().unwrap_or("?")),
341 + },
342 + _ => event.to_string(),
343 + };
344 + if !line.is_empty() {
345 + let mut stdout = std::io::stdout().lock();
346 + let _ = writeln!(stdout, "{line}");
347 + let _ = stdout.flush();
348 + }
349 + }
350 + }
351 + }
352 +
353 + /// Emit the visible part of an assistant turn, skipping empty turns.
354 + fn turn_text(&self, raw: &str) {
355 + let (_think, visible) = crate::chat::strip_think_blocks(raw);
356 + let visible = visible.trim();
357 + if visible.is_empty() {
358 + return;
359 + }
360 + let (text, truncated) = clip(visible, EVENT_FIELD_MAX_CHARS);
361 + let mut event = json!({ "type": "turn_text", "text": text });
362 + if truncated {
363 + event["truncated"] = json!(true);
364 + }
365 + self.event(event);
366 + }
367 +
368 + fn tool_call(&self, call: &backend::ToolCall) {
369 + let (arguments, truncated) = clip(&call.arguments, EVENT_FIELD_MAX_CHARS);
370 + let mut event = json!({
371 + "type": "tool_call",
372 + "id": call.id,
373 + "name": call.name,
374 + "arguments": arguments,
375 + });
376 + if truncated {
377 + event["truncated"] = json!(true);
378 + }
379 + self.event(event);
380 + }
381 +
382 + fn tool_result(&self, call: &backend::ToolCall, output: &str, denied: bool) {
383 + let (clipped, truncated) = clip(output, EVENT_FIELD_MAX_CHARS);
384 + let mut event = json!({
385 + "type": "tool_result",
386 + "id": call.id,
387 + "name": call.name,
388 + "output_chars": output.chars().count(),
389 + "output": clipped,
390 + "denied": denied,
391 + });
392 + if truncated {
393 + event["truncated"] = json!(true);
394 + }
395 + self.event(event);
396 + }
397 +}
398 +
399 +/// Truncate to `max` characters (not bytes — always on a char boundary).
400 +fn clip(text: &str, max: usize) -> (String, bool) {
401 + if text.chars().count() <= max {
402 + (text.to_string(), false)
403 + } else {
404 + (text.chars().take(max).collect(), true)
405 + }
406 +}
407 +
408 +#[cfg(test)]
409 +mod tests {
410 + use super::*;
411 +
412 + fn args(list: &[&str]) -> impl Iterator<Item = String> {
413 + list.iter()
414 + .map(|s| s.to_string())
415 + .collect::<Vec<_>>()
416 + .into_iter()
417 + }
418 +
419 + #[test]
420 + fn parses_prompt_and_defaults() {
421 + let options = parse_args(args(&["--prompt", "fix the tests"])).expect("parses");
422 + assert_eq!(options.prompt, "fix the tests");
423 + assert_eq!(options.max_rounds, DEFAULT_MAX_ROUNDS);
424 + assert_eq!(options.output, OutputMode::Jsonl);
425 + assert!(options.cwd.is_absolute());
426 + }
427 +
428 + #[test]
429 + fn requires_a_prompt() {
430 + let error = parse_args(args(&[])).expect_err("missing prompt");
431 + assert!(error.contains("--prompt"));
432 + }
433 +
434 + #[test]
435 + fn rejects_empty_prompt() {
436 + let error = parse_args(args(&["--prompt", " "])).expect_err("blank prompt");
437 + assert!(error.contains("non-empty"));
438 + }
439 +
440 + #[test]
441 + fn rejects_prompt_and_prompt_file_together() {
442 + let file = std::env::temp_dir().join(format!("sigit-prompt-{}.txt", std::process::id()));
443 + std::fs::write(&file, "task").unwrap();
444 + let error = parse_args(args(&[
445 + "--prompt",
446 + "one",
447 + "--prompt-file",
448 + file.to_str().unwrap(),
449 + ]))
450 + .expect_err("both prompt flags");
451 + assert!(error.contains("once"));
452 + std::fs::remove_file(&file).ok();
453 + }
454 +
455 + #[test]
456 + fn reads_prompt_file() {
457 + let file = std::env::temp_dir().join(format!("sigit-promptf-{}.txt", std::process::id()));
458 + std::fs::write(&file, "task from file\n").unwrap();
459 + let options = parse_args(args(&["--prompt-file", file.to_str().unwrap()])).expect("parses");
460 + assert_eq!(options.prompt, "task from file");
461 + std::fs::remove_file(&file).ok();
462 + }
463 +
464 + #[test]
465 + fn rejects_bad_flags_and_values() {
466 + assert!(parse_args(args(&["--prompt", "x", "--max-rounds", "0"])).is_err());
467 + assert!(parse_args(args(&["--prompt", "x", "--max-rounds", "abc"])).is_err());
468 + assert!(parse_args(args(&["--prompt", "x", "--output", "yaml"])).is_err());
469 + assert!(parse_args(args(&["--bogus"])).is_err());
470 + assert!(parse_args(args(&["--prompt", "x", "--cwd", "/definitely/not/a/dir"])).is_err());
471 + }
472 +
473 + #[test]
474 + fn parses_overrides() {
475 + let options = parse_args(args(&[
476 + "--prompt",
477 + "x",
478 + "--max-rounds",
479 + "7",
480 + "--output",
481 + "text",
482 + ]))
483 + .expect("parses");
484 + assert_eq!(options.max_rounds, 7);
485 + assert_eq!(options.output, OutputMode::Text);
486 + }
487 +
488 + #[test]
489 + fn clip_is_char_boundary_safe() {
490 + let (out, truncated) = clip("héllo wörld", 5);
491 + assert_eq!(out, "héllo");
492 + assert!(truncated);
493 + let (out, truncated) = clip("short", 10);
494 + assert_eq!(out, "short");
495 + assert!(!truncated);
496 + }
497 +}
src/main.rs
+15
@@ -32,6 +32,7 @@ mod account;
32 mod backend;
33 mod chat;
34 mod credentials;
35 +mod headless;
36 mod instructions;
37 mod mcp;
38 mod models;
@@ -3130,6 +3131,20 @@ async fn main() -> anyhow::Result<()> {
3131 println!("{}", account::status_line().await);
3132 return Ok(());
3133 }
3134 + "run" => {
3135 + // Headless one-shot mode: stdout carries JSONL run events, so
3136 + // logs go to stderr exactly like ACP mode.
3137 + init_logging(false);
3138 + setup::setup_shared_model_cache();
3139 + // Best-effort MCP discovery so `mcp__*` tools are offered
3140 + // (`SIGIT_MCP=off` skips this, the cloud-runner default).
3141 + mcp::init().await;
3142 + log::info!(
3143 + "siGit v{} starting (headless run)",
3144 + env!("CARGO_PKG_VERSION")
3145 + );
3146 + return headless::run(std::env::args().skip(2)).await;
3147 + }
3148 _ => {}
3149 }
3150 }
tests/headless_run.rs new
+297
@@ -0,0 +1,297 @@
1 +//! End-to-end `sigit run` (headless mode) against the real binary.
2 +//!
3 +//! Spawns `sigit run` wired to a scripted OpenAI-compatible endpoint via the
4 +//! `OPENAI_BASE_URL` override and asserts on the JSONL event stream. Headless
5 +//! runs use non-streaming completions (no token sink), so the endpoint serves
6 +//! plain JSON chat-completion bodies, not SSE.
7 +
8 +use std::io::{BufRead, BufReader, Read, Write};
9 +use std::net::TcpListener;
10 +use std::path::Path;
11 +use std::process::{Command, Stdio};
12 +use std::sync::{Arc, Mutex};
13 +use std::time::Duration;
14 +
15 +use serde_json::{Value, json};
16 +
17 +/// One scripted JSON completion body.
18 +fn completion_text(text: &str) -> String {
19 + json!({
20 + "choices": [{"message": {"role": "assistant", "content": text}}]
21 + })
22 + .to_string()
23 +}
24 +
25 +fn completion_tool_call(id: &str, name: &str, arguments: &str) -> String {
26 + json!({
27 + "choices": [{"message": {
28 + "role": "assistant",
29 + "content": null,
30 + "tool_calls": [{
31 + "id": id,
32 + "type": "function",
33 + "function": {"name": name, "arguments": arguments},
34 + }],
35 + }}]
36 + })
37 + .to_string()
38 +}
39 +
40 +/// Serves one scripted JSON response per request and records request bodies.
41 +struct FakeEndpoint {
42 + port: u16,
43 + requests: Arc<Mutex<Vec<Value>>>,
44 +}
45 +
46 +fn start_fake_endpoint(responses: Vec<String>) -> FakeEndpoint {
47 + let listener = TcpListener::bind("127.0.0.1:0").expect("bind fake endpoint");
48 + let port = listener.local_addr().unwrap().port();
49 + let requests: Arc<Mutex<Vec<Value>>> = Arc::default();
50 + let recorded = Arc::clone(&requests);
51 + let queue = Mutex::new(std::collections::VecDeque::from(responses));
52 +
53 + std::thread::spawn(move || {
54 + for stream in listener.incoming() {
55 + let Ok(mut stream) = stream else { continue };
56 + let mut reader = BufReader::new(match stream.try_clone() {
57 + Ok(clone) => clone,
58 + Err(_) => continue,
59 + });
60 + let mut content_length = 0usize;
61 + loop {
62 + let mut line = String::new();
63 + if reader.read_line(&mut line).unwrap_or(0) == 0 {
64 + break;
65 + }
66 + let line = line.trim();
67 + if line.is_empty() {
68 + break;
69 + }
70 + if let Some(length) = line.to_ascii_lowercase().strip_prefix("content-length:") {
71 + content_length = length.trim().parse().unwrap_or(0);
72 + }
73 + }
74 + let mut body = vec![0u8; content_length];
75 + if reader.read_exact(&mut body).is_err() {
76 + continue;
77 + }
78 + if let Ok(request) = serde_json::from_slice::<Value>(&body) {
79 + recorded.lock().unwrap().push(request);
80 + }
81 + let payload = queue
82 + .lock()
83 + .unwrap()
84 + .pop_front()
85 + .unwrap_or_else(|| completion_text("out of scripted responses"));
86 + let response = format!(
87 + "HTTP/1.1 200 OK\r\ncontent-type: application/json\r\n\
88 + content-length: {}\r\nconnection: close\r\n\r\n{}",
89 + payload.len(),
90 + payload
91 + );
92 + let _ = stream.write_all(response.as_bytes());
93 + }
94 + });
95 +
96 + FakeEndpoint { port, requests }
97 +}
98 +
99 +/// Run `sigit run` to completion against the endpoint; returns (exit code,
100 +/// parsed JSONL events). A watchdog kills the child if it wedges.
101 +fn run_headless(
102 + endpoint: &FakeEndpoint,
103 + workdir: &Path,
104 + extra_env: &[(&str, &str)],
105 +) -> (i32, Vec<Value>) {
106 + let config_dir = workdir.join("config");
107 + std::fs::create_dir_all(&config_dir).unwrap();
108 +
109 + let mut command = Command::new(env!("CARGO_BIN_EXE_sigit"));
110 + command
111 + .arg("run")
112 + .arg("--prompt")
113 + .arg("do the task")
114 + .arg("--cwd")
115 + .arg(workdir)
116 + .arg("--output")
117 + .arg("jsonl")
118 + .env_remove("SIGIT_PERMISSIONS")
119 + .env_remove("SIGIT_MODEL")
120 + .env(
121 + "OPENAI_BASE_URL",
122 + format!("http://127.0.0.1:{}", endpoint.port),
123 + )
124 + .env("OPENAI_API_KEY", "test-key")
125 + .env("SIGIT_MCP", "off")
126 + .env("SIGIT_CONFIG_DIR", &config_dir)
127 + .env("HOME", workdir)
128 + .stdin(Stdio::null())
129 + .stdout(Stdio::piped())
130 + .stderr(Stdio::piped());
131 + for (key, value) in extra_env {
132 + command.env(key, value);
133 + }
134 +
135 + let child = command.spawn().expect("spawn sigit run");
136 +
137 + // Watchdog: a wedged run must fail the test, not hang CI.
138 + let pid = child.id();
139 + let watchdog = std::thread::spawn(move || {
140 + std::thread::sleep(Duration::from_secs(120));
141 + // Best-effort; on the happy path the process is long gone.
142 + #[cfg(unix)]
143 + unsafe {
144 + libc_kill(pid as i32);
145 + }
146 + let _ = pid;
147 + });
148 +
149 + let output = child.wait_with_output().expect("wait for sigit run");
150 + drop(watchdog); // detached; happy path never joins it
151 +
152 + let stdout = String::from_utf8_lossy(&output.stdout);
153 + let stderr = String::from_utf8_lossy(&output.stderr);
154 + let events: Vec<Value> = stdout
155 + .lines()
156 + .filter(|line| !line.trim().is_empty())
157 + .map(|line| {
158 + serde_json::from_str(line).unwrap_or_else(|error| {
159 + panic!("non-JSONL stdout line {line:?}: {error}\nstderr: {stderr}")
160 + })
161 + })
162 + .collect();
163 + (output.status.code().unwrap_or(-1), events)
164 +}
165 +
166 +#[cfg(unix)]
167 +unsafe fn libc_kill(pid: i32) {
168 + unsafe extern "C" {
169 + fn kill(pid: i32, sig: i32) -> i32;
170 + }
171 + unsafe {
172 + kill(pid, 9);
173 + }
174 +}
175 +
176 +fn event_types(events: &[Value]) -> Vec<&str> {
177 + events
178 + .iter()
179 + .filter_map(|event| event["type"].as_str())
180 + .collect()
181 +}
182 +
183 +#[test]
184 +fn completes_a_tool_free_run() {
185 + let endpoint = start_fake_endpoint(vec![completion_text("All done: nothing to change.")]);
186 + let workdir = std::env::temp_dir().join(format!("sigit-headless-a-{}", std::process::id()));
187 + std::fs::create_dir_all(&workdir).unwrap();
188 +
189 + let (code, events) = run_headless(&endpoint, &workdir, &[]);
190 +
191 + assert_eq!(code, 0, "events: {events:?}");
192 + let types = event_types(&events);
193 + assert_eq!(types.first(), Some(&"run_started"), "events: {events:?}");
194 + assert!(types.contains(&"turn_text"), "events: {events:?}");
195 +
196 + let result = events.last().expect("has a result line");
197 + assert_eq!(result["type"], "result");
198 + assert_eq!(result["status"], "completed");
199 + assert_eq!(result["rounds"], 0);
200 + assert_eq!(result["summary"], "All done: nothing to change.");
201 +
202 + // The request carried the task and offered tools.
203 + let requests = endpoint.requests.lock().unwrap();
204 + let first = &requests[0];
205 + assert_eq!(first["stream"], false);
206 + assert!(
207 + first["tools"]
208 + .as_array()
209 + .is_some_and(|tools| !tools.is_empty())
210 + );
211 + let messages = first["messages"].as_array().unwrap();
212 + assert!(
213 + messages
214 + .iter()
215 + .any(|m| m["role"] == "user" && m["content"] == "do the task")
216 + );
217 +
218 + std::fs::remove_dir_all(&workdir).ok();
219 +}
220 +
221 +#[test]
222 +fn declines_mutating_tools_without_permission_override() {
223 + // Round 1: the model asks to run a mutating tool. With SIGIT_PERMISSIONS
224 + // unset the policy is `ask`, and headless mode cannot prompt — the call
225 + // must be declined (denied: true) and the refusal fed back to the model.
226 + let endpoint = start_fake_endpoint(vec![
227 + completion_tool_call("call_1", "run_command", "{\"command\":\"echo hi\"}"),
228 + completion_text("Understood, stopping."),
229 + ]);
230 + let workdir = std::env::temp_dir().join(format!("sigit-headless-b-{}", std::process::id()));
231 + std::fs::create_dir_all(&workdir).unwrap();
232 +
233 + let (code, events) = run_headless(&endpoint, &workdir, &[]);
234 +
235 + assert_eq!(code, 0, "events: {events:?}");
236 + let tool_result = events
237 + .iter()
238 + .find(|event| event["type"] == "tool_result")
239 + .expect("tool_result event");
240 + assert_eq!(tool_result["name"], "run_command");
241 + assert_eq!(tool_result["denied"], true);
242 + assert!(
243 + tool_result["output"]
244 + .as_str()
245 + .unwrap()
246 + .contains("SIGIT_PERMISSIONS=allow")
247 + );
248 +
249 + // The refusal went back as the tool result of round 1's call.
250 + let requests = endpoint.requests.lock().unwrap();
251 + let second = &requests[1];
252 + let messages = second["messages"].as_array().unwrap();
253 + assert!(messages.iter().any(|m| {
254 + m["role"] == "tool"
255 + && m["content"]
256 + .as_str()
257 + .is_some_and(|content| content.contains("was not executed"))
258 + }));
259 +
260 + let result = events.last().unwrap();
261 + assert_eq!(result["status"], "completed");
262 + assert_eq!(result["rounds"], 1);
263 +
264 + std::fs::remove_dir_all(&workdir).ok();
265 +}
266 +
267 +#[test]
268 +fn executes_allowed_tools_with_permission_override() {
269 + let endpoint = start_fake_endpoint(vec![
270 + completion_tool_call(
271 + "call_1",
272 + "run_command",
273 + "{\"command\":\"echo headless-ok\"}",
274 + ),
275 + completion_text("Command ran."),
276 + ]);
277 + let workdir = std::env::temp_dir().join(format!("sigit-headless-c-{}", std::process::id()));
278 + std::fs::create_dir_all(&workdir).unwrap();
279 +
280 + let (code, events) = run_headless(&endpoint, &workdir, &[("SIGIT_PERMISSIONS", "allow")]);
281 +
282 + assert_eq!(code, 0, "events: {events:?}");
283 + let tool_result = events
284 + .iter()
285 + .find(|event| event["type"] == "tool_result")
286 + .expect("tool_result event");
287 + assert_eq!(tool_result["denied"], false);
288 + assert!(
289 + tool_result["output"]
290 + .as_str()
291 + .unwrap()
292 + .contains("headless-ok"),
293 + "tool output should carry the command's stdout: {tool_result}"
294 + );
295 +
296 + std::fs::remove_dir_all(&workdir).ok();
297 +}