@setoelkahfi / sigit / commits / 240f46c

Add headless one-shot mode: sigit run

A new subcommand executes a single task non-interactively and exits, built for cloud runners (siGit Code Cloud Agent) and scripting: sigit run [--prompt <text> | --prompt-file <path>] [--cwd <dir>] [--max-rounds <n>] [--output jsonl|text] - Progress streams as JSONL events on stdout (run_started, turn_text, tool_call, tool_result, compaction, result); logs stay on stderr. - Exit codes: 0 completed, 1 run failed, 2 usage/config error. - Provider resolves like other modes (OPENAI_BASE_URL/OPENAI_API_KEY override first, signed-in cloud next) and never falls back to on-device inference, so a headless host cannot trigger a multi-GB model download. - The tool loop mirrors the ACP prompt handler: permission gate, auto-compaction between rounds, forced text reply on the last round. A tool that would prompt (Decision::Ask) is declined with a pointer to SIGIT_PERMISSIONS=allow instead of hanging. - Integration tests drive the real binary against a scripted OpenAI-compatible endpoint, covering the tool-free path, the deny-without-override path, and the allow-override path. Co-Authored-By: Claude <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A1xu3s7ey2EqDNjgtt2XFA Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

siGit Code Session committed Jul 6, 2026 at 17:46 UTC 240f46c9e96c716da6ba41dbf185218375b11056
5 files changed +821 -1
CHANGELOG.md
+6
index 19724a1..c090e7c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,11 @@ # Changelog +## Unreleased + +### What changed + +- New headless one-shot mode: `sigit run [--prompt <text> | --prompt-file <path>] [--cwd <dir>] [--max-rounds <n>] [--output jsonl|text]` runs a single task non-interactively and exits. Progress is emitted as JSONL events on stdout (`run_started`, `turn_text`, `tool_call`, `tool_result`, `compaction`, and a final `result`), logs stay on stderr, and exit codes are 0 (completed), 1 (run failed), 2 (usage/configuration error). The provider resolves like other modes (`OPENAI_BASE_URL`/`OPENAI_API_KEY` override first), never falling back to on-device inference, and a tool that would normally prompt for permission is declined with a pointer to `SIGIT_PERMISSIONS=allow` instead of hanging. This is the execution engine for siGit Code Cloud Agent runners + ## 1.3.2 Adds a tool permission system with plan mode, durable sessions with context
CLAUDE.md
+6 -1
index f0a0de8..e9155ef 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -15,7 +15,12 @@ is a TTY: it relies on fd redirection to keep logs out of the TUI, so Windows gets ACP mode only. Before the TTY/ACP split, `main` also dispatches the account subcommands `sigit login`, -`sigit logout`, `sigit whoami` (see `src/main.rs` `main()`). +`sigit logout`, `sigit whoami`, and the headless one-shot mode `sigit run` (see `src/main.rs` +`main()`). `sigit run` (`src/headless.rs`) executes a single task non-interactively — JSONL +progress events on stdout, logs on stderr, exit 0/1/2 — and is what cloud runners (siGit Code +Cloud Agent) drive; it never falls back to on-device inference and declines permission prompts +unless `SIGIT_PERMISSIONS=allow` is set. Its tool loop mirrors `handle_prompt`; keep the two in +sync. ## Working in this repo
src/headless.rs
+497
new file mode 100644 index 0000000..ab7cad8 --- /dev/null +++ b/src/headless.rs @@ -0,0 +1,497 @@ +//! Headless one-shot mode: `sigit run`. +//! +//! Runs a single task non-interactively and exits: the prompt arrives as a +//! CLI flag, progress is emitted as JSONL events on stdout (one object per +//! line), and logs stay on stderr. Built for cloud runners (siGit Code Cloud +//! Agent) and scripting, where nobody is present to answer a permission +//! prompt — on `Decision::Ask` the tool is declined with a pointer to +//! `SIGIT_PERMISSIONS=allow` instead of blocking. +//! +//! The tool loop mirrors the ACP prompt handler (`handle_prompt` in +//! `main.rs`): permission gate → execute → feed results back, with +//! auto-compaction between rounds and a forced text reply on the final +//! round. Keep the two in sync when changing loop semantics. + +use std::io::Write as _; +use std::path::PathBuf; +use std::sync::Arc; + +use serde_json::json; + +use crate::backend::{self, InferenceBackend, OpenAiBackend, ToolResult, ToolSpec}; +use crate::{permissions, provider, tools}; + +/// Headless runs default to a higher round cap than interactive prompts: an +/// autonomous task routinely needs long edit/build/test chains and there is +/// no user present to re-prompt a stopped run. +const DEFAULT_MAX_ROUNDS: usize = 40; + +/// Cap on `arguments`/`output` strings embedded in JSONL events. Full outputs +/// still reach the model; events only need enough for a live transcript. +const EVENT_FIELD_MAX_CHARS: usize = 4_000; + +const USAGE: &str = "usage: sigit run [--prompt <text> | --prompt-file <path>] \ + [--cwd <dir>] [--max-rounds <n>] [--output jsonl|text]"; + +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub enum OutputMode { + Jsonl, + Text, +} + +#[derive(Debug)] +pub struct HeadlessOptions { + pub prompt: String, + pub cwd: PathBuf, + pub max_rounds: usize, + pub output: OutputMode, +} + +/// The value following a flag, or a usage error naming the flag. +fn next_value(args: &mut impl Iterator<Item = String>, flag: &str) -> Result<String, String> { + args.next() + .ok_or_else(|| format!("{flag} needs a value\n{USAGE}")) +} + +/// Parse `sigit run` arguments (everything after the subcommand). +pub fn parse_args(mut args: impl Iterator<Item = String>) -> Result<HeadlessOptions, String> { + let mut prompt: Option<String> = None; + let mut cwd: Option<PathBuf> = None; + let mut max_rounds = DEFAULT_MAX_ROUNDS; + let mut output = OutputMode::Jsonl; + + while let Some(flag) = args.next() { + match flag.as_str() { + "--prompt" => { + let value = next_value(&mut args, "--prompt")?; + if prompt.is_some() { + return Err(format!("give --prompt or --prompt-file once\n{USAGE}")); + } + prompt = Some(value); + } + "--prompt-file" => { + let path = next_value(&mut args, "--prompt-file")?; + if prompt.is_some() { + return Err(format!("give --prompt or --prompt-file once\n{USAGE}")); + } + let text = std::fs::read_to_string(&path) + .map_err(|error| format!("cannot read --prompt-file {path}: {error}"))?; + prompt = Some(text); + } + "--cwd" => { + cwd = Some(PathBuf::from(next_value(&mut args, "--cwd")?)); + } + "--max-rounds" => { + max_rounds = next_value(&mut args, "--max-rounds")? + .parse::<usize>() + .ok() + .filter(|n| *n > 0) + .ok_or_else(|| format!("--max-rounds needs a positive integer\n{USAGE}"))?; + } + "--output" => { + output = match next_value(&mut args, "--output")?.as_str() { + "jsonl" => OutputMode::Jsonl, + "text" => OutputMode::Text, + other => return Err(format!("unknown --output {other}\n{USAGE}")), + }; + } + other => return Err(format!("unknown argument {other}\n{USAGE}")), + } + } + + let prompt = prompt + .map(|text| text.trim().to_string()) + .filter(|text| !text.is_empty()) + .ok_or_else(|| format!("a non-empty --prompt or --prompt-file is required\n{USAGE}"))?; + + let cwd = cwd.unwrap_or_else(|| PathBuf::from(".")); + let cwd = cwd + .canonicalize() + .map_err(|error| format!("--cwd {}: {error}", cwd.display()))?; + + Ok(HeadlessOptions { + prompt, + cwd, + max_rounds, + output, + }) +} + +/// Entry point for `sigit run`. Never returns on failure paths — exits the +/// process with 0 (run completed), 1 (run failed), or 2 (usage/config error). +pub async fn run(args: impl Iterator<Item = String>) -> anyhow::Result<()> { + let options = match parse_args(args) { + Ok(options) => options, + Err(message) => { + eprintln!("sigit run: {message}"); + std::process::exit(2); + } + }; + + // Provider: the explicit override (env / providers.toml) first — this is + // how a cloud runner injects a per-run endpoint and token — then the + // signed-in cloud as a convenience. Never fall back to on-device: a + // headless host should not silently download a multi-GB model. + let Some(config) = + provider::active_provider().or_else(|| provider::cloud_tier_provider("large")) + else { + eprintln!( + "sigit run: no inference provider configured. Set OPENAI_BASE_URL and \ + OPENAI_API_KEY (and optionally SIGIT_MODEL), configure providers.toml, \ + or sign in with `sigit login`." + ); + std::process::exit(2); + }; + + if let Err(error) = std::env::set_current_dir(&options.cwd) { + eprintln!("sigit run: cannot enter {}: {error}", options.cwd.display()); + std::process::exit(2); + } + + let system_prompt = format!( + "{}\n\n{}", + crate::system_prompt_for_model(true), + crate::session_context_message(&options.cwd) + ); + let backend: Arc<dyn InferenceBackend> = Arc::new(OpenAiBackend::new( + config.base_url.clone(), + config.api_key.clone(), + config.model.clone(), + Some(system_prompt), + )); + crate::register_subagent_factory_for(&config); + let tools = crate::agent_tools_as_specs(); + + let emitter = Emitter { + mode: options.output, + }; + emitter.event(json!({ + "type": "run_started", + "cwd": options.cwd.display().to_string(), + "model": config.model, + "max_rounds": options.max_rounds, + })); + + let rounds = match drive_loop(&backend, &tools, &options, &emitter).await { + Ok((summary, rounds)) => { + emitter.event(json!({ + "type": "result", + "status": "completed", + "summary": summary, + "rounds": rounds, + })); + rounds + } + Err((error, rounds)) => { + emitter.event(json!({ + "type": "result", + "status": "failed", + "error": error, + "rounds": rounds, + })); + std::process::exit(1); + } + }; + log::info!("headless run completed after {rounds} tool round(s)"); + Ok(()) +} + +/// The tool loop. Returns `(summary, rounds)` or `(error, rounds)`. +/// +/// Keep in sync with `handle_prompt` in `main.rs`: same permission gate, same +/// auto-compaction trigger, same force-text final round. +async fn drive_loop( + backend: &Arc<dyn InferenceBackend>, + tools: &[ToolSpec], + options: &HeadlessOptions, + emitter: &Emitter, +) -> Result<(String, usize), (String, usize)> { + // Permission decisions are per-session state; a headless process is one + // session. There are no grants to accumulate (nobody can answer "always + // allow"), the id only namespaces the lookup. + let session = format!("headless-{}", std::process::id()); + + let mut result = backend + .send_message_with_tools(&options.prompt, tools, None) + .await + .map_err(|error| (format!("inference failed: {error}"), 0))?; + emitter.turn_text(&result.text); + + let mut round = 0usize; + + while !result.tool_calls.is_empty() && round < options.max_rounds { + round += 1; + + // Auto-compaction: long tool runs grow history fast; fold it into a + // summary before the next round rather than blowing the window. + let estimate = backend::estimate_tokens(&backend.history_snapshot().await); + if estimate > backend::DEFAULT_CONTEXT_TOKEN_BUDGET { + match backend.compact_history(backend::COMPACT_KEEP_LAST).await { + Ok(()) => { + let after = backend::estimate_tokens(&backend.history_snapshot().await); + emitter.event(json!({ + "type": "compaction", + "approx_tokens_before": estimate, + "approx_tokens_after": after, + })); + } + Err(error) => log::warn!("headless compaction failed: {error}"), + } + } + + let mut tool_results = Vec::new(); + for call in &result.tool_calls { + emitter.tool_call(call); + let (output, denied) = match permissions::decision_for(&session, &call.name) { + permissions::Decision::Allow => ( + tools::execute_tool(&call.name, &call.arguments).await, + false, + ), + permissions::Decision::Deny(reason) => { + log::info!("headless: {} denied by policy", call.name); + (reason, true) + } + permissions::Decision::Ask => { + log::info!("headless: {} needs approval, declining", call.name); + ( + format!( + "`{}` was not executed: headless mode cannot prompt for \ + permission. Run with SIGIT_PERMISSIONS=allow to auto-approve \ + mutating tools, or grant this tool in settings.toml.", + call.name + ), + true, + ) + } + }; + emitter.tool_result(call, &output, denied); + tool_results.push(ToolResult { + tool_call_id: call.id.clone(), + content: output, + }); + } + + let next_tools = if round < options.max_rounds { + Some(tools) + } else { + None // last round: force a text reply + }; + result = backend + .send_tool_results(tool_results, next_tools, None) + .await + .map_err(|error| (format!("inference failed: {error}"), round))?; + emitter.turn_text(&result.text); + } + + let (_think, visible) = crate::chat::strip_think_blocks(&result.text); + let summary = if visible.trim().is_empty() { + "The run finished without a final summary.".to_string() + } else { + visible.trim().to_string() + }; + Ok((summary, round)) +} + +// ── Event output ───────────────────────────────────────────────────────────── + +struct Emitter { + mode: OutputMode, +} + +impl Emitter { + /// Write one event. JSONL mode prints the object as-is; text mode renders + /// a human-oriented line per event kind. + fn event(&self, event: serde_json::Value) { + match self.mode { + OutputMode::Jsonl => { + let mut stdout = std::io::stdout().lock(); + let _ = writeln!(stdout, "{event}"); + let _ = stdout.flush(); + } + OutputMode::Text => { + let line = match event["type"].as_str() { + Some("run_started") => format!( + "▶ run started in {} (model {})", + event["cwd"].as_str().unwrap_or("?"), + event["model"].as_str().unwrap_or("?"), + ), + Some("turn_text") => event["text"].as_str().unwrap_or_default().to_string(), + Some("tool_call") => format!( + "→ {}({})", + event["name"].as_str().unwrap_or("?"), + event["arguments"].as_str().unwrap_or_default(), + ), + Some("tool_result") => format!( + "← {} ({} chars{})", + event["name"].as_str().unwrap_or("?"), + event["output_chars"].as_u64().unwrap_or(0), + if event["denied"].as_bool().unwrap_or(false) { + ", denied" + } else { + "" + }, + ), + Some("compaction") => "… compacted conversation history".to_string(), + Some("result") => match event["status"].as_str() { + Some("completed") => format!( + "✔ completed\n{}", + event["summary"].as_str().unwrap_or_default() + ), + _ => format!("✘ failed: {}", event["error"].as_str().unwrap_or("?")), + }, + _ => event.to_string(), + }; + if !line.is_empty() { + let mut stdout = std::io::stdout().lock(); + let _ = writeln!(stdout, "{line}"); + let _ = stdout.flush(); + } + } + } + } + + /// Emit the visible part of an assistant turn, skipping empty turns. + fn turn_text(&self, raw: &str) { + let (_think, visible) = crate::chat::strip_think_blocks(raw); + let visible = visible.trim(); + if visible.is_empty() { + return; + } + let (text, truncated) = clip(visible, EVENT_FIELD_MAX_CHARS); + let mut event = json!({ "type": "turn_text", "text": text }); + if truncated { + event["truncated"] = json!(true); + } + self.event(event); + } + + fn tool_call(&self, call: &backend::ToolCall) { + let (arguments, truncated) = clip(&call.arguments, EVENT_FIELD_MAX_CHARS); + let mut event = json!({ + "type": "tool_call", + "id": call.id, + "name": call.name, + "arguments": arguments, + }); + if truncated { + event["truncated"] = json!(true); + } + self.event(event); + } + + fn tool_result(&self, call: &backend::ToolCall, output: &str, denied: bool) { + let (clipped, truncated) = clip(output, EVENT_FIELD_MAX_CHARS); + let mut event = json!({ + "type": "tool_result", + "id": call.id, + "name": call.name, + "output_chars": output.chars().count(), + "output": clipped, + "denied": denied, + }); + if truncated { + event["truncated"] = json!(true); + } + self.event(event); + } +} + +/// Truncate to `max` characters (not bytes — always on a char boundary). +fn clip(text: &str, max: usize) -> (String, bool) { + if text.chars().count() <= max { + (text.to_string(), false) + } else { + (text.chars().take(max).collect(), true) + } +} + +#[cfg(test)] +mod tests { + use super::*; + + fn args(list: &[&str]) -> impl Iterator<Item = String> { + list.iter() + .map(|s| s.to_string()) + .collect::<Vec<_>>() + .into_iter() + } + + #[test] + fn parses_prompt_and_defaults() { + let options = parse_args(args(&["--prompt", "fix the tests"])).expect("parses"); + assert_eq!(options.prompt, "fix the tests"); + assert_eq!(options.max_rounds, DEFAULT_MAX_ROUNDS); + assert_eq!(options.output, OutputMode::Jsonl); + assert!(options.cwd.is_absolute()); + } + + #[test] + fn requires_a_prompt() { + let error = parse_args(args(&[])).expect_err("missing prompt"); + assert!(error.contains("--prompt")); + } + + #[test] + fn rejects_empty_prompt() { + let error = parse_args(args(&["--prompt", " "])).expect_err("blank prompt"); + assert!(error.contains("non-empty")); + } + + #[test] + fn rejects_prompt_and_prompt_file_together() { + let file = std::env::temp_dir().join(format!("sigit-prompt-{}.txt", std::process::id())); + std::fs::write(&file, "task").unwrap(); + let error = parse_args(args(&[ + "--prompt", + "one", + "--prompt-file", + file.to_str().unwrap(), + ])) + .expect_err("both prompt flags"); + assert!(error.contains("once")); + std::fs::remove_file(&file).ok(); + } + + #[test] + fn reads_prompt_file() { + let file = std::env::temp_dir().join(format!("sigit-promptf-{}.txt", std::process::id())); + std::fs::write(&file, "task from file\n").unwrap(); + let options = parse_args(args(&["--prompt-file", file.to_str().unwrap()])).expect("parses"); + assert_eq!(options.prompt, "task from file"); + std::fs::remove_file(&file).ok(); + } + + #[test] + fn rejects_bad_flags_and_values() { + assert!(parse_args(args(&["--prompt", "x", "--max-rounds", "0"])).is_err()); + assert!(parse_args(args(&["--prompt", "x", "--max-rounds", "abc"])).is_err()); + assert!(parse_args(args(&["--prompt", "x", "--output", "yaml"])).is_err()); + assert!(parse_args(args(&["--bogus"])).is_err()); + assert!(parse_args(args(&["--prompt", "x", "--cwd", "/definitely/not/a/dir"])).is_err()); + } + + #[test] + fn parses_overrides() { + let options = parse_args(args(&[ + "--prompt", + "x", + "--max-rounds", + "7", + "--output", + "text", + ])) + .expect("parses"); + assert_eq!(options.max_rounds, 7); + assert_eq!(options.output, OutputMode::Text); + } + + #[test] + fn clip_is_char_boundary_safe() { + let (out, truncated) = clip("héllo wörld", 5); + assert_eq!(out, "héllo"); + assert!(truncated); + let (out, truncated) = clip("short", 10); + assert_eq!(out, "short"); + assert!(!truncated); + } +}
src/main.rs
+15
index 1631a63..026d98f 100644 --- a/src/main.rs +++ b/src/main.rs @@ -32,6 +32,7 @@ mod account; mod backend; mod chat; mod credentials; +mod headless; mod instructions; mod mcp; mod models; @@ -3130,6 +3131,20 @@ async fn main() -> anyhow::Result<()> { println!("{}", account::status_line().await); return Ok(()); } + "run" => { + // Headless one-shot mode: stdout carries JSONL run events, so + // logs go to stderr exactly like ACP mode. + init_logging(false); + setup::setup_shared_model_cache(); + // Best-effort MCP discovery so `mcp__*` tools are offered + // (`SIGIT_MCP=off` skips this, the cloud-runner default). + mcp::init().await; + log::info!( + "siGit v{} starting (headless run)", + env!("CARGO_PKG_VERSION") + ); + return headless::run(std::env::args().skip(2)).await; + } _ => {} } }
tests/headless_run.rs
+297
new file mode 100644 index 0000000..6a450a4 --- /dev/null +++ b/tests/headless_run.rs @@ -0,0 +1,297 @@ +//! End-to-end `sigit run` (headless mode) against the real binary. +//! +//! Spawns `sigit run` wired to a scripted OpenAI-compatible endpoint via the +//! `OPENAI_BASE_URL` override and asserts on the JSONL event stream. Headless +//! runs use non-streaming completions (no token sink), so the endpoint serves +//! plain JSON chat-completion bodies, not SSE. + +use std::io::{BufRead, BufReader, Read, Write}; +use std::net::TcpListener; +use std::path::Path; +use std::process::{Command, Stdio}; +use std::sync::{Arc, Mutex}; +use std::time::Duration; + +use serde_json::{Value, json}; + +/// One scripted JSON completion body. +fn completion_text(text: &str) -> String { + json!({ + "choices": [{"message": {"role": "assistant", "content": text}}] + }) + .to_string() +} + +fn completion_tool_call(id: &str, name: &str, arguments: &str) -> String { + json!({ + "choices": [{"message": { + "role": "assistant", + "content": null, + "tool_calls": [{ + "id": id, + "type": "function", + "function": {"name": name, "arguments": arguments}, + }], + }}] + }) + .to_string() +} + +/// Serves one scripted JSON response per request and records request bodies. +struct FakeEndpoint { + port: u16, + requests: Arc<Mutex<Vec<Value>>>, +} + +fn start_fake_endpoint(responses: Vec<String>) -> FakeEndpoint { + let listener = TcpListener::bind("127.0.0.1:0").expect("bind fake endpoint"); + let port = listener.local_addr().unwrap().port(); + let requests: Arc<Mutex<Vec<Value>>> = Arc::default(); + let recorded = Arc::clone(&requests); + let queue = Mutex::new(std::collections::VecDeque::from(responses)); + + std::thread::spawn(move || { + for stream in listener.incoming() { + let Ok(mut stream) = stream else { continue }; + let mut reader = BufReader::new(match stream.try_clone() { + Ok(clone) => clone, + Err(_) => continue, + }); + let mut content_length = 0usize; + loop { + let mut line = String::new(); + if reader.read_line(&mut line).unwrap_or(0) == 0 { + break; + } + let line = line.trim(); + if line.is_empty() { + break; + } + if let Some(length) = line.to_ascii_lowercase().strip_prefix("content-length:") { + content_length = length.trim().parse().unwrap_or(0); + } + } + let mut body = vec![0u8; content_length]; + if reader.read_exact(&mut body).is_err() { + continue; + } + if let Ok(request) = serde_json::from_slice::<Value>(&body) { + recorded.lock().unwrap().push(request); + } + let payload = queue + .lock() + .unwrap() + .pop_front() + .unwrap_or_else(|| completion_text("out of scripted responses")); + let response = format!( + "HTTP/1.1 200 OK\r\ncontent-type: application/json\r\n\ + content-length: {}\r\nconnection: close\r\n\r\n{}", + payload.len(), + payload + ); + let _ = stream.write_all(response.as_bytes()); + } + }); + + FakeEndpoint { port, requests } +} + +/// Run `sigit run` to completion against the endpoint; returns (exit code, +/// parsed JSONL events). A watchdog kills the child if it wedges. +fn run_headless( + endpoint: &FakeEndpoint, + workdir: &Path, + extra_env: &[(&str, &str)], +) -> (i32, Vec<Value>) { + let config_dir = workdir.join("config"); + std::fs::create_dir_all(&config_dir).unwrap(); + + let mut command = Command::new(env!("CARGO_BIN_EXE_sigit")); + command + .arg("run") + .arg("--prompt") + .arg("do the task") + .arg("--cwd") + .arg(workdir) + .arg("--output") + .arg("jsonl") + .env_remove("SIGIT_PERMISSIONS") + .env_remove("SIGIT_MODEL") + .env( + "OPENAI_BASE_URL", + format!("http://127.0.0.1:{}", endpoint.port), + ) + .env("OPENAI_API_KEY", "test-key") + .env("SIGIT_MCP", "off") + .env("SIGIT_CONFIG_DIR", &config_dir) + .env("HOME", workdir) + .stdin(Stdio::null()) + .stdout(Stdio::piped()) + .stderr(Stdio::piped()); + for (key, value) in extra_env { + command.env(key, value); + } + + let child = command.spawn().expect("spawn sigit run"); + + // Watchdog: a wedged run must fail the test, not hang CI. + let pid = child.id(); + let watchdog = std::thread::spawn(move || { + std::thread::sleep(Duration::from_secs(120)); + // Best-effort; on the happy path the process is long gone. + #[cfg(unix)] + unsafe { + libc_kill(pid as i32); + } + let _ = pid; + }); + + let output = child.wait_with_output().expect("wait for sigit run"); + drop(watchdog); // detached; happy path never joins it + + let stdout = String::from_utf8_lossy(&output.stdout); + let stderr = String::from_utf8_lossy(&output.stderr); + let events: Vec<Value> = stdout + .lines() + .filter(|line| !line.trim().is_empty()) + .map(|line| { + serde_json::from_str(line).unwrap_or_else(|error| { + panic!("non-JSONL stdout line {line:?}: {error}\nstderr: {stderr}") + }) + }) + .collect(); + (output.status.code().unwrap_or(-1), events) +} + +#[cfg(unix)] +unsafe fn libc_kill(pid: i32) { + unsafe extern "C" { + fn kill(pid: i32, sig: i32) -> i32; + } + unsafe { + kill(pid, 9); + } +} + +fn event_types(events: &[Value]) -> Vec<&str> { + events + .iter() + .filter_map(|event| event["type"].as_str()) + .collect() +} + +#[test] +fn completes_a_tool_free_run() { + let endpoint = start_fake_endpoint(vec![completion_text("All done: nothing to change.")]); + let workdir = std::env::temp_dir().join(format!("sigit-headless-a-{}", std::process::id())); + std::fs::create_dir_all(&workdir).unwrap(); + + let (code, events) = run_headless(&endpoint, &workdir, &[]); + + assert_eq!(code, 0, "events: {events:?}"); + let types = event_types(&events); + assert_eq!(types.first(), Some(&"run_started"), "events: {events:?}"); + assert!(types.contains(&"turn_text"), "events: {events:?}"); + + let result = events.last().expect("has a result line"); + assert_eq!(result["type"], "result"); + assert_eq!(result["status"], "completed"); + assert_eq!(result["rounds"], 0); + assert_eq!(result["summary"], "All done: nothing to change."); + + // The request carried the task and offered tools. + let requests = endpoint.requests.lock().unwrap(); + let first = &requests[0]; + assert_eq!(first["stream"], false); + assert!( + first["tools"] + .as_array() + .is_some_and(|tools| !tools.is_empty()) + ); + let messages = first["messages"].as_array().unwrap(); + assert!( + messages + .iter() + .any(|m| m["role"] == "user" && m["content"] == "do the task") + ); + + std::fs::remove_dir_all(&workdir).ok(); +} + +#[test] +fn declines_mutating_tools_without_permission_override() { + // Round 1: the model asks to run a mutating tool. With SIGIT_PERMISSIONS + // unset the policy is `ask`, and headless mode cannot prompt — the call + // must be declined (denied: true) and the refusal fed back to the model. + let endpoint = start_fake_endpoint(vec![ + completion_tool_call("call_1", "run_command", "{\"command\":\"echo hi\"}"), + completion_text("Understood, stopping."), + ]); + let workdir = std::env::temp_dir().join(format!("sigit-headless-b-{}", std::process::id())); + std::fs::create_dir_all(&workdir).unwrap(); + + let (code, events) = run_headless(&endpoint, &workdir, &[]); + + assert_eq!(code, 0, "events: {events:?}"); + let tool_result = events + .iter() + .find(|event| event["type"] == "tool_result") + .expect("tool_result event"); + assert_eq!(tool_result["name"], "run_command"); + assert_eq!(tool_result["denied"], true); + assert!( + tool_result["output"] + .as_str() + .unwrap() + .contains("SIGIT_PERMISSIONS=allow") + ); + + // The refusal went back as the tool result of round 1's call. + let requests = endpoint.requests.lock().unwrap(); + let second = &requests[1]; + let messages = second["messages"].as_array().unwrap(); + assert!(messages.iter().any(|m| { + m["role"] == "tool" + && m["content"] + .as_str() + .is_some_and(|content| content.contains("was not executed")) + })); + + let result = events.last().unwrap(); + assert_eq!(result["status"], "completed"); + assert_eq!(result["rounds"], 1); + + std::fs::remove_dir_all(&workdir).ok(); +} + +#[test] +fn executes_allowed_tools_with_permission_override() { + let endpoint = start_fake_endpoint(vec![ + completion_tool_call( + "call_1", + "run_command", + "{\"command\":\"echo headless-ok\"}", + ), + completion_text("Command ran."), + ]); + let workdir = std::env::temp_dir().join(format!("sigit-headless-c-{}", std::process::id())); + std::fs::create_dir_all(&workdir).unwrap(); + + let (code, events) = run_headless(&endpoint, &workdir, &[("SIGIT_PERMISSIONS", "allow")]); + + assert_eq!(code, 0, "events: {events:?}"); + let tool_result = events + .iter() + .find(|event| event["type"] == "tool_result") + .expect("tool_result event"); + assert_eq!(tool_result["denied"], false); + assert!( + tool_result["output"] + .as_str() + .unwrap() + .contains("headless-ok"), + "tool output should carry the command's stdout: {tool_result}" + ); + + std::fs::remove_dir_all(&workdir).ok(); +}