@setoelkahfi / sigit / commits / 1a99015

Address review follow-ups: TUI cancel, state pruning, ACP test

Three follow-ups from the permission-system review: The TUI's inference task used to keep running after Ctrl+C — dropping the update channel only silenced it, so it kept burning model rounds and could still execute granted tools in the background. The tool loop now notices the closed channel at each boundary (and when the approval reply channel is dropped), closes out the round in backend history via abandon_round, and stops. Permission state for dead ACP session ids accumulated forever. Session boundaries (new/load/fork) now call permissions::reset_all — the agent drives one shared engine, so only one conversation is live at a time and grants must never cross a boundary anyway. Added tests/acp_permissions.rs: an end-to-end test that spawns the real binary in ACP mode against a scripted OpenAI-compatible SSE endpoint and drives a permission round-trip over stdio — cancel first (asserting the prompt stops with "cancelled" and the next request shows the repaired history), then allow-once (asserting the command really ran and its output reached the endpoint). To make that possible, ACP mode now honors the OPENAI_BASE_URL/OPENAI_API_KEY provider override at startup like the interactive client does; previously the override was silently ignored there. Docs updated to match.

paydii committed Jul 4, 2026 at 19:49 UTC 1a99015f03d8944c95eca7645c2e1e84fb7d588c
7 files changed +468 -21
CLAUDE.md
+4 -4
@@ -63,10 +63,10 @@ Run a single test: `cargo test <test_name>`.
63
64 ## Critical platform constraint: `#[cfg(unix)]` dead code
65
66 -The interactive client, the `InferenceBackend` seam (`backend.rs`), and provider resolution
67 -(`provider.rs`) are wired up **only** through `#[cfg(unix)]` code paths. On Windows the binary
68 -runs ACP-only and drives `onde` directly, so much of `backend.rs` and `provider.rs` is
69 -legitimately unused there and the dead-code lint is suppressed *on non-Unix targets only*.
66 +The interactive client is `#[cfg(unix)]`-only. The `InferenceBackend` seam (`backend.rs`) and
67 +provider resolution (`provider.rs`) are consumed by both the interactive client and the ACP
68 +server, but several of their items are reached only through the Unix-only interactive paths, so
69 +the dead-code lint is suppressed *on non-Unix targets only*.
70
71 Consequence: code can pass clippy on macOS/Linux but fail on the Windows target (or vice versa).
72 When touching `backend.rs`, `provider.rs`, or the interactive path, keep the `cfg` gates intact —
src/backend.rs
+5 -5
@@ -11,11 +11,11 @@
11 //! The trait exposes neither `onde` nor OpenAI types, so the loop does not depend
12 //! on a specific backend.
13 //!
14 -//! The whole backend seam is wired up only through the interactive client, which
15 -//! is `#[cfg(unix)]` (see `run_interactive` in `main.rs` and `mod tui` in
16 -//! `chat.rs`). On non-Unix targets the binary runs ACP-only and drives `onde`
17 -//! directly, so every item here is legitimately unused there. Suppress the
18 -//! dead-code lint on those targets only — Unix builds still get full coverage.
14 +//! The seam is consumed by both surfaces: the interactive client (`#[cfg(unix)]`,
15 +//! see `run_interactive` in `main.rs` and `mod tui` in `chat.rs`) and the ACP
16 +//! server's prompt loop. Some items are still reached only through the
17 +//! Unix-only interactive paths, so the dead-code lint stays suppressed on
18 +//! non-Unix targets only — Unix builds keep full coverage.
19 #![cfg_attr(not(unix), allow(dead_code))]
20
21 use std::sync::Arc;
src/chat.rs
+56 -4
@@ -1606,6 +1606,27 @@ mod tui {
1606 specs
1607 }
1608
1609 + /// Close out a cancelled round in backend history: the results of tools
1610 + /// that already ran this round, plus cancellation notes for `unreached`
1611 + /// calls. Leaving a round's tool calls unanswered breaks strict
1612 + /// OpenAI-compatible endpoints on the session's next request.
1613 + async fn abandon_round(
1614 + backend: &dyn InferenceBackend,
1615 + mut tool_results: Vec<ToolResult>,
1616 + unreached: &[crate::backend::ToolCall],
1617 + ) {
1618 + for pending in unreached {
1619 + tool_results.push(ToolResult {
1620 + tool_call_id: pending.id.clone(),
1621 + content: format!(
1622 + "`{}` was not executed: the user cancelled the turn.",
1623 + pending.name
1624 + ),
1625 + });
1626 + }
1627 + backend.record_cancelled_tool_results(tool_results).await;
1628 + }
1629 +
1630 /// run the tool-calling loop off the main thread, posting updates via `tx`.
1631 /// dropping `tx` signals completion to the event loop.
1632 async fn run_inference_task(
@@ -1667,7 +1688,16 @@ mod tui {
1688
1689 let mut tool_results = Vec::new();
1690
1670 - for tc in &result.tool_calls {
1691 + for (call_index, tc) in result.tool_calls.iter().enumerate() {
1692 + // The UI drops the receiver on Ctrl+C or quit. Stop the turn
1693 + // at the next boundary instead of burning model rounds (and
1694 + // possibly running granted tools) in the background.
1695 + if tx.is_closed() {
1696 + log::info!("turn cancelled by the user — stopping the tool loop");
1697 + abandon_round(&*backend, tool_results, &result.tool_calls[call_index..]).await;
1698 + return;
1699 + }
1700 +
1701 log::info!(
1702 " → {}({})",
1703 tc.name,
@@ -1703,12 +1733,26 @@ mod tui {
1733 permissions::grant_for_session(TUI_SESSION, &tc.name);
1734 crate::tools::execute_tool(&tc.name, &tc.arguments).await
1735 }
1706 - // An explicit "no", or the UI dropped the channel
1707 - // (cancel/quit) — either way, do not run the tool.
1708 - Ok(ApprovalChoice::Deny) | Err(_) => {
1736 + Ok(ApprovalChoice::Deny) => {
1737 log::info!(" ✗ {} denied by user", tc.name);
1738 permissions::user_denial(&tc.name)
1739 }
1740 + // The UI dropped the reply channel (Ctrl+C or
1741 + // quit): the whole turn is over, not just this
1742 + // call. Close out the round and stop instead of
1743 + // continuing rounds in the background.
1744 + Err(_) => {
1745 + log::info!(
1746 + "turn cancelled at the approval prompt — stopping the tool loop"
1747 + );
1748 + abandon_round(
1749 + &*backend,
1750 + tool_results,
1751 + &result.tool_calls[call_index..],
1752 + )
1753 + .await;
1754 + return;
1755 + }
1756 }
1757 }
1758 };
@@ -1720,6 +1764,14 @@ mod tui {
1764 });
1765 }
1766
1767 + // Cancelled while the round's tools ran: record what executed and
1768 + // stop before paying for another model round nobody will see.
1769 + if tx.is_closed() {
1770 + log::info!("turn cancelled by the user — skipping the next model round");
1771 + abandon_round(&*backend, tool_results, &[]).await;
1772 + return;
1773 + }
1774 +
1775 // on the last round, pass no tools so the model must produce text —
1776 // that's also the round we can stream on-device.
1777 let next_tools = if round < MAX_TOOL_ROUNDS {
src/main.rs
+33 -3
@@ -949,9 +949,10 @@ impl SiGitAgent {
949 *guard = Some(args.cwd.clone());
950 }
951
952 - // A reloaded session starts fresh: grants and plan mode from the
953 - // previous life of this session id must not carry over.
954 - permissions::reset_session(&args.session_id.to_string());
952 + // A session boundary: grants and plan mode from the previous life of
953 + // this session id must not carry over — and since one shared engine
954 + // means one live conversation, state for every other id is dead too.
955 + permissions::reset_all();
956
957 // tool calls use relative paths, so we need to match the editor's cwd
958 if args.cwd.is_dir()
@@ -988,6 +989,9 @@ impl SiGitAgent {
989 args: ForkSessionRequest,
990 ) -> agent_client_protocol::Result<ForkSessionResponse> {
991 let new_id = SessionId::new(uuid::Uuid::new_v4().to_string());
992 + // Session boundary: permission grants and plan mode never cross it
993 + // (see handle_load_session), so a fork starts with a clean slate.
994 + permissions::reset_all();
995 log::info!(
996 "fork_session: from={} new={new_id}, cwd={}, additional_directories={:?}",
997 args.session_id,
@@ -1035,6 +1039,9 @@ impl SiGitAgent {
1039 args: NewSessionRequest,
1040 ) -> agent_client_protocol::Result<NewSessionResponse> {
1041 let session_id = SessionId::new(uuid::Uuid::new_v4().to_string());
1042 + // Session boundary: permission grants and plan mode never cross it
1043 + // (see handle_load_session), so stale ids stop accumulating state.
1044 + permissions::reset_all();
1045 log::info!(
1046 "new_session: id={session_id}, cwd={}, additional_directories={:?}",
1047 args.cwd.display(),
@@ -2846,6 +2853,29 @@ async fn run_acp_server() -> anyhow::Result<()> {
2853 needs_download,
2854 ));
2855
2856 + // Honor the explicit provider override (OPENAI_BASE_URL/OPENAI_API_KEY or
2857 + // an active providers.toml profile) in ACP mode too — the interactive
2858 + // client already does. Without this the override was silently ignored here
2859 + // and prompts insisted on a local model. It is also what lets the ACP
2860 + // integration test drive the agent against a scripted endpoint
2861 + // (tests/acp_permissions.rs). The model picker still shows the local
2862 + // selection; overrides are a power-user escape hatch, not a tier.
2863 + if let Some(cfg) = provider::active_provider() {
2864 + log::info!(
2865 + "inference: using {} (model {}) at {}",
2866 + cfg.display_name,
2867 + cfg.model,
2868 + cfg.base_url
2869 + );
2870 + let override_backend: Arc<dyn InferenceBackend> = Arc::new(OpenAiBackend::new(
2871 + cfg.base_url,
2872 + cfg.api_key,
2873 + cfg.model,
2874 + Some(system_prompt_for_model(true).to_string()),
2875 + ));
2876 + *state.backend.lock().await = override_backend;
2877 + }
2878 +
2879 let stdin = tokio::io::stdin().compat();
2880 let stdout = tokio::io::stdout().compat_write();
2881 let transport = ByteStreams::new(stdout, stdin);
src/permissions.rs
+12
@@ -184,6 +184,18 @@ pub fn reset_session(session: &str) {
184 map.remove(session);
185 }
186
187 +/// Drop the recorded state for *every* session. Called at ACP session
188 +/// boundaries (new/load/fork): the agent drives one shared engine, so only one
189 +/// conversation is live at a time and grants must never cross a boundary. This
190 +/// also keeps the map from accumulating entries for session ids that will
191 +/// never be used again.
192 +pub fn reset_all() {
193 + let mut map = sessions()
194 + .lock()
195 + .unwrap_or_else(|poisoned| poisoned.into_inner());
196 + map.clear();
197 +}
198 +
199 /// One-line status summary for `/permissions` and `/status`.
200 pub fn describe(session: &str) -> String {
201 let plan = if plan_mode(session) { "on" } else { "off" };
src/provider.rs
+5 -5
@@ -8,11 +8,11 @@
8 //! endpoint and tier are built in, and the session token is the credential.
9 //! 3. On-device: no login and no override, so inference runs locally.
10 //!
11 -//! Provider resolution is consumed only by the interactive client, which is
12 -//! `#[cfg(unix)]`. The display helpers (`cloud_tier_label`, `CLOUD_TIERS`) are
13 -//! still used cross-platform by `/models`, but the resolution path is dead on
14 -//! non-Unix targets, where the binary runs ACP-only and on-device. Suppress the
15 -//! dead-code lint there only — Unix builds keep full coverage.
11 +//! Provider resolution runs at startup in both modes: the interactive client
12 +//! picks its whole backend from it, and `run_acp_server` applies the explicit
13 +//! override (env or `providers.toml`) the same way. Some items are still wired
14 +//! up only through `#[cfg(unix)]` interactive paths, so the dead-code lint
15 +//! stays suppressed on non-Unix targets only — Unix builds keep full coverage.
16 #![cfg_attr(not(unix), allow(dead_code))]
17
18 use std::path::PathBuf;
tests/acp_permissions.rs new
+353
@@ -0,0 +1,353 @@
1 +//! End-to-end ACP permission round-trip against the real binary.
2 +//!
3 +//! Spawns `sigit` in ACP mode (stdin piped, so not a TTY) wired to a scripted
4 +//! OpenAI-compatible SSE endpoint via the `OPENAI_BASE_URL` override, then
5 +//! drives newline-delimited JSON-RPC over stdio. The scripted model calls
6 +//! `run_command` — a mutating tool — so the agent must send
7 +//! `session/request_permission` mid-turn (the exact path the spawned-handler /
8 +//! `turn_lock` design exists for). The test answers it twice:
9 +//!
10 +//! 1. `cancelled` — the prompt must stop with `stopReason: "cancelled"`, and
11 +//! the *next* request to the endpoint must show the abandoned round closed
12 +//! out with `role: "tool"` results, or a strict OpenAI-compatible endpoint
13 +//! would reject the whole session.
14 +//! 2. `selected: allow_once` — the tool must actually execute and its output
15 +//! travel back to the endpoint as a tool result.
16 +
17 +use std::collections::VecDeque;
18 +use std::io::{BufRead, BufReader, Read, Write};
19 +use std::net::TcpListener;
20 +use std::process::{Child, ChildStdin, Command, Stdio};
21 +use std::sync::mpsc::{Receiver, channel};
22 +use std::sync::{Arc, Mutex};
23 +use std::time::{Duration, Instant};
24 +
25 +use serde_json::{Value, json};
26 +
27 +const TIMEOUT: Duration = Duration::from_secs(60);
28 +
29 +// ── Scripted OpenAI-compatible endpoint ─────────────────────────────────────
30 +
31 +fn sse_body(events: &[Value]) -> String {
32 + let mut body = String::new();
33 + for event in events {
34 + body.push_str("data: ");
35 + body.push_str(&event.to_string());
36 + body.push_str("\n\n");
37 + }
38 + body.push_str("data: [DONE]\n\n");
39 + body
40 +}
41 +
42 +fn sse_tool_call(id: &str, name: &str, arguments: &str) -> String {
43 + sse_body(&[json!({
44 + "choices": [{"delta": {"tool_calls": [{
45 + "index": 0,
46 + "id": id,
47 + "function": {"name": name, "arguments": arguments},
48 + }]}}]
49 + })])
50 +}
51 +
52 +fn sse_text(text: &str) -> String {
53 + sse_body(&[json!({"choices": [{"delta": {"content": text}}]})])
54 +}
55 +
56 +/// Serves one scripted SSE response per request and records each request body.
57 +struct FakeEndpoint {
58 + port: u16,
59 + requests: Arc<Mutex<Vec<Value>>>,
60 +}
61 +
62 +fn start_fake_endpoint(responses: Vec<String>) -> FakeEndpoint {
63 + let listener = TcpListener::bind("127.0.0.1:0").expect("bind fake endpoint");
64 + let port = listener.local_addr().unwrap().port();
65 + let requests: Arc<Mutex<Vec<Value>>> = Arc::default();
66 + let recorded = Arc::clone(&requests);
67 + let queue = Mutex::new(VecDeque::from(responses));
68 +
69 + std::thread::spawn(move || {
70 + // `connection: close` below means one request per connection, so the
71 + // serial accept loop matches the agent's serial completion requests.
72 + for stream in listener.incoming() {
73 + let Ok(mut stream) = stream else { continue };
74 + let mut reader = BufReader::new(match stream.try_clone() {
75 + Ok(clone) => clone,
76 + Err(_) => continue,
77 + });
78 + let mut content_length = 0usize;
79 + loop {
80 + let mut line = String::new();
81 + if reader.read_line(&mut line).unwrap_or(0) == 0 {
82 + break;
83 + }
84 + let line = line.trim();
85 + if line.is_empty() {
86 + break;
87 + }
88 + if let Some(length) = line.to_ascii_lowercase().strip_prefix("content-length:") {
89 + content_length = length.trim().parse().unwrap_or(0);
90 + }
91 + }
92 + let mut body = vec![0u8; content_length];
93 + if reader.read_exact(&mut body).is_err() {
94 + continue;
95 + }
96 + if let Ok(request) = serde_json::from_slice::<Value>(&body) {
97 + recorded.lock().unwrap().push(request);
98 + }
99 + let payload = queue
100 + .lock()
101 + .unwrap()
102 + .pop_front()
103 + .unwrap_or_else(|| sse_text("out of scripted responses"));
104 + let response = format!(
105 + "HTTP/1.1 200 OK\r\ncontent-type: text/event-stream\r\n\
106 + content-length: {}\r\nconnection: close\r\n\r\n{}",
107 + payload.len(),
108 + payload
109 + );
110 + let _ = stream.write_all(response.as_bytes());
111 + }
112 + });
113 +
114 + FakeEndpoint { port, requests }
115 +}
116 +
117 +// ── ACP client over the binary's stdio ──────────────────────────────────────
118 +
119 +struct AgentUnderTest {
120 + child: Child,
121 + stdin: ChildStdin,
122 + incoming: Receiver<Value>,
123 + next_id: u64,
124 +}
125 +
126 +fn spawn_agent(port: u16, config_dir: &std::path::Path) -> AgentUnderTest {
127 + let mut child = Command::new(env!("CARGO_BIN_EXE_sigit"))
128 + .env("OPENAI_BASE_URL", format!("http://127.0.0.1:{port}"))
129 + .env("OPENAI_API_KEY", "test-key")
130 + .env("SIGIT_MODEL", "scripted-model")
131 + .env("SIGIT_CONFIG_DIR", config_dir)
132 + .env("SIGIT_MCP", "off")
133 + // A fresh config dir means the default permission mode, `ask` — make
134 + // sure the environment can't turn the gate off underneath the test.
135 + .env_remove("SIGIT_PERMISSIONS")
136 + .env_remove("SIGIT_LOCAL_INFERENCE")
137 + .stdin(Stdio::piped())
138 + .stdout(Stdio::piped())
139 + .stderr(Stdio::null())
140 + .spawn()
141 + .expect("spawn sigit in ACP mode");
142 +
143 + let stdout = child.stdout.take().unwrap();
144 + let (message_tx, incoming) = channel();
145 + std::thread::spawn(move || {
146 + for line in BufReader::new(stdout).lines() {
147 + let Ok(line) = line else { break };
148 + if let Ok(message) = serde_json::from_str::<Value>(&line)
149 + && message_tx.send(message).is_err()
150 + {
151 + break;
152 + }
153 + }
154 + });
155 +
156 + let stdin = child.stdin.take().unwrap();
157 + AgentUnderTest {
158 + child,
159 + stdin,
160 + incoming,
161 + next_id: 0,
162 + }
163 +}
164 +
165 +impl AgentUnderTest {
166 + fn send(&mut self, message: Value) {
167 + let mut line = message.to_string();
168 + line.push('\n');
169 + self.stdin
170 + .write_all(line.as_bytes())
171 + .expect("write to agent stdin");
172 + self.stdin.flush().expect("flush agent stdin");
173 + }
174 +
175 + fn request(&mut self, method: &str, params: Value) -> u64 {
176 + self.next_id += 1;
177 + let id = self.next_id;
178 + self.send(json!({"jsonrpc": "2.0", "id": id, "method": method, "params": params}));
179 + id
180 + }
181 +
182 + fn respond(&mut self, id: Value, result: Value) {
183 + self.send(json!({"jsonrpc": "2.0", "id": id, "result": result}));
184 + }
185 +
186 + /// Skip notifications and unrelated traffic until `matches` is satisfied.
187 + fn wait_for(&mut self, what: &str, matches: impl Fn(&Value) -> bool) -> Value {
188 + let deadline = Instant::now() + TIMEOUT;
189 + loop {
190 + let remaining = deadline.saturating_duration_since(Instant::now());
191 + match self.incoming.recv_timeout(remaining) {
192 + Ok(message) if matches(&message) => return message,
193 + Ok(_) => continue,
194 + Err(_) => panic!("timed out waiting for {what}"),
195 + }
196 + }
197 + }
198 +
199 + /// The response to one of *our* requests (has our id, no `method`).
200 + fn wait_for_response(&mut self, id: u64) -> Value {
201 + let response = self.wait_for(&format!("response to request {id}"), |message| {
202 + message["id"] == id && message.get("method").is_none()
203 + });
204 + assert!(
205 + response.get("error").is_none(),
206 + "request {id} failed: {response}"
207 + );
208 + response
209 + }
210 +
211 + /// A request *from* the agent (has a `method` and its own id).
212 + fn wait_for_agent_request(&mut self, method: &str) -> Value {
213 + self.wait_for(&format!("agent request {method}"), |message| {
214 + message["method"] == method && message.get("id").is_some()
215 + })
216 + }
217 +}
218 +
219 +impl Drop for AgentUnderTest {
220 + fn drop(&mut self) {
221 + let _ = self.child.kill();
222 + let _ = self.child.wait();
223 + }
224 +}
225 +
226 +// ── The round-trip ──────────────────────────────────────────────────────────
227 +
228 +#[test]
229 +fn permission_round_trip_cancel_then_allow() {
230 + let endpoint = start_fake_endpoint(vec![
231 + sse_tool_call("call_1", "run_command", r#"{"command":"echo sigit-first"}"#),
232 + sse_tool_call(
233 + "call_2",
234 + "run_command",
235 + r#"{"command":"echo sigit-approved"}"#,
236 + ),
237 + sse_text("done"),
238 + ]);
239 +
240 + let scratch = std::env::temp_dir().join(format!("sigit_acp_perm_{}", std::process::id()));
241 + let config_dir = scratch.join("config");
242 + let cwd = scratch.join("cwd");
243 + std::fs::create_dir_all(&config_dir).unwrap();
244 + std::fs::create_dir_all(&cwd).unwrap();
245 +
246 + let mut agent = spawn_agent(endpoint.port, &config_dir);
247 +
248 + let id = agent.request(
249 + "initialize",
250 + json!({"protocolVersion": 1, "clientCapabilities": {}}),
251 + );
252 + agent.wait_for_response(id);
253 +
254 + let id = agent.request("session/new", json!({"cwd": cwd, "mcpServers": []}));
255 + let session_id = agent.wait_for_response(id)["result"]["sessionId"]
256 + .as_str()
257 + .expect("session id")
258 + .to_string();
259 +
260 + // ── Turn 1: cancel at the permission gate ───────────────────────────
261 + let prompt_id = agent.request(
262 + "session/prompt",
263 + json!({
264 + "sessionId": session_id,
265 + "prompt": [{"type": "text", "text": "run the first command"}],
266 + }),
267 + );
268 +
269 + let permission = agent.wait_for_agent_request("session/request_permission");
270 + let params = &permission["params"];
271 + assert_eq!(params["sessionId"], session_id.as_str());
272 + let title = params["toolCall"]["title"].as_str().expect("title");
273 + assert!(
274 + title.contains("run_command") && title.contains("echo sigit-first"),
275 + "the dialog must show the tool and its arguments, got: {title}"
276 + );
277 + assert_eq!(
278 + params["toolCall"]["rawInput"]["command"], "echo sigit-first",
279 + "full arguments must travel as rawInput"
280 + );
281 + let option_ids: Vec<&str> = params["options"]
282 + .as_array()
283 + .expect("options")
284 + .iter()
285 + .map(|option| option["optionId"].as_str().unwrap_or_default())
286 + .collect();
287 + assert_eq!(option_ids, ["allow_once", "allow_session", "reject_once"]);
288 +
289 + agent.respond(
290 + permission["id"].clone(),
291 + json!({"outcome": {"outcome": "cancelled"}}),
292 + );
293 +
294 + let response = agent.wait_for_response(prompt_id);
295 + assert_eq!(response["result"]["stopReason"], "cancelled");
296 +
297 + // ── Turn 2: history must be repaired; then approve once ─────────────
298 + let prompt_id = agent.request(
299 + "session/prompt",
300 + json!({
301 + "sessionId": session_id,
302 + "prompt": [{"type": "text", "text": "run the second command"}],
303 + }),
304 + );
305 +
306 + let permission = agent.wait_for_agent_request("session/request_permission");
307 + agent.respond(
308 + permission["id"].clone(),
309 + json!({"outcome": {"outcome": "selected", "optionId": "allow_once"}}),
310 + );
311 +
312 + let response = agent.wait_for_response(prompt_id);
313 + assert_eq!(response["result"]["stopReason"], "end_turn");
314 +
315 + // ── What the endpoint saw ────────────────────────────────────────────
316 + let requests = endpoint.requests.lock().unwrap();
317 + assert_eq!(requests.len(), 3, "expected exactly three completions");
318 +
319 + // Request 2 replays the full history: the cancelled round's tool call
320 + // must be answered by a `role: "tool"` message, not left dangling.
321 + let messages = requests[1]["messages"].as_array().expect("messages");
322 + let call_position = messages
323 + .iter()
324 + .position(|message| message["tool_calls"][0]["id"] == "call_1")
325 + .expect("cancelled turn's assistant tool call in replayed history");
326 + let repair = &messages[call_position + 1];
327 + assert_eq!(repair["role"], "tool", "dangling tool call not closed out");
328 + assert_eq!(repair["tool_call_id"], "call_1");
329 + assert!(
330 + repair["content"]
331 + .as_str()
332 + .unwrap_or_default()
333 + .contains("cancelled"),
334 + "repair message should say the turn was cancelled: {repair}"
335 + );
336 +
337 + // Request 3 carries the approved call's real output.
338 + let messages = requests[2]["messages"].as_array().expect("messages");
339 + let result = messages
340 + .iter()
341 + .find(|message| message["role"] == "tool" && message["tool_call_id"] == "call_2")
342 + .expect("tool result for the approved call");
343 + assert!(
344 + result["content"]
345 + .as_str()
346 + .unwrap_or_default()
347 + .contains("sigit-approved"),
348 + "the approved command's output should reach the endpoint: {result}"
349 + );
350 +
351 + drop(agent);
352 + let _ = std::fs::remove_dir_all(&scratch);
353 +}