feature/crates-release
md 290 lines 12.3 KB
Rendered Raw
1 # Skill: Tool Calling in siGit Code
2
3 ## Overview
4
5 siGit Code supports **agentic tool calling** — the LLM invokes tools (read/write files, run commands, read websites) to operate on the user's codebase. This works in both **interactive TUI mode** and **ACP server mode** (Zed editor).
6
7 Tool calling spans three layers:
8
9 ```
10 siGit (agent loop + tool execution)
11 → onde (ChatEngine with tool-aware API)
12 → mistral.rs (model inference + tool call parsing)
13 ```
14
15 ---
16
17 ## Model Requirement
18
19 **Only Qwen 3 supports tool calling.** Qwen 2.5 does NOT — mistral.rs only has a parser for Qwen 3's `<tool_call>...</tool_call>` XML format.
20
21 | Model | Constructor | Size | Tool calling | Default |
22 |-------|-----------|------|:---:|:---:|
23 | Qwen 3 8B (Q4_K_M) | `GgufModelConfig::qwen3_8b()` | ~5 GB | ✅ | ✅ **default** |
24 | Qwen 3 4B (Q4_K_M) | `GgufModelConfig::qwen3_4b()` | ~2.7 GB | ✅ | |
25 | Qwen 3 1.7B (Q4_K_M) | `GgufModelConfig::qwen3_1_7b()` | ~1.3 GB | ✅ | |
26 | Qwen 2.5 Coder 3B | `GgufModelConfig::qwen25_coder_3b()` | ~1.93 GB | ❌ | |
27 | Qwen 2.5 Coder 1.5B | `GgufModelConfig::qwen25_coder_1_5b()` | ~941 MB | ❌ | |
28
29 siGit uses **Qwen 3 8B** by default with `max_tokens: 8192` (set in `main.rs` for both TUI and ACP modes).
30
31 ### Why 8B over 4B
32
33 4B can't do `edit_file` reliably. It reads a file, then fails to reproduce the exact `old_text` it just saw. This spirals into 7+ retry rounds that burn through `max_tokens` on `<think>` blocks and return nothing. 8B is the smallest model that actually lands edits.
34
35 ### bartowski GGUF naming convention
36
37 bartowski's repos use the publisher name as a prefix with an underscore:
38
39 | Constant | Value |
40 |----------|-------|
41 | `BARTOWSKI_QWEN3_8B_GGUF` | `"bartowski/Qwen_Qwen3-8B-GGUF"` |
42 | `QWEN3_8B_GGUF_FILE` | `"Qwen_Qwen3-8B-Q4_K_M.gguf"` |
43 | `BARTOWSKI_QWEN3_4B_GGUF` | `"bartowski/Qwen_Qwen3-4B-GGUF"` |
44 | `QWEN3_4B_GGUF_FILE` | `"Qwen_Qwen3-4B-Q4_K_M.gguf"` |
45
46 These constants live in `onde/src/inference/models.rs`.
47
48 ---
49
50 ## Tools (9 total)
51
52 Defined in `sigit/src/tools.rs` via `all_tools()`:
53
54 | # | Tool | Parameters | Behavior |
55 |---|------|-----------|----------|
56 | 1 | `read_file` | `path` | Reads file contents, truncates at 10,000 chars |
57 | 2 | `create_directory` | `path` | Creates directory and all parents |
58 | 3 | `list_directory` | `path` | Lists entries with `[DIR]`/`[FILE]` prefix, dirs first |
59 | 4 | `search_files` | `pattern`, `path` (optional) | Recursive regex search, max 50 matches |
60 | 5 | `read_website` | `url` | Fetches HTTP/HTTPS, strips HTML, returns text |
61 | 6 | `create_file` | `path`, `content` | Creates new file (fails if exists) |
62 | 7 | `edit_file` | `path`, `old_text`, `new_text` | Find-and-replace (must match exactly once) |
63 | 8 | `delete_file` | `path` | Deletes file or empty directory |
64 | 9 | `run_command` | `command`, `cwd` (optional) | Shell command with 120s timeout |
65
66 ### Async handling
67
68 `execute_tool()` is `async`. Most tools run synchronously, except:
69
70 - **`read_website`** — uses `tokio::task::spawn_blocking` because `reqwest::blocking::Client` panics inside a tokio runtime ("Cannot start a runtime from within a runtime")
71
72 ### Tool gating by model
73
74 In TUI mode, `run_inference_task()` takes a `tools_enabled: bool` parameter. When the model's `ModelOption.tool_calling` is `false` (Qwen 2.5), an empty tool list is passed so the model doesn't receive tool schemas it can't use.
75
76 ---
77
78 ## Architecture
79
80 ### Layer 1: mistral.rs (model-level)
81
82 - `RequestBuilder::set_tools(Vec<Tool>)` — attach tool definitions
83 - `RequestBuilder::set_tool_choice(ToolChoice::Auto)` — let model decide
84 - `QwenParser` detects `<tool_call>...</tool_call>` tags in output
85 - Grammar-constrained decoding forces valid JSON inside tool calls
86 - `<think>...</think>` reasoning is separated from tool calls by the reasoning parser
87 - Works identically for GGUF and full-precision models
88
89 ### Layer 2: onde (engine-level)
90
91 #### Key types (`onde/src/inference/types.rs`)
92
93 | Type | Purpose |
94 |------|---------|
95 | `ToolDefinition` | `{ name, description, parameters_schema: String }` |
96 | `ToolCallRequest` | `{ id, function_name, arguments: String }` |
97 | `ToolResult` | `{ tool_call_id, content: String }` |
98 | `ToolAwareResult` | `{ text, tool_calls: Vec<ToolCallRequest>, duration_secs, ... }` |
99
100 #### Key methods (`onde/src/inference/engine.rs`)
101
102 | Method | Purpose |
103 |--------|---------|
104 | `send_message_with_tools(msg, &[ToolDefinition])` | Returns `ToolAwareResult` with possible tool calls |
105 | `send_tool_results(Vec<ToolResult>, Option<&[ToolDefinition]>)` | Feed results back; `None` forces text response |
106
107 #### Internal details
108
109 - `attach_tools()` converts `ToolDefinition` → mistral.rs `Tool`, sets `ToolChoice::Auto` and `strict: Some(true)`
110 - `parse_tool_calls()` extracts tool calls from `choice.message.tool_calls`, generates fallback IDs if empty
111 - `replay_history_with_tools()` uses `.enumerate()` for correct sequential `index` values
112 - Malformed `parameters_schema` JSON logs a warning instead of silently producing empty params
113 - Malformed tool call `arguments` JSON logs a warning for debugging
114
115 ### Layer 3: siGit (agent-level)
116
117 #### ACP session handling (`src/main.rs`)
118
119 All session handlers (`load_session`, `fork_session`, `new_session`) do:
120
121 1. **Store `args.cwd`** in `session_cwd: Mutex<Option<PathBuf>>`
122 2. **`std::env::set_current_dir(&args.cwd)`** — so relative paths in tool calls resolve correctly
123 3. **`engine.clear_history()`** — siGit doesn't persist sessions
124 4. **`engine.push_history(ChatMessage::system(...))`** — injects: *"The user's project working directory is {cwd}. Always use absolute paths..."*
125
126 Without step 4, the model uses the process `cwd` (often `$HOME`) and creates files in the wrong directory.
127
128 #### ACP content block handling (`prompt()`)
129
130 The `prompt()` handler processes all ACP content block types:
131
132 - **`ContentBlock::Text`** — passed through as-is
133 - **`ContentBlock::Resource` (EmbeddedResource)**`TextResourceContents` inlined as `--- {uri} ---\n{text}\n--- end ---`
134 - **`ContentBlock::ResourceLink`**`file://` URIs are read from disk. **Line range fragments** like `#L207:219` are parsed: the `#` fragment is stripped from the path, and only lines 207–219 are extracted and sent to the model
135
136 Example: Zed sends `@ index.html (207:219)` as:
137 ```
138 ResourceLink(name="index.html (207:219)", uri="file:///path/to/index.html#L207:219")
139 ```
140 siGit parses this into path `/path/to/index.html` + lines 207–219.
141
142 ---
143
144 ## The Agentic Loop
145
146 Both ACP mode (`SiGitAgent::prompt()`) and TUI mode (`run_inference_task()`) implement:
147
148 ```
149 1. engine.send_message_with_tools(user_text, &tools) → ToolAwareResult
150 2. while result.tool_calls is non-empty AND round < MAX_TOOL_ROUNDS (10):
151 a. For each tool_call:
152 - Log: → tool_name(arguments)
153 - Execute: tools::execute_tool(name, arguments).await
154 - Log: ← N chars
155 - Collect ToolResult { tool_call_id, content }
156 b. Decide next_tools:
157 - round < MAX_TOOL_ROUNDS → Some(&tools) (allow more calls)
158 - else → None (force text response)
159 c. engine.send_tool_results(results, next_tools) → ToolAwareResult
160 3. Send final result.text to user
161 - Empty reply after tool rounds → log warning (ACP) or show error (TUI)
162 ```
163
164 ---
165
166 ## System Prompt
167
168 The `SYSTEM_PROMPT` in `main.rs` (~122 lines) includes critical instructions:
169
170 - **Never tell the user to run commands** — use `run_command` tool instead
171 - **Can access websites** — use `read_website` tool (overrides RLHF refusal training)
172 - **Prefer absolute paths** in all tool arguments
173 - **Git operations** — always use `run_command` with absolute cwd
174 - **smbCloud domain knowledge** — auth boundaries, deploy flows, project structure
175
176 The session `cwd` is injected as a separate system message at session creation time (not part of the static prompt).
177
178 ---
179
180 ## Model Cache
181
182 Models are stored in the shared Onde App Group container on macOS:
183
184 ```
185 ~/Library/Group Containers/group.com.ondeinference.apps/models/hub/
186 ```
187
188 `setup.rs` sets `HF_HOME` and `HF_HUB_CACHE` to point there at startup, so siGit reuses models downloaded by the Onde desktop app (and vice versa).
189
190 ---
191
192 ## Adding a New Tool
193
194 1. Add an `AgentTool` entry to `all_tools()` in `src/tools.rs`
195 2. Add a match arm to `execute_tool()` — use `spawn_blocking` if the implementation blocks
196 3. Write `exec_your_tool(arguments: &str) -> String`
197 4. Update `test_all_tools_count` test (currently expects 9)
198
199 No changes needed in onde or mistral.rs — tool definitions are passed dynamically.
200
201 ---
202
203 ## Adding a New Model
204
205 1. **`onde/src/inference/models.rs`** — add `pub const` for repo ID and GGUF filename, add to `SUPPORTED_MODELS` array and `SUPPORTED_MODEL_INFO`
206 2. **`onde/src/inference/engine.rs`** — add `pub fn model_name() -> Self` constructor to `impl GgufModelConfig`
207 3. **`sigit/src/chat.rs`** — add `ModelOption` entry to `SIGIT_MODELS` with `tool_calling: true/false`
208 4. **`sigit/src/main.rs`** — update `run_interactive()` and `run_acp_server()` if changing the default
209
210 ---
211
212 ## Debugging
213
214 ### Log locations
215
216 - **TUI mode:** `$TMPDIR/sigit.log` (e.g. `/var/folders/.../sigit.log`)
217 - **ACP mode (Zed):** `~/Library/Logs/Zed/Zed.log` — grep for `agent stderr:.*sigit`
218
219 ### Key log patterns
220
221 ```
222 # Model loaded successfully
223 ChatEngine: model Qwen 3 8B loaded in 6.9s
224
225 # Session cwd captured
226 load_session: id=..., cwd=/path/to/project, additional_directories=[...]
227
228 # Tool call parsed by mistral.rs
229 ChatEngine: tool inference END — 12.3s — tool_calls: 1
230
231 # Tool executed
232 → read_file({"path":"/absolute/path/to/file.rs"})
233 ← 6506 chars
234
235 # Tool result sent back
236 ChatEngine: tool results inference START — 1 results
237
238 # Model returned empty (exhausted max_tokens on thinking)
239 model returned empty reply after 7 tool round(s)
240
241 # ResourceLink received from Zed
242 block[1]: ResourceLink(name=index.html (207:219), uri=file:///path/to/index.html#L207:219)
243
244 # ResourceLink read failed (fragment not stripped — old bug, now fixed)
245 could not read ResourceLink file:///path/to/index.html#L207:219: No such file or directory
246 ```
247
248 ### Common issues
249
250 | Symptom | Cause | Fix |
251 |---------|-------|-----|
252 | Model says "I cannot access websites" | RLHF refusal override not in system prompt | System prompt now has CRITICAL block about `read_website` |
253 | `0 tool call(s)` for every prompt | Wrong model loaded (Qwen 2.5) | Check log for `loading GGUF model` — must be Qwen 3 |
254 | `edit_file` returns `← 161 chars` repeatedly | `old_text not found` — model can't match exact text | Use Qwen 3 8B (not 4B); consider line-based edit tool |
255 | Files created in wrong directory | `cwd` not captured from ACP session | Session handlers must call `set_current_dir` + `push_history` with cwd |
256 | `@ file.html (207:219)` context missing | `#L207:219` fragment not stripped from file path | `prompt()` now parses URI fragments and extracts line ranges |
257 | `read_website` panics/hangs | `reqwest::blocking` inside tokio runtime | `exec_read_website` wrapped in `spawn_blocking` |
258 | Empty reply after many tool rounds | Model exhausted `max_tokens` on `<think>` blocks | Set `max_tokens: 8192`; 8B model wastes fewer tokens on thinking |
259
260 ---
261
262 ## Cargo Dependency Note
263
264 For local development, `sigit/Cargo.toml` must use the path dependency:
265
266 ```toml
267 onde = { path = "../onde" }
268 ```
269
270 For CI/release, switch to the git dependency (after pushing Onde changes):
271
272 ```toml
273 onde = { git = "https://github.com/ondeinference/onde", branch = "development" }
274 ```
275
276 The `qwen3_8b()` constructor only exists in the local Onde SDK until it's pushed to the `development` branch.
277
278 ---
279
280 ## File Map
281
282 | File | What it does |
283 |------|-------------|
284 | `sigit/src/tools.rs` | 9 tool schemas (`all_tools()`), `execute_tool()` dispatch, all `exec_*` implementations |
285 | `sigit/src/main.rs` | `SYSTEM_PROMPT`, `SiGitAgent` struct with `session_cwd`, ACP session handlers (cwd + push_history), `prompt()` with content block parsing, model selection (`qwen3_8b`), `MAX_TOOL_ROUNDS` |
286 | `sigit/src/chat.rs` | `SIGIT_MODELS` array (4 models), `run_inference_task()` with `tools_enabled` gate, TUI tool loop |
287 | `sigit/src/setup.rs` | HF cache setup pointing to shared App Group container |
288 | `onde/src/inference/types.rs` | `ToolDefinition`, `ToolCallRequest`, `ToolResult`, `ToolAwareResult` |
289 | `onde/src/inference/engine.rs` | `send_message_with_tools()`, `send_tool_results()`, `attach_tools()`, `parse_tool_calls()`, `replay_history_with_tools()`, `GgufModelConfig::qwen3_8b()` |
290 | `onde/src/inference/models.rs` | Model constants and `SUPPORTED_MODELS` array |