Shell#
The assistant can execute shell commands with bash by outputting code blocks with shell as the language.
- Configuration:
GPTME_SHELL_TIMEOUT: Environment variable to configure command timeout (set before starting gptme)
Set to a number (e.g., 30) for timeout in seconds
Set to 0 to disable timeout
Invalid values default to 1200 seconds (20 minutes)
If not set, defaults to 1200 seconds (20 minutes)
- GPTME_SHELL_FOREGROUND_TIMEOUT: Soft timeout before a foreground command is
promoted to a conversation-owned background job. Defaults to 120 seconds. Set to 0 to disable promotion. GPTME_SHELL_TIMEOUT remains the hard limit. A command that starts with
timeout DURATIONis not promoted before that duration (plus a short grace) has passed. Promoted jobs are managed with theoutput N,wait N [timeout],kill Nandjobscontrol commands, each given as the entire shell command. POSIX only — Windows has no process-group promotion path, so this setting is ignored there and long commands wait for GPTME_SHELL_TIMEOUT.- GPTME_SHELL_MEMORY_LIMIT: Optional per-shell address-space ceiling (POSIX only,
off by default). Accepts a plain byte count or a binary suffix (e.g. “512M”, “1G”). Applies to the persistent shell and any command it runs via ulimit -v, so a runaway build fails with an allocation error instead of stalling the session.
GPTME_SHELL_TRUNC_PRE_TOKENS / GPTME_SHELL_TRUNC_POST_TOKENS: Override the head/tail token budget for stdout truncation. Defaults: 2000 / 8000. GPTME_SHELL_TRUNC_STDERR_PRE_TOKENS / GPTME_SHELL_TRUNC_STDERR_POST_TOKENS: Same overrides for stderr. Defaults: 2000 / 2000. Lowering these makes the truncation path fire on smaller outputs, which surfaces savings telemetry in context-savings.jsonl and the /context command. Invalid values fall back to defaults.
- GPTME_SHELL_MAX_OUTPUT_BYTES: Hard cap on the total bytes (stdout + stderr
combined) captured into the in-process buffer before the subprocess is killed and the output is truncated. Accepts a plain byte count or a binary suffix (e.g. “32M”, “1G”). Default: 32 MiB. This prevents a runaway
catof a multi-GiB file from exhausting gptme’s RSS. The process receives SIGTERM then SIGKILL; the returned output contains a[output truncated at N MiB, process killed]marker. Token-level truncation (GPTME_SHELL_TRUNC_*) is applied on top as a second stage.
Instructions
The given command will be executed in a stateful bash shell.
The shell tool will respond with the output of the execution.
These programs are available, among others:
- apt-get
- docker
- git
- hyperfine
### When to use the shell
Use the shell to inspect the workspace, search files, check git state, or run
existing commands and tests. Prefer the repo over answering from memory.
### Background Jobs
Use `background: true` for work you already know will run long — dev servers,
builds, test suites. A foreground command that outruns the soft timeout is
promoted to background job N; its exit code and output arrive automatically.
A leading `timeout DURATION` keeps a command in the foreground that long.
Manage jobs with `output N`, `wait N [timeout]`, `kill N`, `jobs`: not bash,
so send each as the entire command, never piped or inside a script.
Examples
| User |
list the current directory |
| Assistant |
To list the files in the current directory, use `ls`: |
| System |
Ran command: `ls` |
The assistant can learn context by exploring the filesystem
| User |
learn about the project |
| Assistant |
Lets start by checking the files |
| System |
|
| Assistant |
Now lets check the README |
| System |
|
| Assistant |
Now we check main.py |
| System |
|
| Assistant |
The project is... |
Create vue project
| User |
Create a new vue project with typescript and pinia named fancy-project |
| Assistant |
Sure! Let's create a new vue project with TypeScript and Pinia named fancy-project: |
| System |
|
Proper quoting for complex content
| User |
add a comment with backticks and special characters |
| Assistant |
When passing complex content with special characters, use single quotes to prevent shell interpretation: |
Long-running commands become background jobs
| User |
run the training script |
| Assistant |
Let me start the training run: |
| System |
Promoted to background shell job #1 after 120s: `python train.py` > > The command is still running. Its exit code and output will be reported automatically when it finishes. > > Manage it with these gptme job-control commands. Each must be the *entire* shell command (they are not bash commands, so they cannot be piped, chained, or used inside a script): > - `output 1` - show recent output (`output 1 --new` for output since the last check) > - `wait 1 [timeout]` - block until it finishes, e.g. `wait 1 10m`, then show its output > - `kill 1` - terminate it > - `jobs` - list background jobs |
| User |
how is it going? |
| Assistant |
Let me check the output from the background job: |
| System |
**Job #1** - Running (185.2s) Command: `python train.py` > > ```stdout > epoch 3/10 loss=0.412 > ``` |
| User |
wait for it to finish |
| Assistant |
I'll block on the job for up to 20 minutes: |
| System |
**Job #1** - Finished (exit code: 0) (912.4s) Command: `python train.py` > > ```stdout > epoch 10/10 loss=0.107 > saved model.pt > ``` |
- class gptme.tools.shell.PromotedJobProcess
Process-like state for a command still running in a detached shell.
- __init__(returncode: int | None = None, _done: ~threading.Event = <factory>, _finish_lock: ~_thread.allocate_lock = <factory>) None
- gptme.tools.shell.close_conversation_shell(conversation_id: str) None
Close and remove the shell session for a conversation.
Called by the SESSION_END hook to clean up shell file descriptors when a conversation’s last session is removed.
- gptme.tools.shell.execute_shell(code: str | None, args: list[str] | None, kwargs: dict[str, str] | None) Generator[Message, None, None]
Execute a shell command, optionally as a harness-owned background job.
- gptme.tools.shell.execute_shell_impl(cmd: str, logdir: Path | None, timeout: float | None = None) Generator[Message, None, None]
Execute shell command and format output.
- gptme.tools.shell.get_shell() ShellSession
Get the shell session for the current context, creating it if necessary.
Uses ContextVar to provide context-local state, allowing each conversation to have its own shell session with independent working directory.
In server contexts (where current_conversation_id is set), also registers the shell in a conversation-level registry for cleanup via SESSION_END hooks.
- gptme.tools.shell.get_shell_command(code: str | None, args: list[str] | None, kwargs: dict[str, str] | None) str
Get the shell command from code/args/kwargs.
- gptme.tools.shell.get_workspace_cwd() str | None
Get the workspace directory for the current context, if set.
- gptme.tools.shell.set_shell(shell: ShellSession) None
Set the shell session for the current context (for testing).
- gptme.tools.shell.set_workspace_cwd(cwd: str) None
Set the workspace directory for the current context (thread-safe).
Call this before any shell creation to ensure the shell subprocess starts in the correct directory, even with concurrent sessions. This is the thread-safe replacement for os.chdir() in server contexts.
- gptme.tools.shell.split_commands(script: str) list[str]
Split at top-level newlines, preserving Bash lists and original source.
Tree-sitter spans include heredoc bodies and use byte offsets, so slicing UTF-8 source preserves quoted delimiters and non-ASCII text without rewrites. When tree-sitter cannot produce a clean tree,
bash -nis the authority: actual syntax errors raise ValueError; valid-but-unparseable scripts are returned as a single command.
- gptme.tools.shell.trim_blank_lines(text: str) str
Trim only leading and trailing blank (whitespace-only) lines.
Interior blank lines are kept. Unlike str.strip(), the first/last contentful lines are returned verbatim, so indentation (e.g. from sed/head/tail of indented code) is preserved.
Command Confirmation#
Not every shell command requires user confirmation. gptme uses a three-tier model:
Allowlisted commands are eligible for auto-confirmation (no prompt). The command name is not enough on its own: arguments, flags, and shell syntax still have to pass the safety checks below. Names considered eligible:
ls stat cd cat pwd echo head find rg ag tail grep
wc sort uniq cut file which type tree du df
Even with one of those names, a prompt is still required when the command has:
File redirections (
ls > files.txt,echo x >> out)Sensitive paths (
cat /etc/shadow,ls /root)Command substitution (
echo $(whoami), backticks)Unpermitted flags (
find . -delete,rg --pre,sort -o)
Transparent wrappers are stripped before the allowlist check, so wrapping an allowlisted command with a timing or resource-limit prefix does not itself require a prompt. Recognised wrappers:
time [−p]
timeout [options] <duration>
nohup
nice [−n <adj>]
stdbuf [−i/−o/−e <mode>]
env [−i / −u VAR] (bare env, without NAME=value)
command [−p]
builtin
For example, timeout 60 grep -r foo /var/log is judged as grep and
runs without a prompt if grep also passes the checks above. Variables
(PATH=. ls) are deliberately not treated as transparent — they
change what the command does and are always gated.
Denied commands are blocked outright and never executed. These match specific command-text patterns, not every equivalent spelling:
git add ./git add -A/git commit -a— bulk stagingDestructive git:
git reset --hard,git clean -f,git push -frm -rf /(also/./,/*, quoted/),sudo rm -rf /, andrm -rf *. Flag-order variants such asrm -fr /orrm -r -f /are not hard-denied; they reach the confirmation prompt.chmod 777pkill/killallPiping to shell interpreters:
... | bash,... | python3
Everything else falls through to the normal confirmation prompt.
Command Parsing#
Since v0.34.0, gptme uses tree-sitter-bash as the primary parser
to split multi-command scripts into individual commands for sequential
execution and to detect background operators (&). bashlex remains a
compatibility fallback when tree-sitter reports a parse error or a generated
split fails syntax validation — which is why some scripts can still receive
the legacy splitting behavior.
The former bashlex-only parser could not handle several common constructs
(time, timeout, process substitution <(…), arithmetic expansion
$(( ))) and reported them as syntax errors to the model.
The allowlist and denylist checks always operate on the raw command text and are unaffected by the parser.