> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mezmo.com/llms.txt
> Use this file to discover all available pages before exploring further.

# AURA CLI Reference

> The aura command-line client. Install, run, configure, and use every feature of the interactive terminal client.

A fast, interactive terminal client for chat completions with tool execution. Built primarily as the command line interface for [AURA by Mezmo](https://mezmo.com/aura), but works with **any OpenAI-compatible API**. Plug in your own models, agents, or LLM endpoints.

***

## Quick Start

### Build from the mono-repo

```bash theme={null}
# Default build (standalone + HTTP  builds agents in-process from TOML config)
cargo build -p aura-cli --release

# HTTP-only mode (lightweight, no agent/MCP dependencies)
cargo build -p aura-cli --release --no-default-features
```

The binary will be at `target/release/aura`.

### Run it

#### HTTP mode (connect to a running `aura webserver`)

```bash theme={null}
# Using environment variables
export AURA_API_URL="https://api.example.com"
export AURA_API_KEY="your-api-key"
aura

# Or pass flags directly
aura --api-url "https://api.example.com" \
         --api-key "your-api-key"
```

#### Standalone mode (no server needed)

Standalone mode is enabled by default. The CLI loads agent configs directly and runs without an HTTP server. When `--api-url` is not set, standalone mode activates automatically:

```bash theme={null}
# Uses config.toml in the current directory by default
aura

# Single TOML config file
aura --config path/to/agent.toml

# Directory of TOML configs (enables /model switching between agents)
aura --config configs/

# One-shot query in standalone mode
aura --config agent.toml --query "hello"

# Select a specific agent from a config directory
aura --config configs/ --model "Math Agent"
```

In standalone mode, the CLI builds agents in-process. MCP tools from the TOML config are available. CLI local tools (Shell, Read, Update, ...) become available when **both** sides opt in. Pass `--enable-client-tools` and set `[agent].enable_client_tools = true` in the loaded TOML config (single-agent configs only; orchestrated configs drop client tools). See [Client-Side Tools](#client-side-tools) for details. The `/model` command works identically. It lists all loaded configs and lets you switch between them.

***

## Generating a starter config (`aura init`)

New to AURA? `aura init` walks you through creating a ready-to-run `config.toml`. No hand-editing TOML required.

```bash theme={null}
aura init                 # interactive; writes ./config.toml
aura init -o my.toml      # choose the output path
```

It will:

* **Sense** your environment for a provider API key (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, …) and suggest that provider.
* **Pick a provider** from a validated list (openai, anthropic, bedrock, gemini, ollama, openrouter).
* **Handle the API key**: if the provider's conventional env var is already set it asks whether to use it; otherwise it prompts for the key (input is masked). The generated config references the key by its native env var (`api_key = "{{ env.OPENAI_API_KEY }}"`). Secrets are never written into `config.toml`.
* **Verify & list models**: it queries the provider's live model list and shows a short, curated shortlist (newest per family: pick by number, accept the default, or type any id). For Ollama it lists whatever you have installed.
* **Write** `config.toml`, plus a `.env` **only when you enter a new key** that isn't already in your environment. The `.env` holds a secret. Add it to your `.gitignore`.

`aura` loads `.env` automatically at startup, so a config generated here runs as-is.

### Flags

| Flag | Description |
| - | - |
| `-o, --output <PATH>` | Output config path (default `config.toml`). |
| `--provider <P>` | Provider: openai, anthropic, bedrock, gemini, ollama, openrouter. |
| `--model <ID>` | Model id (skips the model picker). |
| `--api-key-env <VAR>` | Env var to read the key from (default per provider). |
| `--region <R>` | AWS region (bedrock only). |
| `--base-url <URL>` | Base URL (ollama only; default `http://localhost:11434`). |
| `--name <NAME>` | Agent name written to the config (default `assistant`). |
| `--offline` | Skip live model-list verification. |
| `--non-interactive` | Fail on missing values instead of prompting (automatic when stdin isn't a TTY). |
| `--force` | Overwrite an existing config without asking. |

For scripted/CI use, supply the required values as flags:

```bash theme={null}
aura init --provider openai --model gpt-4o --non-interactive
```

***

## Backends

AURA CLI supports two backends, selected by the presence of `--api-url`:

| Backend | When | Dependencies | Tools |
| - | - | - | - |
| **Direct** (standalone, default) | No `--api-url` flag | Full AURA stack (agents, MCP, providers) | MCP tools from TOML config; CLI local tools when `--enable-client-tools` is set |
| **HTTP** | `--api-url <URL>` | Lightweight. Just HTTP client | Server-side MCP tools always; CLI local tools when both sides opt in (see below) |

Both backends produce identical SSE event streams and share the same stream parser, so all CLI features (stream panel, tool display, orchestration events, etc.) work identically regardless of backend.

### Feature flag: `standalone-cli`

The `standalone-cli` Cargo feature is **enabled by default**. It includes the full agent framework, MCP integration, and provider support. To build a lightweight HTTP-only client with no agent or MCP dependencies, disable default features:

```bash theme={null}
# Default build: standalone + HTTP
cargo build -p aura-cli

# HTTP-only: small binary, fast compile, no agent dependencies
cargo build -p aura-cli --no-default-features
```

***

## Features

* **Interactive REPL** with conversation history, streaming responses, and markdown rendering
* **One-shot mode** for scripting and pipelines (`--query`). Stdout is the raw assistant response, no markers or markdown rendering
* **Local tool execution**: the model can read files, search code, list directories, run shell commands, and edit files on your behalf
* **Standalone mode** (default): run agents directly from TOML config without a web server
* **Conversation persistence**: pick up where you left off with `--resume` or `/resume`
* **Tab completion**: cycle through matching slash commands, models, and conversations with `Tab` / `Shift+Tab`; your typed prefix is preserved as you narrow the matches
* **Model selection**: browse and select models from the server (or loaded configs in standalone mode)
* **Permission system**: control which local tools are allowed, denied, or prompted before execution
* **SSE streaming**: real-time token-by-token output with a toggleable event panel
* **Auto-compaction**: automatic context management when conversations grow large
* **Mid-stream commands**: execute slash commands while the model is still streaming
* **Works with any OpenAI-compatible API**: not locked to a single provider

***

## Environment Variables

| Variable | Description | Default |
| - | - | - |
| `AURA_API_URL` | Base API URL (paths like `/v1/chat/completions` are appended automatically) | `http://localhost:8080` |
| `AURA_API_KEY` | Bearer token for authentication | *(none)* |
| `AURA_MODEL` | Model name to use (in standalone mode, selects agent by name/alias) | *(none: omitted from request, server picks default)* |
| `AURA_EXTRA_HEADERS` | Additional HTTP headers as comma-separated `key:value` pairs (e.g. `x-chat-session-id:foo,authorization:Bearer …`); overrides the auto-injected `x-chat-session-id` | *(none)* |
| `AURA_ENABLE_FINAL_RESPONSE_SUMMARY` | Generate a one-line LLM-based title for each final response (adds an extra round-trip per turn). Set to `true` or `1` to enable. | `false` |
| `AURA_ENABLE_CLIENT_TOOLS` | Advertise CLI local tools to the model and execute them locally (see [Client-Side Tools](#client-side-tools)) | `false` |
| `AURA_LOG_FILE` | Path to a file for diagnostic tracing logs. Unset → no logging. Set → events are appended to the file in both REPL and one-shot mode (see [Logging](#logging)) | *(none: no logging)* |

***

## Command Line Arguments

```
aura [OPTIONS]
```

| Flag | Env Equivalent | Description |
| - | - | - |
| `--api-url <URL>` | `AURA_API_URL` | Base API URL |
| `--api-key <KEY>` | `AURA_API_KEY` | Bearer token for authentication |
| `--model <MODEL>` | `AURA_MODEL` | Model name (HTTP: starting model; standalone: selects agent by name/alias) |
| `--system-prompt <PROMPT>` | - | System prompt (HTTP: see note below; standalone: append/replace agent prompt) |
| `--query <QUERY>` | - | Run a single query and exit (one-shot mode) |
| `--resume <ID>` | - | Resume a previous conversation by ID or prefix |
| `--force` | - | Bypass warnings and non-critical errors (useful in one-shot/query mode) |
| `--enable-client-tools[=<bool>]` | `AURA_ENABLE_CLIENT_TOOLS` | Advertise CLI local tools to the model (default: disabled: see [Client-Side Tools](#client-side-tools)) |
| `--enable-final-response-summary[=<bool>]` | `AURA_ENABLE_FINAL_RESPONSE_SUMMARY` | Generate a one-line LLM title for each final response (adds an extra round-trip per turn; default: disabled) |
| `--standalone` | - | Force standalone mode (default when `--api-url` is absent; mutually exclusive with `--api-url` flag, but overrides `AURA_API_URL` env var) |
| `--config <PATH>` | - | Path to TOML agent config file or directory (standalone mode; defaults to `config.toml`) |
| `--log-file <PATH>` | `AURA_LOG_FILE` | Append diagnostic tracing logs to this file. Omit for no logging (see [Logging](#logging)) |

**Precedence:** CLI flags > environment variables > project `cli.toml` > global `cli.toml` > defaults.

***

## Configuration File

The CLI looks for a TOML preferences file in two places, with the project file
overriding the global one on a per-field basis:

| File | Purpose |
| - | - |
| `~/.aura/cli.toml` | Global defaults (across all projects) |
| `<project>/.aura/cli.toml` | Project-local override, found by walking up from `$PWD` |

The project lookup walks up from the current working directory until it finds a
`.aura/` directory (the same convention used by `.git`, `.editorconfig`, etc.).
Run the CLI from anywhere inside your project tree and the closest `.aura/cli.toml`
wins. `$HOME` is explicitly skipped so the global file is never picked up twice.

<Note>**Renamed from `config.toml`.** Older versions read `~/.aura/config.toml`. The file is still read with a one-time deprecation warning. Rename it to `cli.toml` at your convenience. The old name collided with AURA **agent** config TOMLs and will stop being read in a future release.</Note>

```toml theme={null}
# ~/.aura/cli.toml or <project>/.aura/cli.toml
api_url = "https://api.example.com"
api_key = "your-api-key"
model = "gpt-4o"
system_prompt = "You are a helpful assistant."
enable_client_tools = true   # opt in to local tool execution; default is false
log_file = "/tmp/aura.log"  # append-only; see Logging section below
```

<Note>**Note on system prompts:** In **HTTP mode**, `--system-prompt` is intended for OpenAI-compatible backends that support system messages. **AURA's server ignores system role messages**. The CLI will prompt you to confirm whether you're connecting to AURA or another service. In **standalone mode**, `--system-prompt` can append to or replace the agent's TOML-configured system prompt (you'll be asked which). In one-shot mode (`--query`), standalone silently appends; HTTP mode requires `--force`.</Note>

<Warning>**Don't commit secrets.** A project `cli.toml` checked into source control will share `api_key` with everyone who clones the repo. Keep secrets in `~/.aura/cli.toml`, in `AURA_API_KEY`, or pass them with `--api-key`.</Warning>

Any value set here is overridden by environment variables or CLI flags.

***

## One-Shot Mode

`--query <text>` runs a single round of the conversation and exits. The
output contract is strict: **stdout contains only the raw assistant
response**. Exactly what the model produced, with no bullet markers, no
themed headers, no tool-execution summaries, no markdown rendering, and no
trailing response-summary line.

Anything that isn't the response goes to **stderr**:

* Diagnostic logs from `--log-file` / `AURA_LOG_FILE` (file destination
  unchanged; the file is the same destination in REPL and one-shot mode)
* Permission prompts for local tools (interactive, on TTY)
* Errors and warnings (with `error:` / `warning:` prefixes, no markers)
* The "no result for server tool X. Set `AURA_CUSTOM_EVENTS=true`" hint

This means typical pipe usage works without scrubbing:

```bash theme={null}
aura --query "summarize the README" > summary.md
aura --query "list three ideas as JSON" | jq .
aura --query "what's the version?" 2>/dev/null | tee log.txt
```

Exit code follows the standard contract: `0` ⇒ stdout is the complete
response; non-zero ⇒ stderr explains why and stdout is empty.

The REPL retains its rich formatting (bullet markers, markdown rendering,
tool-call summaries, animated headers). The strict-output rules above
apply only to `--query` mode.

***

## SIGTERM Handling

When the CLI receives SIGTERM in one-shot or REPL mode, it prints this
warning to stderr:

```text theme={null}
warning: received SIGTERM, flushing telemetry before exit
```

The CLI then flushes buffered [product telemetry](/aura/telemetry) and, in
standalone mode, buffered OpenTelemetry spans, so a run you stop with SIGTERM
still reaches your collector. The flush takes up to about 7 seconds. The CLI
then exits with code `143`, the standard exit code for a process ended by
SIGTERM. If you send SIGKILL before the flush finishes, any telemetry still
buffered is lost.

***

## Interactive Commands

Once inside the REPL, the slash commands below are available. All slash commands can be executed while output is streaming.

| Command | Description |
| - | - |
| `/help` | Show available commands and keyboard shortcuts |
| `/clear` | Start a new conversation (saves the current one first) |
| `/expand` | Toggle expanded/compact tool call view |
| `/stream` | Toggle SSE event stream panel |
| `/conversations` | List saved conversations |
| `/resume <filter>` | Resume a saved conversation by ID prefix or name |
| `/rename <name>` | Rename the current conversation |
| `/model <filter>` | Browse and select a model (see [Model Selection](#model-selection)) |
| `/mcp` | List the active agent's configured MCP servers, or set one up with a guided wizard (see [Manage MCP Servers](#manage-mcp-servers)) |
| `/style [name]` | Switch visual style: `normal`, `high-contrast`, `no-colors` |
| `/quit` or `/exit` | Exit the REPL |

## Manage MCP Servers

`/mcp` lists the active agent's configured MCP servers (name, transport, and target). It works in both standalone and HTTP modes. `/mcp add` launches an interactive setup wizard. It isn't available in HTTP-mode builds.

The wizard walks you through the following steps:

1. Pick a server from the built-in catalog (Mezmo, PagerDuty, Datadog, or Kubernetes), or configure a custom server.
2. For a custom server, choose the transport (`http_streamable`, `sse`, or `stdio`) and enter the URL or command, plus an optional authentication header.
3. Name the server and provide credentials. You can reuse an existing environment variable or enter a value with masked input.
4. Confirm the setup. The wizard verifies the connection, then previews the exact `[mcp.servers.<name>]` block before writing the config. The preview also includes a wildcard scratchpad entry that matches every tool, which the wizard adds automatically, without asking: `[mcp.servers.<name>.scratchpad]` with `"*" = { min_tokens = 5120 }`. This entry sends any tool output over `min_tokens` (default `5120`) to scratchpad storage instead of the context window. Because the scratchpad entry is part of the same preview, declining the write declines it too.

The wizard appends the `[mcp.servers.<name>]` block to your loaded single-file agent TOML config. Configuration-directory setups are not supported. Secrets go to a `.env` file next to your config, never into the TOML file, which uses `{{ env.VAR }}` placeholders. For orchestrated configs, the wizard can scope the new server's tools to specific workers through `mcp_filter`. You must restart AURA to activate a newly added server; there is no hot reload. The wizard supplements manual TOML editing rather than replacing it. See [`[mcp]`](/aura/configuration-reference#mcp) in the configuration reference for manual and advanced MCP setup, including the three transports. The scratchpad entry stays inert until agent-level scratchpad is enabled, and the wizard does not enable it for you. If your config lacks `[agent.scratchpad] enabled = true`, the wizard prints a note that interception is configured but inactive. To activate it, add `[agent.scratchpad] enabled = true`, a top-level `memory_dir`, and `[agent.llm].context_window` to your config, then see the [Scratchpad guide](/aura/scratchpad) for the full setup.

***

## Keyboard Shortcuts

| Key | Action |
| - | - |
| `Enter` | Submit input |
| `Ctrl+C` | Cancel the current streaming request, or exit if idle |
| `Ctrl+L` | Clear the current input line; if the line is empty, clear the screen |
| `Tab` | Cycle forward through slash command and argument matches (`/model`, `/resume`) |
| `Shift+Tab` | Cycle backward through matches |
| `Esc` | Cancel tab-completion selection, or exit stream panel focus |
| `Up` / `Down` | Navigate input history |
| `Page Up` / `Page Down` | Jump 10 entries through input history |

When the **stream panel** is visible, arrow keys and page keys scroll through SSE events instead. Press `Esc` to exit stream panel focus.

***

## Model Selection

Use `/model` to browse and select which model to use for requests.

```
/model              # list all available models
/model gpt          # filter models matching "gpt"
/model gpt-4o       # narrow to "gpt-4o"; press Enter to select the unique match
```

`/model` opens a numbered picker with one model per row. A `Name` and `Description` header appears when at least one listed model has a description and the terminal is wide enough for a description column. `❯` marks the row you have Tab-selected, and `✓` marks the model currently in use:

```
     Name                       Description
  1. Mezmo Anthropic SRE Agent  Anthropic-backed config to solve all of your production problems
❯ 2. Mezmo OpenAI SRE Agent ✓   OpenAI-backed config for the same production problems
```

Descriptions are truncated to fit the terminal width. When the terminal is too narrow for a usable description column, the column is dropped rather than truncated to a stub. A listed model without a description shows a blank cell in the `Description` column; the column itself stays, and only a list where no model has a description drops it.

Descriptions come from each agent config's `[agent].description` (see the [configuration reference](/aura/configuration-reference)). In **HTTP mode**, the model list, including ids and descriptions, comes from the server's `/v1/models` endpoint. In **standalone mode**, it comes from the loaded TOML configs, using each config's `alias` or `name` plus its description.

Typing a filter that narrows to exactly one model still shows the unique-match auto-complete row, now including that model's description.

The model list, descriptions included, is remembered per conversation, so it survives restarts.

***

## Client-Side Tools

<Warning>**USE AT YOUR OWN RISK.** Enabling client-side tools gives an LLM the ability to execute commands on your machine. Shell commands, file reads/writes, filesystem search with the same privileges as the user running `aura`. Treat `--enable-client-tools` as functionally equivalent to handing the model a shell prompt. See [Client-Side Tools](/aura/client-side-tools) for the full risk model, protocol mechanics, and server/client configuration.</Warning>

By default, AURA CLI is a **pure chat client**. No local tools are advertised to the model and the REPL never executes anything on the host. Pass `--enable-client-tools` (or set `AURA_ENABLE_CLIENT_TOOLS=true`) to opt in, at which point the model can call tools like `Shell`, `Read`, and `Update` and the REPL runs them locally with permission checks ([Permissions](#permissions)).

```bash theme={null}
# Disabled (default): chat only
aura

# Enabled: local tools available, gated by the permission system
aura --enable-client-tools
AURA_ENABLE_CLIENT_TOOLS=true aura

# Explicitly disable (overrides config file)
aura --enable-client-tools=false
```

**Single-agent configurations only.** Client-side tools are not supported when the selected config has `[orchestration].enabled = true`. Tools advertised to an orchestrated config are dropped with a warning.

**Both the CLI and the server agent must opt in** for local tools to fire. See [Client-Side Tools: Backend symmetry](/aura/client-side-tools#backend-symmetry) for the full mode/effective-behavior table. Precedence for resolving the CLI flag: `--enable-client-tools` argument or `AURA_ENABLE_CLIENT_TOOLS` env var > `<project>/.aura/cli.toml` > `~/.aura/cli.toml` (`enable_client_tools = true|false`) > default (`false`). If local tools never fire when you expect them to, check that the agent's TOML has `enable_client_tools = true` (and the config is single-agent) as well as the CLI flag. In standalone mode the CLI prints a startup warning when the flag is set but no loaded config opts in.

***

## Permissions

AURA CLI includes a permission system that controls which local tools the model is allowed to execute. Permission rules only matter when client-side tools are enabled (see [Client-Side Tools](#client-side-tools)). Configure permissions by creating a `.aura/permissions.json` file in your project directory:

```json theme={null}
{
  "permissions": {
    "allow": ["ListFiles(*)", "Read(*)", "FindFiles(*)", "SearchFiles(*)"],
    "deny": ["Shell(*)"]
  }
}
```

* **Allow rules**: matching tools execute immediately without prompting.
* **Deny rules**: matching tools are blocked with a guidance message.
* **No match**: tools with no matching rule prompt you for approval.

The CLI walks up from the current working directory to find the closest
`.aura/permissions.json`, the same way it finds `cli.toml`. Permissions are
**project-scoped only**: there is no `~/.aura/permissions.json`. Running the
CLI from outside any project directory means no permissions are loaded, and
every local tool call prompts.

<Note>**Renamed from `settings.json`.** Older versions read `.aura/settings.json`. The file is still read with a one-time deprecation warning; new "always allow" rules accepted at the prompt are written to `permissions.json`, migrating any existing rules forward.</Note>

### Available Local Tools

| Tool | Description |
| - | - |
| `Read` | Read file contents (supports chunked reading with offset and limit) |
| `ListFiles` | List directory contents |
| `SearchFiles` | Search file contents with regex or literal patterns |
| `FindFiles` | Find files recursively by glob pattern |
| `FileInfo` | Get file or directory metadata |
| `Shell` | Execute shell commands (last resort) |
| `Update` | Signal intent to modify or create files |
| `CompactContext` | Compact conversation history by discarding older messages |

***

## Conversations

Conversations are automatically saved to `~/.aura/conversations/` and can be resumed:

```bash theme={null}
aura --resume abc123               # from the CLI
/conversations                     # list saved conversations
/resume abc123                     # resume by ID prefix
/rename my chat                    # name the current conversation
```

***

## Logging

The CLI is silent by default. No tracing events are emitted unless you opt
in by pointing the CLI at a log file. When set, fmt events are written to that
path in **both REPL and one-shot mode** so stdout (and any pipe consuming it)
stays untouched.

Three places can supply the path, in precedence order:

| Source | Form |
| - | - |
| CLI flag | `--log-file /tmp/aura.log` |
| Environment | `AURA_LOG_FILE=/tmp/aura.log` |
| `cli.toml` | `log_file = "/tmp/aura.log"` |

The file is opened in **append mode** and created if missing. The default
filter is info-level for the aura crates and rig request handling; override it
with `RUST_LOG` if you need different levels.

<Warning>**Log rotation, truncation, and pruning are your responsibility.** The CLI appends indefinitely. It never truncates, rotates, or compresses the file. Use `logrotate`, a cron job, `truncate -s 0`, or a shell wrapper to keep the file from growing unbounded.</Warning>

### Standalone-mode OpenTelemetry

When running in standalone mode (the default when `--api-url` is absent), the
CLI runs the agent in-process. Set `OTEL_EXPORTER_OTLP_ENDPOINT` and the CLI
will install an OpenTelemetry layer alongside (or instead of) the file fmt
layer, emitting `agent.stream` → `agent.turn` → `mcp.tool_call`, with
`orchestration.*` between them in orchestration mode. See [Tracing & Span Layout](/aura/tracing-spans) for the full span reference.

The CLI omits the `chat_completions` / `streaming_completion` infrastructure
spans because there is no HTTP layer in standalone mode. Those live on a
separate trace in the server. HTTP-mode CLIs skip OTel entirely: your
traces come from the server process.

OTel init is independent of `--log-file`; you can run with traces only
(no log file), logs only (no OTel endpoint), or both.

Other relevant OTel env vars, read by the `aura` crate:

| Variable | Purpose |
| - | - |
| `OTEL_EXPORTER_OTLP_ENDPOINT` | OTLP collector endpoint (gRPC). When unset, no OTel layer is installed even in standalone mode. |
| `OTEL_SERVICE_NAME` | Resource attribute (defaults to `aura`). |
| `OTEL_LOG_LEVEL` | Override the OTel layer's filter. Default captures `aura=trace`, `aura_cli=info`, and rig spans. |
| `OTEL_RECORD_CONTENT` | When `true`, prompt/completion/tool args/results are recorded as span attributes. |
| `OTEL_CONTENT_MAX_LENGTH` | Max bytes for content attributes (default 1000, rounded down to UTF-8 boundary). |

***

## SSE Stream Panel

Toggle with `/stream` to see raw SSE events in real time. Supported event types:

* `aura.tool_requested` / `aura.tool_start` / `aura.tool_complete`
* `aura.usage` / `aura.tool_usage`
* `aura.progress` / `aura.session_info` / `aura.reasoning`
* `aura.orchestrator.*`: multi-agent orchestration events

See the [Streaming API Guide](/aura/streaming-api-guide) for the full event reference.

***

## Compatibility

AURA CLI speaks the standard [OpenAI Chat Completions API](https://platform.openai.com/docs/api-reference/chat) and works with any compatible backend:

* AURA by Mezmo
* OpenAI API / Azure OpenAI
* Local models via Ollama, LM Studio, vLLM, etc.
* Any service implementing `/v1/chat/completions`

### Hybrid Tool Execution

When client-side tools are enabled, both backends execute server-side tools (MCP, RAG) within the agent's stream and pause for client-side tools (Shell, Read, ...) to be executed in the REPL with permission checks. The CLI follows up with a `role: "tool"` result and the agent resumes. Server-side tool results arrive via `aura.tool_complete` SSE events; the CLI uses those rather than executing locally.

When `--enable-client-tools` is off, the CLI only sees server-side tool execution. See [Client-Side Tools](#client-side-tools) for the full flow.
