> ## Documentation Index
> Fetch the complete documentation index at: https://docs.mezmo.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Configuration Reference

> Complete TOML field reference for AURA. Agent identity, LLM providers, MCP, vector stores, scratchpad, skills, orchestration, HITL, and session storage.

AURA is configured with a single TOML file. The file path defaults to `config.toml` in the working directory; override it with the `CONFIG_PATH` environment variable.

```bash theme={null}
CONFIG_PATH=configs/my-agent.toml cargo run --bin aura -- webserver
```

To validate a config file, start the web server or CLI against it. Both validate immediately and exit with a clear error if parsing fails, before binding to any port or entering the REPL:

```bash theme={null}
cargo run --bin aura -- webserver --config your-config.toml   # exits on parse error before binding
cargo run --bin aura -- --config your-config.toml             # exits on parse error before REPL
```

## Environment Variable Interpolation

Any string value in the config can reference an environment variable using `{{ env.VAR_NAME }}`. An optional `| default: 'value'` fallback prevents a hard error when the variable is unset.

```toml theme={null}
api_key = "{{ env.OPENAI_API_KEY }}"
Authorization = "Bearer {{ env.GITHUB_PERSONAL_ACCESS_TOKEN | default: '' }}"
```

***

## Root-Level Fields

| Field | Type | Default | Description |
| - | - | - | - |
| `memory_dir` | string | - | Base directory for scratchpad storage and orchestration artifact persistence. Required when scratchpad is enabled. |

```toml theme={null}
memory_dir = "/tmp/aura-sessions"
```

***

## `[agent]`

Defines the agent identity, system prompt, and behavioral settings.

| Field | Type | Default | Description |
| - | - | - | - |
| `name` | string | *required* | Display name. Doubles as the model identifier clients send in requests. |
| `alias` | string | - | Stable identifier for model selection. Clients send this as the `model` field. Useful when the name contains spaces. |
| `description` | string | - | Optional one-line, human-readable summary of the agent. Use it to distinguish multiple loaded configs that would otherwise differ only by `name` or `alias`. Display-only. It appears in `/v1/models`, `/aura/info`, and the CLI `/model` picker, but is never part of runtime settings and never sent to the model. Omitted from `/v1/models` and `/aura/info` responses when unset. Not to be confused with the required per-worker `description` under `[orchestration]`, which guides planning and is not shown to people. It is free-form text with no enforced length limit; the CLI `/model` picker collapses whitespace and truncates it to fit the terminal. |
| `system_prompt` | string | *required* | The agent's system prompt. Multi-line strings use TOML `"""..."""` syntax. |
| `turn_depth` | integer | `5` | Max tool-call rounds per user turn. Acts as a failsafe to prevent models from spinning out in unbounded tool-call loops. |
| `nudge_last_turn` | bool | `false` | Append a wrap-up warning to tool output on the final turn before `turn_depth` terminates the run (orchestration workers are told to call `submit_result`), rather than silently losing all gathered work. |
| `nudge_turns_remaining` | integer | - | Start wrap-up warnings when N or fewer tool-calling turns remain. Independent of `nudge_last_turn`; both default to off and can be enabled separately. |
| `mcp_filter` | list of strings | - | [Tool-name patterns](#tool-name-patterns) selecting which MCP tools to expose. When omitted, all tools are included. Accepts `<namespace>:<name>` to select tools from one MCP server. |
| `enable_client_tools` | bool | `false` | Allow the agent to invoke client-side tools advertised in the request. See [Client-Side Tools](/aura/client-side-tools) for security implications. |
| `client_tool_filter` | list of strings | - | [Tool-name patterns](#tool-name-patterns) narrowing which client tools are allowed (requires `enable_client_tools = true`). |
| `model_owner` | string | *(LLM provider name)* | Overrides the `owned_by` field in `/v1/models` responses. |
| `created_at` | integer | *(current time)* | Creation timestamp in milliseconds since epoch. Shown in `/v1/models` responses. |
| `hidden` | bool | `false` | Hides this agent from the `/v1/models` list. Useful for agents meant only for internal callers. |

```toml theme={null}
[agent]
name = "DevOps Assistant"
alias = "devops"
description = "General-purpose assistant with tool access"
system_prompt = """
You are a DevOps assistant with access to GitHub.
Help with code review, PR management, and repo exploration.
"""
turn_depth = 10
mcp_filter = ["get_*", "list_*", "search_*"]
model_owner = "acme"
hidden = false
```

***

## Multiple Agents

`CONFIG_PATH` can point to a single TOML file or a directory of `.toml` files. When pointed at a directory, AURA loads every `.toml` file and serves each as a selectable agent:

```
configs/
├── research-assistant.toml
├── devops-agent.toml
└── code-reviewer.toml
```

```bash theme={null}
CONFIG_PATH=configs/ cargo run --bin aura -- webserver
```

Each agent is identified by its `alias` (if set) or `name`. Clients discover available agents via `GET /v1/models` and select one by passing its identifier as the `model` field in chat completion requests. The same field tools like LibreChat, OpenWebUI, and CLI clients use to present a model picker.

Agent selection follows this order:

1. If only one config is loaded, it is always used (the `model` field is ignored).
2. Otherwise, `model` is matched first, then `DEFAULT_AGENT` if `model` is absent.
3. Returns a 400 error if multiple configs are loaded and neither `model` nor `DEFAULT_AGENT` is supplied at all.
4. Returns a 404 error if a `model` or `DEFAULT_AGENT` value is supplied but matches no loaded config.

```toml theme={null}
[agent]
name = "DevOps Assistant"
alias = "devops"             # clients send "model": "devops"
description = "DevOps assistant with GitHub access"
system_prompt = "You are a DevOps expert."
model_owner = "mezmo"        # override owned_by in /v1/models (defaults to LLM provider)
```

Aliases must be unique across all loaded configs. If two configs share the same `name` and neither has an alias, loading fails with a validation error.

**Hidden agents** are excluded from `GET /v1/models` and the CLI's `/model` list but remain fully accessible when a caller targets them by exact name or alias. Set `hidden = true` to hide agents that are in development, restricted to known callers, or should not appear in model pickers (LibreChat, OpenWebUI, etc.):

```toml theme={null}
[agent]
name = "Internal Triage Agent"
hidden = true         # excluded from /v1/models and CLI model list; still callable by name
system_prompt = "..."
```

***

## `[agent.llm]`

Configures the LLM provider. The `provider` field is a discriminant that selects the variant and its required fields.

### Common Fields

These fields are available on all providers except where noted.

| Field | Type | Default | Description |
| - | - | - | - |
| `provider` | string | - | **Required.** One of: `openai`, `anthropic`, `bedrock`, `gemini`, `ollama`, `openrouter`. |
| `model` | string | - | **Required.** Model identifier (e.g. `gpt-4o`, `claude-sonnet-4-20250514`). |
| `max_tokens` | integer | - | Maximum tokens in the response. |
| `context_window` | integer | - | Context window size in tokens. Required when scratchpad is enabled. Used for usage reporting in `aura.session_info` events. |
| `temperature` | float | - | Sampling temperature (0.0–2.0). Higher values increase randomness. |
| `additional_params` | table | - | Provider-specific parameters merged into the API request body. Useful for features like extended thinking. |

### `provider = "openai"`

```toml theme={null}
[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
max_tokens = 8192
context_window = 128000
temperature = 0.7
# base_url = "https://api.openai.com/v1"  # custom endpoint
# reasoning_effort = "medium"             # none, minimal, low, medium, high, xhigh (provider/model support varies)
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `api_key` | string | - | **Required.** OpenAI API key. |
| `base_url` | string | - | Override the API base URL (useful for compatible proxies). |
| `reasoning_effort` | string | - | `none`, `minimal`, `low`, `medium`, `high`, or `xhigh`. Controls reasoning depth. Support varies by model and version. Aura passes the value straight to the OpenAI API, which returns a 400 error for combinations of `model` and `reasoning_effort` that it does not support. |

### `provider = "anthropic"`

```toml theme={null}
[agent.llm]
provider = "anthropic"
api_key = "{{ env.ANTHROPIC_API_KEY }}"
model = "claude-sonnet-4-20250514"
context_window = 200000

# Enable extended thinking:
[agent.llm.additional_params]
thinking = { type = "adaptive", budget_tokens = 8000 }
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `api_key` | string | - | **Required.** Anthropic API key. |
| `base_url` | string | - | Override the API base URL. |

### `provider = "bedrock"`

Uses AWS credentials from the environment (AWS profile, IAM role, or environment variables). No API key field.

```toml theme={null}
[agent.llm]
provider = "bedrock"
model = "us.anthropic.claude-3-5-sonnet-20241022-v2:0"
region = "{{ env.AWS_REGION }}"
profile = "default"  # optional
context_window = 200000
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `region` | string | - | **Required.** AWS region (e.g. `us-east-1`). |
| `profile` | string | - | AWS profile name. Uses the default credential chain if omitted. |

### `provider = "gemini"`

```toml theme={null}
[agent.llm]
provider = "gemini"
api_key = "{{ env.GOOGLE_API_KEY }}"
model = "gemini-2.0-flash"
context_window = 1000000
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `api_key` | string | - | **Required.** Google API key. |
| `base_url` | string | - | Override the API base URL. |

### `provider = "ollama"`

No API key required. Defaults to `http://localhost:11434`.

```toml theme={null}
[agent.llm]
provider = "ollama"
model = "qwen3:30b-a3b"
base_url = "http://localhost:11434"
context_window = 32768
fallback_tool_parsing = true

# Pass Ollama-specific parameters:
[agent.llm.additional_params]
num_ctx = 32768
top_k = 40
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `base_url` | string | `"http://localhost:11434"` | Ollama server URL. |
| `fallback_tool_parsing` | bool | `false` | Parse tool calls from streamed text output. Enable for models that emit tool calls as text (e.g. some qwen3 variants) rather than native tool-call structures. |

### `provider = "openrouter"`

Access 300+ models through a single API key. Uses OpenRouter's reasoning wire format (`reasoning` + `reasoning_details`). For other OpenAI-compatible APIs (Fireworks, Together), use `provider = "openai"` with `base_url` instead.

```toml theme={null}
[agent.llm]
provider = "openrouter"
api_key = "{{ env.OPENROUTER_API_KEY }}"
model = "anthropic/claude-sonnet-4"
# base_url = "https://custom-endpoint/v1"  # OpenRouter-compatible proxies only
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `api_key` | string | - | **Required.** OpenRouter API key. |
| `base_url` | string | - | Override the base URL. |

***

## `[mcp]`

Configures Model Context Protocol (MCP) tool servers.

| Field | Type | Default | Description |
| - | - | - | - |
| `sanitize_schemas` | bool | `true` | Sanitize tool schemas for OpenAI function-calling compatibility. Fixes `anyOf` unions, missing types, and optional parameters that strict mode rejects. |
| `user_agent` | string | `"aura/<version>"` | The identity AURA presents to every MCP server, as a `product/version` token. AURA sends it verbatim as the HTTP `User-Agent` header on `http_streamable` and `sse` connections, and splits it into `clientInfo.name` and `clientInfo.version` for the MCP `initialize` handshake on every transport, including `stdio`. When the token has no version part, AURA reports the running AURA version instead. The value must be printable ASCII with a non-empty product part, or the configuration fails to load. Set it when your MCP servers track clients and you want to tell one AURA deployment from another. A server entry can set its own `user_agent` to replace this value for that server alone. |
| `connect_timeout_secs` | integer | `30` | Seconds to wait for a single MCP server's connect, initialize, and tool-discovery sequence to finish. If the server does not respond within this window, it is marked failed like any other connection error, rather than blocking the agent. Other servers are unaffected. |

```toml theme={null}
[mcp]
sanitize_schemas = true
user_agent = "aura-prod-us/1"  # default is "aura/<version>" (the running AURA version)
connect_timeout_secs = 30
```

A per-server `headers = { "User-Agent" = "..." }` entry replaces the header for that server only. The handshake still announces the name and version from `user_agent`.

### `[mcp.servers.<name>]`

Each server is a named entry under `[mcp.servers]`. The `transport` field selects the connection type.

The server's key (`<name>`) is also the namespace of its tools. Use it in a [tool-name pattern](#scope-a-pattern-to-one-mcp-server) to select tools from this server only. When more than one server advertises the same tool name, see [Tool Names Shared by Multiple Servers](#tool-names-shared-by-multiple-servers).

#### `transport = "http_streamable"` (recommended)

Connects to an MCP server over HTTP using the current MCP streamable transport (post-2025-11-05).

```toml theme={null}
[mcp.servers.my_tools]
transport = "http_streamable"
url = "http://localhost:8081/mcp"
description = "My tool server"

[mcp.servers.my_tools.headers]
Authorization = "Bearer {{ env.MCP_TOKEN }}"
```

#### `transport = "sse"`

Connects using the legacy SSE-based MCP protocol. Supports the same `headers`, `headers_from_request`, and `scratchpad` options as `http_streamable`.

```toml theme={null}
[mcp.servers.legacy_server]
transport = "sse"
url = "http://localhost:8082/sse"
```

#### `transport = "stdio"`

Spawns a local child process. Each agent request creates its own process instance.

`cmd` is a list where `cmd[0]` is the executable and `cmd[1..]` are fixed arguments that are part of the command (e.g. a script path). `args` are additional arguments appended after.

```toml theme={null}
[mcp.servers.everything]
transport = "stdio"
cmd = ["npx"]
args = ["-y", "@modelcontextprotocol/server-everything"]

[mcp.servers.everything.env]
MY_VAR = "value"

# Script-based example:
# cmd = ["python3", "/opt/mcp-servers/weather.py"]
# args = ["--verbose"]
```

#### Per-Server `user_agent`

Every transport accepts an optional `user_agent` that replaces `[mcp].user_agent` for that server. On `http_streamable` and `sse` it also replaces the HTTP `User-Agent` header; on every transport, including `stdio`, it replaces the handshake's `clientInfo.name` and `clientInfo.version`. Use it when one server expects a particular product token, or when you want one server to see a deployment tag that the others should not.

```toml theme={null}
[mcp.servers.my_tools]
transport = "http_streamable"
url = "http://localhost:8081/mcp"
user_agent = "aura-prod-us/1"
```

#### Static Headers

Add static headers to every request to an HTTP or SSE server:

```toml theme={null}
[mcp.servers.my_tools.headers]
Authorization = "Bearer {{ env.MCP_TOKEN }}"
X-Tenant-ID = "acme"
```

#### Header Forwarding (`headers_from_request`)

Forward headers from the incoming API request to the MCP server. The table maps outgoing header name → incoming request header name. Useful for per-user auth delegation.

```toml theme={null}
[mcp.servers.my_tools.headers_from_request]
# Outgoing header = Incoming request header
Authorization = "x-user-token"
X-User-ID = "x-user-id"
```

When `headers_from_request` is set, the forwarded header takes precedence over any matching static header from `headers`.

#### Per-Tool Scratchpad Thresholds

Override when a tool's output gets intercepted by the scratchpad system. Keys are [tool-name patterns](#tool-name-patterns) matched against this server's tool names; the most specific (longest) pattern wins.

```toml theme={null}
[mcp.servers.my_tools.scratchpad]
"get_large_*" = { min_tokens = 100 }  # intercept all large-prefixed tools
"get_small_*" = { min_tokens = 99999 } # effectively disable interception
"get_report"  = { min_tokens = 200 }  # custom threshold for one tool
```

#### Tool Names Shared by Multiple Servers

AURA presents MCP tools to the model by their bare names, so only one server's tool can hold a given name. When more than one server advertises the same tool name, the agent gets the tool from the server whose `[mcp.servers.<name>]` key sorts first, counting only servers whose tool passes the agent's or worker's `mcp_filter`. For example, `[mcp.servers.alpha]` wins over `[mcp.servers.beta]`, whatever transport each server uses. The other servers' tools with that name are unreachable, including through text-fallback tool calls (`fallback_tool_parsing`).

To use the tool from a different server, scope `mcp_filter` to that server, for example `mcp_filter = ["beta:*"]`, or rename the tool on all but one server.

The web server logs a warning for each shared name at startup, and governance catalog sync logs the same warning. The warning names every server that advertises the tool and the server that sorts first. It doesn't account for `mcp_filter`, so an agent's filter can select a different server than the warning names.

To find shared names, the web server connects to each agent's MCP servers once before it starts accepting requests. The servers are checked one at a time, and each connection can take up to [`[mcp].connect_timeout_secs`](#mcp), so leave room for the check in startup and readiness probe timeouts. The check starts each `stdio` server's command and sends only static `headers`. A server that needs headers from `headers_from_request` logs a connection warning at startup, and the web server still starts.

***

## Tool-Name Patterns

Several fields select tools by name, and they all use the same pattern syntax:

* `[agent].mcp_filter` and `[orchestration.worker.<name>].mcp_filter`
* `[hitl].require_approval`
* `[agent].client_tool_filter`
* The keys of `[mcp.servers.<name>.scratchpad]`

AURA checks every pattern when it loads the config. A pattern outside this syntax stops the web server or CLI from starting, and the error names the column of the problem. If you're upgrading a config written for an earlier AURA release, see [Breaking Changes: 5 October 2026](/aura/breaking-changes-20261005-tool-name-patterns).

A pattern is built from these parts:

| Part | Matches | Example |
| - | - | - |
| Letters, digits, `_`, `-`, `/`, and `.` | The same characters | `list_pods` matches only `list_pods`. |
| `*` | Any sequence of characters, including none | `mezmo_*` matches `mezmo_logs` and `mezmo_`. |
| `?` | Exactly one character | `tool_?` matches `tool_a` but not `tool_ab`. |
| `[abc]` | One character from the set | `pod[sx]` matches `pods` and `podx`. |
| `[!abc]` | One character not in the set | `v[!23]` matches `v1` but not `v2`. |
| `{a,b}` | Any one of the comma-separated alternatives | `get_{pods,events}` matches `get_pods` and `get_events`. |
| `<namespace>:<name>` | A tool named `<name>` from an MCP server whose key matches `<namespace>` | `github:*` matches every tool from `[mcp.servers.github]`. |

Patterns follow these rules:

* A pattern matches the whole tool name, not part of it. `list_*` matches `list_pods` but not `k8s_list_pods`.
* Matching is case-sensitive. `List*` doesn't match `list_pods`.
* No other characters are allowed, including spaces and `\`. There is no escape character, and a pattern can't be empty.
* A class holds only letters, digits, `_`, `-`, `/`, and `.`, so `[*]` is invalid.
* An alternative can contain wildcards and classes, but not another `{...}`.
* A run of letters, digits, `_`, `-`, `/`, and `.` can be at most 64 characters long. A wildcard, a class, a `{`, `,`, or `}`, or the `:` ends the run, and a class's contents don't count toward it.

### Scope a Pattern to One MCP Server

Prefix a pattern with a namespace and `:` to match tools from one MCP server. The namespace is the server's key in `[mcp.servers.<name>]`, so `github:create_issue` matches `create_issue` from `[mcp.servers.github]` and no other server. A pattern without `:` matches tools from any server.

Both segments accept the full pattern syntax, so `*:list_pods` and `git{hub,lab}:list_*` are valid. A pattern can contain only one `:`, and neither segment can be empty.

A namespaced pattern never matches a tool that doesn't come from an MCP server, such as a filesystem tool from `[tools].filesystem` or a client-side tool. The namespace form is most useful in these fields:

* **`mcp_filter`:** Select tools from one server, or pick which server provides a [shared tool name](#tool-names-shared-by-multiple-servers).
* **`[hitl].require_approval`:** Gate a tool from one server without gating a tool with the same name from another.

In `client_tool_filter`, a namespaced pattern is accepted but never matches, because client-side tools have no MCP server. In a `[mcp.servers.<name>.scratchpad]` key, the namespace is compared with that table's own server key. Under `[mcp.servers.github.scratchpad]`, the key `github:get_report` matches the same tool as `get_report`, and `gitlab:get_report` matches nothing.

This example assumes the config defines `[mcp.servers.github]` and `[mcp.servers.gitlab]`:

```toml theme={null}
[agent]
name = "Code Review Assistant"
system_prompt = "You review code changes."
# Every GitHub tool, plus only the list_* tools from GitLab
mcp_filter = ["github:*", "gitlab:list_*"]
```

The model never sees the namespace. It always receives the bare tool name, such as `create_issue`. The namespace appears in these places:

* `tool_namespace` on HITL approval webhook items and on `aura.approval_requested` and `aura.approval_pending` events. See [Human-in-the-Loop Approval Gates](/aura/hitl).
* The `tool.namespace` attribute on `mcp.tool_call` spans. See [Tracing & Span Layout](/aura/tracing-spans#mcp-tool-call-spans).
* `namespace` on each tool entry in the governance catalog.

***

## `[agent.scratchpad]`

Controls context window management. When enabled, large MCP tool outputs are saved to disk and replaced with a file pointer. The agent then uses eight exploration tools (`head`, `slice`, `grep`, `schema`, `item_schema`, `get_in`, `iterate_over`, `read`) to selectively read the data it needs. See [Scratchpad](/aura/scratchpad) for the full feature guide.

Requires `memory_dir` to be set at the root level and `context_window` to be set on `[agent.llm]`. Orchestration also accepts the legacy `[orchestration.artifacts].memory_dir` as a fallback when the top-level field is absent.

| Field | Type | Default | Description |
| - | - | - | - |
| `enabled` | bool | `false` | Activate scratchpad interception. |
| `context_safety_margin` | float | `0.20` | Fraction of the context window (0.0–1.0) reserved for reasoning and output. Scratchpad intercepts output before this budget is exhausted. |
| `max_extraction_tokens` | integer | `10000` | Maximum tokens a single exploration tool call may return, preventing a single read from flooding the context. |
| `turn_depth_bonus` | integer | `6` | Extra tool-call turns added when scratchpad is active, to leave room for exploration calls after interception. |

```toml theme={null}
memory_dir = "/tmp/aura-sessions"

[agent]
turn_depth = 10

[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
context_window = 128000

[agent.scratchpad]
enabled = true
context_safety_margin = 0.20
max_extraction_tokens = 10000
turn_depth_bonus = 6
```

***

## `[agent.skills]`

Points the agent at directories of on-demand skills. Skills follow the [Agent Skills specification](https://agentskills.io/specification): each skill is a subdirectory containing a `SKILL.md` file with YAML frontmatter (`name`, `description`). The `name` must match the directory name and consist of lowercase alphanumerics and hyphens only (1–64 characters, no leading/trailing/consecutive hyphens). See [Skills](/aura/skills) for the full feature guide.

```toml theme={null}
[[agent.skills.local]]
source = "skills/"        # relative to process CWD, or absolute

[[agent.skills.local]]
source = "/opt/shared-skills"
```

A `load_skill` tool is exposed to the agent that loads and executes skills on demand, along with a `read_skill_file` tool for fetching individual resource files from a skill's directory. Skills the agent loads stay in context on later turns of the same chat session. See [Skill Persistence Across Turns](/aura/skills#skill-persistence-across-turns).

***

## `[[vector_stores]]`

Configures RAG (Retrieval-Augmented Generation) stores. Each entry creates a `vector_search_<name>` tool the agent can call. Multiple stores create multiple search tools.

| Field | Type | Description |
| - | - | - |
| `name` | string | Unique identifier. Used in worker `vector_stores` lists. |
| `context_prefix` | string | Optional hint shown to the LLM describing what the store contains. |
| `type` | string | Store type: `in_memory`, `qdrant`, or `bedrock_kb`. |

### `type = "qdrant"`

```toml theme={null}
[[vector_stores]]
name = "docs"
type = "qdrant"
url = "http://localhost:6334"
collection_name = "documents"
context_prefix = "Technical documentation and API references"

[vector_stores.embedding_model]
provider = "openai"
model = "text-embedding-3-small"
api_key = "{{ env.OPENAI_API_KEY }}"
```

### `type = "in_memory"`

```toml theme={null}
[[vector_stores]]
name = "knowledge"
type = "in_memory"

[vector_stores.embedding_model]
provider = "openai"
model = "text-embedding-3-small"
api_key = "{{ env.OPENAI_API_KEY }}"
```

### `type = "bedrock_kb"`

Managed RAG uses AWS credentials, so no embedding model is needed.

```toml theme={null}
[[vector_stores]]
name = "company_docs"
type = "bedrock_kb"
knowledge_base_id = "{{ env.BEDROCK_KB_ID }}"
region = "{{ env.AWS_REGION }}"
# profile = "default"  # optional
# managed = true  # optional
context_prefix = "Company documentation"
```

Extra fields:

| Field | Type | Default | Description |
| - | - | - | - |
| `knowledge_base_id` | string | - | **Required.** Bedrock Knowledge Base ID. |
| `region` | string | - | **Required.** AWS region (e.g. `us-east-1`). |
| `profile` | string | - | AWS profile name. Pins the Knowledge Base client to this profile; uses the default credential chain if omitted. |
| `managed` | bool | `false` | Set to `true` when the store points at a Bedrock managed knowledge base; Aura then sends `managedSearchConfiguration` instead of `vectorSearchConfiguration` on retrieve requests. Leave it `false` or omit it for a classic vector-store-backed knowledge base. |

When you set `profile`, the Knowledge Base client uses only that AWS profile's credentials, even if static AWS environment credentials (`AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY`) are also present. When you omit `profile`, it falls back to the default AWS credential chain. This makes profile-based cross-account access reliable alongside static environment credentials used elsewhere in the deployment. For example, an IRSA (IAM Roles for Service Accounts) profile that assumes a role via web identity.

Both managed and classic Bedrock knowledge bases handle embeddings for you. The `managed` flag only selects which kind of knowledge base the store points at. If you leave `managed` unset or `false` for a managed knowledge base, retrieve requests fail with a `ValidationException`. The full error reads: `Incompatible configuration: vectorSearchConfiguration is not supported for managed knowledge bases. Use managedSearchConfiguration instead.` Set `managed = true` to resolve it.

### Embedding Models

Used by `in_memory` and `qdrant` types.

```toml theme={null}
# OpenAI:
[vector_stores.embedding_model]
provider = "openai"
model = "text-embedding-3-small"
api_key = "{{ env.OPENAI_API_KEY }}"

# AWS Bedrock:
[vector_stores.embedding_model]
provider = "bedrock"
model = "amazon.titan-embed-text-v2:0"
region = "{{ env.AWS_REGION }}"
profile = "default"  # optional
```

***

## `[tools]`

Enables built-in server-side tools.

| Field | Type | Default | Description |
| - | - | - | - |
| `filesystem` | bool | `false` | Expose basic read-only filesystem tools (read file, list directory) to the agent. |
| `custom_tools` | list of strings | `[]` | Reserved for future use. |

```toml theme={null}
[tools]
filesystem = true
```

***

## `[hitl]`

Human-in-the-loop (HITL) approval gates let an agent ask for permission before running selected MCP tools. `[hitl]` is the enable bit — there is no separate `enabled` field; presence of the table turns it on, and `route` is required when it's present. See [HITL](/aura/hitl) for the full route contracts, SSE lifecycle events, and current scope (single-agent vs. orchestration worker gating).

| Field | Type | Default | Description |
| - | - | - | - |
| `require_approval` | list of strings | `[]` | [Tool-name patterns](#tool-name-patterns) that gate a matching MCP tool call behind approval. Accepts `<namespace>:<name>` to gate tools from one MCP server. |
| `route` | table | *required* | The decision route: see `[hitl.route]` below. |

### `[hitl.route]`

Tagged by `mode`: `"conversational"` or `"webhook"`.

| Field | Type | Default | Description |
| - | - | - | - |
| `mode` | string | - | **Required.** `"conversational"` (attended, over an open SSE stream) or `"webhook"` (unattended, posts to a URL). |
| `timeout_secs` | integer | `60` (conversational) / `300` (webhook) | Seconds to wait for a decision before failing closed. |
| `url` | string | - | **Required for `webhook` mode only.** Must start with `http://` or `https://`. |

```toml theme={null}
[hitl]
require_approval = ["kubectl_*", "restart_*", "dangerous_*"]

[hitl.route]
mode = "webhook"
url = "https://approvals.example.com/aura"
timeout_secs = 300
```

***

## Session Store (Durable and Multi-Pod Deployments)

Unlike every other section on this page, the session store is **not configured in TOML**. It's deployment infrastructure (one instance per server, not per-agent), so it's configured only via environment variables.

By default, cross-request session state (A2A tasks, parked HITL approvals, and skill-invocation logs) lives in process memory. Correct for a single pod, the CLI, and local dev. Behind a load balancer with multiple replicas, configure a shared Redis/Valkey backend and every cross-request flow works no matter which pod serves each request: A2A `message:send` → poll → `list` → history-by-context, A2A `subscribe`/`cancel` against a task executing on another pod, and conversational HITL approvals resolved by a `POST /v1/approvals/{id}` that lands away from the pod that parked them.

```bash theme={null}
export AURA_SESSION_STORE=redis                          # memory (default) | redis
export AURA_SESSION_STORE_URL=redis://valkey:6379        # redis:// or rediss:// (Valkey compatible)
export AURA_SESSION_STORE_PREFIX=aura:prod               # optional namespace (default "aura")
export AURA_SESSION_STORE_CONNECT_TIMEOUT_SECS=5         # optional
export AURA_SESSION_STORE_TASK_TTL_SECS=86400            # optional; 0 = no expiry
export AURA_SESSION_STORE_SKILLS_TTL_SECS=86400          # optional; 0 = no expiry
```

`AURA_SESSION_STORE_SKILLS_TTL_SECS` sets how long a chat session's skill-invocation log lives after the agent last calls `load_skill` or `read_skill_file`. It defaults to `86400` (24 hours), and `0` disables expiry. The Redis and file backends apply it. The memory backend has no expiry and clears its logs on restart. See [Skill Persistence Across Turns](/aura/skills#skill-persistence-across-turns).

The server pings the backend at startup and fails fast if it is unreachable; `/health` reports the backend and its ping latency.

The Redis backend requires building with the `session-store-redis` cargo feature (`cargo build --release --features aura-cli/session-store-redis`). The in-memory backend is always available. See [the session storage design doc](https://github.com/mezmo/aura/blob/main/docs/design/session-storage.md) for the design and Helm packaging roadmap.

***

## `[orchestration]`

Enables multi-agent orchestration mode. A coordinator agent decomposes user queries into tasks and delegates them to specialized worker agents for parallel execution.

The coordinator's system prompt comes from `[agent].system_prompt`. Workers are defined in `[orchestration.worker.<name>]` sections.

Execution loop:

* `Plan`: coordinator decomposes the request into a task DAG.
* `Execute`: dependency-ready tasks run in parallel waves on worker agents.
* `Continue`: coordinator consolidates worker outputs and routes to a final response, replan, or clarification.

Workers run with isolated task context windows and filtered MCP/vector-store access based on each worker block.

| Field | Type | Default | Description |
| - | - | - | - |
| `enabled` | bool | `false` | Activate orchestration mode. |
| `max_planning_cycles` | integer | `3` | Maximum plan→execute→continue cycles per request. |
| `max_plan_parse_retries` | integer | `3` | Retries on plan parse failure before falling back to single-task execution. |
| `allow_direct_answers` | bool | `true` | Allow the coordinator to answer simple queries directly without delegating. |
| `allow_clarification` | bool | `true` | Allow the coordinator to ask for clarification on ambiguous queries. |
| `tools_in_planning` | string | `"summary"` | How much tool info is shown during planning: `"none"`, `"summary"` (tool names per worker), or `"full"` (names + descriptions). |
| `max_tools_per_worker` | integer | `10` | Truncates the tool list per worker in the planning prompt with a `(+N more)` suffix. |
| `coordinator_vector_stores` | list of strings | `[]` | Names of vector stores the coordinator can access. |
| `worker_system_prompt` | string | - | Custom system prompt injected into generic (non-specialized) workers. |
| `duplicate_call_nudge_threshold` | integer | `3` | Consecutive identical tool calls before appending a guidance annotation. |
| `duplicate_call_block_threshold` | integer | `5` | Consecutive identical tool calls before appending an abort annotation. |

```toml theme={null}
[orchestration]
enabled = true
max_planning_cycles = 3
allow_direct_answers = true
tools_in_planning = "full"
coordinator_vector_stores = ["runbooks"]
memory_dir = "/tmp/aura-orchestration"   # legacy flat field; prefer top-level memory_dir
```

### `[orchestration.timeouts]`

| Field | Type | Default | Description |
| - | - | - | - |
| `per_call_timeout_secs` | integer | `0` (disabled) | Per-call timeout for coordinator and worker LLM calls. Set to a positive value to enable. |
| `stream_inactivity_timeout_secs` | integer | `0` (disabled) | Max seconds of silence between stream items on coordinator and worker orchestration streams before Aura fails the affected task with a no-stream-progress error. Suspends during tool execution and re-arms on each stream item. Must be below `per_call_timeout_secs` (when set) to take effect. Set to a positive value to enable. |

See [request lifecycle](/aura/request-lifecycle) for the operational guidance and caveats on this timeout.

```toml theme={null}
[orchestration.timeouts]
per_call_timeout_secs = 120
stream_inactivity_timeout_secs = 60
```

### `[orchestration.artifacts]`

Controls persistence and artifact promotion.

| Field | Type | Default | Description |
| - | - | - | - |
| `memory_dir` | string | - | Base directory for run artifacts. Alias: `memory_path`. (Prefer top-level `memory_dir`.) |
| `result_artifact_threshold` | integer | `4000` | Character threshold for writing worker results to artifact files instead of inlining them. |
| `result_summary_length` | integer | `2000` | Max inline summary length when a result is promoted to an artifact. |
| `session_history_turns` | integer | `3` | Max prior run manifests injected as session context in the coordinator preamble. Set to `0` to disable. |
| `persistence_drain_timeout_ms` | integer | `2000` | Timeout (ms) for draining in-flight persistence writes between execution phases. |
| `tool_output_artifact_threshold` | integer | `500` | Character threshold for promoting tool outputs to artifact files. |
| `tool_output_duration_threshold_ms` | integer | `5000` | Duration threshold (ms) for promoting tool outputs regardless of size. |
| `show_tool_reasoning_in_continuation` | bool | `false` | Include condensed tool reasoning traces in continuation prompts so the coordinator can see why workers called specific tools. |
| `max_session_runs` | integer | `20` | Max run directories retained per session before oldest runs are pruned. Set to `0` to disable pruning. |

```toml theme={null}
[orchestration.artifacts]
memory_dir = "/tmp/aura-orchestration"
result_artifact_threshold = 4000
result_summary_length = 2000
session_history_turns = 3
max_session_runs = 20
```

### `[orchestration.worker.<name>]`

Defines a specialized worker. Worker names must be unique case-insensitively (to avoid filesystem collisions) and non-empty.

| Field | Type | Default | Description |
| - | - | - | - |
| `description` | string | - | **Required.** One-line description shown to the coordinator during planning. |
| `preamble` | string | - | **Required.** Complete system prompt for this worker (replaces the generic worker template). |
| `mcp_filter` | list of strings | - | [Tool-name patterns](#tool-name-patterns) selecting which MCP tools this worker can access. When omitted, all tools are included; an empty list `[]` grants no tools. Accepts `<namespace>:<name>` to select tools from one MCP server. |
| `vector_stores` | list of strings | `[]` (none) | Names of vector stores this worker can access. Workers have no RAG access by default. |
| `turn_depth` | integer | *(inherits `[agent].turn_depth`)* | Max tool-call rounds for this worker. |

```toml theme={null}
[orchestration.worker.operations]
description = "Logs, pipelines, metrics, and system analysis"
preamble = """
You are an Operations Specialist with access to observability tools.
Use tools to investigate incidents — do not fabricate data.
"""
mcp_filter = ["mezmo_*"]
vector_stores = ["runbooks"]
turn_depth = 8
```

#### Per-Worker LLM Override

Workers inherit `[agent.llm]` by default. Provide `[orchestration.worker.<name>.llm]` to use a different model for a specific worker (e.g. a cheaper model for simple tasks).

```toml theme={null}
[orchestration.worker.summarizer.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o-mini"
context_window = 128000
```

#### Per-Worker Scratchpad Override

```toml theme={null}
[orchestration.worker.analyst.scratchpad]
enabled = true
max_extraction_tokens = 5000
```

#### Per-Worker Skills Override

`None` (field absent) inherits `[agent.skills]`. An explicit empty list disables skills. A non-empty list replaces the agent's skills entirely (no merging).

```toml theme={null}
[[orchestration.worker.researcher.skills.local]]
source = "skills/research"
```

***

## Complete Examples

### Minimal: OpenAI

```toml theme={null}
[agent]
name = "Assistant"
system_prompt = "You are a helpful assistant."

[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
```

### Single Agent with MCP Tools

```toml theme={null}
[mcp]
sanitize_schemas = true

[mcp.servers.github]
transport = "http_streamable"
url = "https://api.githubcopilot.com/mcp/"
description = "GitHub repository operations"

[mcp.servers.github.headers]
Authorization = "Bearer {{ env.GITHUB_PERSONAL_ACCESS_TOKEN }}"

[agent]
name = "DevOps Assistant"
system_prompt = """
You are a DevOps assistant with access to GitHub.
Help with code review, PR management, and repo exploration.
"""
turn_depth = 10

[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
```

### Single Agent with Scratchpad

Large tool outputs are intercepted and saved to disk; the agent uses exploration tools to read what it needs.

```toml theme={null}
memory_dir = "/tmp/aura-sessions"

[agent]
name = "Data Analyst"
system_prompt = """
You are a data analysis assistant. When you see a scratchpad pointer
([scratchpad: file_id=...]), use the exploration tools (head, grep, schema,
get_in, etc.) to selectively read the data you need. Do not re-call the
original tool.
"""
turn_depth = 10

[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
context_window = 128000

[agent.scratchpad]
enabled = true
context_safety_margin = 0.20
max_extraction_tokens = 10000

[mcp.servers.data_api]
transport = "http_streamable"
url = "http://localhost:9000/mcp"

[mcp.servers.data_api.scratchpad]
"get_dataset_*" = { min_tokens = 100 }
```

### Multi-Agent Orchestration with Per-Worker Models

```toml theme={null}
memory_dir = "/tmp/aura-orchestration"

[agent]
name = "SRE Coordinator"
system_prompt = """
You are an SRE coordinator. Decompose observability tasks into sub-tasks
and delegate to the appropriate specialist worker.
- Kubernetes cluster state and workloads: delegate to k8s-specialist
- Prometheus metrics and alerts: delegate to metrics-analyst
"""

[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
context_window = 128000

[mcp.servers.kubernetes]
transport = "http_streamable"
url = "http://localhost:8081/mcp"

[mcp.servers.prometheus]
transport = "http_streamable"
url = "http://localhost:8082/mcp"

[orchestration]
enabled = true
max_planning_cycles = 3
tools_in_planning = "full"
allow_direct_answers = true

[orchestration.timeouts]
per_call_timeout_secs = 120

[orchestration.worker.k8s-specialist]
description = "Kubernetes cluster inspection: namespaces, workloads, pods, events"
turn_depth = 8
mcp_filter = ["namespaces_list", "pods_*", "resources_*", "events_list"]
preamble = """
You are a Kubernetes Specialist. Use tools to inspect the cluster.
Never guess cluster state — always call the relevant tool first.
"""

[orchestration.worker.k8s-specialist.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
context_window = 128000

[orchestration.worker.metrics-analyst]
description = "Prometheus metrics queries, target health, and alert status"
turn_depth = 8
mcp_filter = ["execute_query", "execute_range_query", "list_metrics", "get_targets"]
preamble = """
You are a Prometheus Analyst. Query metrics directly — do not fabricate values.
Report metric names, labels, and exact values from tool results.
"""

[orchestration.worker.metrics-analyst.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o-mini"
context_window = 128000
```

### RAG with Qdrant

```toml theme={null}
[[vector_stores]]
name = "docs"
type = "qdrant"
url = "http://localhost:6334"
collection_name = "documents"
context_prefix = "Technical documentation and API references"

[vector_stores.embedding_model]
provider = "openai"
model = "text-embedding-3-small"
api_key = "{{ env.OPENAI_API_KEY }}"

[[vector_stores]]
name = "runbooks"
type = "qdrant"
url = "http://localhost:6334"
collection_name = "runbooks"
context_prefix = "Operational runbooks and incident response procedures"

[vector_stores.embedding_model]
provider = "openai"
model = "text-embedding-3-small"
api_key = "{{ env.OPENAI_API_KEY }}"

[agent]
name = "Knowledge Assistant"
system_prompt = """
You are a knowledge assistant. Use the vector_search_docs and
vector_search_runbooks tools to ground your answers in documentation.
"""

[agent.llm]
provider = "openai"
api_key = "{{ env.OPENAI_API_KEY }}"
model = "gpt-4o"
```

### Local Ollama (No API Key)

```toml theme={null}
[agent]
name = "Local Assistant"
system_prompt = "You are a helpful assistant running locally."

[agent.llm]
provider = "ollama"
model = "qwen3:30b-a3b"
base_url = "http://localhost:11434"
context_window = 32768
fallback_tool_parsing = true

[agent.llm.additional_params]
num_ctx = 32768
```

### AWS Bedrock with Knowledge Base

```toml theme={null}
[[vector_stores]]
name = "company_kb"
type = "bedrock_kb"
knowledge_base_id = "{{ env.BEDROCK_KB_ID }}"
region = "{{ env.AWS_REGION }}"
context_prefix = "Company internal documentation"

[agent]
name = "Enterprise Assistant"
system_prompt = """
You are an enterprise assistant. Use the company knowledge base
to answer questions grounded in internal documentation.
"""

[agent.llm]
provider = "bedrock"
model = "us.anthropic.claude-3-5-sonnet-20241022-v2:0"
region = "{{ env.AWS_REGION }}"
context_window = 200000
```
