Agents Endpoint Object Structure
This page applies to the agents endpoint.
Example
endpoints:
agents:
modelResponseBodyTimeoutMs: 900000
modelResponseHeadersTimeoutMs: 300000
recursionLimit: 50
maxRecursionLimit: 100
maxSubagents: 20
codeApiUploadConcurrency: 3
codeApiMaxRetryWaitMs: 20000
maxToolCallArgBytes: 65536
maxDeltaEventsPerTurn: 100000
maxToolCallArgBytesByTool:
create_file: 131072
disableBuilder: false
# (optional) Agent Capabilities available to all users. Omit the ones you wish to exclude. Defaults to list below.
# Opt-in capabilities: "programmatic_tools", "stateful_code_sessions", "run_in_background", and "tool_intents".
# capabilities: ["deferred_tools", "execute_code", "file_search", "web_search", "artifacts", "subagents", "actions", "context", "skills", "memory", "ask_user_question", "tools", "chain", "ocr"]
# (optional) File citation configuration for file_search capability
maxCitations: 30 # Maximum total citations in responses (1-50)
maxCitationsPerFile: 7 # Maximum citations from each file (1-10)
minRelevanceScore: 0.45 # Minimum relevance score threshold (0.0-1.0)
titleTiming: immediate
activityLabel: true
activityModel: gpt-4.1-nano
activityPhaseLabel: true
activityPhaseMaxPerRun: 5
skills:
maxCatalogSkills: 20
statefulCodeSessions:
allowedEnvironments: ['user', 'agent-user', 'conversation']
principalWorkers:
enabled: true
maxPerUser: 5
environments:
- id: managed-default
name: Managed Code API
type: managed
baseURL: https://code.example.com/v1
default: true
- id: engineering-vm
name: Engineering VM
type: attached
baseURL: https://code-bridge.example.com/v1
pairing:
workerId: engineering-vm
tokenEnv: CODE_BRIDGE_ADMIN_TOKEN
- id: personal-workers
name: Personal Code Workers
type: attached
baseURL: https://code-bridge.example.com/v1
pairing:
allowPrincipalWorkers: true
tokenEnv: CODE_BRIDGE_ADMIN_TOKEN
# eventDriven:
# selfUrl: https://librechat.internal
backgroundTasks:
completionWakeups: true
ordinaryToolCancellation: false
toolApproval:
enabled: true
mode: default
deny: ['mcp:*:delete_*']
ask: ['mcp:*:*']
checkpointer:
type: mongo
ttl: 86400
managementApi:
auth:
oidc:
enabled: true
issuer: https://identity.example.com/
audience: https://librechat.example.com/agents
# tokenUse: access
# requiredScopes:
# - agents-api/manage
clients:
- clientId: machine-client-id
userId: 507f1f77bcf86cd799439011
tenantId: tenant-id
remoteApi:
auth:
oidc:
enabled: falseThis configuration enables the builder interface for agents.
recursionLimit
| Key | Type | Description | Example |
|---|---|---|---|
| recursionLimit | Number | Sets the default number of steps an agent can take in a run. | Controls recursion depth to prevent infinite loops. When the limit is reached, LibreChat preserves the partial turn and offers Keep going or Answer now. This value can be configured from the UI up to maxRecursionLimit. |
Default: 25
Example:
recursionLimit: 50For more information about agent steps, see Max Agent Steps.
maxRecursionLimit
| Key | Type | Description | Example |
|---|---|---|---|
| maxRecursionLimit | Number | Sets the absolute maximum number of steps an agent can take in a run. | Defines the upper limit for the recursionLimit that can be set from the UI. This prevents users from setting excessively high values. |
Default: If omitted, defaults to the value of recursionLimit or 50 if recursionLimit is also omitted.
Example:
maxRecursionLimit: 100For more information about agent steps, see Max Agent Steps.
maxProviderErrorChars
Controls how much provider error text LibreChat preserves when an Agent failure cannot be classified. The default is 2000 characters. Set it to 0 to omit provider text; the maximum is 1000000.
maxProviderErrorChars: 2000Model response transport timeouts
These settings bound Agent model calls that use OpenAI-compatible transports, including compatible custom endpoints.
| Key | Type | Description | Example |
|---|---|---|---|
| modelResponseBodyTimeoutMs | Integer | Maximum inactivity between provider response-body chunks. Each chunk resets the timer. | 900000 |
| modelResponseHeadersTimeoutMs | Integer | Maximum wait for provider response headers. | 300000 |
Set either value to 0 to disable that timeout. Both accept values from 0 through 86400000 milliseconds (24 hours).
modelResponseBodyTimeoutMs: 900000
modelResponseHeadersTimeoutMs: 300000repositoryInstructions
Controls repository instruction discovery for attached Code environments. timeoutMs bounds the instruction-loading request from 100 through 30000 milliseconds and defaults to 2000.
repositoryInstructions:
timeoutMs: 2000Stream circuit breakers
These settings stop an Agent run when a provider stream appears to be producing runaway tool arguments or an unbounded sequence of delta events.
| Key | Type | Description | Example |
|---|---|---|---|
| maxToolCallArgBytes | Number | Maximum cumulative bytes for one streamed tool call's arguments. Set to 0 to disable the global guard. | 65536 |
| maxDeltaEventsPerTurn | Number | Maximum streamed events for one model generation turn. Set to 0 or omit it to disable this guard. | 0 (disabled) |
| maxToolCallArgBytesByTool | Object | Per-tool argument-byte overrides keyed by the model-facing tool name. A value of 0 disables the guard for that tool. | create_file: 131072 |
LibreChat uses a 64 KiB global argument limit and ships a 128 KiB override for create_file, whose legitimate payloads are often larger. Entries in maxToolCallArgBytesByTool merge over that built-in override. If either guard trips, LibreChat aborts the in-flight provider stream and reports which limit was exceeded.
maxToolCallArgBytes: 65536
maxDeltaEventsPerTurn: 100000
maxToolCallArgBytesByTool:
create_file: 262144
large_import: 524288To disable every argument-size guard, set both the global limit and the shipped create_file override to 0:
maxToolCallArgBytes: 0
maxToolCallArgBytesByTool:
create_file: 0titleTiming
| Key | Type | Description | Example |
|---|---|---|---|
| titleTiming | String | Controls when conversation titles are generated for the agents endpoint. Valid values: "immediate" or "final". | Defaults to "immediate". |
Default: "immediate"
Available Values:
"immediate": Generates the title as soon as the request starts, in parallel with the model response, using the user's first message."final": Defers title generation until the full response completes. This preserves the legacy behavior.
Example:
titleTiming: immediateactivityLabel
Set activityLabel: true to group each block of Agent reasoning and tool calls under a generated one-line header. Use activityModel, activityEndpoint, activityPrompt, activityMaxPerRun, and activityCharLimit to control the additional model call and its cost. See Agent Activity Groups for the full field reference and precedence rules.
activityLabel: true
activityEndpoint: openAI
activityModel: gpt-4.1-nano
activityMaxPerRun: 20
activityCharLimit: 600
activityPhaseLabel: true
activityPhaseMaxPerRun: 5activityPhaseLabel can independently add one collapsed parent summary around each run phase that contains at least two logical activities. Use activityPhaseModel, activityPhaseEndpoint, activityPhasePrompt, and activityPhaseMaxPerRun to tune those calls; activityCharLimit is shared with child activity labels.
disableBuilder
| Key | Type | Description | Example |
|---|---|---|---|
| disableBuilder | Boolean | Controls the visibility and use of the builder interface for agents. | When set to `true`, disables the builder interface for the agent, limiting direct manual interaction. |
Default: false
Example:
disableBuilder: falseallowedProviders
| Key | Type | Description | Example |
|---|---|---|---|
| allowedProviders | Array/List of Strings | Specifies a list of endpoint providers (e.g., "openAI", "anthropic", "google") that are permitted for use with the Agents feature. | If defined, only agents configured with these providers can be initialized. If omitted or empty, all configured providers are allowed. |
Default: [] (empty list, all providers allowed)
Note: Must be one of the following, or a custom endpoint name as defined in your configuration: - openAI, azureOpenAI, google, anthropic, assistants, azureAssistants, bedrock
Example:
allowedProviders:
- openAI
- googlecapabilities
| Key | Type | Description | Example |
|---|---|---|---|
| capabilities | Array/List of Strings | Specifies the agent capabilities available to all users for the agents endpoint. | Defines the agent capabilities that are available to all users for the agents endpoint. You can omit the capabilities you wish to exclude from the list. |
Default: ["deferred_tools", "execute_code", "file_search", "web_search", "artifacts", "subagents", "actions", "context", "skills", "memory", "ask_user_question", "tools", "chain", "ocr"]
programmatic_tools, stateful_code_sessions, run_in_background, and tool_intents are opt-in and must be added explicitly. Programmatic tools and stateful sessions also require execute_code and a compatible Code Interpreter deployment. The Agent Builder disables Programmatic MCP selections without Code Interpreter, and the server strips stale programmatic caller options when code execution is unavailable. run_in_background makes Code Interpreter tools eligible by default and allows per-tool opt-in for eligible MCP, Plugin, and Action tools; background code execution also requires execute_code and Code Interpreter.
stateful_code_sessions is highly experimental. It requires a separate stateful-profile Code Interpreter route configured with LIBRECHAT_CODE_BASEURL_STATEFUL or statefulCodeSessions.environments. Stateful requests fail closed instead of falling back to the normal stateless service. In the Agent Builder, each enabled Agent chooses a user, agent-and-user, or conversation workspace scope and, when named backends are configured, an execution environment. New Agents start with the user's personal scope default and the deployment's default execution environment.
Example:
capabilities:
- 'deferred_tools'
# Optional: enables Programmatic Tool Calling for MCP tools marked Programmatic in the Agent Builder.
# Requires execute_code and a Code Interpreter deployment with the Tool Call Server component.
# - 'programmatic_tools'
- 'execute_code'
- 'file_search'
- 'web_search'
- 'artifacts'
- 'subagents'
- 'actions'
- 'context'
- 'skills'
- 'memory'
- 'ask_user_question'
- 'tools'
- 'chain'
- 'ocr'
# Optional and highly experimental: reuse a user-, agent-and-user-, or
# conversation-scoped workspace. Requires LIBRECHAT_CODE_BASEURL_STATEFUL
# or a named statefulCodeSessions environment.
# - 'stateful_code_sessions'
# Optional: allow eligible Code Interpreter, MCP, Plugin, and Action tools to run as background tasks.
# - 'run_in_background'
# Optional: show live model-written intent labels for native and selected MCP tools.
# - 'tool_intents'Note: This field is optional. If omitted, the default behavior is to include all the capabilities listed in the default.
statefulCodeSessions
Configures the sharing scopes and optional execution backends for highly experimental stateful Code Interpreter workspaces.
Highly experimental
Stateful Code Sessions, named execution environments, attached workers, and pairing are in an early experimentation phase. Their behavior, configuration, persistence characteristics, and integration protocol may change substantially.
| Key | Type | Description | Example |
|---|---|---|---|
| allowedEnvironments | Array/List of Strings | Allowed stateful workspace scopes: `user`, `agent-user`, and/or `conversation`. | ["user", "agent-user", "conversation"] |
| environments | Array of Objects | Optional named managed or attached Code API environments shown in the Agent Builder. | Unset; LIBRECHAT_CODE_BASEURL_STATEFUL remains the stateful endpoint |
| principalWorkers | Object | Optional deployment ceiling for personal worker enrollment. `enabled` defaults to true and `maxPerUser` defaults to 5. | { enabled: true, maxPerUser: 5 } |
| conversationMoves | Object | Opt-in owner-authorized workspace transitions for saved conversations. `allowAttachDetach` additionally permits attaching or leaving an attached workspace. | Disabled |
Attached environments accept environments[].configSchema.limits.maxQueueWaitMs, which controls how long a request may retry waiting for workspace capacity. The default and maximum are 300000 milliseconds (5 minutes). Set it to 0 to disable client retries after the initial admission attempt; it does not cancel work that is already admitted or running. These fields belong under configSchema.limits, not userConfig.bash.
environments[].configSchema.limits.maxRequestTimeoutMs optionally bounds the total HTTP time for one workspace tool call, including admission, execution, settlement, and delivery. It accepts 1 through 610000 ms, but the config rejects budgets too small to leave any command execution time after admission and the 10-second settlement/delivery reserve. Without it, admission still has a 30-second budget per attempt; maxQueueWaitMs remains a separate retry horizon. Do not enable longer admission until every LibreChat API replica understands this field, the Code API route honors per-request queue allowances, and the shortest timeout across the live proxy/ingress path has been measured. An expired mutation may have completed on the worker, so LibreChat does not retry an ambiguous timeout.
With a total request budget, environments[].configSchema.limits.minCommandAdmissionMs reserves time for a command to enter the worker before execution starts. The default is 10000 ms; set 1000–300000 ms. The request budget must exceed this allowance plus 10000 ms of settlement/delivery grace; without a total budget, the setting has no effect. The Bash execution limit advertised to an Agent is the smaller of its configured or upstream ceiling and the time left after those reserves. For example, a 90000 ms request budget with an 80000 ms command ceiling and the default reserve advertises at most a 70000 ms command.
statefulCodeSessions:
allowedEnvironments:
- 'user'
- 'agent-user'
- 'conversation'
principalWorkers:
enabled: true
maxPerUser: 5
environments:
- id: managed-default
name: Managed Code API
type: managed
baseURL: https://code.example.com/v1
default: true
- id: engineering-vm
name: Engineering VM
type: attached
baseURL: https://code-bridge.example.com/v1
owner: deployment
pairing:
workerId: engineering-vm
tokenEnv: CODE_BRIDGE_ADMIN_TOKEN
- id: personal-workers
name: Personal Code Workers
type: attached
baseURL: https://code-bridge.example.com/v1
owner: deployment
configSchema:
limits:
maxCommandTimeoutMs: 120000
# Before enabling, verify Code API queue allowances and the live proxy deadline.
# maxRequestTimeoutMs: 160000
# minCommandAdmissionMs: 10000
# maxQueueWaitMs: 300000
# After upgrading Code API and opting in the worker with --linked-worktree-lanes:
# workspaces:
# linkedWorktrees: true
# Optional: require exact matching even on workers supporting tolerant edits.
# edits:
# tolerantMatching: false
permissions:
fileWrite:
allowed: [allow, ask, deny]
default: ask
commandExecution:
allowed: [ask, deny]
default: ask
pairing:
allowPrincipalWorkers: true
tokenEnv: CODE_BRIDGE_ADMIN_TOKENallowedEnvironments is required whenever the block is present and must include at least one workspace scope. Omit the entire block to allow all three scopes and continue routing stateful work through LIBRECHAT_CODE_BASEURL_STATEFUL. The personal default and Agent Builder only permit allowed values. If an administrator tightens the policy, a saved personal default remains visible but disabled, and new Agents use the first allowed scope. Existing enabled Agents retain their saved scope, but a now-disallowed scope fails closed at runtime until the Agent is reconfigured.
conversationMoves.enabled: true lets the conversation owner explicitly move a sealed attached workspace to a compatible environment and recover a missing workspace without replacing chat history or restoring old workspace files. To additionally allow a saved chat to attach a workspace or leave one for normal chat, set conversationMoves.allowAttachDetach: true. Both switches are off by default; enabling moves alone retains move-only behavior. Upgrade all API replicas, then enable CODE_ENVIRONMENT_DECISION_VERSION=1 before opting in. Ownership, target access, active work, and current policy are rechecked at the transition; a running generation cannot be moved.
statefulCodeSessions:
allowedEnvironments: ['user', 'agent-user', 'conversation']
conversationMoves:
enabled: true
allowAttachDetach: trueprincipalWorkers controls self-service enrollment for personal attached workers. Omitted values resolve to enabled: true and maxPerUser: 5; set enabled: false or maxPerUser: 0 to stop new enrollment. The limit counts every registered personal environment for that user across configured control planes, including offline workers. Lowering the limit or disabling enrollment does not revoke existing environments.
Deployment YAML is the hard ceiling. A role, group, or user configuration override may disable enrollment or lower maxPerUser, but cannot raise the deployment value. A user sees pairing-enabled control planes only when the effective limit is greater than zero, pairing.allowPrincipalWorkers is enabled on that control plane, principal authentication is available, and the user has Code Environment management permission.
When environments is non-empty, every entry requires a unique id, a display name, a type of managed or attached, and an HTTP(S) baseURL without a query string or fragment. If the list contains executable environments, exactly one of those entries must set default: true. The Agent Builder then adds an Execution environment selector; leaving it on Deployment default uses that default entry. A list containing only self-service pairing control planes needs no default. An Agent that names an environment that is later removed or no longer accessible fails closed instead of falling back to another backend.
managedpoints directly to an operator-managed stateful Code API.attachedpoints to a compatible Code APIremote-bridgebackend that leases work to an outbound@librechat/codeworker. The worker does not require an inbound public port, but its local execution endpoint still needs appropriate isolation.workerIdoptionally routes an attached environment to one outbound worker. It is server-only and is not exposed through client startup configuration.ownerdefaults todeployment. Principal-owned environments are created through LibreChat's authenticated Code Environments API, protected by ACLs, and merged only for authorized requests; do not place client-supplied URLs in deployment YAML.pairingis optional and valid only for deployment-ownedattachedentries.pairing.workerIdidentifies an operator-managed worker to enroll, whilepairing.allowPrincipalWorkers: truelets authorized users enroll owner-bound workers against that control plane. At least one of those fields is required whenpairingis present.pairing.tokenEnvnames the environment variable containing the bridge administrator token; it never contains the token itself. Pairing requires HTTPS except for loopback development.configSchema.permissionsoptionally lets authorized owners choosefileWriteandcommandExecutionpolicy for personal environments under Settings > Code environments. Each field requires a non-emptyallowedlist containingallow,ask, and/ordeny;defaultisaskwhen omitted and must appear inallowed. LibreChat rejects stored values outside the administrator's list.configSchema.limits.maxCommandTimeoutMssets the largest Bash timeout the model may request for that environment. It accepts1through300000ms; omission preserves the 30-second default. With a total request budget, the advertised command timeout also fits the remaining time after command admission, settlement, and delivery. This is a deployment ceiling and is not user-configurable.configSchema.workspaces.linkedWorktrees: trueenables per-worktree scheduling only for attached workspaces advertisinggit_linked_worktree; omission keeps all requests in the checkout's root lane. Upgrade LibreChat before a worker advertises this scope, deploy a Code API supporting linked-worktree lanes, then start the worker with--linked-worktree-lanesbefore enabling the YAML setting. File operations targeting a path inside.worktrees/<name>and Bash commands with acwdthere use that lane and return checkout-relative paths. Commands thatcdinside their command text, checkout-root operations, conversation workspace instances, and environment actions stay root-scoped.configSchema.edits.tolerantMatchingpermits whitespace-tolerant fallback by default on attached workers advertisingtolerant_match; set it tofalseto require exact matching, including protected edit previews. Matching must remain unique unless the edit explicitly setsreplace_all. Older workers receive no new matching fields; attachedreplace_allis refused before dispatch unless the worker advertises it. These capabilities require compatible Code API and worker builds implementing Code Interpreter #271.- Do not configure
settingsin deployment YAML. LibreChat supplies that request-scoped field from the authenticated owner's server-validated preferences. - Only file-write and command-execution policy is user-configurable. Isolation, networking, mounts, privileged execution, ingress, egress, and secrets remain administrator-controlled.
- If both top-level
workerIdandpairing.workerIdare present, they must match. An entry withallowPrincipalWorkers: trueand neither worker ID is pairing-only: it is omitted from execution selectors and cannot setdefault: true.
Attached execution automatically enables a safe approval baseline: file writes and command or code execution ask for confirmation, while read-only and search operations continue under the regular tool policy. The baseline is scoped to the Agent using the attached environment. If no user setting is available, LibreChat uses ask; toolApproval.enabled: false is the administrator emergency override that disables this attached-environment baseline. Callers without approval and resume support fail closed unless that explicit override is set.
Attached execution requires a matching experimental Code Interpreter remote-bridge build. Pairing authenticates the outbound worker connection; it does not replace sandbox or VM isolation. Users with Code Environment management permission pair, configure permitted tool policy, and revoke personal workers under Settings > Code environments. See Stateful Code Sessions.
eventDriven
Configures the trusted internal origin used by Agent event delivery. Most deployments should omit this block and use the current process's bound listener.
| Key | Type | Description | Example |
|---|---|---|---|
| selfUrl | URL | Overrides the current process bound listener for trusted internal trigger admission. Use only when delivery must traverse another trusted HTTP origin. | |
| idlePolling | Object | Caps idle MongoDB recovery scans and waiting completion rechecks. Local events can wake a scan sooner. | See defaults below |
endpoints:
agents:
eventDriven:
selfUrl: https://librechat.internalBound child continuations are automatic. Older childTurns, completionWakeups, coalescing, actorMailbox, checkpointForks, and durableReceipts fields under eventDriven are no longer configuration switches and should be removed. The runtime now applies those delivery guarantees directly. The separate backgroundTasks.completionWakeups field controls conversational delivery for completed background tools and detached Subagents.
Leave selfUrl unset for the normal bound-listener path. Set it only when trusted internal trigger admission must traverse another HTTP(S) origin, such as a TLS front door. AGENT_TRIGGERS_SELF_URL remains a compatibility fallback when this field is omitted. See Agents API events and Automatic Parent Continuation.
eventDriven.idlePolling caps per-replica recovery polling when nothing is ready: deliveryMaxIntervalMs defaults to 15000 (valid 1000–300000), queuedTurnMaxIntervalMs and maintenanceMaxIntervalMs each default to 120000 (valid 30000–300000), and completionWaitMaxIntervalMs defaults to 60000 (valid 5000–300000). The latter bounds rechecks while a background or Subagent result is pending or its parent turn is busy. Readiness events expedite delivery; these caps retain cross-replica and crash-recovery fallbacks rather than defining normal completion latency.
eventDriven:
idlePolling:
deliveryMaxIntervalMs: 15000
queuedTurnMaxIntervalMs: 120000
maintenanceMaxIntervalMs: 120000
completionWaitMaxIntervalMs: 60000Event Actor detached Action completion is selected automatically from the built-in generation store. In-memory execution is process-local; Redis generation streams add durable restart recovery and replica handoff. No librechat.yaml switch or environment feature flag is required. See Agent Event Runtime.
backgroundTasks
Controls whether supported completed background tools and detached Subagents automatically resume their saved parent Agent.
| Key | Type | Description | Example |
|---|---|---|---|
| completionWakeups | Boolean | Automatically deliver supported background-task completions as a continuation. Set to false for poll-only behavior. | true |
| ordinaryToolCancellation | Boolean | Allows owners to request cooperative cancellation of ordinary background tools, including attached Bash. | false |
| completionResultMaxChars | Integer | Maximum characters retained from each background completion result before delivery. | 24576 |
| completionResultBatchSize | Integer | Maximum completed background results delivered in one continuation batch. Valid values are 1 through 16. | 8 |
| shutdownInterruptGraceMs | Integer | Graceful-shutdown wait after interrupting a still-running background tool before recording an interrupted result. Valid values are 0 through 60000 ms. | 5000 |
backgroundTasks:
completionWakeups: true
ordinaryToolCancellation: false
completionResultMaxChars: 24576
completionResultBatchSize: 8
shutdownInterruptGraceMs: 5000Automatic completion delivery is enabled even when this block is omitted. Set completionWakeups: false to require the Agent to collect every result through check_background_task.
Set ordinaryToolCancellation: true to expose owner-scoped cancellation for ordinary background tools. Cancellation is cooperative, remains process-local, and keeps the task active and capacity reserved until the invocation settles. It does not cancel detached Subagents.
Ordinary background tool execution remains process-local. During a graceful shutdown, LibreChat waits within the shutdown budget for tasks to settle, then interrupts unfinished work and records a result; a sudden worker loss still cannot preserve the running task. Once a content-only terminal result is persisted, its delivery is durable and may continue on another replica. Tasks with live artifacts still require polling on the owning run. Detached Subagents keep their separate durable transcript and terminal-result path. Manual polling remains available for status, controls, and recovery in either mode. See Background Tool Calls and Automatic Parent Continuation.
toolApproval
Controls human review for Agent tool calls. The deployment-wide policy is disabled when this block is omitted. An Agent using an attached Code environment is the exception: LibreChat enables its scoped ask-by-default safety policy unless the administrator explicitly sets enabled: false. When review applies, matching calls pause until the user submits an allowed decision, such as approving, rejecting, or editing the arguments.
| Key | Type | Description | Example |
|---|---|---|---|
| enabled | Boolean | Enables deployment-wide tool approval. Explicit false also disables the attached Code environment safety baseline. | Omitted |
| mode | String | Sets unmatched-call behavior: `default` asks, `dontAsk` denies, and `bypass` approves unless denied. | default |
| allow | Array of strings | Glob patterns for calls that can run without prompting. | |
| deny | Array of strings | Glob patterns for calls that are always denied. Deny rules win. | |
| ask | Array of strings | Glob patterns for calls that always require review. | |
| reason | String | Optional explanation shown in the approval prompt. Use `{tool}` to insert the tool name. | |
| hooks | Array of objects | Trusted programmatic policy modules for context-aware decisions. Hooks can only tighten the static policy. |
Patterns use glob matching. MCP tools can be scoped with mcp:<server>:<tool>; for example, mcp:github:* matches every tool from the github server. LibreChat applies static rules and hook matchers to MCP runtime names and model-facing aliases, including tools reachable through nested Subagents. Static rules are evaluated in deny, ask, then allow order before the selected mode supplies the fallback. A deny rule therefore always wins, including in bypass mode.
toolApproval:
enabled: true
mode: default
allow:
- 'mcp:analytics:read_*'
deny:
- 'mcp:*:delete_*'
ask:
- 'mcp:*:*'
reason: 'Review {tool} before it runs.'Hooks are loaded once at server startup and execute in the LibreChat process. Each entry accepts a module, an optional tool-name matcher regex, and an optional options object passed to the module's builder. Only load code you trust. Per-role or tenant overrides do not reload hook modules; implement contextual behavior inside the hook.
toolApproval:
enabled: true
hooks:
- module: '@acme/librechat-approval-hook'
matcher: '^mcp:production:'
options:
requireTicket: truecheckpointer
Configures storage for Agent runs paused by tool approval or Ask User. When either feature needs a checkpointer and this block is omitted, LibreChat uses the primary MongoDB with a 24-hour decision window.
| Key | Type | Description | Example |
|---|---|---|---|
| type | String | `mongo` persists paused runs across restarts and replicas; `memory` is process-local and intended for development. | mongo |
| ttl | Number | How long a paused run waits for a decision, in seconds. | 86400 |
| checkpointCollectionName | String | Optional MongoDB checkpoint collection name override. | agent_checkpoints |
| checkpointWritesCollectionName | String | Optional MongoDB checkpoint-write collection name override. | agent_checkpoint_writes |
checkpointer:
type: mongo
ttl: 86400Use type: memory only for a single-process development deployment. A paused run stored in memory cannot resume after a restart or on another replica.
MongoDB checkpoints log a warning when serialized state exceeds 8 MiB and reject a durable pause above 15 MiB, leaving headroom below MongoDB's 16 MiB document limit. Large inlined media or tool outputs are common causes; start a new conversation or reduce its context if a run reports CHECKPOINT_TOO_LARGE.
skills
Controls endpoint-level Skills settings for agents.
| Key | Type | Description | Example |
|---|---|---|---|
| skills.maxCatalogSkills | Number | Caps the number of active accessible Skills exposed in the model-visible catalog. Must be between 1 and 100. | maxCatalogSkills: 20 |
Default: No configured cap beyond the runtime catalog limit.
Example:
skills:
maxCatalogSkills: 20This does not disable Skills. Use the skills capability and per-agent/model-spec skill scoping to control whether Skills are available.
maxSubagents
| Key | Type | Description | Example |
|---|---|---|---|
| maxSubagents | Number | Maximum number of explicit subagents a single agent may reference. | Applies to the `agent_ids` allowlist on agents and on model specs, and to configured subagent `graphs`, which agents support but model specs do not. Must be a whole number between 1 and 50. |
Default: 10
Range: 1-50
Example:
endpoints:
agents:
maxSubagents: 20Raise this when a deployment runs orchestration-heavy agents that need to delegate to more than ten children. The value is read at startup and enforced in three places:
- Agent create, update, and duplicate requests, for
subagents.agent_idsandsubagents.graphs - Model spec
subagents.agent_idsallowlists, validated in the same configuration pass - The Agent Builder panel, which stops accepting new subagent entries at the configured value
Model specs accept only enabled, allowSelf, and agent_ids under subagents. Subagent graphs are an agent-level feature: a graphs key on a model spec is dropped when the config is parsed, without an error.
maxSubagents must be a whole number between 1 and 50. A value outside that range fails
librechat.yaml validation, and LibreChat logs the error and exits at startup instead of falling
back to the default. Starting with CONFIG_BYPASS_VALIDATION=true skips the exit, but the entire
custom config is discarded in that case, so the cap returns to 10. To go back to the default
deliberately, remove the key: an omitted maxSubagents resolves to 10 with no error.
The cap is process-wide and comes from librechat.yaml. A per-user or per-group database override
of endpoints.agents.maxSubagents is reflected in the served configuration but is not applied by
request validation.
Raising this only changes how many subagents one agent may reference. The depth, graph node, and run configuration limits are unaffected. See Subagents.
fileSharing
fileSharing is opt-in, run-scoped sharing for delegated Agent work. Set enabled: true here and
enable subagents.shareFiles on each delegating Agent. Inputs are read-only to recipients; outputs
remain private unless explicitly published, and published files stay downloadable after the run.
endpoints:
agents:
fileSharing:
enabled: true
allowSiblingSharing: false
maxFiles: 100
maxPrivateBytes: 268435456
ttlMs: 3600000allowSiblingSharing permits explicitly named sibling or descendant recipients. maxFiles ranges
from 1 to 1000; maxPrivateBytes bounds retained private output snapshots; and ttlMs is the
active manifest lifetime, capped at 24 hours. Omitting this block leaves sharing disabled.
Code API upload recovery
These settings bound concurrent uploads and retry waiting when Code API responds with rate limiting. They apply per Code API route and authenticated principal, preventing one throttled user or deployment from blocking unrelated upload recovery.
| Key | Type | Description | Example |
|---|---|---|---|
| codeApiUploadConcurrency | Number | Maximum concurrent Code API uploads per route and authenticated principal (1-100). | 3 |
| codeApiMaxRetryWaitMs | Number | Maximum total wall-clock time one operation may spend waiting on Code API rate limits (0-300000 ms). | 20000 |
endpoints:
agents:
codeApiUploadConcurrency: 3
codeApiMaxRetryWaitMs: 20000LibreChat honors Retry-After, reopens persisted streams for each attempt, and removes cancelled callers from the queue. Setting codeApiMaxRetryWaitMs: 0 disables waiting rather than disabling uploads.
maxCitations
| Key | Type | Description | Example |
|---|---|---|---|
| maxCitations | Number | Controls the maximum total number of citations that can be included in a single agent response. | When using file_search capability, limits the total number of source citations returned to prevent overwhelming responses while ensuring comprehensive coverage. |
Default: 30
Range: 1-50
Example:
maxCitations: 30maxCitationsPerFile
| Key | Type | Description | Example |
|---|---|---|---|
| maxCitationsPerFile | Number | Limits the maximum number of citations that can be extracted from any single file. | Ensures citation diversity by preventing any single file from dominating the citations, encouraging representation from multiple sources. |
Default: 7
Range: 1-10
Example:
maxCitationsPerFile: 7minRelevanceScore
| Key | Type | Description | Example |
|---|---|---|---|
| minRelevanceScore | Number | Sets the minimum relevance score threshold for sources to be included in responses. | Filters out low-quality matches based on vector similarity scores. Higher values (e.g., 0.7) ensure only highly relevant sources are cited, while lower values (e.g., 0.0) include all sources regardless of quality. |
Default: 0.45 (45% relevance threshold)
Range: 0.0-1.0
Example:
minRelevanceScore: 0.45File Citation Configuration Examples
Default Configuration (Balanced)
endpoints:
agents:
maxCitations: 30
maxCitationsPerFile: 7
minRelevanceScore: 0.45Provides comprehensive citations while preventing overwhelming responses and filtering out low-quality matches.
Strict Configuration (High Quality)
endpoints:
agents:
maxCitations: 10
maxCitationsPerFile: 3
minRelevanceScore: 0.7Only includes highly relevant citations with strict limits for focused responses.
Comprehensive Configuration (Research)
endpoints:
agents:
maxCitations: 50
maxCitationsPerFile: 10
minRelevanceScore: 0.0Maximum information extraction for exhaustive research tasks, including all sources regardless of relevance.
Agent Capabilities
The capabilities field allows you to enable or disable specific functionalities for agents. The available capabilities are:
- deferred_tools: Allows agents to discover deferred MCP tools at runtime instead of loading every tool into context upfront.
- programmatic_tools: Enables Programmatic Tool Calling for MCP tools marked Programmatic in the Agent Builder. Requires
execute_codeand a Code Interpreter deployment with the Tool Call Server component. This capability is opt-in and is not enabled by default. - execute_code: Allows the agent to execute code.
- stateful_code_sessions: Lets opted-in agents reuse one Code Interpreter workspace at a user, agent-and-user, or conversation scope. Requires
execute_code, a compatible Code Interpreter deployment, and the per-agent Advanced setting. New Agents use the signed-in user's preferred scope. This capability is highly experimental, is not enabled by default, and may change substantially during experimentation; do not treat its current behavior or configuration as a stable production contract. - file_search: Enables the agent to search and interact with files. When enabled, citation behavior is controlled by
maxCitations,maxCitationsPerFile, andminRelevanceScoresettings. - web_search: Enables web search functionality for agents, allowing them to search and retrieve information from the internet.
- artifacts: Enables the agent to generate interactive artifacts (React components, HTML, Mermaid diagrams).
- subagents: Enables isolated-context child agent runs. See Subagents.
- actions: Permits the agent to perform predefined actions.
- context: Enables "Upload as Text" functionality in chat, and "File Context" for agents, allowing users to upload files and have their content extracted and included directly in the conversation.
- skills: Enables Skills in the side panel, manual
$invocation, model-invoked skills, and agent skill allowlists. See Skills. - memory: Lets agents use
set_memoryanddelete_memorywhen memory is configured and the user has access. Enabled by default; removed at runtime when memory is disabled. - ask_user_question: Lets agents pause to ask one to four related questions in a single form and resume with the answers. Enabled by default.
- tools: Grants the agent access to various tools.
- chain: Enables Beta feature for agent chaining, also known as Mixture-of-Agents (MoA) workflows.
- ocr: Optionally enhances "Upload as Text" in chat, and "File Context" for agents, allowing files to be uploaded and processed with OCR. Requires an OCR service to be configured.
- run_in_background: Makes Code Interpreter execution and shell tools background-eligible by default and enables per-tool background opt-in for eligible MCP, Plugin, and Action tools. An Action-level switch opts in every eligible operation; OAuth Actions and operations that already define a
run_in_backgroundparameter are excluded. An agent can explicitly opt Code Interpreter out. The setting permits background execution; the model decides per call whether to use it. Requiresexecute_codeand a configured Code Interpreter deployment for code. This capability is not enabled by default. - tool_intents: Adds live model-written intent labels to eligible tool calls. Native tools, including Code Interpreter tools, opt in automatically; MCP tools can be enabled individually or in bulk in the Agent Builder. Model specs use
describeIntent. This capability is not enabled by default.
By specifying the capabilities, you can control the features available to users when interacting with agents.
Example Configuration
Here is an example of configuring the agents endpoint with custom capabilities and file citation settings:
endpoints:
agents:
disableBuilder: false
# File citation configuration
maxCitations: 20
maxCitationsPerFile: 5
minRelevanceScore: 0.6
# Custom capabilities
capabilities:
# Optional: enables Programmatic Tool Calling for MCP tools marked Programmatic in the Agent Builder.
# - 'programmatic_tools'
- 'execute_code'
# Optional: makes Code Interpreter background-eligible and enables selected MCP, Plugin, and Action tools.
# - 'run_in_background'
# Optional: enables live intent labels for native and selected MCP tools.
# - 'tool_intents'
- 'file_search'
- 'skills'
- 'subagents'
- 'actions'
- 'artifacts'
- 'context'
- 'ocr'
- 'web_search'In this example:
- The builder interface is enabled
- File citations are limited to 20 total, with maximum 5 per file
- Only sources with 60%+ relevance are included
- LibreChat Agents have access to code execution, file search (with citations), Skills, Subagents, actions, artifacts, file context, ocr services if configured, and web search capabilities
- Programmatic Tool Calling remains disabled unless you add the
programmatic_toolscapability alongsideexecute_code - Background tool calls remain disabled unless you add
run_in_background; Code Interpreter tools then become eligible by default, while eligible MCP, Plugin, and Action tools still require selection - Tool intent labels remain disabled unless you add
tool_intents; native tools then opt in automatically, while MCP tools still require per-tool selection
managementApi
The beta Agent Management API lets trusted machine clients create, list, read, update, and delete Agents, manage their files, and discover or edit accessible Skills. It is separate from remoteApi: management routes accept only deployment-bound OIDC machine identities and do not accept LibreChat Remote Agents API keys or browser sessions.
endpoints:
agents:
managementApi:
auth:
oidc:
enabled: true
issuer: https://identity.example.com/
audience: https://librechat.example.com/agents
# jwksUri: https://identity.example.com/.well-known/jwks.json
clients:
- clientId: machine-client-id
# subject: provider-specific-service-principal-subject
userId: 507f1f77bcf86cd799439011
tenantId: tenant-id
enabled: truemanagementApi.auth.oidc
| Key | Type | Description | Example |
|---|---|---|---|
| enabled | Boolean | Enable OIDC machine authentication for Agent Management routes. | false |
| issuer | String | Exact issuer expected in verified access tokens. Required when enabled. | |
| audience | String | Optional audience the access token must contain. Omit for access-token providers that do not issue an `aud` claim. | |
| tokenUse | String | Optional required `token_use` claim, such as `access` for Amazon Cognito access tokens. | |
| requiredScopes | String[] | Optional scopes that must all appear in the token `scope` claim. | |
| jwksUri | String | Optional explicit JWKS endpoint. When omitted, LibreChat discovers it from the issuer. |
For Amazon Cognito machine-to-machine access tokens, omit audience, set tokenUse: access, and list the resource-server scopes required for Agent Management under requiredScopes. Cognito identifies the client with client_id, which LibreChat matches against the configured client allowlist.
managementApi.auth.clients
clients is a deployment-owned allowlist with at most 100 entries. At least one entry is required when OIDC authentication is enabled.
| Key | Type | Description | Example |
|---|---|---|---|
| clientId | String | OAuth client identifier matched against the token `azp` or `client_id` claim. | Required |
| subject | String | Optional exact token `sub` claim. When omitted, `sub` must equal `clientId` or `clientId@clients`. | |
| userId | String | Existing LibreChat MongoDB user ObjectId whose permissions the client inherits. | Required |
| tenantId | String | Exact tenant containing that user. The system tenant is not allowed. | Required |
| enabled | Boolean | Enable this client binding without removing it from configuration. | true |
Tokens must be unexpired and present a single, non-conflicting client identifier. After verification, LibreChat resolves the configured user only inside the bound tenant and rejects missing, inactive, cross-tenant, disabled, or subject-mismatched principals. Every management operation then uses that user's current role capabilities and Agent or Skill ACLs; the binding grants no additional resource access.
Unknown management-auth fields, duplicate client IDs, malformed user or tenant IDs, and enabled authentication without a client binding fail configuration validation. See Agents API - Agent Management for routes and examples.
remoteApi
Configuration for Remote Agent API authentication. Controls how external services authenticate when calling the Agents API endpoints.
remoteApi.auth
| Key | Type | Description | Example |
|---|---|---|---|
| auth | Object | Authentication configuration for the Remote Agent API. | Supports API key and/or OIDC Bearer token authentication. If omitted, only API key auth is active. |
remoteApi.auth.apiKey
| Key | Type | Description | Example |
|---|---|---|---|
| enabled | Boolean | Enable API key authentication for the Remote Agent API. | When true, requests with a valid LibreChat API key are accepted. Can be used alongside or instead of OIDC. |
Default: true
remoteApi.auth.oidc
| Key | Type | Description | Example |
|---|---|---|---|
| enabled | Boolean | Enable OIDC Bearer token authentication. | When true, the middleware validates Bearer tokens against the configured OIDC issuer via JWKS. |
| issuer | String | OIDC issuer URL. | The base URL of your OIDC provider, such as a Keycloak realm URL. Used for token issuer validation and JWKS discovery if jwksUri is not set. |
| jwksUri | String | JWKS endpoint URL. Optional. | If omitted, resolved automatically via {issuer}/.well-known/openid-configuration. You can also set OPENID_JWKS_URL as an alternative. |
| audience | String | Expected token audience. Required when OIDC auth is enabled. | Tokens must contain this value in their aud claim. |
| scope | String | Required scope value. Optional. | If set, the token must contain this value in its scp or scope claim. Use this to distinguish token intent across different APIs. |
Default: enabled: false
Example - OIDC only:
endpoints:
agents:
remoteApi:
auth:
apiKey:
enabled: false
oidc:
enabled: true
issuer: https://auth.example.com/realms/myrealm
audience: my-client-idExample - OIDC with API key fallback:
endpoints:
agents:
remoteApi:
auth:
apiKey:
enabled: true
oidc:
enabled: true
issuer: https://auth.example.com/realms/myrealm
# jwksUri is optional and auto-discovered if omitted
jwksUri: https://auth.example.com/realms/myrealm/protocol/openid-connect/certs
audience: my-client-idJWKS URI resolution priority is explicit jwksUri, then OPENID_JWKS_URL, then
auto-discovery via {issuer}/.well-known/openid-configuration.
OIDC user matching uses the sub claim as primary lookup, with fallback to email,
preferred_username, or upn claims. The matched user must already exist in LibreChat.
Subagents
The subagents field controls which isolated child agents a parent agent can spawn when the subagents capability is available.
| Key | Type | Description | Example |
|---|---|---|---|
| enabled | Boolean | Adds the subagent spawn tool to this agent when true. Default: disabled. | enabled: true |
| allowSelf | Boolean | Allows the agent to spawn itself in a fresh isolated context. Default: true. | allowSelf: true |
| agent_ids | Array/List of Strings | Specific agents this agent may spawn. Capped by `maxSubagents`, which defaults to 10. | agent_ids: ["agent_researcher"] |
| graphs | Array of Objects | Saved Agent teams that can run as one isolated child graph. The list is capped by `maxSubagents` (default: 10); each team can contain up to 32 members. |
subagents:
enabled: true
allowSelf: true
agent_ids:
- 'agent_researcher'
- 'agent_reviewer'Each graphs entry requires a unique type, name, description, agent_ids, direct edges, entry_agent_id, and result_agent_id. Team definitions must form a connected directed acyclic graph, every member must be visible to the user, and the total configuration is bounded to 50 unique Agent targets and 100 expanded run configurations. See the saved-team example and runtime behavior.
For user-facing behavior and limits, see Subagents.
Notes
- It's not recommended to disable the builder interface unless you are using modelSpecs to define a list of agents to choose from.
- File citation configuration (
maxCitations,maxCitationsPerFile,minRelevanceScore) only applies when thefile_searchcapability is enabled. - The relevance score is calculated using vector similarity, where 1.0 represents a perfect match and 0.0 represents no similarity.
- Citation limits help balance comprehensive information retrieval with response quality and performance.
- The
contextcapability works without OCR configuration using text parsing methods. OCR enhances extraction quality when configured. - The
ocrcapability requires an OCR service to be configured (see OCR Configuration). maxSubagentsbounds how many subagents a single agent may reference. It does not change the depth or graph limits documented under Subagents.
How is this guide?