Skip to main content
LibreChat is joining ClickHouse to power the open-source Agentic Data Stack 🎉 Learn more
← Back to changelog

⚙️ Config v1.3.17

v1.3.17
  • Added public Agents API documentation routes

    • openapi.enabled serves the OpenAPI specification at /api/openapi.json and Swagger UI at /api/docs
    • The routes are disabled by default
  • Expanded Agent Management OIDC authentication

    • endpoints.agents.managementApi.auth.oidc.audience is now optional
    • tokenUse can require a token type such as access
    • requiredScopes requires every listed OAuth scope, enabling Cognito machine access tokens that use client_id without an aud claim
  • Added ACL write conflict limits

    • permissions.maxWriteAttempts controls compare-and-swap attempts per ACL document, including the initial write
    • The default is 3; valid values range from 1 through 100
  • Added an xAI Grok 4.7 custom endpoint example

    • Set XAI_API_KEY and use the OpenAI-compatible https://api.x.ai/v1 endpoint
    • Grok 4.7 supports low, medium, high, and xhigh reasoning effort
  • Expanded highly experimental stateful Code session controls

    • endpoints.agents.repositoryInstructions.timeoutMs bounds attached-workspace instruction discovery and defaults to 2000 ms
    • statefulCodeSessions.environments[].configSchema.limits.maxQueueWaitMs bounds attached Bash capacity retries and defaults to 300000 ms
    • Setting maxQueueWaitMs to 0 disables client retries after initial admission without cancelling admitted or running work
  • Added Agent error and background-result limits

    • endpoints.agents.maxProviderErrorChars defaults to 2000; set it to 0 to omit unclassified provider error text
    • endpoints.agents.modelResponseBodyTimeoutMs defaults to 900000 ms and resets whenever a provider response-body chunk arrives
    • endpoints.agents.modelResponseHeadersTimeoutMs defaults to 300000 ms while waiting for provider response headers
    • Set either transport timeout to 0 to disable it; both accept values up to 86400000 ms
    • backgroundTasks.completionResultMaxChars defaults to 24576 characters per delivered result
    • backgroundTasks.completionResultBatchSize defaults to 8 results per continuation and accepts values from 1 through 16
  • Expanded MCP request and OAuth configuration

    • mcpServers.<name>.requestHeaders resolves headers only for an active request, merges over headers, and is omitted during tool discovery
    • oauthRefreshCoordination optionally serializes token refresh across replicas through shared Redis
    • oauthRefreshWaitTimeout and oauthPersistenceWaitTimeout both default to 15000 ms
    • Enable refresh coordination only after every replica is upgraded and uses the same Redis deployment and values
  • Added authenticated RUM proxy exports

    • RUM_PROXY_AUTHORIZATION supplies the complete server-only Authorization header forwarded to the configured collector
    • Authenticated redirects are rejected; configure the final RUM_PROXY_TARGET_URL
  • Updated Redis stream delta coalescing

    • STREAM_DELTA_COALESCE_MS now defaults to 25 ms when unset
    • Set it explicitly to 0 to disable coalescing; values above 1000 are capped
  • Expanded Conversation Trace Viewer configuration

    • interface.traceViewer.showToolNames defaults to false
    • Enabling it performs additional Langfuse observation reads and exposes tool names without exposing tool input or output
  • Added streaming code-highlight throttling

    • interface.codeHighlightThrottleMs controls the minimum interval between syntax-highlighting passes while streamed code changes
    • The default is 300 ms; set it to 0 to disable throttling
  • Added Skill import cleanup concurrency

    • fileConfig.skills.importCleanupConcurrency bounds cleanup after a failed archive import and defaults to 8
    • Set it to 1 for serial cleanup
  • Updated file MIME pattern validation

    • Every supportedMimeTypes pattern is compiled with the active server regex engine during configuration validation
    • Unsupported patterns now reject the configuration instead of being silently skipped
  • Updated Helm credential Secret handling

    • A non-empty global.librechat.existingSecretName is now required to resolve; missing named Secrets block pod startup
    • Set the value to an empty string only when all LibreChat credentials are injected through other supported environment settings
    • Bundled Meilisearch still requires its separately configured master-key Secret
  • Added optional saved-chat workspace transitions after rc4

    • endpoints.agents.statefulCodeSessions.conversationMoves.enabled opts in to moving a conversation's sealed workspace and recovering a missing workspace
    • conversationMoves.allowAttachDetach separately opts in to attaching or leaving a workspace; omission keeps the move-only policy
    • Upgrade every API replica and enable CODE_ENVIRONMENT_DECISION_VERSION=1 before activating the transitions
  • Added optional attached-workspace request budgets

    • statefulCodeSessions.environments[].configSchema.limits.maxRequestTimeoutMs caps total HTTP time for one call at up to 610000 ms
    • The existing 30-second admission budget per attempt remains when this is unset, and maxQueueWaitMs remains an independent retry horizon
    • Do not enable longer admission until the Code API honors per-request queue allowances and the shortest live ingress timeout has been measured
    • configSchema.limits.minCommandAdmissionMs reserves 10000 ms by default before command execution (valid 1000–300000 ms); the total request budget must exceed this reserve plus 10000 ms of settlement/delivery grace
    • Configured and advertised Bash command limits fit within the remaining HTTP budget; a 90000 ms request with an 80000 ms command ceiling and default reserve advertises at most a 70000 ms command
  • Added opt-in linked-worktree scheduling for attached Code environments

    • statefulCodeSessions.environments[].configSchema.workspaces.linkedWorktrees defaults to off; enable only after upgrading LibreChat, deploying Code API linked-worktree support, and starting compatible workers with --linked-worktree-lanes
    • File tools with a worktree path and Bash commands with a worktree cwd use that worker lane; commands that cd internally stay in the checkout lane
  • Added idle Agent recovery and graceful-shutdown controls

    • endpoints.agents.eventDriven.idlePolling.deliveryMaxIntervalMs defaults to 15000 ms
    • queuedTurnMaxIntervalMs and maintenanceMaxIntervalMs default to 120000 ms; completionWaitMaxIntervalMs defaults to 60000 ms
    • Local readiness events can expedite delivery without waiting for an idle MongoDB recovery scan
    • endpoints.agents.backgroundTasks.shutdownInterruptGraceMs defaults to 5000 ms (valid 0 through 60000) before unfinished tools are recorded as interrupted during graceful shutdown
  • Added interface.agentSelectorLimit to cap the unsearched Agent selector at 10 entries by default (valid 1 through 100); searching still reaches all Agents

  • Updated configuration validation and reload behavior

    • Malformed explicitly configured CREDS_KEY or CREDS_IV values now fail startup; migrate data encrypted with an invalid-length key deliberately rather than rotating it blindly
    • rateLimits.conversationsImport, rateLimits.tts, rateLimits.stt, and skill/Agent upload limits from librechat.yaml now apply to their route limiters without matching environment variables
    • Invalid role, group, or user config override writes now return 400 without saving; previously stored overrides that would invalidate the merged config are dropped with a warning
    • Reload failures keep the last good validated configuration active; startup validation still fails closed
  • Hardened MCP OAuth HTTP requests

    • Server-side OAuth discovery, registration, token, and revocation calls reject redirects; configure the final URL when an identity provider redirects
    • Private IP literals are refused unless permitted by the existing admin-trusted host or exact address-and-port policy
  • Added capability-negotiated workspace edit matching

    • statefulCodeSessions.environments[].configSchema.edits.tolerantMatching allows whitespace-tolerant fallback by default when the worker advertises tolerant_match; set it to false to require exact matches
    • Attached edit_file and protected edit previews only send new matching fields after negotiation; replace_all also requires the worker to advertise that feature
    • Requires compatible Code API and worker builds implementing Code Interpreter #271; older deployments keep their current request format and matching behavior
    • Edit conflicts use bounded, host-rendered facts rather than untrusted worker text; a rejected batch writes nothing
  • Added GPT point-release family fallback and GPT-6.1 Sol

    • A GPT point release without its own entry inherits the recognized family's context/output limits, pricing, and reasoning settings; explicit per-release entries win
    • GPT-6.1 Sol joins the OpenAI and Agents catalogs, not Assistants; its limits and prices currently inherit GPT-6 Sol rather than new provider-published figures
    • Native OpenAI Responses routing can inherit the family policy; Azure deployment routing and custom-endpoint opt-in behavior stay unchanged
  • Hardened web-search credential pairing

    • User-provided Firecrawl or SearXNG URLs cannot be combined with the server's API key, even when the destination passes SSRF checks
    • Firecrawl ignores the optional user URL and uses its default endpoint; SearXNG cannot initialize with that mixed-ownership pair because its URL is required
    • User-owned pairs, administrator-owned pairs, and keyless SearXNG remain supported; set both URL and key at the same ownership level
  • Added tenant-scoped YAML custom endpoints

    • endpoints.custom[].tenantId optionally restricts an endpoint and its associated model specs to the authoritative request tenant; omission preserves deployment-wide availability
    • Tenant IDs accept 1–128 letters, digits, hyphens, underscores, or periods; whitespace and the reserved __SYSTEM__ sentinel are invalid
    • Missing tenant context does not receive scoped endpoints; cache, override, runtime augmentation, and fallback paths keep the same boundary
    • Filtering the last model spec restores normal model controls while preserving explicit interface settings
    • Upgrade every API replica before configuring scoped endpoints; older replicas do not enforce the new YAML field
  • Bound per-user MCP API keys to their destinations

    • UI-created/shared servers require each user to enter a key again when the URL path/query, proxy, transport, auth format/header, or configured OAuth client/destination changes
    • Equivalent and cosmetic updates preserve existing keys; legacy keys remain usable at their unchanged boundary without a bulk migration
    • Shared servers resolve only their currently declared credential variables, so old generated names cannot bypass re-entry
    • Upgrade every API replica, including configuration writers, before relying on the new binding
  • Tightened native edit approval and import behavior

    • Empty or malformed edit_file.edits batches now fail closed instead of falling back to top-level replacements
    • Native edit arguments, including valid stringified JSON and hook rewrites, are normalized before approval so the preview matches execution
    • Imports and other bulk conversation writes cannot copy a Code approval mode; new records use the recipient's preference/default and updates preserve the locally stored choice
    • Existing saved modes are not retroactively changed; ordinary user-selected saves remain supported
  • Updated the stable release deployment snapshot

    • The application version is v0.8.8 and the Helm chart source is 2.0.16 with appVersion: v0.8.8; the RAG chart and configuration schema version are unchanged
    • When the chart's librechat.configEnv would otherwise be null, set it to {} to avoid the known template-rendering failure; credentials still belong in the named Secret
  • Added an always-active authenticated 2FA management budget

    • rateLimits.twoFactorManagement.requestsPerFiveMinutes defaults to 7 requests per fixed five-minute window; valid values are integers of at least 3 (the enable → verify → confirm sequence)
    • Successful and failed requests share one tenant/account budget across enable, verify, confirm, disable, and backup-code regeneration; all sessions count toward it
    • Exhaustion returns HTTP 429, Retry-After, and TWO_FACTOR_RATE_LIMITED before verification or mutation; the settings UI explains when to retry
    • Set the limit in deployment YAML; role/group/user configuration overrides do not control this bucket. It is active when omitted and cannot be disabled with 0
    • Redis makes the limit shared across API replicas; without it each process uses its own memory bucket. The temporary-token 2FA login limiter is unchanged
  • Updated the config version to 1.3.17