Commit graph

1399 commits

Author SHA1 Message Date
Danny Avila
2d227dbc3b 🛡️ feat: Configurable Baseline HTTP Security Headers
Adds helmet's CSP-independent headers (HSTS, X-Frame-Options,
X-Content-Type-Options, COOP, CORP, Referrer-Policy) on every response,
with contentSecurityPolicy explicitly disabled. Every header that can
break a deployment is configurable, so there is no allow-list to go
stale the way #7377's hardcoded CSP directives did.

HSTS includeSubDomains defaults off rather than matching helmet's
on-by-default: it would otherwise pin every sibling subdomain to HTTPS
for a year in every visitor's browser, and undoing that requires
serving max-age=0 from each affected host.
2026-07-28 09:37:07 -04:00
Danny Avila
324584552c
⏱️ feat: Configurable HTTP Server Timeouts (#14481)
* http server config added

* Fix TypeScript compatibility by accepting NodeJS.ProcessEnv directly when applying optional HTTP server timeout configuration.

* fix(api): configure HTTP server timeouts for clustered workers

* 🕰️ fix: Warn When HTTP Timeouts Are Not Enforced

Codex review of the rebased contributor work surfaced two ways these settings
silently do nothing. Both reproduce, and neither was reported to the operator.

Bun accepts the four property assignments and reflects them back, but does not
enforce them: with keepAliveTimeout=100 and buffer=1000, Bun 1.3.13 held a
keep-alive connection past 3s where Node 24 closed it at 1101ms. Since `b:api`
runs the server under Bun, the existing info log confirmed a configuration that
was not in effect. Warn instead.

Node sweeps header/request timeouts on `connectionsCheckingInterval`, a
createServer option that `app.listen()` leaves at 30s, so sub-30s values round
up to it: headersTimeout=2000 returned 408 at 30004ms by default versus 2010ms
with a 250ms interval. Warn on values below the sweep interval rather than
restructure server construction, since every documented value and both Node
defaults already sit well above it. keepAliveTimeout is socket-driven and stays
exact, so it is excluded.

Both caveats documented in .env.example.

* 🩹 fix: Inject Runtime Versions Instead of Mutating `process.versions`

The spec deleted `process.versions.bun` to reset between cases, which failed
typecheck with TS2790: `@types/bun` is a packages/api dependency and augments
NodeJS.ProcessVersions with a required `bun: string`, so the property is not
optional and cannot be deleted. Assigning undefined would fail for the same
reason.

That augmentation also made the production check dishonest: TypeScript saw
`process.versions.bun` as always a string, so `!= null` read as a no-op branch
even though it is correct at runtime under Node.

Both resolved by taking runtime versions as a third injectable parameter,
matching the existing `environment` parameter. Callers in api/server are
unchanged, the narrow `{ bun?: string }` type restores honest narrowing, and
the tests no longer mutate global state, so they assert the same behavior
whether the suite runs under Node or `bun jest`.

* 📏 fix: Stop Claiming a Ceiling on Sweep-Delayed Timeouts

The warning added in 9adc3eb1c said the effective timeout is "up to 30000ms",
which promises a bound that does not hold. Node detects header/request expiry
only on the next connection sweep, so the delay is relative to the deadline
rather than capped by the interval: measured against the default 30s sweep, a
2000ms headersTimeout closed at 30004ms, 15x the configured value, and cases
where the timeout is near the interval did not fire within a 9s window at all.

Reworded to state the mechanism without asserting a ceiling, and to point
operators at values of 30000ms or above for predictable enforcement. Same
correction applied to the .env.example note, which claimed short values "round
up" to the interval.

*  fix: Clamp Headers Timeout to the Request Timeout

Setting only HTTP_REQUEST_TIMEOUT_MS below the 60s headers default left
headersTimeout > requestTimeout, a pairing createServer rejects outright with
ERR_OUT_OF_RANGE. Assigning the properties after construction skips that
validation, and the mismatch silently defeats the request timeout for a stalled
body: with requestTimeout=4000 and headersTimeout at its 60000 default, a client
that completed its headers and then stopped mid-body was still connected after
12s. Clamping headersTimeout to 4000 closes the same connection at 4017ms.

An earlier round dismissed this after testing partial *headers*, where
requestTimeout does evict on time. The gap only appears once headers are
complete and the body stalls, which is the case these timeouts exist to bound.

Mirrors Node's own rule rather than its constructor default: zero on either side
means disabled and is left alone, and an explicitly configured
HTTP_HEADERS_TIMEOUT_MS that conflicts is warned about before being clamped
instead of failing startup over a config typo.

---------

Co-authored-by: Peter Rothlaender <peter.rothlaender@ginkgo.com>
2026-07-28 09:10:17 -04:00
Dustin Healy
044c134ecf
🤏 fix: Filter Admin Config Reads by Section-Scoped Read Capability (#14472)
listConfigs, getBaseConfig, and getConfig only checked the broad
read:configs capability, so a caller holding nothing but
read:configs:<section> grants got a blanket 403 on all three instead
of a response filtered to the sections they hold. Any deployment
using section-scoped config grants hits this.

Adds hasAnyConfigReadAccess as a cheap pre-flight check covering
broad and section-scoped read and manage grants (manage implies
read), so a zero-access caller still 403s before a DB fetch while a
section-scoped caller gets the response filtered to exactly what
they hold. The same manage-implies-read rule is fixed at its root in
getParentCapabilities so a manage-only caller sees the section they
manage instead of having it stripped after passing the pre-flight.

Resolves every section for a request in one batched
getHeldCapabilities query via getReadableConfigSections instead of
one round trip per section.

Includes AppConfig field-renaming normalization (interfaceConfig,
turnstileConfig, mcpConfig) so the filter checks the canonical
section name rather than the renamed response field, and stops
availableTools from bypassing the filter by gating it on its
filteredTools/includedTools source sections.
2026-07-28 07:38:37 -04:00
Danny Avila
728fc1276e
🔒 fix: Bound /files/usage TTL Hold Instead of Clearing It (#14470)
* 🔒 fix: Bound `/files/usage` TTL Hold Instead of Clearing It

`POST /files/usage` marks queued attachments so the 1-hour upload-window
TTL cannot reap them before the client queue drains. It did this by
calling `updateFilesUsage`, which unsets `expiresAt` outright, turning
every touched upload into a permanently retained file.

The client queue is ephemeral browser state, so this also leaks in normal
use: a closed tab or cleared queue leaves nothing referencing the files,
but their TTL is already gone. The same mechanism let an authenticated
user pin arbitrary owned uploads indefinitely, and the route was excluded
from the file limiters, so the touch was entirely unmetered.

Make the operation match its intent, a renewable hold rather than a
release:

- Add `extendFilesTTL`, which pushes `expiresAt` forward by a bounded
  window in a single owner-scoped `updateMany`. Two filter guards keep it
  safe under client-supplied ids: `$exists: true` so an already-released
  file never has a TTL re-added (that would schedule a live file for
  deletion), and `$lt` so a hold only ever moves the deadline later.
  The owner scope is a required argument, so an unscoped call is a no-op
  rather than a cross-user update.
- `handleFilesUsageRequest` now holds for 24h instead of clearing, and no
  longer increments `usage`, since a queue touch is not a send. The real
  release still happens at drain, where `updateFilesUsage` marks the
  files used against an actual message.
- Give `/usage` its own per-user limiter. Keeping it off the upload quota
  was intentional, leaving it unmetered was not.

Abandoned queues are now reaped on schedule, and a replayed touch can only
ever re-assert the same bounded window.

* 🔒 fix: Anchor the `/files/usage` hold to upload time

Codex review on b687922.

The hold derived each new deadline from `Date.now()`, so a caller touching
once a day advanced it by another 24h every time, far below the rate limit.
That left indefinite preservation reachable and made the PR's replay claim
wrong: the window was bounded per call but not in aggregate.

Anchor the deadline to the file's immutable `createdAt` instead of the
request clock. `extendFilesTTL` now takes a lifetime and sets
`expiresAt = max(expiresAt, createdAt + holdMs)` in an aggregation
pipeline, so the target is a fixed point per file and replay is inert
rather than merely bounded. `$max` keeps the widen-only property and the
`expiresAt: {$exists: true}` filter still refuses to resurrect a released
TTL; `createdAt: {$exists: true}` fail-closes when the anchor is absent.

The update runs with `timestamps: false`: a hold is TTL bookkeeping, not a
content write, and bumping `updatedAt` also made every re-touch count as a
modification, hiding whether the deadline actually moved.

Also drop four `.node_modules-*` symlinks that `git add -A` swept in from
an npm install. They pointed at absolute paths on one machine, so every
other checkout got dangling entries. Added the pattern to .gitignore so a
workspace install cannot reintroduce them.

* 🔒 fix: Track the configured approval window in the `/files/usage` hold

Codex review on 9277620.

`endpoints.agents.checkpointer.ttl` is a positive int with no upper bound,
and its docs invite raising it for longer review windows. It drives the
pending-action expiry, so a run can legitimately stay paused past 24h. The
fixed 24h lifetime would then let Mongo reap an attachment while its
approval was still live, and the later queue drain would send a file that
no longer exists.

Replace the fixed constant with `resolveFilesUsageHoldMs`, which adds the
configured approval window to a 24h baseline covering upload, enqueue, and
the run reaching its pause. The route reads the window from the same
`getApprovalTtlMs(checkpointerCfg)` the pending action uses, so the two
stay in lockstep.

The replay bound is unaffected: the window is a per-deployment constant and
the deadline is still `createdAt + holdMs`, so a replayed touch re-asserts
the same instant and `$max` skips the write. Only an operator config change
moves it, never a client.

* 🔒 fix: Renew the `/files/usage` hold across queued runs, under a ceiling

Codex review on 2bd3c52.

The drain sends one queued item per run completion, and each item starts a
run that may itself pause for the full approval window. Since the hold was
taken once at enqueue and pinned to the upload time, an item several places
back could sit through multiple approval windows and lose its attachment
while its chip and the live approval were still there. Another regression
from this PR: the old `$unset` made retention permanent, so deep queues
happened to work.

The queue is unbounded, so no fixed lifetime covers it. Split the hold into
a renewable window and a ceiling:

  expiresAt = max(expiresAt, min(now + renewMs, createdAt + maxLifetimeMs))

`renewMs` covers one run's wait and is granted from now, so a queue that is
still draining re-asserts it at each transition; `useQueueDrain` now marks
the remaining items' files whenever it pops one. `maxLifetimeMs` is
measured from the immutable upload time and clamps every renewal, so
repeated touches converge on a ceiling instead of advancing per call, which
keeps the replay bound from the previous round intact.

This also tightens abandonment: a queue nobody drains now lapses one
`renewMs` after its last touch instead of surviving to the ceiling.

`useQueueDrain`'s spec gained a QueryClientProvider, since the renewal goes
through react-query.

* 🔒 fix: Renew queued holds on a heartbeat, and stop dropping batches

Codex review on f616bed.

Three gaps in the renewal added last commit:

- `collectQueuedFileIds` returned early at the server's 10-id cap, so a
  remainder holding more than one batch renewed only its first message and
  left the rest on their enqueue-time hold. Collect everything and split
  into capped requests instead of truncating.
- A refused `ask()` restores the popped item, but renewal ran before the
  send and covered only the pre-existing remainder. Since the run-end signal
  is already consumed, nothing would touch that item again. Renewal now runs
  after `ask` and includes the restored item.
- A single run can interrupt for approval more than once, each pause running
  to the configured window, so renewing only at drain transitions leaves a
  gap longer than `renewMs` with no renewal in it. The ceiling cannot help
  when nothing renews.

The third is the same structural gap as the previous round along a new axis:
renewal tied to discrete events loses the file whenever two events are
further apart than the hold. Rather than hook each transition, renew on a
30 minute heartbeat while anything is queued, which is far below the
smallest hold (24h) and so covers any single gap regardless of cause.

Still bounded: every renewal is clamped against the file's upload time, so
the ceiling is unchanged. A queue nobody has open emits no heartbeat and
lapses one `renewMs` after its last touch, preserving the abandonment
behaviour.

* 🔒 fix: Cover the pre-migration queue, first tick, and `/usage/`

Codex review on 892a27d.

- The heartbeat watched only the active conversation id, but `drainNext`
  merges in the `NEW_CONVO` queue, which outlives the URL update: items
  queued during the first turn stay keyed there until that run ends. It now
  renews the union of both, deduped since they are the same atom before
  migration.
- The interval installed without firing, so returning to a conversation
  whose hold was nearly up waited out a full period before the first
  renewal. It now renews immediately, then on each tick.
- Express's non-strict routing sends `POST /files/usage/` to the same
  handler with `req.path === '/usage/'`, so the exact comparison pushed it
  onto both upload limiters. A trailing-slash client would have spent its
  upload quota, and collected file-upload violations, on metadata
  heartbeats. Matching now tolerates the trailing slash.

Firing on effect start also made the drain-time renewal redundant: popping
an item changes the held set, so the renewal effect re-runs on its own. The
one case it cannot see is a refused send, where restoring the item leaves
the set identical, so that branch keeps an explicit renewal and the rest is
removed. Net one request per transition instead of two.
2026-07-28 07:37:26 -04:00
Danny Avila
250aca375a
🔗 fix: Resolve MCP Tool-Key Boundary Against Configured Server Names (#14448)
* fix: resolve MCP tool-name delimiter collision at invocation time

MCP tool keys are identified internally as `${rawToolName}${mcp_delimiter}${serverName}`
(delimiter `_mcp_`). Several call sites parsed this back apart with a naive
`toolKey.split(Constants.mcp_delimiter)`, assuming the delimiter occurs exactly once.

When the raw upstream tool name itself contains the delimiter substring - which
happens whenever it's exposed through a gateway that prefixes aggregated tool names by
server (e.g. a gateway's own "gitlab-get_mcp_server_version" for GitLab's
"get_mcp_server_version" tool) - the combined key has the delimiter more than once.
`.split()` then produces more than two segments, and destructuring
`[toolName, serverName]` silently keeps only the first two, yielding a bogus server
name that matches no configured server. Tool listing still worked (a different code
path builds keys directly without re-splitting), but invocation failed with
`Tool {name} not found`, and `filterAuthorizedTools` rejected such keys outright as
malformed.

Add `splitMCPToolKey`, which splits on the *last* occurrence of the delimiter instead:
the server-name half is always LibreChat's own normalized suffix (guaranteed not to
contain the delimiter), while the raw tool-name half is untrusted and may legitimately
contain it. This matches `.split()`'s result whenever the delimiter occurs once, and
correctly resolves the collision case. Update the four call sites that parsed this
manually (`handleTools.js`, `MCP.js`, `mcp.js` controller, `filterAuthorizedTools` in
`v1.js`) plus one in the client (`useVisibleTools.ts`) to use it.

Fixes #14440

* fix: resolve MCP tool-key boundary against configured server names

splitMCPToolKey moves to librechat-data-provider so the client and backend
share one parser, and takes the configured server names when the caller has
them: the longest name the key actually ends with wins, which is exact.

Position alone cannot identify the boundary because both halves may contain
the delimiter. lastIndexOf alone fixes gateway-prefixed tool names but
regresses servers whose own name contains it, which ToolService.spec.js
already covered; the last-delimiter path now only serves as the fallback for
callers with no configured set.

Also converts the remaining first-occurrence parsers that the delimiter fix
missed - mcp/auth.ts (custom user vars silently unresolved), mcp/oauth/events.ts,
agents/initialize.ts, and the three client parsers that labelled tool calls
with the wrong server.

* fix: keep client tool-call labels on first-delimiter parsing

The three client parsers had deliberate, tested first-delimiter semantics
(ToolCall.test.tsx asserts the full server name for 'foo_mcp_bar' and the
synthetic 'oauth_mcp_server' call), and the client has no configured server
list in scope to resolve the boundary exactly, so they are left as they were.

Threads the configured names into the event-driven definition loader so it
resolves the same boundary as the authorization filter that admits the key,
and documents the one case that stays undecidable without provenance.

* fix: resolve tool-key boundary against all configured servers

resolveConfigServers only returns lazily-initialized config overrides -
ensureConfigServers skips unmodified YAML servers - so on a stock deployment
the known-name list was empty and suffix resolution never engaged. Adds
resolveMcpServerNames, which keeps every configured server in the normalized
form tool keys carry, and uses it at the loading, auth-map and definition
sites.

Background-tool eligibility now resolves against all configured names before
testing ephemeral membership, so a non-ephemeral server whose name ends in an
ephemeral one is no longer misclassified, and useVisibleTools resolves against
the server map it already receives.

* fix: use resolved server provenance and one app-config read

createMCPTool now uses the serverName loadTools already resolved for the key
and only parses as a fallback, so an unmodified YAML server whose name
contains the delimiter no longer resolves to the wrong server for auth,
reconnection and callTool.

resolveMcpServerContext derives config servers and all configured names from
a single getAppConfigForRequest, replacing two independent lookups on the
chat startup path, and degrades to empty like resolveConfigServers instead of
aborting tool loading when the config lookup fails.

* chore: drop unused resolveConfigServers import

* fix: forward server provenance on the all-tools path and read config once

createMCPTools builds each toolKey from the server name it already has but did
not forward it, so the sys__all__sys path re-derived it by parsing and bound
an unmodified YAML server whose name contains the delimiter to the wrong auth
and invocation context.

loadAgentTools now resolves the MCP server context once and threads it into
loadTools, replacing the second app-config read it had introduced on the
non-event-driven chat startup path.

* fix: carry resolved MCP server name through tool classification

definitions.ts resolves the server for each key and then dropped it when
building loadedTools, so buildToolClassification re-derived it with a
last-segment split and recorded 'Workspace' for a server configured as
'Google_mcp_Workspace'. The resolved name now rides along on the tool
instance and classification prefers it over re-parsing.

* fix: consume carried server name when extracting MCP servers

extractMCPServers re-derived the name with a last-segment split, so a server
configured as Google_mcp_Workspace resolved to Workspace and its instructions
were silently omitted. Prefers the name carried on the tool definition
instance, falling back to the split.

* fix: fail closed on ambiguous MCP keys when persisting server names

Persisted mcpServerNames grant agent-scoped access to a DB server by name
(ServerConfigsDB.getAccessibleServers), so a wrong guess exposes an unrelated
server to everyone who can view the agent. The last-segment split turned
search_mcp_Google_mcp_workspace into 'workspace'; such keys were previously
rejected outright at agent save, so admitting them opened this path.

Derives a name only from unambiguous single-delimiter keys. This is #12250's
guard moved to the boundary it was actually protecting, instead of blocking
tool admission.

* fix: keep DB server access for multi-delimiter tool keys

The fail-closed guard was wrong for the case this PR exists to fix. This index
only grants DB-backed servers, and DB names are slugs that cannot contain the
delimiter (generateServerNameFromTitle strips underscores), so the trailing
segment is always the real server for them - dropping it cost every consumer
of a gateway-prefixed tool their shared-agent access.

Also gates the MCP server-context lookup on the filtered MCP set, so an agent
with no MCP tools no longer pays an app-config read on startup.

* fix: resolve tool-call display names without breaking OAuth calls

The display parsers could not use the shared boundary parser because their
tested behavior depends on first-delimiter semantics. That constraint only
applies to synthetic MCP OAuth calls, whose tool half is always exactly
'oauth', so everything after the first delimiter is the server even when the
server name carries one.

splitToolCallName special-cases that form and defers to splitMCPToolKey for
real tool keys, so a gateway-prefixed tool now renders its own name and
server while oauth_mcp_foo_mcp_bar still resolves to foo_mcp_bar.

* fix: persist resolved MCP server provenance on agents

Deriving mcpServerNames from the tool key cannot tell a config server's
trailing segment from a real DB server name, so a config server named
a_mcp_b indexed an unrelated DB server b and shared the agent's viewers into
it. Neither string rule works: the suffix guess exposes, and failing closed
drops legitimate DB access for gateway-prefixed tools.

filterAuthorizedTools already resolves each tool's server against the merged
registry config, so it now collects those names and create, update and
duplicate persist them. No extra registry queries: the update path unions the
newly resolved names with what the agent already had, and duplicate replaces
the copied list rather than inheriting the source's servers.

Display parsing also takes the configured names, so a real tool call on a
delimiter-bearing server renders the right server and icon.

* test: teach MCP hook mocks about useMCPServerNames

Three specs mock ~/hooks/MCP with a hand-listed factory, so adding the hook
to ToolCall made useMCPServerNames undefined under test and every render
threw. Returns a stable array so the mock cannot perturb render counts.

* fix: rebuild agent MCP server index from surviving tools

Unioning the prior names kept a server indexed after its last tool was
detached, so viewers of a shared agent retained agent-scoped access to it.
The index is now rebuilt from the tools that survive the edit: a prior name
carries forward only while some retained tool still resolves to it, using the
agent's own persisted names as the candidate set, and the rebuild runs on any
tool change rather than only when a new MCP tool is added.

* fix: keep duplicate indexes on registry fallback and harden the oauth split

Duplication blanked mcpServerNames when the registry was unavailable, because
filterAuthorizedTools grandfathers the source's tools without resolving them -
the copy kept tools it could no longer resolve. Source names now carry forward
for the tools that still point at them.

splitToolCallName also treated any oauth_mcp_ prefix as a synthetic OAuth
call, so a genuine upstream tool by that name resolved to the wrong server. A
configured server name now decides when one matches, since a real key always
ends in its server, and the prefix only breaks ties for unconfigured servers.

* fix: thread configured server names through display parsing

parseToolName and getMCPServerName resolved context-free, so a configured
server whose name contains the delimiter showed the wrong server in grouped
tool summaries and subagent tool labels, and stacked icons missed its entry in
the icon map. Both take the configured names now, supplied by the components
that render them.

Adds the hook to SubagentCall's mock factory: the spec renders the real
component, so an unmocked useMCPServerNames would reach the query with no
provider.

* test: cover the auth-map boundary, server provenance and context fallback

Adds regression coverage for three behaviors this PR changed that no test
exercised: customUserVars resolving under the right plugin key for a
gateway-prefixed tool name (the failure that made these tools loadable but
unusable), the resolved server name reaching createMCPTool instead of being
re-parsed, and resolveMcpServerContext degrading to empty rather than
aborting tool loading when the config lookup fails.

Each was checked against a mutated source to confirm it fails when the
behavior is broken.

* fix: normalize server-name candidates and cover the boundary guard

Tool keys embed normalizeServerName's output while the config is keyed by the
raw name, so callers passing raw keys never matched a server whose name needs
normalizing and silently fell back to the last delimiter. filterAuthorizedTools
now maps normalized names back to their config key, and createMCPTool
normalizes its candidates.

Adds the cases an audit found surviving mutation: a configured name that is a
bare but not delimiter-aligned suffix must not match, an empty candidate list
behaves as no list, and splitToolCallName still falls back to the oauth prefix
when a list is supplied but nothing in it matches.

* fix: keep resolved server names when a non-owner retains MCP tools

The shared-agent path keeps an agent's existing MCP tools verbatim but supplied
no mcpServerNames, so persistence re-derived them and reduced a configured
server like Google_mcp_Workspace to Workspace - which ServerConfigsDB then
treats as a DB server, granting the agent's viewers access to an unrelated one.
Carries the existing resolved names across instead, and clears the index on the
owner path where every MCP tool is removed.

* fix: preserve resolved MCP names for every tools update

extractMCPServerNames was reachable from any caller that writes tools without
mcpServerNames - the Action edit path does exactly that - so a configured
Google_mcp_Workspace was reindexed as Workspace and ServerConfigsDB granted
shared-agent viewers an unrelated DB server by that name.

updateAgent now rebuilds the index from the agent's own resolved names: one
carries forward while a retained tool still resolves to it, and only keys
matching none of them fall back to derivation. Callers are safe by default
rather than by remembering to pass the set.

normalizeServerName moves to librechat-data-provider so the client can match
its candidates against tool keys, which embed the normalized form; the icon map
is keyed the same way since it is looked up with a parsed server name.

* refactor: move MCP context resolution into packages/api

New backend logic belongs in the TypeScript workspace per CLAUDE.md, with /api
kept to a thin wrapper. resolveMCPServerContext now lives in
packages/api/src/mcp/context.ts and takes ensureConfigServers by injection,
since the registry accessor is still legacy-only; the /api function is reduced
to loading the request app config and translating failures into the empty
degrade it already promised.

* test: teach the MCP service mock about resolveMCPServerContext

The spec mocks @librechat/api with a hand-listed factory, so moving the
resolver into that package left it undefined and the wrapper degraded into its
own catch, returning empty config servers. The stub mirrors the real resolver
so these tests still cover what the wrapper owns - loading the request config
and degrading on failure - while the resolution logic is unit-tested in
packages/api.

* fix: only persist an authoritative MCP server index on update

Assigning the resolved set unconditionally pinned the index to [] whenever
nothing authoritative was available - a legacy agent holding MCP tools with no
stored mcpServerNames - which suppressed updateAgent's derivation and stripped
agent-scoped access to its DB-backed server.

The field is now supplied only when the result is authoritative: names were
resolved, or no MCP tool survives so the index genuinely is empty. The
retained-tools branch likewise leaves it unset when the agent has none stored.

---------

Co-authored-by: Jens Schumann <schumajs@gmail.com>
2026-07-27 14:45:38 -04:00
Danny Avila
a53936d273
🧭 test: Cover Agent Handoffs End to End (#14428)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Has been cancelled
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Has been cancelled
Sync Helm Chart Tags / Ignore non-main push (push) Has been cancelled
Sync Helm Chart Tags / Sync chart tags (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Has been cancelled
* test: cover agent handoffs end to end

* style: sort handoff imports

* fix: normalize missing agent handoff edges

* chore: update package dependencies and versions in package-lock.json and package.json

* chore: bump agents SDK
2026-07-27 08:47:15 -04:00
Danny Avila
d8427ffc5e
🛂 test: Cover Tool Approval Workflows End to End (#14427)
* test: cover tool approval workflows end to end

* fix: preserve tool approval state across resume

* fix: preserve agent context in mock stream responses

* fix: preserve nested approvals in collapsed groups
2026-07-26 21:58:25 -04:00
Danny Avila
f3159f9891
🧩 fix: Harden Agent Skill Lifecycles End to End (#14429)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Has been cancelled
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
* test: cover agent skill lifecycles end to end

* style: sort agent skill imports
2026-07-25 08:19:12 -04:00
Danny Avila
73699b5c25
perf: Reduce Agent Chat Startup Latency (#14423)
* perf: reduce agent chat startup latency

* test: align Redis stream readiness assertions

* perf: overlap remaining agent startup work

* perf: persist initial agent job metadata atomically

* test: add agent startup latency benchmark

* fix: harden resumable agent stream lifecycle

* fix: isolate replacement stream lifecycles

* fix: preserve terminal stream epochs
2026-07-25 07:58:20 -04:00
Danny Avila
00c5a747e9
🧵 feat: Native Background Execution for Code Interpreter Tools (#14386)
* 🧵 feat: Native Background Execution for Code Interpreter Tools

* 🩹 fix: Address Codex Round 1 (fallback dedupe, harvest failure, handle parsing)

* 🩹 fix: Live Completion Marker + Unkeyed Attachment Dedupe (Codex Round 2)

* 🎨 chore: Sort Imports + Widen Marker Type Comparison (CI)

* 🩹 fix: Stale-Harvest Guard, Error Marker Status, Faster Anchor Retry (Codex Round 3)

* 🧹 refactor: TS Harvest Module, Claim-Neutral Timestamps, Error Parity (Codex Round 4)

* 🩹 fix: Dispatch-Ordered Stale Guard, Foreground Downgrade, Error Wrapper Parity (Codex Round 5)

* 🩹 fix: Retry Past Unfinished Rows + Per-Call Attachment Dedupe (Codex Round 6)

* 🩹 fix: Writer-Dispatch Ordering, Scoped Live Upserts, Reaped-Task Wrapper (Codex Round 7)

* 🩹 fix: Wildcard toolCallId Matching for Bare Attachment Updates (CI)

* 🩹 fix: Claim-Insert Dispatch Stamp (Schema-Backed) + Scoped Status Markers (Codex Round 8)

* 🩹 fix: Pre-Write Ownership CAS + Agent-Scoped Part Patching (Codex Round 9)

* 🩹 fix: Insert-Path Ownership CAS + Agent-Routed Attachments (Codex Round 10)

* 🩹 fix: Agent-Scoped Marker Ids and Attachment Dedupe (Codex Round 11)

* 🩹 fix: Atomic File Commit and Sibling Preview Fan-Out (Codex Round 12)

- Replace the two-step claim-confirm CAS with an atomic conditional updateFile: the ownership predicate (no sourceDispatchedAt, or <= this write's dispatch order) moves into the update filter, removing confirmCodeFileOwnership and the lost-update window between check and write
- Thread agentId through createDownloadFallback so fallback download rows scope to the emitting agent like primary rows
- Fan terminal preview overlays out to every live attachment sharing the file_id in useAttachmentPreviewSync (sibling tool calls no longer stick on pending)
- Restore background artifacts through toStoredArtifact so the size bound applies on re-anchor
- Apply filterAttachmentsForPart to grouped tool-call attachments in ContentParts so handoff agents with colliding provider call ids do not cross-contaminate groups

* 🩹 fix: Agent-Scoped Live Upserts and Monotonic Dispatch Stamps (Codex Round 13)

- Scope the SSE attachment upsert and the useAttachments DB/live merge by agentId with the same wildcard semantics as toolCallId: distinct non-null agentIds stay separate entries, so handoff agents sharing a claimed file_id and a repeated provider tool id (call_0) no longer merge over each other's cards
- Extend the attachment identity key to fileKey::toolCallId::agentId and register less-specific key variants so bare and agent-less live records still dedupe after overlay
- Stamp background task createdAt from a strictly-increasing per-process dispatch counter: raw Date.now() can tie for same-millisecond dispatches and the stale-output guard accepts equal stamps (needed for idempotent re-commits), which would let an older task overwrite a newer task's committed file
2026-07-22 22:13:15 -04:00
Danny Avila
ad5bb477af
🎞️ fix: Surface Clear Error for Unprocessable Gemini YouTube Videos (#14396)
Google rejects a YouTube video it cannot ingest with a generic
`400 INVALID_ARGUMENT` that names no cause, which LibreChat relayed
verbatim. Attribute the failure using request context instead: when a
Google/Vertex turn carried an injected YouTube video part and the
provider returns that generic rejection, map it to a typed error the
client localizes.

Verified against the live API: a public 9h15m video is refused this way
on gemini-2.5-flash, 3.5-flash, 3.5-flash-lite and 3.6-flash, including
at MEDIA_RESOLUTION_LOW, while a short video with an identical payload
succeeds. Duration is the dominant trigger; region and access
restrictions return the same response, so the copy leads with length
without overclaiming.

A duration preflight was evaluated and skipped: oEmbed does not expose
duration, leaving only watch-page scraping — a blocking call against
undocumented markup from rate-limited datacenter IPs that would fail
open and still need this mapping underneath.
2026-07-22 12:11:06 -04:00
Danny Avila
3e9f07976a
🧩 fix: Preserve Deployment Skill IDs on Agents (#14368)
* fix: preserve deployment skills on agents

* fix: expose deployment skills to agent viewers

* refactor: centralize deployment skill ID merging

---------

Co-authored-by: Dennis Schenk <dennis@gridonic.ch>
2026-07-21 19:44:27 -04:00
MarcAmick
ade02054c8
🛟 fix: Keep File Uploads Alive With SSE Heartbeats (#14295)
Some checks are pending
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
* fix: Use SSE to upload files in order to avoid idle timeouts.  Idle timeouts can occur for example from gateways and other services like cloudfare when uploading large files.  For example during rag processing the file is uploaded to librechat which then sends it to rag.  While librechat is waiting for the embeddings to come back from rag the file upload is sitting idle.  Gateways tend to want to cancel the upload with an http 408 , 504, or 524.  This change uses SSE to perform the upload so that while librechat is sending the file to rag, it consistently sends back a heartbeat event to the client to keep the connection alive.  This is especially useful when utilizing  EMBEDDING_BATCH_SIZE in librechat rag which will allow rag to process signifigantly larger files without running out of memory.

* added tests to packages\api\src\files\sse.spec.ts in order to test the new sse.ts

* fix: Harden SSE file upload lifecycle

* style: Sort data provider imports

---------

Co-authored-by: Marc Amick <MarcAmick@jhu.edu>
Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-21 19:27:09 -04:00
Danny Avila
9e245aced4
🎟️ fix: Claim Idempotency Keys to Dedup Retried Generation Requests and Prevent Double Billing (#14344)
* 🐛 fix: Dedup retried start-generation requests to prevent duplicate billing

A lost or reset start-generation response makes the client re-POST the
identical payload (up to 3x on network errors). The resumable-stream
controller had no idempotency: createJob unconditionally overwrote the
running job without aborting the prior one, so both requests ran full
LLM completions and both billed while the UI showed only one (#14339).

Add a stable per-submission clientRequestId (uuid, fresh per ask() so a
regenerate differs, reused across the start-generation retries) and an
atomic claim on the job store keyed by userId:clientRequestId. The first
request wins and generates; a retried POST loses the claim and receives
the original stream, which the client subscribes to and replays - no
second billed generation.

- IJobStore.claimIdempotencyKey/releaseIdempotencyKey (in-memory Map+TTL,
  Redis single-key SET NX PX + GET Lua, cluster-safe)
- GenerationJobManager.claimGeneration/releaseGeneration (20m TTL)
- Controller claims before the concurrency check, dedups with a resumed
  response, releases on start-failure/429
- clientRequestId threaded through TSubmission/TPayload/createPayload

* 🐛 fix: Harden start-generation dedup (Codex review)

Address three P2 findings on the idempotency path:

- Resume replay: a deduped retry now subscribes with resume=true so the
  client replays prior content and any pending-action from the running
  stream instead of only live events (cross-replica / HITL correctness).
  startGeneration returns { streamId, resumed } and the response's
  status:'resumed' drives the subscribe mode.
- Wait for the job record: a duplicate that loses the claim now waits
  briefly for the winner to create the job before returning the stream
  (a stream with no job 404s terminally). If the winner has not
  materialized, return 503 SERVER_NOT_READY so the client retries via the
  existing readiness path instead of attaching to a dead stream.
- Release only owned claims: track whether the request actually won the
  claim; the 429 and init-error paths no longer release a claim owned by
  another in-flight generation (fail-open path could erase it and
  re-enable double billing).

Adds controller tests covering dedup, the 503 race fallback, win-then-
create, and claim-release ownership on 429 / fail-open.

* 🐛 fix: Don't trap deduped retries on missing job records (Codex review)

The previous round returned 503 SERVER_NOT_READY when a deduped retry's
job record was absent. But a missing job usually means the original
generation already completed and was cleaned up (cleanupOnComplete) — the
correct recovery is to return the stream and let the client's subscribe
404 handler refetch the persisted messages. The 503 instead trapped the
send in a readiness-retry loop until the client's window expired.

Keep the bounded wait (it still covers the job-about-to-be-created race)
but always return the resumed stream afterward; a gone/never-created job
recovers via the client's existing 404 path instead of being treated as
indefinitely starting. Updated the controller test accordingly.

* 🐛 fix: Gate deduped resume on claim age, not just job presence (Codex review)

Removing the 503 entirely (previous round) reintroduced the inverse race:
if the winning request stalls between claimGeneration and createJob, a
losing duplicate saw no job, returned status:'resumed' anyway, and the
client subscribed to a stream that did not exist yet — the 404 handler
tore the turn down while the winner went on to generate and bill with no
UI attached.

Distinguish the two missing-job cases by claim age (claimedAt now travels
on the claim value):
- fresh claim, no job yet → winner is still starting → 503 SERVER_NOT_READY
  so the client retries via the readiness path (bounded, not indefinite).
- old claim, no job → the original already completed and was cleaned up
  (or the winner died) → attach; the client's 404 handler refetches.

Tests cover both age branches.

* 🐛 fix: Scope dedup fail-open + keep resumed convos on 404 (Codex review)

- Fail-open only on claim acquisition: a store error while checking an
  already-confirmed existing claim no longer falls through to createJob
  (which would start a second billed generation during a Redis hiccup).
  Once claim.existing is known, a job-lookup error returns 503 retry.
- Don't drop a resumed convo on 404: the optimistic-conversation cleanup
  in useResumableSSE now runs only for fresh (non-resume) subscribes. A
  deduped resume whose original completed and was cleaned up 404s, but its
  conversation is persisted and must stay in the sidebar.

Adds a controller test for the job-lookup-error path (503, no createJob).

* 🐛 fix: Reconcile resumed convos on 404 instead of guessing (Codex review)

Round-4's !isResume guard fixed the completed-and-cleaned case (don't drop
a persisted convo) but left the inverse: a new-conversation retry deduped
to a claim whose original worker died before persisting still resumes,
404s, and — with removal skipped — leaves a phantom /c/<streamId> sidebar
entry.

Stop guessing keep-vs-remove on a resume 404. Reconcile against the
server: invalidate the conversations list so a real (persisted) convo
stays and a phantom is dropped. Fresh (non-resume) optimistic streams
still prune immediately. Adds a client test for the resume path.

* 🐛 fix: Finalize failed job before releasing its claim (Codex review)

In the initialization-error catch, the idempotency claim was released
before completeJob(streamId). A racing retry could win the released key
and createJob() the same streamId while this catch was still running, and
completeJob() (not guarded by the original createdAt) would then abort the
replacement. Finalize the failed job first, then release the claim.

Adds a controller test asserting completeJob precedes releaseGeneration.

* 🐛 fix: Clear claims on destroy + survive completeJob failure (Codex review)

- InMemoryJobStore.destroy() now clears the idempotencyClaims map, so a
  reused/reconfigured store instance doesn't dedup a fresh start against a
  torn-down job's stale claim.
- Init-error cleanup: completeJob() is swallowed so a store-hiccup
  rejection can no longer skip the idempotency-key release and the
  pending-request decrement (which would wedge the retry behind the claim
  and leak the concurrency slot). A failed completeJob finalized nothing,
  so releasing afterward still can't abort a later replacement.

Tests: claims cleared on destroy; release + pending decrement still run
when completeJob rejects.
2026-07-21 08:16:31 -04:00
Danny Avila
3337bde050
🚏 fix: Route Admin-Configured Document Types to RAG /text on Agent Upload (#14345)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* 🩹 fix: Route configured document types to RAG /text on agent upload

Restores pre-#11900 behavior for 'Upload as Text': when an admin narrows
fileConfig.text.supportedMimeTypes to a non-permissive allowlist that includes
a document type (docx/xlsx/pdf/ods/odt) and a RAG API is configured, the file
is sent to RAG /text instead of the built-in document parser.

The permissive default catch-all is excluded via isPermissiveMimeConfig, so RAG
deployments that never customized text handling keep the built-in parser. When
RAG is unreachable, parseText's new allowNativeFallback:false makes it throw so
the upload falls back to the built-in document parser rather than degrading a
docx/pdf to raw native-text bytes.

Fixes #14245

* style: sort imports in text.spec.ts (CI import-order)

* 🩹 fix: Scope RAG fallback catch to extraction only, not persistence

The configured-text branch wrapped both parseText and createTextFile in the
fallback try, so a persistence failure after a successful RAG extraction (size
guard, db.createFile, agent-resource mutation) was misread as RAG-unavailable
and retried with the built-in document parser, masking the real error and
risking a duplicate agent-resource mutation. Only the RAG extraction is now in
the fallback catch; a persistence failure surfaces as itself.

Addresses Codex P2 on #14345.
2026-07-20 22:46:28 -04:00
Danny Avila
74de989bde
🚪 fix: Keep Owners From Being Locked Out of Their Own Resource When Sharing (#14347)
* 🔒 fix: Skip revoke for principals also being granted (owner-lockout guard)

bulkUpdateResourcePermissions flushes grants (upserts) before revokes (deletes).
If a principal appears in both updatedPrincipals and revokedPrincipals, the ACL
entry is granted and then immediately deleted, stripping access the caller just
set. This can strip a resource owner's own grant when the share dialog places
the owner in both lists from a client id/idOnTheSource mismatch (OpenID/Entra).

Add a server-side guard: track principals granted in the same request and skip
any revoke for the same principal, so granting wins and owner lockout is
impossible regardless of how the client computes the share diff. Complements the
client-side keying fix in #14317.

Refs #14316

* 🔒 fix: Exclude PUBLIC from grant-wins guard so public-disable is honored

The grant-wins guard must not apply to PrincipalType.PUBLIC. An explicit
public: false disable adds the public principal to the revoke list; a
contradictory payload that also grants public (public in the updated list) would
otherwise skip the revoke and leave the resource public. Disabling public access
must always win. User/group owner-lockout protection is unchanged.

Addresses Codex P2 on #14347.

* 🔒 fix: Move revoke guard inside per-principal try (tolerate malformed entries)

The grant-wins guard read principal.type before the per-principal try/catch, so
a malformed revoke entry (e.g. removed: [null]) would throw out of
bulkUpdateResourcePermissions after grants were already flushed on
non-transactional MongoDB, leaving partial permission changes. Move the guard
inside the try so a malformed entry is recorded in results.errors and skipped,
matching prior behavior.

Addresses Codex P2 on #14347.
2026-07-20 22:27:25 -04:00
Danny Avila
d02867a3e2
🔍 refactor: Surface primary OpenID JWT failure reason in auth failure log (#14346)
When OPENID_REUSE_TOKENS is enabled and both the openidJwt strategy and the
HS256 jwt fallback fail, the final 'Authentication failed after all strategies'
warn log reported only the fallback's reason. For an RS256 provider (Keycloak,
Auth0, Okta) that surfaces as 'invalid algorithm', which is the HS256 fallback
rejecting the provider token, not the real reason openidJwt did not
authenticate, and it was previously only visible at debug level.

Include the captured primary (openidJwt) failure reason and error name in the
final warn log so reused-token failures are diagnosable without enabling debug
and are not misattributed to the fallback.

Refs #14311
2026-07-20 21:05:20 -04:00
Danny Avila
20cd00c492
🖼️ feat: Return Sandbox Images From read_file as Viewable Artifacts (#14277)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
* 🖼️ feat: Return Sandbox Images From `read_file` as Viewable Artifacts

The code-execution sandbox `read_file` path refused every image
extension because it reads files via `cat` over codeapi's JSON `/exec`
transport, which lossily corrupts non-UTF-8 bytes. The skill-file read
path already surfaced images as artifacts; this brings the sandbox path
to parity so an agent can actually see a chart/screenshot it reads.

- `readSandboxImage` (process.js): a Python base64 reader over `/exec`
  with an in-sandbox size guard so oversize images never cross the wire;
  base64 is ASCII-safe where `cat` corrupts.
- `handleSandboxImageRead` (handlers.ts): byte-integrity check (guards
  against a truncated `/exec` stdout), MIME resolved purely from the
  magic-byte sniff (extension only routes; a mislabeled non-image falls
  back to the bash hint), and graceful degradation on every failure mode.
- Shared `buildImageArtifactResult` used by both read paths; the result's
  `artifact.content` image_url reaches the UI (tool-end callbacks save it
  as an attachment) and the LLM (SDK folds it into the model-visible
  message for Anthropic/OpenAI/Google).

*  test: Sync read_file code-only description assertions with image wording

* 🛡️ fix: Harden sandbox image reads (regular-file guard, completeness check)

Addresses Codex review on PR #14277:

- readSandboxImage now os.stat's the target and rejects non-regular files
  (FIFOs, sockets, /dev/* symlinks) via stat.S_ISREG, and bounds the read at
  limit+1 bytes — a device/FIFO can no longer stream unbounded into memory
  until the request times out.
- handleSandboxImageRead validates completeness (not just the magic header):
  PNG must end with the IEND trailer and WebP's RIFF size must match the byte
  length, so a truncated/interrupted image degrades to the bash hint instead
  of being sent as a corrupt image_url. JPEG/GIF stay header-level (they can
  carry trailing metadata; a strict end-marker would risk false rejections).

* 🩹 fix: Chunk sandbox image reads to fit the runner stdout cap

Inlining any real image failed with "is an image file (.png) and cannot
be read as text". Root cause: readSandboxImage base64-encodes the file to
STDOUT, but the runner caps stdout at SANDBOX_OUTPUT_MAX_SIZE (1024 bytes
by default) and SIGKILLs the job on overflow (status OL), truncating the
JSON mid-base64. The parse then threw and the handler degraded to the
binary hint. The in-sandbox MAX_BINARY_BYTES=5MB guard never fired because
the *transport*, not the file size, is the real ceiling: a 5MB image needs
~6.8MB of stdout. Reproduced against a live MicroVM — a 186KB matplotlib
PNG died with 'stdout length exceeded' at exactly the 65536-byte cap.

Read the file in windows instead: each /exec pulls  raw bytes at an
offset and base64s only that slice, so every response stays under the cap
regardless of how the runner is configured; the chunks are reassembled and
verified against the sandbox-reported total. Verified end-to-end on a real
MicroVM: 25KB and 186KB PNGs both round-trip byte-exact (sha256 match).

Also:
- Detect the truncation explicitly (status OL) and name the fixable cause
  (chunk size / SANDBOX_OUTPUT_MAX_SIZE) instead of "unexpected output".
- Parse the LAST stdout line so a shell banner can't break the read, and
  include a stdout snippet when it genuinely is unparseable.
- LIBRECHAT_CODE_IMAGE_CHUNK_BYTES (default 32KB) tunes the window.
- Tests drive the real reader against a mocked /exec transport rather than
  mocking readSandboxImage, which is why the existing suite stayed green
  through this bug.

* 🎯 fix: Cap sandbox inline images at 1MB, separate from skill-file reads

The sandbox and skill-file image paths shared MAX_BINARY_BYTES (5MB), but
their transports differ: skill files stream from storage, while sandbox
bytes come back base64 over /exec stdout under the runner's output cap, so
the reader windows the file and cost scales in round-trips (~160 at 5MB vs
~32 at 1MB). Nothing is gained by allowing more — vision providers
downsample to ~1.5-2k px regardless, so multi-MB originals buy no fidelity
while grinding through round-trips.

Give the sandbox path its own MAX_SANDBOX_INLINE_IMAGE_BYTES (1MB), used
for both the read cap and the over-limit message (which previously quoted
5MB while the reader enforced something else). Skill-file reads keep 5MB.

Verified against a live MicroVM: a 186KB PNG round-trips byte-exact, and a
1.4MB file returns tooLarge in a single round-trip with zero bytes
transferred, degrading to the existing bash_tool hint.
2026-07-16 07:27:33 -04:00
Danny Avila
7447fddfb2
🙊 refactor: Clarify Ask Question Schema Errors and Retry Guidance (#14279)
* fix(agents): clarify ask question validation errors

* fix(agents): narrow question failure detection

* fix(agents): persist question validation failures

* fix(agents): track question validation failures
2026-07-15 11:06:29 -04:00
Jaime Hidalgo
e813934731
🔖 fix: Preserve Ephemeral Agent Params and Identity for ask_user_question Resume (#14254)
* fix: durable ask_user_question resume for ephemeral agents

* 🤖 refactor: Drop chat.js resume hunks in favor of shared packages/api helpers

* 🤖 fix: Normalize resume thinking param and replay modelLabel (#14253 Bugs 1&2)

* 🤖 fix: Preserve adaptive thinking display and effort across HITL resume

* 🔤 style: Sort load.spec.ts imports (repo import-order)

* 🤖 fix: Replay paused request body params on HITL resume (UI-form source of truth)

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-14 15:12:43 -04:00
Danny Avila
5771bf6e06
♨️ feat: Prewarm Stateful Code Sandboxes with Cold-Boot UX Feedback (#14239)
* ♨️ feat: Prewarm Stateful Code Sandboxes with Cold-Boot UX Feedback

* 🧹 fix: Drain Prewarm Response + Reset Sandbox Atoms on Stream Cleanup

* 🚿 fix: Propagate Prewarm Drain Failures + Warm Marker for Host File Tools

* 🌡️ fix: Decouple Prewarm In-Flight State from Warm Refreshes + Precise Ready Gates

* ☁️ refactor: Redis-Backed Sandbox Prewarm State via standardCache

* 🧪 chore: Hermetic Prewarm Spec + Accurate Signal JSDoc (Copilot review)
2026-07-14 10:25:37 -04:00
Danny Avila
9bb351ad9c
🧭 feat: Mid-Run Steering and Queued Messages for Agent Runs (#14220)
* 🧭 feat: Mid-Run Steering and Queued Messages for Agent Runs

Steering: submit a message while a run is generating; the server queues
it in the job store (cross-instance) and a run-scoped PostToolBatch hook
injects it into graph state at the next tool-batch boundary, records an
inline 'steer' content part on the response (replayed as a user message
on later turns), and streams on_steer_applied to the client.

Queuing: messages composed during a run auto-send as normal follow-up
turns after clean completion (one per final event, FIFO); user aborts
leave them as chips unless armed by interrupt-and-send.

Requires hook injectedMessages support in @librechat/agents
(danny-avila/agents#299); hard-gated via a capability probe so older
SDKs 501 the steer route instead of draining and dropping messages.

* 🧵 fix: Harden Steering Against Finalization Races and Route Guard Gaps

Addresses local Codex review findings on the steering feature:

- Close-and-drain the steer queue atomically at finalization (final event,
  abort) so a steer POST racing teardown is rejected instead of 202-ACKed
  and then silently cleared; the closed flag lives on the job hash and is
  reset when a replacement job reuses the stream id.
- Clear inherited steer queues on createJob — a job replacement must not
  drain the replaced run's messages.
- Keep steers queued across a HITL pause instead of draining them into
  ephemeral client state: resumeState re-seeds chips on reload and the
  resumed run injects them at its first tool boundary (steers key TTL now
  extends to the approval window; on_steers_pending event removed).
- Queue the NO_ACTIVE_RUN steer fallback while the final SSE is still
  settling — a direct send would be dropped by ask()'s in-flight guard.
- Reconcile the 202 ACK against on_steer_applied events that beat it over
  the SSE, so a chip can't be re-minted after its removal event passed.
- Allow the per-send Steer override when the default action is queue.
- Apply the configured message rate limiters and the PII filter to
  POST /chat/steer — a steer is model-bound user text.

*  ci: Assert Steering Capability Probe Against the Installed SDK

CI installs the published @librechat/agents pin (pre-injectedMessages),
where isSteeringSupported() is legitimately false — the probe test now
asserts it mirrors the installed SDK's capability flag instead of
hardcoding the capability-bearing build's value. Verified against both
the published 3.2.61 dist and the agents#299 build.

* 🛟 fix: Preserve Steer Text Across Run-End, Error, and Abort Races

Codex round 2 (4 P2s):
- Applied-steer-id set survives run end (capped at 100) and converted
  ids join it, so a 202 ACK that lands after final/abort drops its chip
  instead of re-minting a stranded pending one.
- Failed runs no longer strand acknowledged chips: both error paths
  convert local pending chips to queued follow-ups (chip text is
  client-local), and the server closes the steer queue before emitting
  the error so a racing steer POST gets 404 fallback instead of a 202
  whose payload dies with the job.
- sendQueuedNow keys on steer availability, not the default action —
  send-now on a queued chip is an explicit override for queue-preferring
  users.
- Stop path consumes pendingSteers from the abort HTTP response as a
  fallback for the SSE final event it may close before processing;
  conversion is deduped so double delivery is a no-op (shared
  useSteerConvert hook).

* 📎 feat: Carry Attachments Through During-Run Queued Messages

Steering stays text-only (SDK injection, inline STEER part, and replay
are all text), so a during-run submit with media now queues the whole
message as one unit instead of silently stranding the files:

- QueuedMessage gains `files`; composer attachments are consumed into
  the queued item at queue time (steerFromComposer / queueFromComposer /
  interruptAndSend), fixing the latent hazard where lingering composer
  files glued onto whatever `ask` vacuumed up next.
- Enter-steer with attachments degrades to queue with an explanatory
  toast; the per-send menu routes through the same composer-aware
  wrappers.
- The drain and sendQueuedNow pass the item's files as `overrideFiles`;
  media items never steer (send as a normal turn when idle, re-front
  otherwise). ask() no longer clears composer state for caller-supplied
  overrideFiles — only regenerate keeps that behavior.
- During-run submits hold while uploads are in flight, mirroring the
  send button's filesLoading gate; queued chips show a paperclip count.

* 🎛️ feat: Rework During-Run Chips into Action Rows

Full-width rows above the composer (reference-UI parity): each queued
message shows a primary Steer/Send-now action, delete, and a "…" menu
with Edit message (restores text + attachments into the composer) and a
Turn on queueing/steering toggle that flips the Enter default. Steer
rows share the layout with status text; failed steers keep retry /
edit / queue-convert. The per-send menu gains the same default toggle.
Queued file refs now retain filename + bytes so edit-restore rebuilds
real composer entries (draft-recovery shape).

* 🖇️ feat: Steer With Attachments (Multimodal Mid-Run Injection)

Steering now carries media end-to-end instead of degrading to queue:

- The steer POST accepts sanitized attachment refs (cap 10; only
  file_id is trusted — the drain re-fetches owner-scoped and re-derives
  everything else). SteerQueueItem/TPendingSteer/SteerContentPart carry
  `files` refs; encoded data is never persisted or queued.
- New api/server/services/Files/steering.js decouples attachment
  building from the request path: encodeSteerContent reuses the exact
  per-turn pipeline (addFileContextToMessage + processAttachments'
  single-pass categorize/encode, SDK formatMessage assembly,
  prependFileContext for extracted text) with zero new encoding code.
  buildSteerMedia feeds the drain hook's new buildMedia seam (any
  failure degrades that steer to text-only — words always land);
  stampSteerPartMedia re-encodes past steer parts per turn with ONE
  batched owner-scoped fetch and stamps a transient `media` array,
  replaced immutably so it can never leak into a save. Replay honors
  resendFiles like regular message media.
- The SDK's formatAgentMessages (the formatter agents actually use)
  gained the steer replay branch on the PR branch; the local
  formatMessages.js branch now mirrors the media preference.
- Client: steerFromComposer consumes composer files into the POST,
  chips/seeding/conversions carry files everywhere (retry, queue
  convert, abort/error recovery), queued media items steer for real,
  and SteerBubble renders the steered attachments inline.

* 🧵 fix: Harden Steer Recovery Races and Drain Isolation

Codex round 3 (7 fixes):
- A 202 ACK landing after the run ended converts straight to a queued
  follow-up (server queue is gone; no event will ever resolve a pending
  chip for a finished run). Covers stream errors with in-flight POSTs.
- A Stop that lands pre-completion can arrive as a final with
  unfinished:true and no aborted flag — runEnd now treats it as aborted
  so queued messages are not auto-sent against the user's Stop.
- Leftover-steer conversion merges chronologically by createdAt instead
  of appending, preserving the order the user composed.
- Auto-drained queued messages pass explicit (possibly empty)
  overrideFiles/overrideQuotes/overrideManualSkills: a drain can no
  longer vacuum up files, quotes, or skill picks staged in the composer
  for the user's NEXT message (ask() treats overrideFiles != null as
  authoritative).
- Failed-steer Retry and resume-on-load chip restoration keep the
  steer's attachments.
- The job-replacement guard moved INSIDE the store's atomic
  drain/close-and-drain (Lua createdAt compare; in-memory equivalent):
  a stale run's hook or finalization can neither consume, close, nor
  steal a replacement job's steer queue, and the drain hook drops its
  separate check-then-drain round trip.

* 🧰 refactor: Typed Steer Controller, Single-Query Media Pass, Round-4 Fixes

Codex round 4 + efficiency tightening in one pass:

- Moved the steer guard ladder (validation, file sanitization via a
  shared toSteerFileRef picker, ownership/tenant checks, status-guarded
  enqueue) into packages/api as handleSteerRequest; api/steer.js is now
  a thin wrapper. Ladder covered against the REAL in-memory job manager
  in request.spec.ts; the api spec pins only the wrapper contract.
- Folded the steer replay stamp into the turn's ONE historical-files
  query: collectHistoricalFileRefs also gathers steer-part refs, the
  owner-scoped doc map rides client state, and stampSteerPartMedia
  consumes it (no second round trip) while encoding parts in parallel.
- Stamped steer media now counts against the run budget (existing
  multimodal counter over the non-text parts, folded into
  indexTokenCountMap/promptTokens after the stamp).
- Steer route runs the PII filter BEFORE moderateText, matching chat.js
  so blocked sensitive text never reaches the external moderation API.
- Interrupt & send survives the abort-response-beats-SSE-final race:
  stopGenerating writes the run-end signal itself when the one-shot
  interrupt flag is armed and no signal landed (double-fire safe).
- Resume reconciles chips against the server's still-queued list even
  when EMPTY, clearing chips for steers applied while disconnected.
- The local formatter's steer flush preserves non-text assistant parts
  (array-content AIMessage) instead of folding to text.

* 🔒 fix: Replay-Aware Capability Gate and Round-5 Race Closures

- isSteeringSupported now requires BOTH halves of the SDK contract:
  injection (HOOK_INJECTED_MESSAGES_CAPABLE) AND replay
  (ContentTypes.STEER, shipped in the same SDK commit as the
  formatAgentMessages steer branch). An SDK that can inject but not
  replay 501s the steer route — no release window can create steer
  parts that would leak into provider-facing assistant content.
- The local formatter mirrors the SDK's anchor reset: a post-steer
  tool_call mints a fresh AIMessage instead of attaching to the
  pre-steer anchor (invalid provider ordering).
- Queued-chip send-now and the NO_ACTIVE_RUN fallback pass explicit
  (possibly empty) overrideFiles so an idle send can't vacuum composer
  files staged for a different draft.
- Redis createJob deletes the stale steer list BEFORE the replacement
  hash is written as running — a steer 202-accepted against the new job
  can never be wiped by the reset.
- Resumed-turn finalization mirrors the normal path's terminal drain:
  createdAt-guarded close-and-drain, leftovers ride the resumed final
  event as pendingSteers instead of being cleared by completeJob.
- buildSteerMedia restores composer order over the $in result so
  multi-attachment steers reach the model in the order the user saw.

* ⚛️ fix: Atomic Job Replacement and Boundary-Clean Steering Module

Codex round 6 (5 fixed, 1 standing deferral):
- createJob resets the steer queue and writes the job hash in ONE
  same-slot Lua script (JOB_CREATE_LUA): a steer POST can no longer
  interleave between them on cluster, so a steer accepted against one
  run can never be drained into another. Redis-validated.
- The steering media pipeline moved to packages/api
  (agents/steering/media.ts) with injected getFiles and a structural
  client interface — /api keeps zero steering logic; specs ported to
  the DI seam.
- handleSteerRequest checks the job BEFORE the capability gate: a steer
  racing completion on an unsupported SDK gets 404 (send-now) instead
  of a 501 queue with no run-end signal left to drain it.
- useQueueDrain binds to the active conversation: navigating away
  between the final SSE and the drain effect leaves the signal
  unconsumed instead of submitting A's follow-up into B; the drain
  fires on return.
- abortJob closes and drains the steer queue BEFORE the content
  snapshot, so a drain-hook apply that lands pre-drain is captured
  inline rather than lost between the snapshot and the terminal drain.

* 🚦 fix: Parked Run-End Signals, Interrupt Priority, Settled-Run Fallbacks

Codex round 7 (5 fixes):
- Run-end signals for a non-active conversation are PARKED per
  conversation instead of squatting the shared index slot: a later run
  finishing on the same pane can no longer overwrite them, and the
  parked drain fires when the user returns.
- "Interrupt & send" front-inserts carry a priority flag that outranks
  createdAt when abort leftovers merge back chronologically — the
  urgent redirect drains first, not the oldest steer.
- STEER_UNSUPPORTED/RUN_PAUSED/QUEUE_FULL rejections landing after the
  run settled mirror the NO_ACTIVE_RUN fallback and send immediately
  (queueing would strand the text with no run-end signal left); on the
  pinned SDK this is the common Enter-near-run-end path.
- A failed abort (e.g. 404 when the run completed first) still signals
  the interrupt drain, so the queued interrupt message can't strand and
  the armed flag can't leak onto a later run.
- Steered-image fallback alt text is localized (com_ui_attached_image).

* 📌 chore: Adopt Published @librechat/agents Types Post-Bump

dev's pin bump to ^3.2.62 (the release carrying injection + steer
replay) landed via merge; the steering runtime now uses the SDK's real
InjectedMessage/hook-output types instead of the local structural
mirrors that bridged the pre-publish window. The two-half capability
probe stays as the defensive gate for mismatched deployments — and the
capability spec now exercises its TRUE path against the published
package in CI.

* 🛅 feat: Park-and-Claim Steer Recovery + Host-View Content Reads

Codex round 8 (6 fixed incl. both P1s, 1 push-back):
- The long-deferred no-subscriber gap is closed: every terminal drain
  (final, aborted-final, error, abortJob, resumed finalize) PARKS
  acknowledged leftovers on the job hash (unrecoveredSteers), and the
  status route claims them exactly once for inactive jobs — a client
  that closed/reloaded past the transient final event restores its
  steers as queued chips within the post-terminal TTL. A replacement
  run clears the parked copy (a live client started it).
- Same-instance content reads are steer-complete: RedisJobStore now
  caches the HOST content array (WeakRef) via setContentParts and
  prefers it over the SDK graph cache, whose view never contains
  host-authored steer parts; the graph fallback splice-INSERTS steer
  chunks at their recorded host-view indices (the graph array is
  unshifted, so assignment would overwrite SDK parts).
- Replay token accounting now counts prepended file-context text: full
  stamped content minus the steer body (already counted), so large
  steered documents hit the budget instead of bypassing pruning.
- The queue drain restores an item when ask() refuses without sending
  (history not yet in cache after navigating back) — text is never
  silently dropped.
- The armed interrupt flag travels WITH a parked run-end signal, so
  another run on the same pane can neither consume nor clear it.
- parseTextParts extracts steer text (search indexing / audio).

* 🎛️ refactor: Single Send Slot + In-Thread Steer Messages

- Merge the during-run send affordance into the send/stop button slot:
  with composer text the send button replaces Stop (Enter = default
  action), hover reveals Steer/Queue/Interrupt rows with shortcuts;
  drop the separate DuringRunActionsMenu chevron
- Add during-run keyboard chords: Cmd/Ctrl+Enter = non-default action,
  Alt+Enter = interrupt & send (plain-Enter submitters only)
- Render steers as standard user messages in the thread: SteerPart
  (icon + author header + user text presentation) replaces the
  SteerBubble, and submitted steers appear immediately at the projected
  injection point via the PendingSteers slot on the streaming message
- Keep composer rows only for recoverable states: failed steers
  (retry/edit/queue) and queued follow-ups

* 🩹 fix: Keep the Replacement Submission Alive Across Abort Settlement

The aborted run's final SSE event fires before the abort HTTP response
resolves, so an armed interrupt & send drains and starts the NEXT
submission while the abort POST is still in flight. The response
handler's unconditional clearAllSubmissions() then reset the new
submission, aborting its stream attach before the subscribe — the
follow-up ran and persisted server-side but the live placeholder
finalized empty (content appeared only after reload).

useAbortCleanup captures the submission before the abort round-trip
and both settlement paths (success and 404-catch) clear only when the
captured submission is still current; a replacement stays untouched.
Plain Stop behavior is unchanged.

* 🧭 test: Playwright E2E for Mid-Run Steering and Queuing

- Add e2e/specs/mock/steering.spec.ts: steer mid-run (202 + immediate
  in-thread pending part + real MCP tool boundary + words survive run
  end), Cmd/Ctrl+Enter queue with auto-send after clean completion,
  and Alt+Enter interrupt & send with the follow-up streaming into the
  live view
- Add the E2E_STEER_TOOL_REPLY fake-model marker: slow preamble, a
  real remember_fact MCP tool call (PostToolBatch boundary), then a
  final turn
- Test 1 pins the run-end degradation contract while the SDK's
  top-level agentId stamping bug blocks live injection; its header
  documents the assertions to flip once the fixed SDK is pinned

* 🧷 fix: Job-Independent Steer Recovery + Expiry and Resume-Gap Parking

Codex round 10: the park-and-claim recovery had lifecycle holes.

- Move parked steers off the job hash onto their own bounded-TTL store
  key (JOB_CREATE_LUA resets it; deleteJob leaves it alone): the default
  completeJob path deletes the job record immediately, and the Redis
  read path never deserialized the old hash field — recovery previously
  worked only with STREAM_KEEP_COMPLETED_JOBS on the in-memory store
- Carry the owner identity inside the parked payload and authorize the
  claim against it, so the status route recovers steers on its jobless
  branch too (the common reload-after-terminal case); a non-owner claim
  returns nothing and re-parks the payload
- Park queued steers on approval expiry: snapshot the frozen queue
  before the requires_action→aborted CAS (whose terminal cleanup drops
  the steers key) and park only when the CAS wins
- Mirror the terminal drain/park block in resume.js's failure path,
  which previously let completeJob's backstop clear 202-accepted steers
- Close the Redis snapshot→subscribe resume gap: re-peek the queue
  after attaching and re-surface missed on_steer_applied events from
  the durable content view (synthesizeAppliedSteerEvents), updating
  resumeState.pendingSteers to the live queue

* 📌 chore: Require @librechat/agents 3.2.63 + Applied-Steer E2E Contract

- Bump the @librechat/agents pin to ^3.2.63 in api/ and packages/api/:
  it scopes the hook agentId marker to subagent child graphs, so the
  steering drain hook fires at top-level tool-batch boundaries and
  mid-run injection is active (danny-avila/agents PR 307)
- Flip e2e steering test 1 from the documented degradation contract to
  the applied-steer contract: the optimistic in-thread part transitions
  to the persisted part at the tool boundary and survives inside the
  response after run end, with no queued follow-up turn

* 🎗️ feat: Steered Messages Join the Message-Nav Ribs

Steers are user messages, so they get their own clickable rib on the
navigation rail, interleaved at their in-thread position inside the
response that absorbed them (one DOM query in document order). SteerPart
anchors itself as #steer-<id> with a steer-render marker — both the
optimistic pending entry and the persisted part — and the rib carries
the user role label with a preview drawn from the steer's text body,
skipping the author header.

*  feat: Cancel a Queued Steer Before Injection + True User-Message Alignment

- Add POST /chat/steer/cancel: removes ONE still-queued steer by id via
  an atomic list rebuild (Redis Lua preserves order and TTL), authorized
  against the job owner; removed:false is advisory — the cancel lost its
  race to the drain or the run end, never an error
- Surface an × on the in-thread pending steer (server-acknowledged
  entries only): optimistic removal, restored if the POST fails since
  the server would still inject the words
- Outdent SteerPart past the response's icon column so steers sit flush
  with top-level message rows, reading as regular user messages

* 🧯 fix: Round-11 Recovery Hardening + Provider-Free Pending Slot

- Reconcile the resume steer gap by steerId SETS, not queue length — a
  steer added in the gap (or an equal-length drain+enqueue swap) now
  refreshes resumeState.pendingSteers and still synthesizes the missed
  on_steer_applied events
- Make completeJob's terminal backstop park: direct error-path callers
  without the controllers' close-and-park no longer silently clear
  202-accepted steers (createdAt-guarded closeAndDrain + owner park
  before the terminal write)
- Persist the steer part BEFORE media encoding in the drain hook: an
  abort inside the encode window can no longer lose a file-steer (the
  part refs come from the enqueue-sanitized item; replay re-encodes
  per turn unchanged)
- Move the parked-claim owner check INSIDE the atomic store claim
  (substring gate in the Lua / in-memory equivalent): a non-owner probe
  can no longer transiently delete the recovery payload; the app-side
  parse stays authoritative
- Park queued steers in BOTH stores' own requires_action expiry
  cleanup, which bypassed the manager-level sweep
- Sweep expired parked steers from the in-memory store's periodic
  cleanup; restore a queued chip when send-now's submit is refused;
  upsert steer ACKs so an SSE reconnect reseed cannot duplicate chips
- Mount the cancel mutation per steer item so the pending slot needs no
  QueryClient on ordinary streaming renders (fixes the CI failure in
  ContentParts.integration.test)
- Skipped delivery-gated parking (finding 8): transport receiver counts
  cannot prove browser delivery, and gating the only durable copy on
  them trades cosmetic chip resurrection for real text loss; the window
  is already bounded by claim-on-read, createJob reset, and the TTL

* 🩺 fix: Annotate PARKED_STEERS_TTL_MS for isolatedDeclarations

tsdown's d.ts generation requires explicit types on exported consts
with computed initializers; tsc --noEmit does not run that check, so
the round-11 export slipped past local verification and broke Build
packages (and every downstream CI job that consumes the built dist).

* 🛟 fix: Round-12 Terminal-Path Recovery + Durable Steer Events

- Park queued steers before the stale-running reap deletes a crashed or
  hung job in BOTH stores — the one terminal path with no controller
  finalization; requires_action expiry parking refactored onto the same
  snapshot/park helpers
- Enqueue instead of dropping when a steer fallback send is refused:
  both the NO_ACTIVE_RUN branch and the settled-run rejection branch
  now observe sendNow's false return
- Recover on the SSE reconnect-404 terminal path: convert local pending
  steers to queued, claim parked steers via /chat/status, and write a
  non-completed run-end signal so interrupt flags release without
  auto-sending an unknown outcome
- Fall back to a positive parked-recovery TTL when completedTtl is 0
  (SET EX 0 is invalid and silently killed recovery)
- Make on_steer_applied durable before publish: emitChunk gains a
  durable option that awaits the chunk-log append (best-effort) ahead
  of the transport publish; the default delta path stays fire-and-forget

* 🔐 fix: Round-13 Steer Authorization + Trusted File Refs

- Resolve client-supplied steer file refs against the DB owner-scoped
  at enqueue and queue only DB-derived shapes (same filter as the
  injection fetch, shared via refs.ts); any unresolved id fails loud
  with 400 — spoofed type/filepath metadata can no longer be persisted
  into assistant content or rendered in chat/share views
- Enforce agent authorization on /chat/steer against the ORIGINATING
  run's job identity: the chat path's role gate (AGENTS:USE, with the
  same non-agents-endpoint skip) plus the per-agent ACL check with the
  capability bypass — revoked access mid-run can no longer inject;
  cancel stays ownership-only (nothing model-bound)
- Mark steered uploads used after a successful enqueue (owner-scoped,
  best-effort) so the upload-window TTL cannot reap a file the
  persisted steer part references
- Consume the parked recovery copy after live delivery: converting
  final/abort/error pendingSteers fires one owner-gated claim-on-read,
  so dismissed chips can no longer resurrect on a later reload

* 🎙️ fix: Round-14 Composer-Context Fidelity + TTS and Queue-State Gaps

- Keep steer text out of generic assistant text extraction:
  parseTextParts excludes STEER parts by default with an includeSteer
  opt-in for the full-record surfaces (Meili indexing, aborted-response
  persistence) — TTS callers no longer speak the user's own mid-run
  words
- Mark queued uploads used at enqueue time via a minimal owner-scoped
  POST /files/usage (fail-closed without a user; upload limiters do not
  apply to a metadata touch), fired once wherever composer files enter
  the queued state — the upload-window TTL can no longer reap a file
  waiting out a long run or approval pause
- Carry quote chips and manual skill picks on queued items: captured
  and consumed from the composer at queue/interrupt time exactly like
  files, threaded through the drain and send-now overrides, and
  restored by the queued row's Edit message
- Key an early-aborted FIRST turn's run-end signal to NEW_CONVO
  (resolveRunEndTarget) so queued follow-ups stay visible on the
  restored new-chat composer instead of parking under an optimistic
  stream id the user never sees again

* 🧿 fix: Round-15 Gap Coverage + Consolidated Sweep (Share Leak, Abort Ids, Chip Hygiene)

- Run the resume steer-gap check for every still-active job: an empty
  snapshot no longer skips the re-peek, and synthesis now keys on the
  FRESH content view so an applied-in-gap steer that was never
  snapshotted still re-surfaces (over-emission is benign — applied-id
  dedupe, index-stable parts)
- Thread queued context through steer degradation: sendQueuedNow passes
  the item's quotes/skills into submitSteer, and every fallback
  (requeue or settled send) restores them instead of dropping to
  text+files
- Stop shared links from leaking steer attachment refs: the share
  snapshot now walks content — files-excluded shares strip steer-part
  files entirely; files-included shares sanitize and share-route them
  like top-level files (copy-on-write, non-steer content by reference)
- Seed pending-steer chips unconditionally on load/return so a steer
  applied while away cannot linger as a stale chip beside its part
- Use the abort response's resolved job id: chips/drain-signal land
  where the user actually is (NEW_CONVO for a new-held first turn,
  consistent with resolveRunEndTarget) while the parked-copy claim hits
  the resolved id instead of a no-op /chat/status/new
- Open steered documents like normal message files (FilePreviewDialog)
- Cap the applied-steer id set on the live path via a shared helper;
  kept surviving run end deliberately (late-ACK race depends on it) and
  fixed the atom comment that claimed otherwise

* 💡 fix: Un-light Steer Ribs When Their Node Is Replaced

Two stacked gaps kept a steer rib lit after scrolling away: the
pending→applied swap replaces the DOM node under the same id, which
produces no IntersectionObserver exit and — because the entry list
dedupes on (id, preview) — no entries change either, so the observer
kept watching a detached node; and the rail's mutation filter only
reacted to .message-render nodes, so steer-node swaps and removals
never triggered a refresh at all.

- reconcileObservedElements re-points the observer at replaced nodes
  from the mutation-driven refresh regardless of entries identity,
  dropping stale visibility until the fresh node reports (the observer
  fires its initial intersection immediately, so a truly visible part
  re-lights within a frame)
- The mutation filter now recognizes steer-render nodes alongside
  message rows

* 🪪 fix: Round-16 Recovery Owner Fields + Context Stickiness + Share Labels

- Park resumed-run leftovers with the manager facade's metadata owner
  fields: a bare job.userId is undefined on that shape, which made
  every parked payload from a resumed HITL run unclaimable
- Keep a queued item's quotes/skills sticky through a successful steer
  ACK: the pending chip carries them (client-only), reseeds preserve
  them across reconnects, and every terminal conversion — local or
  server-list, merged by steerId — restores them onto the queued item
- Convert resumeState.pendingSteers on the inactive status branch
  (deduped against unrecoveredSteers) so steers observed in the
  expired-pause-before-sweeper window convert instead of vanishing
  until a later reload
- Label shared steer parts share-safely via the existing ShareContext:
  a viewer's own name no longer appears on the sharer's steered
  messages

* ✂️ fix: Carry Steer Context Through the Failed-Chip Edit Action

Retry and convert-to-queue already preserve a failed steer's carried
quotes/skills; Edit message dropped them on the way back to the
composer. It now restores them through the same context path.
2026-07-14 10:11:10 -04:00
Danny Avila
520af663bc
🧵 feat: Background Tool Calls for Agents & Model Specs (#14197)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* 🧵 feat: Background Tool Calls for Agents & Model Specs

Opt-in, poll-based background tool execution. The model marks an eligible tool
call with `run_in_background: true`; the host executor registers a task, returns
a handle immediately (so the graph turn resolves), runs the tool as a detached
promise, and the model retrieves the result via a new `check_background_task`
poll tool. Host-side only — no `@librechat/agents` change.

- Opt-in mirrors `deferred_tools`: admin capability `run_in_background`
  (off by default) + per-tool `tool_options.run_in_background`.
- Model specs / ephemeral agents: `TModelSpec.runInBackground` /
  `TEphemeralAgent.run_in_background` synthesize per-tool options; both paths
  converge at `initializeAgent`.
- In-process task registry: scoped per user+conversation, idempotent by
  toolCallId (safe across resume/replay), capped, TTL-swept.
- Excludes direct-path / host-special / code-session tools. Subagents and push
  notifications are deferred follow-ups.

* 🩹 fix: Harden background tool calls (Codex review)

- Reliable per-agent execution gate: thread the injected `run_in_background`
  tool names from `initializeAgent` through `configurable.backgroundToolNames`
  (`toolRegistry` only reaches the executor for PTC/tool_search), fixing the
  silent no-op + unstripped-arg leak for ordinary event-driven tools.
- Enforce the per-tool opt-in at execution (`backgroundToolSet.has(name)`) so a
  non-opted-in tool can't be backgrounded via an extra arg.
- Gate the `check_background_task` interception on the run actually enabling
  background, so a user tool sharing that name still executes.
- Forward `backgroundToolsAvailable` to added-convo (multi-convo) agents.
- Exclude `web_search`/`file_search` from eligibility — their results are turned
  into user-visible attachments/citations only by the foreground toolEndCallback.

* 🩹 fix: Address Codex round 2 on background tool calls

- Idempotency scoped to run+turn: provider tool-call ids repeat across turns
  (e.g. `call_0`), so key the dedupe map by `runId::toolCallId` and sweep
  orphaned mappings — a later turn no longer collides with a retained task.
- Artifacts preserved: a backgrounded tool's artifact is processed through the
  same `toolEndCallback` as the foreground path (images/files/citations no
  longer silently dropped), best-effort/guarded.
- Forward the `run_in_background` capability to connected-agent discovery and
  subagent `processAgent` init, so a child agent's own event-driven tools work
  the same as when it runs as primary.
- Strip the injected flag on foreground calls of background-capable tools
  (the model may emit it as `false`) so strict MCP/action schemas don't reject.
- `check_background_task` list path returns metadata only (result_available /
  result_chars), never full results — prevents context overflow; the full
  result is returned only when a specific id is requested.

* 🩹 fix: Address Codex round 3 on background tool calls

- Exclude background-capable tools from eager execution (run.ts): a speculative
  eager dispatch of a `run_in_background` call could launch the detached task
  with partial/stale args, and that side effect can't be canceled.
- Reserve the `check_background_task` name: overwrite a colliding user/MCP tool
  with the host poll schema (with a warning) so the advertised schema matches
  the executor's interception instead of hijacking a mismatched tool.
- Don't inject background schemas into pure subagents (spawn-tool child graphs)
  whose tools don't reach the host interceptor; keep it for primary/added/
  connected agents. Subagent background is the durable follow-up.
- Thread `backgroundToolsAvailable` + `backgroundToolNames` through the
  OpenAI-compatible and Responses agent routes (was chat-only), so the same
  agent/model spec behaves consistently across surfaces.
- Exclude image-generation built-ins (dalle/flux/gemini_image_gen/image_gen_oai/
  image_edit_oai) — artifact-first tools whose files can't reliably attach to an
  already-saved turn when backgrounded.

* 🩹 fix: Address Codex round 4 on background tool calls

- Sanitize self-spawn subagent inputs: strip `run_in_background` + the
  `check_background_task` def from the parent AgentInputs reused for self-spawn,
  so the isolated child (direct/child-graph path) doesn't advertise a background
  schema it can't honor. The SDK resolver keeps a provided `agentInputs` even
  with `self: true`.
- Exclude `check_background_task` from PTC (`run_tools_with_code`) tool
  definitions — it's host-only and not callable from generated code.
- Parse stringified JSON args before deciding background dispatch and before
  stripping the flag, so string-delivered `run_in_background` is honored and
  never leaks to strict object-schema tools.
- Skip injection for tools that already declare their own `run_in_background`
  param (would otherwise hijack/strip it), and for non-object (string-input)
  schemas (would otherwise rewrite the input contract).

* 🩹 fix: Address Codex round 5 on background tool calls

- check_background_task now parses stringified JSON args, so providers that
  deliver args as a string can retrieve a specific task by id (not just list).
- Include agentId in the background dedupe key (`agentId::runId::toolCallId`):
  two agents in the same run emitting the same provider id (e.g. `call_0`) now
  launch independent tasks instead of colliding.
- Self-spawn sanitization also strips the background entries from the reused
  toolRegistry (not just toolDefinitions), so a child using tool_search/deferred
  loading can't rediscover the host-only run_in_background / check_background_task.

* 🩹 fix: Strip run_in_background from PTC target tool schemas (Codex round 6)

The PTC path already filtered out the host-only check_background_task poll tool
but still exposed target tool schemas with the injected `run_in_background` param
(the shared toolRegistry entries were mutated by applyBackgroundToolCalls). PTC
codegen doesn't go through the host background interceptor, so it could pass the
flag to an MCP/action tool (strict-schema rejection or silent foreground with no
poll). Sanitize the PTC toolDefs like the self-spawn path does.

* 🩹 fix: Sanitize background from explicit subagent inputs (Codex round 7)

A child agent reachable as a top-level/handoff agent is initialized WITH the
background capability, then reused as an explicit subagent via buildSubagentConfigs.
Round 4 only sanitized the self-spawn case; this now applies the same
stripBackgroundFromToolDefinitions/Registry to explicit child agentInputs when
`child.backgroundToolNames` is non-empty, so an isolated child graph doesn't
advertise a run_in_background / check_background_task contract it can't honor.

* 🩹 fix: Reap stuck/expired background tasks (Codex round 8)

- get() now sweeps before returning, so repeatedly polling a known
  background_task_id can't keep an expired completed task (and its retained
  result, up to 100k chars) alive past the one-hour completed TTL.
- sweep() now reaps `running` tasks older than a 30-min running TTL, marking
  them errored. Previously a detached call that never settled (hung network /
  lost MCP connection) held a running slot forever, exhausting the
  per-conversation cap and rejecting every later dispatch.

* 🩹 fix: Evict oldest settled tasks instead of blocking at the cap (Codex round 9)

Only the running-task cap gates dispatch now. The total-tasks cap
(MAX_TASKS_PER_BUCKET) bounds memory but no longer rejects new background calls:
when full, it evicts the oldest settled (completed/error) tasks to make room.
Previously 200 quick background calls in one conversation would block all new
dispatches for up to the completed-task TTL, since polling doesn't remove settled
tasks. Running is already capped, so room always frees.

* 📝 docs: Frame background tool calls as within-turn (Codex P1 contract)

Codex escalated the request-lifecycle findings to P1 on the grounds that the
advertised "poll later" contract can't be honored for genuinely long-running
calls (request-scoped MCP connections + the run abort signal are torn down at
turn end). Align the model-facing contract with what the same-run implementation
actually delivers: the run_in_background param, check_background_task, and the
dispatch handle now instruct the model to collect the result WITHIN THE SAME TURN
(backgrounded work isn't guaranteed to survive past the turn). This is
within-turn parallelism; cross-turn survival of long-running calls remains the
deliberate durable subagent follow-up. Copy/comment-only; no behavior change.

* ♻️ refactor: Cross-turn background tool calls, leak-free

Extend background tool calls from within-turn to cross-turn on a single
process, since the mechanism already supports it: the run's abort signal
never reaches the detached invoke (the graph forwards only configurable/
metadata to the tool-execute handler), so the floating promise keeps
running past turn completion and its result stays in the in-process
registry for a later turn to poll (get/list key only on
user::conversation + id, never the dispatch run/turn).

Guarantee no connection leak: ephemeral request-scoped MCP tools (runtime
{{LIBRECHAT_BODY_*}} placeholders) capture their request-scoped store at
creation and fall back to it, so config manipulation can't redirect them;
their connection is torn down at request end. Tag such tools in
createToolInstance and run them in the foreground instead of backgrounding
them. Pooled/app-level MCP and structured tools are unaffected and survive
cross-turn via their managed pools.

Reword the model-facing contract (run_in_background, check_background_task,
handle message, fileoverview) from within-turn to cross-turn on this server
(not across restart/replica, which stays the durable follow-up).

Tests: cross-turn poll retrieval; ephemeral MCP tool runs foreground.

* 🐛 fix: Guard ephemeral MCP tag against a null server config

createToolInstance can be reached with a null/stale capturedServerConfig
(cached availableTools + getServerConfig returns null, as several MCP unit
tests construct tools). The new unconditional requiresEphemeralUserConnection
call then dereferenced config.source and threw during tool construction
(CI: Tests api shard 2/3). Guard with the same serverConfig ? ... : false
pattern the other callers use; a missing config is not request-scoped.

* 🎨 fix: Deliver backgrounded tool artifacts on the poll turn

A slow backgrounded MCP/action tool resolves after its dispatch turn is
finalized: createToolEndCallback only appends to that turn's artifactPromises
(already awaited) and writes to a closed stream, so the artifact (file/citation/
UI resource) was silently dropped — check_background_task recorded only the
hasArtifact boolean. The cross-turn contract made this the common case.

Hold the artifact on the task and deliver it through the LIVE poll turn's
toolEndCallback the first time check_background_task collects that id (once,
then cleared to free memory), attributed to the original tool. Same-turn and
cross-turn now share this path since the model must poll to collect any result.

Tests: registry claim-once; artifact delivered on poll not dispatch, idempotent.

*  feat: Agent-builder toggle for background tool calls + cap tool descriptions

Add a per-MCP-tool "run in background" toggle in the agent builder, mirroring
the programmatic/deferred pattern: gated on the admin `run_in_background`
capability via useAgentCapabilities, read/written on tool_options[id]
.run_in_background through useMCPToolOptions (per-tool + bulk mark-all), and
rendered as a Zap toggle in MCPToolItem and McpSection with new locale keys.

Also cap the section tool/server descriptions (McpSection, ToolSection,
SkillSection) with max-h-40 overflow-y-auto so a long description scrolls
instead of overflowing the dialog, matching MCPToolItem's existing cap.

Tests: MCPToolItem renders/toggles the background button only when enabled.

* 🧪 fix: Mock new background hook functions in McpSection spec

* 🎨 fix: Restore background artifact when poll-turn delivery fails

* 🛡️ fix: Harden background tool call edges from review findings

- Error immediately (matching foreground) when a background-requested tool
  failed to load, instead of returning a success handle for a dead task
- Exclude ephemeral request-scoped MCP tools at injection time so the model
  never sees a run_in_background param the executor would silently downgrade;
  flip the execute-time tag to fail closed on a missing server config
- Source image-tool background exclusions from the shared imageGenTools set
  (adds missing stable-diffusion, an artifact-first live tool) instead of a
  hand-copied list
- Add check_background_task to the eager-execution exclusion list: artifact
  collection is a one-shot claim that must not fire from a speculative
  snapshot the SDK may discard
- Strip an imitated run_in_background arg on tools the executing agent never
  opted in (multi-agent history bleed), unless the tool's own schema declares
  the parameter
- Truncate oversized stored results with an explicit marker via the shared
  truncateMiddle (moved to utils/text) instead of a silent slice
- Document the at-most-once artifact delivery semantics honestly (the
  callback's downstream persistence is fire-and-forget, as in foreground)

* ♻️ refactor: Deduplicate background tool-call plumbing and tighten types

- Use the SDK's JsonSchemaType instead of a local duplicate; drop all
  as-unknown casts and type the poll-tool serializer explicitly
- Drop derivable BackgroundTask state (progress, hasArtifact) and the dead
  `enabled` param/return on applyBackgroundToolCalls (guarded at the call
  site), which also skips the defs pass when nothing opted in
- Fold the enable expression into synthesizeBackgroundToolOptions so the
  three load/added call sites can't drift
- Throttle the registry's all-buckets sweep and always sweep the accessed
  bucket, so a hot poll loop is no longer O(total tasks server-wide); bound
  retained artifact memory with a size cap
- Single-pass stripBackgroundFromToolDefinitions; pass metadata through to
  the poll-turn callback instead of a no-op reconstruction
- Collapse the client's copy-pasted boolean option families into a keyed
  factory (also removes the shared-object mutation in the bulk toggles) and
  the six toggle-button copies into one OptionToggle component

* 🧪 test: e2e coverage for cross-turn background tool calls

Proves the full contract through the real pipeline (mock harness): an agent
opts an MCP tool in via tool_options.run_in_background, the model dispatches
it detached and receives the synthetic handle while the tool is still running
(status=running in the rendered ack — the non-blocking guarantee without
timing assertions), the tool completes after its turn finalized, and a later
user turn recovers the task id from replayed history, polls
check_background_task, and renders the collected result.

- fake-mcp-server: slow_echo fixture tool (delayed echo)
- fake-model: E2E_BACKGROUND_DISPATCH / E2E_BACKGROUND_COLLECT markers
- e2e yaml: agents capabilities = defaults + run_in_background

* 🔧 fix: Close two background capability gaps from review

- Thread backgroundToolsAvailable through the OpenAI-compatible service
  (derived from app capabilities like codeEnvAvailable/statefulSessions),
  so agents with tool_options.run_in_background keep the feature on that
  route; fold the three capability derivations into one helper
- Index ephemeral MCP servers by normalizeServerName when excluding tools
  from background injection: tool names embed the normalized server name
  while mcpConfig keys the original, so exotic server names previously
  escaped the injection-time exclusion

* 🛂 fix: Fall back to configurable user identity for background task scoping

The in-repo routes merge req into the tool-execute configurable, but external
hosts of the exported OpenAI-compatible service inject their own loadTools and
may not — tasks would then register under an empty user id, collapsing
registry isolation to conversationId alone. Resolve the scoping id from
req.user.id, then configurable.user_id / user, and cover the isolation with a
foreign-user not_found test.

* 🧹 chore: Apply repo import sorter to PR-touched files
2026-07-13 12:51:36 -04:00
Danny Avila
53e369fba8
🧪 feat: stateful_code_sessions capability for warm Code API sandbox sessions (experimental) (#14150)
*  feat: stateful_code_sessions capability for warm Code API sandbox sessions

Wire the @librechat/agents stateful sandbox sub-config behind a new,
off-by-default stateful_code_sessions agent capability. createRun sets
toolExecution.sandbox.statefulSessions when code execution is active in the
run AND the capability is enabled; execute_code and bash_tool factories get
the param so their descriptions hedge toward persistence. Rides the existing
variable-not-literal runConfig pattern, so it no-ops until @librechat/agents
is bumped to the version shipping the sandbox sub-config.

*  feat: per-agent stateful code sessions (builder toggle + init gating)

Stateful sessions now require the agent's own opt-in, not just the admin
capability. New agent field stateful_code_sessions (schema + validation +
types) surfaces as a toggle in Agent Builder Advanced settings, gated on
the app capability and disabled without Code Interpreter. initializeAgent
resolves the per-agent truth (admin capability AND builder opt-in AND
code env) once: the registered bash_tool description, the execute_code
factory, and createRun's toolExecution.sandbox gate all read the same
resolved value. statefulSessionsAvailable threads through the same call
sites as codeEnvAvailable, including handoff discovery and added convos.

* 🐛 fix: propagate runtime_session_hint to sandbox executor in event-driven tool path

The event-driven ON_TOOL_EXECUTE handler built config.toolCall without the
resolved runtime_session_hint, so BashExecutor/CodeExecutor never sent
runtime_session_hint to the Code API. Every conversation then collapsed onto
the server-derived default session (no per-conversation isolation). Copy
tc.runtimeSessionHint onto toolCallConfig._runtime_session_hint, mirroring the
SDK direct-execution path.

* 🐛 fix: address Codex review findings for stateful code sessions

- OpenAI-compatible service (packages/api/src/agents/openai/service.ts) now
  derives and passes statefulSessionsAvailable alongside codeEnvAvailable, so
  the feature activates on that route (previously statefulCodeSessions resolved
  false there and createRun never sent toolExecution.sandbox).
- Thread runtime_session_hint through the host file-authoring tools
  (create_file/edit_file/read_file): those host branches return before the
  generic tool path, so readSandboxFile/writeSandboxFile now forward the
  per-conversation hint instead of falling back to the Code API default session.
- StatefulSessions builder toggle clears its form value when Code Interpreter is
  disabled, so a saved agent matches the disabled UI and re-enabling code doesn't
  silently reactivate stateful sessions.

* 🐛 fix: normalize stateful_code_sessions on save when Code Interpreter disabled

Addresses Codex review (round 2): a stale `stateful_code_sessions` opt-in
could persist when Code Interpreter (`execute_code`) is disabled from the
main agent builder without opening Advanced settings, silently reactivating
warm sessions if code was later re-enabled.

- AgentPanel: normalize in `composeAgentUpdatePayload` (the always-run save
  path) so `stateful_code_sessions` is forced to `false` whenever
  `execute_code !== true`, regardless of whether Advanced was opened.
- StatefulSessions: revert the mount-scoped useEffect (round-1 approach) —
  it only fired while the Advanced panel was mounted, missing this path.
- Add spec coverage for both branches of the normalization.
2026-07-12 08:12:04 -04:00
adamscross04
4182f9094f
🃏 fix: Attach Request-Scoped MCP Servers From the Builder via the mcp_all Wildcard (#14177)
* fix: Attach Request-Scoped MCP Servers from the Agent Builder via mcp_all

Follow-up to #14148 / #14074: request-scoped MCP servers (runtime
{{LIBRECHAT_BODY_*}} placeholder headers) defer their connection on
reinitialize, so their tools are never enumerable in the agent builder
and the attach flow (which waits for isConnected && hasTools) silently
attaches nothing. The runtime already resolves an mcp_all
(sys__all__sys_mcp_<server>) tool entry into the server's full tool set
at chat-turn time - the builder just never writes that token.

- reinitMCPServer returns connectionDeferred: true on the deferred
  branch so clients can distinguish it from a plain empty success
  (server configs are sanitized client-side, so the response is the
  only reliable signal)
- /mcp/:serverName/reinitialize forwards the flag; data-provider
  mutation type includes it
- McpSection attaches [mcp_server, mcp_all] tokens on a deferred
  connect (idempotent) and shows a "tools are resolved at runtime"
  hint instead of "no tools yet" when wildcard-attached
- selectors: mcpAllToken() helper beside mcpServerToken()

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address review — deferred attach via init state; strip stale wildcard

Two review findings:

1. Servers with customUserVars route Connect through the config dialog,
   whose save path calls initializeServer inside the manager — the
   McpSection never awaits that response, so the deferred attach was
   unreachable. Record connectionDeferred in the shared per-server init
   state (MCPServerInitState) on every initialize attempt and key the
   attach off that state in the auto-select effect: one attach site now
   covers both the direct Connect and the config-dialog path.

2. updateFormTools kept an existing mcp_all wildcard when rewriting a
   per-tool selection, so a server that later exposes a normal tool list
   would still grant every tool at runtime while the UI showed a subset.
   The wildcard is now stripped unless explicitly re-passed, making
   per-tool selection always supersede it.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: address review — stale deferred state; fold wildcard into display

Second review round:

1. connectionDeferred persisted across attempts, so a later Connect
   click could attach the wildcard from a stale flag before the new
   attempt reported. Reset it at the start of every initializeServer
   call, and clear it before routing into the customUserVars config
   dialog (resetConnectionDeferred) so only the current attempt's
   outcome can trigger the auto-attach effect.

2. With a wildcard attached and the server's tools later enumerable,
   the dialog showed every tool unchecked while runtime granted all of
   them. getSelectedTools now folds the wildcard into the display (all
   tools selected); any selection interaction rewrites the form with
   concrete ids and drops the wildcard, converting the attachment on
   first touch.

Also sorts imports in McpSection.tsx (CI sort-imports gate).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-12 08:10:01 -04:00
Danny Avila
b753da163e
🤖 feat: Add GPT-5.6 (Sol/Terra/Luna) OpenAI Models (#14206)
*  feat: Add GPT-5.6 (Sol/Terra/Luna) OpenAI models

Adds the GPT-5.6 family (GA 2026-07-09) across the context, output,
pricing, cache, premium, and default-model maps, mirroring gpt-5.5.

- gpt-5.6 (Sol alias), gpt-5.6-terra, gpt-5.6-luna
- 1.05M context / 128K output for all tiers
- Standard + long-context (>272K input) tiered pricing and cache rates

* 🐛 fix: Bill GPT-5.6 cache writes at documented 1.25x input surcharge

OpenAI prices GPT-5.6 cache writes above the base input rate (Sol $6.25,
Terra $3.125, Luna $1.25 per 1M vs $5/$2.50/$1 input). Correct the
cacheTokenValues write rates so explicit prompt-caching usage is billed
and reported accurately, and lock the surcharge with a test.

*  feat: Expose GPT-5.6 max reasoning effort + long-context cache premium

Folds in the two deferred Codex findings:

1. Add `max` to the OpenAI `ReasoningEffort` enum and the reasoning_effort
   parameter options/labels so GPT-5.6 (Sol/Terra/Luna) can request its
   documented highest reasoning setting. Backend passthrough and zod
   validation pick it up via the nativeEnum schema.

2. Apply the long-context (>272K input) premium to cache tokens. Adds
   `premiumCacheTokenValues` + `getPremiumCacheRate`, threads
   `inputTokenCount` into `getCacheMultiplier`, and wires it through both
   structured-spend paths. Covers the gpt-5.4/5.5/5.6 family whose cache
   write/read previously stayed at flat base rates on long-context calls.

* 🐛 fix: Bill GPT-5.6 cache writes + map max effort for OpenRouter Claude

Addresses Codex round-3 findings:

1. (P1) splitUsage only read `cache_creation`/`cache_creation_input_tokens`,
   so OpenAI GPT-5.6's `cache_write_tokens` fell into inputOnly and billed at
   the input rate instead of the 1.25x write rate. Extend UsageMetadata and
   single-source the cache-creation read to also recognize `cache_write_tokens`
   (nested and top-level).

2. (P2) `max` was exposed via OpenRouter (spreads OpenAI settings) but the
   adaptive-Claude verbosity map had no `max`, so it was silently dropped. Map
   max -> 'max' verbosity.

* 🐛 fix: Forward GPT-5.6 cache_write_tokens into emitted usage

Local Codex review (P1): the cache-write fix reached balance billing
(splitUsage/getCacheCreationTokens) but not the emitted-usage pipeline.
ModelEndHandler built the emitted event's cache_creation from only
cache_creation/cache_creation_input_tokens, so GPT-5.6 cache_write_tokens
were dropped and aggregateEmittedUsage classified them as ordinary input —
displayed/persisted cost undercounted and disagreed with the balance charge.
Fold cache_write_tokens (nested + top-level) into the emitted cache_creation.
2026-07-12 07:52:59 -04:00
Danny Avila
6e8f44d72b
🪆 fix: Serve Nested Skill Files Through a Wildcard Route (#14191)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Has been cancelled
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
Nested skill files (references/, scripts/, assets/) returned 404 when
opened, while top-level SKILL.md worked. The file routes matched a single
`:relativePath` segment, so a nested path only resolves when the encoded
`%2F` reaches Node intact. Reverse proxies (nginx, Traefik, k8s ingress,
ALB, Cloudflare) commonly normalize `%2F` to a literal slash before
forwarding, which turns the path into multiple segments the single-segment
route can no longer match.

Switch the GET/DELETE file routes to an Express 5 splat (`*relativePath`)
and reconstruct the path from the decoded segments, so nested files resolve
whether the client sends `%2F` or a proxy decoded it to a literal slash.

Security: add a shared `isSafeSkillFilePath` validator (extracted from the
import validator, single source of truth) and reject traversal/absolute/
empty-segment paths at the route layer. Lookups remain exact DB matches, so
the user-supplied path never touches the filesystem.

Also drops a latent double-decode (Express already decodes route params).

Closes #14190
2026-07-09 16:52:11 -04:00
Danny Avila
3945d293de
🗂️ feat: Per-Agent Memory Partitions (#14084)
* feat: per-agent memory partitions (memory_scope)

Adds an optional agentId partition to MemoryEntry so agents can opt into
isolated memory via a new memory_scope field ('user' | 'agent'). Partition
derives from agentId presence ({agentId: null} matches legacy docs, no
migration). Inline set_memory/delete_memory tools, the post-turn memory
agent, the request-scoped memory cache, and context injection are all
partition-aware; context is only injected into agents whose resolved
partition matches. Memory routes accept the partition param, scope
duplicate/token-limit checks per partition, and enrich entries with agent
names. Memories panel gains a partition filter and agent badges; the agent
builder gains an agent-scoped memory toggle.

* fix: address Codex review findings on memory partitions

- strip runtime ____N id suffixes in getMemoryAgentId so added-conversation
  runs share the persisted agent's partition
- load each agent's own partition in multi-agent context injection instead
  of skipping foreign partitions entirely
- clear memory_scope to 'user' on save when Enable Memory is unchecked
- fall back to 'all' when the selected panel partition no longer exists
- restrict GET /memories agent-name resolution to agents the requester can
  VIEW
2026-07-09 10:48:51 -04:00
Cha
73c43ded25
💂 fix: Enforce ALLOW_EMAIL_LOGIN on the Backend Login Route (#14180)
* 🔒 fix: Enforce ALLOW_EMAIL_LOGIN on Backend Login Route

ALLOW_EMAIL_LOGIN=false previously only hid the login form; POST
/api/auth/login stayed mounted and accepted valid credentials. Add a
validateEmailLogin middleware (mirroring validateRegistration /
validatePasswordReset) that rejects login with 403 when the flag is
disabled, with an ALLOW_EMAIL_LOGIN_OVERRIDE escape hatch for
intentional direct API login (each use logged with request IP).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: Gate admin local login by email login flag

* fix: Move email login gate into api package

* test: Avoid mutating readonly request ip

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-09 09:55:36 -04:00
VictorEPlus
9446f3278c
🔡 fix: Normalize Email Case When Issuing Verification Tokens (#14172)
The user schema lowercases emails on save, but sendVerificationEmail
stored the verification token with the raw registration email. For
mixed-case registrations, verifyEmail's lookup by the lowercased DB
email never matched the token, failing with 'No email verification
data found'. Normalize once and use it for the link, recipient, and
token. resendVerificationEmail already used the DB email and was
unaffected.

Co-authored-by: VictorEPlus <victor.ortega@eplusadvisor.com>
2026-07-09 08:42:39 -04:00
Danny Avila
988a14a405
🙋 feat: ask_user_question - agent-initiated questions with durable pause/resume (#14139)
* feat: ask_user_question tool — agent-initiated questions with durable pause/resume

The HITL runtime merged in #13942/#14024/#14025/#14123 already ships the full
ask_user_question lifecycle (payload-agnostic handleRunInterrupt, resume
validation via mapAskUserAnswer, reconnect rehydration, and the client question
card) — but nothing ever raised the interrupt. This adds the producer:

- packages/api/agents/hitl/askUserQuestionTool.ts: LLM-callable tool whose func
  calls the SDK askUserQuestion() helper (LangGraph interrupt() from the tool
  body); zod schema with length caps mirroring AskUserQuestionRequest, plus a
  JSON-schema twin for the schema-only registry
- Registration: agentToolDefinitions, manifest.json (Tools dialog, admin
  filteredTools/includedTools kill switch), basicToolInstances, handleTools
  constructor branch
- run.ts gating: checkpointer now attaches for hitlCapable runs whose agents
  carry the ask tool even with the tool-approval policy disabled (the interrupt
  needs only durability, not humanInTheLoop/hooks); the tool is stripped
  fail-closed from non-HITL callers (OpenAI-compat/Responses) and subagent
  child configs; excluded from eager event execution (interrupts must be
  raised inside the Pregel task frame)
- resume.js: 16k length cap on the answer wire field
- e2e (real Run + FakeChatModel + LazyMongoSaver + supertest resume): tool-body
  interrupt pauses durably with NO approval policy, answer round-trips as the
  ToolMessage content, tool body re-runs once on resume, sequential questions
  re-pause

* fix: adversarial-review findings — in-graph execution, orphan prunes, endpoint scoping, real kill switch

Pre-PR multi-agent review confirmed 5 defects in the initial commit; all fixed:

1. CRITICAL — the tool never paused on the real agents endpoint: production
   loads tools definitions-only, flipping the SDK ToolNode to event-driven
   dispatch, and the host ON_TOOL_EXECUTE handler runs outside the Pregel task
   frame (under runOutsideTracing), where interrupt() throws and becomes an
   error ToolMessage. Reworked: the ask tool never rides toolDefinitions/
   toolRegistry — on HITL-capable top-level agents a real instance is supplied
   via AgentInputs.graphTools (agents#289, requires @librechat/agents > 3.2.57),
   the SDK's in-graph direct-tool seam; new production-shape e2e pins the
   event-driven mode end to end.
2. CRITICAL — ask-only runs left orphaned interrupted checkpoints (silent
   context duplication on every later turn): both orphan prunes were gated on
   toolApproval.enabled. The pre-turn prune now also fires for ask-capable
   agents (exported agentRequestsAskUserQuestion), and the abort-route prune
   fires when the aborted job carries a pendingAction.
3. MAJOR — self-spawned subagents bypassed the strip (self config resolves from
   the parent's _sourceInputs): fixed SDK-side (buildChildInputs clears
   graphTools) and the tool is now never present on child surfaces host-side.
4. MINOR — the manifest entry leaked into the Assistants tools dialog and the
   legacy plugins endpoint, where tools execute with no run to pause: new
   agentsOnly manifest flag, scoped out of both listings.
5. MINOR — filteredTools/includedTools only hid the tool from the dialog:
   now enforced at run build (strip + no checkpointer), making the admin
   filter a real kill switch for already-saved agents.

* chore: update @librechat/agents dependency to version 3.2.58 in package-lock.json and package.json files

* fix: reject agents-only tools at assistant create/update (Codex round 1)

The tools-dialog scoping keeps ask_user_question out of the assistants
LISTING, but the v1/v2 create/update handlers resolve arbitrary posted tool
strings from the shared getCachedTools map — a REST client or stale saved
payload could still attach it, and the assistants runtime executes tools with
no run to pause, so every call would error. New isAgentsOnlyTool(tool)
(manifest-driven, handles string and function-object shapes) drops such tools
with a warn at all four resolution sites (v1+v2, create+update).

* fix: offset resumed-run content indices past the pre-pause seed

A resumed run rebuilds the graph from the checkpoint, and the fresh graph
numbers content indices from its own empty contentData — starting at 0. The
resume path seeds the (also fresh) content aggregator with the pre-pause
parts at exactly those indices, so the resumed model turn collided with the
seed: type-matching parts silently MERGED (post-resume text appended into a
pre-pause text block), and type-mismatching parts (a reasoning/think part at
index 0 — any Anthropic reasoning agent) dropped EVERY delta with 'Content
type mismatch', losing the entire post-resume output from the live stream
and the saved message.

Latent since #13942 — tool-approval resumes corrupt content the same way
(probe-verified); it surfaced now because ask_user_question makes pausing a
first-class flow and reasoning models make the loss total.

- createContentIndexOffsetHandlers(handlers, offset): wraps ON_RUN_STEP
  (the single point where a content index enters the pipeline — deltas
  resolve through the aggregator's stepMap) and ON_AGENT_UPDATE's inline
  index; every other handler passes through by reference. Probe-validated:
  resumed output now lands as a new part after the paused tool call.
- resumeCompletion wires it with offset = seedContent.length.
- logToolError: a GraphInterrupt unwinding out of a tool body is the HITL
  pause working as designed — no longer logged as a Tool Error.

* fix: unblock live streaming of the resumed segment after an answer

With resume indices now ABSOLUTE (server continues after the pre-pause
parts), the synthetic ask-user-question card was squatting on exactly the
index the resumed segment streams into: applyAskUserQuestion appends the
card at the end of the message content, so on the answering device every
incoming part at that index was blocked and nothing rendered between the
answer submission and the finalize replacing the message.

removeAskUserQuestionPart(message, actionId) strips the pause-scoped card
on successful answer submission (useResumeSubmit onSuccess) — the durable
record of the Q&A is the ask_user_question tool call itself. Pure helper +
specs; same-reference no-op when nothing matches.

* fix: displace the synthetic question card in the streaming content writer

The store-level strip on answer submit wasn't enough: the SSE step handler
keeps its own in-flight copy of the streaming message, so on the answering
device the synthetic ask-user-question card still occupied the ABSOLUTE index
the resumed segment streams into — every delta warned 'Content type mismatch'
(existing ask_user_question vs incoming text) and nothing rendered between the
pending_action and finalize.

Displace the card inside updateContent when any real part claims its slot —
the same displacement pattern as the OAuth prompt part directly above it.
Covers the streaming handler's own copy, reconnecting tabs, and other devices;
once real content streams, the pause is over by definition. Spec drives a
runStep + text delta into the card's index and pins: no mismatch warn, card
gone, text rendered.

* feat: dedicated UI + durable data for completed ask_user_question calls

The completed ask call rendered as a generic tool card labeled 'Cancelled'
with raw (and empty) JSON args. Two layers fixed:

Data: the saved tool_call part had args:'' and no output — streamed arg
chunks carry no tool name so the aggregator drops them (normal tools recover
via the completion event, which never fires for a tool that interrupts
mid-execution and resumes on a rebuilt run with no step id). The resume
controller now stamps the paused ask part with the pendingAction's
authoritative question as args and the user's answer as output
(attachAskUserQuestionAnswer — pure, targets the newest unanswered ask part,
so sequential questions each keep their own answer).

UI: Part.tsx routes ask_user_question tool calls to AskUserQuestionCall — a
compact Q&A record ('Asked a question' header, question, description, 'You
answered: <label>' preferring the picked option's label, or 'No answer was
given' for an abandoned pause) instead of the generic card. New i18n keys;
parseAskUserQuestionArgs degrades to null on malformed model args.

* fix: single question UI per pause + immediate answer display

Two live-turn issues with the new durable Q&A card:

1. Duplicate question on ask: during a live pause the message carries BOTH the
   ask tool_call part (now rendered by AskUserQuestionCall, showing a
   misleading 'No answer was given' while paused) and the synthetic
   interactive card. The durable card now defers while the turn is live and
   unanswered (isSubmitting) — the interactive card owns the question UI until
   it's answered; an abandoned pause still shows its no-answer state once the
   turn settles.

2. 'No answer was given' after answering: the server stamps the answer onto
   the part at resume seed, but the client only received that at finalize. No
   stream emission needed — the client knows the answer it just submitted:
   resolveAskUserQuestionPart (replacing the plain strip on submit success)
   removes the synthetic card AND stamps output/progress onto the newest
   unanswered ask tool_call, seeding args from the synthetic part's question
   when the streamed args were lost — mirroring the server-side
   attachAskUserQuestionAnswer, so the Q&A record shows the answer the moment
   the user submits.

* fix: keep the Q&A record visible while the resumed segment streams

The optimistic output stamp lives in the message store, but the SSE step
handler evolves its own cached copy of the streaming message (created at turn
start) — the first resumed event overwrites the store with that copy, wiping
the stamp, so the Q&A card blinked out during streaming and only returned at
finalize.

Render-layer fallback instead of fighting the handler's copy: submitted
answers are recorded by ask tool_call id when resolveAskUserQuestionPart
stamps the part, and AskUserQuestionCall reads the recorded answer whenever
the part's own output is missing — the record survives any message-copy churn
until finalize delivers the server-stamped part.

* feat: present Ask User as a native builtin in the tools dialog

It ships with the app and pauses the run like a first-class feature, so it
belongs with the builtins (Run Code, Web Search, Memory, ...) rather than in
the third-party plugin list — while its mechanics stay exactly a plugin's:

- BuiltinId += 'ask_user_question' (documented exception: a native TOOL, not
  a capability; selection reads agent.tools, the toggle emits tool-add/remove
  patches instead of a capability field)
- buildCatalog surfaces it as a builtin gated on the same signals as before
  (tools capability on + the server lists the plugin, i.e. not admin-filtered)
  and skips it in the plugin loop so it never double-lists
- On-theme icon: lucide MessageCircleQuestion in a teal chip via the builtin
  icon map, matching the other native entries; the bespoke purple SVG and the
  manifest icon field are gone
- i18n'd name/description keys like the other builtins

* feat: composer popover for answering questions (mentions-style)

Answering moves to the composer, matching the existing mentions/prompts
popover pattern: while an ask_user_question pause is live, a popover anchors
above the textarea with the question as its header, numbered option rows
(hover/click, or ↑/↓ + Enter from the empty composer), and an × to dismiss.
The main textarea doubles as the free-form answer — its placeholder flips to
'Something else...' and form submit routes the text to the paused run as the
answer instead of starting a new turn. Dismissing (× or Escape) restores
normal sends; the inline transcript surfaces stay as before (interactive card
while paused, durable Q&A record after) so the question remains visible in
history.

- findLiveAskUserQuestion (pure, spec'd): newest unanswered synthetic part
  across the conversation IS the popover signal — applied on
  on_pending_action, stripped on answer submit, so visibility tracks the
  pause lifecycle with no extra state
- useLiveAskUserQuestion hook shared by the popover and ChatForm; dismissals
  in a recoil atom so both react
- popover only mounts on the primary composer (index 0), mirroring QuoteButton

* feat: number-key selection + return glyph in the question popover

Pressing 1-9 in the empty composer picks the matching option directly,
mirroring the numbered row chips; the highlighted row shows a return-key
glyph as the Enter affordance. Same empty-composer guard as the arrow keys —
typing a free-form answer is never intercepted.

* refactor: first-class composer answer mode (useAskAnswerMode)

Replaces the bolted-on integration (inline onSubmit interception + raw
capture-phase keydown listeners on the textarea ref) with a single hook that
owns the whole answer mode: live-question derivation, dismissal + highlighted
option (shared recoil state), option selection, free-form submit routing
(submitText returns whether it consumed the submission), and keyboard
handling (handleKeyDown returns whether it consumed the key, composed ahead
of the textarea's normal handler — no more addEventListener).

The popover is now pure rendering off the hook; ChatForm wires placeholder,
onKeyDown, and onSubmit through the same instance. Deliberately scoped to the
composer rather than useSubmitMessage: starters/prompt-commands keep new-turn
semantics (and the existing job-replacement behavior while paused).

* fix: Codex round 2 — inline answer input, approval exemption, pause-time args

F1 (composer submit unreachable while paused — isSubmitting keeps Stop shown
and useTextarea eats Enter): redesigned around it, borrowing Claude Code's
AskUserQuestion semantics. The popover now owns free-form input via an inline
'Other' row (numbered last, 'Something else…'), with select-then-confirm rows
(click/arrows/digits highlight; Submit ↵, Enter, or double-click fires; Skip
dismisses). The composer returns to being a plain composer — no placeholder
swap, no submit interception; Stop keeps meaning stop.

F2: ask_user_question is exempt from the tool-approval prompt unless the
admin explicitly lists it (allow/ask/deny all win) — approving the right to
ask a question was a pure double pause; the tool is side-effect-free.

F3: the question is stamped onto the paused ask tool_call's args at PAUSE
time (attachAskUserQuestionArgs in handleRunInterrupt), so abandoned/expired/
stopped turns persist with the question intact and the record card can render
it — previously only the answer-resume path stamped args.

* fix: fold model-supplied 'Other' options into the inline free-form row

The model can generate its own catch-all option ('Other (type your own)',
value 'other'), duplicating the popover's built-in free-form row — two
other-ish rows, one pickable as a literal answer. Two layers:

- Tool description now tells the model NOT to include catch-all options (the
  answer UI always offers free-form input on its own)
- splitOtherOption (pure, spec'd) folds a catch-all option that arrives
  anyway out of the choice rows and uses its label as the inline input's
  placeholder — conservative match (value 'other', or a label reading as a
  free-form invitation), no false positives on real choices

* fix: single question surface + clean free-form-only popover

Two live-pause confusions: (1) the inline transcript card and the composer
popover both rendered — the card now defers while the popover is up for its
action, returning as the fallback surface when the user dismisses the popover
(and in contexts without a ChatContext, where the popover can't exist);
(2) an options-less question showed a pointless numbered '1 Something else…'
row — free-form-only questions now render the inline input alone, with the
'Type your answer…' placeholder (a folded model 'Other' label still wins).

* feat: the composer is the free-form answer box (like the main chat input)

While a question pause is live, the main chat textarea composes the free-form
answer — placeholder swaps to 'Something else…' (or a folded model 'Other'
label), Enter with text submits the answer through answer-mode key handling
(composed BEFORE useTextarea's submitting-lock, so the lock can't swallow it),
and the Stop button swaps to Send (enabled despite isSubmitting) per the
select-then-confirm design. The popover slims to the question header, numbered
option rows, and Skip/Submit — its inline input is gone since the composer
owns free-form now. Dismissing the popover restores normal composer semantics
(Stop button, normal sends).

* fix: Codex round 3 + real Skip semantics

- Skip now ANSWERS instead of hiding UI (danny): it resumes the run with a
  decline notice ('The user chose not to answer this question.') so the model
  moves on — a client-side dismiss left the run paused until expiry, a hung
  turn. × / Escape remain pure dismiss (switch to the inline card surface).
- P1 (resumed approval tool indices): resumed tool_calls steps whose
  tool_call id matches a seeded UNRESOLVED part now rebind to that seeded
  slot instead of offsetting — the original part resolves in place (output
  attaches) and no duplicate appears; message steps keep the offset, so the
  text-loss fix stands. createContentIndexOffsetHandlers now takes the seed
  array; resolved seeded calls are not rebind targets.
- P2 (stale selection across questions): selection state resets when the
  live actionId changes; the vestigial inline-Other state ('other' selection
  + text atom) is gone — the composer owns free-form.
- P2 (Redis abort path loses the args stamp): the abort route re-stamps the
  question onto the ask tool_call in the reconstructed abort content, so a
  Stop-abandoned question persists with its question intact.
- P2 (malformed args crash): parseAskUserQuestionArgs normalizes untrusted
  shapes (options: {} / non-string entries) instead of throwing in render.

* feat: free-form hint in the question popover footer

Left-aligned in the footer row (opposite Skip/Submit): 'Or type your answer
below' — points open-ended answering at the composer, whose placeholder
already reads 'Something else…'.

* feat: preserve composer drafts across the answer-mode swap

The answer phase gets its own draft key (ask-answer:<actionId>), passed as a
draftId override into useAutoSave — the key change itself drives the existing
save/restore machinery, so the conversation draft (or mid-run PENDING draft)
is stashed when a question pause takes the composer and restored once the
user answers, skips, or dismisses. Ask keys are exempt from the PENDING
migration branch, which would otherwise move-and-delete the stashed draft. A
half-typed answer survives reload/navigation while its question stays live.
Answer submission (option pick, free-form, skip) resets the composer via a
new non-throwing useOptionalChatFormContext, so the swap-back restores into
an empty box even outside ChatView-less render contexts (Share/search).

* fix: rebind resumed steps for ALL seeded tool call ids

The resume controller pre-stamps the user's answer onto the seeded
ask_user_question part, so the unresolved-only rebind predicate treated
it as settled and shifted the tool's re-run step to a fresh offset slot,
leaving a duplicate ask record in streamed/saved content. Tool call ids
are provider-minted per call: a resumed step bearing a seeded id can
only be the interrupted batch re-executing, so rebinding every seeded
id is always correct.

* feat: popover UX round 4 — clickable hint, collapse, click-submit, multiSelect

- Footer hint is a button that focuses the composer; reads 'Type your answer
  below' (no 'Or') when the question has no options.
- Collapse (chevron) hides the popover WITHOUT closing the pause: answer mode
  stays live (placeholder, Enter routing, draft key), the chat card renders
  the question with a ChevronUp affordance to re-expand. x remains dismiss.
- Single-select options submit on a single click; the Submit button renders
  only for multi-select.
- multiSelect end-to-end: tool zod schema + JSON definition twin, wire type,
  client parse, popover check-chips, card toggles, record-card label mapping;
  answer = option values joined ', '; composer Enter and the multi Submit
  button both fold free-form text in with the checked values.
- Hardening from adversarial review: in-flight status guard on every submit
  path (no duplicate resumes on double-click), popover locks while
  submitting, collapsed mode disarms invisible digit/arrow steering, the
  card shares the hook's checked state while the pause is live, the card
  folds catch-all 'Other' options, record mapping is all-or-nothing to avoid
  phantom labels, composer resets only when its text was consumed or the
  draft machinery will restore the stash.

* feat: ask_user_question in model specs and ephemeral agents

A librechat.yaml modelSpec can now equip the tool the same way it equips
webSearch/executeCode/fileSearch/memory:

  modelSpecs:
    list:
      - name: my-spec
        askUserQuestion: true

loadEphemeralAgent pushes the tool name when the spec flag (or the
ephemeralAgent request flag, wired for parity) is set; everything downstream
is the existing persisted-agent machinery — createRun's hitlCapable gating,
graphTools injection, checkpointer attach, subagent strip, and the admin
filteredTools/includedTools kill switch all apply unchanged.

* feat: tense-aware Q&A record label (Asking / Asked)

Shorten the record card header per feedback: 'Asking' while the question is
still unanswered (abandoned/awaiting), 'Asked' once answered — replacing the
single 'Asked a question' label.

* fix: Codex round 4 — added-agent ask parity + preserve answer on failed resume

F1 (added.ts): mirror loadEphemeralAgent's ask_user_question branch in the
added-agent loader so a model spec's askUserQuestion flag (or the ephemeral
request flag) equips added top-level agents too, matching execute_code /
web_search / memory. Two load.spec cases added.

F3 (composer): submitAskAnswer now takes an onSuccess callback and
useAskAnswerMode defers clearing the selection/composer until the resume is
accepted. A failed resume (16k answer-cap 400, expired action, network error)
leaves status re-answerable, so wiping the composer up front lost the user's
only copy of a free-form answer; now it survives for trim/retry.

(F2 — a claimed Tools-capability bypass — was verified NOT reproducible:
agentRequestsAskUserQuestion matches only loaded instances/toolDefinitions/
toolRegistry, all capability-filtered; a raw tools string has no .name and
never triggers the install. Replied on-thread with the probe evidence.)

* fix: Codex round 5 — expired question exits answer mode so its message shows

An expired question (e.g. resume returns the stale-action 409) previously left
the popover open with locked controls and no explanation, because the chat
card — which carries the only 'this action expired' message — was suppressed
by the popover-open guard. Treat 'expired' as no longer active: the popover
closes, the composer reverts to normal, and the card becomes the sole surface
and renders the expired message. 'error' stays active (retryable).

* feat: group ask_user_question calls as their own category

A homogeneous group of ask_user_question tool calls now reads 'Asked N
questions' (present tense 'Asking N questions' while the turn streams) with a
question glyph and no raw-name suffix — mirroring the subagent 'Ran N agents'
category treatment, instead of 'Used N tools — ask_user_question'. Mixed
groups keep 'Used N tools' but humanize the suffix to 'Question' and show a
question icon for the ask entries (TOOL_FRIENDLY_NAME_KEYS + ToolIcon map).
A group only forms at count >= 2, so the plural is always grammatical.
Three ToolCallGroup.test cases cover homogeneous label/icon/suffix, present
tense while streaming, and the mixed-group fallback.

* fix: Codex round 6 — composer submit lock + abort stamp before emit

F7 (composer status lock): the ask submit status lived on ApprovalContext,
a React context mounted only around message content (ContentParts). The
PRIMARY answer surface — the composer in ChatForm — renders outside it, so
useApprovalContext returned the inert FALLBACK: status was always 'idle',
setStatus a no-op. The in-flight double-submit guard (round 4) and the
expired-exits-answer-mode fix (round 5) therefore never engaged for the
composer. Move ask submit status to a global Recoil atom (useAskSubmitStatus)
read/written by the composer, the popover, and the card alike, so a fast
double-click/Enter is actually blocked and expired/error surfaces on every
surface. Tool-approval status stays on the context (unchanged).

F5 (abort stamp before emit): the abort route re-stamped a paused
ask_user_question's args AFTER GenerationJobManager.abortJob had already
emitted the final SSE from the unstamped content, so a Redis/cross-replica
Stop left the live client showing an empty question until reload. abortJob
now takes an optional transformAbortContent applied to the persistable
content BEFORE the final event is built (and returned), so the live client
and the saved message agree. New abort.spec case + updated call assertions.

* feat: gate ask_user_question behind its own agent capability

Add a first-class AgentCapabilities.ask_user_question (in defaultAgentCapabilities,
on by default) so admins can enable/disable questions independently via
endpoints.agents.capabilities, exactly like execute_code / web_search — not
lumped under the generic tools capability.

- ToolService: both filteredTools predicates (definitions-only and instance
  loaders) gate ask_user_question on checkCapability(ask_user_question) before
  the generic tools fallthrough. When off, the tool is dropped from
  toolDefinitions/toolRegistry, so run.ts's agentRequestsAskUserQuestion (which
  keys on the loaded surface) declines to install it and attach a checkpointer —
  the capability is enforced end-to-end at the loader, no run.ts change needed.
- Tools dialog catalog: surface the ask builtin under its own capability rather
  than the generic tools one, so the UI matches the backend gate.
- Tests: ToolService capability on/off filtering + defaults membership; catalog
  builtin visibility keyed on the dedicated capability.

* style: sort imports in ToolCallGroup.test (CI import-order gate)

* fix: Codex round 7 — surface ask-answer errors in the open popover

A failed answer submission (16k reject, network error) sets the ask status to
'error', which — unlike 'expired' — deliberately keeps the question active and
retryable. But the chat card that renders the error message is suppressed while
the popover is open, so a composer/popover answer failed silently. Expose an
'errored' flag from useAskAnswerMode and render a warning line
(com_ui_ask_answer_error) in the popover, so the user gets feedback and retry
guidance without having to collapse/dismiss. It clears automatically on retry
(status flips to 'submitting').

* fix: Codex round 8 — respect IME composition before submitting answers

handleComposerKeyDown runs before useTextarea's composition guard, so with a
CJK/IME keyboard the Enter that commits an in-progress composition was being
intercepted and submitting the partial answer (and the composition buffer can
leave value empty mid-compose, mis-triggering digit/arrow steering too). Bail
at the top when composing — nativeEvent.isComposing, or key==='Process' /
keyCode===229 for Safari's inconsistent reporting — mirroring the existing
composer guard so the character commits normally.

* chore: update `@librechat/agents` to v3.2.60

* 🔧 chore: Update @opentelemetry/core to version 2.9.0 and clean up package-lock.json

* feat: digit shortcuts select options when the popover has focus

Previously a number key (1..N) only selected an option from the empty
composer (handleComposerKeyDown on the textarea) — if focus moved into the
popover (a row/Skip/Submit button clicked or tabbed to), the number keys went
dead. Add handlePopoverKeyDown, wired to the popover container's onKeyDown so
it catches digits bubbling from the focused control: a digit activates its
option exactly like a click (single-select submits, multi toggles). No
highlight/Enter dance on this path — the options are buttons whose action is
the click, and intercepting Enter would fight the focused button. Gated on
active && !locked so it no-ops while a submit is in flight.

* chore: update @librechat/agents to version 3.2.61 and @opentelemetry packages to latest versions
2026-07-08 15:31:05 -04:00
Joseph Licata
fc67416d08
fix: Resolve Local JSON Pointer Refs & Sanitize Tool Search Schema for Gemini (#14161)
*  fix: Resolve Local JSON Pointer Refs & Sanitize Tool Search Schema for Gemini

Gemini/Vertex agents fail when MCP tools are attached, in two ways:

1. MCP servers with OpenAPI-derived schemas can emit local JSON pointer
   refs (e.g. `#/properties/body/properties/start`). `resolveJsonSchemaRefs`
   only resolves `#/$defs` and `#/definitions` refs, so these survive into
   the request payload and Vertex rejects it with 400
   'Invalid JSON payload received. Unknown name "$ref"'. Resolve any local
   `#/...` pointer against the schema root (RFC 6901 unescaping, cycle
   protection), and strip still-unresolvable `$ref` keys in
   `sanitizeGeminiSchema` as a safety net.

2. When deferred tools are enabled, the injected `tool_search` schema
   declares `mcp_server` as `oneOf: [string, string[]]`, which
   `zod_to_gemini_parameters` rejects ('Gemini cannot handle union types').
   `buildToolClassification` now accepts the agent provider and collapses
   the union via `sanitizeGeminiSchema` for Google/Vertex only — the tool's
   runtime (`normalizeServerFilter`) already accepts both forms.

Complements #13623, which sanitizes MCP tool schemas but not the injected
tool_search definition, and did not cover non-$defs local pointers.
Underlying converter limitation: langchain-ai/langchainjs#9691.

* style: format tool search schema assignment

* style: sort tool classification imports

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-08 15:30:37 -04:00
Serhii Zghama
96d213afe6
🧊 fix: Include Conversation Starters in Agent View and List Responses (#14142)
* fix: include conversation_starters in agent view and list responses

* test: cover conversation_starters in agent view and list projections

* test: include conversation_starters in agent list whitelist

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-08 12:54:54 -04:00
Danny Avila
96367828e1
🧷 fix: Align Agent File Attachment Ownership (#14149)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Has been cancelled
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Has been cancelled
Sync Helm Chart Tags / Ignore non-main push (push) Has been cancelled
Sync Helm Chart Tags / Sync chart tags (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Has been cancelled
* fix: Align agent file attachment ownership

* fix: Harden agent file unlink validation

* test: Align file preview agent attachment access

* test: Add agent file ownership e2e regression
2026-07-07 16:23:48 -04:00
Danny Avila
f3692b5d43
fix: Defer MCP Connection for Request-Scoped Placeholders on Reinitialize (#14148)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Co-authored-by: fe2131-art <232701181+fe2131-art@users.noreply.github.com>
2026-07-07 07:54:42 -04:00
Danny Avila
dcdaaeac67
🤐 fix: Exclude Provider Secrets from HITL Pending Actions (#14136)
* 🤐 fix: Exclude Provider Secrets from HITL Pending Actions

Sanitize resolved model parameters before persisting them in the pending
action's resumeContext, and strip resumeContext/requestFingerprint from
every client-facing copy (SSE emit, reconnect gap-fill, resume state,
status route). The full record stays server-side for resume replay.

* 🤐 fix: Treat Header-Carrier Keys as Sensitive Wholesale

Google llmConfig places the Authorization header in customHeaders, and
header names like Ocp-Apim-Subscription-Key defeat name heuristics.
Match 'header' as a key fragment so every header-carrier object is
dropped rather than relying on exact carrier names.

* 🤐 fix: Strip Preconfigured Client Instances from Resume Params

Bedrock stores a BedrockRuntimeClient on llmConfig.client when PROXY or
a bearer token is configured; replaying it would also fold a mangled
client object into additionalModelRequestFields on resume.
2026-07-07 07:22:50 -04:00
Ravi Kumar L
44d1275f36
⚙️ perf: reduce first-load MongoDB round trips (#14101)
* perf(api): reduce first-load database round trips

* docs: move agent guidance to claude docs

* refactor(api): move message validation into api package

* fix(api): narrow active generation job lookup

* fix(api): preserve omitted source identity
2026-07-06 09:36:34 -04:00
Danny Avila
8fcb77fe6f
🧵 fix: Preserve Fenced Markdown Artifacts (#14121)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Has been cancelled
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Has been cancelled
Sync Helm Chart Tags / Ignore non-main push (push) Has been cancelled
Sync Helm Chart Tags / Sync chart tags (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Has been cancelled
* fix: Preserve fenced markdown artifacts

* fix: Satisfy artifact CI checks

* fix: Handle longer artifact fences in updates
2026-07-05 12:04:59 -04:00
Danny Avila
2d4ef52c22
🧮 fix: Prevent String Corruption of Numeric Agent Model Parameters (#14119)
* 🧮 fix: Prevent String Corruption of Numeric Agent Model Parameters

* 🧮 fix: Support Partial Numeric Input and Cover Parameter Aliases
2026-07-05 12:03:41 -04:00
Marco Beretta
edd614bbff
🧰 feat: Redesign Agent Builder with Unified Tools Marketplace, Skills & Orchestration (#13952)
* feat: redesign the agent builder tools, skills, and advanced panels

Replace the stacked capability/MCP/skill/tool/action form sections with a unified tools marketplace, per-item configuration dialogs, and a consolidated Advanced panel.

- unified tools marketplace (catalog, sidebar, polymorphic cards/rows) covering built-in capabilities, plugins, MCP servers, and actions, each with a detail/config dialog
- dedicated Skills picker and a Tools section with selected-item summaries and empty states
- redesigned action editor and authentication dialog (method cards, segmented controls)
- rebuilt Advanced panel: orchestration hub (subagents, handoffs, chain), max steps, skills kill-switch, copyable agent id
- restyled version history (timeline, tool/capability counts, in-app restore confirmation)
- shared component updates (Radio, Input/Textarea, dropdown z-index, dialog primitives) and keyboard-only focus rings via useInputModality
- format-hint placeholders for tool credential fields
- sanitize numeric parameter inputs to prevent comma truncation

* feat: refine agent builder tools, actions, and MCP sections

* feat: restore Memory capability toggle in agent builder tools catalog

* feat: refine agent tools picker (skills, MCP connect/OAuth, web search)

- Skills picker: per-card visibility (public) and shared-author badges,
  category filtering, and an in-place Create skill flow that auto-attaches
  the new skill without leaving the builder
- MCP: inline Connect button in the first dialog plus a dedicated OAuth
  dialog (continue, copyable URL, QR code) shown only when OAuth is required
- Web search: auth-aware affordance, settings cog when user-provided and an
  info icon when system-defined
- Remove orphaned com_ui_unavailable/com_ui_initializing keys and the dead
  Tools/MCPToolItem component

* refactor: streamline MCP OAuth dialog

- Remove the Cancel button (the flow auto-closes on connect / times out)
- Show the URL in a read-only single-line scrollable input (cursor moves
  through it, not fully visible) with the shared CopyButton's smooth
  Copy/Check icon swap, matching the OAuth callback-URL field
- Put the primary Continue with OAuth action (icon trailing) and an
  icon-only QR toggle together in a row at the bottom, below the URL
- The QR reveals between the description and the URL with a smooth height
  animation (grid-rows 0fr to 1fr, matching MCPToolItem's reveal)

* feat: smoothly collapse MCP connect button once connected

* feat: cross-fade MCP tools between loading, list, and empty states

* feat: show MCP server icon in OAuth dialog title

* fix: vertically center OAuth dialog title against the MCP icon

* feat: smoothly animate auth field changes in the MCP server dialog

* feat: match Code Interpreter file upload to the File Search dropzone

Swap Code Interpreter's thin btn-neutral bar for the same dashed dropzone
(DropzoneContent + dropzoneClassName) File Search already uses, so the two
capabilities' upload UIs are consistent.

* feat: show a saving spinner and allow cancelling credential edits

Drive the tool credential Save button from the real mutation state so it
shows a spinner while the request is in flight, and add a Cancel button
when re-editing already-saved credentials so the edit can be dismissed.

* feat: make the skills create button a compact icon button

* fix: restore MCP attach semantics and confirmations in the tools marketplace

Connecting an MCP server from the item dialog now enables all of its tools
once the connection settles, deselect-all keeps the server attached via its
placeholder token instead of detaching it, adding a server writes the token
so a zero-tool attachment survives a save, and removing a server from the
tools list asks for confirmation again. Consume-only servers are excluded
from the catalog, matching the old select dialog.

Also share the catalog/selection pipeline between ToolsSection and the
marketplace through useAgentItems, hoist NEW_ACTION_ID next to ActionItem,
drop unused status/view union members and stale TranslationKeys casts,
document the phase-2 Favorites/Made-by-you views, fix the needs-setup dot
semantics and card focus suppression, remove the redundant close button in
CreateSkillDialog, move useInputModality into @librechat/client so external
consumers can mount it, and delete dead files and orphaned translation keys.

* fix: scope tooltip elevation to dialogs and restore dialog close button size

Tooltips go back to z-150 globally; inside a dialog they now borrow the
depth-aware popover z-index so they still clear nested dialogs (the Tool
Library item dialog) without outranking freshly opened modals everywhere
else. The default dialog close icon returns to its original size, and the
lc-field pointer-focus suppression ships with the package next to Input and
Textarea so external consumers get the whole mechanism from @librechat/client.

* feat: add favorites for marketplace tools, MCP servers, and skills

Reintroduce the favorite star from the old skill picker, generalized to
every marketplace item kind except per-agent actions. Cards in the Tool
Library and Skills dialogs get a hover-revealed star (always visible once
favorited), and the existing Favorites views in both dialogs now filter to
starred items.

Favorites persist in a dedicated ToolFavorite collection, one document per
(user, itemType, itemId) with a unique compound index, exposed through
atomic per-item PUT/DELETE endpoints under /api/user/settings/favorites/
tools. Per-item writes are idempotent and race-free across tabs/devices
(the unique index backstops concurrent toggles), reads are a single
index-backed query capped at 100 favorites per user, and the client keeps
React Query as the source of truth with optimistic updates. Handlers live
in @librechat/api with a thin route wrapper; methods follow the
data-schemas factory pattern with tenant isolation.

The favorites filter now matches on compound kind:id keys instead of bare
ids, closing a cross-kind collision where a tool and a skill sharing an id
would both match. The skill-favorites data-service stubs and the reserved
TUserFavorite.skillId field are replaced by the new tool-favorites service.

* feat: anchor the favorite star at the card's right edge

Swap the ToolCard action-bar order so the star sits rightmost with the
configure/info icon to its left. Every card can be favorited but only some
are configurable, so anchoring the star keeps it in a consistent position
across the grid.

* chore: remove translation keys orphaned by the tool library redesign

* fix: gate marketplace creation entries and resolve off-page selected skills

The Create New menu exposed MCP server creation to users without the
MCP_SERVERS create permission and action creation on deployments with the
actions capability disabled; both entries are now gated like their
pre-redesign counterparts, and the button hides when neither applies.

Selected skills missing from the first catalog page (limit 100) were
dropped from the Skills section entirely, leaving them impossible to
inspect or remove. useResolvedSkills restores the per-id lookup: off-page
skills are fetched individually and confirmed misses (deleted or no longer
shared) stay visible under an Unavailable skill placeholder so the stale
allowlist entry remains removable.

* fix: refetch favorites when toggled before the list loads, lint fixes

An optimistic favorite written over an unpopulated cache seeded the list
with only the toggled item, and cancelQueries killed the initial fetch
that would have corrected it, hiding existing favorites until reload. The
optimistic write now only applies over known data; otherwise onSettled
invalidates so the authoritative list is refetched.

Also unnest the version date-label ternary and drop an unused form watch
flagged by CI.

* fix: sync skills_enabled with selection edits and hydrate agent file entries

skills_enabled is the master opt-in for the skill allowlist, and an empty
allowlist with the flag on means the full accessible catalog. Selection
edits now sync the flag on empty/non-empty transitions via a shared
skillsEnabledTransition helper: picking the first skill enables it so the
choice takes effect on save, and removing the last one disables it so the
agent doesn't silently escalate to every skill. Mid-selection edits leave
the flag alone, preserving the Advanced kill switch's
disable-without-clearing behavior.

Agents loaded from the API carry only tool_resources.*.file_ids; the
client-only context/knowledge/code file entry arrays were read directly,
so existing attachments rendered as empty and could not be removed. A new
useAgentFileEntries hook restores the legacy derivation (agent files query
merged into the file map via processAgentOption) and now feeds AgentConfig,
the item dialog, and the selected-items pipeline.

* fix: hide plugin tools from the marketplace when the tools capability is off

buildCatalog gated built-ins, MCP, and skills on their capabilities and
permissions but pushed regular plugin tools unconditionally, so deployments
that removed the tools capability still offered attachable tool cards in
the marketplace. The loop now requires AgentCapabilities.tools, matching
the old Add Tools gate.

* fix: strip legacy MCP tokens on removal, guard action creation, model button spacing

MCP selection accepts every historical token format (server placeholder,
raw server name, mcp_-prefixed, and per-tool ids in prefix/suffix shapes)
but removal only filtered the new placeholder plus the server's current
tool ids, so a legacy token left the server permanently selected and its
tools still expanded after save. Selection and removal now share a
matchesMcpServer predicate.

Creating an action from the marketplace on an unsaved agent opened an
editor whose save was guaranteed to fail; it now surfaces the existing
save-the-agent-first error, matching the action-removal guard.

The model picker button keeps its tight px-1 with a provider icon but gets
px-3 in the empty Select-a-model state so the placeholder is not flush
against the border.

* fix: strip legacy prefix MCP tokens in useRemoveMCPTool

The hook only filtered the raw server name and suffix-delimiter tokens,
so confirming removal in the selected-tools section left persisted
prefix-format tokens (mcp_<server>, mcp_<server>_<tool>) in the form and
the row reappeared as selected. It now shares the matchesMcpServer
predicate with the selection logic so removal can never lag selection.

* fix: exact MCP token matching and keep errored skill lookups removable

The mcp_<server>_ prefix clause in matchesMcpServer was invented by the
redesign, not a persisted format (mcp_prefix is only ever used as the
exact mcp_<serverName> pluginKey), and it claimed longer server names
sharing a prefix: with servers github and github_extra, removing github
also stripped github_extra's tokens. The predicate now only matches exact
or delimiter-bounded shapes.

An off-page selected skill whose per-id lookup failed with a transient
error (retry disabled) vanished from the selected list until remount. Any
settled lookup failure now keeps the placeholder entry so the allowlist id
stays visible and removable; only in-flight lookups are briefly hidden.

* fix: route file-backed built-in removal to the file manager

Code Interpreter and File Search stay selected while they hold code_files
or knowledge_files, so removing them by flipping the capability flag left
the row visible and unremovable. Their removal now opens the config dialog
where the files are managed, mirroring the file-only context built-in;
with no files attached the flag still toggles off for a clean removal.

* fix: preserve negative values in numeric parameter inputs

sanitizeIntegerInput stripped every non-digit, so typing -1 in a numeric
parameter field became 1. That broke Google thinkingBudget, where -1 is
the dynamic/auto-thinking sentinel (range min is -1): users could no
longer select auto and risked sending a one-token budget. The sanitizer
now takes an opt-in allowNegative flag that keeps a single leading minus,
and DynamicInput passes it when the field's range permits negatives.
Thousands-separator cleanup is unchanged for all other fields.

* fix: keep in-progress negative numeric input and localize the actions heading

Typing a leading minus in a negative-capable numeric parameter (Google
thinkingBudget) sanitized to a lone '-', which was then coerced by
Number('-') to NaN, so the sign could not be typed before the digits. The
lone '-' is now stored as a string until a digit resolves it to a number,
matching how the empty-string case is already handled.

The agent builder actions panel heading hard-coded 'Add'/'Edit actions';
it now uses com_assistants_add_actions and a restored
com_assistants_edit_actions key so non-English locales translate it.

* chore: fix import order drift flagged by CI

* fix: treat pending web-search auth verification as needs_setup

While useVerifyAgentToolAuth is still loading, data is undefined so
web_search was not marked needs_setup, and the marketplace card takes the
direct-enable path only when status is not needs_setup. On a slow
connection a click before the response arrived enabled web_search without
collecting the required user-provided key. The auth map now flags
web_search needs_setup while the query is loading, routing the click to
the config dialog; once verification resolves, a system-defined deployment
or a satisfied key clears the flag for a direct toggle.

* test: update agent builder e2e selectors

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-05 11:30:12 -04:00
Danny Avila
84fa6aa820
🧹 feat: Eager HITL Checkpoint Cleanup (Expiry + Deletion) & Full-Wiring E2E (#14123)
* feat: eager HITL checkpoint cleanup on expiry + deletion, full-wiring e2e

Follow-up to the lazy checkpointer (#14024): two paths still left a paused
run's durable checkpoint to the 24h Mongo TTL, and no test exercised the
whole HITL seam with real components.

1. Approval expiry: GenerationJobManager.setApprovalExpiredHandler(fn) — a
   non-destructive host hook fired after expireApproval's CAS succeeds
   (periodic sweeper AND stale-submit path), safe on startups that run
   constructor defaults. Both startups (index.js configureGenerationStreams,
   experimental.js) register a handler that prunes the checkpoint, resolving
   config lazily per expiry (streamId === conversationId === thread_id).

2. Conversation deletion: deleteConvos now returns the deleted
   conversationIds; the three deletion paths (DELETE /convos, DELETE
   /convos/all, account deletion) prune them via the new bulk
   deleteAgentCheckpoints (one $in deleteMany per collection). The delete
   routes gain configMiddleware for the checkpointer config.

3. Full-wiring e2e (hitlCheckpoint.e2e.spec.js): real SDK Run driven by
   FakeChatModel calling a gated tool, real PreToolUse/humanInTheLoop wiring,
   real LazyMongoSaver over mongodb-memory-server, real GenerationJobManager,
   real /resume controller via supertest. Asserts: clean turn persists
   nothing; error turn persists nothing; pause -> HTTP approve -> gated tool
   executes exactly once -> finalize prunes the checkpoint; expiry prunes the
   abandoned pause eagerly.

Tests: 3 handler unit tests (pendingAction.spec), 2 bulk-prune integration
tests (checkpointer.integration.spec), convos route + deleteUser specs
updated, 4 e2e scenarios. 207 tests green across changed areas.

* fix: tenant-scoped expiry prune, store-won expiry relay, resilient deleteConvos ids

Codex round 1 on #14123 — all three valid:

1. The approval-expired handler now receives the expired JOB so both startups
   resolve config in the paused job's tenant/user scope (getAppConfig({userId,
   tenantId})) — a tenant checkpointer override no longer sends the prune to
   the base config's collections. expireApproval fetches the job best-effort.

2. Multi-replica: when RedisJobStore.cleanupRequiresActionIndex wins the
   expiry CAS on another replica, this replica's sweeper relay branch now runs
   the approval-expired cleanup too (prune is idempotent) — store-driven
   expiry no longer bypasses the hook.

3. deleteConvos: post-delete cleanup (deleteMessages, project stats refresh)
   is now best-effort — the conversations are already gone, so throwing hid
   the deletion and dropped the conversationIds the checkpoint prune needs,
   unrecoverable on retry. Updated the existing tag-decrement-on-failure test
   to the new contract (ids still returned).

Tests: handler-receives-job, store-won relay path, ids-survive-cleanup-failure.
134 tests green across changed suites.

* fix: relay cleanup independent of cached errorEvent; enter tenant ALS context

Codex round 2 on #14123:

1. The sweeper's relay branch gated BOTH the terminal-error emit and the new
   checkpoint cleanup on !runtime.errorEvent — but a reconnect seeds errorEvent
   from the aborted job (runtime-state creation), which then suppressed the
   cleanup entirely. The emit stays gated; the idempotent cleanup now runs
   independent of the cached error, once per runtime lifetime
   (approvalCleanupRan flag — the aborted job is swept repeatedly).

2. Passing userId/tenantId to getAppConfig only keys the config cache; the
   Config query is ALS-scoped by the tenant-isolation plugin. Both startup
   handlers now ENTER the paused job's tenant context via tenantStorage.run
   before resolving config + pruning, so a tenant checkpointer override is
   honored in strict and non-strict modes.

Tests: relay-cleanup-with-cached-error (reconnect simulation), repeated sweeps
run the cleanup once. 27+4 green.

* fix: dedup expiry cleanup across winner and relay paths

Codex round 3 (P3): expireApproval ran the handler without marking the
runtime's approvalCleanupRan flag, so the next sweep's relay branch (the
aborted job outlives expiry for the completed-job TTL) ran the cleanup a
second time. The dedup now lives inside runApprovalExpiredHandler — the
single choke point both paths call — set-before-run, once per runtime
lifetime. Test: local expiry followed by a sweep fires the handler once.
2026-07-05 11:29:30 -04:00
Danny Avila
7b7fa496aa
🗃️ perf: Cache Group Memberships for ACL Principal Resolution (#14075)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* 🗃️ perf: Cache Group Memberships for ACL Principal Resolution

Continues #14069: caches resolved group-membership ids per member key in a
new USER_PRINCIPALS namespace so repeated permission checks skip the group
query, with exact invalidation on every membership mutation, same-process
build dedup, and optional Redis cross-container build locks.

Co-authored-by: Peter Rothlaender <peter.rothlaender@ginkgo.com>
Co-authored-by: Joachim Keltsch <joachim.keltsch@daimlertruck.com>

* 🧵 fix: Defer Transactional Invalidation and Gate No-Op Bulk Clears

Codex round-2 findings on the principals cache: membership writes inside a
transaction now defer cache invalidation until the session ends so concurrent
readers cannot re-cache pre-commit state with no later correction, and
bulkUpdateGroups skips invalidation entirely when MongoDB reports no modified
or upserted documents. userGroup.spec.ts now runs on a single-node replica set
so the deferral is covered by a real transaction.

* 🌐 fix: Evict Cross-Process Stale Rewrites and Scope Entra Sync to Tenant

Codex round-3 findings: a delayed second invalidation pass (lock wait plus a
short grace) now evicts membership entries that a build in another container
re-cached after the first delete, and syncUserEntraGroupMemberships establishes
the user's tenant ALS context when the OAuth callback runs it pre-middleware,
so tenant-scoped principal keys are invalidated (and sync queries and created
groups are scoped) exactly like authenticated reads. Deferred transactional
invalidations snapshot the caller's ALS so they keep tenant scoping.

---------

Co-authored-by: Peter Rothlaender <peter.rothlaender@ginkgo.com>
Co-authored-by: Joachim Keltsch <joachim.keltsch@daimlertruck.com>
2026-07-05 08:40:14 -04:00
Danny Avila
ed8547018c
perf: Persist HITL checkpoints only on pause (lazy checkpointer) (#14024)
*  feat: Persist HITL checkpoints only on pause (skip clean-exit writes)

With `durability: 'exit'` (set by the SDK whenever a checkpointer is active) LangGraph
persists ONE checkpoint at the exit boundary on EVERY run — paused or not. So a non-paused
HITL turn writes a dead checkpoint whose only fate is to be pruned by deleteAgentCheckpoint:
pure write+delete churn on the common path, given HITL only ever resumes an *interrupt*
checkpoint.

`InterruptOnlyMongoSaver` (a MongoDBSaver subclass) persists only interrupt checkpoints and
discards clean-exit ones, so a non-paused turn writes nothing.

How it tells them apart (verified empirically against @langchain/langgraph, not docs):
when a run interrupts, the runner calls `putWrites` with the `INTERRUPT` ("__interrupt__")
channel for the checkpoint it's about to create, and that write's `config.checkpoint_id`
equals the `checkpoint.id` of the `put` that immediately follows. A clean exit calls `put`
with no preceding interrupt `putWrites`. So we record the checkpoint id of any interrupt
`putWrites` and persist a `put` only when its `checkpoint.id` was so marked. Keying on the
globally-unique checkpoint id (not thread_id) keeps this correct even when two runs race on
the same conversation (the job-replacement scenario).

Correctness is preserved end-to-end: interrupt checkpoints + their pending writes persist
exactly as before (resume unchanged); clean checkpoints were only ever written-then-pruned,
so not writing them is observationally equivalent. The eager prune stays as the backstop.

Tests (mongodb-memory-server): a bare put() is discarded; an interrupt-seeded checkpoint is
persisted with its __interrupt__ pending write; and an end-to-end real-graph run writes 0
checkpoints on a clean completion and a resumable one on interrupt.

NOTE: a non-paused turn's deleteAgentCheckpoint now finds nothing to delete (a 0-match
no-op) — a follow-up can skip that call entirely once the lingering-abandoned-pause cleanup
role is reassigned to the TTL + expiry sweeper.

*  feat: Drop the redundant clean-path checkpoint prune

With the lazy checkpointer (InterruptOnlyMongoSaver) a non-paused turn no longer writes a
clean-exit checkpoint, so the post-completion prune in chatCompletion's finally had nothing
left to delete. It was also already redundant: every fresh turn runs a pre-run prune
(`deleteAgentCheckpoint` before `processStream`) that clears any checkpoint orphaned by a
prior abandoned pause — verified empirically that a lingering interrupt checkpoint WOULD
otherwise poison a fresh turn (LangGraph continues the abandoned state + re-interrupts), and
that the pre-run prune is what prevents it. The Mongo TTL remains the backstop, and the
resume path still prunes after a successful finalize.

Removing the clean-path prune also deletes its job-replacement race surface (round-17 F21):
an older run's late finally can no longer delete a newer paused run's checkpoint, because
there is no longer a clean-path prune to race. Dropped the now-dead F21 predicate test.

Net per non-paused HITL turn: from {pre-run prune + checkpoint write + post-run prune} down
to {pre-run prune} — no write, no post-completion delete.

* 🛡️ fix: Anchor any pending-write checkpoint; stale-only eviction (Codex)

Broaden the lazy saver's keep-rule from "interrupt-only" to "persist any checkpoint that
carries pending writes" (renamed InterruptOnlyMongoSaver → LazyMongoSaver). This makes it
robust to delta-channel graphs without changing behavior for LibreChat's graph:

- K1 (P1): a delta-channel graph can write a synthetic PARENT/anchor checkpoint (no
  __interrupt__ mark) that the interrupt checkpoint then points at, with the delta writes
  stored under the parent id. The old rule discarded that parent, breaking delta-state
  resume. Now any checkpoint that received putWrites is persisted, so the anchor parent and
  its writes survive and resume can walk the chain.
- K3 (P2): for the same reason, clean delta-write rows are no longer orphaned — their
  checkpoint is persisted alongside them. (For LibreChat's standard Annotation/messages
  graph a clean run makes no putWrites at all — verified empirically — so the common path
  still writes nothing and the optimization is unchanged.)
- K2 (P2): the 1024 FIFO cap could evict a valid in-flight id whose put() was just behind
  Mongo I/O, mis-classifying its interrupt checkpoint as a clean exit. Replaced with
  time-based eviction: only ids older than 5 min (a put always follows its putWrites within
  ms) are swept; a recent in-flight id is never dropped, and the map grows rather than evict
  a valid id if nothing is stale.

New integration test: a checkpoint anchored by a NON-interrupt write is persisted. Full
agents/HITL suites green (108).

* style(checkpointer): fix import order to satisfy sort-imports CI

* fix(checkpointer): don't persist failed-turn (error-only) checkpoints

LazyMongoSaver anchored on ANY pending write, so a non-paused turn that
errors (LangGraph records an __error__ write then a put) was persisted and,
with the clean-path prune removed, lingered until the next fresh-turn prune
or the Mongo TTL. Anchor only on resumable writes — INTERRUPT or a real
(non-__-prefixed) state/delta channel — so error/bookkeeping-only checkpoints
are discarded at the source. Addresses Codex P3.

Codex P2 (delta-stub parent orphan) is not reachable: the SDK graph uses
standard Annotation/MessagesAnnotation channels (no DeltaChannel), and under
durability:'exit' putWrites precedes put with a parentless boundary
checkpoint — probe-confirmed against @langchain/langgraph@1.4. Documented the
durability:'exit' invariant the saver depends on.

Tests: error-only put discarded; e2e throwing graph persists 0 checkpoints.

* fix(checkpointer): drop bookkeeping-only write batches, not just the checkpoint

The prior fix stopped the failed-turn CHECKPOINT from persisting, but putWrites
still forwarded the __error__ batch to MongoDBSaver.putWrites — writing a row to
agent_checkpoint_writes whose parent checkpoint is then discarded. With the
post-run deleteThread removed, that orphan row lingered until the Mongo TTL or
the conversation's next pre-run prune. putWrites now drops a non-resumable
(bookkeeping-only) batch entirely instead of forwarding it.

Probed against a real MongoDBSaver (mongodb-memory-server): a throwing graph now
leaves 0 checkpoints AND 0 write rows (was 0 + 1 orphan), while interrupt->resume
is unaffected — the __interrupt__ write is resumable so it is still forwarded.
Addresses Codex P2 (round 3).

Tests: error-only put leaves no checkpoint and no write row; e2e throwing graph
leaves both collections empty; new e2e interrupt->resume completes with the
approval value.

* fix(checkpointer): un-anchor a checkpoint whose putWrites failed; freshen comments

Self-review findings on the converged PR:

1. LangGraph dispatches put() concurrently with putWrites (probe-confirmed on
   1.4.5), and put() still completes when putWrites rejects — so a transient
   Mongo failure during the interrupt write could persist a checkpoint whose
   __interrupt__ row is missing (an unresumable phantom pause). putWrites now
   deletes the write anchor on rejection (best-effort) and rethrows, so that
   put() discards the checkpoint instead. The pre-recorded anchor stays where
   it is — recording after the await would drop slow-I/O interrupts on the
   success path, which the same probe showed is reachable.

2. Renamed leftovers: two comments still said InterruptOnlyMongoSaver; the
   class is LazyMongoSaver.

3. Documented why the pre-run prune is deliberately unconditional per HITL
   turn (any cheaper gate can go stale across replicas and skip the prune
   exactly when an orphaned interrupt exists).

Test: failed putWrites → subsequent put persists nothing (14/14 green).

* fix(checkpointer): bookkeeping write batches follow their checkpoint's fate

The round-3 rule dropped bookkeeping-only putWrites batches (__error__/
__resume__/__no_writes__) unconditionally — batch-scoped, when the decision
must be checkpoint-scoped. Probe-confirmed (langgraph 1.4.5, durability:'exit'):
a Send fan-out that pauses on one sibling records the completed siblings as
pure __no_writes__ batches on the RETAINED interrupt checkpoint; dropping those
markers makes resume re-execute the completed siblings (side effects measured
twice). Addresses Codex M2 (P2).

putWrites now PARKS a bookkeeping-only batch in memory until the checkpoint's
fate is known: forwarded when the checkpoint is anchored (or was just
persisted — put is dispatched concurrently), dropped when put discards it. Net:
an errored turn still leaves nothing durable (0 checkpoints, 0 write rows),
and a retained checkpoint stores byte-for-byte what a plain MongoDBSaver would.

Codex M1 (__resume__ lost on re-pause) did not reproduce: the re-pause emits
[__interrupt__,__resume__] as ONE batch (anchored, forwarded whole) and a
second resume on a rebuilt graph replays both answers correctly — but the
fate-scoped buffering now covers a lone __resume__ batch in any ordering too.

Tests: bookkeeping preserved on a retained checkpoint in either arrival order;
e2e Send-sibling pause/resume with side-effect counters (was {a:2,c:2} under
the drop rule, now {a:1,c:1}); error-only turn still leaves both collections
empty. 16/16 green.
2026-07-05 08:30:06 -04:00
Danny Avila
424ccffd83
🪝 feat: Configurable Tool-Approval Policy via Programmatic Hooks (#14025)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Has been cancelled
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
* 🪝 feat: Programmatic tool-approval hook seam (configurable beyond on/off)

Adds a process-wide registry so host code can plug context-aware PreToolUse decision
hooks into the tool-approval policy, composing with the static
`endpoints.agents.toolApproval` config instead of replacing it.

- `registerToolApprovalHook(factory, { matcher? })` — register a factory that builds a
  PreToolUse hook per run from a ToolApprovalHookContext (userId, conversationId, tenantId,
  appConfig); return undefined to opt the run out. Returns an unregister fn.
- `buildHITLRunWiring(policy, context)` now registers the static-config policy hook as the
  baseline, then layers each resolved host hook after it. Decisions fold in the SDK as
  deny > ask > allow, so a host hook can only TIGHTEN a configured ask/deny — it can never
  silently auto-approve past policy (to loosen, change the static policy). updatedInput /
  allowedDecisions follow the SDK's last-writer-wins, so host hooks win over the baseline.
- `createRun` threads the per-run context (user / conversation / tenant / appConfig) into
  the wiring; non-HITL and HITL-disabled runs never invoke any factory.

This unlocks dynamic policy the static name-lists can't express — per-args (e.g. ask before
write_file outside a workspace, the SDK's createWorkspacePolicyHook shape), per-agent,
per-user. Inert until tool approval is enabled and the caller is hitlCapable.

Tests: registry register/unregister/opt-out/order (hooks.spec.ts) + wiring composition,
context passthrough, and disabled-path inertness (runtime.spec.ts). Full HITL suite green.

* 🪝 feat: Config-driven tool-approval hook loader (librechat.yaml → hook modules)

Lets operators declare programmatic tool-approval hooks in config instead of code, so the
registerToolApprovalHook seam is usable without a custom build.

- Config (data-provider): `endpoints.agents.toolApproval.hooks[]`, each entry
  `{ module, matcher?, options? }`. `module` is a bare package name or a path (resolved
  against the app root); its default export is a builder `(options?) => ToolApprovalHookFactory`.
- Loader (@librechat/api `loadToolApprovalHooks`): imports each module, builds the factory
  with the entry's options, and registers it (with its optional tool-name matcher). Reload-
  safe (each call first unregisters its previous batch, leaving code-registered hooks alone)
  and robust — an unimportable module / non-function export / throwing builder is logged and
  skipped, never crashing startup or blocking the other hooks. Importer is injectable for tests.
- Startup (api/server/index.js): loads the configured hooks once after appConfig resolves.

SECURITY: modules are dynamically imported + executed in-process; this is admin-level config,
documented as trusted-code-only.

Tests: 9 loader cases (default/no-default export, options passthrough, bad-export skip,
builder-returns-non-function skip, import-failure resilience, continue-past-bad-entry,
reload de-dup). Full HITL suite green (80).

* 💄 style: Sort imports in HITL hook spec files (CI sort-imports:check)

* 🛡️ fix: Harden tool-approval hook loader (Codex review)

Six P2 findings on the hook loader / startup wiring:

- CJS/transpiled interop: unwrap a nested `default` (TS/Babel `exports.default = fn`
  surfaces through import() as `{ default: { default: fn } }`) before rejecting a module,
  so documented default-export hook modules actually load.
- Validate the matcher regex at load time and skip invalid ones — the SDK compiles it with
  `new RegExp` at run-build time, where a bad pattern would throw out of buildHITLRunWiring
  and break EVERY HITL run instead of just skipping the one bad hook.
- Honor the `enabled` kill switch: startup now passes hooks to the loader only when
  toolApproval is enabled, so a disabled endpoint imports/runs nothing (and unregisters any
  prior batch).
- Resolve app-root-relative paths without a leading dot: a bare specifier that is a real
  file under basePath (e.g. `config/hooks/workspace.js`) resolves as a path; scoped/other
  bare names still import as packages.
- Base-config-only: documented that hooks register once process-wide at startup and are NOT
  reloaded from per-role/user/tenant overrides — encode per-tenant logic inside the hook.
- Wire the loader into the clustered startup path (api/server/experimental.js) too, not
  just the standard server.

Tests: CJS-interop unwrap, invalid-matcher skip (+ sibling still loads), and specifier
resolution (app-root file / bare package / ./relative). Full HITL suite green.

* 🛡️ fix: Read tool-approval hooks from base config in clustered startup (Codex)

The clustered experimental.js path read toolApproval from getAppConfig() (which merges DB
__base__ overrides with no principal), so a DB override could enable/disable/replace
toolApproval.hooks and import those modules in every worker — violating the base-config-only
contract and diverging from the standard server. Fetch getAppConfig({ baseOnly: true })
specifically for the hook loader, matching api/server/index.js.
2026-07-02 11:51:24 -04:00
Danny Avila
53b3e166d8
📉 perf: Skip Redundant Permission Queries on MCP Servers List (#14077)
Some checks failed
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Has been cancelled
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Has been cancelled
Sync Helm Chart Tags / Ignore non-main push (push) Has been cancelled
Sync Helm Chart Tags / Sync chart tags (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Has been cancelled
2026-07-02 11:45:23 -04:00
Danny Avila
b6c0bc7c0d
perf: Skip Role Route Capability Probe for Own and Default Roles (#14073) 2026-07-02 10:45:41 -04:00
Arjun Vijay
89931baf22
🚪 fix: Support Admin Redirect Detection for Same-Origin Subpaths (#14040) 2026-07-01 11:40:02 -04:00
matt burnett
b20abb2593
fix: bound peak memory of concurrent base64 attachment encoding (#14023)
* fix: bound peak memory of concurrent base64 attachment encoding

* chore: sort encode imports

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-07-01 08:22:16 -04:00
Alexey Korepanov
ac759ef2f7
🥷 feat: Add showInMenu Option to Model Specs (#14034)
Add an optional `showInMenu` flag to model specs. When set to false, the
spec is dropped from the model selector menu and from the client startup
config (GET /api/config), but remains resolvable server-side by name — a
request that sends `spec: "<name>"` still works, since server-side
resolution uses the full, unfiltered list.

Unlike `showIconInMenu` (which only hides the icon), this hides the whole
entry. The flag is optional and defaults to listed, so existing specs are
unaffected.

Adds an `excludeHiddenModelSpecs()` helper (applied before
`sanitizeModelSpecs`) plus unit tests.
2026-06-30 19:32:59 -04:00
Danny Avila
6dbf9d5ad3
🪝 feat: Human-in-the-Loop Runtime - Tool Approval + Ask-User-Question (Slice B) (#13942)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* chore: add @langchain/langgraph-checkpoint-mongodb for HITL durable resume

* feat: HITL tool approval runtime — backend (Slice B)

- endpoints.agents.checkpointer config + durable Mongo checkpointer (seam over the app
  connection; SDK MemorySaver fallback) with a TTL index + deleteThread pruning
- HITL run wiring (PreToolUse policy hook + humanInTheLoop) attached in createRun, fully
  inert when toolApproval.enabled is off
- interrupt gate (pause job -> requires_action + emit on_pending_action) and a resume
  route that rebuilds the run from the durable checkpoint and run.resume()s it
- atomic single-winner resolve; agent-consistency guard; expireStaleApprovals terminal
  event; checkpoint pruned on every non-paused completion (thread_id == conversationId)

* feat: HITL tool approval UI — frontend (Slice B)

approve/reject/edit/respond + ask-user controls in the tool card (OAuth-button precedent),
batch-aware single submit, live + reconnect (resumeState.pendingAction) wiring, and resume
mutations posting to /agents/chat/resume.

* fix(hitl): decouple ApprovalProvider from chat context

ApprovalProvider is now pure state (safe to mount in provider-less / shared / test
renders); the context-dependent submit moved to a useResumeSubmit hook the cards call.
Part imports getAskUserQuestionPart from ~/utils/approval directly so suites that
partial-mock ~/utils render Part without throwing.

* fix(hitl): address Codex review — backend

- P1: enforce per-tool allowed_decisions on resume (reject a crafted decision the
  policy disallows) via findDisallowedDecisions
- prune the durable checkpoint on user-abort of a paused run, and before a fresh
  HITL turn, so a new turn cannot rehydrate an expired/aborted interrupt (thread_id
  is the stable conversationId)
- persist + use isTemporary and the original parentMessageId on resume (temporary
  chats stay temporary; initializeAgent scopes thread files off the right parent)
- generate a deferred first-turn title BEFORE completeJob so its event reaches the
  client and the final event carries the real title
- moderateText: skip when there is no text (tool-approval resume) and moderate the
  ask-user answer, instead of denying on an empty input

* fix(hitl): address Codex review — frontend

- render ToolApproval for ANY paused agent tool card (bash/code/file/etc.), not just
  the generic ToolCall, by wrapping the tool-card branch in Part (moved the rendering
  out of ToolCall)
- findPendingActionMessageIndex only matches an assistant message, never the user
  message (the underscore-strip could target the user bubble before the assistant
  placeholder exists)

* fix(hitl): address Codex re-review

- title eligibility checks the user message’s parent (first turn), not the response’s
  parent — the previous check could never be true and skipped title generation
- use client.buildResponseMetadata() for the resumed message so contextUsage /
  thoughtSignatures survive (the abort-only helper dropped them)
- moderate decisions[].responseText (the respond action’s user text)
- give /chat/abort req.config (configMiddleware) so the HITL checkpoint prune on abort
  actually runs
- read resume state BEFORE setContentParts so the in-memory store does not lose the
  pre-pause seed content
- count resumes against LIMIT_CONCURRENT_MESSAGES (increment/decrement) so paused-then-
  resumed turns cannot bypass the limit
- require actionId on resume so a body without it cannot resolve the current action

* fix(hitl): address Codex re-review (round 3) — resume fidelity

Bring the lean resume path to parity with sendMessage for things it bypassed:
- carry userMCPAuthMap into the rebuilt run so approved MCP tools keep the user's creds
- seed initialSessions (buildInitialToolSessions) so approved code/file/skill tools have
  the pre-pause uploaded-file context (esp. cross-replica / after restart)
- await client.artifactPromises and persist them as response attachments (else tool
  artifacts created after the pause vanish on reload / for late subscribers)
- merge metadata: cumulative usage (+ summary marker) from the job, contextUsage /
  thoughtSignatures from the client — fixes the round-2 regression that underreported
  post-resume cost

* fix(hitl): address Codex re-review (round 4) — resume hardening

- resume: require an EXACT paused agent_id match (reject omitted/ephemeral
  agent_id, not just a different one) and reject an endpoint mismatch, so a
  request can't rebuild the claimed checkpoint on a different graph
- moderateText: also moderate a tool-approval decision's reject `reason` and
  stringified `editedArguments`, not just `responseText`
- request: re-mark the paused response `unfinished:true` after BaseClient saves
  it as completed, so an expired / never-resumed approval doesn't leave a
  "finished" response in history; the resume path overwrites it on success

* test(hitl): route-level integration test for the resume controller

Adds api/server/controllers/agents/__tests__/resume.spec.js, a supertest
integration test that drives the real ResumeAgentController over the full
pause -> approve -> resume -> finalize lifecycle with the SDK run, durable
checkpointer, Mongo, and concurrency cache mocked. The pure decision/liveness
helpers run for real via requireActual, so the guard ladder is exercised end to
end rather than stubbed.

25 cases covering:
- the authorization / staleness / agent-and-endpoint / actionId guard ladder
- tool_approval validation (undecided tool call, policy-disallowed decision)
- ask_user_question answer requirement
- the concurrency gate (429) and the atomic single-winner claim (409)
- the happy path: ACK, run reconstruction, decision->SDK mapping, finalize
  (save the now-finished response, emit done, complete job, prune checkpoint)
- first-turn title generation before stream completion
- re-pause (no double finalize), abort-during-resume (no double finalize),
  and the resume-failure terminal path (emitError + completeJob + prune)

* test(hitl): strengthen resume coverage + add approval util tests

Acts on a self-audit of the new resume integration test.

resume.spec.js (25 -> 32 cases):
- replace the tautological emitDone assertion (it only checked the hardcoded
  `final: true`) with a structural check of the finalEvent payload —
  responseMessage content/id/unfinished, requestMessage identity, title
- cover the previously-unwalked finalize branches: tool-artifact attachments
  (null-filtered), the aggregatedContent fallback when live content is empty,
  and client response-metadata attachment
- add guard cases: unsupported pending-action type (400) and the
  pre-multi-tenancy null-tenantId pass-through (must not 403)
- add error-path cases: first-turn title generation throwing must still
  finalize, and a completeJob failure during a resume error must force a
  terminal job state via the last-resort updateJob

client/src/utils/approval.spec.ts (new, 15 cases):
- applyPendingAction tool_approval: join by tool_call_id not position,
  skip completed calls, default allowed_decisions to [], referential
  stability when nothing changes
- applyPendingAction ask_user_question: append, idempotent replace on replay,
  non-array content coercion
- getAskUserQuestionPart type guard; findPendingActionMessageIndex
  assistant-only resolution (never resolves to the user bubble)

* fix(hitl): address Codex re-review (round 5)

Five findings verified against the code before fixing:

- resume: require an EXACT endpoint match (like agent_id) — a resume that OMITS
  endpoint must not fall through, since the shared chat middleware treats a
  missing/non-agents endpoint as the ephemeral agent and could rebuild the
  claimed checkpoint on a different graph
- resume: filter malformed content parts before saving the finished response,
  matching the normal AgentClient path (a resumed turn could otherwise persist
  an empty/invalid tool_call part that breaks reload/rendering)
- resume: accumulate tool artifacts across pause segments — persist them on
  re-pause and MERGE (not overwrite) at finalize, so artifacts produced before
  a second approval pause aren't dropped by the next rebuilt client
- approval (client): findPendingActionMessageIndex returns -1 when a provided
  responseMessageId isn't found, so the caller retries instead of attaching the
  prompt/approval to a prior assistant reply; fall back to the last assistant
  only when no responseMessageId is given
- RedisJobStore: make appendChunk extend-only (XADD + EXPIRE-if-shorter via a
  single eval) so the on_pending_action chunk emitted after a pause can't reset
  the chunk-stream TTL back to the running window and evict pre-pause content
  before the approval is resolved

Tests: +endpoint-omitted/unsupported-type/malformed-filter/attachment-merge/
re-pause-persist cases in resume.spec.js (36); ask-retry -1 semantics in
approval.spec.ts (16); extend-only TTL assertion in the RedisJobStore Redis
integration spec.

* test(hitl): mongodb-memory-server integration test for the checkpointer seam

The checkpointer unit spec covers config/selection with no DB connection; this
exercises the durable Mongo seam against a real (in-memory) MongoDB — the part
correctness actually depends on:

- getAgentCheckpointer builds a real MongoDBSaver when Mongo is connected and
  setup() creates the TTL index (expireAfterSeconds) on the checkpoint collection
- memory type returns undefined (SDK MemorySaver fallback) even when connected
- saver is memoized per resolved config
- deleteAgentCheckpoint prunes a thread's persisted checkpoint (the cross-turn
  isolation guarantee: turn N+1 on the same conversationId can't rehydrate it)
- pruning is thread-scoped — deleting one conversation leaves others intact
- undefined threadId is a no-op

* fix(hitl): address Codex re-review (round 6)

Four findings verified against the code before fixing:

- messageFilterPii: scan the resume payload's user-authored text (ask-user
  `answer`, and a tool-approval decision's `respond` text, `reject` reason, and
  edited tool arguments) — the shared /resume route ran through the PII filter
  but it only inspected req.body.text, so a blocked token rode the resume
  payload back into the model/tool (mirrors the earlier moderateText fix)
- resume: re-prime skill files invoked in the pre-pause segment before rebuilding
  the run, so an approved code/file-backed tool keeps the injected skill-file
  session refs instead of running without them (mirrors the normal path's
  primeInvokedSkills; the pre-pause content stands in for the message payload)
- hitl: pin the graph identity. Persist a fingerprint of the graph-determining
  request fields (endpoint, agent_id, model, spec, ephemeralAgent — normalized)
  on the pending action at pause, and reject a resume whose recomputed
  fingerprint differs. This closes the ephemeral-agent gap, where agent_id is
  undefined so the id guard can't tell two ephemeral configs apart
- resume: reject incomplete edit/respond decisions (findIncompleteDecisions) —
  an `edit` without an object editedArguments or a `respond` without non-empty
  responseText is 400'd before mapping, rather than defaulting to {} / '' and
  resuming with behavior the user never approved

Tests: incomplete-decision + fingerprint match/mismatch cases in resume.spec.js
(41); findIncompleteDecisions + computeAgentRequestFingerprint unit tests; and
resume-field PII cases in messageFilterPii.spec.ts.

* fix(hitl): address Codex re-review (round 7)

Four findings verified against the code before fixing:

- RedisJobStore: clear `agent_id` on createJob (add it to staleHitlFields). The
  job hash is keyed by conversationId and reused across turns; updateMetadata
  only writes agent_id when truthy, so a conversation that switched from a saved
  agent to an ephemeral/no-agent turn kept the old id and the resume guard
  rejected the valid pause as a different agent. (real correctness bug)
- fingerprint: include `promptPrefix` in computeAgentRequestFingerprint, and
  re-send it on resume (ResumeAgentFields + buildResumeFields). Ephemeral agents
  derive their system instructions from promptPrefix, so a resume changing it
  previously passed the pin and rebuilt different instructions. (completes the
  round-6 fingerprint)
- resume: the re-pause branch now persists the segment's accumulated CONTENT
  (filtered), not just artifacts, so an approval that expires/reaps without a
  final resume no longer loses everything streamed during the resumed segment.
- request: carry `manualSkills`/`alwaysAppliedSkills` on the persisted user
  message so a resumed turn's reconstructed requestMessage keeps its skill pills
  instead of dropping them until a full reload.

Deferred (narrow, no safe contained fix yet — see PR thread replies):
- resume rebuild without `addedConvo` for a multi-conversation/added-agent pane
- cross-replica re-prime of manually-selected (not model-invoked) skill files

Tests: stale-agent createJob clearing (Redis integration), promptPrefix
fingerprint match/mismatch (resume.spec.js + policy.spec.ts), re-pause content
persistence (resume.spec.js).

* fix(hitl): address Codex re-review (round 8)

Five findings verified against the code before fixing; the headline is a durable-
resume correctness fix (the fingerprint had surfaced it as a 403):

- resume durability (the important one): persist the graph-determining request
  fields (endpoint, agent_id, model, spec, promptPrefix, ephemeralAgent) on the
  pending action as `resumeContext`, and REPLAY them onto the resume request via
  a router-level middleware that runs before buildEndpointOption. The client
  can't reconstruct the ephemeral-agent config after a reload/cross-session, so
  the round-6/7 fingerprint would 403 a valid durable resume — and even without
  it the rebuilt agent would lose its tools. Replaying server-side rebuilds the
  SAME graph regardless of client state (and a crafted resume can't swap it; the
  fingerprint still matches because the body is restored first).
- RedisJobStore: also clear `isTemporary` on createJob (same class as agent_id):
  a prior temporary turn's flag would otherwise survive a reused conversation
  hash and a later non-temporary resume would save its response as temporary.
- resume: persist `contextMeta` (context-window calibration) onto the saved
  response like BaseClient does, so the next turn can seed its pruner.
- request: carry manualSkills/alwaysAppliedSkills into the onStart metadata
  update (not just the preliminary one it overwrites), so a resumed turn's
  requestMessage keeps its skill pills.

Deferred (narrow — see thread reply):
- saved-agent edited WHILE a run is paused: agent_id matches but the definition
  changed; needs an agent version/config hash, which is a larger change for a
  narrow window.

Tests: resumeContext pick/apply + round-trip (policy.spec.ts), contextMeta +
manualSkills-on-requestMessage (resume.spec.js), isTemporary clearing (Redis
integration).

* style(hitl): prettier line-wrap in policy.spec.ts (R8 lint fix)

* fix(hitl): address Codex re-review (round 9)

Five findings, all fixed (addedConvo — deferred in rounds 7/8 — is now trivial
thanks to the round-8 replay):

- replay addedConvo: add it to RESUME_CONTEXT_KEYS so the resume middleware
  restores the parallel/secondary-agent config from the paused request; the
  client can't reconstruct it, and it determines the rebuilt graph.
- skill pills (the real fix this time): the round-8 onStart metadata write was
  overwritten by trackUserMessage (the authoritative userMessage writer). Carry
  manualSkills/alwaysAppliedSkills in the emitted `created` message and persist
  them in trackUserMessage; widen UserMessageMeta + SerializableJobData.userMessage.
- execute-code files on resume: seed the paused user message's own files onto
  req.body.files before initializeClient — they're excluded from the
  parent-walk code-session rebuild, so an approved code/read-file tool would
  otherwise resume without them.
- in-memory pending-action UI: route ApprovalEvents.ON_PENDING_ACTION in the
  resume replay/pending-event loops to applyPendingActionToMessages (mirror the
  live handler), so a pause that lands in the snapshot window still renders its
  approval controls instead of sitting paused with no UI.
- abort isTemporary: the /chat/abort partial-save now sources isTemporary from
  the job metadata, not req.body (the stop button posts only conversationId), so
  aborting a paused temporary chat no longer persists an orphaned partial.

Tests: addedConvo in pickResumeContext (policy.spec.ts), file-restore on resume
(resume.spec.js), abort-from-job-isTemporary (abort.spec.js).

* fix(hitl): address Codex re-review (round 10) — resume/expiry races

Three concurrency/coherence findings, verified against the code before fixing:

- expiry-sweep CAS scope: both stale-approval sweeps (GenerationJobManager
  expireStaleApprovals and the RedisJobStore requires_action cleanup) called
  expire()/transitionStatus WITHOUT the observed pendingAction.actionId, so the
  CAS only checked status===requires_action. Between the read and the CAS a user
  could resolve the observed action and the run re-pause on a FRESH action; the
  stale sweep would then abort that valid new pause. Now both pass the observed
  actionId as expectActionId, so the CAS only fires for the action read as stale
  (a re-paused action has a different id → no-op).
- resume graph cache: resumeCompletion cached the rebuilt graph (created with
  messages:[]) via setGraph; RedisJobStore.getContentParts prefers a cached
  graph over reconstructing from the chunk log, so a same-replica reload/status
  poll mid-resume returned aggregatedContent missing the pre-pause content. Skip
  setGraph on resume so introspection falls back to the complete chunk
  reconstruction (setContentParts still seeds the in-memory store).
- pending-action UI: applyPendingActionToMessages scheduled a SINGLE
  animation-frame retry then dropped the pending action; Recoil/React updates can
  take several frames under load, leaving a valid requires_action run with no
  approval controls. Retry across frames (bounded at 120) until the target
  message commits.

Test: expire() with a mismatched expectedActionId no-ops while the matching id
expires (pendingAction.spec.ts).

* chore(deps): update @librechat/agents to version 3.2.53 and @langchain/langgraph to version 1.4.7 in package-lock.json and related package.json files

* refactor(hitl): add resolveToolApprovalPolicy seam for layered policy

Extract the single point where tool-approval policy is resolved for a turn
(`resolveToolApprovalPolicy`) and route the run call site through it instead
of reading `endpoints.agents.toolApproval` inline.

Behaviour-preserving: only the `endpoint` layer is wired today, so the result
is identical to reading the app policy directly. The `agent` and `skills`
layers are reserved seams with documented precedence (endpoint owns the
`enabled` kill switch; agent overrides mode/allow/deny/ask/reason; skills may
only tighten), so future per-agent and per-skill policy plumbing lands in one
function rather than at the `createRun` site. Adds focused unit tests.

* fix(hitl): address Codex re-review (round 11) — resume hardening

F1 (P2, security) — applyResumeContext now DELETES any RESUME_CONTEXT_KEY
absent from the persisted context, so the resume body carries exactly the
graph-determining fields the pause had. Previously only defined keys were
overwritten, leaving a client-supplied `addedConvo` (which the request
fingerprint does not cover) in place — a crafted resume could rebuild a
single-agent checkpoint as a different multi-agent graph/tool set.

F3 (P2) — the resume route ACKs (res.json) before initializeClient, so a
post-ACK getMCPRequestContext(req, res) saw the response as finished and
returned undefined, leaving the resumed run without its run-scoped MCP
connection store (approved MCP / OAuth-overlay tools then ran without their
request-scoped connections). Pre-seed the store with a null res +
cleanupOnResponse:false before the ACK and tear it down in the finally,
mirroring the normal stream path (request.js). userMCPAuthMap was already
preserved separately, so credentials were not lost — only the connection store.

Declined: the ApprovalContext NEW_CONVO guard (P2) is a false positive — the
`created` SSE event updates the conversation atom before any pause renders, so
the id is concrete by click time (details in the PR thread).

Tests: policy.spec (absent-key delete) + resume.spec (MCP context pre-seed/cleanup order).

* fix(hitl): address Codex re-review (round 12) — resume fidelity + multi-tool UI

F4 (P2) — temporal prompt vars: resume rebuilt the agent without restoring
req.conversationCreatedAt or req.body.timezone, so {{current_datetime}}-style
vars compiled a different system prompt than the paused graph (resume wall-clock,
unzoned). Add 'timezone' to RESUME_CONTEXT_KEYS (persisted at pause, replayed by
the resume middleware) and restore conversationCreatedAt from the convo before
initializeClient — mirroring the normal path's resolveConversationCreatedAt.

F5 (P2) — multi-tool approval: applyPendingActionToMessages stopped retrying once
ANY tool-call part was tagged, so siblings that rendered on later frames never got
approval controls and the resume route 400'd the partial batch. Add
countTaggedApprovalParts and keep the bounded RAF retry going until every
action_request is tagged (ask_user_question unchanged — one synthetic part).

F6 (P3) — Edit accepted `null`/`[]` (valid JSON, non-object), enabling Submit for
a value the resume route rejects via findIncompleteDecisions. Mirror the server's
plain-object check in the client (store + editIsValid) so Submit only enables for
an accepted value.

Tests: policy.spec (timezone round-trip), resume.spec (conversationCreatedAt
restore), approval.spec (countTaggedApprovalParts).

* fix(hitl): address Codex re-review (round 13) — recurse into subagent approvals

F9 (P2) — a tool paused INSIDE a subagent has its tool_call_id in the parent
subagent tool_call's nested `subagent_content`, not as a top-level message part.
applyToolApproval and countTaggedApprovalParts only scanned top-level content, so
the approval never attached and the round-12 retry loop counted 0 tagged parts and
spun to its frame cap with no controls. Both now recurse into `subagent_content`
(immutably, so React refs update): the nested call gets tagged and is counted, so
the retry terminates. Added approval.spec cases for the nested tag + count.

Note: surfacing the interactive approve/reject controls inside the subagent view is
a deliberate follow-up — ToolApproval -> useResumeSubmit -> useChatContext crashes
when rendered in the portaled subagent dialog (outside the chat/approval providers),
so that needs the controls scoped to the in-provider inline render (or the dialog
wrapped with the providers). This commit fixes the data/traversal layer only.

F7 (discovered-tool history on resume) and F8 (redis chunk TTL pause race) were
verified false positives — see the PR threads.

* fix(hitl): address Codex re-review (round 14) — resume fidelity + expiry relay

F13 (P2) — manualSkills are graph-determining (skill allowed-tools union into the
tool set before tools load) but weren't replayed, so a reload lost the skill tools
and a crafted resume could inject a different skill past the fingerprint. Add
'manualSkills' to RESUME_CONTEXT_KEYS (same replay-only pattern as timezone/
addedConvo; the delete-absent half blocks injection). Not alwaysAppliedSkills —
that's resolved server-side from the DB, not req.body.

F12 (P2) — the resume final SSE built requestMessage from job.metadata.userMessage
(persisted without files), so attachments vanished from the user bubble on resume.
Spread the already-restored req.body.files onto it, matching the normal path.

F11 (P2) — multi-replica approval expiry: RedisJobStore.cleanupRequiresActionIndex
on another replica can win the requires_action->aborted CAS (it sets the hash error
but has no event transport), and the local sweep then skips because the job is no
longer requires_action, so a client subscribed here never gets the terminal error
until the reap path. expireStaleApprovals now relays APPROVAL_EXPIRED_ERROR for a
locally-subscribed job already aborted FOR approval expiry (error-string gated,
idempotent via the errorEvent flag). emitError already publishes cross-replica.

Tests: policy.spec (manualSkills round-trip + inject-drop), resume.spec (final
requestMessage carries restored files).

* fix(hitl): render approval controls for subagent-nested tool pauses (F10)

Round-13 made applyToolApproval/countTaggedApprovalParts recurse into
subagent_content (data), but SubagentDialogPart rendered nested TOOL_CALL parts
with <ToolCall> only and never mounted <ToolApproval>, so a tool paused inside a
subagent showed no controls and the run was unresolvable.

Render <ToolApproval> in SubagentDialogPart's TOOL_CALL branch when the nested
tool_call carries an approval and isn't yet resolved, mirroring the top-level
Part.tsx render. The subagent dialog portals (OGDialog → ReactDOM.createPortal),
but React context flows through the React tree, not the DOM tree, so ToolApproval
resolves ApprovalProvider/ChatContext and the controls work + submit.

Also harden useResumeSubmit: read ChatContext via useContext (non-throwing)
instead of the throwing useChatContext wrapper, so the cards never crash when
rendered outside a ChatContext.Provider (e.g. a search/citation render that passes
chat context as a prop) — they degrade to inert (buildResumeFields returns null).

* style(hitl): re-sort run.ts imports after dev rebase

* fix(hitl): address Codex re-review (round 15) — resume content fidelity

F14 (P2) — hide_sequential_outputs was applied in chatCompletion before
saving/emitting content but not on resume, so a sequential-agent chain that
pauses for HITL and resumes persisted/emitted intermediate outputs the setting
is meant to hide. Extracted the filter into applyHideSequentialOutputsFilter()
and call it from both chatCompletion and resumeCompletion (after handleRunInterrupt,
covering the finalize + re-pause reads of client.contentParts).

F16 (P2) — on a reloaded HITL pause, the DB already holds the paused user row +
partial assistant row; useResumeOnLoad fed those as submission.messages, then
finalHandler/createdHandler appended the same pair via requestMessage/responseMessage,
duplicating the turn (buildTree doesn't dedupe children by messageId). buildSubmission-
FromResumeState now strips the paused user/response rows (by messageId, incl. the
padded/unpadded response id) from submission.messages — they're re-supplied by the
placeholders + final event. Frontend-only; live (non-reload) pause path untouched.

Deferred: F15 (collapsed-card subagent approval registration/visibility) — see thread.

Tests: client.test (filter keeps last + tool_call parts / no-op when off),
useResumeOnLoad.spec (paused pair stripped from submission.messages).

* fix(hitl): address Codex re-review (round 16) — chunk TTL, slot, job replacement

F17 (P2) — chunk-stream TTL on pause-before-chunk. CHUNK_APPEND_LUA derived its
ceiling only from the chunk key's current TTL, so when the chunks key didn't exist
at pause (fire-and-forget append in flight, or an ask-user pause before any chunk),
the on_pending_action append created the stream with only the 20m running TTL while
the approval window is 24h — content evicted before resume. The Lua now also reads
the job key (KEYS[2]); when status == requires_action it takes max(running, TTL(jobKey))
(the approval window transitionStatus set), else the running TTL. Extend-only preserved;
gated on paused status so normal runs never inflate. Both keys share {streamId} (cluster-safe).

F19 (P2) — with LIMIT_CONCURRENT_MESSAGES, the approval prompt was emitted before the
original request released its slot, so a fast Approve got /resume 429'd. handleRunInterrupt
now releases the slot (idempotent via pendingRequestReleased) right after the pause, before
the prompt; the request.js pause branch and resume.js finally only release if it didn't
(no double-release).

F20 (P2) — finalizeResumedTurn never checked the job wasn't replaced before emitDone/
completeJob/saveMessage, so a stale resume could clobber a newer turn that reused the
conversationId. Added the createdAt guard the normal request path uses (skip finalization
when the live job's createdAt != the paused job's).

Deferred: F18 (subagent_content not reconstructed on Redis resume) — joins the subagent
cluster (F15). See thread.

Tests: RedisJobStore integration (pause-before-chunk gets approval TTL; running stays short),
resume.spec (skip finalization on replacement; no double slot release on re-pause).

* 🛡️ fix: Guard HITL terminal side-effects against job replacement

Jobs are keyed by streamId == conversationId, so a new request REPLACES the
running one on the same conversation. The replaced generation's tail must not
clobber the live generation's state. Each path now re-reads the live job and
compares createdAt against the generation's captured identity before acting.

- Thread the generation's createdAt onto the client (request.js + resume.js)
  as client.jobCreatedAt — the identity every guard compares against.
- handleRunInterrupt: skip approvals.pause when this run is no longer the live
  job, so a stale interrupt can't flip the NEWER job to requires_action.
- chatCompletion finally: skip the checkpoint prune when replaced, so an older
  run's late finally can't delete the newer run's resume checkpoint.
- resume catch-path: gate emitError/completeJob/prune behind a stillLive check
  (fail-open if the read throws), mirroring finalizeResumedTurn's success guard.
- Persist the turn's uploaded files on job.metadata.userMessage (authoritative
  trackUserMessage writer) and prefer them on resume over the user DB row, whose
  save can still be racing a fast /resume.

Tests: 13 guard-predicate cases in jobReplacement.spec.js.

* 🔁 fix: Harden HITL resume — ownership re-check, file seeding, deferred-tool replay

Three follow-ups to the round-17 job-replacement guards (Codex review 4594099963):

- G1 (resume.js): the success-path ownership guard runs at the START of
  finalizeResumedTurn, but saveMessage + first-turn title generation await long
  enough for a new request to replace the job on the same conversationId. Re-read
  the live job immediately before emitDone/completeJob/prune so the terminal writes
  can't tear down the REPLACEMENT job — mirrors the catch-path guard.

- G2 (request.js): onStart's metadata/chunk writes that persist the turn's files
  are fire-and-forget, so a fast approval could read job.metadata.userMessage before
  files landed. Seed files into getPreliminaryUserMessage instead — that write is
  AWAITED before the run starts, so files are durable before any interrupt can emit.

- G3 (run.ts + client.js + resume.js + IJobStore.ts): the resumed graph is rebuilt
  with messages: [], so createRun's tool_search-discovery scan finds nothing. A
  deferred tool discovered earlier in the turn (and targeted by the paused call) was
  therefore absent from the rebuilt schema-only toolMap — resume would throw "unknown
  tool" (no loadRuntimeTools fallback is wired). Capture discovered tool names at
  pause via extractDiscoveredToolsFromHistory(run.getRunMessages()), persist them on
  job.metadata.discoveredTools, and replay them into createRun's new discoveredToolNames
  input (merged with message-extracted names, gated on hasAnyDeferredTools — inert
  otherwise). A new createRun test proves the deferred tool is promoted with the replay
  and absent without it (reproducing the bug).

Tests: real createRun deferred-replay suite (run-summarization.test.ts) + G1/G2/G3
guard predicates (jobReplacement.spec.js). Full suite green.

* 🔒 fix: Close HITL resume metadata + file-substitution + pause-race gaps

Four findings on the round-18 commit (Codex review 4594430222):

- H1 (P1, regression in round-18 G3): the discoveredTools captured at pause never
  reached resume — three metadata allowlists dropped it: GenerationJobManager
  .updateMetadata, RedisJobStore.deserializeJob, and buildJobFacade (plus the
  GenerationJobMetadata type). Added discoveredTools to all four, so the deferred-tool
  replay actually works end-to-end (in-memory store already kept it via Object.assign).

- H2 (P2, security): /resume honored a client-supplied `files` array, letting a crafted
  client resume an approved code/read-file tool against a DIFFERENT file set than the one
  approved (files aren't in the resume fingerprint/context). Resume now ALWAYS sources
  files from the paused job (metadata → DB row), clearing any client-supplied set.

- H3 (P2, ephemeral fidelity): non-default model parameters (temperature, max tokens,
  custom endpoint params) were lost on resume — ephemeral agents derive them from the
  request body, which the resume payload omits. Capture the resolved model_parameters in
  resumeContext at pause and replay them onto the body on resume (excluding `model`, which
  is replayed via the fingerprinted RESUME_CONTEXT_KEYS path). Saved agents already source
  these from the DB.

- H4 (P2, Redis race): a pause landing between the resume snapshot and the Pub/Sub
  subscription reached neither resumeState.pendingAction nor (Redis) pendingEvents, and
  approval events aren't persisted to replayEvents — the client attached to a paused job
  with no approval UI. subscribeWithResume now re-reads the live job AFTER subscribing and
  surfaces the pending action if the snapshot missed it (live read, no staleness).

Tests: discoveredTools metadata round-trip + subscribeWithResume re-read (pendingAction
.spec.ts); client-file substitution rejection (resume.spec.js); model-parameter replay
predicate (jobReplacement.spec.js).

* 🧹 fix: Clear stale discovered tools, release slot on claim error, extend run-step TTL

Three follow-ups on the round-19 commit (Codex review 4594783691):

- I1 (P2): the round-19 discoveredTools field wasn't cleared on Redis streamId reuse.
  HSET only overwrites listed fields and handleRunInterrupt only writes discoveredTools
  when THIS turn discovers a deferred tool — so a replacement turn that pauses without its
  own discovery inherited the prior run's tool names and force-loaded undiscovered deferred
  tools on resume. Added discoveredTools to createJob's staleHitlFields HDEL list (the
  in-memory store already builds a fresh object, so it was Redis-only).

- I2 (P2): with LIMIT_CONCURRENT_MESSAGES, approvals.resolve runs after the slot increment
  but before the run's try/finally, so a store/Redis error there leaked the slot until the
  counter TTL expired (spurious 429s on retry of the still-paused approval). Wrapped the
  claim in try/catch that decrements the slot and returns 500.

- I3 (P3): saveRunSteps did SET ... EX running unconditionally, resetting the run-steps key
  to the 20-min running TTL even while the job is paused for the longer approval window —
  a reload after that window lost the tool timeline. Now uses a paused-window TTL script
  mirroring the chunk-stream no-shrink behavior (extends to the approval window when the
  job hash is requires_action).

Also fixes a latent strict-tsc cast error in the round-19 pendingAction test.

Tests: claim-throws-releases-slot (resume.spec.js); discoveredTools cleared on reuse +
saveRunSteps preserves the paused TTL (RedisJobStore integration, USE_REDIS).

* 🛡️ fix: Guard fast-resume save race, gate HITL to resumable routes, expire on stale submit

Three findings on the round-20 commit (Codex review 4595045652):

- J2 (P1): a fast /resume can claim + finalize the COMPLETED response while the original
  request's pause branch is still awaiting `response.databasePromise`; the later
  unfinished-save then overwrites the completed content. Re-check the job is still paused on
  THIS generation's action (a claim leaves requires_action; a replacement bumps createdAt)
  before marking the row unfinished; fail open on a read error.

- J3 (P1): the tool-approval wiring (humanInTheLoop + PreToolUse hook + checkpointer) was
  applied to EVERY createRun caller when toolApproval.enabled, but the OpenAI-compatible and
  Responses controllers never inspect run.getInterrupt() or persist a pending action — an
  approval-gated tool would pause there with no approval surface or resume endpoint and the
  route would emit a normal final response / [DONE] with the tool call dangling. Gate the
  wiring on a new createRun `hitlCapable` flag, set only by AgentClient (chat + resume).

- J4 (P2): a stale-action 409 on submit returned without driving expiry, leaving the job
  requires_action with a dead action until the periodic sweeper ran — any attached SSE client
  got no terminal event and the stream appeared to hang. Extracted GenerationJobManager
  .expireApproval(streamId, actionId) (expire CAS + terminal SSE, shared with the sweeper) and
  call it from the resume route when the observed action is stale.

J1 (nested subagent approval controls not mounting while the details dialog is closed) is a
valid frontend issue in the deferred subagent-HITL path — tracked separately (replied on the
thread) since the fix touches the shared dialog primitive and needs UI verification.

Tests: HITL-gate both directions (run-summarization.test.ts); expire-on-stale-submit
(resume.spec.js); fast-resume unfinished-save guard predicate (jobReplacement.spec.js).

* 💄 style: Wrap captureAgents signature to satisfy prettier (CI lint)
2026-06-29 16:56:41 -04:00