Commit graph

1602 commits

Author SHA1 Message Date
Danny Avila
7132fdcf32 🐛 fix: Hydrate deferred attachments for later tool calls
An attachment accepted on one turn whose tool never ran carried neither an
embedding nor a code reference, and every hydration query matches only files
that already have one: getToolFilesByIds requires embedded, getUserCodeFiles
requires a codeEnvRef. The file was therefore absent from later turns entirely,
so asking to search or run code against it found nothing.

Hydrating it back into attachments would have re-delivered earlier uploads to
the model on every turn, so provisioning and delivery are now separate inputs:
deferred records are fetched by their own query and reach the provisioning
computation alone, never the returned attachments.

Also provisions for the host create_file tool, whose name is distinct from
write_file, and falls back to the primary context when a tool batch omits its
agent id, matching what the tool loaders already do.
2026-08-31 14:22:50 -04:00
Danny Avila
b8c7e0b109 🐛 fix: Create the code API axios instance on first use
Sharing the provisioning callback pulled provision.js into the OpenAI-compatible
controllers, where it called createAxiosInstance at module load. Those suites
partially mock @librechat/api, so every test in both files failed on import
before any provisioning was requested. The instance is now created on first use,
which removes the load-time side effect rather than widening the mocks.
2026-08-31 12:42:55 -04:00
Danny Avila
637ba51b12 🐛 fix: Provision attachments on the OpenAI-compatible agent APIs
The chat path provisioned queued attachments at tool execution, but the
/v1/chat/completions and /v1/responses controllers assembled their tool execute
options without a provisionFiles callback, so those surfaces loaded execute_code
or file_search and ran the tool without the attachment. Eager provisioning used
to make files available everywhere, so lazy provisioning regressed them. The
callback moves into a shared factory that all three call sites use; the agent
tool contexts these controllers already build carry the provisioning state.
2026-08-31 12:35:24 -04:00
Danny Avila
b5a274cbb3 🐛 fix: Provision search files for programmatic tool calls
A file_search configured with allowed_callers: ['code_execution'] is reachable
only through run_tools_with_code, whose nested tool names never reach the
provisioning predicate, so the nested search ran before queued attachments were
embedded and silently searched without them. Programmatic tool calls now mark
the turn as needing search.
2026-08-31 11:24:16 -04:00
Danny Avila
1ca246b3bd 🧪 test: Assert denial for permanent agent uploads without permission
This case asserted that an upload naming an agent but no tool_resource skipped
the permission check, which is the behavior the unified path made exploitable.
It also described itself as a message attachment while sending none; the
adjacent case already covers real message attachments. It now asserts denial.
2026-08-31 11:06:12 -04:00
Danny Avila
f804930954 🐛 fix: Provision for file-editing tools and keep failed entries queued
edit_file and write_file read their target from the code environment but did not
mark a turn as needing code, so an edit issued straight after an upload reported
the file missing. Provisioning failures also cleared the queue and continued into
tool loading; failed entries are now preserved for a later retry, and a code
provisioning failure aborts the preflight rather than answering from input the
sandbox never received. Vector failures still only narrow results, so they do
not abort.
2026-08-31 10:58:28 -04:00
Danny Avila
57609ba137 🐛 fix: Fail tool preflight when code provisioning cannot persist
primeCodeFiles reads the persisted record and skips files with no code
environment ref, so an unpersisted upload leaves the tool running against an
attachment it cannot see. Persistence now retries once, and a still-unpersisted
code reference aborts the preflight instead of executing code without the file.
2026-08-31 10:35:35 -04:00
Danny Avila
2cd5076f50 🐛 fix: Retain auto-routed context uploads and isolate vector temp files
An agent upload with no explicit tool_resource that resolves to the text path is
promoted to a context resource, but retention still received the original
undefined tool_resource, so in retention modes that expire conversation files
while keeping persistent agent resources these files were given an expiration
and swept.

Vector provisioning also derived its temp path from file_id alone, so two
concurrent requests for the same file shared one path and could unlink or
truncate it mid-stream, producing failed or corrupted embeddings.
2026-08-31 10:11:03 -04:00
Danny Avila
bc2e9c868f 🐛 fix: Only stream provisioning sources with the standard contract
Provisioning called getDownloadStream(req, filepath) for every source, but that
contract is not universal: openai takes (file_id, client) and execute_code takes
(fileIdentifier, identity, req) and returns an Axios response, so files backed by
those sources were mis-invoked rather than provisioned. Streaming now goes
through an allowlist of storage-backed sources, so an unfamiliar source is
skipped with a warning instead of called incorrectly.
2026-08-31 09:49:03 -04:00
Danny Avila
230d2a7908 🐛 fix: Treat failed liveness probes as unknown, not expired
A timeout or 5xx while probing a code env session left its files out of the
alive set, so the staleness path cleared a potentially live ref and forced a
re-upload the same outage would likely fail, losing the file for that turn. Only
a successful response that omits the file now marks it expired.

Also corrects TProvisionToCodeEnv, which still declared the pre-merge codeEnvRef
result while the implementation and its consumer use referenceSet, and logs
rejected provisioning persistence instead of discarding it.
2026-08-31 08:52:21 -04:00
Danny Avila
57ed43b51c 🐛 fix: Gate provider media defaults on endpoint capability
The system default sent every PDF, video, and audio file down the provider path
even for endpoints with no native document encoder, so the upload succeeded and
the model received neither the file nor extracted text. The system default now
falls back to text for a known endpoint that cannot encode documents; explicit
endpoint or global config still wins, images are unaffected, and an unknown
endpoint keeps the existing behavior.
2026-08-31 08:38:05 -04:00
Danny Avila
f0978e7d70 🔒 fix: Keep generated code artifacts user-scoped when re-provisioning
Code outputs are recorded with kind: 'user' at generation time, but the
provisioning scope predicate treated any context other than message_attachment
as agent-scoped. An expired artifact re-provisioned on a later turn was
therefore uploaded into the agent's shared sandbox, exposing one user's private
conversation artifact to every user of a shared agent. Both the provisioning
writer and the tool-resource reconstruction now share one predicate that treats
execute_code outputs as user-scoped alongside chat attachments.
2026-08-31 08:38:05 -04:00
Danny Avila
e13226608d 🧪 test: Fix planner assertion for auto-routed text uploads
The new test stubbed getStrategyFunctions with only handleFileUpload, so the
document-parser path it exercises threw before the assertion ran. It now
mirrors the passing OCR strategy tests and asserts only the planner argument,
which is what the fix changed.
2026-08-31 08:28:07 -04:00
Danny Avila
590698e5ee 🐛 fix: Plan extraction with the promoted context resource
An auto-routed text upload promotes its effective resource to context, but the
extraction planner still received the original undefined tool_resource and
returned no document-parser plan. On deployments without RAG a text-routed
DOCX/ODT/XLSX then reached parseText with native fallback and was decoded as
UTF-8 instead of parsed.
2026-08-31 08:14:06 -04:00
Danny Avila
9c6cb14e01 🐛 fix: Reconcile lazy code-env provisioning with execution routes
Code env pointers are deployment-local, but the liveness probe always queried
the default Code API. With the staleness repair now live, a ref belonging to a
configured stateful route would fail that probe and be cleared, re-uploading a
still-valid file to the wrong deployment. Only default-route refs take part in
the check and only they can be cleared.

Lazy provisioning also wrote a bare codeEnvRef while the eager upload path
persists through mergeCodeEnvRef. Both paths now write the same shape, so the
legacy pointer and the route-keyed map stay in sync and pointers for other
routes survive re-provisioning.

Provisioning computation moves into a helper that both return paths call, so a
turn carrying no request attachments still queues the agent's persistent
context files instead of skipping them at the early return.
2026-08-31 07:37:58 -04:00
Danny Avila
7e458649ef 🐛 fix: Name provisioned images by their stored MIME type
Image uploads are converted to appConfig.imageOutputType while the file
record keeps the original filename, so code-env provisioning shipped webp
bytes as photo.jpg and extension-sniffing tools mis-handled them. Uploads now
rename known converted image types to match the persisted MIME type.
2026-08-30 22:33:47 -04:00
Danny Avila
106f8d7fcb 🐛 fix: Resolve agent provider for image upload delivery paths
processImageFile resolved llmDeliveryPath from the request endpoint (agents),
ignoring metadata.agent_id, so provider-specific defaultLLMDeliveryPath and
legacyFileUploadUX config never applied to images even though
processAgentFileUpload already honored it. Both paths now share
resolveUploadEndpoint; the raw endpoint still drives image resize/storage.
2026-08-30 21:37:16 -04:00
Danny Avila
6a1a67242b 🐛 fix: Provision under JWT code auth + repair stale codeEnvRef re-check
Lazy code-env provisioning was gated on a loaded LIBRECHAT_CODE_API_KEY, so
JWT-auth deployments (which mint bearer tokens via getCodeApiAuthHeaders and
need no legacy key) silently never provisioned attachments. The gate now
accepts either auth mode and checkSessionsAlive composes X-API-Key with the
minted bearer headers.

Separately, pre-categorization added every codeEnvRef file to
processedResourceFiles before the provisioning loop ran, so the staleness
branch was unreachable and expired sandbox refs were never cleared or
re-provisioned. Staleness is now repaired ahead of the processed guard,
clearing both the legacy ref and its route entry so getCodeEnvRefs cannot
resolve the dead session.
2026-08-30 21:37:12 -04:00
Danny Avila
781e6679e7 🐛 fix: Persist provisionedAt + scope lazy file_search to agent resources
Round-3 Codex follow-ups completing the round-2 fixes:
- The file schema/type dropped codeEnvRef.provisionedAt, so the liveness fast-path
  was dead after reload. Add provisionedAt to CodeEnvRef + the Mongoose subschema.
- Lazily provisioned agent-scoped file_search files were only added to
  tool_resources.file_search.files, so primeFiles treated them as user attachments
  and queried without entity_id, missing the agent-scoped vectors. Add agent-scoped
  files to file_ids instead so they are queried with entity_id.
2026-08-30 20:33:55 -04:00
Danny Avila
665912350c 🧪 test: Assert provisionedAt on codeEnvRef in process.spec
Finding B stamps codeEnvRef with a provisionedAt timestamp at code-env upload
time; update the four exact-shape assertions to expect it (expect.any(Number)).
2026-08-30 20:12:23 -04:00
Danny Avila
90f2e45587 🧪 test: Cover agent-provider file config resolution
Mock getAgent in the process.spec ~/models mock (my C fix calls db.getAgent) and
add a test: an agent upload (endpoint=agents) whose provider has a none fallback
resolves llmDeliveryPath from the provider config, not the generic agents config.
2026-08-30 20:12:23 -04:00
Danny Avila
81f7f122b4 🐛 fix: Code-env freshness marker + agent-provider file config
Addresses two Codex round-2 P1 findings:
- Liveness: checkSessionsAlive trusted file.updatedAt to skip the code-env live
  check, but updateFilesUsage bumps updatedAt on resend, so a usage-touched file
  with an expired sandbox session was wrongly treated as alive. Stamp codeEnvRef
  with provisionedAt at upload time and gate the fast-path on that instead; refs
  without the marker fall through to a live check.
- Routing: an agent upload carries endpoint=agents, so provider-specific
  defaultLLMDeliveryPath overrides (endpoints.<Provider>) were ignored and a
  none/tool-only file could be persisted as text/provider. Resolve the file config
  from the agent's own provider when agent_id is present.
2026-08-30 20:12:23 -04:00
Danny Avila
3f167089a0 🐛 fix: Correct scope + tool visibility for lazily provisioned files
Addresses Codex P1 findings on the lazy provisioning path:
- Scope per file like the direct upload path: current-message chat attachments
  (context=message_attachment) provision to the user's code sandbox / unscoped
  vector index (entity_id undefined); only agent setup files use entity_id=agentId.
  Previously every lazily provisioned file was agent-scoped, so a user's chat
  attachment landed in the agent sandbox and file_search queries (unscoped for
  fromAgent=false) missed the agent-scoped embedding.
- Surface provisioned files to the tool loaded immediately after by adding them to
  ctx.tool_resources.<resource>.files, which primeCodeFiles/primeFiles read. Before,
  a freshly provisioned unified upload was invisible to the first code/file_search
  call even though provisioning had completed.
2026-08-30 20:12:23 -04:00
Danny Avila
0774cc2d8b 🐛 fix: Provision unified uploads for bash_tool/read_file code execution
provisionFiles only treated execute_code / run_tools_with_code as code execution,
but that capability now expands into bash_tool / read_file / run_tools_with_bash.
So a unified upload in provisionState.codeEnvFiles was never provisioned before the
first bash_tool/read_file call — the file was missing from the sandbox (Codex P1).

Broaden the needsCode guard to the current tool names. The lazy-provisioning e2e
now emits the actually-advertised code-exec tool (bash_tool) so it exercises the
real path. Removes the temporary diagnostics.
2026-08-30 20:12:23 -04:00
Danny Avila
b84594d450 🔍 chore: TEMP diagnostics for lazy code provisioning (revert) 2026-08-30 20:12:23 -04:00
Danny Avila
79e3c07c85 🐛 fix: Load code API key from LIBRECHAT_CODE_API_KEY
EnvVar.CODE_API_KEY was removed from @librechat/agents, so loadCodeApiKey
resolved authFields to [undefined], threw on undefined.split in loadAuthValues,
and lazy code-env provisioning silently bailed with a warning. Read the
canonical LIBRECHAT_CODE_API_KEY (symmetric with LIBRECHAT_CODE_BASEURL) and
degrade gracefully when unset.
2026-08-30 20:12:09 -04:00
Danny Avila
e621535665 🔀 chore: align lazy provisioning with codeEnvRef schema
Rebase onto current dev brought in the metadata.fileIdentifier →
metadata.codeEnvRef migration (HEAD uploadCodeEnvFile now returns
{ storage_session_id, file_id } and requires kind/id). Update the
unified-upload code paths to match:

- provision.js: provisionToCodeEnv now derives kind/id from entity_id,
  calls uploadCodeEnvFile with the new signature, and returns codeEnvRef
- checkSessionsAlive/checkCodeEnvFileAlive: read storage_session_id and
  remote file_id from metadata.codeEnvRef instead of parsing the legacy
  fileIdentifier string
- resources.ts: primeResources gates on metadata.codeEnvRef and clears
  it on staleness; TProvisionToCodeEnv reflects the new return shape
- initialize.js: provisionFiles closure destructures codeEnvRef
- process.spec.js: align two legacyFileUploadUX tests with the
  endpoint-level check landed in 7384947 and update the execute_code
  expectation to the codeEnvRef metadata shape
- resources.test.ts: import FileSources for the typed source field and
  guard the optional attachments map
2026-08-30 20:12:09 -04:00
Atef Bellaaj
19872a2537 🔧 fix: legacy file upload UX checks in processFileURL and processAgentFileUpload 2026-08-30 20:12:09 -04:00
Atef Bellaaj
7eb0428917 🔧 feat: Unified file upload — per-mime-type routing with lazy provisioning 2026-08-30 20:12:09 -04:00
Danny Avila
b4bac55422 🔧 feat: Lazy file provisioning — defer uploads to tool invocation time
Move file provisioning from eager (at chat-request start) to lazy
(at tool invocation time via ON_TOOL_EXECUTE). Files are now only
uploaded to code env / vector DB when the LLM actually calls the
respective tool.

- resources.ts: primeResources no longer provisions; computes
  provisionState (which files need code env / vector DB uploads)
  with staleness check and single credential load
- handlers.ts: add provisionFiles callback to ToolExecuteOptions,
  called once per tool-call batch before execution
- initialize.ts: pass provisionState through InitializedAgent
- initialize.js: implement provisionFiles closure that provisions
  files in parallel, batches DB updates, clears state after use;
  store provisionState in agentToolContexts for all agent types
2026-08-30 20:12:09 -04:00
Danny Avila
1201de11e2 🧹 chore: Optimize provisioning — single credential load, deferred DB writes
- loadCodeApiKey: load CODE_API_KEY once per request, pass to both
  checkSessionsAlive and provisionToCodeEnv (was N+1 lookups)
- provisionToCodeEnv/provisionToVectorDB now return fileUpdate objects
  instead of writing to DB immediately
- primeResources batches all DB updates via Promise.allSettled after
  provisioning completes
- Remove updateFile import from provision.js (no longer writes directly)
2026-08-30 20:12:09 -04:00
Danny Avila
1f72ea28bf 🔧 feat: Unified file experience — schema, deferred upload, lazy provisioning
Phase 2 fixes for the unified file experience:

- Add code env file staleness detection via batch session checks
  (checkSessionsAlive) — groups files by session_id, one API call per
  session, skips files updated within 6h safe window
- Parallelize file provisioning across files using Promise.allSettled
- Surface provisioning failures as warnings on InitializedAgent
- Fix temp file path safety (use file_id + extension, not raw filename)
- Fix inconsistent return types (normalize to [] instead of undefined)
- Wire checkSessionsAlive through initialize.js → initialize.ts →
  primeResources
2026-08-30 20:12:08 -04:00
Danny Avila
af164ae7ce 🔧 feat: Unified file experience — schema, deferred upload, lazy provisioning
Introduces the foundation for a unified file upload experience where users
upload files once without choosing a tool_resource upfront. Files are stored
in the configured storage strategy and lazily provisioned to tool environments
(execute_code, file_search) at chat-request time based on agent capabilities.

Phase 1 - Schema + Server-Side Unified Upload:
- Add FileInteractionMode enum (text/provider/deferred/legacy) to fileConfigSchema
- Add defaultFileInteraction field to EndpointFileConfig and FileConfig types
- Update mergeFileConfig/mergeWithDefault to propagate the new field
- Modify processAgentFileUpload to support uploads without tool_resource
  using effectiveToolResource resolved from config (default: deferred)

Phase 2 - Lazy Provisioning + Multi-Resource Support:
- Create provision.js with provisionToCodeEnv and provisionToVectorDB
- Extend primeResources with lazy provisioning step that provisions
  deferred files to enabled tool environments at chat-request start
- Remove early returns in categorizeFileForToolResources so files can
  exist in multiple tool_resources simultaneously
- Wire provisioning callbacks through initializeAgent dependency injection
2026-08-30 20:12:08 -04:00
Danny Avila
fcae1025c0
🧳 feat: Register Principal-Owned Code Environments (#15365)
* feat: add principal-owned code environments

* fix: address code environment CI coverage

* fix: expose principal code environments in endpoint config

* fix: reject missing code environment bodies

* fix: harden code environment lifecycle

* fix: revalidate principal code environments

* fix: synchronize code environment authorization

* fix: bind code environments to current principals

* fix: fail closed on code ACL cache errors

* fix: prevent code environment override shadowing

* fix: preserve principal environment defaults

* fix: suppress revoked code environment aliases

* fix: fail closed on environment augmentation

* fix: narrow code environment defaults

* fix: isolate code environment fallback

* style: sort code config imports
2026-08-30 19:33:19 -04:00
Danny Avila
a9ccac8656
🧲 feat: Enable Secure Attached Environment Pairing (#15355)
* feat: add secure code environment pairing

* fix: satisfy code environment type checks

* fix: secure code environment administration

* fix: isolate code pairing control plane

* fix: validate code pairing control responses

* fix: secure code pairing transport

* fix: validate code pairing wire format

* fix: harden pairing secret lookup
2026-08-30 17:12:20 -04:00
Danny Avila
7533d138fa
🧬 perf: Evolve Compaction Guidance on Warm Turns (#15371)
* perf: evolve compaction guidance on warm turns

* style: sort compaction adapter imports
2026-08-30 17:11:20 -04:00
Danny Avila
30124f21b2
🎻 refactor: Orchestrate Agent Runs Through a Request-Free Host (#15366)
* refactor: decouple agent initialization from HTTP

* refactor: centralize remote agent execution lifecycle

* fix: preserve pre-settlement error rendering

* fix: preserve request-backed tool loading

* style: sort agent execution imports

* fix: adapt public agent tool loaders
2026-08-30 15:38:57 -04:00
Danny Avila
c1cb591d49
🗿 feat: Add Attached Stateful Code Environments (#15352)
* feat: add attached stateful code environments

* fix: include code environment in lazy agent type

* fix: harden stateful environment routing

* fix: complete code environment route isolation

* fix: declare code environment map type

* fix: preserve configured code execution routes

* fix: harden stateful environment updates

* test: assert route-scoped sandbox readiness

* fix: harden stateful environment lifecycle

* fix: isolate migrated code sessions

* test: assert route-qualified code sessions
2026-08-30 15:38:45 -04:00
Danny Avila
cd3768ed1f
🍵 feat: Continue Late Steers in Warm Agent Runs (#15357)
* feat: continue late steers in warm agent runs

* chore: bump agents sdk to v3.7.9

* fix: guard terminal steer admission
2026-08-30 11:54:17 -04:00
Danny Avila
2ae6c8aea9
🛎️ feat: Wake Agents on Background Tool Completion (#15350)
* feat: wake agents for background tool completion

* fix: isolate background completion contracts

* fix: anchor background completion identity

* fix: close background completion delivery gaps

* fix: preserve manual background polling

* fix: preserve legacy tool group identity

* fix: close background wakeup identity gaps

* test: type background wakeup enqueue mock

* fix: preserve background completion identity

* test: type activity phase fixture

* test: mock phase media query

* fix: arbitrate background result ownership

* fix: retain tool step routing helper

* test: expand phase groups before identity checks

* fix: preserve background completion ownership

* fix: bound durable background results

* fix: type wakeup input budget export

* chore: sort background handler imports

* fix: persist timed-out background completions

* fix: reconcile background completion ownership

* fix: type background completion capabilities

* fix: harden background completion terminalization

* test: type missing completion evidence

* fix: require evidence for completion retirement

* chore: satisfy background completion static checks

* test: cover automatic background completion wakeups

* test: assert background wakeup agent identity

* fix: wake capability-fenced trigger deliveries

* test: keep capability shield fixture public

* test: assert capability worker wakeup

* fix: close background completion ownership gaps

* test: satisfy completion lease static checks

* fix: preserve artifact completion wakeups

* fix: close background delivery recovery gaps

* fix: simplify background receipt guidance

* fix: recover dead background completion batches

* fix: fence background completion recovery

* fix: recheck background recovery ownership

* fix: fence unpublished background continuations

* test: await background message hydration
2026-08-30 08:56:41 -04:00
Danny Avila
e3ccaba5af
🛄 feat: Restore Compaction Guidance Across Continuations (#15356)
* 🧭 feat: preserve compaction guidance across continuations

* 🔧 fix: narrow persisted compaction fields
2026-08-30 08:13:24 -04:00
Danny Avila
70f735336d
🛎️ fix: Enroll Remote Agent Runs in the Generation Lifecycle (#15349)
* fix: enroll remote agent runs in generation lifecycle

* fix: close remote lifecycle ownership gaps

* fix: close remote conversation drain races

* fix: fence remote runs during conversation deletion

* fix: reconcile remote deletion and settlement races

* fix: complete owner deletion recovery

* fix: preserve remote cleanup receipts

* fix: consume deletion receipts before cleanup

* fix: expose idempotent deletion option

* chore: sort remote lifecycle imports
2026-08-30 07:14:13 -04:00
Dustin Healy
8fcab7e44f
🔄 fix: Recover Missing MCP Marketplace Catalogs (#15323)
* fix: recover missing MCP marketplace catalogs

* fix: make MCP catalog recovery passive

* test: type MCP catalog recovery fixtures

* fix: bound and back off passive MCP catalog recovery

Passive recovery runs inline on `GET /api/mcp/tools` and its results are
request-local by design, so every list request re-dialed the same cold
servers with the default connection timeout. Three limits keep that cost
proportional to what recovery can actually recover:

- Cap the discovery timeout at 5s instead of inheriting the connection
  default (`initTimeout ?? 30s`); a server configured to connect faster
  keeps its own shorter limit.
- Skip a server the config tier already marked `inspectionFailed`, leaving
  it to that tier's retry window rather than re-dialing it per request.
- Skip a server whose declared `customUserVars` are unset, matching the
  gate `reinitMCPServer` applies for issue #10969 — connecting without them
  fails auth, so the attempt is spent for nothing.

Servers that still fail discovery enter a one-minute per-process cooldown,
which is what stops an unreachable server from being re-dialed by every
subsequent list request. A server that recovers clears its own entry, and
expired entries are swept at most once per window so the map stays bounded.

Skipped servers render exactly as they did before recovery existed: present
in the catalog with an empty tool list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xr1Dabvdn1mzzyYgpJgU5B

* fix: bound passive MCP recovery by deadline, key cooldowns by config

Both follow-ups address the same mistake: recovery expressed its own
request-level constraints in terms borrowed from other layers.

`connectionTimeout` bounds one connection attempt, and
`MCPConnectionFactory.discoverToolsInternal` spends it twice — once on the
authenticated connection, then again in `attemptUnauthenticatedToolListing`
— so capping it bounded no total this layer could reason about. Recovery now
enforces its own wall-clock deadline per server with `withTimeout`, which
holds however many attempts the factory makes; `connectionTimeout` is left to
do only its own job, still honouring a shorter operator `initTimeout`. An
attempt abandoned by the deadline disposes its own connection when it
settles, and `Promise.race` keeps a handler on it, so a late rejection is
not unhandled.

A per-request budget now caps total recovery regardless of server count.
A server is dialed only if the remaining budget can fund a full deadline;
never dialing one is not evidence against it, so a skipped server records no
cooldown and a later request reaches it once those ahead are cached or
cooling down.

Cooldown identity now includes the publication generation — the same
effective-config identity the tool caches fence on — instead of just user and
server name. Correcting a server's URL or transport keys a new entry, so the
refetch the client issues on update is no longer skipped for up to a minute
by the previous configuration's failure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xr1Dabvdn1mzzyYgpJgU5B

* refactor: keep passive MCP recovery stateless and bounded by its own work

Reverts the cooldown, request budget and deadline race added in 257d5cf and
fc3e3c9, and keeps only the three stateless limits.

The tool cache refuses unfenced writes (`tools.ts`), and a discovery
connection owns no publication generation and is disposed, so a recovered
catalog cannot be retained by design. Those commits responded by building a
cache-shaped memory in front of it — per-process failure state, a scheduling
budget, an identity, an eviction sweep — and each round of review found
another way that hand-rolled cache differed from a real one: wrong identity
for configuration, wrong identity for credentials, no fairness across
requests, and a limiter slot released while its network operation was still
running. None of that machinery was asked for; all of it was compensation for
a result the architecture does not allow keeping.

Recovery is now stateless. It skips only what configuration alone proves
pointless — a server the config tier already marked `inspectionFailed`, and
one whose declared `customUserVars` are unset — and bounds the work itself
rather than racing it, so a limiter slot is held for exactly as long as its
network operation runs and the concurrency limit of three is real.

The attempt timeout is not a compromise: recovery exists for a server that is
reachable and authorized but whose catalog cache expired, and such a server
answers tools/list well inside 1.5s. Anything slower cannot be rescued here,
so failing fast costs nothing. The factory spends that value per attempt, so
a server's ceiling is it times the attempts made; the constant documents that
rather than hiding it behind a number tuned to today's attempt count.

Consequences that were bugs are now gone by construction: every cold server
is attempted on every request, so none is starved by those ahead of it, and
correcting a server's configuration or credentials takes effect on the next
refetch instead of waiting out a stale cooldown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xr1Dabvdn1mzzyYgpJgU5B

* fix: correct the inspection-failure skip and bound catalog fan-out

Three fixes that belong to this layer; a fourth issue does not, and is
described below.

The `inspectionFailed` skip was too broad. `MCPServersInitializer` stores a
YAML server that was unreachable at startup via `addServerStub`, which stamps
`source: 'yaml'`, and only config-tier entries get the timed retry in
`ensureSingleConfigServer`. Skipping every failed stub therefore hid a
recoverable server from the marketplace permanently — the exact state this
recovery exists to escape. It now defers only `source === 'config'`, matching
what `reinitMCPServer` already does.

Plugin auth is read only when some cold server actually declares
`customUserVars`, and only for those servers. The common unauthenticated case
no longer pays a MongoDB round trip whose result nothing can consume.

Snapshot refreshes are now bounded by the same limiter as discovery. They are
not local reads: both connection paths reach `fetchOrderedToolsSnapshot` and
issue a real `tools/list`, so a cache reset across many servers previously
burst unbounded outbound requests while discovery was capped at three.

Not fixed here, because it cannot be: `connectionTimeout` does not bound
discovery. It covers `connection.connect()` only, and `fetchToolsSnapshot`
then applies its own `TOOLS_LIST_TIMEOUT_MS` (30s) to `tools/list`, so a
server that connects fast and stalls while listing still holds its slot for
that window. The factory also does not cancel a timed-out connect before
starting the unauthenticated fallback. Bounding this end to end needs a
deadline threaded through `MCPConnectionFactory` into both `connect()` and
`fetchToolsSnapshot()`, which is a change to shared connection machinery
rather than to this caller. The constant's comment now states what it does
and does not bound instead of implying an end-to-end guarantee.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xr1Dabvdn1mzzyYgpJgU5B

* fix: thread live-session OBO context into passive catalog discovery

The merge of #15334 sources OBO tokens from the live OpenID session via
request-boundary closures. Passive catalog recovery is a discovery call
site too; without these options an OBO server whose stored token went
stale fails recovery — the exact cold-catalog class this PR fixes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xr1Dabvdn1mzzyYgpJgU5B

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-30 06:55:30 -04:00
Danny Avila
16c2bb4149
📬 feat: Establish Background Continuation Admission (#15348)
* feat: establish background continuation admission

* fix: preserve routed subagent wakeup controls

* fix: probe configured subagent task store
2026-08-30 06:51:36 -04:00
Danny Avila
fa913148fb
🔒 fix: Refresh MCP OBO Tokens From the Live OpenID Session (#15334)
* 🧊 fix: Inline-refresh OpenID session tokens at MCP OBO call time

Resolves the walk-away failure mode where MCP tool calls using OBO auth
fail with "No valid OpenID access token is available for OBO exchange"
after a user idles past their access-token lifetime. The strategy-time
snapshot on `user.federatedTokens` could expire mid-stream before
`resolveOboToken` ran, while `req.session.openidTokens` carried a still-
valid (or refreshable) token that nothing read.

- New OpenIDSessionRefresh service: per-user single-flighted closure that
  reads `req.session.openidTokens` at OBO time and inline-refreshes via
  `openid-client.refreshTokenGrant` when expired (30s skew), persisting
  via `req.session.save()`. No cookie writes (headers already flushed).
- `resolveOboToken` gains a required UpstreamTokenProvider parameter
  (typed as `() => Promise<OIDCTokens | null>`, reusing the shared shape
  from @librechat/data-schemas). Compile-time guarantee that every call
  site is updated.
- New `session_refresh_failed` OboTokenResolutionReason distinguishes
  "session expired and IdP rejected refresh" from "no upstream token
  ever existed."
- `req` threaded through createMCPTool/createMCPTools/createToolInstance
  to construct the closure with captured request, plus fail-closed
  guards in MCPConnectionFactory.getOboTokens and MCPManager.callTool
  when the closure isn't plumbed.
- Startup warning in MCPServersInitializer when OBO is configured but
  OPENID_REUSE_TOKENS is unset (the strategy populating
  user.federatedTokens is only registered under reuse, so OBO would
  fail every call without it).

Tests: 16 new in OpenIDSessionRefresh.spec.js; obo.spec.ts extended
for the new param + error reason; wiring smoke tests in MCPManager,
MCPConnectionFactory, MCPServersInitializer, and MCP.spec.js.

* 🛡️ fix: Harden OBO inline-refresh against token type and session edge cases

- Token-preference asymmetry: live-token reuse and expires_at derivation
  now strictly gate on the access_token, not the id_token. Added a
  required `tokenPreference` parameter on isLiveSessionTokenStillValid,
  buildOIDCTokensFromSession, and createOpenIDSessionTokenProvider
  so every call site is explicit. Dropped the bogus id_token-exp
  fallback in performIdpRefresh — id_token TTL is governed by IdP
  session policy and would mark a short-lived access_token reusable
  past its real lifetime.
- Missing req in /reinitialize route: the manual reconnect
  endpoint now forwards req into reinitMCPServer, so OBO servers can
  build a session-aware upstream-token closure instead of failing with
  missing_upstream_token.
- Single-flight key collisions: composed key as
  tenantId:openidIssuer:openidId:sessionId via getSingleFlightKey.
  Concurrent calls in the same session still coalesce; separate sessions
  never share an in-flight refresh, preventing refresh-token rotation
  from breaking sibling sessions and preventing cross-tenant token
  crossover when distinct users share an IdP sub.
- Opaque access token reuse): persist accessTokenExpiresAt
  (unix seconds, from tokenset.expires_in) on each refresh AND on initial
  login / SPA refresh in setOpenIDAuthTokens. New getAccessTokenExp
  helper falls back to it when the access token isn't a JWT, avoiding
  redundant inline refreshes for Microsoft Graph and Auth0 default
  audiences.
- Log hygiene: the single-flight key (containing sessionId,
  openidId, openidIssuer, tenantId) is now SHA-256-hashed in the
  "Joining in-flight refresh" debug log. Preserves cross-line correlation
  via a 12-char prefix without leaking credential or PII material.

Documented req.session.openidTokens shape contract via JSDoc typedef so
the new accessTokenExpiresAt field has a discoverable home alongside the
existing accessToken/idToken/refreshToken/expiresAt/lastRefreshedAt.

Tests: OpenIDSessionRefresh.spec.js up to 30 passing (added coverage for
opaque-token reuse, JWT-access-token-exp fallback, no-id_token-fallback
regression, cross-session no-coalesce, persistence on refresh, and a
guard against stale accessTokenExpiresAt carryover). AuthService.spec.js
adds two cases covering accessTokenExpiresAt persistence on login.
mcp.spec.js (route) gains a regression test asserting req flows into
reinitMCPServer.

* 🔍 fix: Detect OBO-only MCP admin config overrides

Admin Config overlays for YAML-defined MCP servers compare only
ADMIN_CONFIGURABLE_FIELDS to decide whether to lazy-init a config-tier override.
The OBO config field was added after that fingerprint list, so an override that
only added or changed `obo` was treated as unchanged YAML and skipped.

Include `obo` in the admin-configurable field list and add a regression test for
an OBO-only override.

* 🔊 fix: Mock MCP OAuth timeout in SDK integration test

MCPConnectionFactory.attemptToConnect reads mcpConfig.OAUTH_HANDLING_TIMEOUT
when building the OAuth connection timeout. The SDK OAuth integration test
mocked mcpConfig without that field, which made the timeout calculation produce
NaN and caused the test to fail before the OAuth refresh/start path completed.

Add OAUTH_HANDLING_TIMEOUT to the test mock.

* ♻️ refactor: Pass OBO upstream-token closure into MCP instead of req

Build the OpenID upstream-token provider at the request boundary and thread
only the closure through MCP handling, so the MCP service layer no longer
receives the raw Express request. The closure still reads/refreshes the live
session at tool-call time, preserving the walk-away recovery.

- Drop `req`/`capturedReq` from createMCPTools, createMCPTool, reconnectServer,
  createToolInstance, and reinitMCPServer; forward `upstreamTokenProvider`
  instead. Closure is constructed in loadTools, loadToolDefinitionsWrapper, and
  the reinitialize route, where req/res are in scope.
- OBO: fall back to user.federatedTokens when the provider yields no live
  session, so OIDC remote-agent calls (verified bearer, no session) still work.
- Inline refresh: mirror a rotated refresh token to the refreshToken cookie via
  a shared setRefreshTokenCookie helper, guarded by !res.headersSent (no-op on
  the streaming path; session copy stays authoritative).
- Single-flight: hydrate a joining request's own session from the resolved
  tokens so a later OBO call doesn't replay a rotated-away refresh token.

Addresses owner feedback and three review findings.

* 🔒 fix: Recover OIDC refresh-token rotation after SSE OBO refresh

When an inline OBO refresh rotates the OpenID refresh token after SSE headers
have already been sent, the browser refreshToken cookie cannot be updated. Store
a short-lived encrypted bridge from the stale cookie token to the rotated token
so /api/auth/refresh can recover after express-session loss.

Use the signed openid_user_id cookie to load user context for bridge validation,
retry only on invalid_grant, and delete the bridge only after the bridged refresh
succeeds.

* 🔨 fix: hydrate joined OIDC refresh sessions with stable refresh tokens

Update single-flight OIDC refresh joiners whenever refreshed access token
state changes, even if the IdP keeps the refresh token unchanged.

This prevents joined requests from retaining stale accessToken or
accessTokenExpiresAt values and redundantly refreshing later in the same run.

* 🌉 Persist OIDC refresh-token recovery bridges in MongoDB

Store SSE OBO refresh-token recovery bridges in MongoDB instead of
process-local memory so /api/auth/refresh can recover after worker
restarts or cross-worker routing.

Derive bridge expiry from REFRESH_TOKEN_EXPIRY so the recovery window
matches the stale refreshToken cookie it repairs, and delete bridges
after successful recovery.

* 🤝 Coordinate OIDC inline refreshes across workers

Add a short-lived Mongo-backed refresh-flight record so concurrent
OBO refreshes for the same OpenID session do not redeem the same
rotating refresh token on different workers.

The winning worker performs the IdP refresh and stores an encrypted
result; joiners wait for that result, hydrate their request session,
and return without calling the IdP.

*  Keep OpenID marker cookies aligned on inline refresh

Refresh token_provider and openid_user_id with the same expiry as the
rotated refreshToken cookie when an inline OBO refresh can still write
headers.

Share the marker-cookie writer with the normal OpenID auth refresh path
so the fallback /api/auth/refresh branch continues to recognize valid
OpenID refresh tokens after session expiry.

* 🔑 fix: include refresh token in OIDC local refresh flight key

Key the process-local OIDC refresh coalescing by the current session
refresh token, matching the Mongo-backed flight key. This prevents a
request with a newly rotated token from joining an older pending refresh
and inheriting its failure/result.

* 🌉 fix: store OIDC refresh bridge without cookie response

Treat missing or non-cookie responses like headers-sent streaming
responses during inline OIDC refresh. When the IdP rotates the refresh
token and cookies cannot be written, persist a recovery bridge so a later
/auth/refresh can recover after session expiry.

* 🫙 fix: preserve stale OIDC cookie bridge key

Track the refresh token last written to the browser cookie separately
from the current session refresh token. When inline OIDC refreshes rotate
tokens without a writable response, keep bridging from the browser-stale
token directly to the latest session token.

* 🙌 fix: keep OIDC bridge recovery success on cleanup failure

Make refresh-token bridge cleanup best-effort after a bridged OIDC
refresh succeeds. A transient delete failure now logs a warning but does
not convert the already-refreshed session and cookies into a 403 response.

* 📦 test: Exclude RefreshTokenBridge from tenant-isolation coverage

Add RefreshTokenBridge to the tenant-isolation coverage allowlist because
refresh bridge lookups run during unauthenticated OpenID refresh recovery.
The controller first recovers user context from the signed OpenID marker
cookie, then the bridge methods apply explicit user and tenant filters.

Ambient tenant isolation would bind this recovery path to request-local
tenant context that is not available at the point the stale cookie is being
resolved

*  Fix OpenID refresh flight retry and marker hydration

Allow failed OpenID refresh flights to be reclaimed immediately instead of pinning transient errors.

Preserve the browser refresh-token marker when joined refreshes hydrate session tokens from a shared flight result.

Stabilize AuthService tests by isolating mocked module imports from prior suites.

* 🛠️ fix: centralize OBO identity scoping

Add shared auth identity helpers for app user ids, OpenID subjects,
tenant ids, and normalized OpenID issuers.

Thread a non-placeholder-visible OBO identity context from the real
request user through MCP connection, tool-call, reinit, and refresh
paths. Keep tenantId and openidIssuer out of createSafeUser so MCP
user placeholders do not expose those fields.

Scope OBO token cache and in-flight exchange keys by tenant, issuer,
OpenID subject, scopes, and a SHA-256 hash of the upstream assertion.
This prevents cross-tenant/cross-issuer collisions and avoids reusing
tokens minted from stale rotated assertions.

Use the shared identity helpers for OpenID refresh-flight keys and
refresh-token bridge recovery records so related OBO refresh paths share
the same identity normalization rules.

The helper is intended for auth-boundary and credential-cache code, not
as a blanket replacement for ordinary app user id ownership checks.

* 🛠️ fix: preserve OIDC refresh-token sync on save failures

Sync OpenID refresh-token cookie/bridge state before persisting the
session so a transient session-store failure cannot lose an IdP-rotated
refresh token.

Also trigger sync when the session refresh token differs from the
browser refresh-token marker, not only when the current grant rotates
the token. This lets later writable refreshes repair stale browser
cookies left behind by SSE refreshes.

Route refresh bridge identity through the shared identity helper with
the threaded OBO identity context, falling back to request/user context
when needed.

Add regression coverage for session-save failures, stale browser cookie
repair, non-writable bridge storage, and shared-helper identity fallback.

* 🛠️ fix: keep OIDC refresh bridge during recovery grace

After successful bridged refresh recovery, re-store the stale-cookie bridge
with a short grace TTL instead of deleting it immediately. This lets parallel
/api/auth/refresh requests that already sent the stale browser cookie recover
before they can observe the first response's Set-Cookie.

Retarget the bridge to the refresh token returned by the bridged retry so
B-to-C refresh-token rotation remains recoverable. The grace TTL is parsed with
math() and defaults to 60s, which shrinks the replay window from the original
REFRESH_TOKEN_EXPIRY bridge lifetime to the short recovery grace period.

Remove the now-unused explicit bridge delete path from the service and
data-schemas method surface. Add regression coverage for grace re-store,
identity symmetry, retry failure behavior, and same-key upsert replacement.

* 🛠️ fix: fail closed on OBO MCP user identity mismatch

Add an OBO-specific guard before MCP tool execution that requires the
effective invocation user and captured request user to both have ids and
to match. This prevents OBO tool calls from falling back to a separate
configurable.user_id identity after request-bound OBO context has already
been captured.

Keep the existing user id fallback behavior for non-OBO MCP calls.

Tests cover mismatched OBO users, missing user ids, and the matching-user
path ignoring a conflicting configurable.user_id.

* 🛠️ fix: Guard OpenID bridge retry user identity

Extract the shared OpenID refresh/user-resolution flow in AuthController
so the normal refresh path and bridge-recovery retry use the same grant,
claims, issuer, user lookup, and diagnostic logging code.

Preserve the existing path-specific behavior: the normal path still owns
migration updates and 401 login redirects, while the bridge retry still
falls through to the existing 403 invalid-token response.

Add a bridge-recovery guard that rejects retry results whose resolved
user id differs from the signed openid_user_id cookie before issuing
tokens or re-storing the grace bridge. Cover both the successful
matching-user recovery and the mismatched-user rejection.

* 🛠️ fix: type-safety polish on OBO data layer

Replace refresh token bridge query/update Record<string, unknown> usage
with typed Mongoose FilterQuery and UpdateQuery definitions.

Harden OpenID marker cookie JWT expiry handling by converting refresh
expiry milliseconds to integer seconds and rejecting invalid or
non-positive durations.

Add focused CSRF tests for fractional refresh expiry values and invalid
expiry configuration.

* 🛠️ fix: Bind OpenID session tokens to authenticated identity

Stamp OpenID session token state with the LibreChat user id, OpenID subject,
tenant id, and normalized issuer when tokens are stored.

Fail closed before OBO inline token reuse/refresh when the session token
identity does not match the current authenticated identity, preventing a stale
or mixed Express session from supplying another user's upstream assertion.

Also validate the normal /api/auth/refresh session-token reuse shortcut against
the signed marker-cookie user before returning cached session tokens.

Note: sessions created before this change carry no identity stamp and are
treated as a mismatch. This is self-healing — the reuse path forces a full IdP
refresh (which re-stamps the session) and the OBO path throws, surfacing as a
one-time re-authentication for active OBO users at deploy time. The session
re-stamps within one session lifetime (SESSION_EXPIRY, default 15 min).

* 🛠️ fix: Recover OpenID refresh token drift

Prefer the browser refresh-token cookie when it differs from the
server-side OpenID session state, and force a real IdP refresh in that
case instead of reusing stale session tokens.

Store a short-lived refresh-token bridge when inline OBO refresh writes
a rotated browser cookie but session persistence fails, so follow-up
refreshes can still recover from the old token.

Keep the bridge grace TTL centralized in RefreshTokenBridge so both
recovery paths use the same env-backed value.

Note: drift is measured against the last-synced browserRefreshToken
marker, so the SSE path (intentionally stale cookie, authoritative
session) does not false-positive. Sessions predating the marker have no
browserRefreshToken; for those, drift falls back to comparing the cookie
against the session refresh token and prefers the cookie on difference.
This is the same self-healing pre-change-session window as the identity
binding fix and re-syncs within one session lifetime.

Tests cover cookie/session drift selection, reusable-session bypass on
drift, bridge storage after session-save failure, and the shared bridge
constant wiring.

* 🛠️ fix: Harden OBO token caching and expiry handling

Reject malformed OBO grant responses before writing them to the exchanged-token cache so a missing access_token cannot poison the cache.

Store absolute expires_at values with cached OBO tokens and ignore legacy cache entries without usable expiry metadata. This keeps cached-token freshness based on the token’s real
remaining lifetime instead of reusing the original relative expires_in on cache hits.

Move OBO expiry normalization and skew helpers into packages/api and use them from both the JS exchange service and the TS MCP resolver. Apply a 30-second safety margin with a one-
second floor for short-lived tokens, covered by direct helper tests and caller-level regression tests.

Tests:
- packages/api: npm run build
- packages/api: npx jest src/mcp/oauth/expiry.spec.ts src/mcp/oauth/obo.spec.ts
- api: npx jest server/services/OboTokenService.spec.js

* 🛠️ fix: Harden OBO refresh-token bridge lookup and indexing

Reuse getValidOpenIDReuseUserId for the bridge-recovery user lookup in
refreshController instead of re-verifying openid_user_id inline. The shared
helper enforces the JWT_REFRESH_SECRET presence check and a strict
typeof payload.id === 'string' guard, rejecting tokens whose id claim is
present but not a string (e.g. a numeric id) that the inline check accepted.

Fail closed on issuer mismatch in getRefreshTokenBridge. Both the stored and
the expected issuer are now normalized and compared for equality, so a bridge
is recovered only when both sides agree (both absent, or both present and
equal after normalization). Previously the check was skipped whenever the
stored issuer was absent, allowing recovery across mismatched issuer context.

Drop the unused {oldRefreshTokenHash, userId, tenantId, openidIssuer} index
and the openidIssuer field on RefreshTokenBridgeQuery. The data-layer filter
only queries the 3-field {oldRefreshTokenHash, userId, tenantId} index; the
issuer is verified in application code, not the query. Hoist the repeated
model accessor into getRefreshTokenBridgeModel.

Note: issuer is now load-bearing for recovery. A bridge stored with an issuer
recovers only when the lookup supplies a matching issuer; the recovery lookup
reads user.openidIssuer via AUTH_REFRESH_USER_PROJECTION (an exclusion
projection that retains the field). If a user's persisted openidIssuer is
empty while the stored bridge has one, recovery fails closed (falls through to
normal re-authentication) until the bridge TTLs out — no security regression.

Tests cover invalid signed-cookie payloads bypassing the bridge, both
asymmetric issuer-presence cases, issuer normalization before comparison, and
an index-alignment assertion guarding against re-adding the dropped index.

* 🛠️ fix: Degrade OBO discovery on token resolution failures

Catch expected OboTokenResolutionError failures during MCP tool discovery and
fall back to unauthenticated tool listing instead of aborting discovery. This
keeps discovery aligned with the existing unauthenticated listing behavior while
preserving unexpected errors as real failures.

Also correct OBO tool-call freshness comment and tighten the OBO trust-check
permissions type to the existing role permission shape.

Tests:
- npx jest src/mcp/__tests__/MCPConnectionFactory.test.ts --runInBand --coverage=false
- npx jest src/mcp/oauth/obo.spec.ts --runInBand --coverage=false

* 🛠️ fix: tighten OBO tool-call errors, bridge logging, and flight typing

Move resolveToolCallUserId inside the tool-call try/catch so an OBO
identity mismatch surfaces with serverName/toolName context and the
standard tool-call-failed message instead of an opaque bare Error.

Raise the refresh-token bridge lookup failure log from debug to warn so
transient infrastructure failures on the unauthenticated /api/auth/refresh
path are observable, and guard the message access against non-Error values.

Replace the unknown+cast in isDuplicateKeyError with a hasErrorCode type
predicate so the duplicate-key check reads error.code without an assertion.

Preserve real math/isEnabled in the MCPConnectionFactory test mock (mock
only processMCPEnv) so mcpConfig timeouts no longer resolve to NaN, fixing
the TimeoutNaNWarning that masked slow OAuth retry behavior.

* 🧪 fix: Restore the Flight Uniqueness Index and Buffer the Graph Cache TTL

Two CI failures on the merge, both in suites this environment cannot run
(their MongoDB binary download is blocked).

`GraphApiService.spec.js` still asserted the unbuffered TTL. Graph tokens
route through the same `getTokenCacheTtlMs` as the OBO and openidStrategy
caches, so the entry now expires 30s before the credential does.

`openidRefreshFlight.spec.ts` dropped the database between tests, which takes
the indexes with it, and Mongoose builds them only once when the model is
compiled. Whether the unique `key` index survived into a test was a race with
that one-time build. Without it a second `create` inserts instead of raising a
duplicate-key error, so every worker believes it won the flight — the
mutual exclusion the file exists to prove. Indexes are now rebuilt after each
drop, which also makes the reclaim and complete cases reach those paths for
the right reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: address OBO review findings

* 🔐 fix: Install Bridge Indexes and Carry OBO Through Assistant Recovery

Two findings from the Codex pass on d08c82f0d.

The refresh-token bridge relied on Mongoose auto-indexing for both of its
indexes, and `MONGO_AUTO_INDEX=false` is a supported deployment setting. A
bridge holds an encrypted refresh token and the TTL index is the only thing
that ever deletes one, so under that setting they would accumulate for the
life of the collection while concurrent upserts lost the compound uniqueness
the filter assumes. Installed before the first write, matching the flight
methods and the session and schedule methods before them.

`recoverServerTools`, the assistant create/update path that reruns
`reinitMCPServer` when a referenced server's catalog and connection snapshot
are both missing, was the last reinit site not carrying the upstream-token
closure. For an OBO server the factory rejects the connection outright, so the
assistant write failed with unavailable MCP definitions. It now builds the
provider at that request boundary like the other entry points; assistant
writes have no `res`, so a rotation there falls back to the recovery bridge.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: harden OBO refresh coordination

*  fix: Keep Elapsed Expiries Elapsed and Revoke the Superseded Session

Two of the five findings from the Codex pass on fcdc15885 — the two that are
defects in code this branch introduced rather than design questions about the
bridge.

`getSkewedTokenExpiresAtMs` floored every result at a second in the future,
including an expiry the provider had already declared elapsed. An exchange
answering `expires_in: 0` or a past `expires_at` was handed to the MCP
connection stamped valid for another second, which only moves the failure
downstream. The floor now applies to a lifetime that is still live, which is
what it was for; an elapsed one stays elapsed so the caller rejects it. Same
for the cache TTL, which falls back to the elapsed-credential floor.

Bridge recovery left the stale token's durable Session behind. Only the token
it recovered through was passed as `existingRefreshToken`, so that one's
session was replaced while the token the browser actually presented kept its
record until its original expiry. That record, with the marker cookie still
bound to it, is what authorizes local image access for OpenID users — so a
copy of the stale cookie outlived the rotation it had lost. Revoked
explicitly on successful recovery.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: close OBO refresh review findings

* 🎟️ fix: Carry the Bridged Token Through a Non-Rotating Recovery

Bridge recovery passes the browser's stale token as `existingRefreshToken` so
the durable Session naming it is the record replaced. That also makes the
stale token the fallback the installed session and the refresh cookie use when
a tokenset carries no `refresh_token` of its own, which holds only while the
recovery grant rotates.

An IdP that answers that grant without rotating sends the browser back to the
very token the bridge exists to retire: `storeOpenIDSession` installs and
deletes the same stale record in one call, and the cookie is rewritten to a
token the IdP already rejected — a sign-out on the next refresh. The grace
bridge one line above already guards this with `|| bridgedRefreshToken`; the
resolved tokenset now does the same, so leader and followers alike publish the
recovered token.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* 🚫 fix: Reject an OBO Exchange That Returns an Expired Credential

Preserving an elapsed expiry through the skew helper only helps if something
acts on it, and nothing did: `MCPManager.callTool` checks the access token and
nothing else before setting the Authorization header, so a credential the IdP
declared spent still went downstream to fail there. It is rejected at the
exchange now, where the reason is known, and retryably — the exchange itself
worked, so a fresh grant can succeed.

Completes the elapsed-expiry change in cb26a6f7d, which made the stamp honest
without giving anyone a reason to look at it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: close OpenID refresh review findings

* fix: coordinate OpenID refresh entry points

* chore: sort OpenID flight imports

* fix: fence OpenID refreshes during logout

* fix: close OpenID logout publication races

* fix: narrow completed refresh flight

* test: cover bridge cleanup failure after ownership loss

The compensating delete in storeRefreshTokenBridgeWithLease swallows its own
failure so the lease error stays the one the caller sees. Nothing asserted
that, so removing the inner catch left every suite green while callers began
receiving the cleanup error instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: compensate OpenID bridges only on proven ownership loss

The post-write lease assertion deletes the bridge it just published when it
throws, but it threw for two different reasons: a coordination record that is
no longer ours, and a coordination read that simply failed. Treating the second
as the first destroys the only mapping from the token the browser still holds
to the one the IdP already rotated to, so a transient Mongo error on the
headers-already-sent path signed the user out.

Tag the ownership error where the lease raises it and compensate only for that,
preserving the bridge whenever ownership is merely undetermined. A preserved
bridge stays behind the logout revocation fence, so the safe default costs
nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: reject spent OpenID refresh results

Two ways a refresh could report success while handing back a credential
nothing can use.

An inline refresh carries the previous id_token forward when the IdP omits
one on rotation, so tokenset.id_token is not necessarily freshly issued.
setOpenIDAuthTokens applied its freshness guard only to the session copy and
took tokenset.id_token unconditionally, so /refresh returned an expired
bearer even though the grant produced a usable access token. Skip it only
when it is provably expired: an id_token whose expiry cannot be read stays
preferred, since access_token may be opaque or scoped to another audience.

normalizeExpiresIn preserves a zero or negative lifetime rather than
discarding it, so a grant declaring an already-spent access token still
published, rotating the refresh token and returning a token every freshness
check rejects. Each OBO call then repeated the grant. Reject an elapsed
lifetime before publishing; an unknown lifetime still publishes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* fix: harden OpenID refresh publication

* fix: keep identity and results intact through OpenID refresh cleanup

Two follow-ons from the last round's fixes.

Stripping an expired carried-forward id_token from the refresh result removed
the only identity material a rotation without id_token leaves behind. The
result is rebuilt by buildOIDCTokensFromSession, so it carries no provider
claims() either, and getTokenClaims accepts only those two — bridge recovery
failed with "no usable identity claims" before setOpenIDAuthTokens could hand
back the fresh access token. The stripped token now travels in a
non-enumerable marker, alongside the existing browser and predecessor markers,
which identity resolution reads and the authentication response never sees.

The lease drained a pending renewal by awaiting it in finally, so a transient
coordination failure there threw from finally and replaced the operation's
result. The refresh had already settled and published, so the caller saw a
failure on credentials that had rotated. Proven ownership loss is recorded on
ownershipLost and checked before the return, so the drain has nothing to add
but noise; it now absorbs and logs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SxKWxwqxAGckYpRsYTqx3F

* refactor: fence OpenID recovery publication

* fix: close OpenID publication transaction

* refactor: make OpenID publication transactional

* fix: satisfy OpenID publication type checks

* fix: fence OpenID session publication

* fix: authorize OpenID refresh publication

* fix: bind OpenID replay generations

* fix: fence OpenID response generations

* fix: authorize OpenID token delivery

* fix: linearize OpenID publication delivery

* chore: sort OpenID refresh flight imports

---------

Co-authored-by: J.C. Bartle <jcbartle@users.noreply.github.com>
Co-authored-by: jbartle <jbartle@rand.org>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: jcbartle <7274202+jcbartle@users.noreply.github.com>
2026-08-29 23:55:58 -04:00
Danny Avila
10981b8637
🍱 feat: Guide Compaction With a Bounded Semantic Index (#15340)
* 🧭 feat: guide compaction with semantic context

* 🛡️ fix: fail closed on semantic intent collisions
2026-08-29 17:53:04 -04:00
Danny Avila
7a0061507e
🪪 fix: Admit Confirmed Generation Retries Before Message Limits (#15341)
* fix: admit idempotent retries before message limits

* fix: bound generation retry admission

* fix: keep retry probe store-compatible

* fix: exclude normalized resume routes

* fix: preserve trusted retry exemptions

* fix: bound retry claim admission

* style: satisfy generation retry static checks
2026-08-29 13:41:26 -04:00
Danny Avila
773127bff2
🎠 refactor: Route Every Event Actor Turn Through One Lifecycle (#15325)
Some checks are pending
Backend Unit Tests / Build packages (push) Waiting to run
Backend Unit Tests / Codegraph select (push) Waiting to run
Backend Unit Tests / TypeScript type checks (push) Blocked by required conditions
Backend Unit Tests / Circular dependency checks (push) Waiting to run
Backend Unit Tests / Tests: api (shard 1/3) (push) Blocked by required conditions
Backend Unit Tests / Tests: api (shard 2/3) (push) Blocked by required conditions
Backend Unit Tests / Tests: api (shard 3/3) (push) Blocked by required conditions
Backend Unit Tests / Tests: data-provider (push) Blocked by required conditions
Backend Unit Tests / Tests: data-schemas (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 1/4) (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 2/4) (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 3/4) (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 4/4) (push) Blocked by required conditions
Codegraph E2E Votes / vote (full suite) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
* refactor: unify Event Actor turn lifecycle

* fix: retain Event Actor fence ownership

* fix: preserve mixed-version actor suspension safety
2026-08-28 17:30:23 -04:00
Danny Avila
3fa33b740b
🛫 refactor: Promote Generation Protocol V2 Automatically (#15324)
* refactor: promote generation protocol v2 automatically

* fix: remove unused protocol import
2026-08-28 17:17:09 -04:00
Danny Avila
77c2a51cf3
🪧 fix: Advertise Detached Event Actor Support from the Generation Store (#15322)
Some checks failed
Backend Unit Tests / Build packages (push) Waiting to run
Backend Unit Tests / Codegraph select (push) Waiting to run
Backend Unit Tests / TypeScript type checks (push) Blocked by required conditions
Backend Unit Tests / Circular dependency checks (push) Waiting to run
Backend Unit Tests / Tests: api (shard 1/3) (push) Blocked by required conditions
Backend Unit Tests / Tests: api (shard 2/3) (push) Blocked by required conditions
Backend Unit Tests / Tests: api (shard 3/3) (push) Blocked by required conditions
Backend Unit Tests / Tests: data-provider (push) Blocked by required conditions
Backend Unit Tests / Tests: data-schemas (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 1/4) (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 2/4) (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 3/4) (push) Blocked by required conditions
Backend Unit Tests / Tests: @librechat/api (shard 4/4) (push) Blocked by required conditions
Codegraph E2E Votes / vote (full suite) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
Frontend Unit Tests / Codegraph select (push) Has been cancelled
Frontend Unit Tests / Build packages (push) Has been cancelled
Frontend Unit Tests / TypeScript type checks (client) (push) Has been cancelled
Frontend Unit Tests / Tests: @librechat/client (push) Has been cancelled
Frontend Unit Tests / Tests: Ubuntu (shard 1/2) (push) Has been cancelled
Frontend Unit Tests / Tests: Ubuntu (shard 2/2) (push) Has been cancelled
Frontend Unit Tests / Vite build verification (push) Has been cancelled
* fix: activate detached event actions automatically

* fix: shield detached terminal generations

* perf: share capability availability index

* style: flatten capability status selection

* fix: preserve mixed-version lifecycle shells

* fix: complete detached action store adapters

* fix: close mixed-version capability races

* fix: honor legacy capability success
2026-08-28 16:33:23 -04:00