Commit graph

5081 commits

Author SHA1 Message Date
Danny Avila
c7e355b219
🛑 fix: Confirm Scheduled Stops and Separate Terminal Scheduler Failure from Startup (#15051)
* 🛑 fix: Confirm Scheduled Stops and Separate Terminal Scheduler Failure from Startup

Fixes #15042, fixes #15043.

`resume.js` inferred a confirmed stop from the ABSENCE of `failureReason`, but
`abortJob` had four `success: false` paths that returned no reason at all. Those
settled the occurrence as `interrupted` and pruned the checkpoint on aborts that
never landed — including one where a REPLACEMENT generation owned the
conversation, which pruned the successor's checkpoint.

Every `success: false` return now names itself (`job_not_found`,
`already_settled` added alongside the existing `generation_replaced` /
`job_still_active`), and a single canonical `isStopConfirmed` predicate decides
whether durable state may be settled. `already_settled` confirms a stop —
`awaitProviderDrain` has proven the provider segment can no longer persist — so a
permanently terminal generation is not answered with a retry loop.

Separately, a schedule engine that failed to arm advertised its permanent outage
as a transient 503 with `Retry-After`, so a client obeying it would poll forever.
Readiness is now tri-state (`starting` / `armed` / `unavailable`): the retry
contract applies only while arming is genuinely pending, and a failed arm returns
a terminal `SCHEDULES_UNAVAILABLE` with no `Retry-After` and an error-level log.

* 🏷️ fix: Declare Schedule Write Gate Return Types

`--isolatedDeclarations` requires an explicit return type on the exported factory
and on the middleware it returns (TS9007). Adds a named `ScheduleWriteGate` type
matching the existing `ShareMiddleware` shape.
2026-08-20 18:33:51 -04:00
Danny Avila
9c9696de8b
🧵 fix: Hide Child Threads From Navigation (#15041)
* fix: hide subagent threads from conversation lists

* fix: exclude child threads from search indexes

* fix: fence child search index cleanup

* fix: preserve private search exclusion markers

* fix: wait for cleanup through Meili client

* fix: preserve lean message update results

* fix: make search cleanup acknowledgments retryable

* fix: complete child search cleanup migration
2026-08-20 18:33:25 -04:00
Danny Avila
276f5f88fe
🗓️ fix: Hide Unsupported Schedule Variables (#15053) 2026-08-20 18:33:00 -04:00
Alexandre Beguel
6baeff801c
📐 docs: Make CLAUDE.md Portable and Restore AGENTS.md Parity (#15046)
CLAUDE.md pointed at /home/danny/agentus for the @librechat/agents source.
That path resolves on one machine, so for every other contributor the line
occupies context on every turn while pointing nowhere. Replaced with the
public repository URL.

AGENTS.md opens with "See CLAUDE.md", which makes CLAUDE.md the source of
truth, but it carried an auth cache invalidation rule that appears nowhere
in CLAUDE.md. Claude Code reads CLAUDE.md and not AGENTS.md, so contributors
using it never received a rule about serving stale req.user.

Moved that rule into a Backend Rules section in CLAUDE.md, mirroring the
existing Frontend Rules section, and left AGENTS.md pointing at it the same
way its theming paragraph already does.

Closes #15045
2026-08-20 17:28:49 -04:00
Danny Avila
868cfbb343
🎡 fix: Scheduled Chat Slot Accounting and Reconciliation Rotation (#15040)
* fix(schedules): harden limits and reconciliation

* docs(schedules): align reconciliation invariants

* fix(schedules): preserve prompt editor contract
2026-08-20 17:18:41 -04:00
Danny Avila
b4593f80b7
📦 chore: bump @librechat/agents to v3.6.9 (#15037) 2026-08-20 12:44:21 -04:00
Danny Avila
81783b2d52
🧲 refactor: Consolidate Claude Prompt Cache and Context Checks (#15008)
* 🧭 fix: Unify Future Claude Capabilities

* 🐛 fix: Cover Claude Capability Edge Cases
2026-08-20 12:18:50 -04:00
James Todaro
e49e264487
fix: Resolve axe Violations in Sidebar, Tools Dropdown and Footer (#14979)
*  fix: Resolve axe Violations in Sidebar, Tools Dropdown and Footer

- give virtualized conversation rows the row/gridcell roles their grid and rowgroup parents require
- make the conversation row a non-interactive container, moving its accessible name, aria-current and focus ring onto the title control so it no longer wraps the options button
- open the tools menu non-modally and portal it into the main landmark, dropping Ariakit's injected dismiss button and keeping menu content inside a landmark
- expose aria-valuenow, aria-valuemin and aria-valuemax on the sidebar resize handle
- drop role="contentinfo" from the chat footer, which is never rendered outside main
- scan the loaded app in a11y.spec.ts, and cover seeded conversation rows, a hovered row and the open tools menu

*  fix: Keep the Resize Handle's ARIA Range Valid at Every Viewport

- floor the announced maximum at the aside's own min-width, so viewports where 40% falls under it no longer report a maximum below the minimum
- track the viewport so the announced range follows a resize instead of a render-time snapshot
2026-08-20 11:52:53 -04:00
Danny Avila
c5276fc63d
⏱️ feat: Run Scheduled Chats Through Durable Agent Triggers (#14939)
* feat: Scheduled Chats — agent-centric scheduled runs creating real conversations

Squash of the full review-hardened branch (PR 14540, supersedes 14373) onto
latest dev, preserving the exact verified tree. History prior to this commit
lived on the pre-squash branch; every invariant below survived 25 Codex review
rounds plus two external audit rounds (R26) with regression tests that fail
without their fixes.

Feature:
- Schedules CRUD + side-panel UI (cadence dialog, run cards, Run Now), roles/
  permissions (SCHEDULES:USE), interface.schedules availability, per-user limits
  and capacity slots, timezone-aware cadence with DST-conservative floors and
  misfire grace.
- Engine: single-process claim/fire loop with leases, loopback POST dispatch
  (signed schedule-fire JWT claims, per-occurrence idempotency key inside the
  route's clientRequestId charset), overlap/balance/capacity/duplicate skip
  policies with auto-disable streaks (too_many_failures, insufficient_balance),
  reconciliation from retained terminal-job evidence, erasure sweep.
- Scheduled runs create real conversations through the resumable agents chat
  path: HITL pauses surface on the card (requires_action), resumes re-apply the
  fire boundary's admission policy (revision fence, enabled, global kill switch,
  SCHEDULES:USE, availability) before continuing a billed generation.

Correctness invariants (the audit surface):
- Settlement discipline: every persistence-producing write happens-before a
  run's terminal outcome write; Stop/complete/pause race through single-winner
  terminal CAS claims (dev's TerminalJobClaim substrate) with retained,
  completedAt-less evidence for scheduled fires plus an owner-intended outcome
  stamp (scheduleOutcome) the reconciler prefers over re-derived success —
  round-tripped through the Redis hash mapper.
- Swallowed generation failures (client error content parts) classify to
  error/skipped_balance instead of success on both initial and resumed paths;
  stale stamps are refreshed evidence-first when persistence plus the Mongo
  outcome write both fail.
- Abort honesty: delivery judged by generation ownership and the CAS's actual
  from-status; republication escalates to the transport's acknowledged variant;
  the Stop route settles only genuinely paused runs.
- Account deletion: one-way barrier (deletionRequestedAt) with auth-cache
  tombstone-before-stamp, boundary rechecks across all auth strategies, quiesce
  of scheduled + interactive work with durable per-stream abort fences
  (positive-evidence acknowledgement only), owner-side finalization markers for
  post-terminal billed writes, deferred-deletion sweep (explicitly ensured
  partial index) that completes cascades autonomously. Remote OpenAI-compatible/
  Responses requests are documented as outside the quiesce and tracked in issue
  14594.
- Store compatibility: finalization markers optional on the legacy IJobStore
  contract with coherent degradation (registration fails -> synchronous title
  fallback; count reads 0).

* fix: fence terminal response persistence from deletion; retain stale-pause evidence

Two blockers from the third external review round.

Terminal persistence visible to account deletion:

- The finalization marker was registered only when post-terminal TITLE work was
  possible, but every persistence-owning terminal CAS opens the same window: the
  claim drops the job out of the active set BEFORE the response save (and the
  background user-message/convo saves), so a deletion quiesce landing there saw
  neither an active job nor a marker and could cascade while the admitted request
  could still recreate messages. Both controllers now register the marker before
  every persistence-owning claim — the fresh-turn path and the HITL resume — and
  release it only after their pending saves (and any post-terminal title) have
  landed, on success and failure paths alike. The TTL bounds a crash.
- settleAbortFence no longer clears a complete/error fence while
  `terminalPersistencePending` is set: terminal at the CAS is not settled while
  the owner is still persisting. The job facade now surfaces the flag.
- The marker trio is REQUIRED by the runtime store contract (assertJobStoreV2
  refuses a store without it at configure time, keeping the failure loud and
  deterministic) while remaining optional on the legacy public IJobStore type for
  source compatibility. The silent degrade path from the previous round is gone —
  it was not deletion-safe.

Stale-pause recovery retains scheduled evidence:

- Three crash/timeout recovery paths — ApprovalLifecycle.failStalePausePersistence
  and the InMemory/Redis stale-pause cleanups — unconditionally stamped
  `completedAt`, putting a scheduled fire's failed-pause error terminal on the
  short completed TTL. A Mongo outage longer than that TTL erased the evidence and
  the reconciler recovered the run as `interrupted` instead of `error`. All three
  now follow the controller-observed path from the previous round: scheduled jobs
  omit `completedAt` (retained-evidence TTL) and stamp the error outcome.

The PR description now explicitly narrows the deletion guarantee for the remote
OpenAI-compatible/Responses paths (tracked in issue 14594).

* fix: deletion-fence marker protocol — generation-scoped, atomic, fail-closed, all terminal paths

One consolidated pass over the finalization-marker mechanism, per the fourth
external review round. The invariant it establishes: NO persistence-owning
terminal CAS runs without a durable, generation-qualified marker covering the
window it opens, and every consumer treats a pending terminal as unsettled.

- Generation-scoped markers. Entries were keyed (userId, streamId), so a
  Stop-superseded generation finishing late could clear the marker its
  replacement registered on the same conversation. Marker fields are now
  qualified by the generation's createdAt; clears must present the same
  identity, and an unqualified legacy clear cannot drop a qualified entry.
- Atomic Redis registration. HSET-then-EXPIRE loses the fresh marker when the
  user's existing hash expires between the two commands (or the process dies
  there); registration is now a single Lua script carrying both.
- Fail closed everywhere. Registration failure (after one retry) now REFUSES
  the terminal CAS instead of proceeding uncovered: the completion claim throws
  into the error path, the error path skips completeJob and leaves the job
  ACTIVE — deletion-visible by itself, recovered by the stale-running reaper —
  and abortJob returns a new retryable `fence_unavailable` failure the Stop
  route answers with 503 and the deletion quiesce treats as an unacknowledged
  stop (fence kept). The previous round's log-and-proceed is gone.
- Every terminal path enrolled. abortJob now owns its window (register before
  the abort CAS, clear in its finally — the Stop route's checkpoint prune and
  partial save run inside beforePublish, between CAS and publication); the
  interactive and background generation-error paths register before their
  completeJob; the resume controller's error finalization registers before its
  completeJob. A lost or thrown claim releases the marker after pending saves
  flush instead of holding the user's deletion behind the TTL.
- Every pending terminal unsettled. settleAbortFence defers on
  terminalPersistencePending for ALL statuses — including `aborted`, whose
  route-side persistence the previous guard missed.

Also: a direct Redis regression for scheduled stale-pause retention (the P2
test gap), and the PR description no longer claims to carry every commit.

* fix: lease-token lifecycle fences — same-generation isolation, admission fence, undelivered-Stop retention, legacy abort enrollment

Fifth external review round; four P1s, handled as the requested consolidated
lifecycle-fence pass.

- Lease tokens. Marker fields were (streamId, createdAt), shared by every
  contender on the same generation — completion and Stop, or two racing Stops —
  so a losing contender's cleanup erased the winner's still-live marker. Every
  registrant now carries a unique lease token in the field and may only ever
  clear its own lease; unqualified legacy clears cannot touch qualified entries.

- Admission fence. Authentication can pass before the deletion barrier goes up,
  and the durable createJob is several async steps later — a deletion quiesce in
  that window saw neither an active job nor a marker and could cascade before
  the admitted request created its job. The controller now registers an
  admission lease and THEN rereads the deletion barrier: the ordering guarantees
  either this request observes the barrier (403, lease released, slot/claim
  cleanup) or the quiesce observes the lease and defers. Held until createJob is
  durable; released on every refusal and initialization-error path. Fail closed
  when the lease itself cannot be registered (503 retryable).

- Undelivered-Stop retention. abortJob released its lease in a finally even when
  delivery AND publication had provably failed — the job reads terminal
  (invisible to active-set scans), a user Stop writes no durable abort fence,
  and the remote owner keeps generating and will persist its abort-catch writes
  whenever the signal finally lands. The lease is now retained in exactly that
  case, and each resignal attempt heartbeats a fresh lease so the fence outlives
  the TTL for as long as delivery is still being driven. The abort-winning
  turn's own loser-side pending saves are additionally fenced in the controller
  catch (best-effort — those writes are already in flight).

- Legacy abort enrollment. abortMiddleware (assistants abort route fallback for
  non-assistants endpoints) awaited abortJob and then spent usage and saved the
  stopped response AFTER the abort's lease was released. Both writes now run
  inside `beforePublish`, between the abort CAS and publication, covered by the
  same lease as every other abort.

Barrier tests, each verified to fail without its fix: same-generation lease
isolation (store), racing two-Stop loser cleanup (manager, stale-read forced
CAS race), undelivered-Stop lease retention, resignal heartbeat, admission
refusal with lease-before-reread ordering plus release-on-durable-create, and
legacy-abort persistence inside beforePublish.

* fix: heartbeat-backed owner-lifecycle leases close the settlement handoff races

Sixth external review round: the remaining P1 interleavings were one structural
problem — lease handoffs that were not atomic — resolved as the requested
consolidated lifecycle-lease pass.

- Quiesce reads leases BEFORE the active-job scan. The admission-lease -> durable
  -job handoff is only atomic against a reader in the OPPOSITE order of the
  writer: writers hold the lease strictly until the job is active-set visible,
  so leases-first shows every interleaving either the lease or the job.
  Jobs-first allowed a request to create its job after the scan and release its
  lease before the count — hiding both, cascading, and letting the new
  generation persist into a deleted account.

- The abort acknowledgement is fenced by the owner-lifecycle lease. Redis ACKed
  the moment the owner's AbortController tripped; the stopping side released its
  lease on that ACK while the owner's asynchronous abort-catch persistence was
  still ahead. The transport now awaits a manager-installed pre-ACK hook that
  registers a DETERMINISTIC owner lease (exactly one owner exists per
  generation, and determinism is what lets the signal-time registrant and the
  owner's catch-side release agree across processes) before the acknowledgement
  is persisted or published; an owned same-replica abort bridges to the same
  lease before tripping its local controller. Both generation-owner catches
  (fresh turn, resume) release it once their writes land.

- A failed replacement handoff no longer orphans the predecessor. The atomic
  replacement removes it from active storage, and an unconfirmed handoff
  terminalizes the replacement too — leaving nothing a quiesce could discover
  while the predecessor's provider may still be generating. Its owner lease is
  now retained at the point the receipt fails delivery; the owner replica renews
  it through the pre-ACK fence when the signal finally lands.

- Leases HEARTBEAT while held. The five-minute store TTL only bounds a crashed
  holder; live persistence — a stalled save, a long deferred title — must never
  outlive its own fence. holdUserFinalization registers and renews every minute
  until released; the controllers' completion/error/admission leases all hold.
  The undelivered-Stop retention moved to a deterministic `stop` lease that
  every resignal attempt renews and the first successful one clears (no more
  opaque leases accumulating to TTL), and a THROWN abort transition releases the
  contender lease instead of leaking it.

- The user-document abort-fence mutations now invalidate the auth user-doc
  cache, matching every other user-doc write.

Barrier tests, each verified to fail without its fix: quiesce lease-scan
ordering (plus the observed-lease defer), pre-ACK fence ordering at the
transport, owned-abort owner-lease bridging, replacement-handoff predecessor
retention, held-lease heartbeat past the TTL, and failed-then-successful
resignal reaping the retained stop lease.

* fix: one manager-owned owner-lease span across every abort delivery path

Seventh external review round; four lifecycle-fence gaps, closed by making the
owner-lifecycle lease a single manager-owned, heartbeat-held span.

- Fail-closed acknowledgements. The pre-ACK hook registered a one-shot lease and
  the transport ACKed even when it failed; a same-replica owned abort likewise
  swallowed registration failure. The hook now acquires a HELD owner lease
  (heartbeat until the owner's catch releases it via releaseOwnerLease) and a
  rejection SUPPRESSES the acknowledgement — the stopping side stays retryable
  behind its retention lease, and every resignal re-drives the handler. The hook
  also stopped gating on `job.createdAt === generationId`: during a replacement
  handoff the store holds the replacement while the abort targets the
  predecessor, and that gate silently skipped exactly the generation being
  acknowledged (the store job is owner identity, never a generation gate). A
  local owned abort acquires the same held lease before tripping its provider;
  post-CAS the trip cannot be withheld, so acquisition failure downgrades
  delivery and the retention handoff keeps the user fenced.

- Committed-but-lost-reply disambiguation. A thrown abort transition released
  the contender lease as if nothing had happened, but a Lua CAS can commit and
  lose its reply — an aborted job invisible to active-set scans whose provider
  was never signalled, with no fence left. The throw path now re-reads the exact
  generation: only a job still live under the caller's identity proves no
  commit; committed or ambiguous outcomes hand the fence to the deterministic
  stop lease (kept on the contender lease if even that fails) before rethrowing.

- Replacement handoff covered end to end. A LOCAL replacement abort acquires the
  predecessor's held owner lease before the trip (failure reports the receipt
  undelivered, engaging retention). Failed-handoff retention is no longer a
  swallowed one-shot: it heartbeats with the durable acknowledgement proof as
  its renewal predicate — acquisition failures keep retrying for as long as the
  fence is needed, and the retainer stands down (without clearing the shared
  field) once the remote owner ACKs and thereby holds its own lease.

- A LOCAL resignal delivery hands off to the owner lease BEFORE clearing the
  retained stop lease, and keeps the stop lease when that handoff fails.

Barrier tests: hook rejection suppressing the ACK, commit-then-lost-reply
retention with its provably-uncommitted counterpart, local-resignal owner
handoff ordering, pre-ACK owner lease held past the store TTL until release
(and provably stopped after), and failed-handoff retention retrying on its
heartbeat — verified fail-before/pass-after by stashing the fixes.

* fix: finish scheduled chat lifecycle hardening

* fix: close scheduled chat review follow-ups

* fix: generation-fence abort recovery evidence

* test: wait for settled approval tool output

* test: preserve scheduled init reconciliation option

* refactor: rebuild scheduled chats on durable agent triggers

* test: reset MCP cache mock between cases

* test: isolate scheduler startup in server specs

* fix: harden scheduled run lifecycle

* fix: normalize schedule capacity conflicts

* test: type schedule collision fixture

* fix: re-fence scheduled resume and expiry

* fix: fence scheduled resume handoffs

* fix: release superseded manual schedule leases

* fix: release failed run-now claims

* fix: release superseded engine claims

* fix: repair schedule dialog interaction and rework its form

The agent picker was unusable: ControlCombobox portals its popover to the
body by default, which lands it outside the dialog's Radix focus trap. Clicks
passed through it, it could not be tabbed into, and the trap fighting Ariakit
for focus locked the page up on selection. The prop is documented for exactly
this case — pass `portal={false}` and give the dialog `overflow-visible`, as
ProjectButton already does. The time and day dropdowns defaulted the same way.

Alongside that:

- Extract the agent builder's instructions editor (special-variable menu plus
  expand-to-fullscreen) into a controlled `VariableEditor` and use it for the
  schedule prompt. Insertions now route through `onChange`, so react-hook-form
  sees them — the schedule PATCH is built from `dirtyFields`, and a `setValue`
  that skipped dirty tracking would have dropped an inserted variable silently.
- Replace the hand-rolled frequency buttons with the shared `Radio`. They marked
  the selection with `bg-surface-hover` on an outline button whose hover is the
  same token, so the selected option was indistinguishable from a hovered one;
  `Radio` is also a real radiogroup rather than four `aria-pressed` toggles.
- Wrap the fields in a real `<form>` and associate the footer button by id, so
  Enter submits. Group the frequency, day and time controls in fieldsets.
- Add placeholders for name and prompt, match the textarea fill to the other
  fields, and label the hourly case as minutes past the hour.
- Widen the dialog to `md:max-w-3xl` and pair name with agent so the form fits
  without scrolling on desktop.
- Move scheduled chats below skills and above prompts in the side nav.

The new dialog spec fails when the portal fix is reverted.

* test: cover scheduled and subagent deletion drains

* refactor: own the form-control appearance in the client primitives

Addresses the codex finding on ScheduleDialog: a feature-local `FIELD_CLASS`
restated the `Input` primitive's border, radius, height and background so it
could be pasted onto the schedule dropdowns, leaving those controls with no
connection to the primitive they were imitating.

Move that appearance into `packages/client/src/components/Field.ts` as the
single source `Input` and `Textarea` now compose, and give `Dropdown` and
`ControlCombobox` a `variant="field"` that applies it. The schedule dialog
passes the variant and carries no class strings of its own.

This also repairs a break the dev merge would otherwise have introduced: the
newer `Dropdown` splits `className` (wrapper) from `triggerClassName`, so the
old pasted classes would have landed on the wrapper and left the triggers
unstyled.

The semantic-token guard now watches the shared module and asserts each
primitive still composes it, which covers more than the two files it read
before.

* fix: keep an explicit schedules disable from becoming an opt-in

`use` is two things at once for a dual-purpose runtime interface field: a
permission bit, which DB overrides strip, and the runtime disable signal that
`getLimits` reads. Stripping it from `{ use: false, maxPerUser: 2 }` leaves an
object, and `getLimits` treats any object without `use: false` as enabled — so
an admin override written to stop scheduled billing for a role or user started
it instead.

Collapse an explicit disable to the boolean form before the strip, on both
paths that accept it: the `interface.schedules` field patch, which admitted the
object wholesale because bare runtime paths deliberately bypass the permission
gate, and the overrides merge, which reached the composite-field branch and
kept `maxPerUser`. Objects that only narrow limits are untouched, so a
principal can still be given a smaller cap.

Both regressions fail without the normalizer.

* fix(schedules): Wave A — null-balance CAS, atomic paused-card clear, clustered erasure sweep

Slice 1 (thread r3804518381): route the existing-null balance initialization
through a { user, tokenCredits: null } compare-and-set instead of a blind $set,
so a concurrent initializer/charge landing between the preflight read and the
write is never handed back its spent starting balance. On a CAS miss the
preflight re-reads the winner. The absent-record $setOnInsert path and the
credited-record refill-config sync are unchanged. Adds the initializeNullBalance
adapter (no upsert) and regression coverage for winner/miss/sync cases.

Slice 2 (thread r3804518388): updateScheduleById now drops a `requires_action`
lastRun projection atomically with the configRevision bump. Any pause present at
edit time was projected under the pre-edit revision and can never be replaced by
its own revision-fenced terminal outcome — a disabling edit would strand the
card on "Needs approval" forever. Implemented as classic-operator CAS branches
(DocumentDB rules out a conditional pipeline $unset), fenced on the card STILL
being the pause so a terminal outcome or newer occurrence that races in is
preserved. Terminal history survives untouched.

Slice 4 (thread r3803826204): expose initializeScheduleErasureSweep from the
schedule runtime facade and start it in every clustered (experimental) worker
after Mongo is up. It re-drives eraseScheduleIfDrained for soft-deleted rows so
a hidden prompt cannot outlive its drain when the delete/erase-on-settle
attempts miss. It arms nothing else and never infers owner death from a
process-local missing job (isTopologySafeToArm gates that).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): Wave B.1 — reversible account-deletion schedule suspension

Thread r3804518383. Account-deletion quiesce marked every schedule `deleting`,
disabled it, and cleared nextRunAt destructively. When a later cascade step (or
the drain itself) failed, the controller cancelled the user-deletion fence —
restoring the user — but the schedules stayed `deleting` and were erased by the
sweep, silently losing all of a live user's scheduled prompts.

Replace the destructive marking with a REVERSIBLE, token-fenced suspension:

- suspendUserSchedulesForDeletion(userId, token) snapshots each schedule's prior
  enabled/nextRunAt under a per-attempt token, then fences firing (disable, clear
  nextRunAt, rotate claimToken). It never sets `deleting`, so a suspended row is
  not erasure-eligible. Snapshotting reads then bulkWrites (a classic update
  cannot copy field values under DocumentDB), fenced per row so an already-
  suspended/soft-deleted/edited row is left alone; idempotent per token.
- restoreUserSchedulesFromDeletion(userId, token) reverses it, re-enabling and
  re-arming only rows still carrying the exact attempt token and not independently
  deleted — so an owner-deleted or newer-attempt-suspended schedule is never
  resurrected.
- deleteUserController generates the attempt token, passes it to quiesce, and on
  any failure that cancels the user-deletion fence restores the suspended rows. A
  successful deletion hard-deletes them (and their snapshots) via the existing
  cascade and never restores.

Adds a `deletionSuspension` embedded field (excluded from the wire schedule),
data-method regression tests (suspend/restore/fence/idempotency/no-resurrect),
and controller tests (restore on drain-false and post-quiesce cascade failure,
no restore on success).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): Wave B.2 — wire the interactive Stop persistence protocol

Thread r3804255932. The durable Stop primitives (requestRunAbort 'stop',
getScheduleRunAbortState, markRunAbortPersisted) already existed but production
only used requestRunAbort(..., 'deletion'). An interactive Stop flipped the job
to `aborted` and then persisted its partial message + checkpoint inside
`beforePublish`, without ever stamping the schedule Stop or acknowledging it —
so reconciliation, the generation owner, or a concurrent schedule/account
deletion could terminalize the run and release its capacity (and erase data)
mid-write.

Wire the request -> persist -> acknowledge -> settle barrier through the
schedule runtime (the route never touches raw Mongo):

- Expose beginScheduledStop / acknowledgeScheduledStopPersistence on the service.
- The abort route stamps the Stop BEFORE signalling abortJob (a serialized
  'in_progress' loser returns 409 STOP_IN_PROGRESS without a second abort),
  acknowledges only after beforePublish persistence succeeds, and on a
  persistence failure leaves the barrier unresolved so the run stays preserved
  (client retries; stale-owner timeout is the bounded recovery). A failed abort
  releases the stamp it placed so a replacement/retry is never blocked through
  its predecessor.
- recordScheduleOutcome (the owner settlement path) now waits, bounded, for the
  Stop acknowledgement before terminalizing; a resolved/non-stop/stale marker
  proceeds immediately. The paused Stop settles only after its own ack.

Adds service-layer barrier tests and route-level ordering/persistence-failure/
in-progress tests; the data-layer serialization is already covered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): Wave C.1 — reconcile durable trigger delivery with the reservation

Threads r3803826192 (manual limiter) and r3804255924 (PII/moderation), plus the
independently-found long-Retry-After race. fireSchedule reserves a `started` run
and a global capacity slot BEFORE the durable trigger delivery reaches the chat
route, where an interactive limiter (manual Run Now), PII, or moderation can
reject it before any generation job exists — dead-lettering the delivery while
the run sat `started` until the 30-minute orphan sweep mislabeled it interrupted.
And a valid delivery deferred by Retry-After (up to 24h) could be orphan-settled
and have its capacity released, then fire anyway.

Translate durable delivery state into the schedule outcome:

- Store the deterministic trigger deliveryKey on the ScheduleRun reservation
  (computed from the envelope BEFORE enqueue, so an ambiguous commit still has it).
- Add a getTriggerDelivery engine dep (wired to the merged trigger service's
  getDelivery) that reads the durable delivery by key.
- Schedule reconciliation, for a jobless `started` run: staging/pending/leased →
  admission is live, never orphan; dead → record `error` from the durable
  lastError and release capacity promptly (no 30-minute wait), through the
  ordinary outcome/auto-disable path; succeeded or no record → the existing
  legacy orphan policy (interrupted only past the cutoff); a delivery lookup
  failure defers rather than orphaning a possibly-live delivery. Limiter/PII/
  moderation middleware writes no schedule state.

Adds reconcile state-mapping tests (dead/pending/leased/staging/succeeded/none/
lookup-failure) and a fire test that the reservation's deliveryKey equals the
enqueued delivery's idempotency key.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(stream): Wave C.2 — durable retry/ack for terminal host lifecycle actions

Thread r3804518375. Approval expiry won the `requires_action → aborted` CAS and
then invoked the host hook best-effort: `runApprovalExpiredHandler` swallowed a
failure, and because later sweeps enumerate only `requires_action` jobs, the now-
aborted job was never offered again. In the clustered entrypoint (no schedule
reconciler) the ScheduleRun stayed `requires_action` and its retained job
persisted indefinitely.

Make the host lifecycle work durable rather than schedule-specific:

- Add a generic `terminalHostActionPending` marker, set ATOMICALLY in the same
  terminal transition (ApprovalLifecycle.expireWithIdentity), only when a host
  adapter is installed.
- Retain and index such jobs: both stores keep them out of terminal reaping and
  expose getTerminalHostActionJobs(); Redis adds a set + extended (24h-bounded)
  TTL, in-memory a bounded 24h retention so a permanently-failing hook cannot leak.
- The manager clears the marker only after the adapter acknowledges success,
  fenced by generation identity (clearTerminalHostAction), so a replacement
  generation can neither clear nor execute its predecessor's action.
- cleanup()/expireStaleApprovals() enumerates unacknowledged terminal host actions
  across restarts and replicas and retries the idempotent hook; the relay only
  re-invokes while the marker is unacknowledged, so a successful ack prevents
  duplicate work. Store-won expiry marks it too, so a loser-replica relay still
  crosses the hook.
- Terminal SSE notification continues regardless of host-hook outcome.

Covers in-memory behavior (retry after failure, restart/other-replica retry, ack
prevents duplicates, identity fence, terminal notification on failure, no marker
accumulation for non-scheduled jobs) and updates the Redis cluster-membership
contract test for the new index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* style(schedules): satisfy import sorting in fire.ts and fire.spec.ts

CI "Static checks" failed on IMPORT_SORT for the two files Wave C.1 added imports
to (the AgentTriggerEnvelope type import and getAgentTriggerIdempotencyKey).
Applied scripts/sort-imports.mts to exactly those files — imports-only reordering,
no behavior change. Other files reported by a repo-wide check are pre-existing on
dev and deliberately left untouched so this PR is not widened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): repair the deletion CLI and order restore before the fence release

Addresses three findings from the fresh Codex review.

P1 — config/delete-user.js called methods.disableUserSchedulesForDeletion, which
Wave B removed in favor of suspendUserSchedulesForDeletion. The file is
@ts-nocheck and its spec mocked the removed name, so neither typecheck nor tests
caught it; the real CLI would throw a TypeError before deleting anything and then
only unwind the fence. The CLI now uses the tokenized protocol: it mints a
suspension token, suspends with it, and restores that exact attempt's rows in its
finally block when the deletion does not commit. Its spec mocks the real methods,
so the breakage can no longer hide.

P2 — both the HTTP controller and the CLI released the user-deletion fence BEFORE
restoring schedules. That fence is what refuses new schedule writes/claims, so the
gap let an owner PATCH edit a still-suspended row and have its enabled/next-run
state overwritten by the older snapshot, and let a second deletion attempt
re-suspend under a new token — making the first restore a no-op and stranding the
disabled snapshot permanently. Restore now runs first, while writes are still
fenced.

Hardening for the same defect class across a crash: suspendUserSchedulesForDeletion
now ADOPTS an existing suspension's snapshot when re-suspending a row abandoned by
an earlier attempt, instead of re-capturing the row's current (already-suspended)
state. Without this, an attempt that died before restoring would have its
successor snapshot "disabled, no next run" and permanently strand the schedule.

Tests: CLI restore-before-fence ordering and no-restore-on-success; the same
ordering assertion on both controller post-quiesce failure paths; a data-method
regression that a second attempt adopts the abandoned snapshot and restores the
original enabled/next-run state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): converge dead deliveries in topology-safe maintenance

Codex finding: the `dead` delivery mapping added in Wave C.1 lives only inside
startScheduleEngine's reconciler, but the clustered entrypoint arms no engine — it
runs erasure-only maintenance. A delivery queued before a restart into clustered
mode and then rejected before generation creation (interactive limiter, PII,
moderation) dead-letters while its ScheduleRun stays `started`, holding a global
capacity slot indefinitely for an ordinary non-deleting schedule.

Add a dead-delivery convergence pass to the erasure sweep, so every topology that
runs schedule maintenance settles it. The pass is POSITIVE-EVIDENCE-ONLY and is
therefore safe where absence-based reconciliation is not: a `dead` delivery is
durable shared state proving no generation owns the reservation. It settles only
when the job is confirmed absent or identity-mismatched (an identity-matched job
still owns the run), defers on an unknown job lookup, on an in-flight abort, and
on an in-flight resume hand-off, ignores legacy reservations with no deliveryKey,
and applies a short grace so an accepted delivery still creating its generation is
never settled mid-handoff. Auto-disable policy is deliberately left to the armed
engine; this path records the failure and frees the slot.

Deliberately does NOT touch api/server/experimental.js — the clustered entrypoint
already starts this sweep, so the convergence arrives through the existing
initializer and the shared-file footprint stays as-is.

Tests: settles a dead delivery as error under an explicitly UNSAFE topology,
leaves live deliveries alone, never settles under an identity-matched running
generation (delivery is not even consulted), defers an in-flight abort, and
ignores a reservation with no deliveryKey.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(stream): refresh host-action retention on each retry attempt

Codex finding: unacknowledged terminal host-action evidence was capped at the
24h pause TTL, so a host dependency (Mongo) unreachable for longer than that let
the Redis key — and its pending marker — expire with no generation-fenced
acknowledgement, stranding the ScheduleRun where no reconciler is armed.

Measure retention from the LAST retry rather than from the terminal transition:
enumerating a pending host action IS the retry attempt, so both stores refresh
its retention as they hand it to the hook (Redis re-EXPIREs the job key; the
in-memory store stamps terminalHostActionRefreshedAt and bounds from it). Evidence
therefore survives as long as some replica is still actively retrying, while a
deployment that stops sweeping entirely still lets it age out — so this does not
reintroduce the unbounded leak the cap existed to prevent.

Test: after a failed hook, a later cleanup pass keeps the marker pending and moves
its retention basis forward.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): defer settlement when the Stop barrier times out

Codex finding: waitForStopPersistence returned after its 5s poll budget even when
the Stop was still fresh and unacknowledged, and recordScheduleOutcome then went
straight on to terminalize the run — releasing its capacity, deletion, and erasure
barriers while beforePublish may still have been writing. Slow checkpoint cleanup
is indistinguishable from a dead route on that signal, so the timeout was being
treated as if the barrier had been satisfied.

The poll budget now means "undecided", not "clear". On timeout with a fresh,
unacknowledged Stop the barrier DEFERS: recordScheduleOutcome returns false
without recording, leaving the run active/preserved. Settlement then happens
either when the route acknowledges, or once the existing stale-owner cutoff
(ABORT_OWNER_PRESUMED_ALIVE_MS) authorizes a later attempt — which the loop
already treats as clear-to-settle. Callers with durable retry (the approval-expiry
host action, reconciliation) re-drive it, so a deferral converges rather than
stranding the run.

Test: a fresh Stop that never acknowledges within the budget reports not-settled
and records no outcome.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): require definite delivery failure, extract its message, converge deferred Stops

Third Codex round. Three findings in code from this closeout, plus one pre-existing
P1 that is a one-line operator-safety fix.

P1 — dead-delivery settlement demanded too little evidence. `dead` does not prove a
request was rejected: the trigger host marks response timeouts and invalid success
responses `certainty: 'ambiguous'`, and the engine dead-letters those once retries
are exhausted. The erasure sweep treated every dead letter as positive evidence, so
an ambiguous one sitting over a generation a peer had accepted could terminalize the
run and release its capacity mid-flight. It now settles only on a DEFINITE rejection,
unless job absence is deployment-authoritative (safe topology), where the
confirmed-absent job is itself the evidence.

P1 — `lastError` is an `AgentTriggerDeliveryFailure` object, not a string. A
duplicated local interface declared it `string` (against CLAUDE.md's no-duplicate-
types rule), so both the sweep and the engine reconciler passed the object into the
String-typed run/schedule `error` fields; Mongoose would reject the cast, the per-row
catch would swallow it, and the run would keep its global capacity slot. The dep type
now reuses the canonical `AgentTriggerDeliveryFailure` and both call sites pass
`.message`. Re-typing immediately surfaced a stale test that had asserted a string.

P1 (pre-existing) — the base-config global stop is honored in `getLimits` via
`isRuntimeDisabled` rather than a literal `=== false`. The stop has two shapes, and
deepMerge turns base `{ use: false }` plus a principal override of `true` into
`{ use: true }`, so the literal check reported the feature enabled and Run Now
dispatched straight through fireSchedule, bypassing the operator's emergency stop.
Now the same predicate the engine gate already uses.

P2 — a Stop whose settlement DEFERRED past the poll budget had no convergence path
where no reconciler is armed. `acknowledgeScheduledStopPersistence` now optionally
re-drives the terminal outcome once the barrier clears; `recordRunOutcome` is
match-guarded and idempotent, so an owner that already settled makes it a no-op. The
abort route passes it for a running generation; a paused job still settles explicitly.

Tests: ambiguous dead letters refused under unsafe topology but settled when absence
is authoritative, definite rejections settled either way, the failure message carried
through, and the abort route's re-drive present for running / absent for paused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): converge terminal runs in clustered workers, retry suspension restore

Fourth Codex round on 74f7af8. Both findings are in code from this closeout.

P2 - a clustered worker never settled a run whose generation finished but whose
outcome write failed. `recordScheduleOutcome` retries three times and then returns
false; the owner honors that by PRESERVING the terminal job as the only surviving
evidence, and the armed engine's reconciler replays exactly that. The clustered
entrypoint arms no engine, and its sweep only covered deleting schedules
(settleAbandonedRuns) and dead deliveries - an identity-matched job was skipped
outright. So an ordinary live schedule's run stayed `started`, held its GLOBAL
capacity slot, and kept a preserved job that carries no `completedAt` and is
therefore invisible to the store's finished-job sweep: both leaked until store
expiry.

The live-schedule pass now also converges from a retained terminal job, mirroring
the reconciler branch it stands in for: honor the owner's stamped outcome over the
generic status (so a balance refusal still walks its streak rather than resetting
it), clear the reserved conversationId when the generation never emitted its
created event, and delete the retained job only AFTER the outcome write is durable.

This stays positive-evidence-only and safe in every topology. Presence, not
absence, is the evidence: an identity-matched job is authoritative wherever it is
observed - a shared store shows the real generation, a process-local store can
only be showing this process's own - which is why it needs no
canInferOwnerDeathFromMissingJob fence, unlike the absence-based paths. The
in-flight abort/resume fences still defer, so an `aborted` job cannot settle a run
whose owner may still be persisting. Both cases now share one pass over one window
rather than two, and `retainedOutcome` moved to types.ts so the sweep reuses the
engine's mapping instead of duplicating it (and stays independent of the engine).

P2 - a failed restore stranded a live account's schedules. Cancelling an account
deletion restores the suspended rows while the deletion fence is still armed and
then releases that fence; nothing re-drives the restore afterwards, so one
transient write failure left the user with silently disabled, next-run-less
schedules. The restore is now retried at the single choke point both the HTTP
controller and the CLI share. Retrying is safe because each attempt re-reads only
the rows STILL carrying the token: a partially-applied unordered write converges
on exactly the stragglers, and a fully-applied one finds nothing.

The fence is still released when every attempt fails, deliberately: retaining it
would refuse the live account's schedule writes AND make beginAgentTriggerUserDeletion
report `in_progress` forever, blocking the retry that is the convergence path (a
later attempt adopts this snapshot, so its cancel restores these exact rows). Both
callers now log the user id and suspension token so the state stays recoverable by
hand if that never happens.

Tests: retained terminal job settled and its evidence released, stamped outcome
preferred over the generic status, stamped failure reason carried, reserved
conversationId cleared for a never-created conversation, identity-mismatched
terminal job ignored, abort-in-flight deferred; restore retried past a transient
write failure and converging on a partially applied restore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): converge unprojected pauses and release replayed bookkeeping jobs

Two of the four findings I flagged as pre-existing to my closeout commits. Both are
in code this PR introduces, so merging would have shipped them; both are the same
capacity/evidence-leak class the last four rounds have been closing. The remaining
two (slotless rows escaping the per-user cap, and the non-rotating `started`
reconciliation bucket) are genuinely latent and deliberately left alone.

An UNPROJECTED PAUSE held a global capacity slot forever. The pause projection is
what moves a run row off `started`; `recordScheduleOutcome` already retries it, but
the request controller discarded the result, so three failed attempts were dropped
silently. The armed engine's reconciler replays that state, but the clustered sweep
did not: a paused job is not terminal, so the retained-job path returned without
settling, and the dead-delivery path never inspects an identity-matched job at all.
The row stayed `started` with no cutoff that would ever clear it.

The call site now surfaces the failure, and the sweep converges it, mirroring the
reconciler's pause branch: project `requires_action` (which frees the slot) but do
NOT release the job's evidence — unlike a terminal job it is still live, awaiting an
approval. The resume hand-off fence still defers, so re-projecting cannot release a
slot a continuation just claimed.

The BOOKKEEPING REPLAY pass leaked its retained job. A run reaches that pass only
because its owner crashed before bookkeeping — which is also before it could release
the job it retained for exactly this recovery. The active-run pass clears its own;
this one never did, and a preserved job is deliberately kept WITHOUT `completedAt`
so the store's finished-job sweep cannot reap it early, so nothing else ever would.
It now clears after `finalizeBookkeeping` succeeds — identity-guarded, a no-op when
no job is retained, and skipped entirely when the replay itself failed, since the
retained job is the only surviving evidence in that case.

Tests: pause projected and the slot freed while the live job's evidence is kept,
paused job deferred during a resume hand-off, retained job released once replayed
bookkeeping is durable, and retained job kept when the replay fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): fence the clustered pause replay against a concurrent resume claim

Codex round 5, on code I added in f735341. The finding is correct and the exposure
is one I introduced.

The pause replay lives in a sweep that runs in EVERY clustered replica, so several
sweepers can observe the same unprojected pause. Its guards were all derived from an
in-memory row SNAPSHOT: hasResumeHandoffInFlight reads the snapshot's
resumeClaimedAt, and recordRunOutcome's pause branch matched any row currently in
`started`/`requires_action` with no fence of its own. The race: sweeper A projects
the pause and frees the slot; the owner's approval then claims a fresh one
(markRunResumeClaimed takes the row to `started` WITH resumeClaimedAt in one write);
sweeper B, still holding the pre-projection snapshot, passes its hand-off check and
replays — `$unset: { capacitySlot, resumeClaimedAt }` under a continuation that is
already running. The run reverts to `requires_action` while its generation proceeds
outside global capacity.

The engine's reconciler makes the same call and has the same snapshot-derived guard,
but v1 arms exactly one engine, so it has no concurrent racer; the sweep is the first
thing to run this transition in parallel. Left the engine alone rather than widening
the change: its `requires_action` re-affirmation is deliberate and single-writer.

Fixed where the race is, in the write itself. `recordRunOutcome` takes an optional
`requireNoResumeClaim`, which adds `resumeClaimedAt: { $exists: false }` to the pause
filter, and the sweep sets it. Because the stamp is written in the SAME update that
moves the row to `started`, its absence is atomic proof no resume owns the row. The
flag is deliberately NOT set by the generation owner: its own re-pause legitimately
clears the stamp as the hand-off's completion signal.

The fence cannot block the recovery it exists to enable: markRunResumeClaimed only
matches `requires_action`, so a genuinely stuck `started` row can carry no resume
claim — it is in fact blocking its own approval until this replay frees it.

Tests: a replay from a stale snapshot leaves a resume-claimed row's status, slot, and
claim stamp intact, while a stuck row with no claim still recovers; and the sweep is
asserted to send the fence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): fence stale resume claims by age, not existence

Codex round 6, on the fence I added in ee00c60. Correct again, and the defect is
the mirror image of the one it fixed.

The fence was an EXISTENCE check (`resumeClaimedAt: { $exists: false }`) while its
caller's guard is a FRESHNESS check (hasResumeHandoffInFlight, bounded by
RESUME_HANDOFF_STALE_MS). They agree while a claim is fresh and disagree once it is
abandoned: a worker that dies after markRunResumeClaimed takes the row to `started`
and stamps resumeClaimedAt — but before the continuation resumes or
releaseRunResumeClaim rolls it back — leaves the stamp set forever. Past the bound
the sweep correctly stops deferring and tries to recover the row, but the write
rejected it purely because the field still existed. The row stayed `started` holding
its global capacity slot, and its approval was unresumable for good, since
markRunResumeClaimed only matches `requires_action`. That is exactly the stuck state
this replay exists to clear, so the fence had reintroduced it for the crashed-resume
case.

`requireNoResumeClaim: boolean` becomes `resumeClaimStaleBefore: Date`, and the
filter matches a row with no claim OR a claim older than that cutoff. The sweep
passes the SAME bound its in-flight check uses, so the two can no longer disagree.
A genuinely racing claim is by construction fresh — it is created after the sweeper's
snapshot — so the race from round 5 stays closed.

Tests: a row whose resume claim was abandoned by a dead worker now recovers (status,
slot and stamp all cleared), alongside the existing two — a fresh claim still repels
a stale-snapshot replay, and an unclaimed stuck row still recovers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-20 11:51:30 -04:00
Marco Beretta
bf3fb17b58
fix: Keep Portaled Dropdown Menus Clickable Inside Modal Dialogs (#15026)
* fix: keep portaled dropdown menus clickable inside modal dialogs

Modal OGDialog layers set pointer-events: none on body and re-enable it
only on their own content. A portaled DropdownPopup menu is a body-level
sibling of those layers, so it inherited pointer-events: none: hover
never reached the items and every click passed through to nothing, which
Ariakit treated as an outside interaction and closed the menu.

This made the Agent File Search and File Context upload menus dead when
SharePoint was enabled, since that flag is what switches them from a
plain in-dialog button to the portaled menu.

Restore pointer-events: auto on the menu element so it stays clickable
regardless of the surrounding modal layers.

Fixes #14487

* fix: sort imports in DropdownPopup spec
2026-08-20 11:20:42 -04:00
Marco Beretta
f1fbaeb6d8
🔗 feat: Open Footer Site Links in a New Tab (#15019)
The footer's LibreChat site link navigated away from the chat in the
same tab. Footer content links (default version link and custom footer
markdown) now open in a new tab with noopener noreferrer.

Privacy policy and terms of service links keep same-tab navigation,
following the WCAG decision in #10997.
2026-08-20 11:17:52 -04:00
Marco Beretta
7569404a7c
🎛️ feat: Make Max Subagents Configurable via librechat.yaml (#15023)
* feat: make max subagents configurable via endpoints.agents.maxSubagents

The per-agent subagent cap was hardcoded at 10 in MAX_SUBAGENTS, leaving
orchestration-heavy deployments no option but patching limits.ts and
rebuilding. Add an optional endpoints.agents.maxSubagents key to
librechat.yaml (default 10, hard ceiling 50) that drives request
validation, model spec presets, and the agents panel UI cap.

* style: fix import order in OrchestrationHub
2026-08-20 11:16:08 -04:00
Danny Avila
986b1218ac
🔓 fix: Unblock Detached Subagent Preparation (#15016)
* fix: retain detached subagent execution lifetime

* test: verify detached timer lifecycle

* fix: settle Meilisearch query middleware

* fix: preserve detached child failure diagnostics

* test: select active preparation watchdog

* test: keep watchdog assertion target-compatible
2026-08-20 11:14:27 -04:00
dymux
49083d0adf
✂️ fix(agents): warn when skill descriptions are truncated in the model catalog (#14878)
formatSkillCatalog caps every catalog entry at 250 characters and truncates
silently, while the import validator allows descriptions up to 1024 chars.
Skill authors therefore get no signal that the trigger phrases at the end
of a long description will never reach the model, and the skill silently
stops firing on the prompts it was written for.

Log a warning per truncated skill so operators can see which descriptions
need tightening, and pass the cap explicitly so the warning and the
catalog stay in sync if @librechat/agents changes its default.

Closes #14657

Co-authored-by: dymux <putramkti@users.noreply.github.com>
2026-08-20 06:04:09 +02:00
Marco Beretta
16e4d14191
refactor: Presets, Skills Motion and Model Selector Polish (#14953)
* refactor: presets, skills motion and model selector polish

Four surfaces that had drifted from the rest of the app, plus the CI
fragility that surfaced while getting them green.

Two were functional bugs rather than styling:

Keyboard focus was invisible in the model selector. The highlight rule
existed and the background was painted, but it used surface-secondary and
the menu sits on bg-presentation, which resolve to the same value in dark
and to within 3/255 in light, so only the thin indicator bar ever showed.
Keyboard focus now uses the same surface a pointer gets.

Importing a malformed preset raised com_ui_upload_invalid, which talks
about image size limits, and FileUpload's JSON.parse had nothing catching
it at that call site. The overflow menu owns the input and reports the
existing preset import error instead.

The rest is polish: preset surfaces use the theme radius roles rather than
raw values; the edit dialog stops nesting a fixed 350px scroll box inside
an already scrolling dialog and pins its title and actions, with the
endpoint picker moved to ControlCombobox and kept out of any clipping
ancestor; Clear all and Import move into a three-dots menu matching the
conversation row; the Skills sections and pinned chats adopt the Collapse
that Projects already used; the rendered/source toggle slides between
states, is extracted rather than duplicated, and gains the accessible name
and RTL mirroring it lacked; the header toggle loses its fill and the
mobile new chat button hides when you are already in a new chat.

The CI changes are unrelated to the UI but blocked it: the MCP and Redis
cache jobs installed Redis with a bare apt-get and lost a race against the
runner's own apt-daily work, failing four times and once hanging for 30
minutes. They now stop that background work and wait for the lock.
DPkg::Lock::Timeout alone does not help, since it covers the dpkg frontend
lock and not the lists lock.

* refactor: move the section label appearance into the Label primitive

The preset dialog reached into the agent panel's private `Advanced/ui` for
its field eyebrow, so an agent-only refactor could change the dialog.

Give the shared `Label` a `section` variant and export the recipe for the
agent id row, which heads its value on a span and must not inherit the
label's block layout. Each variant carries its own size, leading and color:
the recipe output reaches that span unmerged, and a font size declared after
`leading-none` drops it.

* fix: derive the mobile new chat action from the route

The context conversation still holds the previous chat for a render after a
history or link navigation, a lag ChatView already guards against, so the
action could show on /c/new or hide while an existing chat loaded.

* style: sort imports in the touched files

* fix: return focus to the menu item after the clear dialog

The dialog is controlled and has no trigger, so Radix restored focus to
whatever held it when the content mounted, the menu's own focus trap, and a
keyboard user was left on the document. The menu stays open behind the
dialog, so the invoking item is still there to take focus back.

* fix: fall back to the trigger when clearing removes the invoking item

Confirming empties the presets optimistically, so React commits the removed
menu item together with the dialog close and the saved invoker is already
disconnected when focus is handed back.

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-08-19 15:49:47 -04:00
Danny Avila
f431b01d1f
🚚 chore: Bump @librechat/agents to v3.6.8 For Token, Trace, and Stream Fixes (#15012)
* chore: bump `librechat/agents` to v3.6.7

* v3.6.8
2026-08-19 15:48:14 -04:00
Danny Avila
ec5283452e
🪪 fix: Preserve Detached Subagent Owner Context (#15015)
* fix: preserve detached subagent owner context

* style: sort detached subagent imports
2026-08-19 15:40:27 -04:00
Danny Avila
2f50e38217
📎 fix: Never Let a Stalled Attachment Disable the Composer (#15013)
* 🖼️ fix: Keep Composer Send Enabled When an Attachment Stalls

The composer's send button is gated on `hasIncompleteFiles(files)`, so any
attachment that can never reach `progress: 1` reads as "still uploading" and
disables send for the rest of the session — draft text intact, no error, no
way out but removing the chip or reloading. Two paths could park an
attachment there:

- `loadImage` starts the upload from `img.onload` and had no `onerror`, so an
  image the browser refuses to decode (unsupported codec, truncated bytes, a
  revoked object URL) never uploaded at all and stranded the file at
  `progress: 0.2`. Drop the file and surface the error instead.
- Upload completion reconciled against `temp_file_id`, the server's echo of
  the id the request was sent with, while every client-side handle for that
  upload — file map key, delayed-toast timer, recovery callbacks — is keyed by
  the id the client owns. A mismatch applied the completion update to a key
  that does not exist, leaving the attachment at `progress: 0.9`.

Covered by unit regressions in the file-handling suite and a composer-level
spec that drives a real upload through `ChatForm`, plus a render-bound guard
on typing (react-scan measures one ChatForm render per keystroke in a browser;
the guard fails on a multiplier).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cq3tev2rPbc2pWyVJwnh6u

* 🧹 fix: Stop the Draft Restore From Clobbering Live Composer Attachments

`restoreFiles` runs on every `QueryKeys.files` write — an upload landing, an
SSE attachment mid-run — not just on a conversation swap, and it was written
as if the draft were always the whole truth:

- An empty draft cleared the composer outright. On the swap path that is
  redundant (the effect already clears explicitly one line earlier); on the
  cache path an empty draft only means the draft write has not caught up, so
  clearing there discards an attachment the user just added — and with no text
  typed, the send button has nothing left to submit. Restoring now only adds.
- A match replaced the composer's entry with the persisted record, dropping the
  local `File`, the blob preview the chip renders from (`FileRow` falls back to
  refetching `filepath`), and the tool resource the upload was staged under,
  and stamping `attached: true` so removing a chip the composer still owns
  leaves the file orphaned server-side. It now layers the record over the live
  entry and leaves `attached` to files actually adopted from a draft.

Confirmed against a real browser run: the entry is at `progress: 0.9` when this
restore fires, so it — not the upload's own completion — is what was re-enabling
send. react-scan render counts are unchanged (typing 20 keystrokes: 111 renders,
ChatForm=20; attaching an image: 1373, FileRow=6).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cq3tev2rPbc2pWyVJwnh6u

* 🔗 fix: Keep an Attachment's Stored Temporary Id Equal to Its Map Key

Two follow-ups from review on the upload reconciliation.

Completion stored the server's `temp_file_id` echo in the entry's value while
keying the map by the id the request was sent with. `useFileDeletion` deletes
map entries by the value's own `file_id` and `temp_file_id`, so where the two
disagreed — the exact case the reconciliation exists to tolerate — Remove would
delete the file server-side and leave the chip behind, and the draft restore
could not correlate its saved key with the cached record. Store the request id.

A refused image decode also left its `uploadScope.recent` reservation behind:
reservations are released by the render that observes the file in the shared
state, which a decode failing before that render never reaches, and once the
file is deleted no later render can either. The ghost is merged into every
later batch's validation, so re-picking the same file reads as a duplicate and
its size keeps counting against the composer's limits.

Both covered; both new guards fail without their fix. Also sorts the composer
spec's imports, which the static-checks import-order gate flagged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cq3tev2rPbc2pWyVJwnh6u

* 🧷 fix: Normalize an Upload's Temporary Id at the Cache Boundary

The composer keys its file map — and the draft it saves — by the `file_id` the
upload request was sent with; `temp_file_id` is only the server's echo of that
id. The previous commit reconciled the composer's own entry against the request
id but left the record the mutation inserts into `QueryKeys.files` carrying the
raw echo, and `restoreFiles` can only correlate a saved draft id by matching a
cached record's `file_id` or `temp_file_id`. Where the echo disagreed the draft
matched neither, so the attachment was silently dropped on the next conversation
switch or reload — the same class of loss, one layer further out.

Normalize once where the response enters client state, and hand the normalized
record to the mutation's callers, so the cache, the composer entry and the draft
all agree on one id. An agreeing response is passed through untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Cq3tev2rPbc2pWyVJwnh6u

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-19 14:24:11 -04:00
Danny Avila
c8953b8f32
🪂 fix: Land Navigation Auto-Scroll on the Rendered Thread (#15014)
The "auto scroll to latest message" setting stopped taking readers to the
newest message when opening a conversation, most visibly on long threads.

`useMessageScrolling` fired its landing on the conversation id alone. That id
reaches the hook a commit or more before the tree does, so `scrollIntoView` ran
against the OUTGOING conversation's rows: it scrolled that thread to its end,
and — having no dependency on the tree — never ran again once the requested
thread mounted. The reader was left at whatever offset the old thread's bottom
happened to be, which on a long thread is the top.

Key the landing on the conversation that owns the RENDERED rows instead, using
the same `messagesTree[0].conversationId` fallback `MessagesView` already uses
to key the mount window, and land once per conversation so the tree identities
a stream mints cannot haul back a reader who scrolled away.

This is independent of the progressive row mounting: that window only ever
grows upward from the newest row, so the end of the mounted content is already
the end of the thread, and the landing needs no full mount to be correct.
Measured against the real client (react-scan render tallies over a 10-message
to 120-message navigation), render counts are unchanged at ~16k and the thread
still mounts progressively; distance from the bottom on arrival goes 841px to
0. With progressive mounting disabled the same navigation landed 15421px from
the bottom, confirming the anchoring was masking this rather than causing it.

Also moves the `autoScroll` setting from Recoil to Jotai, keeping the same
`autoScroll` localStorage key so a stored preference survives, and matching the
`showThinking`/`smoothStreaming` atoms already served through `ToggleSwitch`.


Claude-Session: https://claude.ai/code/session_01BDQSLdbwvtSqCmQSw7Nz91

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-19 14:23:08 -04:00
Dustin Healy
7c71d6dc1a
🛂 fix: Preserve OpenID Re-Auth Errors Through MCP Tool Error Classification (#15010)
The MCP tool-call catch block classifies errors by message substring, so an OpenIDReauthRequiredError raised during header resolution was rewritten into an MCP OAuth configuration prompt and its class identity discarded. The typed error now passes through ahead of the heuristic, so the actionable re-authentication message reaches the caller intact.
2026-08-19 14:21:39 -04:00
Danny Avila
b91691937e
🙋 fix: Free the Composer When a Question Pause Collapses (#15011)
* 🙋 fix: Free the Composer When a Question Pause Collapses

Collapsing a live `ask_user_question` left the user with nothing to do.
A batch of questions disables the composer, the send button, and the stop
button for as long as the pause is active — and `collapse` deliberately
keeps it active, while hiding the popover that carried the only dismiss.
After the chevron there was no way to type, send, or stop the run short
of reloading the page.

Split the composer's role out of `active`: `composerAnswers` (a single
question, answered IN the composer) and `composerLocked` (a batch,
answered in its own card — and only while the popover is up). Collapsing
a batch now hands the composer back to the thread; the stop button
follows `composerAnswers`, so a paused run stays stoppable.

Both collapsed cards also carry the popover's ×, so dismiss survives the
handover, and `submitText` declines a batch's composer text instead of
claiming it — the old `return true` reported success and dropped
whatever was staged when the pause began.

Contrast, per feedback that the questions were hard to read: the answer
options, the answer textarea, and the digit chips all drew their edge
from `border-light`, which measures 1.20:1 against the panel (WCAG
1.4.11 wants 3:1 for a UI component boundary) — a column of choices read
as flat text. Adds a `choice` Button variant carrying its own fill and a
`border-xheavy` edge (5.49:1 dark / 6.54:1 light), at `font-normal` so
the question above stays the heading, and replaces the single-question
popover's hardcoded `bg-white`/`dark:bg-gray-700` with the semantic
surface role it should have been using.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UPtmUb6VLBhhxXkS3PfV6r

* 🧹 refactor: Render Popover Answers Through the Choice Variant

The popover's option rows re-stated the shared `choice` variant's border,
fill, weight, and hover on a raw `<button>` — the same answer control as
the cards', so a later fix to the variant would have drifted the live
popover away from them.

Renders `Button variant="choice"` instead, keeping only what the popover
actually owns: the full-width row layout and the keyboard highlight.
Locked rows now take the primitive's `disabled:` styling rather than a
local `cursor-not-allowed opacity-60`, matching the cards.
-----
2026-08-19 14:21:20 -04:00
Danny Avila
5e3c680761
🪃 feat: Wake Parent Agents on Child Completion (#14975)
* feat: wake parent agents on child completion

* wip: harden child completion wakeup lifecycle

* fix: close the completion-wakeup static failures

Type the durable-claim store fixture, the continue-envelope test helper,
and the terminal message's task metadata so the wakeup suites compile
against the shapes they actually exercise. Replace `Array.prototype.at`,
which the package target library does not provide.

Capture the prepared child thread in a non-optional local before the
provider callback closes over it, and narrow the trigger envelope itself
on `mode === 'continue'` rather than a separately copied mode, so reading
the continue target is sound. Lift the parent-message fallback out of a
nested ternary into a named resolver.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: cover the active-predecessor admission fence

The Redis job-creation call gained a thirteenth scalar argument, so the
spec helper reconstructed the HSET pairs one slot early and rebuilt an
invalid job hash; three creation tests failed on that alone.

Give the fence itself direct coverage in both store adapters, which it
had none of despite deciding whether an automatic continuation may
replace a live parent turn. Each proves a running and a requires_action
predecessor are refused with the state a controller needs for a finite
409, that an absent or settled predecessor is admitted, and that an
ordinary user turn without the policy still replaces its predecessor.
The Redis case also asserts a refused continuation leaves the parent's
durable job and chunks untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: harden completion wakeup rollout and claims

* fix: close completion wakeup race windows

* test: keep the child store fixture exact

* fix: close final subagent wakeup gaps

* fix: preserve ambiguous completion claims

* fix: release pre-admission wakeup claims

* fix: stabilize subagent completion recovery

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-19 12:18:31 -04:00
Danny Avila
d175741010
🪢 ci: Prevent Playwright Apt Lock Leakage (#14993) 2026-08-19 10:21:12 -04:00
Danny Avila
259f1e0c32
🛰️ feat: Route Live Subagent Controls Across Replicas (#14971)
* feat: route live subagent controls across replicas

* fix: initialize task routing in cluster workers

* fix: harden cross-replica task routing

* fix: expire routed task owners independently

* fix: close cross-replica routing edge cases

* fix: bound owner refresh and close routed cancellation gaps

Refresh owned task registrations in bounded parallel batches so a full
heartbeat pass stays well inside the 30-second directory lease instead of
serializing one Redis EVAL per registration.

Route conversation-deletion cancellation through a dedicated owner-side
scope operation. The owner applies the deletion predicate to its complete
local task set, so a scope holding more children than the model-facing
list cap no longer leaves live executors running after their parent is
removed.

Key a consumed claim's retained response by its operation rather than by
one caller's correlation id, so a later poll recovers a terminal result
whose responses were all lost. Live claim statuses stay uncached so a
poll always observes the task's current state.

Type the model-facing `maxLength` bounds with a narrow local string
schema; the SDK's JsonSchemaType does not declare the keyword, and the
runtime checks continue to enforce the same limits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: retain claimed results apart from control replays

A consumed claim is the only routed response whose loss destroys data, so
it no longer shares one bounded cache with control replays that unrelated
command traffic can evict. Claims are retained under their own budget, and
the requester acknowledges a result it received so the owner releases the
copy immediately instead of holding it for the full replay window.

Resolve the post-delete cancellation pass from durable leases. The deleted
conversations cannot be read back, so re-reading each one only scaled the
cascade while probing the owner directory once per removed id; one lease
read now resolves every live child address instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: never consume a result the owner cannot replay

Retention for consumed claims is bounded, so a burst of undelivered results
could evict an earlier one and lose it for good. The owner now admits a claim
only while it can retain a worst-case result, and refuses the routed claim
otherwise instead of consuming it, leaving the result on the task for a later
poll. Retained claims are never displaced; control replays keep evicting.

Key a control replay by the command itself rather than by one caller's
correlation id. The transport's own retry reuses a single envelope, but a
caller that saw the owner as unavailable reissues the command under a new id,
which steered, queued, or interrupted the child a second time.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: own a claimed result until it is acknowledged

A consumed terminal result is task-owned state, not a cache entry. It now
carries no expiry at all: the owner holds it until a caller acknowledges
receipt, and only then is it released. Retention stays bounded by the
existing admission gate, which refuses a claim the owner could not keep
rather than consuming a result it might drop.

Identify a control by the caller's invocation instead of by its content.
The tool mints one id per invocation and routing carries it, so a routed
retransmission of that invocation replays the owner's result while two
deliberate identical commands arrive under distinct ids and both apply.
Content-derived identity could not tell those apart and would have
answered the second from a stale snapshot.

Wait for the dpkg frontend lock in the best-effort Playwright font step.
Its timeout kills npx while the apt-get it spawned keeps the lock, which
then failed the fatal Redis install and ended the MCP replica jobs before
any test ran (#14983).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: treat acknowledgement as part of delivering a result

Publishing an acknowledgement once and ignoring the outcome meant a result
could be reported as delivered while the owner never learned it could let
go, and since that retention neither expires nor evicts, enough lost
acknowledgements would fill it and refuse every later remote claim.

An acknowledgement is now confirmed: publishing to zero subscribers is not
success, it retries inside the ordinary request window, and a claim whose
acknowledgement cannot be confirmed reports the retryable unavailable path
instead of handing back a result the owner still holds. A later poll
recovers that result and acknowledges it, and releasing is idempotent.
Owner registration also outlives the task while a result is unacknowledged,
so the retained result cannot become unreachable.

Take the control invocation identity from the provider's tool-call id
rather than minting one per execution, so replaying the same tool call
stays idempotent while two distinct calls with identical payloads both
apply.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* style: sort the widened node:crypto import

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: own control invocations and cancellation plans at the task seam

Applies one logical control exactly once for its owning task rather than
in the transport, so a local caller and a routed caller of the same
invocation agree, and reusing an invocation id for different content is
refused instead of silently applied. Invocation identity now comes from
the run, agent, and provider tool-call id hashed to a bounded 32
characters, so a repeated `call_0` never bleeds across tasks and no id can
overrun the routed bound.

Cancellation for conversation deletion is now resolved into a plan while
those rows are still readable, then replayed against the owner directory
after the cascade is deleted. Owner registration is awaited before any
provider work, so a child that cannot be addressed never starts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* style: separate the control invocation map from the next member

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: close subagent deletion, claim, and control invocation gaps

Bulk conversation deletion now runs behind a durable owner admission fence.
Draining alone could not close the race: a child admitted on another replica
after the drain read its leases would start provider work against a parent
about to disappear. The fence is written before any lease is read and each
child revalidates it after its own lease is written, so one of the two always
observes the other. It expires on its own, so a process lost mid-deletion
cannot leave an account unable to run subagents.

A terminal child result is no longer kept alive in the owning replica's
memory until someone acknowledges it. Collection is recorded durably on the
child's own message against the polling invocation, so the poll whose
response was lost recovers its own result while a different invocation is
told the result was already collected. Owner-side retention returns to an
ordinary bounded cache that expires, which is what abandoned polls needed:
they can no longer occupy claim capacity until the process restarts.

The deletion drain now cancels each task under one invocation held for the
whole drain, stops re-sending once the owner answers, and retries only
deliveries it could not confirm. A routed control replay also validates the
command fingerprint, so one invocation id carrying different content reaches
the owner to be refused instead of collecting the earlier command's success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: assert the drain's calls before restoring its spies

Restoring a spy also clears its recorded calls, so the drain assertions
ran against an emptied mock. Formats the durable claim method tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: close the follow-on gaps in the deletion fence and result claim

The admission fence now carries an ownership token, so an overlapping
deletion's fence is never lifted by the one that finishes first, and both
fence writes invalidate the cached auth user document. It also covers the
other bulk-delete path: `DELETE /` with no conversation filter removes every
conversation, so it runs behind the same fence rather than a bare drain.

The durable record now decides who holds a one-shot result. An owner replaying
a retained response could hand the same terminal claim to a second invocation;
that invocation is told the result was already collected, while the one that
consumed it still recovers its own. A task with no durable record to arbitrate
keeps whatever the owner answered.

Drain cancellation treats `not_found` as unconfirmed: a missing registration
while the durable lease is still live means the child may be running, so the
command is retried under its invocation once the owner republishes itself.
Control fingerprints are hashed, so retaining one per invocation costs a fixed
few bytes instead of a bounded message.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: hold every deletion fence and keep live idempotency records

An owner now holds one admission fence per concurrent bulk deletion instead
of one at a time, so admission reopens only when the last deletion finishes
regardless of completion order. Expired fences are pruned as new ones arrive
and the set is bounded, so an abandoned fence cannot accumulate or lock an
account out.

A failed durable claim write is no longer read as an absent record. Handing a
terminal result over without recording its claimant would let another
invocation collect the same one-shot output once the database recovered, so
the collection reports the retryable unavailable path and leaves the result
for a later poll.

Control invocation records now evict tasks the store no longer holds before
live ones, over a bounded scan. Dropping a live task's record would let a
caller retry apply its queue, steer, or interrupt a second time once the
transport replay had also expired.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: keep the deletion fence portable and never drop a live record

The admission fence is written with plain update operators again. DocumentDB
rejects pipeline-form updates, and this runs before any deletion, so the
pipeline form would have failed both bulk-delete endpoints outright on a
supported database target.

An excess deletion is now refused rather than silently displacing the oldest
active fence, which would have reopened admission for a deletion still
running. Expired fences are pruned before the cap is tested, so only genuinely
concurrent deletions count against it.

Control invocation records now sweep every settled task's entry when the
window fills, and a window of entirely live records refuses the new control
before touching the child instead of evicting one. Applying a command with no
room to record it would let the caller's own retry apply it twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: hold the fence, bound recovered results, and expire stale commands

The admission fence is renewed for as long as its deletion runs, so a very
large account or a stalled database cannot let it lapse while conversations
are still being removed. Only the deletion's own fence is renewed, and the
renewal stops with the operation.

Cancellation now covers every conversation the cascade removed, not only the
ones a plan named: a grandchild lives in its own parent's scope, which a plan
naming the deleted root never reaches.

A routed request carries the deadline its caller waits for, and an owner drops
one that arrives past it. A publisher disconnected mid-request queues the
envelope offline and delivers it after the caller was told the owner was
unavailable, which would otherwise steer a child the caller believes untouched.

A result recovered from its durable child message is bounded like a routed one.
The message keeps the child's untruncated output, so recovery could otherwise
return far more than the routed result limit allows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: size the fence window so a renewal can be observed

The renewal test set a 30ms drain timeout but the five-minute grace window
dominates it, so the interval was 100 seconds and no renewal could fire
inside the test's deletion. The grace window is an option now, matching the
store's other timings, and the test sizes the window to 90ms.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix: wire the durable claim method and close the fence follow-ons

The production store never received `claimSubagentTaskResult`, so every
terminal result would have surfaced as unavailable once a task settled. The
host wires that object from JavaScript, where the factory's parameter type
checks nothing, so the factory now refuses a store missing any method it
calls rather than failing at the first claim.

The routing transport takes a dedicated publisher with the offline queue
disabled. The shared client held commands issued during a disconnect and
delivered them after the caller had given up, which the request deadline
narrowed but could not close inside the clock-skew allowance.

Fence renewal invalidates the cached auth document like the fence and release
paths, and a renewal reporting its entry gone re-takes the fence instead of
letting the deletion run on unfenced. The post-delete cancellation retries a
transiently unreachable owner: the conversations are already gone, so it is
the only pass that can still stop a late-admitted child.

A replaced replay entry no longer leaves its bytes counted, which would have
inflated the cache's total until unrelated responses were evicted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: wait on observed lease renewal instead of a fixed delay

The shared-lease renewal test held a 60ms lease and slept 100ms before
asserting an overlapping worker was refused, so a loaded runner that
starved the 10ms heartbeat past the TTL let the lease lapse and the
second worker run. Spy on acquisition and renewal, then wait until a
renewal succeeds past the acquired lease's own deadline — direct
evidence the heartbeat carried it past expiry, with no timing
assumption — and give the lease enough headroom that a stalled timer
no longer decides the outcome.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix(agents): close the routing, fence, and cache gaps found in review

Five separate seams, each with its own failure:

`Cluster.duplicate` reads its first argument as a startup-node list and
its second as the overrides, unlike `Redis.duplicate`, so the publisher's
`enableOfflineQueue: false` was silently dropped under
`USE_REDIS_CLUSTER` and a command issued mid-disconnect could still
reach a child after its caller was told `unavailable`. Route both
through `duplicateIoRedisClient`.

The control window's capacity refusal ran before the store knew whether
it owned the task, so unrelated local load could veto a cancellation
bound for another replica. Establish that the task is local first and
leave a remote one to its owner's window.

`clearInterval` stops only future fence renewals. One already waiting on
the database could resolve after the release, read its own lifted fence
as expiry, and write a replacement that nothing remained to lift —
closing subagent admission for the account until it aged out. Track the
in-flight renewal, refuse overlapping passes, and await it before
releasing.

Every owner bounds its own task list, but the aggregation appended each
batch whole, so the model-facing list grew with the number of replicas
holding the scope. Cap the merged list while still reading every reply
for the stale-registration sweep.

The admission-fence prune commits independently of the fence that
follows it, so a refused or failed push left the cached auth document
describing entries the collection no longer held. Invalidate whichever
way the second write goes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix(agents): cap the merged task list the poll tool actually reads

Each owner bounds its own reply and the remote aggregation bounds their
sum, but `listTasks` merged that bounded remote list with however many
children this replica owns and returned it whole. `check_background_task`
could therefore still receive roughly twice the advertised cap. Bound the
deduplicated, sorted result and export the cap so both seams share one
number.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: admit every task the merged-list cap test starts

The base store admits ten concurrent runs per scope by default, so
starting 150 at once left most refused for capacity and the assertion
never reached the merge it was written to check. Raise the cap for this
store only; admission is a different invariant with its own tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix(agents): let a deletion notice its admission fence lapsing

Renewal failures were logged and swallowed, so a run of rejected writes
let the last confirmed `fencedUntil` pass while the deletion carried on
believing admission was still closed — long enough for another replica
to admit a child against conversations about to be removed. Track the
deadline only a confirmed write advances, and check it after the drain,
before anything is deleted: nothing has been removed at that point, so
the operation fails closed and the caller retries once the fence can be
held. A lapse detected after the rows are gone is logged instead, since
reporting failure there would invite a retry against conversations that
no longer exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: raise both concurrency caps the merged-list test trips

Raising the per-scope limit left the store-wide `maxRunningTotal` at its
default hundred, so fifty of the hundred and fifty starts were still
refused. Verified against the base store directly this time: with only
the per-scope cap raised it admits a hundred, and with both raised it
admits all hundred and fifty.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix(agents): close the fence renewal gap and keep running tasks listed

A renewal that started before its deadline but landed after it was still
credited with extending the fence from its own start time, so a window in
which admission stood open was papered over: a child could take a lease
the drain had already read past and the deletion would proceed without
cancelling it. The deadline now only advances when the write lands while
the previous one still holds; anything later records a lapse the fence
cannot be restored backwards over.

The model-facing cap sorted oldest-first and sliced, which dropped the
newest tasks — including children that had only just started running,
and which the poll tool offers no other way to discover. Bound by status
instead: running children first, then the most recent settled results.
Both caps share one helper, and the routed aggregation now bounds after
its loop so the choice is made across every owner's reply rather than by
whichever answered first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix(agents): finish the cap and the fence at the seams they still missed

The status-aware cap only reached the requester: an owner's own reply
still sliced positionally, so a replica holding more than the cap dropped
its running children before the requester could bound anything. Both
sides now share `boundedTaskList`.

A fence that lapsed during the deletion itself was only logged. The rows
are gone by then, so failing is still wrong, but the child another
replica admitted while the fence was down is not: the fence is retaken
and the drain repeated to cancel it.

A child's lease renewal had the same retroactive hole the admission fence
had — Mongo filters on the `now` captured before the call, so a write
landing after the lease expired still moves the row forward, while an
owner drain reading active leases in that gap saw the thread as free. The
lease now carries its own deadline and a late renewal stops the executor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* test: cover the lease lapse and the post-deletion re-drain

The owner-side cap shipped with a regression test; these two did not.
One drives a lease renewal that succeeds only after the lease it was
extending had expired and asserts the executor stops; the other lets the
fence lapse during the deletion itself and asserts a second drain runs
while the request still reports success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01H1cCMDrTWaRNkmtKjpWELZ

* fix(agents): close live-task lifecycle gaps

* test(redis): exercise cluster node discovery

* fix(test): type cluster discovery seam

* fix(ci): wait for orphaned apt processes

* fix(ci): reserve time for apt drain

* fix(ci): skip optional fonts in MCP jobs

* fix(agents): recover tasks after owner loss

* fix(agents): preserve local task discovery

* fix(agents): initialize fail-fast cluster publisher

* style(agents): sort routing test imports

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-19 02:16:35 -04:00
Ravi Kumar L
5f95631283
🪢 fix: Harden Langfuse Media Upload Targets (#14974) 2026-08-18 22:06:33 -04:00
Danny Avila
e4d6bb71f9
📁 feat: Surface Stateful Workspace Downloads (#14984)
* feat: surface stateful workspace downloads

* fix: sort workspace change imports

* fix: reuse workspace button primitives

* fix: hide collapsed workspace actions
2026-08-18 22:05:13 -04:00
Dustin Healy
a33b128c47
🪪 fix: Preserve Stored Access Token Expiry Over ID Token Exp (#14982)
* 🪪 fix: Preserve Stored Access Token Expiry Over ID Token Exp

extractOpenIDTokenInfo let the ID token exp claim overwrite the token set's stored expires_at. The ID token is minted at login and never refreshed, so once a session outlives the ID token TTL, isOpenIDTokenValid reports the access token as expired even when expires_at is hours in the future, and OpenID placeholder substitution silently stops: MCP headers configured with {{LIBRECHAT_OPENID_ACCESS_TOKEN}} ship the literal placeholder string as the bearer credential and the receiving server rejects every connection with an unparseable JWT until the user fully logs out and back in.

The ID token exp now only fills a missing expiresAt instead of overriding a stored one. Identity claim enrichment from the ID token is unchanged, and the exp fallback for token sets without expires_at is preserved.

* 🪪 fix: Validate ID Token Expiry Before ID Token Placeholder Substitution

The precedence fix made isOpenIDTokenValid track only the access token expiry, so an MCP header using {{LIBRECHAT_OPENID_ID_TOKEN}} could substitute an ID token that had already expired. The ID token exp is now preserved separately as idTokenExpiresAt and checked at the ID token substitution site, so an expired ID token substitutes empty rather than a stale credential while access token substitution is unaffected.

* 🪪 fix: Address OpenID Expiry Review Round

Fix expires_at at the source in the OpenID JWT strategy. The stored value described the
incoming bearer's exp even when access_token came from the session or a cookie, so it could
describe a different credential entirely. A new decodeJwtExpiry helper reads the exp of the
token actually stored, and payload.exp is kept only when the raw bearer is the resolved
access token. Opaque session or cookie tokens now store no expiry rather than a wrong one.

Apply a 30 second clock skew buffer in isOpenIDTokenValid and isIdTokenCurrent via a new
exported OPENID_EXPIRY_BUFFER_SECONDS, mirroring OPENID_REUSE_EXPIRY_BUFFER_SECONDS in
AuthController. Tokens that would expire in transit are treated as already expired.

Make isIdTokenCurrent fail closed when idTokenExpiresAt is absent. exp is REQUIRED in an ID
token, so a missing value means the token is malformed or the claims parse threw. The check
uses == null so an exp of 0 counts as present and therefore expired.

Read the ID token exp with a numeric type check so an exp of 0 records idTokenExpiresAt and
fails closed downstream while a non-numeric exp is ignored, and compare the stored expiry
with != null so a gap filled expiry of 0 reads as expired instead of as no expiry at all.

Raise an actionable re authentication error for the ID token placeholder instead of
substituting an empty string. An empty substitution produced a malformed Authorization
header and a 400 downstream rather than a clean signal that the user must re authenticate.

Raise the same re authentication error from processSingleValue when a user has an OpenID
identity, the stored token set is no longer valid, and the value still contains a credential
bearing OpenID placeholder, so the expired access token case that motivated this PR signals
re auth instead of silently shipping or stripping the placeholder. Only the access token, ID
token, and generic token names raise: identity metadata resolves from the user document and
an expiry hint never needed a token, so those keep their existing literal then strip
behaviour. Unknown placeholder names also stay literal and diagnosable, matching the
existing resolvable placeholder policy.

Add the comments the review asked for on the exp fallback heuristic, the EXPIRES_AT
placeholder semantics, why stale ID token claims stay usable for identity fields, and the
advisory nature of the freshness check.

* 🪪 fix: Honour Opaque Access Tokens And Type The OpenID Re-Auth Error

Drop the ID token exp fallback in extractOpenIDTokenInfo. Storing the access token expiry
honestly means an opaque access token now records no expiry, and the fallback then handed the
ID token exp authority over a credential it does not describe. A deployment issuing opaque
access tokens alongside a short lived ID token saw isOpenIDTokenValid go false and the
credential guard reject a perfectly good access token, which worked before this branch. An
unknown access token expiry is now treated as no expiry, and the ID token exp only ever gates
ID token substitution through idTokenExpiresAt.

Give the re-authentication signal a type. OpenIDReauthRequiredError is raised at both the ID
token placeholder and the credential placeholder guard, ErrorController maps it to a 401
carrying the actionable message, and the class exposes statusCode so the agent generation
path answers 401 instead of a bare 500 for the same condition.

Omit rather than blank a header whose credential placeholder is still unresolved on a final
resolution pass, since an empty bearer credential is malformed under RFC 6750 while an absent
header lets the upstream answer its own challenge. Identity placeholders keep stripping to an
empty string.

Move the resolvable placeholder docblock onto the pattern it describes, resolve an EXPIRES_AT
of 0 as the string 0 for consistency with the neighbouring null checks, and let
AuthController consume the exported OPENID_EXPIRY_BUFFER_SECONDS so the 30 second skew
allowance has a single definition.
2026-08-18 22:02:33 -04:00
Ravi Kumar L
da0491d5db
💻 fix(agents): require Code Interpreter for programmatic MCP tools (#14977)
* fix(agents): require code interpreter for programmatic MCP tools

* test(data-provider): fix tool options fixture type

* fix(agents): address programmatic tool review feedback

* fix(agents): avoid no-op update on version revert
2026-08-18 17:16:58 +02:00
Danny Avila
6daafda86f
🧯 ci: Disarm the Graceful-Shutdown Force-Exit Timer on Test Reset (#14972)
* fix(shutdown): disarm the force-exit timer when test state is reset

`shutdown()` arms a 60s timer that calls `process.exit(1)` as a safety net for
drains that never finish. It is cleared only when the drain runs to completion,
so a drain that never settles — an HTTP server whose close callback never
fires, a task that hangs — leaves it armed.

`__resetShutdownStateForTests()` clears the task list, the shutting-down flag
and the server reference, but not that timer. A suite that triggers a signal
therefore leaves a live self-destruct behind: `unref` keeps it from holding the
process open, but it still fires if anything else keeps the process alive to
the timeout, and `process.exit(1)` then takes down whatever is running a minute
later. Jest reports that as a bare `process.exit called with "1"` with no
failing test, because the run dies before it can print a summary.

Track the timer at module scope, clear it from the reset helper, and clear it
from a `finally` so a throwing drain step cannot leak it either.

The new test fails without the reset change: it starts a drain that never
settles, resets state, advances 120s, and asserts the process was not exited.

* Scope the force-exit timer to the shutdown that armed it

Hoisting the timer to module scope introduced an aliasing hazard: a drain that
settles late runs its `finally` against whatever `forceExitTimer` points at by
then. If state was reset and a second shutdown armed its own timer in the
meantime, the late `finally` cleared the second shutdown's safety net instead
of its own.

Keep a local handle per shutdown, always clear that, and null the module
reference only while it still identifies the same timer.

The added test fails without this: it starts a drain whose close callback is
withheld, resets state, starts a second shutdown, then releases the first
callback and asserts the second net still force-exits.
2026-08-18 08:57:37 -04:00
Heinz-Alexander Fuetterer
0ab3414c01
🪙 chore: GPT-5.6 Terra and Luna Token Cost Rates (#14964)
* fix: update openai gpt-5.6 family token costs

* test: cover GPT-5.6 pricing and float ratios

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-08-18 08:12:32 -04:00
Danny Avila
a0fda2cab1
🗳️ ci: Vote on Fail-Open Merges Too (Tiers Are Bridge-Derived) (#14969) 2026-08-18 08:10:09 -04:00
Ravi Kumar L
006e421cd2
💡 feat: add DB-backed admin insights (#14898)
* feat: add Mongo-backed admin insights

* feat: gate insights with environment variable

* fix: tighten insights access and activity metrics

* fix: preserve insights date selections

* perf: parallelize insights search aggregation

* test: wait for MCP conflict recovery

* test: satisfy strict MCP recovery typing

* fix: disable insights pagination while loading

* fix: localize insights range shortcuts

* fix: bound insights search input
2026-08-18 07:51:51 -04:00
Jackson Riding
389cfebea1
🔒 fix: Hide Password Input in Reset CLI (#14923)
Prevent reset-password prompts from echoing sensitive input and add regression coverage for terminal output.
2026-08-18 07:50:47 -04:00
Danny Avila
547bd8c4bf
🧵 feat: Persist View-Only Subagent Threads (#14957) 2026-08-18 07:41:42 -04:00
Danny Avila
20e2f78490
🩹 ci: Quote Colon in Codegraph Votes Workflow (Invalid YAML) (#14952) 2026-08-17 22:42:59 -04:00
Marco Beretta
7480e93181
🌐 fix: Scrollable Language Dropdown in Shared Chat Settings (#14954)
* fix: make the shared chat language dropdown scrollable and use available height

The language dropdown in the shared chat settings dialog could not be
scrolled with the wheel and was capped at 256px, so most of the language
list was unreachable.

Radix wraps the dialog overlay in RemoveScroll with its shards limited to
DialogContent, so wheel events over a popover portaled to document.body
were cancelled. That same portal placement also left the popover inside
Radix's aria-hidden treatment, hiding the whole option list from assistive
technology. Render the popover inside the dialog and let that dialog's
content overflow so the popover is not clipped by it.

Drop the hardcoded max-height so the popover uses the available height
reported by the positioner. This also restores flipping, because the
positioner can now see that the natural height overflows and place the
popover above the trigger when there is more room there.

Remove declarations that never took effect: max-h-[80vh] and
overflow-y-auto on the popover, both shadowed by .popover-ui later in the
same stylesheet, and the --anchor-max-height and --anchor-max-width custom
properties, which nothing reads.

Move the theme and language selectors into their own directory so the
public share page no longer imports through the Nav settings tabs.

* chore: drop the redundant nested winston entry from the lockfile

packages/data-schemas declares winston as a peer dependency of ^3.17.0,
which the root winston 3.19.0 already satisfies, so npm deduped the
nested 3.17.0 copy.

* refactor: give Dropdown separate wrapper, trigger and popover class props

className was spread onto three elements at once: the positioning wrapper,
the trigger button and the popover. A caller styling the trigger silently
restyled the popover as well, and because className was merged after
sizeClasses it also beat the popover's own sizing. LangfuseConnection asked
for a popover the width of its anchor and got a full width one instead.

className now applies to the wrapper only, triggerClassName styles the
trigger and sizeClasses continues to style the popover. Call sites that
relied on the old spread pass the class to the part that needs it, so the
rendered result is unchanged apart from the LangfuseConnection width.

Also add portalElement so a caller can render the popover into a specific
container rather than document.body.

* fix: align the packaged popover radius with the app stylesheet

.popover-ui is declared both in the component's own stylesheet and in the
app's, and the two had drifted: the packaged copy used a 1rem radius while
the app used 0.7rem. The app copy wins inside LibreChat, so consumers of
@librechat/client saw a different corner radius from the app itself.

* fix: keep the shared chat settings dialog scrollable

The dialog content was made overflow visible so the language popover would
not be clipped, which meant the dialog itself could no longer scroll. If it
ever grew past the viewport its content would have been unreachable.

Move the scroll onto an inner region and portal the popover into the dialog
content, outside that region. The popover still sits inside DialogContent,
so it stays within the scroll lock shard and out of the aria-hidden subtree,
while the rows above it can scroll on their own.

* style: format the locales README

Applies the repository Prettier style, which the file did not satisfy.
Formatting only, no content changes.

* chore: remove the unused DropdownNoState component

The file defined a HeadlessUI based dropdown that nothing imported. It was
absent from the package barrels and from the generated type declarations, so
it was never part of the published API and no consumer can be relying on it.

It carried the same defect the Ariakit Dropdown just had, spreading className
onto the wrapper, the trigger and the popover, so deleting it is preferable
to fixing code that never runs.

* fix: declare the dependencies packages/client imports

InputNumber imports the ValueType type from @rc-component/mini-decimal and
the generated declarations re-export that import, but the package never
declared it. It resolved only because npm hoists it as a transitive
dependency of rc-input-number, so a consumer on a strict or nested layout
would fail to resolve the type. Declare it as a peer alongside the other
externals, using the same range rc-input-number asks for.

The theme test requires tailwindcss directly, so add it to devDependencies
rather than relying on hoisting there too.

Also mark the ValueType import as a type import, matching the convention
used elsewhere.

* style: group the ValueType import with the package imports

Type-only imports belong before local imports, as in Avatar.tsx.
2026-08-17 22:42:28 -04:00
Danny Avila
a47ba7168f
🪢 feat: Custom Request Headers For Langfuse (#14945)
*  feat: Custom Request Headers For Langfuse

Self-hosted Langfuse behind an authenticating proxy or gateway could not
be reached: every outbound Langfuse request hardcoded `Authorization` and
nothing else. Adds `langfuse.headers`, mirroring `endpoints.custom`
headers, and applies it to all four request surfaces — trace/media export
(via the agents run config), feedback scores, central project-identity
lookup, and admin credential verification.

Values resolve through the same pipeline as endpoint headers, so
`${ENV_VAR}` interpolation and header-safe encoding come along.
`extractEnvVariable` continues to refuse infrastructure secrets, so a
config cannot exfiltrate `MONGO_URI` through a header. A header whose
variable is unset is dropped with a one-time warning rather than sent as
a literal `${...}`, which a gateway would read as a wrong credential
instead of a missing one.

Headers merge beneath LibreChat's own `Authorization` on the REST
surfaces, matching `mergeHeaders`, so a custom header can never displace
the Langfuse credential.

These are deployment-level and documented as such: trace export batches
spans from every user through a single exporter, so unlike endpoint
headers they cannot carry per-user placeholders.

The central project-id cache key now includes the headers, so the
header-less module warm-up cannot record a proxy rejection against the
entry the request path later reads.

* 🔒 fix: Keep Langfuse Headers Out Of Stored Overrides

The generic admin config API accepts any field path inside an allowed
section, so `langfuse.headers` could be written through it. Unlike
`langfuse.secretKey`, headers are a map rather than one scalar path, so
the config secret registry cannot encrypt them at rest or mask them on
read — an admin-written map would sit in Mongo in plaintext and come
back in plaintext, widening exposure of what are gateway credentials.

Rejects them on both the dotted-patch and object-upsert routes, the same
way process-backed MCP servers are held to librechat.yaml. This is what
makes "deployment-level" true rather than merely documented.

* 🐛 fix: Wire Config Middleware And Header Collisions For Langfuse

Two codex review findings.

P1 — `api/server/routes/admin/langfuse.js` never mounted
`configMiddleware`, so `req.config` was undefined in production and
credential verification silently ran without the deployment's proxy
headers: exactly the deployments this feature targets could not save a
connection. The handler unit tests injected `config` into their mock
requests, so they stayed green. Mounts the middleware after the access
checks (unauthorized callers still short-circuit first) and adds
route-level tests that assert the handler actually receives a resolved
config — the composition root, not the component.

P2 — spreading custom headers under `Authorization` only replaced an
exact-case collision. A configured `authorization` survived alongside
the managed `Authorization` and fetch appends rather than replaces,
sending both credentials in one combined value. All four request sites
now use `mergeHeaders`, which already merges case-insensitively with the
override winning; tests cover the lower- and upper-case variants.

* 🔒 fix: Mask Langfuse Headers On Read And Harden Value Handling

Three codex round-2 findings.

P1 — the write guard blocked storing `langfuse.headers` in Mongo but did
nothing for the read path: `GET /api/admin/config/base` serves the
resolved AppConfig through `redactConfigSecrets`, which only knows
registered scalar secrets, so a yaml-configured gateway credential was
returned in full to any admin with Langfuse read access. Adds a
secret-map registry that masks values while keeping key names, so an
admin can still see which headers a deployment sets. Masking is safe
precisely because these are yaml-only — a masked read cannot be
round-tripped back over the real values. A malformed non-object value at
that path is dropped rather than serialized.

P2 — `mergeHeaders` indexes one spelling per lowercase name, so a config
holding both `authorization` and `AUTHORIZATION` had only one displaced;
the survivor was then appended by `Headers` into a combined value. Case
variants are now collapsed at resolution, before any consumer sees them.

P2 — `resolveHeaders` encodes only values it substitutes a user field
into, and no user is supplied here, so a literal or interpolated
character above U+00FF reached `Headers` unencoded and threw. Final
values now go through `encodeHeaderValue`; Latin-1 still passes verbatim.

* 🔒 fix: Keep Langfuse Header Credentials Out Of Logs And Validate Names

Three codex round-3 findings, plus a documented boundary for the fourth.

P1 — `loadCustomConfig` logs the parsed config at startup (`printConfig`
defaults true), so a literal gateway credential in `langfuse.headers` was
copied into application logs on every boot, undoing the masking the admin
read path had just gained. The printed copy now goes through
`redactConfigSecretMaps`, reusing the same registry. Scoped to map-valued
secrets so scalar-secret log behavior is unchanged; the live config keeps
its real values.

P2 — a nonempty but invalid field name (` X-Token`, `X Proxy Token`)
passed the emptiness check and then threw in the `Headers` constructor,
which would break export, verification, lookup, and feedback for the whole
deployment rather than that one header. Names are trimmed and validated
against the RFC 7230 token grammar, and dropped with a warning otherwise.

P2 — unresolved `${VAR}` detection tested the *resolved* value, so a
credential legitimately containing `${...}` was mistaken for a failed
substitution and dropped. Detection now inspects the configured text and
checks the referenced variables directly, which also drops references to
denylisted infrastructure secrets instead of forwarding them verbatim.

The fourth (fanout gateway forwards only `Authorization`, so a tenant
Langfuse behind its own proxy is not covered) is a real limitation in a
separate component. Documented on the schema field and in the example
config rather than left implied.

* 🔒 fix: Scope Langfuse Headers To Configured Origins

Three codex round-4 findings.

P1 — one header map was attached to every destination a run resolves to.
Under fanout that means a credential meant for an internal gateway was
also sent to the central destination, typically Langfuse Cloud: an
unrelated third-party origin. Headers are now attached only when the
destination's origin is one the deployment explicitly configured (a
self-hosted base URL, the fanout collector, or a tenant destination set
by env). The built-in `*.cloud.langfuse.com` defaults are excluded
precisely because nobody pointed at them. For trace export this also
means attaching after the export branch settles on a `baseUrl` rather
than before, since which destination wins depends on the branch.

P2 — `encodeHeaderValue` only encodes above U+00FF, so a newline, CR, or
NUL passed through and threw in `Headers`, breaking every request rather
than the one header. Values are trimmed (the common trailing-newline
case) then validated against the legal field-value bytes; an embedded
CRLF is a request-splitting attempt and is dropped, not stripped.

P2 — the write guard matched only the exact `headers` property, so
`{ langfuse: { "headers.X-Token": "..." } }` and root-level dotted
variants slipped through into the Mixed overrides document, where the
nested-map redactor never walks them and a later read returns them in
plaintext. All dotted spellings are now rejected.

* 🔒 fix: Bind Langfuse Headers To One Configured Origin

Four codex round-5 findings.

P1 — the round-4 allowlist still authorized every configured origin, so a
deployment with both a collector and an explicit central host sent the
same credential to both. `langfuse.headers` is one map with no way to say
which endpoint it authenticates to, so it is only unambiguous when the
deployment configures exactly one Langfuse origin. Iterating on which
origins to guess was the wrong axis; headers are now sent only when there
is a single configured origin and the destination is it, with a warning
when several make the intent unresolvable. That covers the self-hosted
case this feature exists for; multi-destination deployments need
per-destination headers the schema cannot yet express.

P1 — `fetch` defaults to following redirects, and Node strips
`Authorization` across origins but keeps arbitrary headers, so a redirect
off an allowed origin would hand the gateway credential to a host that
passed no check. Requests carrying custom headers now refuse redirects;
requests without them keep the default, so nothing changes for existing
deployments.

P2 — `extractEnvVariable`'s whole-string branch is anchored and greedy, so
`${CLIENT_ID}:${CLIENT_SECRET}` parsed as one variable name and the raw
template was sent as the credential. References are expanded here now, so
only literal values reach that path.

P2 — a valid token name is not necessarily usable: `Transfer-Encoding`
makes `fetch` throw and a fixed `Content-Length` misdescribes the body of
every other request sharing the map. Request-framing names are dropped.

* 🐛 fix: Expand Langfuse Header References Exactly Once

Codex round 6 (P2). After expanding `${VAR}` references myself I still
handed the result to `resolveHeaders`, which runs `extractEnvVariable`
over it again — so a credential containing `${PATH}`, or any other name
that happens to be set, was silently rewritten on export, verification,
lookup, and feedback. The round-3 test only used an *unset* embedded
name, which the second pass leaves alone, so it could not catch this.

Resolution no longer round-trips through `resolveHeaders`. The only part
still wanted from it was stripping `{{...}}` user placeholders, which is
now applied directly; expansion, encoding, and validation were already
local. Adds a test whose embedded variable is set, which fails against
the previous pipeline.

* 🐛 fix: Process Langfuse Header Templates Before Substitution

Codex round 7 (P2), the mirror of round 6. Having stopped re-expanding
the resolved credential, the placeholder strip was still running over it:
a token containing `{{LIBRECHAT_USER_ID}}` had that span deleted and
`abc{{...}}ghi` went out as `abcghi`.

Establishes the invariant the last two rounds were circling. Every
template operation — placeholder strip, unresolved-reference check,
expansion — now runs on the operator's configured text, and the
credential is substituted last and never touched again. Gateway
credentials are arbitrary strings, so none of their bytes are syntax.

Also moves the unresolved-reference check after the strip, so it no
longer reports a variable inside a `{{...}}` span that the strip removes.
2026-08-17 20:39:23 -04:00
Danny Avila
7d62be2ad3
🕸️ feat: Run Saved Agent Teams as Subagents (#14944)
* feat: Add graph subagent integration

* style: Sort response usage test imports

* fix: Preserve lazy graph runtime context

* fix: Use isolated graph input helper

* test: Align graph integration fixtures

* fix: Preserve lazy graph runtime capabilities

* fix: Bound lazy graph metadata preload

* fix: Harden lazy graph resolution lifecycle

* fix: Coalesce lazy graph member resolution

* fix: Snapshot initialized graph members only

* fix: Preserve lazy agent runtime context

* fix: Preserve batched lazy context preparation

* fix: Preserve graph member capability bounds

* fix: reconcile graph subagents with execution profiles

* style: align graph subagent types with formatter
2026-08-17 18:02:52 -04:00
Danny Avila
7ee9e4e363
🚏 fix: Preserve Endpoint Routing on HITL Resume Replay (#14948) 2026-08-17 18:02:16 -04:00
Danny Avila
68ea03ffb2
🗳️ ci: Codegraph E2E Vote Runs on Dev Merges (Observe-Only) (#14947) 2026-08-17 16:22:59 -04:00
Danny Avila
fe4615591b
📦 chore: bump @librechat/agents to v3.6.3 (#14941) 2026-08-17 13:39:59 -04:00
Danny Avila
09148efe6a
🩹 ci: Codegraph Probe Summary Rendering (#14942) 2026-08-17 13:39:48 -04:00
Danny Avila
c14b8c54eb
🔭 ci: Codegraph Test-Selection Probe (Observe-Only) (#14936)
* ci: codegraph test-selection probe (observe-only)

Asks the codegraph service which test files and matrix jobs the PR
needs and writes the decision to the job summary. Gates nothing —
every path exits 0, forks without secrets no-op. Companion to the
shadow-mode evaluation: the decision CI would act on, made visible
next to the runs it would have replaced.

* chore: remove GitNexus CI and deployment configs

Superseded by the codegraph service: the index workflow spent ~45min
per invocation building an artifact the PR flow never served, while
the replacement indexes incrementally in ~1.4s per commit server-side.
Removes the four workflows (index, deploy, cleanup-pr, pr-command) and
the .do/gitnexus deployment bundle. No remaining references.

* ci: render playwright spec tiers in the codegraph probe summary
2026-08-17 12:16:32 -04:00
Danny Avila
f9876eaaf0
🪜 style: Step Through Batched Questions One at a Time (#14935)
A batched `ask_user_question` interrupt rendered every question stacked in
one scrolling form, which reads as a wall on mobile and desktop alike. Show
one question per step instead, with clickable progress dots, Back/Next, and
Submit only on the last step.

The batch contract is untouched: one interrupt, one answer map, Submit still
gated on every question having an answer, Skip still declines the whole
batch from any step. Single-question batches render exactly as before.
2026-08-17 12:16:19 -04:00
Danny Avila
df294fa474
🧩 refactor: Resolve Tool-Card State Once (#14934)
* 🧩 refactor: Resolve Tool-Card State Once (AI-1810)

Each tool card derived its state several times over — the visible label
from one expression, the `aria-live` announcement from another, the
icon and shimmer from a third, and since #14906 the follow-scroll from
a fourth. Nothing tied them together; they agreed only because each was
written to agree. Thirteen of the seventeen review findings on #14873
were instances of one derivation being updated and another left behind,
and #14892 added more.

`resolveToolCallPhase` is now the single source: one function encoding
the precedence rules, each of which a specific review finding
established, returning `running | completed | cancelled | failed`.
Everything the card shows reads that value.

`ProgressText` takes `phase` in place of the `error` + `errorSuffix`
pair, which encoded three terminal states in two booleans — `error`
meant cancelled, a present `errorSuffix` meant failed — and made every
consumer reconstruct the distinction. That shape is precisely what let
a duration render beside "failed" (Codex round 1 on #14892).

Two things fell out once the state had one home, both dead code rather
than deletions of behaviour:

- `progress` left `ProgressText` entirely; the phase already carries
  everything it was used to decide.
- The `useProgress` mask went with it. Passing 1 in still matters — it
  stops the 200ms interval — but masking the output no longer does,
  because the phase treats an explicit close as terminal outright. The
  "both halves are load-bearing" subtlety is now one half.

Scope: the nine cards that render the shared `ProgressText`. The three
with bespoke layouts (`WebSearch`, `SubagentCall`, `OpenAIImageGen`)
still resolve their own state and are the natural follow-up — they can
adopt the resolver without adopting the component.

Refactor-only. 4891/4891 client tests pass unchanged, including the
suites that encode the cancelled/failed precedence in both directions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

* 🐛 fix: Infer Cancellation From Reported Progress, Not The Animation

`useProgress` holds below 1 for ~200ms after a call reports completion:
it emits the previous value, then `0.99`, then `1` on a timeout. The
resolver read that animated value for its cancellation inference, so a
successful call whose submission ended inside that window rendered —
and announced — as "Cancelled".

The input is now split. `reportedProgress` is what the stream said and
drives the inference; `displayProgress` is the animated value and drives
`running` vs `completed`, so the label and shimmer still follow the
animation rather than snapping.

This restores `ToolCall` and `RetrievalCall`, whose previous predicates
used `initialProgress` and were immune, and additionally fixes
`useToolCallState`, which inferred from `rawProgress` and therefore
carried the bug already — every card the hook backs was exposed to it
before this PR.

Three tests cover the window: a reported-complete call mid-settle is
`running`, a genuinely unfinished one is still `cancelled`, and the card
settles to `completed` without a cancelled frame in between.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

* 🧹 chore: Drop Unused Phase Predicates; Correct A Stale Comment

`isFailedPhase` and `isRunningPhase` had no callers — every consumer
compares the phase directly, which reads better than a wrapper. An
unused abstraction is the thing this PR argues against, so it should not
ship one.

The comment above the hook's resolver call still described "the raw
progress the legacy heuristic was written against", which stopped being
true when the input split into reported and display progress.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-17 12:05:17 -04:00
Danny Avila
aa35cd42b1
📬 feat: Add Durable Agent Trigger Delivery (#14925)
* feat: wire trusted agent trigger dispatch

* feat: add durable agent trigger delivery

* fix: annotate trigger envelope byte limit

* test: isolate trigger startup in server specs

* fix: fence trigger delivery during account deletion

* test: isolate trigger service in user controller specs

* fix: close trigger deletion admission race

* fix: harden account deletion fences

* fix: close durable trigger review gaps

* fix: require offline stale-fence recovery

* fix: type trigger lane sequence ids

* fix: fence admin user deletion triggers

* fix: make trigger deletion recovery durable

* fix: harden offline user deletion

* fix: serialize trigger lane publication

* style: sort trigger delivery imports

* fix: recover orphaned trigger publications

* fix: preserve trigger recovery ordering

* fix: fence trigger publication during purge

* fix: defer remote trigger deletion fences

* fix: close durable delivery cleanup races

* fix: drain CLI generation owners before deletion
2026-08-17 09:25:08 -04:00
github-actions[bot]
b00a6717e7
🌍 i18n: Update translation.json with latest translations (#14919)
Co-authored-by: danny-avila <110412045+danny-avila@users.noreply.github.com>
2026-08-17 09:21:14 -04:00
Danny Avila
2ab0a6d93f
📡 feat: Add Generic Agent Trigger Dispatch Seam (#14916)
* feat: add generic agent trigger dispatch seam

* refactor: harden trigger dispatch contract

* fix: annotate envelope depth alias

* style: sort trigger envelope imports

* fix: reject unknown trigger dispatch modes

* fix: reject unknown trigger envelope versions

* refactor: validate complete trigger envelopes
2026-08-17 09:03:23 -04:00
Danny Avila
f8f118ef29
🛰️ feat: Execute Generic Agent Trigger Deliveries (#14921)
* feat: add generic agent trigger dispatch seam

* refactor: harden trigger dispatch contract

* fix: annotate envelope depth alias

* style: sort trigger envelope imports

* fix: reject unknown trigger dispatch modes

* fix: reject unknown trigger envelope versions

* refactor: validate complete trigger envelopes

* feat: add agent trigger execution host

* fix: enforce trigger delivery contracts

* fix: harden trigger admission path

* fix: finish trigger cancellation handling

* fix: retry strict steer rollout gaps

* fix: retry paused trigger steers

* fix: parallelize trigger admission setup
2026-08-17 08:45:49 -04:00
Danny Avila
6b9fe97990
🎬 style: Reveal the Chat on Programmatic Drawer Closes (#14930) 2026-08-17 08:45:16 -04:00