Commit graph

127 commits

Author SHA1 Message Date
Danny Avila
8969ee4b18
🎚️ feat: Configure Agent Event Runtime in YAML (#15128) 2026-08-23 02:37:33 -04:00
Danny Avila
67b7b441b2
🛂 feat: Filter Model-Bound Content by Source (#14425)
* feat: introduce optional content protection seam

* feat: enforce source-aware content filters

* feat: complete source-aware content enforcement

* test: activate skill file-text fail-close fixtures

* fix: harden source-aware content filters

* fix: harden model-bound content filtering

* fix: preserve legacy filters and generated files

* fix: inspect shared scalar metadata

* test: align mocks with current dev dependencies

* feat: add persisted content filter safeguards

* feat: complete source-aware content filter enforcement

* fix: move resume content preflight into TypeScript

* fix: close content inspection edge cases

* fix: harden content protection boundaries

* fix: complete content protection safeguards

* test: align persisted memory filter coverage

* fix: reconcile content protection with current dev

* fix: reconcile content protection with latest dev

* fix: close content protection review gaps

* fix: enforce source-aware provider boundaries

* fix: preserve legacy PII preflight semantics

* test: stabilize stored branch preflight fixture

* fix: defer agent writes until protected model admission

* perf: harden source-aware model-bound filtering

* fix: canonicalize provider lineage before validation

* fix: satisfy model-bound callback type checks

* perf: Bound content protection filtering work

* fix: Bound submission array traversal

* fix: Stabilize bounded content snapshots

* fix: Scope model-bound traversal overflows

* fix: Preserve scoped content inspection

* fix: Accumulate aggregate traversal scopes

* fix: centralize content policy boundaries

* test: align deferred tool policy context

* test: align controller policy mocks

* style: normalize content protection imports

* fix: close content policy review gaps

* fix: narrow active skill policy config

* fix: address content protection review boundaries

* fix: retain exact provenance overflow sentinel

* fix: preserve literal and scoped provenance updates

* fix: narrow persisted edit provenance

* fix: isolate exact overflow attribution

* fix: centralize stored prompt protection

* fix: fail closed on incomplete transcript evidence

* fix: align canonical transcript routing

* refactor: centralize content policy preflights

* fix: isolate upload policy error typing

* style: sort policy preflight imports

* refactor: centralize content policy boundaries
2026-08-21 22:43:32 -04:00
Marco Beretta
33e42e6d5d
🎛️ feat: Configurable SearXNG Search Options (#14987)
* feat: configurable SearXNG search options

SearXNG queries were hardcoded to google,bing,duckduckgo with no way to
change the engine list, the result language, or the request timeout. Most
self-hosted instances get served CAPTCHAs by DuckDuckGo, so a third of
every query silently returns nothing and operators have no lever to pull.

Add a searxngSearchOptions block to the webSearch config that accepts
engines (as a comma-separated string or a list), language, timeRange, and
timeout, and thread it through to the search tool. Engines are normalized
to the comma-separated form SearXNG expects, with blank entries dropped so
a stray comma cannot produce an empty engines parameter.

Refs #14117

* fix: normalize SearXNG engines on the runtime config path

The engines transform lived only on the zod schema, but loadCustomConfig
returns the raw YAML object rather than result.data, so nothing downstream
ever saw the transformed value. A YAML list reached the SDK as an array and
threw "options?.engines?.trim is not a function" when the search tool was
built, taking web search down entirely for the exact block the example yaml
documents. An untrimmed string reached SearXNG with spaces still in it.

Extract the normalization into normalizeSearxngEngines and apply it in
loadWebSearchConfig as well as the schema, so both the parsed and the raw
path produce the same comma-separated value. Widen the loader's parameter to
TWebSearchConfigInput, which models engines as the list or string an operator
actually writes, and cover the raw path with tests that call the loader rather
than the schema.

* chore: drop unused RerankerTypes import in web config loader
2026-08-21 11:36:23 -04:00
Marco Beretta
29f6ec6eae
🙈 feat: Config Option to Hide Response Feedback Buttons (#15085)
* feat: add interface option to hide response feedback buttons

Adds `interface.feedback` to librechat.yaml. When set to false, the
thumbs up/thumbs down buttons are removed from the message action row
and the feedback endpoint rejects writes with 403, so deployments that
do not consume the data can stop collecting it. Defaults to true.

* refactor: hoist the feedback gate out of the message row and into typed middleware

Reading startup config inside HoverButtons put a query observer and two Recoil
subscriptions on every message row, and rows never unmount, so the cost grew
with the conversation. Resolve the flag once per chat in useChatHelpers and
carry it on TMessageChatContext; useMessageActions withholds handleFeedback
when it is off, which the action row already treats as "no feedback controls".
The flag now stays false until the config resolves, so a disabled deployment
never flashes controls whose writes are rejected.

Move the server-side policy into requireFeedbackEnabled under packages/api so
the route keeps no policy of its own.

* test: stub the feedback gate in specs that replace the api package

The messages router now imports requireFeedbackEnabled, and express rejects an
undefined handler at require time, so every spec that mocks @librechat/api
wholesale has to carry the export.
2026-08-21 11:16:44 -04:00
Danny Avila
8c14f03432
🗂️ feat: Scope Scheduled Chats to Chat Projects (#15056)
* 🗂️ feat: Scope Scheduled Chats to Chat Projects

Adds an optional chat-project destination to a schedule, plus the operator
config to require one — or to pin every scheduled run to a specific project.

Feature:
- `chatProjectId` on the schedule row, accepted on create/update, projected on
  the wire, and carried into the run's conversation through the durable trigger
  envelope's `run` context.
- `interface.schedules.requireProject` refuses schedules that are not filed
  under a project; `interface.schedules.projectId` pins every run to one
  project and implies the requirement.
- Dialog gains a project picker (required when configured, a read-only row when
  pinned); the card shows the destination and the new disabled reasons.

Invariants:
- ONE resolver (`resolveScheduleProjectId`) decides the destination for the
  write handler, the fire path, and the wire projection alike, and an operator
  pin OUTRANKS the stored id in all three. Tightening the config therefore
  redirects — or stops — existing schedules instead of grandfathering where
  their runs land. A pin implies `requireProject` for the same reason: without
  it, a row created before the pin would keep firing with no project at all.
- Create/fire precheck symmetry, mirroring `resolveAgentFireAccess`: a write
  this handler accepts is one the next fire also accepts. Any edit leaving a
  schedule ENABLED re-validates its EFFECTIVE (possibly stored) project, like
  the existing stored-agent and cadence-floor rechecks. A DISABLING edit skips
  the requirement, or a schedule auto-disabled for `project_required` could
  never be turned off.
- Fire-time enforcement auto-disables rather than filing runs loose, matching
  agent_deleted: new `project_required` (requirement raised after creation) and
  `project_deleted` (gone, or pinned to a project this owner does not have)
  reasons, both refused BEFORE a billed generation is dispatched, and both
  advancing so a schedule can never wedge on the occurrence.
- `computeCreateDigest` appends the field only when present, so a payload
  without a project digests byte-identically to one from before this change —
  a create retried across the upgrade still matches its own row instead of
  reading as key reuse.
- Project reads are scoped to the owner, so ownership and existence are the
  same lookup; a read ERROR propagates instead of failing closed, so a Mongo
  blip retries the fire rather than auto-disabling the schedule.

The trigger idempotency key hashes principal/event/target and never
`envelope.run`, so the added run field cannot destabilize delivery identity.

* 🗜️ fix: Keep the Schedule Dialog Inside Its Height Budget

The project picker landed as a new ROW in the schedule dialog, which broke the
e2e edit spec: `md:overflow-visible` turns off the template's scrolling from
`md` up, so the dialog's content must fit the viewport. The extra row pushed the
footer's Save button below a 720x1280 window, where Playwright reported a
visible, enabled button it could never click — 226 scroll-into-view retries and
a 2-minute timeout, on all three attempts.

Measured against `dev` at 1280x720 (Save button's bottom edge, viewport 720):
  dev              688   (32px slack)
  project row      ~790  (off-screen, CI failure)
  3-column row     704   (16px slack — half the budget spent)
  this commit      686   (34px slack, 2px better than dev)

The identity row is now three columns — name, agent, project — and its caption
moved out of the agent cell to sit full width beneath the row: at a third of the
dialog that sentence wraps an extra line, and the row is the tallest thing
competing for the budget. The caption is grouped with the row rather than left
to the form's own 4-unit rhythm, which spent more height on the gap than the
caption occupies.

The e2e spec now asserts the button is in the viewport before clicking it, so
the next field that overflows this dialog says so in one line instead of a
two-minute timeout on a visible element. FOLLOWUPS.md records what the planned
dialog controls (multi-day weekly, timezone, attachments) need first: give
`ControlCombobox` the `portalElement` prop `Dropdown` already has, portal the
popovers into the dialog content, and let the form scroll again.

* 🧹 fix: Address Codex Review on Scheduled Chat Project Scope

Four P2 findings, all real.

Unreachable clearing path (ScheduleDialog). The picker only held live projects, so
`com_ui_schedule_project_none` was a PLACEHOLDER — nothing selectable. Once a
schedule had a project the owner could never take it away, leaving the server's
`chatProjectId: null` path reachable only by API. The picker now carries a real
"No project" option whenever a project is optional, and omits it when one is
required, where there is nothing valid to select.

Placeholder shown for a real project (ScheduleDialog). A stored or pinned project
outside the first loaded page had no name in the paged map, and the combobox
renders its placeholder for an empty display value — telling the owner a scoped
schedule had no project. That one project is now read by id, with the raw id as a
last resort: a poor label, but an honest one.

Project policy skipped at the resume boundary (service.ts). `claimScheduleResume`
re-applied the schedules gate, the revision fence, the kill switch and
SCHEDULES:USE, but not the project policy this PR added. Approving a paused run
whose project was deleted — or whose owner now sits under a requirement or a pin
it no longer satisfies — billed a continuation the very next scheduled fire would
refuse and auto-disable the schedule for. The effective-project resolution now
runs there too, refused before the lease and the capacity slot so a policy refusal
costs nothing and leaves no state to unwind. NOTE: agent access and balance are
still not rechecked on resume; that gap predates this PR and is left alone.

Per-card project derivation (ScheduleCard). Every card ran the projects hook and
rebuilt the full option array, name map, and one icon element per project, to use
a single name — O(schedules x projects) per render and per project-list refresh.
The hook is split: `useChatProjectNames` (map only, for the panel, which resolves
every card's name once and passes it down) and `useChatProjectPicker` (options and
pagination, for the dialog's one combobox). The panel skips the query entirely
until some schedule actually has a scope.

Tests: four at the resume boundary (verified failing without the gate) and four in
the dialog spec. The picker selection in one existing test now goes through the
search field — the popover's VIRTUALIZED renderer sizes its window from a scroll
height jsdom always reports as 0, so with three options it materialized only two.
Full schedules e2e re-run green against the rebuilt client.

* 🎯 fix: Settle Project-Policy Refusals and Keep Create Retries Idempotent

Second Codex round, four P2s. Three were consequences of the resume gate added in
the previous commit, which was half-built: it admitted where it should not and
stranded the run where it refused.

Project policy moves from `claimScheduleResume` into `isScheduleLive`'s `policy`
branch. Both entry points consult that branch FIRST, and both already route its
refusal through abort-and-settle — so a policy stop now settles the occurrence
instead of answering a bare 409 while the job stays `requires_action`, the card
keeps reading "Needs approval", and every retry repeats the same 409 until expiry.
No change to resume.js: the existing branch does the work.

The rule is deliberately NARROW. It refuses only where no valid destination is
left — the requirement is on with nothing satisfying it, or the schedule's own
project is gone (which also unset it on the conversation). It does NOT refuse
because an operator's pin moved: the paused conversation cannot be rebound
(`chatProjectId` is excluded from the resume context and the continuation reuses
the same conversationId), so refusing would strand a pending approval over a
destination it can never reach, for a pin that governs only where the NEXT run
lands — which the fire path already redirects.

Create retries are idempotent again. Project policy had been applied BEFORE the
`clientRequestId` replay lookup, so a raised requirement, a deleted project, or a
moved pin could answer 400 for a create that already committed — pushing the
client to rotate its key and create a DUPLICATE schedule, the exact failure the
key exists to prevent. Policy now applies only to a genuinely new insert, and the
digest is computed from the CLIENT's payload rather than the resolved destination,
so today's policy can no longer re-digest a genuine retry into a mismatch.

An explicit `chatProjectId: null` under a pin is refused rather than silently
resolved to the pin. The payload contract defines `null` as clearing the scope;
answering 201 while filing under the pin reported success for the opposite of what
was asked. Only an OMITTED field takes the pin silently.

Tests: five on the policy branch (both refusals verified failing without it, plus
guards that a moved pin and a live project still admit) and two on the handlers
(the pinned explicit clear, and a committed create recovered by retry after the
policy tightened — verified answering 400 without the reordering).

* 🧭 fix: Converge the Stored Project on the Destination a Fire Resolved

Third Codex round, four P2s.

The root confusion behind the resume findings: an operator pin outranks the stored
id at fire time, `fireSchedule` sends the pin in the trigger envelope, and the row
keeps its old value. The row therefore LIED about where that occurrence's
conversation went, and every later re-validation — the resume boundary above all —
checked a project the conversation was never filed under. A schedule storing A,
pinned to B, with B later deleted and A still live, was admitted for resume into a
conversation that had just been unscoped.

Fixed at the source rather than at each reader: a fire that resolves a destination
different from the stored one writes it back, claim-token fenced like every other
worker-side write and deliberately WITHOUT a configRevision bump — this is the
server reconciling itself to policy, not an owner edit, and a bump would fence an
in-flight occurrence off its own run. Written only AFTER the destination validates,
so an unusable pin never lands in the row, and best-effort: the envelope already
carries the right destination, so a failed write costs accuracy on a later recheck,
never the run itself. The wire projection already reported the pin, so this also
stops the row and the UI disagreeing.

An explicit `chatProjectId: null` under a pin is now refused on the DISABLING edit
path too. The pin check and the requirement are independent rules, and folding them
together let `{enabled: false, chatProjectId: null}` skip the pin check entirely,
unset the row, and answer with a wire projection still naming the pin. Only the
requirement is waived for a disabling edit.

The dialog no longer requires a project for an edit that leaves a schedule DISABLED.
The server waives the requirement there precisely so a row auto-disabled for
`project_required` can still be renamed or tidied up; requiring it in the form made
that unreachable, and an owner with no projects could not edit the stopped schedule
at all.

FOLLOWUPS.md records the two residual gaps with their exact triggers: the sub-second
deletion race inside the resume claim window (which needs the effective project
persisted per OCCURRENCE plus a distinct policy conflict routed through
abort-and-settle), and the fact that convergence happens only when a schedule fires.

Tests: four on convergence (pin written, no write when unchanged, no write for a
destination that failed validation, fire survives a failed write), two on the
disabling-edit pin rules, two in the dialog. Schedules e2e re-run green.

* 🔑 fix: Keep an Explicit Project Clear Out of an Omitted Field's Digest

`computeCreateDigest` appended `chatProjectId` on `!= null`, so an OMITTED field and
an explicit `null` produced the same digest. Because the replay lookup and
`matchesCreateIntent` deliberately run before project policy, a request could reuse a
pinned create's `clientRequestId` while explicitly sending `chatProjectId: null` and
receive 201 describing the pinned row — success reported for the opposite of what it
asked, and the pinned-clear refusal the normal create path applies never reached.

Now `!== undefined`: an omitted field still digests byte-identically to a payload from
before project scope existed, so a create in flight across the upgrade still matches
its own row, while an explicit clear is a distinct intent and digests differently. A
pre-scope client never sent the field at all, so nothing legacy can carry an explicit
null.

* 📍 fix: Validate a Paused Run Against the Project Its Own Occurrence Used

The schedule-wide convergence from the previous commit was not enough, and the
reason is the single-active run index: it covers `status: 'started'` only, so a
PAUSED run does not block the next occurrence. While run 1 sat paused in project A,
a pin move plus one later fire rewrote the schedule row to B — and the resume
policy then validated B while run 1's conversation was still filed under A. Delete
A and the continuation was admitted into a conversation that had just been unscoped.
That window lasts as long as the pause, not the sub-second race the previous commit
documented.

The reservation now records the destination THIS occurrence used, and
`isScheduleLive` validates that record when given the occurrence's `scheduledFor` —
which `resume.js` already reads two lines above the call. No new conflict type and
no settlement-path surgery: the refusal rides the abort-and-settle branch that check
already has.

An ABSENT record falls back to the schedule-level resolution rather than reading as
unscoped. A pre-scope occurrence, or one whose row is gone, must never be treated as
evidence to stop a run — the fallback keeps legacy paused runs behaving exactly as
they do today.

Schedule-wide convergence stays: it keeps the row honest for the UI and for every
check that has no occurrence in hand.

FOLLOWUPS.md now describes the one remaining gap accurately — a deletion inside the
claim window, which needs a distinct policy conflict routed through abort-and-settle
rather than the bare 409 an `inactive` conflict produces.

Tests: three on occurrence-vs-row precedence (the decisive one verified failing
without the lookup), two on what the reservation records, and the resume controller
spec now pins `scheduledFor` in the policy call. Schedules e2e re-run green.

* 🏷️ fix: Tell a Deliberately Unscoped Occurrence From an Unrecorded One

The occurrence fallback added in the previous commit conflated two different
absences. A post-upgrade run that deliberately went unscoped omitted the field
exactly like a row written before the field existed, so both took the fallback —
and a paused unscoped run was then validated against the schedule's CURRENT
project. Under a requirement or a pin added while it sat paused, that admitted a
billed continuation into a conversation satisfying no present policy.

The reservation now ALWAYS records its decision, `null` for unscoped, and the read
reports `recorded` from key PRESENCE rather than truthiness. Only an unknown record
— a pre-scope row, or no row at all — falls back to the schedule-level resolution;
a recorded null is the genuinely unscoped occurrence it says it is, and is refused
once a project becomes required.

The distinction rests entirely on a stored `null` surviving as a present key while a
never-written field stays absent, so that is asserted against real Mongo rather than
assumed: if it ever stopped holding, unscoped runs would silently start being
validated against the schedule's current project again.

Tests: two against mongodb-memory-server (recorded null vs never-written vs missing
row), plus refusal of a recorded-unscoped occurrence under a new requirement and the
preserved fallback for a pre-scope one. Schedules e2e green.

* 🔁 fix: Validate an Initial Scheduled Start Against Its Own Occurrence

The initial-start policy check in `request.js` called `isScheduleLive(..., { policy:
true })` without `scheduledFor`, so it fell back to the schedule-level resolution
even though the run row already carries the occurrence's recorded scope. An
occurrence reserved unscoped, with a pin introduced while its loopback request sat
queued, was therefore admitted against the new pin — producing a billed unscoped
conversation under a requirement it does not satisfy, from an envelope already built
without a project.

`scheduledFor` was already in scope there. Passing it makes the initial start and the
resume validate the same way: against the destination the occurrence itself recorded.
2026-08-21 11:09:04 -04:00
Danny Avila
c5276fc63d
⏱️ feat: Run Scheduled Chats Through Durable Agent Triggers (#14939)
* feat: Scheduled Chats — agent-centric scheduled runs creating real conversations

Squash of the full review-hardened branch (PR 14540, supersedes 14373) onto
latest dev, preserving the exact verified tree. History prior to this commit
lived on the pre-squash branch; every invariant below survived 25 Codex review
rounds plus two external audit rounds (R26) with regression tests that fail
without their fixes.

Feature:
- Schedules CRUD + side-panel UI (cadence dialog, run cards, Run Now), roles/
  permissions (SCHEDULES:USE), interface.schedules availability, per-user limits
  and capacity slots, timezone-aware cadence with DST-conservative floors and
  misfire grace.
- Engine: single-process claim/fire loop with leases, loopback POST dispatch
  (signed schedule-fire JWT claims, per-occurrence idempotency key inside the
  route's clientRequestId charset), overlap/balance/capacity/duplicate skip
  policies with auto-disable streaks (too_many_failures, insufficient_balance),
  reconciliation from retained terminal-job evidence, erasure sweep.
- Scheduled runs create real conversations through the resumable agents chat
  path: HITL pauses surface on the card (requires_action), resumes re-apply the
  fire boundary's admission policy (revision fence, enabled, global kill switch,
  SCHEDULES:USE, availability) before continuing a billed generation.

Correctness invariants (the audit surface):
- Settlement discipline: every persistence-producing write happens-before a
  run's terminal outcome write; Stop/complete/pause race through single-winner
  terminal CAS claims (dev's TerminalJobClaim substrate) with retained,
  completedAt-less evidence for scheduled fires plus an owner-intended outcome
  stamp (scheduleOutcome) the reconciler prefers over re-derived success —
  round-tripped through the Redis hash mapper.
- Swallowed generation failures (client error content parts) classify to
  error/skipped_balance instead of success on both initial and resumed paths;
  stale stamps are refreshed evidence-first when persistence plus the Mongo
  outcome write both fail.
- Abort honesty: delivery judged by generation ownership and the CAS's actual
  from-status; republication escalates to the transport's acknowledged variant;
  the Stop route settles only genuinely paused runs.
- Account deletion: one-way barrier (deletionRequestedAt) with auth-cache
  tombstone-before-stamp, boundary rechecks across all auth strategies, quiesce
  of scheduled + interactive work with durable per-stream abort fences
  (positive-evidence acknowledgement only), owner-side finalization markers for
  post-terminal billed writes, deferred-deletion sweep (explicitly ensured
  partial index) that completes cascades autonomously. Remote OpenAI-compatible/
  Responses requests are documented as outside the quiesce and tracked in issue
  14594.
- Store compatibility: finalization markers optional on the legacy IJobStore
  contract with coherent degradation (registration fails -> synchronous title
  fallback; count reads 0).

* fix: fence terminal response persistence from deletion; retain stale-pause evidence

Two blockers from the third external review round.

Terminal persistence visible to account deletion:

- The finalization marker was registered only when post-terminal TITLE work was
  possible, but every persistence-owning terminal CAS opens the same window: the
  claim drops the job out of the active set BEFORE the response save (and the
  background user-message/convo saves), so a deletion quiesce landing there saw
  neither an active job nor a marker and could cascade while the admitted request
  could still recreate messages. Both controllers now register the marker before
  every persistence-owning claim — the fresh-turn path and the HITL resume — and
  release it only after their pending saves (and any post-terminal title) have
  landed, on success and failure paths alike. The TTL bounds a crash.
- settleAbortFence no longer clears a complete/error fence while
  `terminalPersistencePending` is set: terminal at the CAS is not settled while
  the owner is still persisting. The job facade now surfaces the flag.
- The marker trio is REQUIRED by the runtime store contract (assertJobStoreV2
  refuses a store without it at configure time, keeping the failure loud and
  deterministic) while remaining optional on the legacy public IJobStore type for
  source compatibility. The silent degrade path from the previous round is gone —
  it was not deletion-safe.

Stale-pause recovery retains scheduled evidence:

- Three crash/timeout recovery paths — ApprovalLifecycle.failStalePausePersistence
  and the InMemory/Redis stale-pause cleanups — unconditionally stamped
  `completedAt`, putting a scheduled fire's failed-pause error terminal on the
  short completed TTL. A Mongo outage longer than that TTL erased the evidence and
  the reconciler recovered the run as `interrupted` instead of `error`. All three
  now follow the controller-observed path from the previous round: scheduled jobs
  omit `completedAt` (retained-evidence TTL) and stamp the error outcome.

The PR description now explicitly narrows the deletion guarantee for the remote
OpenAI-compatible/Responses paths (tracked in issue 14594).

* fix: deletion-fence marker protocol — generation-scoped, atomic, fail-closed, all terminal paths

One consolidated pass over the finalization-marker mechanism, per the fourth
external review round. The invariant it establishes: NO persistence-owning
terminal CAS runs without a durable, generation-qualified marker covering the
window it opens, and every consumer treats a pending terminal as unsettled.

- Generation-scoped markers. Entries were keyed (userId, streamId), so a
  Stop-superseded generation finishing late could clear the marker its
  replacement registered on the same conversation. Marker fields are now
  qualified by the generation's createdAt; clears must present the same
  identity, and an unqualified legacy clear cannot drop a qualified entry.
- Atomic Redis registration. HSET-then-EXPIRE loses the fresh marker when the
  user's existing hash expires between the two commands (or the process dies
  there); registration is now a single Lua script carrying both.
- Fail closed everywhere. Registration failure (after one retry) now REFUSES
  the terminal CAS instead of proceeding uncovered: the completion claim throws
  into the error path, the error path skips completeJob and leaves the job
  ACTIVE — deletion-visible by itself, recovered by the stale-running reaper —
  and abortJob returns a new retryable `fence_unavailable` failure the Stop
  route answers with 503 and the deletion quiesce treats as an unacknowledged
  stop (fence kept). The previous round's log-and-proceed is gone.
- Every terminal path enrolled. abortJob now owns its window (register before
  the abort CAS, clear in its finally — the Stop route's checkpoint prune and
  partial save run inside beforePublish, between CAS and publication); the
  interactive and background generation-error paths register before their
  completeJob; the resume controller's error finalization registers before its
  completeJob. A lost or thrown claim releases the marker after pending saves
  flush instead of holding the user's deletion behind the TTL.
- Every pending terminal unsettled. settleAbortFence defers on
  terminalPersistencePending for ALL statuses — including `aborted`, whose
  route-side persistence the previous guard missed.

Also: a direct Redis regression for scheduled stale-pause retention (the P2
test gap), and the PR description no longer claims to carry every commit.

* fix: lease-token lifecycle fences — same-generation isolation, admission fence, undelivered-Stop retention, legacy abort enrollment

Fifth external review round; four P1s, handled as the requested consolidated
lifecycle-fence pass.

- Lease tokens. Marker fields were (streamId, createdAt), shared by every
  contender on the same generation — completion and Stop, or two racing Stops —
  so a losing contender's cleanup erased the winner's still-live marker. Every
  registrant now carries a unique lease token in the field and may only ever
  clear its own lease; unqualified legacy clears cannot touch qualified entries.

- Admission fence. Authentication can pass before the deletion barrier goes up,
  and the durable createJob is several async steps later — a deletion quiesce in
  that window saw neither an active job nor a marker and could cascade before
  the admitted request created its job. The controller now registers an
  admission lease and THEN rereads the deletion barrier: the ordering guarantees
  either this request observes the barrier (403, lease released, slot/claim
  cleanup) or the quiesce observes the lease and defers. Held until createJob is
  durable; released on every refusal and initialization-error path. Fail closed
  when the lease itself cannot be registered (503 retryable).

- Undelivered-Stop retention. abortJob released its lease in a finally even when
  delivery AND publication had provably failed — the job reads terminal
  (invisible to active-set scans), a user Stop writes no durable abort fence,
  and the remote owner keeps generating and will persist its abort-catch writes
  whenever the signal finally lands. The lease is now retained in exactly that
  case, and each resignal attempt heartbeats a fresh lease so the fence outlives
  the TTL for as long as delivery is still being driven. The abort-winning
  turn's own loser-side pending saves are additionally fenced in the controller
  catch (best-effort — those writes are already in flight).

- Legacy abort enrollment. abortMiddleware (assistants abort route fallback for
  non-assistants endpoints) awaited abortJob and then spent usage and saved the
  stopped response AFTER the abort's lease was released. Both writes now run
  inside `beforePublish`, between the abort CAS and publication, covered by the
  same lease as every other abort.

Barrier tests, each verified to fail without its fix: same-generation lease
isolation (store), racing two-Stop loser cleanup (manager, stale-read forced
CAS race), undelivered-Stop lease retention, resignal heartbeat, admission
refusal with lease-before-reread ordering plus release-on-durable-create, and
legacy-abort persistence inside beforePublish.

* fix: heartbeat-backed owner-lifecycle leases close the settlement handoff races

Sixth external review round: the remaining P1 interleavings were one structural
problem — lease handoffs that were not atomic — resolved as the requested
consolidated lifecycle-lease pass.

- Quiesce reads leases BEFORE the active-job scan. The admission-lease -> durable
  -job handoff is only atomic against a reader in the OPPOSITE order of the
  writer: writers hold the lease strictly until the job is active-set visible,
  so leases-first shows every interleaving either the lease or the job.
  Jobs-first allowed a request to create its job after the scan and release its
  lease before the count — hiding both, cascading, and letting the new
  generation persist into a deleted account.

- The abort acknowledgement is fenced by the owner-lifecycle lease. Redis ACKed
  the moment the owner's AbortController tripped; the stopping side released its
  lease on that ACK while the owner's asynchronous abort-catch persistence was
  still ahead. The transport now awaits a manager-installed pre-ACK hook that
  registers a DETERMINISTIC owner lease (exactly one owner exists per
  generation, and determinism is what lets the signal-time registrant and the
  owner's catch-side release agree across processes) before the acknowledgement
  is persisted or published; an owned same-replica abort bridges to the same
  lease before tripping its local controller. Both generation-owner catches
  (fresh turn, resume) release it once their writes land.

- A failed replacement handoff no longer orphans the predecessor. The atomic
  replacement removes it from active storage, and an unconfirmed handoff
  terminalizes the replacement too — leaving nothing a quiesce could discover
  while the predecessor's provider may still be generating. Its owner lease is
  now retained at the point the receipt fails delivery; the owner replica renews
  it through the pre-ACK fence when the signal finally lands.

- Leases HEARTBEAT while held. The five-minute store TTL only bounds a crashed
  holder; live persistence — a stalled save, a long deferred title — must never
  outlive its own fence. holdUserFinalization registers and renews every minute
  until released; the controllers' completion/error/admission leases all hold.
  The undelivered-Stop retention moved to a deterministic `stop` lease that
  every resignal attempt renews and the first successful one clears (no more
  opaque leases accumulating to TTL), and a THROWN abort transition releases the
  contender lease instead of leaking it.

- The user-document abort-fence mutations now invalidate the auth user-doc
  cache, matching every other user-doc write.

Barrier tests, each verified to fail without its fix: quiesce lease-scan
ordering (plus the observed-lease defer), pre-ACK fence ordering at the
transport, owned-abort owner-lease bridging, replacement-handoff predecessor
retention, held-lease heartbeat past the TTL, and failed-then-successful
resignal reaping the retained stop lease.

* fix: one manager-owned owner-lease span across every abort delivery path

Seventh external review round; four lifecycle-fence gaps, closed by making the
owner-lifecycle lease a single manager-owned, heartbeat-held span.

- Fail-closed acknowledgements. The pre-ACK hook registered a one-shot lease and
  the transport ACKed even when it failed; a same-replica owned abort likewise
  swallowed registration failure. The hook now acquires a HELD owner lease
  (heartbeat until the owner's catch releases it via releaseOwnerLease) and a
  rejection SUPPRESSES the acknowledgement — the stopping side stays retryable
  behind its retention lease, and every resignal re-drives the handler. The hook
  also stopped gating on `job.createdAt === generationId`: during a replacement
  handoff the store holds the replacement while the abort targets the
  predecessor, and that gate silently skipped exactly the generation being
  acknowledged (the store job is owner identity, never a generation gate). A
  local owned abort acquires the same held lease before tripping its provider;
  post-CAS the trip cannot be withheld, so acquisition failure downgrades
  delivery and the retention handoff keeps the user fenced.

- Committed-but-lost-reply disambiguation. A thrown abort transition released
  the contender lease as if nothing had happened, but a Lua CAS can commit and
  lose its reply — an aborted job invisible to active-set scans whose provider
  was never signalled, with no fence left. The throw path now re-reads the exact
  generation: only a job still live under the caller's identity proves no
  commit; committed or ambiguous outcomes hand the fence to the deterministic
  stop lease (kept on the contender lease if even that fails) before rethrowing.

- Replacement handoff covered end to end. A LOCAL replacement abort acquires the
  predecessor's held owner lease before the trip (failure reports the receipt
  undelivered, engaging retention). Failed-handoff retention is no longer a
  swallowed one-shot: it heartbeats with the durable acknowledgement proof as
  its renewal predicate — acquisition failures keep retrying for as long as the
  fence is needed, and the retainer stands down (without clearing the shared
  field) once the remote owner ACKs and thereby holds its own lease.

- A LOCAL resignal delivery hands off to the owner lease BEFORE clearing the
  retained stop lease, and keeps the stop lease when that handoff fails.

Barrier tests: hook rejection suppressing the ACK, commit-then-lost-reply
retention with its provably-uncommitted counterpart, local-resignal owner
handoff ordering, pre-ACK owner lease held past the store TTL until release
(and provably stopped after), and failed-handoff retention retrying on its
heartbeat — verified fail-before/pass-after by stashing the fixes.

* fix: finish scheduled chat lifecycle hardening

* fix: close scheduled chat review follow-ups

* fix: generation-fence abort recovery evidence

* test: wait for settled approval tool output

* test: preserve scheduled init reconciliation option

* refactor: rebuild scheduled chats on durable agent triggers

* test: reset MCP cache mock between cases

* test: isolate scheduler startup in server specs

* fix: harden scheduled run lifecycle

* fix: normalize schedule capacity conflicts

* test: type schedule collision fixture

* fix: re-fence scheduled resume and expiry

* fix: fence scheduled resume handoffs

* fix: release superseded manual schedule leases

* fix: release failed run-now claims

* fix: release superseded engine claims

* fix: repair schedule dialog interaction and rework its form

The agent picker was unusable: ControlCombobox portals its popover to the
body by default, which lands it outside the dialog's Radix focus trap. Clicks
passed through it, it could not be tabbed into, and the trap fighting Ariakit
for focus locked the page up on selection. The prop is documented for exactly
this case — pass `portal={false}` and give the dialog `overflow-visible`, as
ProjectButton already does. The time and day dropdowns defaulted the same way.

Alongside that:

- Extract the agent builder's instructions editor (special-variable menu plus
  expand-to-fullscreen) into a controlled `VariableEditor` and use it for the
  schedule prompt. Insertions now route through `onChange`, so react-hook-form
  sees them — the schedule PATCH is built from `dirtyFields`, and a `setValue`
  that skipped dirty tracking would have dropped an inserted variable silently.
- Replace the hand-rolled frequency buttons with the shared `Radio`. They marked
  the selection with `bg-surface-hover` on an outline button whose hover is the
  same token, so the selected option was indistinguishable from a hovered one;
  `Radio` is also a real radiogroup rather than four `aria-pressed` toggles.
- Wrap the fields in a real `<form>` and associate the footer button by id, so
  Enter submits. Group the frequency, day and time controls in fieldsets.
- Add placeholders for name and prompt, match the textarea fill to the other
  fields, and label the hourly case as minutes past the hour.
- Widen the dialog to `md:max-w-3xl` and pair name with agent so the form fits
  without scrolling on desktop.
- Move scheduled chats below skills and above prompts in the side nav.

The new dialog spec fails when the portal fix is reverted.

* test: cover scheduled and subagent deletion drains

* refactor: own the form-control appearance in the client primitives

Addresses the codex finding on ScheduleDialog: a feature-local `FIELD_CLASS`
restated the `Input` primitive's border, radius, height and background so it
could be pasted onto the schedule dropdowns, leaving those controls with no
connection to the primitive they were imitating.

Move that appearance into `packages/client/src/components/Field.ts` as the
single source `Input` and `Textarea` now compose, and give `Dropdown` and
`ControlCombobox` a `variant="field"` that applies it. The schedule dialog
passes the variant and carries no class strings of its own.

This also repairs a break the dev merge would otherwise have introduced: the
newer `Dropdown` splits `className` (wrapper) from `triggerClassName`, so the
old pasted classes would have landed on the wrapper and left the triggers
unstyled.

The semantic-token guard now watches the shared module and asserts each
primitive still composes it, which covers more than the two files it read
before.

* fix: keep an explicit schedules disable from becoming an opt-in

`use` is two things at once for a dual-purpose runtime interface field: a
permission bit, which DB overrides strip, and the runtime disable signal that
`getLimits` reads. Stripping it from `{ use: false, maxPerUser: 2 }` leaves an
object, and `getLimits` treats any object without `use: false` as enabled — so
an admin override written to stop scheduled billing for a role or user started
it instead.

Collapse an explicit disable to the boolean form before the strip, on both
paths that accept it: the `interface.schedules` field patch, which admitted the
object wholesale because bare runtime paths deliberately bypass the permission
gate, and the overrides merge, which reached the composite-field branch and
kept `maxPerUser`. Objects that only narrow limits are untouched, so a
principal can still be given a smaller cap.

Both regressions fail without the normalizer.

* fix(schedules): Wave A — null-balance CAS, atomic paused-card clear, clustered erasure sweep

Slice 1 (thread r3804518381): route the existing-null balance initialization
through a { user, tokenCredits: null } compare-and-set instead of a blind $set,
so a concurrent initializer/charge landing between the preflight read and the
write is never handed back its spent starting balance. On a CAS miss the
preflight re-reads the winner. The absent-record $setOnInsert path and the
credited-record refill-config sync are unchanged. Adds the initializeNullBalance
adapter (no upsert) and regression coverage for winner/miss/sync cases.

Slice 2 (thread r3804518388): updateScheduleById now drops a `requires_action`
lastRun projection atomically with the configRevision bump. Any pause present at
edit time was projected under the pre-edit revision and can never be replaced by
its own revision-fenced terminal outcome — a disabling edit would strand the
card on "Needs approval" forever. Implemented as classic-operator CAS branches
(DocumentDB rules out a conditional pipeline $unset), fenced on the card STILL
being the pause so a terminal outcome or newer occurrence that races in is
preserved. Terminal history survives untouched.

Slice 4 (thread r3803826204): expose initializeScheduleErasureSweep from the
schedule runtime facade and start it in every clustered (experimental) worker
after Mongo is up. It re-drives eraseScheduleIfDrained for soft-deleted rows so
a hidden prompt cannot outlive its drain when the delete/erase-on-settle
attempts miss. It arms nothing else and never infers owner death from a
process-local missing job (isTopologySafeToArm gates that).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): Wave B.1 — reversible account-deletion schedule suspension

Thread r3804518383. Account-deletion quiesce marked every schedule `deleting`,
disabled it, and cleared nextRunAt destructively. When a later cascade step (or
the drain itself) failed, the controller cancelled the user-deletion fence —
restoring the user — but the schedules stayed `deleting` and were erased by the
sweep, silently losing all of a live user's scheduled prompts.

Replace the destructive marking with a REVERSIBLE, token-fenced suspension:

- suspendUserSchedulesForDeletion(userId, token) snapshots each schedule's prior
  enabled/nextRunAt under a per-attempt token, then fences firing (disable, clear
  nextRunAt, rotate claimToken). It never sets `deleting`, so a suspended row is
  not erasure-eligible. Snapshotting reads then bulkWrites (a classic update
  cannot copy field values under DocumentDB), fenced per row so an already-
  suspended/soft-deleted/edited row is left alone; idempotent per token.
- restoreUserSchedulesFromDeletion(userId, token) reverses it, re-enabling and
  re-arming only rows still carrying the exact attempt token and not independently
  deleted — so an owner-deleted or newer-attempt-suspended schedule is never
  resurrected.
- deleteUserController generates the attempt token, passes it to quiesce, and on
  any failure that cancels the user-deletion fence restores the suspended rows. A
  successful deletion hard-deletes them (and their snapshots) via the existing
  cascade and never restores.

Adds a `deletionSuspension` embedded field (excluded from the wire schedule),
data-method regression tests (suspend/restore/fence/idempotency/no-resurrect),
and controller tests (restore on drain-false and post-quiesce cascade failure,
no restore on success).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): Wave B.2 — wire the interactive Stop persistence protocol

Thread r3804255932. The durable Stop primitives (requestRunAbort 'stop',
getScheduleRunAbortState, markRunAbortPersisted) already existed but production
only used requestRunAbort(..., 'deletion'). An interactive Stop flipped the job
to `aborted` and then persisted its partial message + checkpoint inside
`beforePublish`, without ever stamping the schedule Stop or acknowledging it —
so reconciliation, the generation owner, or a concurrent schedule/account
deletion could terminalize the run and release its capacity (and erase data)
mid-write.

Wire the request -> persist -> acknowledge -> settle barrier through the
schedule runtime (the route never touches raw Mongo):

- Expose beginScheduledStop / acknowledgeScheduledStopPersistence on the service.
- The abort route stamps the Stop BEFORE signalling abortJob (a serialized
  'in_progress' loser returns 409 STOP_IN_PROGRESS without a second abort),
  acknowledges only after beforePublish persistence succeeds, and on a
  persistence failure leaves the barrier unresolved so the run stays preserved
  (client retries; stale-owner timeout is the bounded recovery). A failed abort
  releases the stamp it placed so a replacement/retry is never blocked through
  its predecessor.
- recordScheduleOutcome (the owner settlement path) now waits, bounded, for the
  Stop acknowledgement before terminalizing; a resolved/non-stop/stale marker
  proceeds immediately. The paused Stop settles only after its own ack.

Adds service-layer barrier tests and route-level ordering/persistence-failure/
in-progress tests; the data-layer serialization is already covered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): Wave C.1 — reconcile durable trigger delivery with the reservation

Threads r3803826192 (manual limiter) and r3804255924 (PII/moderation), plus the
independently-found long-Retry-After race. fireSchedule reserves a `started` run
and a global capacity slot BEFORE the durable trigger delivery reaches the chat
route, where an interactive limiter (manual Run Now), PII, or moderation can
reject it before any generation job exists — dead-lettering the delivery while
the run sat `started` until the 30-minute orphan sweep mislabeled it interrupted.
And a valid delivery deferred by Retry-After (up to 24h) could be orphan-settled
and have its capacity released, then fire anyway.

Translate durable delivery state into the schedule outcome:

- Store the deterministic trigger deliveryKey on the ScheduleRun reservation
  (computed from the envelope BEFORE enqueue, so an ambiguous commit still has it).
- Add a getTriggerDelivery engine dep (wired to the merged trigger service's
  getDelivery) that reads the durable delivery by key.
- Schedule reconciliation, for a jobless `started` run: staging/pending/leased →
  admission is live, never orphan; dead → record `error` from the durable
  lastError and release capacity promptly (no 30-minute wait), through the
  ordinary outcome/auto-disable path; succeeded or no record → the existing
  legacy orphan policy (interrupted only past the cutoff); a delivery lookup
  failure defers rather than orphaning a possibly-live delivery. Limiter/PII/
  moderation middleware writes no schedule state.

Adds reconcile state-mapping tests (dead/pending/leased/staging/succeeded/none/
lookup-failure) and a fire test that the reservation's deliveryKey equals the
enqueued delivery's idempotency key.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(stream): Wave C.2 — durable retry/ack for terminal host lifecycle actions

Thread r3804518375. Approval expiry won the `requires_action → aborted` CAS and
then invoked the host hook best-effort: `runApprovalExpiredHandler` swallowed a
failure, and because later sweeps enumerate only `requires_action` jobs, the now-
aborted job was never offered again. In the clustered entrypoint (no schedule
reconciler) the ScheduleRun stayed `requires_action` and its retained job
persisted indefinitely.

Make the host lifecycle work durable rather than schedule-specific:

- Add a generic `terminalHostActionPending` marker, set ATOMICALLY in the same
  terminal transition (ApprovalLifecycle.expireWithIdentity), only when a host
  adapter is installed.
- Retain and index such jobs: both stores keep them out of terminal reaping and
  expose getTerminalHostActionJobs(); Redis adds a set + extended (24h-bounded)
  TTL, in-memory a bounded 24h retention so a permanently-failing hook cannot leak.
- The manager clears the marker only after the adapter acknowledges success,
  fenced by generation identity (clearTerminalHostAction), so a replacement
  generation can neither clear nor execute its predecessor's action.
- cleanup()/expireStaleApprovals() enumerates unacknowledged terminal host actions
  across restarts and replicas and retries the idempotent hook; the relay only
  re-invokes while the marker is unacknowledged, so a successful ack prevents
  duplicate work. Store-won expiry marks it too, so a loser-replica relay still
  crosses the hook.
- Terminal SSE notification continues regardless of host-hook outcome.

Covers in-memory behavior (retry after failure, restart/other-replica retry, ack
prevents duplicates, identity fence, terminal notification on failure, no marker
accumulation for non-scheduled jobs) and updates the Redis cluster-membership
contract test for the new index.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* style(schedules): satisfy import sorting in fire.ts and fire.spec.ts

CI "Static checks" failed on IMPORT_SORT for the two files Wave C.1 added imports
to (the AgentTriggerEnvelope type import and getAgentTriggerIdempotencyKey).
Applied scripts/sort-imports.mts to exactly those files — imports-only reordering,
no behavior change. Other files reported by a repo-wide check are pre-existing on
dev and deliberately left untouched so this PR is not widened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): repair the deletion CLI and order restore before the fence release

Addresses three findings from the fresh Codex review.

P1 — config/delete-user.js called methods.disableUserSchedulesForDeletion, which
Wave B removed in favor of suspendUserSchedulesForDeletion. The file is
@ts-nocheck and its spec mocked the removed name, so neither typecheck nor tests
caught it; the real CLI would throw a TypeError before deleting anything and then
only unwind the fence. The CLI now uses the tokenized protocol: it mints a
suspension token, suspends with it, and restores that exact attempt's rows in its
finally block when the deletion does not commit. Its spec mocks the real methods,
so the breakage can no longer hide.

P2 — both the HTTP controller and the CLI released the user-deletion fence BEFORE
restoring schedules. That fence is what refuses new schedule writes/claims, so the
gap let an owner PATCH edit a still-suspended row and have its enabled/next-run
state overwritten by the older snapshot, and let a second deletion attempt
re-suspend under a new token — making the first restore a no-op and stranding the
disabled snapshot permanently. Restore now runs first, while writes are still
fenced.

Hardening for the same defect class across a crash: suspendUserSchedulesForDeletion
now ADOPTS an existing suspension's snapshot when re-suspending a row abandoned by
an earlier attempt, instead of re-capturing the row's current (already-suspended)
state. Without this, an attempt that died before restoring would have its
successor snapshot "disabled, no next run" and permanently strand the schedule.

Tests: CLI restore-before-fence ordering and no-restore-on-success; the same
ordering assertion on both controller post-quiesce failure paths; a data-method
regression that a second attempt adopts the abandoned snapshot and restores the
original enabled/next-run state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): converge dead deliveries in topology-safe maintenance

Codex finding: the `dead` delivery mapping added in Wave C.1 lives only inside
startScheduleEngine's reconciler, but the clustered entrypoint arms no engine — it
runs erasure-only maintenance. A delivery queued before a restart into clustered
mode and then rejected before generation creation (interactive limiter, PII,
moderation) dead-letters while its ScheduleRun stays `started`, holding a global
capacity slot indefinitely for an ordinary non-deleting schedule.

Add a dead-delivery convergence pass to the erasure sweep, so every topology that
runs schedule maintenance settles it. The pass is POSITIVE-EVIDENCE-ONLY and is
therefore safe where absence-based reconciliation is not: a `dead` delivery is
durable shared state proving no generation owns the reservation. It settles only
when the job is confirmed absent or identity-mismatched (an identity-matched job
still owns the run), defers on an unknown job lookup, on an in-flight abort, and
on an in-flight resume hand-off, ignores legacy reservations with no deliveryKey,
and applies a short grace so an accepted delivery still creating its generation is
never settled mid-handoff. Auto-disable policy is deliberately left to the armed
engine; this path records the failure and frees the slot.

Deliberately does NOT touch api/server/experimental.js — the clustered entrypoint
already starts this sweep, so the convergence arrives through the existing
initializer and the shared-file footprint stays as-is.

Tests: settles a dead delivery as error under an explicitly UNSAFE topology,
leaves live deliveries alone, never settles under an identity-matched running
generation (delivery is not even consulted), defers an in-flight abort, and
ignores a reservation with no deliveryKey.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(stream): refresh host-action retention on each retry attempt

Codex finding: unacknowledged terminal host-action evidence was capped at the
24h pause TTL, so a host dependency (Mongo) unreachable for longer than that let
the Redis key — and its pending marker — expire with no generation-fenced
acknowledgement, stranding the ScheduleRun where no reconciler is armed.

Measure retention from the LAST retry rather than from the terminal transition:
enumerating a pending host action IS the retry attempt, so both stores refresh
its retention as they hand it to the hook (Redis re-EXPIREs the job key; the
in-memory store stamps terminalHostActionRefreshedAt and bounds from it). Evidence
therefore survives as long as some replica is still actively retrying, while a
deployment that stops sweeping entirely still lets it age out — so this does not
reintroduce the unbounded leak the cap existed to prevent.

Test: after a failed hook, a later cleanup pass keeps the marker pending and moves
its retention basis forward.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): defer settlement when the Stop barrier times out

Codex finding: waitForStopPersistence returned after its 5s poll budget even when
the Stop was still fresh and unacknowledged, and recordScheduleOutcome then went
straight on to terminalize the run — releasing its capacity, deletion, and erasure
barriers while beforePublish may still have been writing. Slow checkpoint cleanup
is indistinguishable from a dead route on that signal, so the timeout was being
treated as if the barrier had been satisfied.

The poll budget now means "undecided", not "clear". On timeout with a fresh,
unacknowledged Stop the barrier DEFERS: recordScheduleOutcome returns false
without recording, leaving the run active/preserved. Settlement then happens
either when the route acknowledges, or once the existing stale-owner cutoff
(ABORT_OWNER_PRESUMED_ALIVE_MS) authorizes a later attempt — which the loop
already treats as clear-to-settle. Callers with durable retry (the approval-expiry
host action, reconciliation) re-drive it, so a deferral converges rather than
stranding the run.

Test: a fresh Stop that never acknowledges within the budget reports not-settled
and records no outcome.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): require definite delivery failure, extract its message, converge deferred Stops

Third Codex round. Three findings in code from this closeout, plus one pre-existing
P1 that is a one-line operator-safety fix.

P1 — dead-delivery settlement demanded too little evidence. `dead` does not prove a
request was rejected: the trigger host marks response timeouts and invalid success
responses `certainty: 'ambiguous'`, and the engine dead-letters those once retries
are exhausted. The erasure sweep treated every dead letter as positive evidence, so
an ambiguous one sitting over a generation a peer had accepted could terminalize the
run and release its capacity mid-flight. It now settles only on a DEFINITE rejection,
unless job absence is deployment-authoritative (safe topology), where the
confirmed-absent job is itself the evidence.

P1 — `lastError` is an `AgentTriggerDeliveryFailure` object, not a string. A
duplicated local interface declared it `string` (against CLAUDE.md's no-duplicate-
types rule), so both the sweep and the engine reconciler passed the object into the
String-typed run/schedule `error` fields; Mongoose would reject the cast, the per-row
catch would swallow it, and the run would keep its global capacity slot. The dep type
now reuses the canonical `AgentTriggerDeliveryFailure` and both call sites pass
`.message`. Re-typing immediately surfaced a stale test that had asserted a string.

P1 (pre-existing) — the base-config global stop is honored in `getLimits` via
`isRuntimeDisabled` rather than a literal `=== false`. The stop has two shapes, and
deepMerge turns base `{ use: false }` plus a principal override of `true` into
`{ use: true }`, so the literal check reported the feature enabled and Run Now
dispatched straight through fireSchedule, bypassing the operator's emergency stop.
Now the same predicate the engine gate already uses.

P2 — a Stop whose settlement DEFERRED past the poll budget had no convergence path
where no reconciler is armed. `acknowledgeScheduledStopPersistence` now optionally
re-drives the terminal outcome once the barrier clears; `recordRunOutcome` is
match-guarded and idempotent, so an owner that already settled makes it a no-op. The
abort route passes it for a running generation; a paused job still settles explicitly.

Tests: ambiguous dead letters refused under unsafe topology but settled when absence
is authoritative, definite rejections settled either way, the failure message carried
through, and the abort route's re-drive present for running / absent for paused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): converge terminal runs in clustered workers, retry suspension restore

Fourth Codex round on 74f7af8. Both findings are in code from this closeout.

P2 - a clustered worker never settled a run whose generation finished but whose
outcome write failed. `recordScheduleOutcome` retries three times and then returns
false; the owner honors that by PRESERVING the terminal job as the only surviving
evidence, and the armed engine's reconciler replays exactly that. The clustered
entrypoint arms no engine, and its sweep only covered deleting schedules
(settleAbandonedRuns) and dead deliveries - an identity-matched job was skipped
outright. So an ordinary live schedule's run stayed `started`, held its GLOBAL
capacity slot, and kept a preserved job that carries no `completedAt` and is
therefore invisible to the store's finished-job sweep: both leaked until store
expiry.

The live-schedule pass now also converges from a retained terminal job, mirroring
the reconciler branch it stands in for: honor the owner's stamped outcome over the
generic status (so a balance refusal still walks its streak rather than resetting
it), clear the reserved conversationId when the generation never emitted its
created event, and delete the retained job only AFTER the outcome write is durable.

This stays positive-evidence-only and safe in every topology. Presence, not
absence, is the evidence: an identity-matched job is authoritative wherever it is
observed - a shared store shows the real generation, a process-local store can
only be showing this process's own - which is why it needs no
canInferOwnerDeathFromMissingJob fence, unlike the absence-based paths. The
in-flight abort/resume fences still defer, so an `aborted` job cannot settle a run
whose owner may still be persisting. Both cases now share one pass over one window
rather than two, and `retainedOutcome` moved to types.ts so the sweep reuses the
engine's mapping instead of duplicating it (and stays independent of the engine).

P2 - a failed restore stranded a live account's schedules. Cancelling an account
deletion restores the suspended rows while the deletion fence is still armed and
then releases that fence; nothing re-drives the restore afterwards, so one
transient write failure left the user with silently disabled, next-run-less
schedules. The restore is now retried at the single choke point both the HTTP
controller and the CLI share. Retrying is safe because each attempt re-reads only
the rows STILL carrying the token: a partially-applied unordered write converges
on exactly the stragglers, and a fully-applied one finds nothing.

The fence is still released when every attempt fails, deliberately: retaining it
would refuse the live account's schedule writes AND make beginAgentTriggerUserDeletion
report `in_progress` forever, blocking the retry that is the convergence path (a
later attempt adopts this snapshot, so its cancel restores these exact rows). Both
callers now log the user id and suspension token so the state stays recoverable by
hand if that never happens.

Tests: retained terminal job settled and its evidence released, stamped outcome
preferred over the generic status, stamped failure reason carried, reserved
conversationId cleared for a never-created conversation, identity-mismatched
terminal job ignored, abort-in-flight deferred; restore retried past a transient
write failure and converging on a partially applied restore.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): converge unprojected pauses and release replayed bookkeeping jobs

Two of the four findings I flagged as pre-existing to my closeout commits. Both are
in code this PR introduces, so merging would have shipped them; both are the same
capacity/evidence-leak class the last four rounds have been closing. The remaining
two (slotless rows escaping the per-user cap, and the non-rotating `started`
reconciliation bucket) are genuinely latent and deliberately left alone.

An UNPROJECTED PAUSE held a global capacity slot forever. The pause projection is
what moves a run row off `started`; `recordScheduleOutcome` already retries it, but
the request controller discarded the result, so three failed attempts were dropped
silently. The armed engine's reconciler replays that state, but the clustered sweep
did not: a paused job is not terminal, so the retained-job path returned without
settling, and the dead-delivery path never inspects an identity-matched job at all.
The row stayed `started` with no cutoff that would ever clear it.

The call site now surfaces the failure, and the sweep converges it, mirroring the
reconciler's pause branch: project `requires_action` (which frees the slot) but do
NOT release the job's evidence — unlike a terminal job it is still live, awaiting an
approval. The resume hand-off fence still defers, so re-projecting cannot release a
slot a continuation just claimed.

The BOOKKEEPING REPLAY pass leaked its retained job. A run reaches that pass only
because its owner crashed before bookkeeping — which is also before it could release
the job it retained for exactly this recovery. The active-run pass clears its own;
this one never did, and a preserved job is deliberately kept WITHOUT `completedAt`
so the store's finished-job sweep cannot reap it early, so nothing else ever would.
It now clears after `finalizeBookkeeping` succeeds — identity-guarded, a no-op when
no job is retained, and skipped entirely when the replay itself failed, since the
retained job is the only surviving evidence in that case.

Tests: pause projected and the slot freed while the live job's evidence is kept,
paused job deferred during a resume hand-off, retained job released once replayed
bookkeeping is durable, and retained job kept when the replay fails.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): fence the clustered pause replay against a concurrent resume claim

Codex round 5, on code I added in f735341. The finding is correct and the exposure
is one I introduced.

The pause replay lives in a sweep that runs in EVERY clustered replica, so several
sweepers can observe the same unprojected pause. Its guards were all derived from an
in-memory row SNAPSHOT: hasResumeHandoffInFlight reads the snapshot's
resumeClaimedAt, and recordRunOutcome's pause branch matched any row currently in
`started`/`requires_action` with no fence of its own. The race: sweeper A projects
the pause and frees the slot; the owner's approval then claims a fresh one
(markRunResumeClaimed takes the row to `started` WITH resumeClaimedAt in one write);
sweeper B, still holding the pre-projection snapshot, passes its hand-off check and
replays — `$unset: { capacitySlot, resumeClaimedAt }` under a continuation that is
already running. The run reverts to `requires_action` while its generation proceeds
outside global capacity.

The engine's reconciler makes the same call and has the same snapshot-derived guard,
but v1 arms exactly one engine, so it has no concurrent racer; the sweep is the first
thing to run this transition in parallel. Left the engine alone rather than widening
the change: its `requires_action` re-affirmation is deliberate and single-writer.

Fixed where the race is, in the write itself. `recordRunOutcome` takes an optional
`requireNoResumeClaim`, which adds `resumeClaimedAt: { $exists: false }` to the pause
filter, and the sweep sets it. Because the stamp is written in the SAME update that
moves the row to `started`, its absence is atomic proof no resume owns the row. The
flag is deliberately NOT set by the generation owner: its own re-pause legitimately
clears the stamp as the hand-off's completion signal.

The fence cannot block the recovery it exists to enable: markRunResumeClaimed only
matches `requires_action`, so a genuinely stuck `started` row can carry no resume
claim — it is in fact blocking its own approval until this replay frees it.

Tests: a replay from a stale snapshot leaves a resume-claimed row's status, slot, and
claim stamp intact, while a stuck row with no claim still recovers; and the sweep is
asserted to send the fence.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

* fix(schedules): fence stale resume claims by age, not existence

Codex round 6, on the fence I added in ee00c60. Correct again, and the defect is
the mirror image of the one it fixed.

The fence was an EXISTENCE check (`resumeClaimedAt: { $exists: false }`) while its
caller's guard is a FRESHNESS check (hasResumeHandoffInFlight, bounded by
RESUME_HANDOFF_STALE_MS). They agree while a claim is fresh and disagree once it is
abandoned: a worker that dies after markRunResumeClaimed takes the row to `started`
and stamps resumeClaimedAt — but before the continuation resumes or
releaseRunResumeClaim rolls it back — leaves the stamp set forever. Past the bound
the sweep correctly stops deferring and tries to recover the row, but the write
rejected it purely because the field still existed. The row stayed `started` holding
its global capacity slot, and its approval was unresumable for good, since
markRunResumeClaimed only matches `requires_action`. That is exactly the stuck state
this replay exists to clear, so the fence had reintroduced it for the crashed-resume
case.

`requireNoResumeClaim: boolean` becomes `resumeClaimStaleBefore: Date`, and the
filter matches a row with no claim OR a claim older than that cutoff. The sweep
passes the SAME bound its in-flight check uses, so the two can no longer disagree.
A genuinely racing claim is by construction fresh — it is created after the sweeper's
snapshot — so the race from round 5 stays closed.

Tests: a row whose resume claim was abandoned by a dead worker now recovers (status,
slot and stamp all cleared), alongside the existing two — a fresh claim still repels
a stale-snapshot replay, and an unclaimed stuck row still recovers.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZwZPDBgxy4mCWWg6zQSjq

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-20 11:51:30 -04:00
Marco Beretta
7569404a7c
🎛️ feat: Make Max Subagents Configurable via librechat.yaml (#15023)
* feat: make max subagents configurable via endpoints.agents.maxSubagents

The per-agent subagent cap was hardcoded at 10 in MAX_SUBAGENTS, leaving
orchestration-heavy deployments no option but patching limits.ts and
rebuilding. Add an optional endpoints.agents.maxSubagents key to
librechat.yaml (default 10, hard ceiling 50) that drives request
validation, model spec presets, and the agents panel UI cap.

* style: fix import order in OrchestrationHub
2026-08-20 11:16:08 -04:00
Danny Avila
a47ba7168f
🪢 feat: Custom Request Headers For Langfuse (#14945)
*  feat: Custom Request Headers For Langfuse

Self-hosted Langfuse behind an authenticating proxy or gateway could not
be reached: every outbound Langfuse request hardcoded `Authorization` and
nothing else. Adds `langfuse.headers`, mirroring `endpoints.custom`
headers, and applies it to all four request surfaces — trace/media export
(via the agents run config), feedback scores, central project-identity
lookup, and admin credential verification.

Values resolve through the same pipeline as endpoint headers, so
`${ENV_VAR}` interpolation and header-safe encoding come along.
`extractEnvVariable` continues to refuse infrastructure secrets, so a
config cannot exfiltrate `MONGO_URI` through a header. A header whose
variable is unset is dropped with a one-time warning rather than sent as
a literal `${...}`, which a gateway would read as a wrong credential
instead of a missing one.

Headers merge beneath LibreChat's own `Authorization` on the REST
surfaces, matching `mergeHeaders`, so a custom header can never displace
the Langfuse credential.

These are deployment-level and documented as such: trace export batches
spans from every user through a single exporter, so unlike endpoint
headers they cannot carry per-user placeholders.

The central project-id cache key now includes the headers, so the
header-less module warm-up cannot record a proxy rejection against the
entry the request path later reads.

* 🔒 fix: Keep Langfuse Headers Out Of Stored Overrides

The generic admin config API accepts any field path inside an allowed
section, so `langfuse.headers` could be written through it. Unlike
`langfuse.secretKey`, headers are a map rather than one scalar path, so
the config secret registry cannot encrypt them at rest or mask them on
read — an admin-written map would sit in Mongo in plaintext and come
back in plaintext, widening exposure of what are gateway credentials.

Rejects them on both the dotted-patch and object-upsert routes, the same
way process-backed MCP servers are held to librechat.yaml. This is what
makes "deployment-level" true rather than merely documented.

* 🐛 fix: Wire Config Middleware And Header Collisions For Langfuse

Two codex review findings.

P1 — `api/server/routes/admin/langfuse.js` never mounted
`configMiddleware`, so `req.config` was undefined in production and
credential verification silently ran without the deployment's proxy
headers: exactly the deployments this feature targets could not save a
connection. The handler unit tests injected `config` into their mock
requests, so they stayed green. Mounts the middleware after the access
checks (unauthorized callers still short-circuit first) and adds
route-level tests that assert the handler actually receives a resolved
config — the composition root, not the component.

P2 — spreading custom headers under `Authorization` only replaced an
exact-case collision. A configured `authorization` survived alongside
the managed `Authorization` and fetch appends rather than replaces,
sending both credentials in one combined value. All four request sites
now use `mergeHeaders`, which already merges case-insensitively with the
override winning; tests cover the lower- and upper-case variants.

* 🔒 fix: Mask Langfuse Headers On Read And Harden Value Handling

Three codex round-2 findings.

P1 — the write guard blocked storing `langfuse.headers` in Mongo but did
nothing for the read path: `GET /api/admin/config/base` serves the
resolved AppConfig through `redactConfigSecrets`, which only knows
registered scalar secrets, so a yaml-configured gateway credential was
returned in full to any admin with Langfuse read access. Adds a
secret-map registry that masks values while keeping key names, so an
admin can still see which headers a deployment sets. Masking is safe
precisely because these are yaml-only — a masked read cannot be
round-tripped back over the real values. A malformed non-object value at
that path is dropped rather than serialized.

P2 — `mergeHeaders` indexes one spelling per lowercase name, so a config
holding both `authorization` and `AUTHORIZATION` had only one displaced;
the survivor was then appended by `Headers` into a combined value. Case
variants are now collapsed at resolution, before any consumer sees them.

P2 — `resolveHeaders` encodes only values it substitutes a user field
into, and no user is supplied here, so a literal or interpolated
character above U+00FF reached `Headers` unencoded and threw. Final
values now go through `encodeHeaderValue`; Latin-1 still passes verbatim.

* 🔒 fix: Keep Langfuse Header Credentials Out Of Logs And Validate Names

Three codex round-3 findings, plus a documented boundary for the fourth.

P1 — `loadCustomConfig` logs the parsed config at startup (`printConfig`
defaults true), so a literal gateway credential in `langfuse.headers` was
copied into application logs on every boot, undoing the masking the admin
read path had just gained. The printed copy now goes through
`redactConfigSecretMaps`, reusing the same registry. Scoped to map-valued
secrets so scalar-secret log behavior is unchanged; the live config keeps
its real values.

P2 — a nonempty but invalid field name (` X-Token`, `X Proxy Token`)
passed the emptiness check and then threw in the `Headers` constructor,
which would break export, verification, lookup, and feedback for the whole
deployment rather than that one header. Names are trimmed and validated
against the RFC 7230 token grammar, and dropped with a warning otherwise.

P2 — unresolved `${VAR}` detection tested the *resolved* value, so a
credential legitimately containing `${...}` was mistaken for a failed
substitution and dropped. Detection now inspects the configured text and
checks the referenced variables directly, which also drops references to
denylisted infrastructure secrets instead of forwarding them verbatim.

The fourth (fanout gateway forwards only `Authorization`, so a tenant
Langfuse behind its own proxy is not covered) is a real limitation in a
separate component. Documented on the schema field and in the example
config rather than left implied.

* 🔒 fix: Scope Langfuse Headers To Configured Origins

Three codex round-4 findings.

P1 — one header map was attached to every destination a run resolves to.
Under fanout that means a credential meant for an internal gateway was
also sent to the central destination, typically Langfuse Cloud: an
unrelated third-party origin. Headers are now attached only when the
destination's origin is one the deployment explicitly configured (a
self-hosted base URL, the fanout collector, or a tenant destination set
by env). The built-in `*.cloud.langfuse.com` defaults are excluded
precisely because nobody pointed at them. For trace export this also
means attaching after the export branch settles on a `baseUrl` rather
than before, since which destination wins depends on the branch.

P2 — `encodeHeaderValue` only encodes above U+00FF, so a newline, CR, or
NUL passed through and threw in `Headers`, breaking every request rather
than the one header. Values are trimmed (the common trailing-newline
case) then validated against the legal field-value bytes; an embedded
CRLF is a request-splitting attempt and is dropped, not stripped.

P2 — the write guard matched only the exact `headers` property, so
`{ langfuse: { "headers.X-Token": "..." } }` and root-level dotted
variants slipped through into the Mixed overrides document, where the
nested-map redactor never walks them and a later read returns them in
plaintext. All dotted spellings are now rejected.

* 🔒 fix: Bind Langfuse Headers To One Configured Origin

Four codex round-5 findings.

P1 — the round-4 allowlist still authorized every configured origin, so a
deployment with both a collector and an explicit central host sent the
same credential to both. `langfuse.headers` is one map with no way to say
which endpoint it authenticates to, so it is only unambiguous when the
deployment configures exactly one Langfuse origin. Iterating on which
origins to guess was the wrong axis; headers are now sent only when there
is a single configured origin and the destination is it, with a warning
when several make the intent unresolvable. That covers the self-hosted
case this feature exists for; multi-destination deployments need
per-destination headers the schema cannot yet express.

P1 — `fetch` defaults to following redirects, and Node strips
`Authorization` across origins but keeps arbitrary headers, so a redirect
off an allowed origin would hand the gateway credential to a host that
passed no check. Requests carrying custom headers now refuse redirects;
requests without them keep the default, so nothing changes for existing
deployments.

P2 — `extractEnvVariable`'s whole-string branch is anchored and greedy, so
`${CLIENT_ID}:${CLIENT_SECRET}` parsed as one variable name and the raw
template was sent as the credential. References are expanded here now, so
only literal values reach that path.

P2 — a valid token name is not necessarily usable: `Transfer-Encoding`
makes `fetch` throw and a fixed `Content-Length` misdescribes the body of
every other request sharing the map. Request-framing names are dropped.

* 🐛 fix: Expand Langfuse Header References Exactly Once

Codex round 6 (P2). After expanding `${VAR}` references myself I still
handed the result to `resolveHeaders`, which runs `extractEnvVariable`
over it again — so a credential containing `${PATH}`, or any other name
that happens to be set, was silently rewritten on export, verification,
lookup, and feedback. The round-3 test only used an *unset* embedded
name, which the second pass leaves alone, so it could not catch this.

Resolution no longer round-trips through `resolveHeaders`. The only part
still wanted from it was stripping `{{...}}` user placeholders, which is
now applied directly; expansion, encoding, and validation were already
local. Adds a test whose embedded variable is set, which fails against
the previous pipeline.

* 🐛 fix: Process Langfuse Header Templates Before Substitution

Codex round 7 (P2), the mirror of round 6. Having stopped re-expanding
the resolved credential, the placeholder strip was still running over it:
a token containing `{{LIBRECHAT_USER_ID}}` had that span deleted and
`abc{{...}}ghi` went out as `abcghi`.

Establishes the invariant the last two rounds were circling. Every
template operation — placeholder strip, unresolved-reference check,
expansion — now runs on the operator's configured text, and the
credential is substituted last and never touched again. Gateway
credentials are arbitrary strings, so none of their bytes are syntax.

Also moves the unresolved-reference check after the strip, so it no
longer reports a variable inside a `{{...}}` span that the strip removes.
2026-08-17 20:39:23 -04:00
Danny Avila
57ea1137f6
🛡️ feat: Let Admins Restrict Stateful Workspace Scopes (#14910)
* feat: let admins restrict stateful workspace scopes

* fix: enforce stateful scope policy across agent paths

* fix: close stateful scope policy activation gaps
2026-08-17 01:29:19 -04:00
Danny Avila
7d850c308a
🧠 feat: Add Live Reasoning Labels (#14893)
* feat: add live reasoning labels

* fix: Stabilize reasoning label checks

* fix: Address reasoning label review findings

* chore: Bump Agents SDK for reasoning labels

* fix: Reset reused reasoning step evidence

* fix: Reconcile cleared reasoning labels

* fix: Fence reasoning label resets

* fix: Reset reasoning ownership before gap labels

* fix: Preserve THINK type through label reset

* test: Expect run-global reasoning revision
2026-08-16 18:11:55 -04:00
Marco Beretta
3bd2358805
🗜️ feat: Let Users Toggle Client-Side Image Resizing (#14883)
* feat: allow users to toggle client-side image resizing

Client-side image resizing could only be configured in librechat.yaml and
defaulted to off, so users had no way to enable it for themselves.

Add a "Resize images before upload" toggle in Settings > Chat > Sending,
stored per device in localStorage. When librechat.yaml sets
clientImageResize.enabled, mergeFileConfig marks the value as enforced and
the toggle renders the admin value read-only. Admin resize parameters still
apply without locking the toggle when enabled is omitted.

Also fix shouldResizeImage, which compared file size against 10% of a 512MB
fallback limit and so only triggered above roughly 51MB. It now uses a 512KB
floor, which lets the setting affect everyday photos.

* fix: harden client image resizing

* fix(client): restrict image resizing and localize toast

* fix: harden client image resizing

* fix(client): recognize static WebP image chunks

* fix(client): clamp resized image dimensions

* fix(client): enforce safe image resize uploads

* fix(client): recheck duplicates after transforming uploads

* fix(client): keep selected files after input reset

* fix(client): disable image resizing when file config fails

* fix(client): decode resize candidates without a base64 copy

* fix(client): coordinate upload batches across hook instances

* test(client): drop the untyped conversation from shared upload state

* fix(client): start the config wait before queueing uploads

* fix(client): disable the resize switch while file config is pending

The switch stayed clickable during the initial file-config load even
though the checked state cannot update until that query settles.

* fix(client): only track upload reservations for observable state
2026-08-16 18:04:03 -04:00
Danny Avila
eaef87fa26
🚀 chore: Prepare v0.8.8-rc1 (#14394)
* 🚀 chore: Prepare v0.8.8-rc1 release

* 📚 docs: Complete v0.8.8-rc1 operator references

* 📚 docs: Mark stateful sessions experimental

* 📚 docs: Clarify background code capability

* 📚 docs: Refresh v0.8.8-rc1 operator guidance

* 📚 docs: Highlight v0.8.8-rc1 features in README

* 📦 chore: Bump publishable packages again

* 📚 docs: Add streaming question progress

* 📦 chore: Bump publishable packages again

* 📚 docs: Refresh v0.8.8-rc1 release highlights

* 📦 chore: Bump publishable packages again

* 📚 docs: Refresh v0.8.8-rc1 release guidance

* 📦 chore: Bump publishable packages again

* 📚 docs: Highlight batched Agent questions

* 📦 chore: Bump publishable packages again

* 📦 chore: Bump publishable packages again

* 📦 chore: Bump publishable packages again

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📄 docs: Note PowerPoint template support

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📄 docs: Note latest provider and file support
2026-08-14 03:24:59 -04:00
Danny Avila
a3cec67e08
🪆 feat: Add Parent Activity Phase Summaries (#14721)
* feat: add activity phase summaries

* fix: preserve activity phase lifecycle semantics

* fix: satisfy activity phase type checks

* fix: simplify activity phase status mapping

* style: format activity phase changes

* fix: rebase activity phase bounds after shaping

* fix: link activity phase trace ancestry

* fix: reconcile activity phase bounds

* style: format activity phase reconciliation test

* style: align activity phase assertion

* fix: retain reasoning across commentary

* fix: preserve activity phase boundary state

* fix: detect renderable phase children

* test: type parallel phase assertion

* chore: bump agents SDK for activity phases

* fix: retain unphased lane reasoning

* fix: preserve tool group expansion across phases

* style: format phase expansion regression

* fix: preserve phase interaction state efficiently

* perf: skip sparse phase segment holes

* perf: partition phase segments with offsets

* fix: preserve phase boundaries and cursor state

* test: align activity phase regressions with CI

* test: keep phase context mock hoist-safe
2026-08-10 13:41:37 -04:00
Dustin Healy
87a8b9aa12
📡 feat: Route Web Search and Scrape Egress through the SSRF-safe Agent (#14606)
* feat(web-search): route outbound search and scrape requests through the SSRF-safe agent

Build the SSRF-safe agents at the web-search tool-assembly site and pass them into the
search tool config so outbound search and scrape connections are validated at connect
time against their resolved IP, on every hop including redirects, consistent with the
other outbound clients.

Add allowedAddresses to webSearchSchema, reusing allowedAddressesSchema, so self-hosters
can permit a deliberately-private search or scrape endpoint (for example a private
SearXNG instance). The field is resolved directly from the webSearch config at the
createSearchTool call site, not through loadWebSearchAuth, because it is config and not
an auth credential. webSearchSchema is flat (providers are chosen by enums, not by
counting keys), so the field is inert with respect to provider selection.

Document the field and its operator warning in librechat.example.yaml, and assert the
wiring in handleTools.test.js: the SSRF-safe agents are threaded into the search tool
config, allowedAddresses is passed through when set, and omitting it still threads the
agents with no exemptions.

TODO awaits @librechat/agents release with the httpAgent hook: this consumes optional
httpAgent/httpsAgent fields on the search-tool config that are not yet in a published
@librechat/agents. package.json is intentionally left at the current version; bump it to
the release that ships the hook before this lands. Validated locally against a revendored
@librechat/agents build, not a published release.

* 🛡️ fix: Apply allowedAddresses to the Web Search SSRF Preflight

The connect-time SSRF agent already honors webSearch.allowedAddresses, but
loadWebSearchAuth ran the isSSRFUrl preflight without it, so an admin-permitted
private search or scrape URL was stripped before the agent could ever use it.

Thread allowedAddresses and the URL's effective port through isSSRFTarget and
resolveHostnameSSRF so the exemption is consistent across both SSRF layers.

* 🛡️ fix: Validate Web Search Destinations and Defer to Configured Proxies

Handing agents to createSearchTool covered only the connect-time DNS lookup,
which Node skips for IP-literal hosts, and a configured proxy connects on our
behalf without running that check. A literal private target such as
http://169.254.169.254 could therefore reach the network.

Route every resolved web-search destination through the existing
applySSRFSafeAgentIfDirect contract so a blocked literal target throws before
any request is made, and withhold the agents when a proxy owns egress, since one
agent pair is shared by every provider and a direct-connect agent on a proxied
connection would break the request while asserting protection the proxy's
network context cannot provide.

* 🛡️ fix: Keep Web Search SSRF Agents Under a Proxy and Restore Pooling

Withholding the agents whenever a proxy was configured removed protection from
every direct and NO_PROXY destination in exchange for preventing a failure that
cannot occur: for an https target Axios substitutes its own CONNECT tunnel, so
the injected agent is never used for the proxy connection. Only a plaintext http
target keeps our agent and repoints it at the proxy, and only a proxy whose
hostname resolves private then trips the connect-time check.

Always pass the agents and exempt the proxy endpoint instead, deriving host:port
from the same PROXY, HTTP_PROXY, and HTTPS_PROXY resolution the rest of LibreChat
uses so the proxy hop stays reachable while destinations remain guarded. Axios
already applies NO_PROXY per request, so bypassed routes keep enforcement with no
extra logic.

Drop the load-time destination validation. It duplicated the isSSRFTarget
preflight for user-provided URLs, rejected admin values that were previously
legal, and threw from inside loadTools, where both loader wrappers swallow the
error and drop every tool for the turn rather than degrading web search alone.

Build the agents with keepAlive and cache them per exemption list. A bare
http.Agent does not pool, so the previous code replaced the pooled global agents
for every search, scrape, and rerank call and allocated a fresh pair per turn.

* 🛡️ fix: Reject IP-Literal Private Targets on Web Search Connections

Node resolves nothing for a literal host, so the connect-time lookup never saw
one: a destination or a redirect target given as http://169.254.169.254 reached
the network. Redirect hops pass through the same createConnection, so checking the
literal there covers both cases and removes the need for a maxRedirects control
that createSearchTool cannot accept.

Gate it behind blockLiteralHosts so only web search opts in. A caller that
reaches a proxy or a deliberate private service by literal address must exempt it
first, and the merged consumers of createSSRFSafeAgents have no such exemption, so
enabling this everywhere would break configurations that work today.

* 🛡️ fix: Keep IPv6 Brackets on Derived Proxy Exemptions

The exemption parser accepts an IPv6 entry only as [ipv6]:port, so stripping the
brackets produced fd00::1:3128, which carries three colons and is dropped as
malformed. An IPv6 proxy therefore stayed unexempted and the connect-time check
rejected it, failing every web-search request routed through it.

Use the URL hostname as parsed, which already carries the brackets.

* 🛡️ fix: Exempt Proxies Configured Through ALL_PROXY

Axios resolves a proxy through proxy-from-env, which falls back to all_proxy in
either case after <protocol>_proxy, so ALL_PROXY on its own is enough to route a
request through a proxy. Exemptions were derived from PROXY, HTTP_PROXY, and
HTTPS_PROXY only, leaving such a proxy unexempted and rejected with ESSRF.

Derive the exemptions from the full set of variables that can put a proxy in
front of these requests instead. The installed proxy-from-env 2.1.0 reads no
npm_config variables, so those are deliberately not included.

* 🛡️ fix: Drop the Unearned PROXY Exemption and Harden the Web Search Guard

Nothing on this path consumes PROXY: Axios resolves proxies through
proxy-from-env, which reads only <protocol>_proxy and all_proxy, and web search
never calls applyAxiosProxyConfig. Exempting it therefore granted a bypass rather
than preserving a working route, and a user-settable search URL that redirects to
that address reached it and returned the body. Remove PROXY and proxy, and skip a
socks endpoint for the same reason, since Axios cannot proxy through one.

Tolerate a non-array allowedAddresses instead of spreading it, which threw out of
loadTools and dropped every tool for the turn. The YAML path is schema-validated
but the admin override path merges without parsing, so the value is reachable.

Separate cache keys with NUL rather than a newline, so an entry containing a
newline cannot collide with two separate entries, and bound the cache. Give the
agents the idle timeout the global agents carry, which keepAlive alone did not
restore. Reject a unix socket, which carries no host to validate. Also treat
fec0::/10 site-local as private, matching the fe80::/10 handling beside it.

Exercise the real resolver in handleTools.test.js rather than mocking it, so the
wiring test now fails if the agents it threads do not actually block a private
target.

* 🛡️ fix: Derive Proxy Exemptions Through Axios's Own Resolver

Unioning every populated proxy variable exempted addresses that never carry a
request. proxy-from-env picks a protocol-specific variable before all_proxy and
lowercase before uppercase, so an ignored value became a trusted host:port that a
redirect onto a direct route could reach. It also normalizes a scheme-less value
such as proxy.internal:3128 to an http URL, where parsing the raw string yielded
an empty hostname and no exemption at all, breaking the proxy hop.

Resolve through getProxyForUrl, the entry point Axios itself calls, so precedence,
scheme normalization, and NO_PROXY match exactly and cannot drift. NO_PROXY
covering everything now yields no exemption, since nothing is proxied. Declared
locally rather than adding a types package, alongside the existing declaration in
the same directory.

Also revert the fec0::/10 site-local change. domain.spec asserts that boundary
deliberately to prove the fe80::/10 mask does not over-reach, and the shared
address schema still classifies fec0 as public, so a runtime block there would
leave operators unable to configure the exemption. It belongs with those two
together, not in this PR.

* 🛡️ fix: Resolve Proxy Exemptions Against the Real Destinations

Resolving against placeholder probe hosts applied destination-specific NO_PROXY
rules to a host nobody dials. With NO_PROXY matching the probe domain but not a
real provider, no exemption was derived even though Axios still proxied the actual
request, so the agent rejected the private proxy hop with ESSRF.

Resolve per configured destination instead, passing the values loadWebSearchAuth
already resolved. Only plaintext http destinations are considered, since for an
https destination Axios substitutes its own CONNECT tunnel and never uses the
injected agent for the proxy connection, which is also why provider defaults need
no exemption: every one of them is https.

* 🛡️ fix: Accept Embedded-IPv4 IPv6 Forms in the Address Exemption Schema

The runtime guard blocks 6to4, NAT64, and Teredo addresses whose embedded IPv4 is
private, but the schema's local copy recognized only ULA, link-local, and the
dotted IPv4-mapped form, so an entry such as [64:ff9b::a00:1]:8080 was dropped as
a public literal. An operator reaching a private endpoint that way could not
configure the exemption at all.

Mirror hasPrivateEmbeddedIPv4 in the schema helper, which the surrounding comment
already asks to keep in sync. Public embedded addresses stay rejected, since an
exemption there has no defensive purpose.
2026-08-10 10:33:54 -04:00
Danny Avila
33c8801ba6
🌊 feat: Wire Adaptive Stream Smoothing Across Google, Bedrock, and Zero-Disable Semantics (#14660)
* 🌊 feat: Wire Adaptive Stream Smoothing Across Google, Bedrock, and Zero-Disable Semantics

With @librechat/agents 3.4.0 smoothing defaults ON (25ms adaptive) for
every provider; this completes the LibreChat side:

- google: fix the unguarded endpoints.all clobber and wire streamRate
  into llmConfig._lc_stream_delay — previously read from config and
  silently dropped, making google streamRate a no-op end to end
- bedrock: wire streamRate (endpoint + endpoints.all) — previously
  absent entirely, so bedrock smoothing was unreachable from config
- anthropic/openai: nullish guards so streamRate: 0 survives as the
  explicit smoothing disable; drop the azure 30/17 hardcoded fallback
  the SDK default now supersedes
- delete dead createHandleLLMNewToken (no call sites since #6886, and
  LangChain backgrounds callbacks so a sleep there never paced anything)
- schema: streamRate gains .min(0) and docs; example docs updated,
  including pairing STREAM_DELTA_COALESCE_MS with the smoothing tick

* 🩹 fix: Address Review — Keep Published Shim, Typed Delay Access, Scoped Docs

- restore createHandleLLMNewToken as a @deprecated compatibility shim:
  it ships in the public @librechat/api root, so removal is reserved for
  a major release
- assign llmConfig._lc_stream_delay via the SDK's typed property (3.4.0
  StreamSmoothingOptions) instead of Record<string, unknown> casts at
  all five sites
- scope the 25ms-default wording to agents SDK-backed providers (legacy
  Assistants/Ollama still per-chunk sleep at DEFAULT_STREAM_RATE=1) and
  clarify coalescing guidance when streamRate: 0
2026-08-06 12:39:12 -04:00
Dustin Healy
96404bfc73
🛡️ fix: SSRF-Guard Speech (STT/TTS) and OCR Outbound Requests at Connect Time (#14560)
* fix: SSRF-guard speech (STT/TTS) and OCR outbound requests at connect time

Speech (STT/TTS) and OCR issued outbound HTTP to operator-provided target URLs
with only proxy config attached, so a target that resolves to a private, loopback,
link-local, or cloud-metadata address was reachable by the server. This is the same
class already guarded for the custom models fetch, Actions, avatar, MCP, and the
OpenAI/Anthropic endpoint clients.

Add one helper applySSRFSafeAgentIfDirect(config, url, allowedAddresses) that rejects
non-http(s) or unparseable target URLs, sets maxRedirects to 0 so a redirect cannot
bypass the connect-time check, and attaches createSSRFSafeAgents when no proxy or agent
is already set (proxy precedence preserved). Wire it into the six speech/OCR call sites
and thread each section's allowedAddresses exemption from pr-01.

Scope is private/internal SSRF only. It does not restrict forwarding a credential to a
public host, which is a separate egress-allowlist concern. Absent an allowedAddresses
entry, private targets now fail closed, matching endpoints, actions, and MCP. Document
the new fields in librechat.example.yaml with the operator warning that allowedAddresses
hostnames are trusted before the private-IP check.

* fix: block literal private-IP hosts and thread STT exemptions through generic uploads

applySSRFSafeAgentIfDirect only attached the DNS-lookup agents, but Node
skips the custom lookup for IP-literal hosts, so a literal private IP such
as http://127.0.0.1 connected unchecked. Reject literal private IPs
synchronously in the helper, reusing the same allowedAddresses exemption
logic as the lookup path.

The generic audio-upload path did not forward the section-level
allowedAddresses to sttRequest, so a private STT endpoint permitted via
speech.stt.allowedAddresses failed with ESSRF outside the speech route.
Thread the exemption through files/audio.ts and widen the STTService type.

* fix: derive effective SSRF port for literal IPs and validate OCR target before opening its stream

The literal-IP precheck normalized an empty URL.port to '', so
allowedAddresses exemptions on a default port (127.0.0.1:80, [::1]:443)
never matched. Derive 80/443 from the scheme when the port is omitted.

uploadDocumentToMistral opened the upload file stream before the SSRF
check, so a blocked or malformed target threw with the descriptor still
open. Run the proxy/SSRF setup before fs.createReadStream.

Reword the librechat.example.yaml allowedAddresses guidance to prefer a
private IP literal over a hostname, and reconcile the speech/OCR SSRF
specs to assert the synchronous literal-IP block instead of driving the
connect-time lookup that Node skips for IP literals.

* fix: block literal private IPs before the proxy return, destroy OCR stream on failure, canonicalize IPv6 exemptions

Move the literal-IP check above the proxy/agent early return so a literal
private IP is rejected even when a proxy is configured; document that a
forward proxy must be SSRF-enforcing.

Wrap the Mistral upload post in try/finally and destroy the file stream,
so an async connect-time block does not leak the descriptor.

Canonicalize IP literals in normalizeAddressCandidate through the same URL
serialization targets use, so IPv4-mapped and expanded IPv6 exemptions match.
2026-08-06 09:06:04 -04:00
Dustin Healy
3f0a1ec8d9
🛡️ fix: Run message-filter PII patterns on a linear-time regex engine (ReDoS) (#14554)
* 🛡️ fix: Run message-filter PII patterns on a linear-time regex engine

The messageFilter.pii middleware compiled admin-configured customPatterns with the native RegExp engine and ran them synchronously against every message on the shared event loop, so a catastrophic-backtracking pattern such as (a+)+$ could stall the entire process (native RegExp takes tens of seconds at roughly 32 characters) and take the instance down for every user.

Compile these patterns with RE2JS, a linear-time RE2 port with no native addon, so catastrophic backtracking is impossible regardless of the pattern rather than something the code tries to detect. Patterns using features RE2 does not support, such as backreferences, fail to compile and are dropped and logged exactly as an invalid pattern already is. The filter only tests for a match, so this is a drop-in engine swap with no behavior change for valid patterns.

* 🛡️ fix: Reject RE2-incompatible messageFilter patterns at config load

The customPatterns regex was validated with native RegExp at config load, but the runtime now compiles it with a linear-time engine (RE2) that does not support backreferences or lookaround. Such a pattern passed validation, then failed to compile and was silently dropped at request time, quietly removing PII protection after upgrade.

Reject backreferences and lookaround during config validation with an explicit message, and document RE2 syntax in the example config instead of "JavaScript-flavor". The runtime engine remains the authoritative boundary and still drops-and-logs anything this load-time check misses.

* 🧹 test: Use direct MessageFilterPiiConfig annotations in the PII specs

The added ReDoS cases satisfy the exported MessageFilterPiiConfig type directly, so the `as unknown as` assertions were unnecessary. Annotate the config objects directly, matching the repo's type-safety guidance.

* 🧹 fix: Reject named backreferences in messageFilter patterns at config load

Extend the config-load check to also reject named backreferences (\k<name>), which are valid JavaScript regex but unsupported by the linear-time runtime engine, so they surface at load rather than being dropped at request time. Together with the existing numeric-backreference and lookaround checks this covers the RE2-incompatible construct set; the runtime engine remains authoritative.

* 🛡️ fix: Preserve Unicode whitespace matching in messageFilter starter patterns

RE2's \s is ASCII-only, so after the engine swap the built-in api-key and Bearer starters no
longer matched a secret separated by non-ASCII whitespace (e.g. a non-breaking space), which
native RegExp did match. Broaden the whitespace classes to [\s\p{Zs}] so those patterns keep
their original coverage, and add a regression test for a non-breaking-space separator.

* 🛡️ fix: Validate messageFilter patterns with the RE2 engine at config load

Replace the syntax blacklist (numeric/named backreferences, lookaround) with authoritative
validation: config load now compiles each custom pattern with the same linear-time engine the
runtime uses, so any RE2-incompatible construct (including control escapes like \cA) is rejected
at load with a clear error instead of being silently dropped at request time.

The validator is swappable and defaults to native RegExp so browser builds add no engine; the
server wires the RE2-backed check at startup via configureMessageFilterRegexValidator in both
entry points.

* 🛡️ fix: Match the full whitespace set in messageFilter starter patterns

RE2's `\s` omits the vertical tab and `\p{Zs}` omits U+2028, U+2029, and
U+FEFF, so a separator built from one of those characters slipped past the
`api-key` and `Bearer` starter patterns and reached the model. Broaden the
starter whitespace class to the full JavaScript whitespace set so those
separators are covered again.

* fix: fail closed when messageFilter.pii compiles to zero patterns

DB and admin config overrides bypass the RE2 schema validation (it only
runs at YAML load), so an override whose only pattern is RE2-incompatible
was dropped at compile time, left zero patterns, and let the request
through. compile() now returns a failClosed flag when a config declared
patterns but every one failed to compile; the middleware returns 400 and
findPiiMatchInMessages returns a distinct misconfigured match that the
OpenAI and Responses controllers surface with an admin-facing message.

* 🛡️ fix: Fail closed when any messageFilter.pii custom pattern drops

compile() previously set failClosed only when every pattern dropped (patterns.length === 0 && dropped > 0). With the default starters present, a single RE2-incompatible custom override incremented dropped but left patterns.length > 0, so the filter silently enforced only the surviving subset and text matching only the dropped rule passed.

failClosed now keys off dropped > 0, so any dropped custom pattern blocks with the misconfigured 400. YAML patterns are RE2-validated at load, so dropped stays 0 for valid configs and only unvalidated DB or admin overrides can trip it. Reframed the two keeps-others-active specs to assert fail-closed and added a default-starters partial-drop regression.

* 🧹 fix: Correct the misconfigured JSDoc and drop redundant casts in the PII specs

The misconfigured flag now means any configured custom pattern failed to compile, not that every pattern failed, so its JSDoc on PiiMatch is updated to match. The partial-drop regressions now use direct MessageFilterPiiConfig annotations instead of as-unknown-as casts, keeping the specs type-checked, consistent with the rest of the suite.
2026-08-05 13:42:18 -04:00
Danny Avila
ccf4301093
🚦 feat: Configurable Circuit Breakers for Runaway Streamed Tool Args (#14613)
* 🚦 feat: Configurable Circuit Breakers for Runaway Streamed Tool Args

* docs: forewarn create_file about the streamed tool-argument limit

The breaker failing a near-limit write should not be the model's first
exposure to the bound. Both create_file variants now state the default
64 KB per-call limit and the incremental pattern (create the first
section, extend with edit_file) in the tool description and the content
parameter description.

* fix: keep skill create_file description under the provider advisory cap

The limit-guidance paragraph pushed the skill-aware description to 1169
chars, past the 1024-char advisory bound where providers may truncate.
The skill variant now carries the guidance only in its content parameter
description, which sits closest to the generated payload and is not at
truncation risk; the shorter code-sandbox variant keeps the full
paragraph.

* 🚦 feat: per-tool streamed-arg limits with a create_file default

Thirty days of production data show create_file is the only tool class
with legitimate near-limit arguments (p99 80.6 KiB; every other tool
p99 under 10 KiB). Rather than loosening the global 64 KiB cap for all
tools, the yaml gains maxToolCallArgBytesByTool (per-tool overrides,
keyed by model-facing tool name, 0 disables that tool's guard) and
LibreChat ships { create_file: 131072 } by default; yaml entries merge
over and can replace it. Pairs with maxToolCallArgBytesByTool support
in the agents SDK and stays inert until the dependency bump.

* test: pass per-tool spec configs as plain Partial literals

The as-TAgentsEndpoint casts fail TS2352 for object-valued fields:
comparability does not grant nested literals the implicit index
signature that plain assignability does, so casts carrying
maxToolCallArgBytesByTool never sufficiently overlap. The mapper
already accepts Partial<TAgentsEndpoint>, so the new cases pass
uncast literals instead.

* chore(deps): bump @librechat/agents to 3.3.12
2026-08-04 23:05:05 -04:00
Danny Avila
cc813f430e
🎯 feat: Tool Intent Label Capability (tool_intents) (#14499)
* 🎯 feat: Tool Intent Label Capability (tool_intents)

Adds the fourth member of the per-tool capability family (defer_loading,
allowed_callers, run_in_background): an admin capability
AgentCapabilities.tool_intents plus a per-tool
tool_options[name].describe_intent flag. Opted-in tools get an optional
intent string injected as the FIRST property of their schema — one
model-authored sentence per call, streamed to the client as the call's
live status label (args already reach the client verbatim, so no new
event plumbing). Native host tools (web_search, create_file/edit_file,
set_memory/delete_memory, ask_user_question) default on while the
capability is enabled; explicit false opts out. SDK-native intent
schemas (@librechat/agents coding suite) are recognized and left alone.

- packages/api/src/agents/intent.ts: structural sibling of
  background.ts — first-key non-mutating injection with registry
  parity (covers deferred/tool_search discovery), eligibility and
  PTC-only skips, arg read/strip helpers, self-spawn strip for defs and
  registry, ephemeral/model-spec synthesis with a tool_options merge so
  the background and intent toggles compose.
- handlers.ts: intent runs BEFORE background injection so the label
  stays the first streamed key when a tool carries both (pinned by
  test); the arg is stripped before invocation unless the tool's own
  schema declares it, on both the foreground and background-dispatch
  paths; PTC target schemas are sanitized like background's.
- Capability plumbing through all four routes (endpoint initialize,
  openai + responses controllers, the exported OpenAI-compatible
  service) plus handoff discovery and added-convo agents, and the
  intentToolNames execution channel via configurable.
- describe_intent on toolOptionsSchema (all three written-out Zod
  annotations), ToolOptions, TEphemeralAgent, TModelSpec (+ zod), and
  data-schemas doc comments (tool_options is Mixed — no migration).
- intent.spec.ts: 28 tests cloned from background.spec.ts structure,
  including the intent+background key-order composition.

* 🧯 fix: Codex Review — Opt-Out Strips SDK-Native Intent, Skip mcp_all Placeholders

- An explicit describe_intent: false now REMOVES an SDK-native intent
  property from the definition and registry entry, so the per-tool
  opt-out actually disables the arg's token cost for tools like
  web_search that carry the schema natively (SDK bodies tolerate its
  absence). Previously the early return left the property in place.
- synthesizeIntentToolOptions skips lazily-expanded mcp_all
  placeholders instead of recording options under names that
  applyIntentLabels' exact-name matching can never match, and documents
  the limitation (parity with synthesizeBackgroundToolOptions).

The P1 about the client not rendering the label is the documented
slicing: the UI streaming-label PR follows once #14391's ToolCallGroup
changes merge — args already reach the client, so that slice is purely
rendering.

* 🧯 fix: Codex Re-Review — Label Marker Guard, Capability Kill Switch, Late Defs, Service Threading

- removeIntentParam is now marker-guarded (the label contract's opening
  instruction discriminates it), so an MCP/action tool's own business
  `intent` parameter is never stripped by an opt-out or the disabled
  path — previously an explicit false could remove a real, possibly
  required argument.
- New sanitizeIntentLabels pass runs AFTER every registration step
  (the skill catalog appends its SDK definition post-injection): with
  tool_intents disabled it strips SDK-native intent labels from all
  definitions and registry entries, making the capability a real kill
  switch over their token cost; with it enabled it enforces explicit
  per-tool opt-outs on late-registered definitions.
- ask_user_question removed from the native default-on set: its graph
  tool is rebuilt in run.ts from its own Zod schema (also the HITL
  card's wire shape), so definition-level injection never reached the
  model. Its intent support lands with the HITL slice, which threads
  the label into the interrupt payload deliberately.
- The exported OpenAI-compatible service now threads intentToolNames
  into the run configurable, so the executor's PTC path can strip
  host-injected intent schemas on that route like the in-repo
  controllers do.

* 🧯 fix: Codex Round 2 — Post-Skill Injection, PTC Native Strip, Service Boundary, Honest Docs

- Intent injection now runs LAST in initializeAgent, after the skill
  catalog — which both appends its own definition and REPLACES upgraded
  ones (skill-aware read_file), clobbering an earlier injection while
  intentToolNames still listed the tool. Injection PREPENDS while
  background APPENDS, so intent stays the first schema property under
  the new ordering (pinned by a reverse-order composition test).
- The PTC target-schema strip is now marker-guarded strip-ALL: SDK-
  native intent labels (which are deliberately never in intentToolNames)
  are removed from sandbox-advertised schemas alongside host-injected
  ones; business intent params survive.
- toolIntentsAvailable on the exported service documents the loader
  boundary: a custom LoadToolsFn returning only structured instances
  bypasses definition/registry injection and sanitize by construction.
- librechat.example.yaml describes tool_intents as backend groundwork
  with UI rendering in an upcoming release rather than promising a live
  label today.

* 📦 chore: bump `@librechat/agents` to v3.3.6

Brings in the SDK half of tool intent labels (danny-avila/agents#347,
#349): intent-first schemas on the coding suite across all three
engines, plus web_search / subagent / skill / tool_search, and the
outcome / outcome_patch result channel.

Activates three host paths that were inert while no SDK tool shipped an
`intent` property — verified against the real 3.3.6 schemas:
- capability OFF now strips SDK-native labels (a real admin kill switch)
- explicit `describe_intent: false` removes them per tool
- host injection stays idempotent against an SDK schema, keeping
  `intent` first and never double-injecting

* 🔬 test: Real-Provider Verification for Tool Intent Labels

Adds the live check the unit tests structurally cannot perform: whether a
real model actually authors the injected arg, places it FIRST, and gives
sibling calls to one tool distinct labels. Reuses the existing
real-provider harness (in-memory Mongo, seeded user, credential
neutralizer) and the existing stdio MCP fixture as a genuine tool, so no
external service is involved.

- e2e/config/librechat.real.yaml: adds the e2e-memory MCP server and the
  tool_intents capability, giving the real model something to call. The
  sibling spec asserts only relative token growth, so the extra schemas
  do not perturb it.
- e2e/playwright.config.real.ts: optional Langfuse passthrough. The
  LANGFUSE_* keys match the credential-neutralizer pattern and were being
  blanked before the server booted; they are preserved explicitly, read
  from the invoking environment only, and never written to the generated
  config.
- e2e/specs/real/tool-intents.spec.ts: two facts stored in one turn, both
  through the same tool, asserting intent is the first key of each call
  and that the two labels differ. Args are read from persistence rather
  than the DOM deliberately — no UI renders the label yet, and
  persistence is what a reloaded conversation and the trace both read.

First run against claude-haiku-4-5 produced 'Recording the location of
the OAuth callback router' and 'Recording the location of the MCP
connection pool configuration' — distinct, first-position, no tool name.

Also updates tool-intent-spec.md: records the 3.3.7 removal of the tense
verb map with the evidence that motivated it, the trimmed description and
the marker's role as an API, and a new mandatory requirement that
client-side label rendering be gated on a server-sent signal rather than
the presence of an intent key (a tool's own business 'intent' parameter
would otherwise render as a status label).

* 📦 chore: bump `@librechat/agents` to v3.3.7 and dedupe the intent contract

Picks up danny-avila/agents#353: the tense verb map is gone (a bare
intent now displays unchanged, with completion carried by UI state), the
model-facing description is trimmed 502 → 289 chars, and both the marker
and the description are exported.

Stops redeclaring the SDK contract here:
- INTENT_LABEL_MARKER is imported instead of duplicated as a string
  literal. Every removal path in this module keys on it, and a local copy
  that drifted from the SDK's would make them all stop recognizing
  SDK-native labels — failing OPEN, with labels left in schemas and
  per-tool opt-outs silently inert.
- INTENT_DESCRIPTION is imported too, so host-injected tools and
  SDK-native tools present the model with one identical instruction.
  Keeping the old local copy would also have meant host-injected tools
  still paying ~126 tokens per schema while SDK tools paid ~72.

Verified live against real Anthropic after the trim: two sibling calls to
one MCP tool produced 'Storing the OAuth callback router file location'
and 'Storing the MCP connection pool configuration file location' —
first-position and distinct, so the shorter description holds compliance.
2026-07-29 15:40:52 -04:00
Danny Avila
7b6900d556
🏷️ feat: Activity Groups With Fast-Model Headers (#14391)
*  feat: Activity Groups with Fast-Model Labels

Groups each contiguous block of reasoning + tool calls into a collapsible
unit headed by a fast-model label (claude.ai-style hierarchy), off the
critical path: a PostToolBatch hook claims a live content slot at the
batch boundary (steering index-offset pattern), renders a deterministic
counts phrase instantly, and swaps in the generated label ~1s later while
the next model call streams. Labels are UI-only — stripped before the SDK
formatter and skipped in the legacy formatter — and reach live clients
via a dedicated on_activity_label SSE event (live/replay/pending paths).

Grouping preserves legacy rendering byte-for-byte when no label part is
present. Generation bridges to Run.generateActivityLabel() when the SDK
ships it (session-grouped Langfuse tracing); falls back to a direct call
today. Env-gated: ACTIVITY_LABELS_POC=true, ACTIVITY_LABEL_MODEL.

* 🧷 fix: Address Codex and Copilot Review Findings for Activity Labels

- Settle in-flight label fills (bounded 3s) before finalization on both
  the main and resume paths, so a label resolving during the final batch
  still reaches the durable log and saved message.
- Overlay on_activity_label chunks in RedisJobStore content
  reconstruction (splice path last-wins per index; replay path
  chronological overwrite), matching steer handling.
- Wire activity labels into the HITL resume createRun so post-resume
  batches keep claiming slots.
- Guard against out-of-order publishes: fill() awaits the claim emit
  before emitting the resolved label, and the client applier ignores a
  stale pending placeholder once a resolved label is present.
- Stamp the batch's groupId onto label parts so parallel-column runs
  place them inside their group instead of filtering them out.
- Localize the counts fallback phrase (10 keys, singular/plural) through
  useLocalize across chat rendering and exports.
- Type the hook with Providers/ClientOptions instead of stringly types;
  drop the unknown cast in the spec; add a dedicated rAF retry ref for
  label events with effect cleanup.

* 🛡️ fix: Address Independent Review — Abort, Usage, Lane Context, Redis Test

- Propagate the run abort signal into label generation (both wiring call
  sites; runtime combines host + dispatch signals with the timeout) so a
  user abort cancels in-flight label calls instead of paying to timeout.
- Record label-call usage like titles: the SDK bridge aggregates via
  chainOptions callbacks, the fallback path via a per-generation callback
  factory; both feed recordCollectedUsage under context 'activity-label'.
- Scope block-context capture: reasoning collection stops at the previous
  block's label part and filters by executingAgentId, so consecutive or
  parallel batches can no longer bleed another block's thinking into the
  payload; intent text still scans past labels (persists across batches).
- Forward the effective charLimit to the SDK call so host and SDK prompts
  agree (SDK default aligned to 600 in agents#327).
- Add a Redis integration test proving last-write-wins reconstruction of
  on_activity_label chunks per claimed index.
- Rebased onto main: only the two activity commits replay (the nine
  steering commits belonged to the old base branch), zero conflicts,
  steering suites green.

* 📐 refactor: Move Activity-Label Wiring to TypeScript, Address Codex Round 2

- [P1] Slot claiming, lane stamping, emit ordering, context capture, and
  settle tracking now live in packages/api (createActivityLabelWiring +
  captureActivityBlockContext); client.js is a thin closure wrapper.
- Register the activity-label hook BEFORE the steer drain so a steer
  draining at the same batch boundary cannot flush the tool block and
  orphan the label outside its group.
- Resolve request-based header placeholders in resolveActivityLabelLLM
  (titleConvo parity) so metadata-keyed proxies work on label calls.
- Trim labels centrally before filling so whitespace-only output from
  either generation path keeps the deterministic counts fallback.

* 🧭 fix: Codex Round 3 — Capture Order, Shared Strip, Token Estimator, Hide Filter

- Capture block context BEFORE pushing the label part: the scan stops at
  ACTIVITY_LABEL parts, so post-push capture hit the just-inserted label
  and silently collected no reasoning excerpts (regression test added).
- Share stripActivityLabelParts from packages/api and apply it in the
  Responses and OpenAI-compatible controllers, closing the replay leak
  for entry points still running SDKs without the formatter skip.
- Exclude activity_label parts from the fallback response-token estimator
  (UI-only parts must not inflate no-usage provider billing).
- Keep label parts explicitly under hide_sequential_outputs — they
  summarize exactly the outputs that mode hides.

* 🔁 fix: Codex Round 4 — Resume Gap, Delta Flush, Agent-Scoped Intent, Token Counter

- Synthesize on_activity_label events for labels claimed or filled in the
  snapshot→subscribe window (the publish is fire-and-forget, so Redis-mode
  reconnects missed them). Feature-gated so the default path adds no
  content re-read; the client applier already ignores duplicates.
- Flush queued deltas before applying a label part, matching the pending-
  action and steer appliers — without it the handler read a stale message
  cache and syncStepMessage pushed a pre-delta copy back.
- Skip another agent's tail text when resolving intent, so parallel runs
  cannot seed a label prompt with a sibling agent's narration.
- Exclude activity_label parts from countFormattedMessageTokens (the
  agent-path counter), not just the legacy BaseClient one.

* 🏗️ refactor: Codex Round 5 — Extract Label Host Logic, Report Usage, Icon Strip

- Move provider/model resolution, usage-metadata mapping, and the settle
  loop into packages/api (activityLabels/host.ts); client.js keeps only
  thin delegations, per the repo's TypeScript-implementation convention.
- Fold label usage into the response rollup with an 'activity-label' tag
  (subagent precedent) so metadata.usage and the live cost gauge account
  for it; tagged, so it stays out of PRIMARY usage/context pairing.
- Narrow tool metadata once in ToolCallGroup so THINK parts in a labeled
  block no longer render phantom generic icons in the stacked strip.
- Import the activity-label helpers by deep path in GenerationJobManager:
  the package barrel now reaches provider-config/cache modules that
  import back into the stream layer, and the cycle broke suite loading.

Declined: resetting steerOffsetState before HITL resume — resume builds a
FRESH AgentClient via initializeClient (initialize.js:978), so the offset
is already zero; the seed wrapper alone accounts for pre-pause parts.

* 🚦 fix: Codex Round 6 — Stream Label Usage, Close Late Fills

- Emit an on_token_usage chunk for label calls (sink push alone left the
  live session gauge blind); retained in pendingSubagentEmits so job
  cleanup cannot race the persist, tagged 'activity-label' as before.
- Close the label scope when settle times out: the wiring gates fill() on
  isClosed and the client fires a label-scoped AbortController, so a
  straggling generation can neither mutate a saved response nor emit into
  a job whose runtime is gone. The controller also chains to the run
  signal, so a user abort still cancels label work.

* 🩹 fix: Repair CI — Package Typecheck and Module Mocks

Local runs covered the client tsconfig and jest, but never packages/api's
own tsconfig, so nine type errors in the extracted host module shipped.

- Type host.ts against the real contracts: ServerRequest, EndpointDbMethods,
  AppConfig from @librechat/data-schemas, IUser for createSafeUser, and a
  MaybeAzureConfig view for the azure instance-name probe and configuration.
- Widen resolveConfigHeaders' llmConfig to Partial<RunLLMConfig>: it only
  reads the three provider header carriers, so auxiliary generations with a
  bare ClientOptions can resolve headers without assembling a run config.
  Type-only widening; every existing caller still satisfies it.
- Add stripActivityLabelParts to the @librechat/api mock in the OpenAI and
  Responses controller specs — those mocks enumerate exports, so a new
  import read as undefined and threw before the assertions ran.
- Use the real activity-label helpers in the ToolCallGroup spec's ~/utils
  mock; stubbing them out would hide the header logic under test.

* ⚙️ feat: Configure Activity Labels via librechat.yaml, Drop Env Vars

Replaces the ACTIVITY_LABELS_POC / ACTIVITY_LABEL_MODEL env gate with
per-endpoint settings, following the title options convention rather than
a top-level block — each endpoint picks its own cheap label model.

- Add activity, activityModel, activityEndpoint, activityPrompt,
  activityMaxPerRun, and activityCharLimit to the endpoint schema, and to
  the endpoints.all pick list (enumerated, so 'all:' would otherwise drop
  them silently).
- resolveActivityConfig reads them with title-style precedence:
  endpoints.all > named endpoint > custom endpoint config.
- Model precedence is now activityModel > titleModel > the agent's model.
  activityEndpoint runs labels on another endpoint's credentials, with
  titleConvo's fallback-on-unknown-name behavior.
- Thread activityPrompt/MaxPerRun/CharLimit through the wiring into the
  hook and the SDK bridge; they were hardcoded defaults.
- The resume gap-repair gate keyed on the env var; it now keys on the
  snapshot actually containing label parts, so deployments without the
  feature still perform no extra content read.
- Document the fields in librechat.example.yaml; add host.spec.ts
  covering precedence, custom-endpoint fallback, and opt-out.

* 📝 refactor: Rename Enable Flag to activityLabel, Document Schema Inheritance

- Rename the boolean from `activity` to `activityLabel`, matching the
  titleConvo/titleModel shape: a verb-object toggle whose prefix matches
  its modifiers (activityModel, activityPrompt, ...). `activity: true`
  alone read ambiguously — it could mean tracking or logging activity.
- Document the two endpoint-schema inheritance paths, which behave
  oppositely and are ~900 lines apart:
  * `endpoints.all` omits from baseEndpointSchema, so new options are
    inherited automatically — nothing to maintain.
  * `azureEndpointSchema` enumerates via .pick(), so a new option is
    silently unavailable on Azure endpoints until listed there.
  The activity block now carries a pointer to the Azure caveat.

* 🔍 fix: Address Codex Findings on the Config Rework

- Pass the matched custom-endpoint config into the label gate. Custom
  endpoints live in the `endpoints.custom` ARRAY, so without it every
  custom endpoint resolved as disabled — including the example this PR
  added to librechat.example.yaml.
- Give label usage a unique `runId:seq`. Label usage is billed but never
  appended to `collectedUsage`, so its length was static: every label
  event reused the last primary usage's pair and collided with itself,
  and the client dedupes on exactly that.
- Attach `cost` to label usage when `interface.contextCost` is on;
  aggregateEmittedUsage treats coverage as all-or-nothing, so an event
  without it suppressed the whole response's cost.
- Honor `activityPrompt` on the direct fallback path, not just the SDK
  bridge — it previously always used the built-in instruction.
- Seed the per-response label cap from labels already on the response so
  a HITL resume cannot mint a fresh quota after every approval.
- Reconcile label gaps on resume via a durable per-job `activityLabels`
  flag instead of probing the snapshot: the FIRST label of a run can be
  claimed inside the snapshot->subscribe window, which the old signal
  missed. The flag is read from a job record already fetched there, so
  runs without the feature still add no content read.
- Auto-collapse labeled single-tool groups; one-call batches are common
  in agent runs and rendering them expanded defeats the grouping.

* 🎯 fix: Correct Label Usage Seq, Cross-Endpoint Pricing, Close Scopes

- Give label usage a NEGATIVE seq namespace. The previous fix was wrong:
  seq is a position in `collectedUsage` (push, then emit with the new
  length), so sink-length + array-length still lands on a real position —
  primary emits 1, the label computes 2, the next primary also emits 2.
  Labels have no position at all (billed separately, never appended), so
  they now occupy a namespace positional sequences cannot reach. The
  client key is a string used for Set membership, so the sign is inert.
- Price cross-endpoint labels with the LABEL endpoint's token config:
  resolveActivityLabelModel now returns the resolved endpointTokenConfig,
  and both the streamed cost and recordCollectedUsage use it instead of
  the agent endpoint's rates.
- Make close state per-wiring rather than per-client. A HITL resume
  rebuilds the wiring, and resetting a shared flag re-opened closures from
  the pre-pause segment whose provider call ignored the abort; settle now
  closes every retained scope, past generations included.

* 🎯 fix: Make the Activity Header Say Something the Cards Cannot

The header read "ran 1 command" next to a card already labeled "Code" —
it restated the UI beneath it instead of adding to it. Two causes, both
about content rather than timing:

- A deterministic tool-type tally was the primary display and also fed
  the prompt, so the best case was a tally and the worst case was a
  tally dressed as prose. Removed from the metadata, the prompt, the
  part type, and the client.
- The instruction only ever reached the fallback path. The wiring
  passed a prompt only when  was configured, so the
  preferred SDK path silently used the published package default. The
  wiring now always supplies one and the hook forwards it on both
  paths.

The register is rewritten around what the cards cannot show:
past-tense git-commit-subject, leading with the distinctive noun,
outcome over attempt, tool names and counts and arguments explicitly
forbidden. The batch entries are labeled as reference material so the
model stops transcribing them.

Claiming a slot no longer emits. The slot still reserves its index so
streamed parts never collide, but with nothing to say there is nothing
to render: until a description exists the block looks exactly as it
does without the feature.

* 🧹 fix: Drop the Localize Hook Left Unused by the Counts Removal

*  test: Add Activity-Label e2e Coverage with a Recording Label Server

Activity labels are the one model call a mock run does not already fake:
fake-model.js swaps the GRAPH model via overrideTestModel, while
run.generateActivityLabel() calls the endpoint resolved client options
over HTTP. The custom endpoints already point baseURL at 127.0.0.1:8889,
so serving that port exercises the real path with no production seam.

fake-label-server.js answers it in both JSON and SSE form, records each
prompt, and can inject blank/error responses. Recording is what lets the
spec assert the CONTRACT rather than the rendering: that this repo
register and the tool OUTPUTS actually reach the model. That is the bug
class that produced unusable labels before, and rendered text looks
identical whether or not the instruction arrived.

Labels get a dedicated endpoint (Mock Provider E). A labeled block
auto-collapses even at one tool call, which hides the tool cards other
specs assert on -- enabling this on a shared endpoint broke
steering.spec.ts. Provider D is the unlabeled control.

Request-count assertions are scoped to a per-test token: a 5xx label
response is retried by the provider client, and a retry can land after
the next test has reset the server.

* 🩹 fix: Address Review Findings on Activity-Label Indexing and Pricing

Replay index (P1). Reserving the slot only in server memory left no event
for it, so a cross-instance replay rebuilt content as [tool, hole, later],
compacted the hole away, and the fill for the reserved index landed on the
following part and overwrote it. The claim now publishes the empty,
pending part so the index is real for every consumer, and fill publishes
even when generation returned nothing so the client cannot stay pending.

It stays invisible: an empty label still DELIMITS its batch in
groupSequentialToolCalls but is not attached as the header, so grouping
does not re-shuffle when the text lands and the block renders exactly as
it does with the feature off.

Edited-response index (P1). Edit-and-resubmit replays the kept prefix and
the server indexes only new content, so run steps offset by that prefix.
Labels are claimed in the same space and now take the identical shift;
without it a label could land inside the prefix and overwrite it.

Redis flag. deserializeJob never read activityLabels back, so every Redis
reload left it undefined and resume skipped label gap reconciliation.

Executing agent. RunActivityLabelOptions.agentId selects the executing
agent tracing metadata AND its tool-output redaction policy; omitting it
let a handoff be redacted under the default agent configuration.

Label pricing. An undefined endpointTokenConfig is meaningful for a
built-in label endpoint (priced from the shared table), so the nullish
fallback billed those labels at a custom primary rates. Inherit only when
the label runs on the agent own endpoint.

HITL usage sequence. runId is the response message id and the counter was
instance-local, so a resume restarted at -1 and the client runId:seq
deduper discarded the post-approval label usage. Seeded past the labels
already on the response.

Also distinguishes "cannot serve" (undefined) from "no label" (null) in
the SDK bridge, so a missing run falls back to the direct call instead of
filling the slot empty. Version gating already happens at wiring time via
the sdkCapable prototype probe.

* 🩹 fix: Keep Unfilled Activity Labels Invisible and Unmask Endpoint Settings

Follow-up review round. Publishing the reservation on every batch made two
latent rendering paths reachable on every run, and both are fixed here.

Empty labels no longer change grouping. The previous pass still formed a
tool-group for a textless label, which wrapped even a single tool call and
pulled THINK parts inside it — and since a reservation is published the
moment each batch ends, that applied during every generation and
permanently after a blank or failed fill. An empty label now flushes the
legacy way instead: it still delimits its batch, but the block re-splits
exactly as it renders with the feature off.

Parallel lanes no longer show a blank line. Lanes render raw parts, so an
unfilled label had nothing to draw; empty ones are dropped. Making labels
act as collapsible headers inside lanes is still a separate gap.

Edited responses no longer offset on resume. The sync replaces
initialResponse.content with the server's aggregatedContent, which already
contains the kept prefix AND everything generated since — so its length is
not the prefix length, and indices reconciled from that snapshot are
already absolute. Offsetting again pushed the label past its slot onto a
later part. The shift now applies only to a fresh edited submission.

Activity settings resolve field by field. Selecting one config object
whole meant any endpoints.all block — even one carrying nothing but
headers — shadowed the named or custom endpoint and silently disabled
activity labels everywhere. Global still wins per field.

Adds groupToolCalls coverage for the invisible-while-empty contract, which
is the part most likely to regress: it is normal state on every run, not
an edge case.

* 🔒 fix: Scope Detached Label Writes to Their Generation Epoch

Epoch scoping (P1). Label generation is detached and can outlive the
generation that started it. emitChunk only proves that SOME runtime is
current, not that the caller belongs to it, so an aborted generation's
fill(null) -- and its usage event -- could be attributed to whichever
generation replaced it, landing an index from the abandoned response on
top of the new one. Because an empty label renders nothing, that
overwrote content silently. emitChunk now takes an optional jobCreatedAt
and drops the event when the runtime epoch differs, mirroring the
existing setGraph/setContentParts convention, and both label emitters
pass it.

An abort now CLOSES the label scope instead of only cancelling the call:
the rejected generation still runs its catch and calls fill(null), which
would otherwise emit into a stream the next generation may already own.

Edited-response indexing (P1). The previous pass skipped the prefix
offset on resume, which was the wrong half of the problem: a sync
replaces initialResponse.content with the server's aggregatedContent,
which is completion-local, so after a reconnect its length is not the
kept-prefix length and the offset is wrong -- but it is wrong for run
steps in exactly the same way. Tool cards and the label that heads them
must share one index space; a label shifting differently from its tools
lands on another part. The label path now uses the identical expression
as useStepHandler, with no resume special-case. Correcting the
post-resume prefix length belongs in calculateContentIndex, where it
fixes both at once.

titleModel masking. The activity settings were made per-field last pass,
but the titleModel fallback a few lines below still selected an entire
config object, so a partial endpoints.all (for example one carrying only
headers) hid a named endpoint's titleModel and quietly fell the label
back to the main agent model. Both now read through one shared per-field
helper.

Resume reconciliation no longer depends solely on markActivityLabels,
which is best-effort yet had come to gate correctness: a lost flag write
silently dropped a label. The snapshot is consulted as a fallback.

The exported host type for generateLabel now admits undefined, which is
the documented "cannot serve, fall back to the direct call" signal the
hook keys on -- distinct from null, meaning it ran and produced nothing.

* 🧷 fix: Keep Group Identity Stable and Memoize Label Endpoint Resolution

Group remount. Tool-group identity was keyed on the first part in the
block. An activity label absorbs the block's leading THINK part the moment
its text lands, so the key flipped from tool:<id> to fallback:<scope>:<idx>
mid-run, remounting the group and discarding whatever the user had
expanded. The key now scans for the first tool call, which does not move
when the block re-forms.

Label endpoint resolution is memoized per response. It reads provider
config and can hit the database for user keys, yet nothing it depends on
changes between batches of one run — and it ran twice per batch, once for
generation and once for usage accounting. The promise is cached rather
than the value so concurrent batches share a single in-flight resolution,
and a rejection is evicted so one transient credential failure cannot
disable labels for the rest of the response.

* 🎯 fix: Offset Edited Resubmissions by a Prefix Length That Survives Resume

The server indexes only NEW content for an edited resubmission, so the
client offsets incoming indices by the prefix it retained. That prefix was
read as initialResponse.content.length, which is correct only until a
resume: the sync replaces that array with the server's completion-local
snapshot, whose length is unrelated to the prefix. After a reconnect every
offset was therefore wrong -- run steps and activity labels alike -- and
could write over content the edit kept. For a label the symptom is worse
than a bad position: the fill misses its own reservation, so the pending
placeholder is never resolved.

The prefix length is now captured when the submission is built, while
initialResponse.content still IS the retained prefix, and carried on the
submission as editPrefixLength. calculateContentIndex takes that length
instead of deriving it from an array that a resume may have replaced, so
run steps and labels share one index space by construction rather than by
both happening to read the same field.

Note the prefix is the FULL original content with the edited part
substituted in place (useChatFunctions clones latestMessage.content and
mutates one entry) -- it is not a slice, so the length cannot be inferred
from editedContent.index.

Group identity no longer changes when a label fills. Tool-group keys were
derived from the first part in the block; an activity label absorbs the
leading THINK part when its text lands, flipping the key mid-run and
remounting the group, which discarded the user's expansion state. The key
now scans for the first tool call, which does not move.

Label endpoint resolution is memoized per response. It reads provider
config and can hit the database for user keys, yet ran twice per batch --
once to generate, once for usage accounting -- while nothing it depends on
changes within a run. The promise is cached so concurrent batches share one
in-flight resolution, and rejections are evicted so a transient credential
failure cannot disable labels for the rest of the response.

The resume gap passes for steers and activity labels now share a single
lazy content read instead of each issuing its own. The label pass stays
gated on the run flag with a snapshot fallback: reconciling
unconditionally would also close the residual first-label window, but it
would bill a read to every resume of every run, including deployments with
the feature off -- which the steer pass deliberately avoids. That residual
requires a lost flag write, which shares fate with the content writes the
labels live in.

* 💵 fix: Bill Cross-Endpoint Labels at Their Own Rates

recordCollectedUsage never accepted an endpointTokenConfig, so the value
the activity-label caller passed was dropped and the balance transaction
was written at the primary agent's rates. Only the UI cost honored the
label endpoint, so a custom primary pointing activityEndpoint at another
endpoint showed one price and charged another. The parameter is now
accepted, and an explicit config wins outright over per-agent resolution:
that map is keyed by AGENT, so it cannot describe usage that ran on a
different endpoint.

Group identity is stable for id-less tool calls too. The previous pass
anchored the key to the first tool call ID; where a supported tool call
carries no id the fallback still used the block's first part index, which
shifts when a filled label absorbs the leading THINK part. The fallback now
anchors to the first TOOL entry's index, so only a block containing no
tool call at all keys off parts[0].

markActivityLabels is retried rather than fire-and-forget. It gates resume
gap reconciliation and is a SEPARATE write from the durable label append,
so a single lost write silently drops a label the content itself recorded.
The earlier "shared fate with content writes" reasoning was wrong. One
retry at run setup costs nothing and removes the only realistic way the
gate goes stale, without billing a content read to every resume.

* 🧮 fix: Stop Offsetting Once SYNC Drops the Edited Prefix

The edit offset was applied unconditionally, but whether it is correct
depends on which branch SYNC took. SYNC either preserves the content
already loaded for the response -- which still contains the retained
prefix, so the offset is required -- or replaces it with the server's
aggregatedContent, which is completion-local and indexed from zero, after
which any offset writes past the end of a now shorter array.

That is why the two previous attempts each fixed half of it: skipping the
offset on resume was right for the replace branch, applying it
unconditionally was right for the preserve branch, and neither holds on its
own. The offset now tracks the actual state of the rendered content.

For an activity label the replace branch was worse than a bad position:
the fill landed past its own reservation, so the pending placeholder was
never resolved and the block kept its generic header for the rest of the
run.

Applied to run steps as well, not just labels. useStepHandler reads the
prefix from the same submission and had the same unconditional offset, so
after a mid-session resume of an edited response tool cards were misplaced
too. Normalizing at the dispatch boundary keeps both in ONE index space by
construction: a label that shifted differently from the tools it heads
would land on another part.

Note the reload path was already coherent -- useResumeOnLoad rebuilds the
submission without editedContent or editPrefixLength, giving no offset
against server-supplied content -- so only the mid-session SYNC path was
inconsistent.

* 🧾 fix: Keep Label Accounting Out of the Primary Usage Slot

Label usage no longer owns getStreamUsage(). recordCollectedUsage assigned
its result to this.usage unconditionally, so when the primary provider
reported no usage metadata but the label provider did, BaseClient took the
label's output tokens as the assistant response's authoritative count. The
later primary call returns early on an empty collectedUsage and never
replaced it, so the wrong value stood, the text-based fallback was skipped,
and the real generation went unbilled. Secondary usage is still billed but
no longer writes that slot.

Cross-endpoint pricing keys off an explicit discriminator rather than the
presence of a value. A built-in label endpoint prices from the shared
table, so an undefined endpointTokenConfig is its MEANINGFUL value --
reading that as "no override" fell back to the primary's custom rates and
restored the exact mismatch the previous pass set out to fix. The caller
already knows whether the label ran elsewhere and now says so.

markActivityLabels rejects on failure instead of swallowing it. The flag
gates resume gap reconciliation and the caller retries it, but the internal
catch resolved successfully and made that retry unreachable -- so the two
changes cancelled out and a transient write failure still left the flag
absent.

Late label accounting is suppressed with the same gate as the late fill. A
straggler that outlived the settle timeout still ran its finally block, so
it charged the balance and appended to usageEmitSink after the response had
passed its usage flush and metadata snapshot: a cost the user pays but is
never shown.

The cleared-prefix state is scoped to one generation. It was set on a
resume SYNC that replaced the response and then never reset, so a later
edited resubmission in the same mounted hook dispatched run steps and
labels with no offset against content that still held its retained prefix.
Reconnects pass isResume and keep the state; a new generation clears it.

* 🔑 fix: Key Prefix State to the Stream and Honor current_model for Labels

The cleared-prefix reset keyed on isResume, which skips exactly the case it
was added for: a submission whose POST succeeded server-side but lost its
response is retried, comes back resumed: true, and subscribes in resume
mode even though it is a NEW generation. A previous generation's cleared
state then survived into it, and incoming run steps and labels applied no
offset against content that still held its retained prefix. The state is
now keyed to the stream id, which changes with the generation and stays put
across reconnects of one.

activityModel now honors current_model. The options are documented as
title-shaped and the titleModel fallback already excludes the sentinel, but
the higher-precedence activity override passed the literal string through to
getOptions and the provider, so an endpoint following that convention failed
every label instead of using the agent model.

* 🎯 fix: Key Prefix State to the Generation and Resolve the Run Model

The cleared-prefix state was keyed to the stream id, which never changes
within a conversation: request.js sets streamId = conversationId, so once a
reconnect cleared the state every later edited resubmission in that
conversation dispatched run steps and labels with no offset and could
overwrite the prefix it retained. It is now keyed to the response message
id, the only per-generation identity available here -- minted per
submission and carried through a resume unchanged.

That is the third identity tried for this state. isResume missed the
deduplicated-retry path (a lost response returns resumed: true for a new
generation); the stream id is conversation-scoped. The response id is the
boundary that actually matches a generation.

current_model labels now resolve the model the run is really using.
initializeAgent merges the request's endpointOption override into
model_parameters and the run gives it precedence, so preferring the saved
agent.model could send labels to a different, potentially unavailable or
more expensive model than the conversation is on.

* 🆔 fix: Key Prefix State to the Submission and Keep the Origin Title Model

Editing an assistant response reuses that response's messageId as
editedMessageId, and useChatFunctions carries it onto
initialResponse.messageId -- so re-editing the same response produced two
generations with the same key and the cleared-prefix state survived between
them, leaving run steps and labels with no offset against content the edit
retained. Keyed now to clientRequestId, the per-submission uuid, which is
minted fresh per edit attempt and forwarded unchanged on retries.

That is the fourth key this state has had, and each earlier one failed at a
real boundary: isResume missed the deduplicated-retry path, the stream id is
the conversation id, and the response message id is reused across edits of
one response. clientRequestId is the identity that actually means "this
submission".

The titleModel fallback is read from the ORIGINATING endpoint again, matching
how titleConvo captures its config before switching credentials. Reading it
after an activityEndpoint switch meant an OpenAI endpoint configured with
titleModel claude-haiku and activityEndpoint anthropic fell through to the
OpenAI run model and sent that name to Anthropic, failing every label. The
destination endpoint supplies credentials, not the model choice.

* 🧷 fix: Close the Remaining Edit, Epoch, and Scope Gaps for Labels

SYNC clears the edit prefix on the new-row branch too. When a resumed
edited submission cannot match an existing assistant row, that branch
builds the response straight from the server's completion-local
aggregatedContent, so it holds no retained prefix -- but the reset lived
only in the matched branch, leaving later steps and labels adding an
offset to indices that were already absolute.

Label usage is keyed per GENERATION. Editing one assistant response reuses
its responseMessageId while each fresh generation restarts
activityLabelUsageSeq, so a second edit re-emitted the same runId:seq and
the client discarded the newer usage while its balance transaction was
still written. The key now carries jobCreatedAt, the run's own epoch:
stable across reconnects and HITL resumes, distinct between generations.

The scope is revalidated at commit time. Checking once before the await let
a scope that closed mid-flight still charge the balance after finalization,
while the matching fill saw the closed scope and dropped the label --
billed but never surfaced, the exact outcome the guard exists to prevent.

The titleModel fallback no longer reaches the destination endpoint. With
activityEndpoint set and no titleModel on the originating endpoint, it
picked up the destination's, so changing only the credential target
silently changed the model and its cost. Precedence is activityModel, then
the originating endpoint's titleModel, then the run model; the destination
supplies credentials only.

* ✂️ refactor: Confine the Edit-Prefix Offset to Activity Labels

useStepHandler is now byte-identical to dev again. The resume-aware prefix
offset was applied there too, which was more correct in principle -- the
post-resume prefix length is genuinely wrong for run steps as well -- but it
changed index math that EVERY run step flows through, for every user,
including everyone who never enables activityLabel.

That shared correction needed five revisions in two days (isResume, the
stream id, the response message id, clientRequestId, and the SYNC new-row
branch), each passing the full suite and each failing at a boundary only
review found. Carrying it inside an opt-in feature put every user behind
logic with that track record. It belongs in its own change, with tests that
construct the edit-plus-resume states none of the current suites reach.

The offset now applies only where the label handler places its part, so
this PR cannot alter rendering for anyone with the feature off. The known
consequence is recorded in the description: with activity labels ENABLED,
an edited response that reconnects mid-generation can place its label and
its tool cards in different index spaces. That is a bug for opt-in users
rather than a regression for everyone, and it disappears once the shared
fix lands.

submission.editPrefixLength stays: the label path still needs a prefix
length that survives a SYNC replacing initialResponse.content.

* 🧾 fix: Commit Labels Before Billing and Keep Blank Slots Invisible

Round-nine review (all P2, feature-scoped):

- Billing ordering (client.js:409, runtime.ts): usage accounting ran
  BEFORE the slot commit on both generation paths, so the settlement
  deadline could expire during the balance write — charged, then the
  fill dropped as out-of-scope: billed, never shown. `slot.fill` now
  resolves a commit flag, generators register their accounting via
  `deferUsage`, and the hook runs it only after a committed fill.
- Scope gates (client.js:757): the direct-fallback `collect` omitted
  `scopeOpen`; both paths now gate on the OWNING wiring's scope, so a
  pre-pause straggler cannot bill because the resumed generation's
  scope is still open.
- Blank-label grouping (groupToolCalls.ts:81): a blank slot forced a
  flush, splitting adjacent single-call batches into standalone cards
  where the feature-off path merges them. Blank labels now only mark
  the claim boundary — structurally invisible, while a later filled
  label still cannot claim an earlier batch.
- Stale fill indices (wiring.ts:301): the skill-card unshift and the
  hide-sequential filter reshape contentParts before the finalization
  settle, so an in-flight fill emitted its claim-time index against a
  shifted array. Both completion paths now settle label fills before
  any post-run content reshaping (the finally settle stays as the
  error-path net; the second call sees an empty pending list).
- Bounded serialization (runtime.ts:238): `JSON.stringify` fully
  materialized unbounded tool results to keep 200/600 chars per entry.
  A budget-bounded serializer stops at the limit (which also bounds
  cyclic values) and preserves the exact truncate-with-ellipsis output.

Tests: fill/bill ordering + suppression on dropped fills (runtime.spec),
blank-slot merging and claim boundaries (groupToolCalls.test), bounded
serialization equivalence and giant-output truncation (runtime.spec).

* 🧮 fix: Keep Deferred Label Billing Inside the Settle Window

Self-review follow-up to the billing reorder: deferring usage until
after the commit moved it PAST the fill's resolution, so a settle keyed
on fills alone could let finalization flush the usage sink and snapshot
metadata while the label's billing was still in flight — the usage row
would silently miss the message rollup even on the happy path.

The hook now reports its whole detached task (generate → fill →
deferred usage) via a `trackTask` option, wired to the same settle
tracker as the fills, so finalization waits for billing exactly as it
did when accounting preceded the fill. The task never rejects. Pinned
in runtime.spec: the tracked task resolves only after usage collection.

* 🧰 fix: Harden Label Resolution, Output Bounds, and Cache Billing

Round-ten review (all P2, feature-scoped); the sixth finding is the
documented edited+reconnect index-space limitation, answered on-thread
as deliberately out of scope for this PR.

- Rejected-LLM memoization (runtime.ts): the hook cached a rejected
  `resolveLLM()` promise permanently, failing every later batch and
  silently defeating the host resolver's own rejected-cache eviction.
  The memo now evicts on rejection so the next batch retries.
- `current_model` precedence (host.ts): an explicit
  `activityModel: current_model` resolved to `undefined` and then lost
  to a configured `titleModel`. The sentinel now resolves straight to
  the run model; the title fallback applies only when `activityModel`
  is absent.
- Output bounds (runtime.ts): label text was persisted verbatim; a
  model ignoring the 4–9-word instruction (or steered by injection in
  untrusted tool output) could emit thousands of tokens duplicated
  through SSE, the chunk log, persistence, and the UI.
  `normalizeLabelOutput` keeps the first non-empty line, collapses
  whitespace, and hard-caps at 200 chars on both generation paths.
- Cache-token billing (host.ts, client.js): the usage mapper dropped
  cache fields, vanishing Anthropic cache tokens from billing and
  charging OpenAI cache reads at the full input rate. The mapper now
  normalizes Anthropic/OpenAI/LangChain cache shapes into
  `input_token_details`, and the emit + cost path carries them with the
  label endpoint's `provider` (additive-provider adjustment).
- Usage-type union (runs.ts): `TTokenUsageEvent.usage_type` now
  includes the emitted `activity-label` literal; the lone consumer
  keys on `usage_type != null`, so this is type-level completion.

Tests: sentinel/title/explicit model precedence and all three cache
shapes (host.spec), transient-resolution retry and output normalization
with truncation (runtime.spec), the new usage literal (runs.spec).

* 🪗 fix: Let Settled Labels Collapse Void Tools and Keep the Tail Cursor

Round-eleven review (all P2, client-side). Two fixed; the other two
findings restate documented Known limitations (edited+reconnect run-step
index space; parallel-lane collapsible headers), answered on-thread.

- Void-tool auto-collapse (ToolCallGroup.tsx): `allCompleted` keyed
  solely on output truthiness, so a tool that legitimately returns an
  empty string kept its labeled group expanded forever. A settled,
  filled label is itself a completion proof — the PostToolBatch claim
  only happens after every output in the batch returned — so it now
  satisfies `allCompleted`; pending labels keep the group live.
- Trailing-reservation cursor (ContentParts.tsx): a blank label
  reservation at the content tail renders nothing but still counted as
  the last part, stripping the streaming cursor and last-item
  affordances from the last VISIBLE part until the next delta.
  `lastContentIdx` now walks back past empty label slots.

Tests: labeled void-tool group auto-collapses, pending-label group
stays expanded (ToolCallGroup.test).

* 💳 fix: Price Label Cache Correctly, Honor endpoints.agents, Cancel Every Retry

Round-twelve review: four fixed here; the remaining P1 (move the
client.js bridge into packages/api) is an architecture call answered
on-thread for the maintainer.

- Provider on billed entries (client.js, P1): round ten added cache
  details to label usage entries but not `provider`, and `splitUsage`
  treats an unknown provider as additive — re-adding cache_read and
  cache_creation on top of an input count that already contains them,
  double-charging Anthropic/OpenAI cached label calls while the
  streamed cost (which carried the provider) disagreed. Every mapped
  entry now carries the label endpoint's provider.
- endpoints.agents honored (host.ts, client.js): `initializeAgent`
  rewrites `agent.endpoint` to the backing provider, so activity
  settings under the PUBLIC `agents` endpoint — valid config, inherited
  by `agentsEndpointSchema` — were silently ignored. Field resolution
  is now `all` > public endpoint > backing provider/custom, applied to
  both the enable gate and the model/titleModel resolution.
- E2E_LABEL_PORT reaches the YAML (playwright.config.mock.ts): an
  overridden port moved the fake label server and its health check but
  not the generated config's hard-coded 8889 baseURLs, so readiness
  passed while every label request targeted the wrong port. The
  override is now substituted into the generated copy.
- Every retry frame cancelled (useResumableSSE.ts): concurrent label
  retry chains (reservation + fill per slot) overwrote one rAF handle,
  so cleanup cancelled only the newest chain; the rest ran up to 120
  frames past unmount and could apply a stale label to a replacement
  generation reusing the same response id. Outstanding frame ids now
  live in a Set that cleanup drains.

Tests: public-endpoint gate/precedence/all-above-public (host.spec).

* 🖱️ fix: Keep the Last-Part Cursor in Parallel Lanes Too

Round-thirteen review (single P2): `ParallelContentRenderer` computed
`lastContentIdx` from the unfiltered array, so a trailing blank label
reservation — filtered out of every lane — left NO rendered part
carrying the last-part cursor and running-subagent affordances until
the label filled.

The sequential renderer's walk-back is extracted into a shared
`lastVisibleContentIdx` helper (utils/activityLabels) used by both
`ContentParts` and `ParallelContentRenderer`, so the two index spaces
cannot drift again. Behavior pinned in activityLabels.spec: trailing
blank skipped, consecutive blanks skipped, filled label counts,
label-free content unchanged.

* 🧹 chore: Alias the Retry-Frame Set for the Effect Cleanup Lint Rule

* 📏 fix: Let activityCharLimit Reach Tool Inputs

Round-fifteen review: `activityCharLimit` is documented as the
per-entry truncation for tool input AND output, but `buildPrompt`
hard-coded inputs at 200 characters — so raising the setting could
never surface a distinguishing path, query, or operation that appears
past the first 200 characters of a long argument. Inputs now truncate
at the configured limit alongside outputs; the 200-char constant
remains only for the intent line (renamed INTENT_CHAR_LIMIT to match).
Config fidelity pinned in runtime.spec: a 400-char argument survives a
450 limit and truncates under a 50 limit.

The round's other finding is the fifth restatement of the documented
edited+reconnect index-space limitation, answered on-thread with the
prior four cross-references.

* 🤝 fix: No Labels for Pure Handoff Batches

Round-sixteen review: a PostToolBatch containing only `transfer_to_*`
calls claimed a label slot, but transfer parts are never groupable —
the client flushed the handoff card standalone and the label orphaned
into a stray line after it, restating what the card already says.

Two-sided fix:
- Hook (runtime.ts): a batch whose every entry is a transfer call
  claims nothing — no slot, no model call, no `maxPerRun` consumption.
  Mixed batches still label (the header describes the real work).
- Renderer (groupToolCalls.ts): an orphan label whose `tool_call_ids`
  are all transfer calls is dropped instead of rendered standalone,
  covering content persisted before the hook-side skip.

The round's two P1s are repeats answered on-thread: the packages/api
extraction (maintainer-decided follow-up, recorded in the description)
and the sixth restatement of the edited+reconnect index limitation.

Tests: transfer-only batch claims nothing, mixed batch still claims
(runtime.spec); transfer-only orphan label dropped, real-batch orphan
label still renders (groupToolCalls.test).

* 🎛️ fix: Sanitize Label Client Options and Bound the Batch Prompt

Round-seventeen review: two fixed; the other two findings repeat the
maintainer-decided packages/api extraction (follow-up) and the
edited+reconnect index limitation (seventh instance), answered
on-thread.

- Primary-option strip (host.ts): the label client copied the resolved
  `llmConfig` wholesale, so an endpoint whose defaults enable extended
  thinking or carry model-specific output caps forwarded them to the
  (often cheaper) label model — unsupported options failed every label,
  and supported thinking spent real tokens and the settlement window on
  a 4–9 word header. The copy now strips `omitTitleOptions` keys and
  the `modelKwargs` output caps exactly like the title path, restoring
  the Anthropic `clientOptions` carrier by reference so proxy
  `defaultHeaders` still reach label requests.
- Batch prompt budget (runtime.ts): per-entry truncation left the batch
  dimension unbounded — hundreds of parallel calls could build a prompt
  past the fast model's window. The entries section now has a total
  budget (8k chars, scaling with `activityCharLimit` so a raised limit
  still fits several entries); entries past it are skipped without
  paying their serialization cost, and the list notes how many were
  omitted. The first entry always renders in full.

Tests: option strip with header-carrier survival (host.spec); giant
batch bounded with omission marker, small batch untouched
(runtime.spec).

* 🛡️ fix: Keep SSRF Guards on Label Calls, Skip Mixed Handoff Batches

Round-eighteen review: four fixed; the fifth repeats the
maintainer-decided packages/api extraction (eighth instance), answered
on-thread.

- SSRF-safe carrier (host.ts, P1): the sanitize step restored the
  Anthropic `clientOptions` carrier only when `defaultHeaders` existed,
  but for user-provided base URLs `getLLMConfig` stores the guarded
  Undici dispatcher and `redirect: 'error'` there — dropping it
  reopened DNS-rebinding/redirect paths on label calls to
  user-controlled URLs. The carrier (client CONSTRUCTION options, not
  generation params) is now restored whenever present, same reference.
- Primary maxTokens (host.ts): top-level `maxTokens` is not in
  `omitTitleOptions` and survived the strip; the title path deletes it
  explicitly, and a cap sized for the primary model can be rejected by
  the substitute. Deleted on the copy.
- Bounded keys (runtime.ts): the object branch materialized every key
  via `Object.keys` and quoted oversized keys in full before the budget
  check. Enumeration is now lazy (`for..in` + own-property guard) and
  keys slice to the budget before quoting, like string values.
- Mixed handoff batches (runtime.ts, groupToolCalls.ts): the client
  flushes the block at the transfer card, so a mixed batch's label
  orphaned exactly like a pure one. The hook now skips ANY batch
  containing a transfer call, and the renderer drops orphan labels
  covering one (legacy content).

Tests: carrier survival without headers by same reference, maxTokens
strip (host.spec); mixed batch claims nothing (runtime.spec); mixed
orphan dropped, real-batch orphan kept (groupToolCalls.test).

* 🧢 fix: Cap Label Generation, Order the Flag Persist, Detach Settled Listeners

Round-nineteen review: three fixed; the fourth is the ninth instance of
the edited+reconnect index limitation, answered on-thread.

- Generation cap (host.ts): stripping the primary output caps left
  label calls with NO cap at all — `normalizeLabelOutput` bounds what
  persists, not what the provider generates and bills, so a model
  ignoring the 4–9-word instruction (or steered by injected tool
  output) could emit its provider-default output per batch. The
  sanitize step now installs a 256-token label cap (per provider
  family: `maxOutputTokens` for Google-style wrappers, `maxTokens`
  otherwise), after the filter so the omit set cannot remove it.
- Flag-persist ordering (client.js): the `markActivityLabels` write was
  fire-and-forget, so an immediate cross-replica reconnect could read
  the job between the write and the first claim, see neither flag nor
  snapshot label, and skip gap reconciliation. Label emission now
  awaits the (settled-on-failure) persist chain, making "a label event
  exists" imply "the flag is durable" — the race window is gone; only
  the documented double-write-failure residual remains.
- Listener detach (client.js): each HITL approval cycle's wiring adds a
  `once` abort listener to the shared job signal that only an actual
  abort removes; settled segments now detach theirs in
  `settleActivityLabels`, so long multi-approval runs cannot accumulate
  dead closures toward the listener-limit warning.

Tests: the primary cap is REPLACED by the 256-token label cap
(host.spec).

* 🎯 fix: Route the Label Cap Per Model Family

Round-twenty review: the 256-token label cap set maxTokens
unconditionally, but GPT-5+ rejects max_tokens (the OpenAI builder
routes its cap into modelKwargs.max_completion_tokens /
max_output_tokens) and o-series models reject it with no stable kwargs
alternative — every label on those models would have failed. The cap
now mirrors the builder: modelKwargs for GPT-5+ (responses-API aware),
no cap for o-series (title parity; the 200-char persistence bound
still applies), maxOutputTokens for Google, maxTokens otherwise.
Pinned in host.spec for both reasoning families.

The round's other finding is the tenth instance of the documented
edited+reconnect index limitation, answered on-thread.

* ⏱️ fix: Persist the Label Flag at Run Start, Not on the Emit Path

Round-twenty-one review: two fixed; the other three repeat the
maintainer-decided packages/api extraction, the edited+reconnect index
limitation, and the parallel-lane header limitation — all answered
on-thread with their standing decisions.

- Flag ordering, corrected (client.js): sequencing label emission
  behind the flag persist (previous round) delayed the claim-time
  reservation while the shared index offset had ALREADY shifted
  subsequent SDK chunks — reopening the cross-instance
  hole-compaction overwrite the reservation emit exists to prevent.
  The reservation emits immediately again; instead, run start
  (processStream and resume alike) awaits the settled-on-failure
  persist chain, so the flag is durable before any batch can claim a
  label. Same guarantee, zero latency on the emit path.
- Tail-label cursor (ContentParts.tsx): a filled label at the content
  tail is consumed into the group header rather than listed in
  `group.parts`, so the `isLast` check missed it and nothing held the
  streaming cursor until the next delta. The check now includes
  `labelPart.idx`.

* 🔌 fix: Detach Label Abort Listeners Even Without Claims

A segment with labels enabled can end without a single claim (text-only, or handoff batches, which skip labels); the early return in settleActivityLabels skipped the detach added for HITL listener accumulation. The detach now runs on both paths.

* ⚖️ fix: Make the Commit Flag the Sole Billing Authority

Round-twenty-three review: a committed fill racing a late scope close
(user abort or settle timeout during the durable emit) stayed visible
— the part is mutated and persisted before the close — yet the
deferred accounting's scope gates then skipped the charge: a completed
provider call escaping both the label charge and the primary abort
accounting.

The scope gates on the deferred-usage path are removed; the hook's
commit flag is now the single billing authority in BOTH directions. A
dropped fill never reaches the accounting callback (billed-never-shown
stays impossible), and a committed fill bills regardless of when its
scope closed (shown-never-billed now impossible too). The dead
`scopeOpen` payload threading is removed with it; the
`recordActivityLabelUsage` parameter survives, defaulting open, for
callers that own no commit signal.

The round's other finding is the twelfth instance of the documented
edited+reconnect index limitation, answered on-thread.

* 🧮 feat: Bill Labels by Estimate When Providers Omit Usage

Maintainer decision: follow the title convention rather than leaving
label calls unbilled when a provider returns no usage metadata.

The hook now passes a LAZY estimate thunk with the deferred accounting
on the success path — the EXACT prompt the direct path sent (or the
locally built equivalent for the SDK path: same entries, context,
instruction, truncation contract, and continuity headers) plus the
final normalized label. `recordActivityLabelUsage` invokes it only
when no collected entry carries a real token count, counts both texts
with the shared o200k_base tokenizer, and feeds the synthesized entry
through the SAME pipeline (provider-tagged, streamed event, cost,
balance transaction). Real provider usage always wins when present.
The failure path passes NO estimate: a throw before a response bills
only real collected metadata, never a full phantom prompt.

Tests: the estimate thunk carries the exact invoked prompt and final
label; the failure path defers with no estimate (runtime.spec).

* 💵 fix: Estimate From the Raw Completion, Not the Normalized Label

The fallback estimate counted the normalized label (first line, 200-char cap) while the provider generated and would bill the raw output up to the 256-token generation cap — under-recording verbose replies. The estimate thunk now carries the raw pre-normalization text; the persisted label is unchanged. Pinned with a multi-line reply test.

* 🧾 fix: Commit Label Text Only After the Durable Emit, Estimate the Real SDK Prompt

Round review on the billing work: two fixed; the third is the
fourteenth instance of the edited+reconnect index limitation, answered
on-thread.

- Copy-first fill (wiring.ts): the fill mutated the shared content part
  BEFORE its durable emit, so a failed emit left the label text on
  `contentParts` anyway — persistence could save and display a label no
  client ever received and billing (keyed on the commit flag) never
  charged. The new state is staged on a copy; the shared part mutates
  only after the emit succeeds, so content, delivery, and billing move
  together.
- Real SDK prompt for estimates (client.js): the estimate thunk carried
  this module's locally built prompt, but the SDK path frames entries
  differently — the estimated input count was for a prompt never sent.
  Chain-start callbacks (handleLLMStart/handleChatModelStart) now
  capture the prompt the SDK actually rendered, and the deferred
  accounting substitutes it into the estimate when capture succeeded,
  falling back to the local approximation otherwise.
2026-07-29 14:05:47 -04:00
Danny Avila
cd215150cc
✳️ feat: Claude Opus 5 Support (#14422)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* ✳️ feat: Claude Opus 5 Support

- Add claude-opus-5 to Anthropic/Bedrock model lists, token maps, and pricing
- Extend requiresExplicitThinkingDisabled to Opus 5 so thinking-off sticks
- Clamp xhigh/max effort to high when thinking is disabled (Opus 5 400)

* 🪣 fix: Use Bedrock Inference Profiles and Add Vertex Opus Models

Bare `anthropic.` Claude 4+ IDs are not invocable on-demand via Converse:
Bedrock rejects them with "Retry your request with the ID or ARN of an
inference profile that contains this model." Verified live against
us-west-2 for Fable 5, Opus 5, Opus 4.8, Sonnet 5, Sonnet 4.6, Opus 4.6,
Sonnet 4.5, Haiku 4.5, and Opus 4.1. Switch those defaults to the
`global.` profile (no regional pricing premium); Opus 4.1 has no global
profile, so it uses `us.`.

Also add the modern Opus family to the Vertex defaults. `loadEndpoints`
swaps the shared Anthropic list for the Vertex model names, so Opus was
invisible to every Vertex deployment that did not enumerate models by hand.

* 📋 chore: Cover Opus 5 Gaps From PR #14420

Picks up items from the parallel community PR by @jona7o:

- Add claude-opus-5 to the librechat.example.yaml Vertex example (both the
  legacy array and the deploymentName map), which already lists Fable 5
  and Opus 4.8
- Mention Opus 5 in the configureReasoning doc comment, and note that its
  early return is why the effort cap is enforced by the caller
- Assert Opus 5 carries no long-context premium pricing
- Cover the Sonnet 5 negative case for the effort cap, and the persisted
  disabled-object round-trip carrying an effort

* 🌍 docs: Warn That Vertex Regional Endpoints Reject Modern Models

Anthropic serves Sonnet 4.6 and earlier on specific Vertex regional
endpoints; newer models (Opus 4.7+, Opus 5, Sonnet 5, Fable 5) require
`global` or a multi-region location and 404 on a specific region. The
`us-east5` default therefore cannot serve the Opus models added here, nor
the Sonnet 5 entry that predates this branch.

Documents the constraint at all three places an operator sets the region,
and at the fallback itself. Leaves the default unchanged: switching it to
`global` would silently alter data routing and residency for existing
deployments, which is a separate call.

* 🩹 fix: Restore PDF Exemption for Undated IDs and Gate Vertex Defaults

Two issues raised in review:

- BEDROCK_CLAUDE_4_PLUS_RE required a `-` after the major version, so it
  matched `claude-opus-4-8` but not undated IDs like `claude-opus-5`.
  Those models silently lost the Claude 4+ PDF exemption and fell back to
  the 4.5 MB limit. Sonnet 5 and Fable 5 were already affected before this
  branch; Fable/Mythos were also missing from the family alternation.

- The Vertex defaults advertised models that only `global` and the
  multi-region locations serve, so a default `us-east5` deployment listed
  Opus choices that 404 on first request. Filter the built-in defaults by
  configured region instead of changing the region default, which would
  alter data routing for existing deployments. An explicit `vertex.models`
  list is the operator's choice and is never pruned.

* 🧩 fix: Match Bare Claude IDs in the Bedrock PDF Exemption

An application inference profile maps a LibreChat model ID with no
`anthropic.` segment, so `claude-opus-5` failed the Claude 4+ check and
fell back to the 4.5 MB PDF limit. Make the prefix optional and accept
both segment orders, mirroring BEDROCK_CLAUDE_4PLUS_THINKING in
librechat-data-provider, which matches on the family token for exactly
this reason.

Only reached for the Bedrock provider, so the looser prefix cannot leak
into other endpoints. Verified Claude 3.x, Nova, Llama, Cohere, and
Mistral IDs still fall through to the default limit.

* 🧹 fix: Drop Retired Claude 3.5 Models From Bedrock Defaults

The three Claude 3.5 entries reached end of life at AWS and return
ResourceNotFoundException in every prefix form (bare, `us.`, `global.` —
verified live against us-west-2), so selecting one was a hard error.
Their modern equivalents are already in the list: Sonnet 5 / Sonnet 4.6
supersede the 3.5 Sonnets, and Haiku 4.5 supersedes 3.5 Haiku.

Every remaining Anthropic default is now live-verified invocable.
`.env.example` swaps its retired example ID for Haiku 4.5.

* 🔒 refactor: Narrow Effort-Clamp Types Instead of Asserting

Both clamp sites reached into loosely-typed containers with assertions:
llm.ts used an `as unknown as { type?: string }` double assertion to read
the thinking type, and the Bedrock parser cast `output_config` to
`{ effort?: unknown }` before confirming it was an object. CLAUDE.md's
type-safety rules call for narrowing over both.

Adds `isThinkingDisabled` and `clampOutputConfigEffort` to
librechat-data-provider, using `in`-operator narrowing and a type
predicate so no assertion is needed at all. Both call sites now share one
implementation rather than duplicating the clamp.

Behavior is unchanged; existing clamp tests cover it.
2026-07-24 21:45:32 -04:00
Danny Avila
e515063ffe
🔗 feat: Snapshot Files for Shared-Link Attachments (#13740)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* 🔗 feat: Snapshot Files for Shared-Link Attachments

Shared-link viewers could read a shared conversation snapshot but not its
attachments: file preview/download still went through the owner-scoped file
ACL (the /api/files router sits behind requireJwtAuth + owner/agent checks),
so anonymous viewers got 401s and authenticated non-owners got 403s — the
repeated `[fileAccess] denied` warnings seen for the preview poller.

Capture an immutable per-share file snapshot (embedded on the SharedLink
document, referencing the original stored object — no byte copy) at share
create/update, and serve those files through new share-scoped routes
authorized by the existing shared-link view permission (public/ACL) plus
snapshot membership, never the owner's live file ACL.

- data-schemas: fileSnapshots on the share doc; capture in create/update;
  read-time rewrite of filepath/preview to /api/share/:id/files/:fileId;
  getSharedLinkFile + lazy backfillSharedLinkFiles for legacy links
- api: GET /api/share/:shareId/files/:file_id[/download|/preview]; route
  context added to fileAccess denial logs
- packages/api: isFileSnapshotEnabled resolver (env + yaml)
- data-provider: interface.sharedLinks.snapshotFiles (default on) + client
  endpoints/services
- client: ShareContext.shareId wired to Image, preview hook, and downloads
- config: SHARED_LINKS_SNAPSHOT_FILES env override (default on)

* 🔒 fix: Address Codex review on shared-link file snapshots

Triage of the Codex review on PR #13740 (2 P1, 7 P2 — all valid):

- P1 (cross-user access): scope the snapshot lookup to the sharing user's own
  files so a message referencing another user's file_id can't widen access.
- P1 (stored XSS): the inline share-file route now serves only safe preview
  types inline (raster images/pdf); everything else is forced to attachment with
  X-Content-Type-Options: nosniff.
- Stream shared downloads by default; redirect to a signed URL only on
  ?direct=true (blob/XHR callers work without bucket CORS).
- Read preview status live from the file record (always current for deferred
  previews) and stop embedding extracted text in the share doc (16MB-limit risk).
- Only lazily backfill when the fileSnapshots field is absent (legacy), not on
  every snapshot miss.
- Backfill legacy shares before rewriting message URLs, and gate URL rewriting
  to public shares so non-public (ACL) shares keep prior behavior (img/anchor
  can't carry the bearer token).
- Frontend: only route a download through the share path when the file was
  actually snapshotted (rewritten href / filepath), else fall back.

* 🔑 feat: Authorize shared-link files for non-public shares via cookie

Extends shared-link file access to non-public (ACL) shares (Codex finding 5).
`<img>`/anchor requests can't carry the bearer access token, so non-public
shares previously 401'd on file loads. Add an optional cookie-auth fallback on
the share file routes that resolves the viewer from the `refreshToken` cookie
(or signed `openid_user_id` cookie) — the same mechanism secure image links use
(validateImageRequest) — then let canAccessSharedLink run the viewer's ACL check.

- new middleware optionalShareFileAuth (+ unit spec); applied to the three
  share file routes after optionalJwtAuth
- URL rewriting in getSharedMessages is no longer gated to public shares (the
  route now authorizes header-less requests), so files work uniformly across
  public and non-public shares; revert the now-unused req.sharePublic plumbing

* 🔒 fix: Second Codex pass on shared-link file snapshots

Addresses the follow-up Codex findings on PR #13740:

- Don't snapshot transient text-source files: FileSources.text filepaths are
  Multer temp paths the upload route deletes, so they can't be streamed —
  removed from the streamable allowlist.
- Unset stale snapshots on a disabled-feature update: updateSharedLink now
  $unsets fileSnapshots when snapshotFiles is false, so an opted-out update
  can't keep serving file ids the update dropped.
- Load tenant config after share resolution: configMiddleware now runs after
  canAccessSharedLink (which enters the share's tenant ALS context), so
  per-tenant interface.sharedLinks.snapshotFiles overrides apply to anonymous
  public views.
- Return a clean 404 when the snapshotted object is gone: resolveShareFile now
  requires the live file record and 404s if it's been deleted/expired, instead
  of letting the stream error after headers are sent (ENOENT / 500).

(The re-flagged P1 about private-viewer rewriting was already fixed in the prior
commit's cookie-auth change.)

* 🔒 fix: Third Codex pass on shared-link file snapshots

Addresses the third Codex review pass on PR #13740:

- P1: keep shared previews/files pinned to the snapshotted version. Snapshot the
  small previewRevision; resolveShareFile 404s when the live file's revision no
  longer matches (file_id reused/overwritten by a later turn), so old links can't
  surface post-share content — covers both preview text and streamed bytes.
- Honor the toggle as a kill switch: resolveShareFile 404s when snapshotFiles is
  disabled, instead of only skipping backfill, so disabling stops serving
  already-snapshotted file URLs.
- Lazy-sweep orphaned 'pending' previews to 'failed' in the share preview route
  (mirrors the owner route) so the client poller reaches a terminal state.
- Resolve the cookie-fallback user in runAsSystem so strict tenant isolation
  doesn't throw before canAccessSharedLink establishes the share tenant context.

*  feat: Per-link "share files" checkbox for shared links

Add a checkbox to the share-link dialog (checked by default) letting the user
choose whether to include the conversation's files in the shared link, with
copy explaining images/files won't be visible to viewers otherwise. Opting out
skips snapshot creation/serving for that link.

- client: ShareButton renders the checkbox gated on the new
  startupConfig.sharedLinksSnapshotFilesEnabled flag; state threads through
  SharedLinkButton into the create/update mutations as `snapshotFiles`.
- data-provider: createSharedLink/updateSharedLink send `snapshotFiles` in the
  body; TStartupConfig gains `sharedLinksSnapshotFilesEnabled`.
- api: POST/PATCH /api/share compute snapshotFiles as
  isFileSnapshotEnabled(req.config) && body.snapshotFiles !== false (admin gate
  AND per-link opt-out); config.js exposes the effective enabled flag to clients.
- en locale: com_ui_share_files (+ _description).

* 🐛 fix: Make the "share files" opt-out actually hide files

Unchecking "share files" at creation didn't hide anything: the shared message
JSON still carried each file's original (e.g. static-served) path, and because
opting out only meant "no fileSnapshots field" — indistinguishable from a legacy
link — getSharedMessages would backfill snapshots on first view whenever the
admin feature was on, re-enabling files entirely.

Fix by persisting and honoring the per-link choice:
- Store `snapshotFiles` (boolean) on the SharedLink so opt-out is distinct from a
  legacy link; set it on create and update.
- getSharedMessages computes includeFiles = adminEnabled && link not opted out;
  when excluded it strips files/attachments from the payload (no original-path
  leak) and never backfills the opted-out link.
- Surface the stored choice via getSharedLink so the dialog checkbox reflects an
  existing link's actual setting instead of always defaulting to checked.

Note: changing the checkbox on an already-created link still applies only when
the link is refreshed (which regenerates the URL) — a UX follow-up.

* 🔒 fix: Close remaining shared-link file opt-out leaks (Codex)

Follow-up to the per-link opt-out, addressing the third Codex pass:

- Honor the opt-out on the file route too: getSharedLinkFile now returns the
  link's `optedOut` choice; resolveShareFile 404s (and never backfills) an
  opted-out link, so a direct /files/:id request can't re-create snapshots.
- Make read/serve viewer-independent: the gate no longer uses the viewer's
  resolved config (isFileSnapshotEnabled(req.config)) — it uses the link's stored
  choice plus a global env-only kill switch (isFileSnapshotKillSwitchActive). A
  viewer's own interface.sharedLinks.snapshotFiles can no longer hide a link's
  files. Create/update still use the creator's config to set the per-link choice.
- Neutralize render URLs for non-snapshotted files: applyShareFileRoute now
  strips filepath/preview for any file/attachment not in the snapshot, so the
  owner's original (e.g. static) path can't be loaded through the share.

* 🔒 fix: Harden shared-file version pinning and local path handling (Codex)

- Refuse reused/overwritten file snapshots more broadly: resolveShareFile now
  refuses to serve when either previewRevision OR `bytes` changed vs the
  snapshot. `bytes` catches non-office reused outputs (e.g. code-exec
  same-filename images that lack previewRevision) and is stable across S3 URL
  refresh and the pending->ready transition. Same-size content swaps remain a
  best-effort gap inherent to the no-byte-copy design.
- Strip cache-busting query strings before local streaming: code-output images
  add `?v=...` to filepath; the share route now splits it off so getLocalFileStream
  resolves the real filename instead of a literal `*.png?v=...` path.

* 💬 fix: Clarify that file-sharing changes apply on link refresh

For an already-created shared link, changing the "share files" checkbox only
takes effect when the link is refreshed (which regenerates the snapshot). Add a
note under the checkbox, shown only when a link already exists, so the behavior
isn't surprising: "Refresh the link to apply this change — files are snapshotted
when the link is refreshed."
2026-06-20 23:05:13 -04:00
Danny Avila
f76a5faa9e
📌 feat: Seed Default Pinned Tools and MCP Dropdown via Interface Config (#13865)
*  feat: Add `defaultPinnedTools` interface config for default tool & MCP pinning

Adds an `interface.defaultPinnedTools` string array letting admins pin tools and the MCP servers dropdown to the prompt bar by default for all users.

- Tool keys (artifacts, execute_code, web_search, file_search, skills) pin their badge via `useToolToggle`.

- The keyword `'mcp'` or a configured MCP server name pins the MCP dropdown via `useMCPSelect`.

- Only seeds initial state; a user's stored pin preference always wins. When unset, tools start unpinned and the MCP dropdown keeps its legacy default (pinned).

Unifies the approaches from #11646 (pinnedTools) and #9251 (defaultPinMcp) into one config key.

* 🐛 fix: Apply defaultPinnedTools pin once startupConfig resolves

On a cold load, useToolToggle can mount before useGetStartupConfig() resolves, so defaultPinned starts false and useLocalStorageAlt eagerly persists it; its init effect never re-runs for the later config-driven default. Fresh users would then miss the admin-configured default pin whenever startup config was not already cached.

Capture whether a pin preference existed before mount (pre-seed) and, once startupConfig arrives, apply the real default for users with no prior preference. Runs once and never overrides an existing stored pin, so the conservative behavior for existing users is preserved.

* 🐛 fix: Preserve pin clicks made before startupConfig resolves

The cold-load default-seeding effect captured the stored-pin state only at mount, so a pin toggled before startupConfig resolved was treated as no-preference and overwritten when the admin default applied.

Track explicit pin toggles via a ref (set through the returned setter) and skip the default application when the user has interacted in-session — in addition to the existing stored-preference guard.
2026-06-20 13:40:10 -04:00
Danny Avila
b917e0418b
v0.8.7-rc1 (#13592)
* chore: Bump LibreChat to v0.8.7-rc1

* docs: Sync Chinese README
2026-06-15 13:10:30 -04:00
Danny Avila
9efe4878e7
🅰️ feat: Native Anthropic Provider for Custom Endpoints (#13748)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
* 🅰️ feat: Native Anthropic provider for Custom Endpoints

Let a custom endpoint declare `provider: anthropic` to use the native Anthropic
`/v1/messages` client (the agents SDK's ChatAnthropic) against its own
`baseURL`/`apiKey`/`headers`, instead of being forced through the
OpenAI-compatible client. Enables Anthropic itself and Anthropic-compatible
gateways (AI gateways, OpenCode Zen, etc.) as custom endpoints — including for
agents and role-scoped model access.

Closes #10655 (Option 1: explicit provider).

- Schema: add optional `provider` (currently `anthropic`) to the custom
  `endpointSchema` in data-provider.
- Routing: `getProviderConfig` maps a custom endpoint with `provider: anthropic`
  to `Providers.ANTHROPIC` (was always `Providers.OPENAI`).
- Config: `initializeCustom` builds the native Anthropic config via the Anthropic
  `getLLMConfig` (custom baseURL/apiKey/headers) and returns `provider: anthropic`;
  `useLegacyContent` is left unset to match the built-in Anthropic endpoint. The
  OpenAI-compatible path is unchanged for endpoints without `provider`.
- Summarization: `resolveSummarizationProvider` builds an Anthropic config for a
  cross-endpoint native-Anthropic summarization target (self-summarize already
  reuses the agent's client options).

Title generation already resolves via `agent.endpoint`, and provider-specific
handling (tool conflicts, content/PDF validation, token counting, streamUsage)
already branches on `Providers.ANTHROPIC`, so it applies automatically.

Note: model auto-fetch (`models.fetch`) uses the OpenAI `/models` convention and
is not used for this provider — list models explicitly under `models.default`.

* 🅰️ fix: Anthropic custom-endpoint param parity (Codex review)

Address Codex P2 findings — the native Anthropic path must match the
OpenAI-compatible path's parameter handling:

- UI param set: `loadCustomEndpointsConfig` now surfaces `provider` as the
  client `customParams.defaultParamsEndpoint`, so the Agents model panel shows
  Anthropic fields (`maxOutputTokens`/`thinking`) instead of OpenAI `max_tokens`
  (which the native initializer ignored). An explicit non-default
  `defaultParamsEndpoint` still wins.
- Provider override: `getProviderConfig` re-applies `provider: anthropic` after
  all `customEndpointConfig` resolution, so it also wins when the endpoint name
  collides with a known custom provider (e.g. `openrouter`) — fixing the
  token/context budget derived from `overrideProvider`.
- Default params: the native path (and cross-endpoint Anthropic summarization)
  now apply `customParams.paramDefinitions` defaults via `extractDefaultParams`,
  matching what `getOpenAIConfig` does for the OpenAI-compatible path.

Adds tests for each.
2026-06-14 18:23:48 -04:00
Danny Avila
2350ebb24a
📨 feat: Custom Headers on Built-in Provider Endpoints (#13742)
* 📨 feat: Custom Headers on Built-in Provider Endpoints

Add a `headers` config option to the built-in `openAI`, `anthropic`, and
`google` endpoints (incl. Anthropic/Google Vertex), mirroring the custom
endpoint header mechanism. Values support the same placeholder resolution
(env vars, `{{LIBRECHAT_USER_*}}`, `{{LIBRECHAT_BODY_CONVERSATIONID}}`) and
are resolved at request time so dynamic values like conversationId resolve
against the live request — without losing provider-native request shaping.

Closes #13082. Covers #13713: forwarding conversationId to a reverse proxy
is now `X-Conversation-Id: '{{LIBRECHAT_BODY_CONVERSATIONID}}'` — an unknown
header is ignored by the native Anthropic API, so no 400 and no metadata
gating needed.

- Schema: `headers` on `baseEndpointSchema` (openAI/google/anthropic/all).
- New `mergeHeaders`/`resolveConfigHeaders` utils centralize the per-provider
  header locations (`configuration.defaultHeaders`, Anthropic
  `clientOptions.defaultHeaders`, Google `customHeaders`); provider-managed
  headers (auth, `anthropic-beta`) always win on collision.
- Each initializer threads configured headers (endpoint over `all`) into the
  right place; request-time resolution runs across all locations in the main
  and title flows.

* 🩹 fix: Cast endpoints.all to TEndpoint for headers DeepPartial widening

Adding `headers` (a Record) to `baseEndpointSchema` makes `DeepPartial<TCustomConfig>`
widen its value type to `string | undefined`, which is not assignable to the
concrete `TEndpoint['headers']: Record<string, string>` at the `loadedEndpoints.all`
assignment. Cast at the assignment site, mirroring the existing
`anthropicConfig as TAnthropicEndpoint` cast in the same function.

* 🛡️ fix: Harden built-in endpoint custom headers (Codex review)

Address Codex P2 findings on the custom-headers feature:

- Anthropic title requests: `omitTitleOptions` strips the `clientOptions`
  carrier, which dropped its `defaultHeaders`. Preserve just the header carrier
  so gateway/reverse-proxy metadata still reaches title generation.
- mergeHeaders: match header names case-insensitively so an override (e.g. a
  provider-managed `Authorization`/`anthropic-beta`) replaces/uniones a
  case-variant from the base instead of emitting two names a client may collapse.
- OpenAI: withhold admin-configured headers when the user supplies the base URL
  (`user_provided`), since values may carry `${SECRET}`/token placeholders that
  must not reach a user-controlled endpoint — mirrors the custom-endpoint guard.
- Azure: honor global `endpoints.all` headers (same OpenAI carrier) while keeping
  Azure-managed `api-key`/version headers authoritative.

Adds tests for each.

* 🔐 fix: Resolve-once + provider-managed header safety (Codex review round 2)

Address Codex P2 findings:

- Azure: keep global `endpoints.all` headers unresolved at init and let
  request-time `resolveConfigHeaders` resolve them once, avoiding a
  second-order env expansion of already-substituted user values.
- Google: `resolveConfigHeaders` no longer template-resolves the
  provider-managed `Authorization` header (built from a possibly user-provided
  key), so a user key like `${ENV}` can't leak server environment values.
- Model fetches: thread configured headers (endpoint over `all`) + user object
  through `getOpenAIModels`/`getAnthropicModels` → `fetchModels`, so a
  gateway-fronted built-in provider receives the header on `/models` too. Fixed
  `fetchModels` to merge custom headers for Anthropic instead of overwriting
  them (managed `x-api-key`/version still win).

Adds/updates tests for each.

* 🧯 fix: Header provenance, memory/title coverage, idempotency (Codex round 3)

Address Codex P2 findings, including two regressions from the prior round:

- Google auth (findings 6 & 8): move native Google header resolution to init
  (`initializeGoogle`), resolving admin templates BEFORE the key-derived auth
  header is built. resolveConfigHeaders no longer touches Google `customHeaders`,
  so admin `Authorization` templates resolve again (fixes the round-2 regression)
  while the SDK auth header (possibly a user-provided key) is never env-expanded.
- Memory runs: memory extraction now calls `resolveConfigHeaders`, so native
  Anthropic (and OpenAI) headers resolve for memory requests too.
- Vertex titles: restore the ORIGINAL `clientOptions` object reference (not a
  copy) when preserving headers across `omitTitleOptions`, so the Vertex
  `createClient` closure and the resolved headers stay on the same object.
- Reuse: `resolveConfigHeaders` is now idempotent (resolve-once per header map),
  preventing a second pass from env-expanding values already substituted with
  user/body data when an agent object flows through buildAgentInput twice.

Adds/updates tests for each.
2026-06-14 17:02:04 -04:00
Danny Avila
197a1dc4e2
🧬 feat: Add GitHub Skill Sync (#13293)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
Publish `librechat-data-provider` to NPM / pack (push) Has been cancelled
Publish `librechat-data-provider` to NPM / publish-npm (push) Has been cancelled
* feat: Add GitHub skill sync

* fix: Address GitHub skill sync CI

* fix: Harden GitHub skill sync review paths

* fix: Prevent overlapping skill sync runs

* fix: Address GitHub skill sync review findings

* fix: Satisfy Git ref lint rule

* fix: Address GitHub sync review follow-ups

* fix: Match skill frontmatter closing fence

* fix: Address GitHub sync review cycle

* fix: Address GitHub sync review follow-ups

* fix: Harden GitHub skill sync worker

* fix: Format GitHub sync rollback log

* fix: Address GitHub sync review feedback

* fix: Format skill import parse handling

* fix: Coerce scalar skill frontmatter and correct scheduler timer clear

- parse: coerce numeric/boolean name and description scalars to strings instead of dropping them to empty (restores pre-refactor behavior; preserves absent-vs-empty distinction for the when-to-use fallback)
- scheduler: clear the setTimeout handle with clearTimeout rather than clearInterval
- test: cover non-string scalar frontmatter coercion

* fix: Tolerate trailing whitespace after SKILL.md opening frontmatter fence

extractFrontmatterBlock required the opening fence to be exactly '---\n', so an opener with trailing spaces/tabs (e.g. '---   \n') silently dropped all frontmatter even though the closing-fence regex already tolerates it. Match the opener with /^---[ \t]*\n/ for symmetry. Addresses Codex P3 (parse.ts:24).

* feat: Run GitHub skill sync under a per-source tenant context

Under TENANT_ISOLATION_STRICT, the sync ran with no async tenant context, so the tenant-isolation mongoose hooks threw on every Skill/SkillFile/AclEntry operation; in non-strict mode synced skills were written tenant-less and never matched tenant-scoped reads. Add an optional per-source tenantId to the skillSync config; when set, each source sync runs inside tenantStorage.run({ tenantId }) so skills, files, and public ACL grants are created and listed within that tenant, and the skill row is stamped with the tenantId for correct dedup. Sources without tenantId keep the prior single-tenant behavior. Avoids runAsSystem. Addresses Codex P2 (sync.js:70).

Lock/status/credential bookkeeping stays outside the tenant context (those collections are intentionally global).

* test: Restore dropped tenant-context coverage for GitHub skill sync

The prior commit shipped the getTenantId import in github.spec.ts without the tenant tests that use it (lost in an interrupted edit), which failed the eslint --max-warnings=0 CI job on an unused import. Restore both github.spec.ts tenant tests (tenant-scoped run stamps tenantId and executes inside the tenant ALS context; no-tenant run stays ambient) and the two config-schemas tenant tests (accepts tenantId, rejects __SYSTEM__).

* test: Restore dropped github.spec tenant-context tests

The previous commit's github.spec.ts edit did not apply (anchor mismatch), so the getTenantId import remained unused and failed eslint --max-warnings=0. Add the two tenant tests that use it: a tenant-scoped run stamps tenantId and executes inside the tenant ALS context, and a no-tenant run stays ambient.

* feat: Scope synced skill author to tenant and harden tenant-context sync

Addresses the latest Codex review on the per-source tenant change:
- makeSourceAuthorId now folds tenantId into the synthetic author hash so the
  same source mirrored into different tenants gets distinct author ids (clearer
  audits, no cross-tenant author collisions). Single-tenant author ids stay
  stable (suffix omitted when tenantId is absent).
- syncSourceInTenantContext uses an async callback per the tenant-context
  contract so the ALS store propagates across awaited Mongoose calls.
- Tests: same-source/different-tenant yields distinct authors; mirror cleanup
  is scoped to the source and deletes only its absent-upstream skills.

* fix: Repair tsc error and guard external edits in github skill sync

- Fix TS2352 in github.spec mirror-cleanup test: build the existing-skill mock via makeSkill with authorName instead of an under-typed 'as CreateSkillInput' cast (this was the failing TypeScript CI check on f00ce3c5a).
- 808: commitExistingRemoteSkillAfterFileSync re-reads to clear our own file-sync version bumps, but now compares refreshed content against the pre-sync snapshot (body/name/description/always-apply) and throws SKILL_CONFLICT on a concurrent external edit instead of overwriting it.

* docs: Note skillSync source tenantId is effectively immutable

Changing/adding/removing a source's tenantId orphans previously mirrored skills in the old tenant (a tenant-scoped sync cannot clean another tenant's data without runAsSystem, which is intentionally avoided).

* fix: Key GitHub skill upstream identity on source id and path only

Addresses Codex finding (github.ts:217): makeUpstreamId previously included owner/repo, so repointing a source to a renamed or replacement repository (same source id) changed the upstreamId, made findSkillBySourceIdentity miss the existing mirror, and then collided on the (name, author, tenantId) uniqueness constraint — leaving the source stuck failing. Identity now keys on the stable source id + root path only. The feature is unreleased, so there is no stored-id migration. Updated spec upstreamId fixtures to the new format; the existing ref-independent identity test now also covers repo moves.

* fix: Scope GitHub skill mirror deletion to the source tenant

Addresses Codex P1 (github.ts:1047/1057): an ambient source (no tenantId) runs listSkillsBySource without tenant context, which under non-strict isolation returns github-synced skills across all tenants. The mirror-deletion pass then treated other tenants' skills as absent-upstream and could delete them. Filter existingSyncedSkills to rows whose tenantId matches the source's configured tenantId (absent = its own ambient bucket) before deleting, so a sync never removes another tenant's mirrored skills. Covered by a test where an ambient run leaves a tenant-b-owned skill untouched.

* fix: Apply tenant-scoped mirror deletion implementation

The prior commit (75ccfa3fc) added the test but the source change to github.ts was lost in an interrupted edit, leaving a failing test with no implementation. This adds the actual guard: the mirror-deletion pass skips skills whose tenantId does not match the source's configured tenantId (absent = ambient bucket), so an ambient source whose listSkillsBySource returns cross-tenant rows under non-strict isolation cannot delete another tenant's mirrored skills.

* fix: Resolve global access role outside tenant context for synced skill grants

Addresses Codex P2 (github.ts:1166): default access roles (incl. skill_viewer) are seeded globally with no tenantId under runAsSystem, but a tenant-scoped sync wraps ensurePublicViewer in the source's tenant context. The PermissionService grantPermission resolved the role via a tenant-isolated AccessRole query, so the global role did not match and tenant-scoped syncs failed with 'Role skill_viewer not found'. The sync adapter now resolves the role inside runAsSystem (matching the global seed) and writes the ACL entry in the active tenant context, so the AclEntry is tenant-scoped (visible to tenant users) while the role lookup still succeeds. Covered by service tests for the resolve-vs-write split and the missing-role failure.

* fix: Strip placeholder frontmatter booleans and check skill conflict before file sync

- 1083 (github.ts:759): toCleanFrontmatter now drops a non-boolean always-apply (e.g. the 'always-apply:' / 'always-apply: # TODO' placeholder, which js-yaml yields as null). The boolean is already captured in the dedicated alwaysApply field; persisting null left ambiguous frontmatter on the synced skill.
- 1080 (github.ts:1057): for an existing mirrored skill, check for an external content edit (via getSkillById + hasExternalSkillEdit) BEFORE syncSkillFiles mutates the bundled files, so a concurrently edited skill fails fast with SKILL_CONFLICT without partial file rewrites. The post-file-sync check still guards edits that land during the file sync window.
Tests: placeholder always-apply is dropped from synced frontmatter; concurrent-edit conflict leaves files unmutated (no upsert/delete).

* fix: Harden GitHub skill sync review paths

* fix: Reuse moved GitHub skill mirrors

* fix: Scope GitHub sync identity conflicts

* test: Fix GitHub sync conflict mock typing

* fix: Support nested env-backed skill sync

* fix: Keep skill sync config base-only

* fix: Scope GitHub skill identity lookup by tenant

* fix: Harden GitHub skill sync admin gates

* fix: Guard existing skill sync permission grants

* feat: Trigger skill sync from resolved config

* fix: Scope resolved skill sync by tenant

* test: Allow manual skill sync status tenant scoping

* refactor: Extract skill sync trigger orchestrator

* test: Complete orchestrator status fixture

* chore: Bump data provider version

* fix: Restrict skill sync server credentials

* test: Complete admin skill sync status fixtures

* fix: tighten skill sync trigger safeguards

* fix: preserve alwaysApply skill sync alias

* chore: sort skill sync imports

* fix: preserve skill sync request scope

* fix: harden skill sync review edges

* refactor: move skill sync admin access to api package

* fix: add skill sync declaration return types

* fix: satisfy skill sync type checks

* fix: resolve codex skill sync review findings

* fix: harden skill sync review edges

* fix: resolve codex skill sync edge findings

* fix: satisfy API declaration build after rebase
2026-06-10 21:05:54 -04:00
Dustin Healy
5867f1a065
🛡️ feat: Configurable Message PII Filter (#13602)
* 🛡️ feat: Reject chat messages matching configured credential patterns

Adds an opt-in `messagePiiFilter` middleware mounted on the agent
chat route ahead of `moderateText`. When the configured patterns
match the user's input the request is refused with 400, so the
credential never reaches OpenAI moderation, the model, or MongoDB.
Three starter patterns ship by default and operators can subset
them or add their own regex via `customPatterns` in librechat.yaml.

* 🧪 test: Memoize compiled patterns + add middleware spec

Memoize the compiled pattern array via a WeakMap keyed by the
messagePiiFilter config object so repeat requests against the same
config skip the per-request RegExp construction. Cache entries are
released automatically when the config object itself rotates.

Adds packages/api/src/middleware/messagePiiFilter.spec.ts covering
the default-starter rejections, the starterPatterns subset and
empty-array semantics, customPatterns matching layered on top of and
in place of the starters, the no-config and empty-text pass-through
paths, and a memoization regression check.

* 🛡️ fix: Skip invalid customPattern regexes instead of crashing the request

Admin DB overrides for `messagePiiFilter.customPatterns` reach
`req.config` via `mergeConfigOverrides`, which deep-merges raw
override values without re-running `configSchema`. A typo'd regex
like `(` would slip past the YAML-load validation and throw inside
`new RegExp(...)` during `compile()`, returning 500 for every chat
request until the operator rolled the override back.

Wrapped the per-pattern compile in a try/catch that logs the
invalid pattern id + reason and skips it, so other valid patterns
(starters and other custom entries) keep filtering. Added a
regression test alongside the existing spec.

* 🛡️ feat: Extend PII filter to OpenAI-compatible and Responses agent APIs

The chat-route middleware operates on `req.body.text`, but the remote
agent API endpoints (`/api/agents/v1/chat/completions`,
`/api/agents/v1/responses`) accept the same prompt content as a
`messages` array or an `input` field. A caller using their API key
could send a credential-shaped value through either route and bypass
the configured PII filter even though they share the same agent and
model backbone the middleware is meant to guard.

Factored out `findPiiMatchInMessages`, a tolerant walker that handles
both `content: string` and `content: ContentPart[]` user-message
shapes against the same compiled, cached pattern list. Wired it into
the OpenAI-compat controller after agent lookup and into the
Responses controller right after `convertToInternalMessages`. Each
returns the endpoint's native 400 error shape
(`sendErrorResponse` / `sendResponsesErrorResponse`) with the
`message_pii_filter_block` code when a user message matches.

* 🩹 test: Add findPiiMatchInMessages to OpenAI + Responses controller mocks

The OpenAI-compat and Responses controller specs mock `@librechat/api`
with a hand-listed object. The new `findPiiMatchInMessages` export
wired into both controllers in 3ea35af9a was missing from those
mocks, so the production lookup returned undefined and the controllers
threw at request time under jest. Added the missing entries (default
mock: returns null so the handlers fall through to the existing happy
paths). All 278 agents-controller tests pass locally.

* 🧹 refactor: Namespace messagePiiFilter under messageFilter.pii + fix import order

Renames the yaml field `messagePiiFilter` to `messageFilter.pii`, the
module to `messageFilterPii`, the factory to `createMessageFilterPii`,
the type to `MessageFilterPiiConfig`, and the error code to
`message_filter_pii_block`. The wrapper `messageFilter` namespace
gives future safety filters (e.g. `messageFilter.toxicity`) a place
to plug in without restructuring the config later. The
`findPiiMatchInMessages` helper kept its name because it already
describes what it does at the value level.

Also fixes import order Danny flagged on the OpenAI-compatible and
Responses controllers: `findPiiMatchInMessages` was appended at the
bottom of two `require('@librechat/api')` destructures rather than
placed in the length-sorted slot the house style expects.

* 🧹 chore: Length-sort the general require destructure in responses.js

Reorders the general sub-group inside the `require('@librechat/api')`
destructure shortest to longest so the whole block conforms to the
length-sort rule the file's `// Responses API` sub-group already
follows. Pure reorder, no other changes.

* 🧹 chore: Length-sort the defaultConfig block in AppService

Reorders the `defaultConfig` keys in `packages/data-schemas/src/app/service.ts`
shortest-line to longest-line, with the explicit-value entries
(`mcpConfig`, `fileStrategies`, `cloudfront`) trailing the shorthand
ones. Pure reorder, no behavior change.
2026-06-10 09:03:05 -04:00
Danny Avila
2aea5f4a3a
📖 feat: Add Claude Fable 5 Support (#13628)
* 📖 feat: Add Claude Fable 5 Support

Claude Fable 5 (`claude-fable-5`) is Anthropic's most capable widely
released model (GA 2026-06-09). Its naming drops the opus/sonnet/haiku
tier, so LibreChat's name-parsing helpers miss it; this teaches them the
Mythos-class family (Fable / Mythos) and registers the model.

- Add `parseMythosClassVersion` and route Fable/Mythos through
  `supportsAdaptiveThinking`, `omitsThinkingByDefault`,
  `omitsSamplingParameters`, and `supportsContext1m`
- Extend the Bedrock detection regexes (beta headers + adaptive-thinking
  branch) and `checkPromptCacheSupport` to match `claude-(fable|mythos)`
- Return 128K max output for Fable/Mythos in `maxOutputTokens.reset`/`set`
- Register `claude-fable-5` in shared Anthropic + Bedrock model lists,
  1M context / 128K output token maps, and $10/$50 pricing with 12.5/1
  cache rates (`claude-mythos-5` added to token + pricing maps only,
  since it is limited-availability)
- Update `.env.example` and the Vertex `librechat.example.yaml` examples
- Add parallel tests across tokens, Anthropic llm config, the Bedrock
  parser, and tx pricing

* 🧹 refactor: Centralize Mythos-class detection; address review feedback

- Add `isMythosClassModel` + `MYTHOS_CLASS_FAMILIES` in schemas.ts as the
  single source of truth for the Fable/Mythos family; route every gate
  (adaptive thinking, omit-thinking, omit-sampling, 1M context, prompt cache,
  128K max-output reset/set) through it. A future sibling class is now a
  one-line edit.
- [Codex P2] Exclude Mythos-class from getBedrockAnthropicBetaHeaders: Fable/
  Mythos ship 128K output + fine-grained tool streaming by default, and the
  legacy output-128k-2025-02-19 beta is 3.7-Sonnet-only on Bedrock and risks
  request rejection. They still get adaptive thinking + effort.
- [Copilot] Add Mythos 5 test parity (name variations, cache rates, pinned
  $10/$50) in tx.spec; add Mythos context/max-output/name-match in tokens.spec;
  fix the stale claude-3-7-sonnet-only comment in bedrock.ts.
- Add isMythosClassModel unit tests covering all declared families.

* 📝 docs: Clarify Mythos-class Bedrock requirements; correct beta-omit rationale

Verified live against Bedrock (acct 951834775723, us-west-2):
- anthropic.claude-fable-5 IS a real Bedrock catalog model, INFERENCE_PROFILE-only
  exactly like the existing anthropic.claude-opus-4-7/4-8 and claude-sonnet-4-6
  default entries (refutes the "invalid model id" review claim).
- Mythos-class also requires opting into Anthropic data sharing (Bedrock Data
  Retention API) before invocation.

Changes:
- .env.example: note that Mythos-class (Fable/Mythos) is inference-profile-only on
  Bedrock and needs the data-sharing opt-in.
- bedrock.ts: reword the beta-omit comment to the verified rationale — output-128k /
  fine-grained-tool-streaming are built-in/no-op for the 4.7+ generation, so omitting
  them is lossless (dropped the unverified "Bedrock may reject" wording).

* 🔄 refactor: Reorganize imports in schemas.ts and tx.spec.ts

- Moved `TFeedback` and `Tools` imports to the top of `schemas.ts` for better readability.
- Adjusted import order in `tx.spec.ts` to maintain consistency and improve clarity.
2026-06-09 16:22:39 -04:00
Danny Avila
8fc2314208
🧠 fix: Bound Memory Agent Input (#13606) 2026-06-09 14:38:21 -04:00
Danny Avila
fb87abe773
🧩 feat: Enable Model Spec Subagents (#13598)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
Publish `@librechat/client` to NPM / pack (push) Has been cancelled
Publish `librechat-data-provider` to NPM / pack (push) Has been cancelled
Publish `@librechat/client` to NPM / publish-npm (push) Has been cancelled
Publish `librechat-data-provider` to NPM / publish-npm (push) Has been cancelled
2026-06-08 11:43:08 -04:00
Danny Avila
83bdd3d65d
🌱 feat: Support Soft Default Model Spec (#13554)
* feat: add soft default model spec

* chore: sort ChatRoute imports
2026-06-06 14:20:59 -04:00
Danny Avila
6357ea10c1
🧭 feat: Scope Model Spec Skills (#13522)
* feat: scope model spec skills

* style: format skill catalog limit

* fix: serialize model spec skill resolution

* test: satisfy model spec load config typing

* fix: apply model spec skills to added conversations

* fix: support alwaysApply frontmatter alias

* fix: address model spec skills review
2026-06-05 10:22:02 -04:00
Atef Bellaaj
86fe79c37d
🔗 feat: Add Granular Access Control to Shared Links via ACL System (#13051)
* feat: Add granular access control to shared links via ACL system

* fix(shared-links): preserve isPublic on failed migration grants

Transient ACL failures during auto-migration permanently stranded
links — $unset ran unconditionally, removing the legacy flag that
triggers retry. Now only $unset isPublic after all grants succeed.

* fix(config): skip isPublic unset for failed ACL grants

Bulk migration unconditionally removed isPublic from all links,
even those whose ACL writes failed. Failed links then lost the
legacy marker needed for auto-migration retry. Now tracks failed
link IDs per-batch and excludes them from the $unset step.

Also adds sharedLink to AccessRole resourceType schema enum —
was missing, only worked because seedDefaultRoles uses
findOneAndUpdate which bypasses validation.

* ci(config): add jest config and PR workflow for migration tests

config/__tests__/ specs depend on api/jest.config.js module
mappings but had no dedicated runner. Adds config/jest.config.js
extending api config with absolutized paths, npm test:config
script, and a GitHub Actions workflow triggered by changes to
config/, api/models/, api/db/, or packages/ ACL code.

* fix(permissions): honor boolean sharedLinks config

SHARED_LINKS has no USE permission, so boolean config produced
an empty update payload — gate conditions only matched object
form, making `sharedLinks: false` a no-op on existing perms.

* fix(share): resolve role before creating shared link

Role lookup between create and grant left an orphaned link
without ACL entries if getRoleByName threw — retry then hit "Share already exists" with no recovery path.

* fix: Restore Public ACL Access Checks

* fix: Type Public ACL Lookup

* fix: Preserve Private Legacy Shared Links

* chore: Promote Shared Link Permission Migration

* fix: Address Shared Link Review Findings

* fix: Repair Shared Link CI Follow-Up

* fix: Narrow Shared Link Mongoose Test Mock

* fix: Address Shared Link Review Follow-Ups

* fix: Close Shared Link Review Gaps

* fix: Guard Missing Shared Link Permission Backfill

* test: Add Shared Link Mock E2E

* test: Stabilize Shared Link Mock E2E

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-06-03 14:17:17 -04:00
Danny Avila
2ef7bdfbc2
feat: Immediate Conversation Title Generation (#13395)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Sync Helm Chart Tags / Ignore non-main push (push) Waiting to run
Sync Helm Chart Tags / Sync chart tags (push) Waiting to run
*  feat: Immediate Conversation Title Generation

Generate conversation titles as soon as the request is made (in parallel
with the response, from the user's first message) as the new default,
fixing the #13318 race where a transient /gen_title 404 left new chats
stuck on "New Chat".

- Add per-endpoint `titleTiming` ('immediate' | 'final') to baseEndpointSchema;
  `endpoints.all` acts as the global default, unset = immediate. Resolve via
  a new `resolveTitleTiming` helper (`all` takes precedence).
- Fire title generation in parallel with `sendMessage`; `titleConvo` waits
  (bounded, abortable) for the agent run and titles from the user input only.
  Persist after the conversation row exists; defer `disposeClient` until the
  title settles.
- Expose `titleGenerationTiming` via startup config; `useTitleGeneration`
  fetches eagerly in immediate mode with a bounded 404 retry and never treats
  a transient 404 as final. Skip title queueing for temporary conversations.
- Supersedes #13329 while incorporating its bounded 404-retry.

* 🩹 fix: Address Copilot review findings on title timing

- Guard against an undefined conversationId in addTitle (skip + warn) so the
  gen_title cache key can't collide as `userId-undefined` and saveConvo is
  never called without a conversationId.
- Gate the title `useQueries` on `enabled` so no /gen_title request fires while
  unauthenticated (e.g. after logout) even if the module queue holds IDs.
- Drop the stale `conversationId` param from the titleConvo JSDoc.
- Add a regression test for the undefined-conversationId guard.

* 🧵 fix: Harden immediate-title edge cases from codex review

- Cancel in-flight immediate title generation when the request aborts: thread
  job.abortController.signal through addTitle so pressing Stop on a new chat
  neither consumes the title model nor surfaces a title for a cancelled turn.
- Preserve a locally-applied title when the final SSE event's conversation
  carries no title yet (built before the title was saved), so long immediate-mode
  responses no longer revert the chat to "New Chat" until reload.
- Guarantee one full post-completion gen_title fetch cycle before giving up, so a
  `final`-mode title (generated only after the stream ends) is still fetched under
  a global `immediate` default instead of being stranded.
- Add regression tests for the abort propagation and the undefined-conversationId guard.

* 🔁 fix: Correct title abort, post-completion refetch, and replacement ordering

Follow-up to codex review of the immediate-title fixes:

- Use a dedicated title AbortController instead of `job.abortController`. The
  latter is also aborted by `completeJob` on *successful* completion, which
  cancelled any title slower than a short response. The title is now cancelled
  only on a real user Stop or when the stream is replaced; a completed-then-
  aborted title is discarded (no save, cache cleared) rather than persisted.
- Reset (not remove) the post-completion title query: `resetQueries` refetches
  the mounted observer with a fresh retry budget, whereas `removeQueries` left it
  stuck in its error state, so the promised post-completion cycle never ran.
- Run the job-replacement check before resolving `convoReady`, and on a replaced
  stream cancel/discard the stale title so a discarded prompt can't persist a title.

* 🧷 fix: Tighten title abort ordering and endpoint-level timing resolution

Follow-up to codex review:

- Abort the title controller before resolving `convoReady` on a stopped turn, so
  the title task can't resume and persist before the later abort.
- Cancel the title and unblock its waits on ANY send failure (not just user
  aborts): a preflight/quota failure before the run exists otherwise hangs
  `_waitForRun`, deferring client disposal until the 45s title timeout.
- Resolve `titleTiming` for custom endpoints via `getCustomEndpointConfig`
  (their config lives under `endpoints.custom[]`, not `endpoints[endpoint]`).
- Derive the startup `titleGenerationTiming` via `resolveTitleTiming` for the
  agents endpoint so an endpoint-level `final` (without `endpoints.all`) is honored
  client-side instead of defaulting to immediate and burning eager gen_title polls.

* 🪢 fix: Per-agent title timing and safer abort/replacement handling

Follow-up to codex review:

- Resolve `titleTiming` from the agent's actual endpoint after initialization, so a
  per-endpoint `final` override on a custom/provider endpoint backing an (ephemeral)
  agent is honored instead of always using the `agents` endpoint's value.
- Don't preserve a locally-fetched title on a stopped (unfinished) turn: the server
  cancels and discards that title, so keeping it client-side would diverge from
  server state and leave the stopped chat titled until reload.
- On abort/replacement, only delete the cached title if it still holds THIS task's
  value — a replacement stream shares the `userId-conversationId` key and may have
  already cached its own valid title that must not be removed.

* 🪞 fix: Mirror AgentClient title-config resolution for titleTiming

Per maintainer guidance, keep titleTiming resolution identical to how
`AgentClient#titleConvo` already resolves the endpoint config — `endpoints.all`
is the intended global override and the agent's actual provider endpoint is used:

- Resolve via `endpoints.all ?? endpoints[endpoint] ?? getProviderConfig(endpoint)
  .customEndpointConfig` (was using `getCustomEndpointConfig` directly). Going
  through `getProviderConfig` picks up its case-insensitive fallback for normalized
  provider names (e.g. `openrouter` → `OpenRouter`), so a custom endpoint's
  `titleTiming` is honored like its other title settings.
- Add `titleTiming` to the Azure endpoint schema `.pick()` so
  `endpoints.azureOpenAI.titleTiming` is no longer silently stripped by Zod.

Note: per-endpoint title settings being skipped when `endpoints.all` is present is
the existing, intended global-override behavior — not changed here.

* 🧪 test: Cover useTitleGeneration effect logic (integration)

Adds a deterministic white-box integration test that drives the real hook's
React effects with a controllable react-query surface, locking down the
stateful decisions that previously had no coverage:

- immediate mode fetches a queued conversation while its stream is still active
- final mode gates until the stream completes, then becomes eligible
- success applies the fetched title to the conversation caches
- a 404 while active defers (removeQueries) instead of giving up
- a 404 after completion forces a fresh fetch via resetQueries (post-completion remount)

* feat: Stream immediate title events

* style: Format title SSE handler

* test: Preserve data-provider exports in OAuth mock

* test: Isolate OAuth route API mock

* test: Keep OAuth callback factory capture

* fix: Replay streamed title events on resume

* fix: Honor agents title timing precedence

* style: Format title timing fixes
2026-06-02 16:40:57 -04:00
Danny Avila
8ba0249f1e
🗃️ feat: Retain Agent Files During All-Data Retention (#13477)
* feat: add agent file retention exemption

* refactor: centralize agent file retention policy
2026-06-02 15:04:10 -04:00
Danny Avila
566e20b613
v0.8.6 (#13302)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Publish `librechat-data-provider` to NPM / pack (push) Waiting to run
Publish `librechat-data-provider` to NPM / publish-npm (push) Blocked by required conditions
Publish `@librechat/data-schemas` to NPM / pack (push) Waiting to run
Publish `@librechat/data-schemas` to NPM / publish-npm (push) Blocked by required conditions
Docker Dev Images Build / build (Dockerfile, librechat-dev, node) (push) Waiting to run
Docker Dev Images Build / build (Dockerfile.multi, librechat-dev-api, api-build) (push) Waiting to run
Sync Locize Translations & Create Translation PR / Sync Translation Keys with Locize (push) Waiting to run
Sync Locize Translations & Create Translation PR / Create Translation PR on Version Published (push) Blocked by required conditions
Publish `@librechat/client` to NPM / pack (push) Has been cancelled
Publish `@librechat/client` to NPM / publish-npm (push) Has been cancelled
2026-05-31 17:36:47 -04:00
Danny Avila
62dff69300
🧠 feat: Add Claude Opus 4.8 Support (#13380)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* feat: add Claude Opus 4.8 support

* fix: omit sampling params for Claude Opus 4.8

* fix: flatten Bedrock beta header merge

* fix: strip Bedrock sampling params for Opus 4.8
2026-05-28 13:50:39 -07:00
Danny Avila
05a3d1ed81
🛣️ feat: Add MCP Remote Proxy Support (#13076)
* feat: add MCP remote proxy support

* fix: Harden MCP Proxy Review Findings

* fix: Honor MCP Proxy Env Precedence

* fix: Harden MCP proxy routing

* fix: Align MCP proxy bypass semantics

* test: Pin MCP proxy admin scope
2026-05-21 15:28:54 -04:00
Danny Avila
9dd062e42e
🧯 fix: Harden Data Retention Semantics (#13049)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* feat: support data retention for normal chats

Add retentionMode config variable supporting "all" and "temporary" values.
When "all" is set, data retention applies to all chats, not just temporary ones.
Adds isTemporary field to conversations for proper filtering.

Adapted to new TS method files in packages/data-schemas since upstream
moved models out of api/models/.

Based on danny-avila/LibreChat#10532

Co-Authored-By: WhammyLeaf <233105313+WhammyLeaf@users.noreply.github.com>
(cherry picked from commit 30109e90b0)

* feat: extend data retention to files, tool calls, and shared links

Add expiredAt field and TTL indexes to file, toolCall, and share schemas.
Set expiredAt on tool calls, shared links, and file uploads when
retentionMode is "all" or chat is temporary.

(cherry picked from commit 48973752d3)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: lint/test

(cherry picked from commit 310c514e6a)

* fix: address code review feedback for data retention PR

Critical:
- Fix BookmarkMenu crash: restore optional chaining on conversation
- Fix migration hazard: backward-compatible sidebar filter that also
  checks expiredAt for documents without isTemporary field

Major:
- Add logging to getRetentionExpiry error path, align with tools.js
- Add tests for retentionMode: ALL in saveConvo and saveMessage
- Fix share route: apply expiredAt for temporary chats too by
  querying the conversation's isTemporary flag server-side
- Add assertions for getRetentionExpiry mocks in process tests

Minor:
- Fix ChatRoute isTemporaryChat to be strictly boolean via Boolean()
- Fix stale test description (expired -> temporary)
- Comment out retentionMode default in example yaml
- Simplify verbose if/else to isTemporary === true
- Add compound index on { user: 1, isTemporary: 1 }
- Remove narrating comment from process.spec.js

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6bad535f90)

* chore: fix typescript

(cherry picked from commit 826527a46b)

* fix: lint

(cherry picked from commit 77817e80ea)

* fix: use mockSanitizeArtifactPath in retention test

The 'getRetentionExpiry is called with the request object' test
referenced an undefined `mockSanitizeFilename` identifier, breaking
both lint (no-undef) and the test suite. Use the existing
`mockSanitizeArtifactPath` mock that the surrounding tests already
use, since `processCodeOutput` calls `sanitizeArtifactPath` (not
`sanitizeFilename`) before invoking `getRetentionExpiry`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
(cherry picked from commit 52ea2da66d)

* fix: forward isTemporary from client for retention on file uploads and tool calls

Server-side `getRetentionExpiry` (file uploads) and the tool-call
controller both read `req.body.isTemporary`, but the file upload
multipart form and the tool-call payload did not include that field.
In `retentionMode: temporary` (default), files uploaded and tool
calls created from temporary chats were therefore retained
indefinitely.

Forward the Recoil `isTemporary` flag in both client paths so the
existing server checks can fire correctly. `ToolParams` gains an
optional `isTemporary` field.

Addresses Codex P1 review feedback on PR #29.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
(cherry picked from commit 7e937df05a)

* test: stub store.isTemporary in useFileHandling test mocks

Previous commit added `useRecoilValue(store.isTemporary)` to the
hook. The test file mocks `~/store` with only `ephemeralAgentByConvoId`
and does not stub `useRecoilValue`, so all 7 cases threw
"Invalid argument to useRecoilValue: expected an atom or selector but
got undefined". Add a stub default export with `isTemporary` and a
`useRecoilValue` mock returning `false`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
(cherry picked from commit eb1609537d)

* fix: harden data retention semantics

* fix: provide sweep request context for expired files

* fix: preserve temporary flags in all-retention updates

* fix: honor assistant versions in retention sweeps

* fix: retain non-temporary flags in all mode

* fix: hide expired retained records

* fix: propagate retained conversation expiry

* fix: refresh meili retention cutoff

* fix: prevent overlapping file sweeps

* fix: show legacy retained conversations

* fix: index legacy retained records

* fix: harden retention cleanup edge cases

* fix: count failed file storage sweeps

* fix: preserve legacy temporary retention

* fix: assign retention sweep worker deterministically

* fix: hide expired shared links on reads

* fix: prevent retention refresh after parent expiry

* fix: break code output retention import cycle

* fix: harden retention review findings

* fix: ignore expired share duplicates

* fix: reject expired retained share creation

* fix: harden retention review edge cases

* fix: address retention audit findings

* fix: enforce expired conversation shares in all retention

* fix: scope temporary upload flag to chat files

* fix: address retention review findings

* fix: address codex retention review findings

* fix: tighten missing storage detection

* test: remove unused file process spec bindings

---------

Co-authored-by: WhammyLeaf <233105313+WhammyLeaf@users.noreply.github.com>
Co-authored-by: Aron Gates <aron@muonspace.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-05-19 21:58:42 -04:00
Danny Avila
738ed005b6
🏷️ feat: Hide Model Spec Badge Rows (#13124)
* feat: hide model spec badge row

* chore: import order

* feat: hide model spec badge row
2026-05-14 09:39:55 -04:00
Danny Avila
68d80f3324
v0.8.6-rc1 (#13094) 2026-05-12 21:40:23 -04:00
Danny Avila
c3ec23f9b8
🌐 feat: Support Vertex AI Multi-Region Endpoints (#13044)
* feat: support Vertex AI multi-region endpoints

* fix: sync Vertex endpoint with final location
2026-05-10 13:41:58 -04:00
Danny Avila
9441563b95
🛡️ refactor: Scope allowedAddresses By Port (#13022)
* fix: Scope allowedAddresses by port

* test: Fix SSRF agent spec typing
2026-05-08 12:28:34 -04:00
Danny Avila
1bc2692a15
🌥️ feat: Add Optional Region-aware S3/CloudFront Storage Keys (#12987)
* feat(files): add optional region-aware storage keys

* test(files): fix region storage CI fixtures

* feat(files): finalize inline CloudFront asset namespaces

* fix(files): allow wildcard region CloudFront cookies

* fix(files): preserve legacy storage key compatibility

* fix(files): align CloudFront clear cookie cleanup

* fix(files): clear legacy CloudFront cookie scopes

* chore(files): clean up storage review nits

* fix(files): keep inline namespaces CloudFront-only
2026-05-06 23:16:56 -04:00
Danny Avila
9c81792d25
🔐 feat: Add Signed CloudFront File Downloads (#12970)
* feat: add signed CloudFront downloads

* fix: preserve local IdP avatar paths

* fix: address signed download review findings

* fix: harden CloudFront cookie scope validation

* fix: preserve URL save API compatibility

* fix: store CDN SSO avatars under shared prefix

* fix: Harden CloudFront tenant file access

* fix: Preserve CloudFront download compatibility

* fix: Address CloudFront review follow-ups

* fix: Preserve file URL fallback user paths

* fix: Address download review hardening

* fix: Use file owner for S3 RAG cleanup

* fix: Address final download review nits

* fix: Clear stale avatar CloudFront cookies

* fix: Align download filename helpers with dev

* fix: Address final CloudFront review follow-ups

* fix: Stream S3 URL uploads

* fix: Set S3 stream upload length

* fix: Preserve download metadata filepath

* fix: Avoid remote content length for stream uploads

* fix: Use bounded multipart URL uploads

* fix: Harden S3 filename boundaries
2026-05-06 19:48:30 -04:00
Atef Bellaaj
187ab787da
🌩️ feat: CloudFront CDN File Strategy (#12193)
* 🌩️ feat: CloudFront CDN File Strategy + signed cookies

Squashed from PR #12193:
- feat(storage): add CloudFront CDN file strategy
- feat(auth): add CloudFront signed cookie support

Note: package.json/package-lock.json dependency additions are intentionally
omitted from this commit and will be re-added via `npm install` after rebase
to avoid lock-file merge conflicts. The two new peer deps that need to be
re-installed are:
  - @aws-sdk/client-cloudfront@^3.1032.0
  - @aws-sdk/cloudfront-signer@^3.1012.0

Also fixes 4 missing destructured names in AuthService.spec.js
(getUserById, generateToken, generateRefreshToken, createSession) that
were referenced in tests but not imported from the mocked '~/models'.

* 📦 chore: install CloudFront SDK deps for PR #12193

Adds the two AWS CloudFront packages required by the rebased
CloudFront CDN strategy:
  - @aws-sdk/client-cloudfront
  - @aws-sdk/cloudfront-signer

Following the @aws-sdk/client-s3 pattern:
  - api/package.json: regular dependency (runtime resolution)
  - packages/api/package.json: peerDependency

Generated by `npm install` against the freshly rebased lock file
to avoid the merge conflicts that came from the original PR's
lock-file edits being made against an older base of dev.

* 🐛 fix: CI failures + review findings on CloudFront PR #12193

CI fixes
- Rename packages/data-provider/src/__tests__/cloudfront-config.test.ts
  → src/cloudfront-config.spec.ts. Jest's default testMatch picks up
  __tests__/ directories even inside dist/, so the compiled .d.ts shell
  was being executed as an empty test suite. Moving to .spec.ts (matching
  the rest of the package) avoids the dist/ pickup.
- Add cookieExpiry: 1800 to CloudFront crud.test makeConfig: the schema
  applies a default so CloudFrontFullConfig requires it.

Review findings addressed
- #1 (Codex + comprehensive): Normalize CloudFront domain with /\/+$/
  regex (and key with /^\/+/ regex) in buildCloudFrontUrl, matching the
  cookie code so resource policy and file URLs stay aligned even when
  the configured domain has multiple trailing slashes. Added tests.
- #2: Move DEFAULT_BASE_PATH out of s3Config into shared
  packages/api/src/storage/constants.ts. ImageService no longer imports
  S3-specific config.
- #3: getCloudFrontConfig() returns Readonly<CloudFrontFullConfig> | null
  to discourage mutation of the cached signing config.
- #4: Add cross-field refinement tests for cloudfrontConfigSchema
  (invalidateOnDelete-without-distributionId,
  imageSigning="cookies"-without-cookieDomain).
- #6: Revert unrelated MCP comment re-indentation in
  librechat.example.yaml.
- #7: Add azure_blob to the strategy list comment.

Skipped
- #5 (extractKeyFromS3Url with CloudFront URLs): existing
  deleteFileFromCloudFront tests already cover the path-equivalence
  assumption; renaming the helper is real refactor work beyond this
  PR's scope.
- #8, #9 (NIT, low confidence): leaving for author judgement.

* 🧹 chore: drop dead DEFAULT_BASE_PATH from s3Config test mock

After moving DEFAULT_BASE_PATH to ~/storage/constants, crud.ts no longer
reads it from s3Config — so the entry in the s3Config jest mock was
misleading dead config. The tests still pass because the unmocked real
constants module provides the value.

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-05-05 13:21:05 -04:00
Danny Avila
2c4a78094a
🛂 refactor: Avoid Default Tavily Safe Search (#12939) 2026-05-04 12:01:23 +09:00
Yashwanth Alapati
3da1d8c961
🔍 feat: add Tavily as Search and Scraper Provider (#12581)
* feat: add Tavily integration as search provider and scraper provider

* chore:update tavily web search parameters

* chore:tavily paramer update

* chore:update data-schemas test for tavily

* fix: allow Tavily string option modes

* fix: align Tavily config options

* fix: scope Tavily scraper timeout

* fix: use resolved scraper provider timeout

* fix: widen Tavily search provider types

* fix: harden Tavily web search config

* fix: cap Tavily option timeouts

---------

Co-authored-by: Danny Avila <danny@librechat.ai>
2026-05-04 11:29:13 +09:00
Danny Avila
4cce88be42
🪟 feat: Add allowedAddresses Exemption List For SSRF-Guarded Targets (#12933)
* 🪟 feat: Add allowedAddresses Exemption List For SSRF-Guarded Targets

LibreChat already blocks SSRF-prone targets (private IPs, loopback,
link-local, .internal/.local TLDs) at every server-side fetch site
that consumes user-controllable URLs — custom-endpoint baseURLs, MCP
servers, OpenAPI Actions, and OAuth endpoints. The only existing
escape hatch is `allowedDomains`, but that flips the field into a
strict whitelist: adding `127.0.0.1` to permit a self-hosted Ollama
also blocks every public destination that isn't in the list.

Introduce `allowedAddresses` as the orthogonal primitive: a private-
IP-space exemption list. When a hostname or its resolved IP appears
in the list, the SSRF block is bypassed for that target. Public
destinations remain reachable. Operators can now run self-hosted
LLMs / MCP servers / Action endpoints on private addresses without
weakening the default-deny posture for everything else.

Schema additions in `packages/data-provider/src/config.ts`:
- `endpoints.allowedAddresses` (new — gates `validateEndpointURL`)
- `mcpSettings.allowedAddresses` (parallel to `allowedDomains`)
- `actions.allowedAddresses` (parallel to `allowedDomains`)

Core changes in `packages/api/src/auth/`:
- New `isAddressAllowed(hostnameOrIP, allowedAddresses)` — pure,
  case-insensitive, bracket-stripped literal match.
- Threaded the list through `isSSRFTarget`, `resolveHostnameSSRF`,
  `isDomainAllowedCore`, `isActionDomainAllowed`, `isMCPDomainAllowed`,
  `isOAuthUrlAllowed`, and `validateEndpointURL`.
- Extended `createSSRFSafeAgents` and `createSSRFSafeUndiciConnect`
  to accept the list, building an SSRF-safe DNS lookup that exempts
  matching hostnames/IPs at TCP connect time (TOCTOU-safe).

Wiring:
- Custom and OpenAI endpoint initialize sites pass
  `endpoints.allowedAddresses` to `validateEndpointURL`.
- `MCPServersRegistry` stores `allowedAddresses` and exposes it via
  `getAllowedAddresses()`. The factory, connection class, manager,
  `UserConnectionManager`, and `ConnectionsRepository` all thread
  it through to the SSRF utilities.
- `MCPOAuthHandler.initiateOAuthFlow`, `refreshOAuthTokens`, and
  `validateOAuthUrl` accept the list and consult it on every URL
  validation along the OAuth chain.
- `ToolService`, `ActionService`, and the assistants/agents action
  routes pass `actions.allowedAddresses` to `isActionDomainAllowed`
  and to `createSSRFSafeAgents` for runtime action calls.
- `initializeMCPs.js` reads `mcpSettings.allowedAddresses` from the
  app config and forwards it to the registry constructor.

Documentation:
- `librechat.example.yaml` shows the new field next to each existing
  `allowedDomains` block, with a note clarifying that
  `allowedAddresses` is an exemption list (not a whitelist).

Tests:
- Unit tests for `isAddressAllowed` covering literal IPs, hostnames,
  IPv6 brackets, case insensitivity, and partial-match rejection.
- Exemption tests for every entry point: `isSSRFTarget`,
  `resolveHostnameSSRF`, `validateEndpointURL`, `isActionDomainAllowed`,
  `isMCPDomainAllowed`, `isOAuthUrlAllowed`.
- Existing tests updated to reflect the new optional parameter.

Default behavior is unchanged: omitted = empty list = no exemptions.

* 🩹 fix: Plumb allowedAddresses Through AppConfig endpoints Type

The initial PR added `endpoints.allowedAddresses` to the
data-provider config schema and consumed it in the endpoint
initialize sites, but the runtime `AppConfig.endpoints` shape in
`@librechat/data-schemas` was a hand-maintained subset that didn't
include the new field — so `tsc` rejected `appConfig.endpoints.allowedAddresses`.

Add the field to `AppConfig['endpoints']` in
`packages/data-schemas/src/types/app.ts` and forward it from the
loaded config in `packages/data-schemas/src/app/endpoints.ts` so the
runtime config carries the value.

Update `initializeMCPs.spec.js` to expect the third positional
argument (`allowedAddresses`) on the `createMCPServersRegistry` call.

* 🩹 fix: Enforce allowedDomains Before allowedAddresses In isOAuthUrlAllowed

The initial implementation checked the address exemption first, so a
URL whose hostname appeared in `allowedAddresses` would return true
even when the admin had configured `allowedDomains` as a strict bound
on OAuth endpoints. A malicious MCP server could advertise OAuth
metadata, token, or revocation URLs at any address the admin had
permitted for an unrelated reason (a self-hosted LLM at `127.0.0.1`,
for example) and pass validation, expanding SSRF reach beyond the
configured domain whitelist.

Reorder: when `allowedDomains` is set, treat it as authoritative —
return true only if the URL matches a domain entry, otherwise fall
through to false. The address exemption only applies when no
`allowedDomains` is configured (mirrors how the downstream SSRF check
in `validateOAuthUrl` consults `allowedAddresses`).

Add a regression test asserting that an `allowedAddresses` entry does
not broaden a configured `allowedDomains` list.

Reported by chatgpt-codex-connector on PR #12933.

* 🩹 fix: Forward allowedAddresses To Remaining OAuth Callers

Two `MCPOAuthHandler` callers still used the pre-feature signatures and
were silently dropping the new `allowedAddresses` argument:

- `api/server/routes/mcp.js` invoked `initiateOAuthFlow` with the old
  5-argument shape, so OAuth flows initiated through the route handler
  ignored the registry's `getAllowedAddresses()` and would reject any
  metadata/authorization/token URL on a permitted private host.
- `api/server/controllers/UserController.js#maybeUninstallOAuthMCP`
  invoked `revokeOAuthToken` without the address exemption, so
  uninstalling an OAuth-backed MCP server on a permitted private host
  would fail at the revocation step even though the rest of the MCP
  connection path now permits it.

Both sites now read `allowedAddresses` from the registry alongside
`allowedDomains` and forward it. Reported by Copilot on PR #12933.

* 🩹 fix: Update Test Mocks And Assertions For OAuth allowedAddresses

The previous commit started passing `allowedAddresses` to
`MCPOAuthHandler.initiateOAuthFlow` from `api/server/routes/mcp.js`
and to `MCPOAuthHandler.revokeOAuthToken` from
`api/server/controllers/UserController.js`, but the corresponding
test files mocked the registry without `getAllowedAddresses` (causing
`TypeError`s) and asserted the old positional shape on
`toHaveBeenCalledWith`.

Update the mocks and assertions to match the new arity:

- `api/server/routes/__tests__/mcp.spec.js`: add
  `getAllowedDomains`/`getAllowedAddresses` to the registry mock and
  expect the additional positional args on `initiateOAuthFlow`.
- `api/server/controllers/__tests__/maybeUninstallOAuthMCP.spec.js`:
  add a `getAllowedAddresses` mock alongside the existing
  `getAllowedDomains` and seed it in `setupOAuthServerFound`.
- `api/server/controllers/__tests__/UserController.mcpOAuth.spec.js`:
  add `getAllowedAddresses` to the registry mock and expect the
  trailing `null` arg on the three `revokeOAuthToken` assertions.

* 🛡️ fix: Address Comprehensive Review — Scope allowedAddresses To Private IP Space

Major findings from the comprehensive PR review (severity → fix):

**CRITICAL — `validateOAuthUrl` SSRF fallback bypass.** When `allowedDomains`
is configured and a URL fails the whitelist, the SSRF fallback in
`validateOAuthUrl` was still passing `allowedAddresses` to `isSSRFTarget` /
`resolveHostnameSSRF`, letting a malicious MCP server advertise OAuth
endpoints at any address the admin had permitted for an unrelated reason.
Suppress `allowedAddresses` in the fallback when `allowedDomains` is active —
the address exemption is opt-in for the no-whitelist mode only.

**MAJOR — WebSocket transport SSRF check ignored exemptions.** The
`constructTransport` WebSocket branch called `resolveHostnameSSRF(wsHostname)`
without `this.allowedAddresses`, so a permitted private MCP server would
pass `isMCPDomainAllowed` but be blocked at transport creation. Forward
the exemption.

**Scope `allowedAddresses` to private IP space only (operator directive).**
The exemption list is for permitting private/internal targets; it must not
be a back-door to broaden trust to public destinations.
- Schema (`packages/data-provider/src/config.ts`): new
  `allowedAddressesSchema` rejects URLs (`://`), paths/CIDR (`/`),
  whitespace, and public IPv4/IPv6 literals at config-load time. Wired
  into `endpoints`, `mcpSettings`, and `actions`.
- Runtime (`packages/api/src/auth/domain.ts`): `isAddressAllowed` now
  drops public-IP candidates and public-IP entries on the match path —
  defense in depth so a misconfigured runtime list never grants exemption.
- Hot path (`packages/api/src/auth/agent.ts`): `buildSSRFSafeLookup`
  pre-normalizes the list into a `Set<string>` once at construction and
  applies the same scoping filter, so the connect-time DNS lookup is an
  O(1) Set membership check instead of a full re-iterate-and-normalize on
  every outbound request.

**Test coverage for the connect-time and OAuth-fallback paths.**
- `agent.spec.ts`: new describe block exercising `buildSSRFSafeLookup` and
  `createSSRFSafe*` with `allowedAddresses` — hostname-literal exemption,
  resolved-IP exemption, public-IP scoping, URL/CIDR/whitespace rejection,
  and the default no-list block.
- `handler.allowedAddresses.test.ts` (new): integration tests for
  `validateOAuthUrl` — covers both the no-domains-set "permit private"
  path and the strict-bound regression where `allowedAddresses` must NOT
  bypass `allowedDomains`.

**Documentation & cleanup.**
- `connection.ts` redirect SSRF check: explicit comment that
  `allowedAddresses` is intentionally NOT consulted for redirect targets
  (server-controlled, must not inherit the admin's exemption).
- `MCPConnectionFactory.test.ts`: replaced an `eslint-disable` with a
  proper `import { getTenantId } from '@librechat/data-schemas'`. The
  disable was added to make a pre-existing `require()` quiet — the cleaner
  fix is to use the existing top-level import.

Updated `MCPConnectionSSRF.test.ts` WebSocket SSRF assertions to match the
new two-argument call shape (`hostname, allowedAddresses`).

* 🩹 fix: Require Absolute URL Before allowedAddresses Trust Bypass In isOAuthUrlAllowed

`parseDomainSpec` is lenient — it silently prepends `https://` to
schemeless inputs so it can match patterns like bare `example.com`.
That leniency leaked into `isOAuthUrlAllowed`'s new
`allowedAddresses` short-circuit: a value like `10.0.0.5/oauth` (no
scheme) would parse successfully via the prepended default, hit the
address-exemption path, return `true`, and skip `validateOAuthUrl`'s
strict `new URL(url)` parse-or-throw — only to fail later in OAuth
discovery with a less clear runtime error.

Add a strict `new URL(url)` gate at the top of `isOAuthUrlAllowed`.
Schemeless inputs now fall through to `validateOAuthUrl`'s explicit
"Invalid OAuth <field>" rejection. Tests added in both
`auth/domain.spec.ts` (unit) and the OAuth handler integration spec
(end-to-end).

Reported by chatgpt-codex-connector (P2) on PR #12933.

* 🛡️ fix: Address Follow-Up Comprehensive Review — Schema Tests, Shared Normalization, host:port

Auditing the second comprehensive review:

**F1 MAJOR — schema validation untested.** `allowedAddressesSchema` had
zero coverage, so a regression in the three refinement stages or the
three wiring locations (`endpoints` / `mcpSettings` / `actions`) would
silently let invalid entries reach the runtime. Added a dedicated
`describe('allowedAddressesSchema')` block in `config.spec.ts` covering:
valid private IPs (v4 + v6, including the previously-missed 192.0.0.0/24
range), accepted hostnames, all rejection categories (URLs, CIDR, paths,
whitespace tabs/newlines, host:port, public IP literals), and full
`configSchema.parse()` integration at each of the three nesting points.

**F2 MINOR — `isPrivateIPv4Literal` divergence.** The schema reimpl in
`packages/data-provider` was discarding the `c` octet, so the
`192.0.0.0/24` (RFC 5736 IETF protocol assignments) range that the
authoritative `isPrivateIPv4` accepts was being rejected with a
misleading "public IP" error. Destructure `c` and add the missing range
check; covered by the new schema tests.

**F3 MINOR — DRY violation across `domain.ts` and `agent.ts`.** Both
files had independent normalization implementations with a subtle
whitespace-check divergence (`/\s/` vs `.includes(' ')`). Extracted the
shared logic into a new `packages/api/src/auth/allowedAddresses.ts`
module that both consumers import:
  - `normalizeAddressEntry(entry)` — single-entry shape check
  - `looksLikeHostPort(entry)` — host:port detector (used by F4)
  - `normalizeAllowedAddressesSet(list)` — pre-normalized Set for the
    connect-time hot path
  - `isAddressInAllowedSet(candidate, set)` — membership check that
    enforces private-IP scoping on the candidate

Both `isAddressAllowed` (preflight) and `buildSSRFSafeLookup` (connect)
now go through the same primitives; the whitespace divergence is gone.

To break the import cycle (`allowedAddresses` needs `isPrivateIP`,
`domain` previously owned it), extracted IP private-range detection
into a leaf `auth/ip.ts` module. `domain.ts` re-exports `isPrivateIP`
for backward compatibility with existing call sites.

**F4 MINOR — `host:port` silently misclassified.** Entries like
`localhost:8080` previously slipped through the URL/path guard, were
mis-detected as IPv6, failed `isPrivateIP`, and were silently dropped
with a misleading "public IP" schema error. Added an explicit
`looksLikeHostPort` check with a clear error: "allowedAddresses
entries must not include a port — list the bare hostname or IP only."
Bare `::1`, `[::1]`, and other valid IPv6 literals are intentionally
not matched (regex distinguishes by colon count and the bracketed
`[ipv6]:port` form).

**F5 MINOR — hostname-trust documentation gap.** Hostname entries
short-circuit `resolveHostnameSSRF` before any DNS lookup — that's a
deliberate design (admin trusts the name) but it means the exemption
follows whatever the name resolves to at runtime. Added an explicit
note in `librechat.example.yaml` for both `mcpSettings.allowedAddresses`
and `endpoints.allowedAddresses`: "a hostname entry trusts whatever IP
that name resolves to. Only list hostnames whose DNS you control.
Prefer literal IPs when you can."

**F6** (8 positional params) is flagged for follow-up; refactor to an
options object is a breaking-API change deferred to a separate PR.
**F7** (redirect/WebSocket asymmetry, NIT, conf 40) — skipping; the
existing inline comment is sufficient.

* 🧹 chore: Address Follow-Up NITs — Import Order And Mirror-Function Naming

Three NITs from the latest comprehensive review:

**NIT #1 (conf 85) — local import order.** AGENTS.md requires local
imports sorted longest-to-shortest. Both `domain.ts` and `agent.ts`
had `./ip` (shorter) before `./allowedAddresses` (longer). Swapped.

**NIT #2 (conf 60) — missing cross-reference.** The schema-side
`isHostPortShape` in `packages/data-provider/src/config.ts` had no
note pointing at the canonical runtime mirror. Added a JSDoc paragraph
explaining the mirror relationship and why a local copy exists (the
data-provider package can't import from `@librechat/api` without
creating a circular dependency).

**NIT #3 (conf 50) — naming inconsistency.** Renamed
`isHostPortShape` → `looksLikeHostPort` so the schema mirror matches
the runtime helper exactly. Kept as a separate function (not a shared
import) for the same circular-dependency reason; the matching name
makes it obvious they should stay in lockstep.
2026-05-03 21:43:59 -04:00