Commit graph

4984 commits

Author SHA1 Message Date
Danny Avila
fe71ffdf42
🛂 ci: Grant Multi-Convo Permission in E2E Mock Config (#14839)
`agent-skills-added.spec.ts` drives the composer's `+` command, which opens the
added-model popover. That path is gated on MULTI_CONVO.USE:

    if (!hasMultiConvoAccess || !plusCommandEnabled || isAssistantsEndpoint(endpoint)) return;

The mock config never sets `interface.multiConvo`, so the permission falls
through to the seeded role default and `handlePlusCommand` returns before
opening the popover. The spec then fails on a popover that is absent from the
DOM entirely, which reads as a selector or timing problem rather than a missing
permission.

Set it explicitly, the same way `contextCost` is set just above for the usage
gauge — the mock config's job is to make each exercised feature's gate explicit
rather than inherit a default.
2026-08-15 10:47:35 -04:00
Danny Avila
88747f0ad8
🩺 fix: Render Stopped Run Steps From Explicit Status (#14871)
* 🩺 fix: Render Stopped Run Steps From Explicit Status

Tool calls decided "still running" vs "stopped" with a whole-message
heuristic:

  const cancelled = !isSubmitting && progress < 1 && !hasError;

That inference cannot tell which step actually stopped. An aborted step
keeps spinning while `isSubmitting` is still true, and when submitting
ends, every unfinished part flips to "Cancelled" at once regardless of
which one died.

`@librechat/agents` v3.4.6+ emits `on_run_step_closed`, a terminal
per-step signal carrying `status` and timestamps — including for steps
swept at end-of-run because the caller aborted. The pinned 3.5.1 already
ships it; nothing consumed it.

- `StepEvents.ON_RUN_STEP_CLOSED` plus `RunStepClosedEvent` /
  `RunStepStatus` types mirroring the SDK payload.
- `PartMetadata.runStepStatus` — a dedicated field, since `status` is
  already claimed by activity-label and question-form parts.
- Server handler forwards the event without the visibility gating the
  other step handlers apply: a step whose open reached the client must
  get its close, or the client is left inferring again.
- `useStepHandler` writes the terminal status onto the tool call part.
- Both decision points (`ToolCall`, the shared `useToolCallState`) prefer
  explicit status, keeping the heuristic as fallback for messages saved
  before this and endpoints that do not emit the event.

Threaded through the five cards sharing `useToolCallState`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

* 🩹 fix: Address Codex Review On Run Step Closure Rendering

- Persist the terminal status server-side. The handler emitted the
  closure without folding it into `contentParts`, so the status existed
  only on the live React message: a reload or resumable reconnect
  dropped it and fell back to the very heuristic this fixes. Now stamped
  onto the aggregated tool-call part (via `stepMap`, falling back to the
  event's own index) before forwarding.

- Honor terminal status independently of output parsing. Gating on
  `hasError` meant a `failed` step with unparseable output rendered as
  "cancelled", while a `failed`/`cancelled` step whose output did parse
  as an error was not terminal at all and shimmered indefinitely when no
  completion event arrived. A closed step now forces progress complete
  and reports `failed` as an error state on its own authority.

- Pass the status to the second `BashCall` branch, which rendered the
  same updated component without it.

- Reuse `Agents.RunStepClosedStatus` in `PartMetadata` instead of
  redeclaring the union, so a future SDK status cannot diverge between
  the event and the persisted part.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

*  fix: Replay Closed Status On Redis Resume, Announce Failures

- Apply closure events during Redis reconstruction. The stamp added in
  the previous commit mutates only the originating process's in-memory
  `contentParts`; a resumable reconnect landing on another replica
  rebuilds from `RedisJobStore.getContentParts`, whose allowlist omits
  `on_run_step_closed`. The status was therefore absent from the sync
  snapshot and, being snapshot-covered, never redelivered as pending —
  so multi-replica resume fell back to the whole-message heuristic.
  Handled as a host-authored event alongside `on_steer_applied` and
  `on_activity_label`, since the SDK aggregator has no notion of it.

- Announce terminal failures in the live region. Forcing terminal
  progress for a closed step meant a `failed` tool reached the
  `aria-live` region through `getFinishedText()`, which only special-
  cased cancellation and otherwise announced "completed function" —
  telling screen-reader users the opposite of what the card showed. A
  regression introduced by the previous commit; error states now
  announce failure before any completion string.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

* 🎯 fix: Resolve Closed Steps By ID, Never By Index

The steer and HITL offset wrappers clone and shift only `ON_RUN_STEP`
and `ON_AGENT_UPDATE`; every other event passes through untouched. A
stored `on_run_step_closed` therefore carries the SDK's unshifted index,
while the part it belongs to was rebuilt at the shifted one. Any run
containing a steer insertion or HITL resume would stamp the status onto
an earlier tool card, or none — leaving the real card on the fallback
heuristic while mislabeling a different one.

- Redis reconstruction builds a step ID -> index map from the replayed
  `on_run_step` payloads (which carry the shifted index) and resolves
  closures against it, mirroring what the live callback does via
  `stepMap`.

- The live handler drops its `?? data.index` fallback for the same
  reason. Skipping is the safe failure: a missing status degrades to the
  old heuristic, whereas a misplaced one actively mislabels the wrong
  card.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-15 10:14:16 -04:00
Danny Avila
cd4511038d
🚏 feat: Central Trace Destination Opt-Out for Langfuse (#14838)
* feat(langfuse): let callers opt out of the central trace destination

Adds `centralTraceExportEnabled` to the score-destination options and threads
it through `getScoreDestinations`, `getLangfuseTraceDestinationIds` and
`getLangfuseTraceMessageFields`.

Deployments that route traces per tenant may want a given turn's spans to reach
only the tenant destination — for example when central export is a per-tenant
setting rather than a deployment-wide one. Today the central project is included
whenever env credentials exist, with no way for a caller to decline it for a
single trace.

Defaults to `true` everywhere, so existing callers are unaffected: the option is
additive and every current call site resolves exactly as before. The three
public helpers gain an optional trailing parameter and nothing else.

While here, `getScoreDestinations` destructures its options with defaults
instead of repeating `options?.waitForCentralProjectId !== false` at both call
sites, which is what made adding a second flag awkward.

Verified: no new tsc errors (one pre-existing cacheFactory error is unchanged),
99 langfuse tests pass, ESLint and Prettier clean.

* fix(langfuse): keep the central opt-out intact when destinations resolve

Addresses two review findings on the new `centralTraceExportEnabled` option.
Both are cases where opting out of central export was silently discarded by
destination resolution, letting later feedback reach a project the trace never
went to.

1. Non-fanout deployments with no central env credentials still returned the
   configured connection, because only the central-credential branch was gated.
   `resolveLangfuseExportPlan` reports `disabled` for that shape — without a
   fanout route there is nowhere for a central-suppressed trace to go — so
   return no destinations and match it.

2. `getLangfuseTraceDestinationIds` returned `undefined` whenever any
   destination lacked an id, and `sendFeedbackScore` reads `undefined` as
   unrestricted. A tenant destination has no id when its optional `projectId`
   is unset, so a suppressed-central trace could resolve back to the central
   project at feedback time. Fail closed with an empty list, which stays
   restricted, instead.

Both paths now have regression tests, each verified to fail without the
corresponding change: the first returned 3 destination ids, the second returned
`undefined`.

* fix(langfuse): carry the central opt-out into the feedback path

`getLangfuseTraceDestinationIds` returns `undefined` when an eligible
destination has no stable id, which a tenant route hits whenever the
optional `langfuse.projectId` is unset. `sendFeedbackScore` read that as
"unrestricted" and re-resolved destinations with central export enabled,
so a suppressed-central trace still drew central feedback.

Persisting an empty list instead only traded the leak for a drop: the
destination filter rejects every id-less destination, discarding ratings
the tenant should receive. The id list restricts feedback to destinations
that survived reconfiguration; it cannot also encode deployment policy.

Thread `centralTraceExportEnabled` through `sendFeedbackScore` so the
policy is evaluated the same way at trace time and feedback time, and
restore `undefined` for unidentifiable destinations.
2026-08-15 08:41:18 -04:00
Marco Beretta
c4357fc9e3
📐 feat: Match the Message Column to the Composer (#14851)
* feat: Match the message column to the composer width

Give messages the same max-width and horizontal padding as ChatForm,
reserve the scrollbar gutter on the composer wrapper, and drop the
65ch prose cap so the body fills that column.

* style: Drop the assistant avatar gutter

Keep the icon and provider name on the same left edge as the message
body. Mid-message author headers and steer bubbles no longer outdent
past a column that no longer exists.

* style: Reveal the timestamp on the message header bar

Put the icon, provider name, and datetime on one full-width row, and
show the time only when the message is hovered or focused. CSS on
.message-render wins the hover hide that Tailwind group-hover lost.

* style: Align scroll-to-bottom with the chat column

Sit the control in the same padded column as the composer, swap the
hard-coded disc for a themed outline Button, and fade it in on an
8px rise instead of a scale pop.

* feat: Crossfade the provider name to the model on hover

Swap the assistant header label to the real model name when the
provider is hovered or focused. Skip agent_* document ids so the
hover text is only a model name.

* fix: Reserve the message column gutter without clipping the composer

`scrollbar-gutter: stable` only holds its band back while the element is a
scroll container, and a scroll container clips. Wrapping the composer in one
put the in-flight steer overlay outside the clip: it is painted above the
composer's top edge, so for the whole run a submitted steer was invisible and
its cancel unreachable. The scroll-to-bottom wrapper had the same problem in
a smaller way, cutting off the button's focus ring.

Reserve the same band with padding instead, sized by the width the app already
gives its own scrollbars, so both columns still line up with the messages
without either becoming a scroll container.

Also scopes the header label crossfade to the two labelled spans, and drops
its `:focus-within` rules, which no focusable descendant can ever trigger.

* fix: Keep document ids out of the header and name the model to screen readers

An Assistants-endpoint message keys the assistant map by `assistant.id`, so
its `model` field holds an `asst_` id, not a model name. The header label only
skipped `agent_`, so it crossfaded the assistant's name into an internal id.
Skip both prefixes, and offer `assistant.model` ahead of the message field in
the callers that already resolved the assistant.

The crossfade itself is pointer-only: the model span is `aria-hidden` and
nothing in the label can take focus, so keyboard and screen reader users had
no path to the value at all. Carry the model in text that never hides, which
puts it in the header's accessible name alongside the author and the time.

* refactor: Own the header crossfade in the component

The provider-to-model crossfade lived in global CSS even though HeaderLabel
is its only consumer. Tailwind expresses the whole effect: a named group for
the hover scope, one grid cell shared by both labels, and the existing resize
duration and easing variables. Reduced motion now follows the same
motion-reduce convention as the rest of the client.

* fix: Return a defined model name from the header lookup

Array.find over nullable candidates widens the return to include null, which
tsc rejects against the declared string | undefined. Narrow with a predicate
and sort the imports the pre-commit hook rewrote.

* fix: Keep the model reachable by keyboard and the scroll button inert

The header crossfade was pointer-only, so a sighted keyboard user never saw
the model name; the screen-reader copy covered announcement but not sight.
Focusing anything in the message row now swaps the label too, the same hook
the timestamp already reveals itself with.

The scroll-to-bottom wrapper spans the column and stays inert so it never
swallows clicks meant for the thread, which left the button to opt back into
pointer events. A descendant that opts in is hit-testable however its parent
paints, so the transition classes could not hold the control inert as they
claimed: the button took clicks while invisible. Gate the opt-in on the enter
transition settling and drop the declarations that never applied.

* fix: Measure the scrollbar gutter and keep the scroll button unreachable

The spacer assumed the gutter was the `::-webkit-scrollbar` width. Blink and
WebKit honour that rule, Firefox ignores it and sizes the band itself, and an
overlay scrollbar reserves nothing at all, so on those the composer and the
scroll-to-bottom control sat off the messages they are supposed to line up
with. Measure what the message column actually holds back and publish it for
the spacer to read, leaving the token as the pre-measurement fallback.

Gating the scroll button on pointer events alone also left it enabled, so it
kept its place in the tab order and answered Enter while invisible. Disable it
until the same gate opens, and hold its opacity so being briefly unreachable
does not dim it on top of the wrapper's own fade.
2026-08-15 07:19:34 +02:00
Marco Beretta
73812adca1
📐 fix: Size the Model Picker to Its Content (#14859)
The header trigger was w-full inside a max-w-md wrapper, so it rendered
as a fixed 448px pill no matter how short the model name was. Let the
wrapper shrink-wrap the button and cap it at 60vw (20rem from sm up), so
the pill hugs its label and long names truncate instead of claiming the
whole row.
2026-08-15 06:21:36 +02:00
Marco Beretta
8d99fd16fc
🌗 fix: Keep Code Block Header Visible in User Messages (#14856)
* fix: Keep Code Block Chrome Visible on the User Message Bubble

The code bar dropped its background in dark mode and inherited whatever
sat behind it. That works on the chat background, but a user message
bubble is surface-tertiary, which resolves to the same gray-700 as the
code block's border-light outline, so both the header bar and the
outline disappeared into the bubble and only the code body showed.

Paint the bar with surface-secondary instead. It resolves to the same
value as the old pair on the chat background (gray-50 in light,
gray-800 in dark, matching the presentation background), so assistant
messages are unchanged, while the bar keeps a surface of its own inside
the bubble. The execution output panel used the same pattern and gets
the same treatment.

* fix: Derive the Code Block Surface So Custom Themes Keep Their Colors

Painting the code bar with surface-secondary only reproduced the old
appearance because the built-in themes happen to give surface-secondary,
surface-primary-alt and presentation coinciding values. A custom theme
sets those three independently, so the bar could shift on the chat
background where nothing was meant to change.

Add a surface-code role that derives from whatever the bar used to show:
surface-primary-alt in light, where the bar was already opaque, and
presentation in dark, which is exactly what the transparent bar
inherited. Every theme therefore renders the bar as before on the chat
background, while the bubble no longer bleeds through it.

Give ResultSwitcher the same surface. It had no background of its own,
so once the output panel above it gained one it became a detached band
of bubble color, the same defect one element lower.

* fix: Scope the Opaque Code Surface to User Message Bubbles

CodeBar renders on more than the chat background. The terms dialog and
the subagent panel put MarkdownLite on surface-dialog and
surface-primary, and dark:bg-transparent was what let the bar sit on any
of them. Painting it unconditionally gave those surfaces a header that
contrasts with their own background.

Restore the original declarations and scope the opaque role with a
.user-turn ancestor selector, which MessageRow and SteerPart both
already set, so every user bubble is covered without threading context
through react-markdown. Outside a user bubble the classes are byte for
byte what they were, so no other surface and no custom theme can drift.
2026-08-15 05:39:02 +02:00
Danny Avila
530a935a74
🎨 feat: Color the Context Gauge by Category and Collapse its Breakdown (#14855)
* 🎨 feat: Color the Context Gauge by Category and Collapse its Breakdown

The context window bar becomes a stacked meter — one hue per category — and
the breakdown collapses behind a disclosure so the gauge alone is the default
view. The collapse choice persists per user.

Adds a categorical series scale (`rgb-series-1`…`rgb-series-7`) to the
versioned theme registry, so themes and `REACT_APP_THEME_SERIES_*` can retint
it. Hues are anchored on LibreChat's own brand tokens; every step was computed
rather than picked, by enumerating slot orderings and snapping each step until
all gates passed in both modes:

  worst adjacent CVD ΔE          12.4 light / 13.0 dark  (target 8)
  worst adjacent normal-vision   19.0 light / 19.0 dark  (floor 15)
  contrast                       all 14 steps ≥ 3:1 on both the popover
                                 surface and the meter track

Slot order is the colour-vision-deficiency safety mechanism, not cosmetics.
Reserved status colors are never reused for series identity, and the circular
composer gauge is deliberately untouched — it answers "how close am I to the
limit", which stays a status question.

- `SegmentedMeter` + `MeterSwatch` land beside `Progress` in `@librechat/client`,
  owning the 2px surface gaps, rounded ends, the min-width floor, and the hatch.
  The category-to-slot mapping stays feature-local: the palette is theme data,
  the mapping is not.
- Every present category gets a 2px floor so a 251-token row cannot render as
  0.09px; the shortfall comes out of free space, never another category.
- Deferred tools keep their family's hue and add a 135° hatch, so a hue never
  means two things. Segments are reordered to put each deferred pair beside its
  parent, which is also the adjacency the palette was validated on.
- Messages is drawn as a translucent fill with a solid edge: it is the only
  category the user grows, and the form difference doubles as secondary encoding.
- A row carries a swatch if and only if it is a segment. The estimate path knows
  the total but not the composition, so it keeps a single unsegmented fill.
- Usage totals gain a "Totals" heading, and row text lifts to primary ink on
  hover/focus.
- The popover widens 256px → 288px to absorb the chevron and the legend swatches.

Guardrails: the series scale is held to the 3:1 mark floor on both surfaces, the
app CSS defaults are held in step with the runtime themes, and each slot is
asserted to resolve to a Tailwind utility backed by its CSS variable.

* 🐛 fix: Address Codex Review on the Segmented Context Gauge

Three P2 findings, all confirmed.

**Gaps inflated the fill.** Segment widths were percentages of the whole track
while `gap-[2px]` was added on top, so the gaps ate into the free-space
remainder instead of living inside the filled region. Measured on the real
component: a window at 47.2% painted 55.6% full, and the bar read full at ~94%.
Each segment now surrenders its share of the gap budget, so fills plus gaps span
exactly the used fraction. Same case now paints 50.2%.

The residual 3.0pp is the `SEGMENT_MIN` floor doing its job — five sub-pixel
categories rounded up to 2px each. That overshoot is deliberate and bounded, it
comes out of free space rather than a neighbouring category, and the doc comment
now states the magnitude instead of leaving it implicit.

**No reference-theme test.** The suite only exercised the bundled token tables,
so it could not detect the shared component becoming coupled to LibreChat's
values. Adds a deliberately different reference `ThemeDefinition` and asserts the
registry accepts it, the values reach the applied CSS variables, and every
rendered mark takes its colour from those variables — no literal colours in the
tree. `SegmentedMeter.tsx` also joins the shared-primitive colour guardrail.

**Series tokens missing from the public maps.** `IThemeVariables` and
`IThemeColors` are exported for downstream consumers to type their CSS-variable
and Tailwind maps, and would have rejected the new keys. Adds the series entries
to both, plus a compile-time guard in the registry so a slot added to one map and
missed in another fails the build.

The guard deliberately lives in `registry.ts`, not the spec: `tsconfig.json`
excludes `*.spec.ts`, so an assertion there is never checked by the build —
verified by removing a key from each map in turn and confirming the error.

*  fix: Expand the Breakdown in the Context Gauge e2e Specs

`e2e/specs/mock/usage.spec.ts` asserts on rows that now sit behind the
disclosure, so four tests failed on the collapsed default. My miss — I updated
the component spec and never grepped for e2e coverage.

`openBreakdown` now expands the detail after opening, so every caller that
reads a row keeps working; the helper is idempotent, since a reload restores an
already-expanded preference. The one inline `gauge.click()` that duplicated the
helper now uses it.

Adds the case the regression should have been caught by, and which only e2e can
reach: the popover opens to the gauge alone with no detail mounted, expanding
reveals the labelled Totals section, and the choice survives a real reload
through localStorage without a second click.

`e2e/specs/real/usage.spec.ts` reads the totals the same way. It also hovered
rather than clicked, which never opened the popover at all — hover surfaces only
the compact snapshot tooltip, as the mock spec asserts.
2026-08-14 22:46:10 -04:00
Danny Avila
58f0ab7f62
➡️ style: Flush the In-Flight Steer Stack to the Composer Edge (#14854)
The stack carried its own `mx-auto max-w-3xl` cap, a second width constant
beside the composer's. They agree only at `md`. Past that the composer widens
to `xl:max-w-4xl` (or to the full pane under maximized chat space) while the
stack stays pinned at 768px and centered, so the steer bubble drifts inboard —
about 64px short of the composer's right edge at `xl`, and much further when
chat space is maximized.

`inset-x-0` already spans the composer wrapper, so the cap was never adding a
bound the composer didn't already impose; it was overriding one. Dropping it
leaves a single source of truth for the column width, and `p-2` then lands the
send-now arrow in the same column as the composer's own send button (`mr-2`
plus the 1px border). Below `md` nothing moves — the wrapper was already
narrower than the cap.

This puts the stack in the scroll-to-bottom button's lane at every desktop
width, where before `xl` kept them 64px apart by accident. That is what
`steerOverlayHeightFamily` is for: the button offsets by the published overlay
height, measured from the same edge the stack grows from, so it rests 20px
above the top of the stack (#14844).
2026-08-14 22:45:14 -04:00
Marco Beretta
8640bf89ef
fix: copy only assistant response text (#14853)
Restore the pre-#14770 copy path so the message copy button
serializes TEXT parts only and skips tools, thinking, and errors.
2026-08-15 03:24:21 +02:00
Danny Avila
6f05f2427b
👷 ci: Stop Optional Playwright Fonts From Failing E2E (#14852)
`npx playwright install-deps chrome` is the third-most-common e2e failure:
three of the last twenty-five Playwright runs died on it, taking the whole
aggregate gate with them. The step is not installing anything CI needs.

The runner's Chrome is an apt package, so apt has already satisfied every
library Playwright lists — the log shows each one "already the newest version".
All `install-deps` adds are decorative CJK/Thai/Cyrillic font packages, ~21MB
pulled from azure.archive.ubuntu.com by seven jobs on every PR. No CI assertion
depends on them: the only spec that screenshots gates its comparison behind
`E2E_VISUAL_SNAPSHOTS`, which no workflow sets, and no baselines are committed.

Keep the install, but demote it. `google-chrome --version` becomes its own
fatal step so a genuinely missing browser still fails loudly and immediately,
while the font install retries with a per-attempt cap and degrades to a warning.
The Redis install in the list_changed job stays fatal — that one is required.
2026-08-14 20:35:37 -04:00
Danny Avila
af7e890b14
🐛 fix: Give the Header's Sidebar Toggle Its Own Test Id (#14850)
The header now branches on CSS instead of `useMediaQuery`, so its mobile
`OpenSidebar` stays mounted at every breakpoint. The sidebar rail already
publishes `open-sidebar-button` for its own collapsed toggle, so both held
the id at once and `getByTestId` resolved to two elements.

Scope the header's copy to `header-open-sidebar-button` and assert the count
in the spec that broke — the click there failed only once the header had
mounted, so a count assertion pins the collision deterministically.
2026-08-14 20:34:03 -04:00
Danny Avila
0e160d2ba0
📱 style: Consolidate the Mobile Chat Header (#14843)
* ♻️ refactor: Extract `useNewChat` as the Single New-Chat Path

The new-chat sequence (clear the outgoing conversation's cached messages,
invalidate the messages query, reset the conversation atom) existed in
three places: the sidebar's `NewChatButton`, the `newChat` keyboard
shortcut, and an unrendered `Nav/NewChat` component.

Consolidate into `hooks/Chat/useNewChat`. The panel switch stays an
optional `onNewChat` callback rather than living in the hook, because
`useActivePanel` throws outside `ActivePanelProvider` and the chat header
sits outside it — the upcoming header button needs this seam.

`useKeyboardShortcuts` consumes the returned `newConversation` so the file
still instantiates `useNewConvo` exactly once.

Delete `Nav/NewChat`: it was reachable only through its own barrel export,
and carried a stale `max-md:hidden` plus a `data-testid` that collided
with the sidebar's button.

`handleNewChatClick` now also defers on shift-click, so shift-click opens a
new window like any other link.

`ExpandedPanel.spec` mocks the new hook — it reaches `useNewConvo` by deep
path, which escapes the spec's `~/hooks` barrel mock.

* ♻️ refactor: Split Header Action Logic Out of Its Buttons

Lift the behaviour behind the compare and temporary-chat header buttons
into `useMultiConvo` and `useTemporaryChat`, leaving each component as a
thin trigger. The upcoming mobile overflow menu needs the same actions as
menu items, and the visibility rules (assistants have their own comparison
surface; temporary chat can't be toggled mid-thread) have to stay in one
place rather than being restated per surface.

Add the header's new-chat button, consuming `useNewChat`. It renders as an
anchor to `/c/new` so modified clicks still open a tab, and uses a distinct
`data-testid` from the sidebar's button so queries can't match both.

`useTemporaryChat` toggles through a functional updater, dropping the
`useRecoilCallback` that existed only to close over the current value.

No visual change yet — the header layout lands next.

* 📱 style: Fold the Mobile Header Into Four Targets

The mobile header was a horizontally scrolling strip of up to seven
controls, each in its own outlined box, so nothing grouped and nothing
receded. The overflow was hidden rather than solved: ModelSelector alone is
capped at 70vw (273px) and the side clusters need ~130px, which does not
fit a 390px phone.

Mobile now reads: sidebar toggle, model selector, new chat, ellipsis.

Lift the bookmark and export/share menu items into `useBookmarkItems` and
`useExportShare`, each returning the items plus the dialog instance the
surface must render. Both menus already built `MenuItemProps[]` internally,
so the desktop buttons keep their exact markup and simply consume the hook
— the two surfaces cannot drift apart.

`HeaderMenu` composes those with the compare and temporary-chat actions.
Bookmarks nests through `subItems` rather than flattening every tag to the
top level, permission gates decide membership, and the trigger does not
render when nothing survives — reachable, since export/share self-hides on
a new conversation.

Layout is one DOM order serving both breakpoints; hidden items generate no
flex gap, so each collapses without reordering. Branching is CSS-only:
`useMediaQuery` resolves after paint, and the old
`isSmallScreen ? <OpenSidebar/> : null` popped the row a frame late on
every mount.

`overflow-x-auto` is gone. It hid the overflow instead of fixing it, and it
is a horizontal-swipe sink the later edge-swipe work needs removed.

Presets stays a visible mobile icon for now: `PresetItems` uses Radix's
`Close`, which throws outside a Popover root, so folding it needs a
controlled + anchored menu and browser verification.

Also drops two stray `console.log` calls carried along from the bookmark
mutation handlers.

* 🩹 fix: Address Codex Findings on the Mobile Header Menu

`separate: true` marks an entry as *being* a divider — `DropdownPopup`
returns only a `MenuSeparator` for it and drops the item. Setting the flag
on Share and on temporary chat therefore deleted those actions whenever an
earlier group existed. Push standalone divider entries instead.

The spec missed this because its `DropdownPopup` mock rendered every label
regardless of the flag. It now mirrors the real contract — dividers replace
items, `show: false` entries are dropped — and asserts both actions survive
alongside their dividers.

Gate the bookmark tags query on the bookmark permission. `HeaderMenu`
mounts unconditionally and called `useBookmarkItems` before the permission
result applied, so users without `BOOKMARKS:USE` issued a forbidden request
on every chat header mount; the old header dodged this by mounting
`BookmarkMenu` only after the check.

Restore two states the collapsed trigger had dropped: the shared-link
indicator and its active-link label, and a visible checked state for
temporary chat, which previously only reached assistive tech through
`aria-checked` while the old button switched to `bg-surface-active`.

Compose both new controls from the shared `Button` primitive with the same
override `OpenSidebar` already uses, rather than restating the bordered
icon-button recipe locally.

* 🐛 fix: Give the Overflow Menu's Share Indicator Its Own Test Id

Restoring the shared-link indicator on the mobile trigger reused the id
`ExportAndShareMenu` already owns. Both headers stay mounted and are only
hidden by CSS, so `getByTestId('header-shared-link-indicator')` matched two
elements and `shared-links.spec.ts` failed on a strict-mode violation.

Distinct id, matching the new-chat button, which was already separated from
the sidebar's for the same reason.
2026-08-14 18:14:05 -04:00
Danny Avila
f44ce0bb5d
⬇️ fix: Keep Scroll-to-Bottom Clear of the In-Flight Steer Stack (#14844)
Both elements claim the same corner. `ScrollToBottom` is `bottom-5`,
right-aligned, anchored to the message scroll region. `InFlightSteers` is
`bottom-full`, right-aligned, stacking upward from the composer's top edge.
They overlap at every breakpoint, and because they sit in different
stacking contexts, which one paints on top depends on ancestor DOM order
rather than intent.

The reservation mechanism already exists: `InFlightSteers` measures itself
into `steerOverlayHeightFamily` and `MessagesView` reads it to pad the
thread so the newest message clears the overlay. The scroll button was
never included. Thread that same height through and offset the button by
it — no new state, no second measurement.

Also gives the button a mobile gutter. Its width was `md:max-w-3xl` with no
base value, so on a phone it escaped the message column and pinned to the
viewport edge while the steer bubbles inset by 8px. `px-4 md:px-0` aligns
it with the message content and leaves desktop untouched.

The overlay is capped at `max-h-[35vh]`, so the button can rise at most a
third of the screen.
2026-08-14 16:16:01 -04:00
Danny Avila
b0ed8524d4
fix: avoid modal attachment menu overhead (#14847) 2026-08-14 16:15:47 -04:00
Danny Avila
e1178d3c65
🏘️ fix: Scope OpenID User Cache Keys to Signed User Identity (#14837)
* fix(auth): scope OpenID user cache by tenant

* fix(auth): preserve pre-auth cache scope

* fix(auth): type OpenID reuse secret
2026-08-14 15:51:18 -04:00
Danny Avila
ee21066590
🏎️ ci: Focus Redis E2E Coverage (#14842) 2026-08-14 12:54:30 -04:00
Marco Beretta
336703fe48
🔔 fix: Report Agent Saves That Reuse the Newest Version Entry (#14824)
* fix: report agent saves that reuse the newest version entry

An update whose result matches the newest version is written without
recording a version entry. The Agent Builder derived its success message
from the version count, so every such save reported "No changes were
made" while the edit had in fact been persisted. Base the message on
whether the submission carried an edit of its own instead, and keep the
version count for the version history panel.

Also stop suppressing the version entry when the update carries an
atomic operator. isDuplicateVersion compares direct updates only, so it
cannot speak for the operator half; suppressing there applied a change
that no version entry recorded, leaving the document diverged from every
entry in its own history.

Closes #14809

* fix: count an avatar reset as a persisted edit

An avatar upload uses its own endpoint, but a reset rides the update
payload as avatar: null, so classifying every avatar-only submission as
non-persisted was wrong for resets. Clearing an avatar the newest
version never recorded reads as a duplicate to isDuplicateVersion, since
it skips a field when both sides are falsy, so the reset landed with the
version count unchanged and reported "No changes were made".

* fix: skip the version entry when an atomic operator changes nothing

An update carrying $push, $pull or $addToSet bypassed duplicate suppression on
the operator's mere presence. Re-attaching a resource file an agent already holds
makes $addToSet a no-op, so an agent with actions recorded a version entry for a
write that never touched the document, and its version count climbed on retries.

Resolve the operators against the current document instead. $push always appends
and $pull matches arbitrary criteria, so both still count as mutating; $addToSet
counts only when some value it adds is missing. Whatever cannot be compared
cheaply counts as mutating, since over-reporting costs a redundant version while
under-reporting would apply a change no version records.

* fix: confirm the submitted edit survived before claiming a save changed anything

Treating a dirty form as proof of a persisted edit reports success for a save
that stored nothing. The server can normalize a submission straight back to the
stored value: an MCP tool the user added is dropped when authorization rejects
it, and a skill is pruned when it no longer exists. Neither moves the version
count, so the toast claimed the agent was updated when it was untouched.

Capture the agent as it stands before the write, since the mutation replaces
that cache entry on success, and compare it against the one the server returns
across the fields the submission carried. Keep the dirty check alongside it: an
agent loaded through the basic projection carries fewer fields than the update
endpoint returns, and pairing the two keeps an untouched save honest either way.

* fix: compare a save against the expanded agent, not a basic projection

The panel falls back to the basic agent query whenever the expanded one has not
resolved, and that projection drops instructions, tools, edges, skills and the
rest while reducing model_parameters to a single flag. Comparing a submission
against it made every one of those fields read as changed, so a rejected MCP
tool or a pruned skill still reported success.

Compare against the expanded agent, the only projection carrying every field a
submission sends. When it is unavailable the comparison reports true and leaves
the dirty check to decide, since claiming nothing changed for a save that did is
the worse of the two errors. Renamed to say what it now answers.

* fix: drop the operator a suppressed update judged a no-op

Suppression reads whether $addToSet would add anything from a document fetched
before the write, and that reading cannot bind a concurrent one. A $pull landing
in between leaves the operator re-adding the value while the version entry has
already been suppressed, which is the one outcome this path exists to prevent: a
change applied with nothing in the history recording it.

Drop what was judged a no-op instead of racing it. Only $addToSet reaches here,
and only once every value it adds was found stored, so removing it makes the
suppression true by construction rather than true if nothing else writes first.

* fix: leave a suppressed update carrying no operator at all

Dropping only $addToSet left the invariant resting on which operators callers
happen to send. A present but empty $push or $pull counts as no operator when
deciding suppression, yet survived into the write, so the suppressed update was
operator-free by convention rather than by construction.

Drop all three. Reaching suppression already means none of them can change the
document, so removing them states that outright and keeps the write consistent
with the history it declines to record.
2026-08-14 12:20:51 -04:00
Danny Avila
69ce4b7b00
📱 style: Reclaim the Assistant Avatar Gutter on Mobile (#14836)
Assistant content sat 36px from the left edge on mobile (a 24px avatar
column plus `gap-3`) against a 16px gutter on the right, costing ~13% of
the reading width on every response.

Move the avatar into the assistant heading and restore the gutter as
`md:pl-9` on the content column, so the column only exists from `md` up.
Geometry is unchanged on desktop: 768 - 24 - 12 and 768 - 36 both leave a
732px content box, and an absolutely positioned child resolves against the
padding box, so `md:left-0` lands the avatar where the column started.

`AuthorHeader` and `SteerPart` hard-coded `-ml-9` to reach back past that
column; both are gated to `md` so they no longer outdent off-screen.

Pure `md:` variants rather than `useMediaQuery`, which resolves
desktop-first after paint and would reflow every row on mobile.

Covers all five surfaces sharing `MessageRow`: the three message paths,
the shared-conversation view, and search results.
2026-08-14 11:57:58 -04:00
Danny Avila
5d3edeb383
🪄 feat: Smooth Activity Phase Transitions (#14832)
* feat: Animate activity phase transitions

* style: Match activity phase formatting

* 🪄 fix: Fold activity phase entrance in one direction, flush-left label

The phase header replaced <summary> with <button>, which brought the UA
`text-align: center` with it — the label span is `flex-1`, so the text
filled the row and centered inside it. Left-align it and drop the leading
glyph: the card's border and fill already carry the weight, and the child
tool groups keep their own icons.

The entrance also read as two movements. The card, header and inset all
hard-cut in at full size, displacing the transcript below by ~57px, then
folded back up past the header that had just pushed it down. The card now
mounts in the shape of what was already on screen — zero-height header,
transparent chrome, no inset — and grows the header as the panel collapses,
so the block's height only ever decreases. Chrome, padding and both heights
share one curve.

The collapse also waits for a painted start value; a single rAF can land
before paint, and a start value the compositor never saw snaps rather than
transitions.

- Restore the e2e parent-phase selectors, which still matched `summary`
- Memoize the hoisted `groupActivityPhases` pass and its phase-index set
- Finish the amber -> `text-text-warning` sweep in ToolCallGroup and Part

* 🩹 fix: Scope phase-entrance history and resolve media queries at mount

Addresses both Codex findings on #14832.

`MultiMessage` renders siblings without a key, so `ContentParts` survives a
sibling switch with its refs intact. The recorded phase-marker set outlived
the message it described, and any phase in the newly selected sibling whose
index was absent from the previous sibling's set was read as a live arrival —
already-loaded history mounted expanded and collapsed itself. Scope the set
to its messageId and treat a mismatch as a fresh mount.

`useMediaQuery` initialized to `false` and resolved only in a passive effect,
so the first render always reported "no match". Anything branching once at
mount — the frozen entrance flag here, and every other first-paint decision
across its call sites — never saw the correction, which is how a
`prefers-reduced-motion: reduce` user still got the fold. Read the query
synchronously in the state initializer and guard both paths for environments
without `matchMedia`.

*  fix: Honor reduced motion on manual phase disclosure

The entrance already respected the preference, but manually opening or
closing a phase did not: `useExpandCollapse` writes its transition as an
inline style, which cannot carry a `prefers-reduced-motion` media query,
and there is no global reduced-motion reset in the stylesheet. Before this
PR the phase used `<details>`, which had no animation at all — so the swap
to an animated disclosure handed reduced-motion readers a 300ms fold they
did not have.

Resolve the preference in the hook and drop the transition outright. Every
expanding panel in the message content shares it, so tool calls, thinking
blocks, attachments and web-search sources are covered by the same change.
The chevron and the fold's own utility classes get `motion-reduce`
overrides, which the inline styles cannot express.

* 🩹 fix: Keep the collapse completion signal under reduced motion

`transition: none` emits no `transitionend`, and ToolCallGroup waits on
that event to drop `shouldRenderBody`. Removing the transition therefore
left every collapsed tool subtree mounted indefinitely — expensive and
stateful children retained for exactly the readers who asked for less
work, not more.

Shorten the duration to 0.01ms instead. It is imperceptible, still fires
the event, and keeps the hook the single place that knows about the
preference. Caught by Codex on 3b9bd2181d.
2026-08-14 11:57:48 -04:00
Danny Avila
c06fbff475
📦 chore: bump @librechat/agents to v3.5.1 (#14830)
* 📦 chore: bump `@librechat/agents` to v3.5.0

* chore: bump agents sdk to v3.5.1
2026-08-14 11:14:07 -04:00
Ravi Kumar L
bc6392d05b
🪢 fix(langfuse): mark provider-backed agent traces (#14833)
* fix(langfuse): mark provider-backed agent traces

* fix(langfuse): mark stored response traces

* test(langfuse): isolate provider marker setup
2026-08-14 10:28:25 -04:00
Danny Avila
d170ecf481
🧹 ci: Remove Obsolete Test Server Deployment (#14823) 2026-08-14 09:59:35 -04:00
Danny Avila
0ce4c3374b
⏲️ test: Give ServerConfigsDB Mongo Hooks a 60s Timeout Budget (#14831)
`beforeAll` boots a real mongod via MongoMemoryServer, resets the module
registry and re-imports data-schemas, ServerConfigsDB and the MCP OAuth handler
before a single test runs. That exceeds the 15s global `testTimeout` once the
runner is busy: the suite finishes in ~4.5s on its own but has been observed at
16.8s under a loaded `@librechat/api` shard, failing every test in the file with
"Exceeded timeout of 15000 ms for a hook".

Give both mongo hooks an explicit 60s budget, matching
`checkpointer.integration.spec.ts`, the other MongoMemoryServer suite that
already opts out of the global default. `afterAll` gets the same treatment since
`mongoServer.stop()` is subject to the same contention.

No behaviour change — the timeout only bounds setup, and the suite still
completes well inside it.
2026-08-14 09:55:42 -04:00
Danny Avila
eaef87fa26
🚀 chore: Prepare v0.8.8-rc1 (#14394)
* 🚀 chore: Prepare v0.8.8-rc1 release

* 📚 docs: Complete v0.8.8-rc1 operator references

* 📚 docs: Mark stateful sessions experimental

* 📚 docs: Clarify background code capability

* 📚 docs: Refresh v0.8.8-rc1 operator guidance

* 📚 docs: Highlight v0.8.8-rc1 features in README

* 📦 chore: Bump publishable packages again

* 📚 docs: Add streaming question progress

* 📦 chore: Bump publishable packages again

* 📚 docs: Refresh v0.8.8-rc1 release highlights

* 📦 chore: Bump publishable packages again

* 📚 docs: Refresh v0.8.8-rc1 release guidance

* 📦 chore: Bump publishable packages again

* 📚 docs: Highlight batched Agent questions

* 📦 chore: Bump publishable packages again

* 📦 chore: Bump publishable packages again

* 📦 chore: Bump publishable packages again

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📄 docs: Note PowerPoint template support

* 📦 chore: Refresh v0.8.8-rc1 package versions

* 📄 docs: Note latest provider and file support
2026-08-14 03:24:59 -04:00
Danny Avila
d4c64d485f
ci: gate the e2e activity-phase DOM assertions on the persisted phase (#14821)
`activity-phases` asserted the parent `summary` was visible immediately after
`sendMessage` resolved. A parent phase only exists once the turn completes, the
phase closes, and its summary round-trips to the phase-label model — so that
assertion raced the entire pipeline and only survived on Playwright's retries.
It shows as `1 flaky` on the memory lane of a green dev run, and fails all three
attempts on slower hardware.

Gate the DOM on the durable projection instead. The test already fetched
/api/messages twice; the first fetch now also waits for the persisted phase part
before any DOM assertion runs, so the client is only asked about a phase the
server has already written.

Also drops the duplicate fetch. The two poll blocks queried the same endpoint
for the same message and both asserted
`finalTextIndex === activity_end_index`; the removed copy left `liveAssistant`,
`livePhase` and `liveFinalTextIndex` shadowing their durable equivalents.

No coverage removed — every assertion is preserved, reordered to follow the
dependency chain: persisted shape, then DOM, then label-model requests, then
the reload round-trip.
2026-08-14 01:42:39 -04:00
Danny Avila
2f0cd2eb75
🔌 chore: Bump the MCP SDK to 1.30.0 and Parse Content-Type Instead of Searching It (#14820)
`@modelcontextprotocol/sdk@1.30.0` is a small maintenance release on the 1.x line
(upstream's active line is now the 2.0.0 scoped packages). The range was already
`^1.29.0`, so only the lockfile pinned the old version; the manifests move too so
the floor matches what we test against.

Nothing in it is breaking. The four changed type declarations are additive —
optional `maxBufferSize` on `StdioServerParameters`, an optional third
constructor argument on `StdioServerTransport`, optional options on `ReadBuffer`,
optional `keepAliveMs` on the server transport — and the only manifest change is
`@hono/node-server` widening to `^1.19.9 || ^2.0.5`. No new dependencies.

Two behavior changes are worth knowing about even though neither is an API break.
`ReadBuffer` now caps a single stdio message at 10 MB (previously unbounded) and
errors the transport instead of growing, which is reachable through
`StdioClientTransport` if a stdio server returns a very large single result; it
takes `maxBufferSize` if that ever needs raising. And Content-Type handling
switched from substring search to parsed media types, client and server.

Most of the release is Streamable HTTP server hardening we do not run — a 15s SSE
keep-alive, `X-Accel-Buffering: no` on SSE responses, guards so a stale stream's
cancel cannot tear down its successor, and `_closed` checks so a transport closing
mid-request stops registering streams into swept maps. None of it changes how we
behave as a client. In particular it does not address the stale-stream 409 in
#14816: that keep-alive runs in whichever server we connect to, not here.

The same substring-vs-parse mistake the SDK corrected exists in our streamable
HTTP response guard, which classified a response as SSE with
`contentType.includes('text/event-stream')`. A `Content-Type` naming the SSE type
in a parameter — `text/plain; boundary=text/event-stream` — is not an event
stream, but matched. The guard then took `canEmitFallbackSSEError`, so an
oversized body was answered with a synthetic SSE error frame the caller reads as
a well-formed response body, rather than the throw a non-SSE response gets. The
check now compares the parsed media type, via a `mediaTypeEssence` helper added
to the header utils where `mergeHeaders` already lives.

Verified against 1.30.0 rather than assuming: the package was staged into the
worktree's own `node_modules` so it shadowed the shared install, and
`packages/api` `src/mcp` ran green on it — same four pre-existing red suites as
on 1.29.0 (`MCPReinitRecovery` plus three Redis `cache_integration` suites that
need a live Redis), no new failures.
2026-08-14 01:12:56 -04:00
Danny Avila
24d111fde9
feat: Add Gemini 3.7 Flash Support (#14818)
*  feat: Add Gemini 3.7 Flash Support

Adds first-class support for Google's Gemini 3.7 Flash (`gemini-3.7-flash`)
for both the Gemini API (AI Studio) and Google Cloud Gemini Enterprise Agent
Platform, following the Gemini 3.6 Flash integration (#14369).

- Context window (1,048,576) in googleModels; API + cache pricing in tx.ts.
- Model dropdown (config.ts) and GOOGLE_MODELS examples for both integrations.
- Register the model in the Flash-family handler so it inherits the existing
  strip of deprecated sampling params (temperature/topP/topK), rejected
  penalty params, and thinkingBudget, and defaults to `medium` thinking.
- Generalize that handler's enumerated table from a [id, level] tuple to a
  rule object, so a model can also declare thinking levels it rejects. Gemini
  3.7 Flash errors on `minimal` (which the Google endpoint offers in its
  thinkingLevel slider), so an explicit `minimal` is substituted with the
  nearest supported level, `low`. Explicit low/medium/high pass through
  unchanged.
- Apply Google's introductory pricing ($0.75 in / $3.75 out / $0.075 cached,
  per 1M) to Gemini 3.7 Flash and correct Gemini 3.6 Flash to the same rates.
  Both revert to $1.50 / $7.50 / $0.15 on 2027-01-01; noted at both call sites.

Resolves #14802

Ref: https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash
Ref: https://ai.google.dev/gemini-api/docs/pricing

* 📝 docs: Match the House Style for Promotional Rate Comments

Align the Gemini 3.6/3.7 Flash introductory-pricing notes with the existing
Sonnet 5 convention in the same file: one comment per group, naming the models
and the exact values to restore, so the manual follow-up is unambiguous.

No rate changes.

* ⬆️ chore: Bump `@librechat/agents` to 3.4.7 for Gemini 3.7 Flash Prefill

Unblocks this PR. `NO_PREFILL_GEMINI_MODELS` is model-enumerated in the agents
SDK, so 3.4.6 does not know `gemini-3.7-flash` forbids a trailing `model`-role
turn — editing an assistant reply and resubmitting would reach Google as a
prefill and return HTTP 400 on a model this PR adds to the default list.

3.4.7 (danny-avila/agents#412, released via #413) adds it. Verified the
published tarball: `3.4.6...3.4.7` touches only
`dist/{cjs,esm}/llm/google/utils/common.*` — the prefill array and its comment.
`dist/types` is byte-identical, so there is no API surface change.

Raises the declared range in both workspaces alongside the lock. `^3.4.6`
already permitted 3.4.7, but the fix is required rather than merely compatible,
so the floor should say so.
2026-08-14 01:12:42 -04:00
Danny Avila
5e464bc930
📎 fix: Alias Shell Script MIME Variants to application/x-sh (#14817)
* 📎 fix: Alias Shell Script MIME Variants to `application/x-sh`

Chrome on Linux reports `.sh` files as `application/x-shellscript`
(freedesktop shared-mime-info) and libmagic reports `text/x-shellscript`.
Neither string appears anywhere in the source, so uploads were rejected
even though `application/x-sh` is in the default allowlist and
`codeTypeMapping` maps `sh` to it — `inferMimeType` only consults the
extension map when the client sends no type at all, so a non-empty
browser value passed straight through to the allowlist check.

Alias both variants to the canonical `application/x-sh`, matching the
existing treatment of `text/x-markdown` and `application/x-zip-compressed`.

Also attach `statusCode`/`body` to multer file-filter rejections. Without
them the error misses the `isCustomError` branch in `ErrorController` and
falls through to a bare `500 An unknown error occurred.`, so the rejection
reason was logged server-side but never reached the client. The upload
hook already surfaces `error.response.data.message`, so a rejected file
now explains itself instead of showing a generic upload failure.

* 🔁 refactor: Move Upload Error Contract Into `packages/api`

Addresses codex P1 on #14817.

The producer of the `statusCode`/`body` pair now sits beside its consumer:
`isCustomError` and `ErrorController` are already in
`packages/api/src/middleware/error.ts`, and `CustomError` is already in
`packages/api/src/types/error.ts` — only the construction of that pair was
stranded in legacy JS. `createCustomError` is exported from the same module
as the guard that recognizes it, and `multer.js` is back to a thin caller.

Also pins the `.sh` back-compat claim with tests: configs from the
documented workarounds (`application/x-sh` per #4660/#5689/#6297, and the
broad patterns from #14804) still accept a `.sh` upload after the alias
rewrites the type. A negative control confirms the endpoint config is
genuinely in play rather than falling back to the default allowlist.
2026-08-14 01:12:23 -04:00
snapydziuba
6c46fd1252
📄 feat: accept PowerPoint template MIME type (#14761) 2026-08-14 00:22:01 -04:00
Danny Avila
6cbfd82772
🔌 fix: Recover Quietly From Stale MCP SSE Stream Conflicts (#14816)
A Streamable HTTP server allows one standalone `GET` SSE stream per session and
releases its mapping from the response stream's cancel callback. That callback
never runs when the connection dies at a proxy rather than at the client, so the
server keeps holding a stream nobody is reading while the client knows its stream
is gone. Every reconnect carrying that session id then gets a 409:

    SSE stream disconnected: TypeError: terminated
    Transport error (may require manual intervention):
      Streamable HTTP error: Failed to open SSE stream: Conflict
    Transport error (may require manual intervention):
      Maximum reconnection attempts (2) exceeded.

Nothing there requires manual intervention. The connection recovers on its own in
a few seconds, because the rebuild the first 409 escalates to sends the
spec-mandated `DELETE`, which drops the server's session along with the stream it
leaked. Two things made a self-healing event read as a fatal one.

`extractSSEErrorMessage` classified status by scanning the message text for
digits, but `StreamableHTTPError` and `SseError` carry the status on `code` and
their messages do not always repeat it. "Failed to open SSE stream: Conflict"
has no digits at all, so a 409 never reached the status branch and fell through
to the terminal `isTransient: false` — the same verdict as a DNS typo. A 5xx
arriving on `code` alone had the same blind spot. The status is now read from
`code` when it is in HTTP range, with the message scan kept as a fallback, and
409 joins 5xx as transient: the stale session it reports is cleared by the
rebuild, with nothing for an operator to do.

The second is volume. Each SDK retry fires `onerror` twice — once with the raw
throw out of `_startOrAuthSse`, once with the `Failed to reconnect SSE stream`
wrapper. Only the wrapper matched the existing suppression, so every doomed retry
logged at error level, and the retries are doomed by construction: nothing about
the same session id can stop conflicting. The first conflict now escalates for
rebuild and the rest are logged as the echo they are, along with the SDK's
out-of-retries announcement when a rebuild is already underway. The non-conflict
path for that announcement is untouched, so an exhausted budget still falls
through to our reconnection everywhere else.

`extractSSEErrorMessage` moves to `errors.ts` alongside `isOAuthAuthenticationError`.
It had no test: `MCPConnection.test.ts` held a hand-copied clone marked "keep in
sync with the actual implementation", so 66 assertions were exercising the copy.
The clone is deleted and the suite now imports the real function, which it turns
out had not drifted.

`MCPConnectionSseConflict.test.ts` drives a real client transport against a real
in-process `StreamableHTTPServerTransport` reproducing the sequence above: the
stream opens, its socket is destroyed underneath the client, and every later
`GET` on that session id conflicts while a rebuilt session gets a healthy stream.
2026-08-14 00:01:13 -04:00
Danny Avila
0654efb7ed
🔌 fix: Preserve MCP serverInstructions Declaration Through Inspection (#14815)
`MCPServerInspector` overwrote the operator's `serverInstructions` declaration
with the text fetched from the server. That made a YAML server's cached entry
differ from its own raw config on an admin-configurable field, so
`isUnmodifiedYamlServer` misclassified it as admin-modified and re-inspected it
on the first user-scoped resolve.

The second inspection produced a config with a newer `updatedAt`, which:

- flipped `getServerConnectionStatus` to `disconnected` permanently, since the
  healthy app connection was then measured against the newer timestamp; and
- made `isAppServerConfig` reject the effective config, gating off the app
  connection so `GET /api/mcp/tools` returned zero tools and cached nothing.

Fetched instructions now land on a separate `resolvedInstructions` field,
matching how every other inspector-derived value is stored, so the declaration
survives inspection and the guard compares like with like.

Bumps `REGISTRY_STORAGE_SCHEMA_VERSION` so Redis-backed deployments rewrite
entries whose `serverInstructions` still holds fetched text.

Fixes #14798
2026-08-13 23:37:58 -04:00
Danny Avila
7694428ca9
💬 style: Right-Align In-Flight Steer Bubbles to the Message UI (#14814)
* 💬 style: Right-Align In-Flight Steer Bubbles to the Message UI

The chat surface reads as message bubbles now — user turns on the right,
assistant turns on the left — but the in-flight steer bubbles anchored above
the composer were still left-aligned, so a steer sat on the opposite side from
the words the user had just sent, then jumped across on `on_steer_applied`
when the persisted `SteerPart` landed in-thread on the right.

Align the overlay with the user turn it belongs to:

- The bubble stack right-aligns and is constrained to the message column
  (`max-w-3xl`), so the in-flight bubble sits where its applied twin lands
  instead of ~52px further right (the composer runs wider than the message
  column at `xl`).
- The bubble adopts the same theme-token geometry as `SteerPart` and every
  user turn (`rounded-theme-surface rounded-br-theme-control`,
  `px-theme-normal`), replacing the raw `rounded-3xl`/`pl-3 pr-4`. It keeps its
  outline: an in-flight steer is still provisional.
- The controls flank the bubble — overflow menu outboard-left, send-now arrow
  outboard-right — so neither reads as belonging to the other. DOM order
  matches visual order, so focus order stays coherent.

Also drops the thin `bg-border-medium` divider that bound the arrow to its
message: with the arrow now outboard on the far side of the bubble it has
nothing to separate, and `EscalateNowButton` no longer needs its fragment.

* 💬 style: Center the Steer Controls on the Bubble's First Line

The flanking controls read as neither top-aligned nor centered, because their
resting position was an accident of `sticky top-2`: the topmost rail trips the
sticky inset at rest and is shoved 8px below the row top, while every rail
below it clears the inset and stays at the top. So the controls sat 3.8px above
the bubble's centre — and stacked steers did not even agree with each other.

Give each rail a `py-3` band that reproduces the bubble's own first line (its
`py-2.5`, its 1px border, and half the gap between the 24px control and the
taller text line box), and pad the overlay evenly so the topmost rail already
clears the sticky inset instead of being displaced by it.

A 24px control now centres on the first line: measured at 722.0 against the
text's 721.8, versus 718.4 before. Beside a one-line steer that reads as
centred; on a tall one it aligns to the opening line rather than drifting to
the middle, and sticky still carries it while the stack scrolls.
2026-08-13 23:37:18 -04:00
Mihidum
da390fa919
🩹 fix: apply agent updates that match the newest version entry (#14810)
`updateAgent` returned early when the resulting state matched the newest
`versions` entry, so `findOneAndUpdate` never ran and the caller's update
was discarded behind a 200 response.

Suppressing a redundant version entry is correct; suppressing the write is
not. The document is regularly not equal to its newest version entry:
`$push`/`$pull`/`$addToSet` updates snapshot the pre-update state (as
`addAgentResourceFile` does on every file attach), `skipVersioning` writes
snapshot nothing, and `removeAgentResourceFiles` bypasses `updateAgent`
altogether. Any update that moved the document back onto that entry's
content was then dropped, leaving the drifted state in place.

Keep the version entry suppressed, apply the write, and still report the
unchanged `versions` count as `version` so callers keep their existing
"no new version" signal.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 23:18:34 -04:00
Danny Avila
abc669ab58
🩹 fix: Restore the @librechat/api Build and Remove Legacy Code (#14808)
* 🧹 chore: Remove Dead Legacy Agent Controller

`_LegacyAgentController` has been unreachable since the resumable path became
the only route: it is unreferenced, unexported, and untested. It had also
drifted out of compilability against the live file — line 2009 called
`attachConversationCreatedAt(req, { userId, conversationId, isNewConvo })`
against the 3-argument signature declared at line 97, which would await
`undefined` and then throw dereferencing `resolved.createdAt`.

Keeping it was not free. It carried a third independent copy of the response
message-id wiring (`getReqData`, `onStart`, four `updateMetadata` calls), so
every change to how a generation identifies its response row had a dead third
site to keep in step, and no test to say whether it had been kept in step.

Removing the block leaves `createCloseHandler` and the `sendEvent`,
`clientRegistry`, `requestDataMap` and `handleAbortError` imports with no
remaining callers, so those go too. `AgentController` was a three-line
passthrough to `ResumableAgentController`; the real controller is now exported
directly, which also matches the `[ResumableAgentController]` prefix every log
line in the file already uses. `server/routes/agents/chat.js` binds the export
to its own local name and passes the same five arguments, so the route is
unchanged.

No behavior change: 379 lines removed, 2 added.

* 🩹 fix: Remove Duplicated Anchor Block Breaking the `@librechat/api` Build

`dev` does not build. `packages/api/src/agents/activityPhases/runtime.ts`
carries two byte-identical 98-line copies of the same block (former lines
516-613 and 614-711), so rolldown fails to parse it:

    [PARSE_ERROR] Identifier `AnchorFields` has already been declared

The duplicated block is the anchor-construction work from #14805:
`AnchorFields`, `laterDefinedIndex`, `foldedAgentIds`, `boundedAnchor` and
`mergeAnchors`. #14807 was squashed from a branch that predated #14805 and
re-included that commit, so both copies landed. Only the `type` produced an
error — the four function declarations simply redeclare.

This removes the first copy. The two blocks were verified byte-identical
before the cut, and the resulting file has no duplicate top-level
declarations, is missing nothing that #14805 introduced, and retains
everything new to #14807 (`ResolvedPosition`, `resolvePosition`).

Verified: `tsdown` builds, `tsc --noEmit` clean, `config/circular-deps.mjs`
green across all five graphs (it was reporting `✗ @librechat/api` purely
because the build it shells out to was failing), and the 68 tests in
`activityPhases/runtime.spec.ts` pass.

Carried here rather than in a separate PR because this PR's checks cannot go
green until it lands: the failed `packages/api` build cascades into e2e, MCP
list_changed, bombadil and the Docker image jobs.
2026-08-13 19:57:32 -04:00
Lizzy
5bd745783c
🖌️ style: Use correct text colour for answer textarea in dark mode (#14791) 2026-08-13 19:37:13 -04:00
Marco Beretta
d920328bfa
💬 style: Unify Message Row Layout and Edit Surfaces (#14770)
* style: Unify message row layout and edit surfaces

Route chat, share, and search messages through a shared MessageRow so
user turns render as right-aligned bubbles and assistant turns keep a
visible identity column.

Replace per-part text editors with one edit surface that keeps tools,
errors, and artifacts visible. Preserve non-text fields when saving
content parts, copy the full serialized message, and hide hover actions
that do not apply during streaming or errors.

* style: Align edit footer and lighten editor field in dark mode

Drop the divider above the user edit footer so both edit surfaces share
the same footer treatment.

Move the editor fields to surface-tertiary-alt. Light mode is unchanged
at #fff, while dark mode lifts from #0d0d0d to #2f2f2f so the field sits
above the #212121 panel instead of sinking into near-black.

* style: Drop focus border and ring from message editors

The editor fields changed border color and added a ring on focus. Keep
the border static and rely on the app-level focus handling instead.

* fix: Keep a triggered message action visible when the row is not hovered

Hover actions fade out on non-last rows, and mobile.css only restored
display and visibility for an active button, never opacity. Opening the
fork popover therefore left it anchored to an invisible trigger once the
pointer left the row. Skip the fade entirely while a button is active.

Extract the recipe the three toolbars repeated so the rule has one home.

Rework the streaming guard to the contract the toolbar now implements:
edit and fork are omitted from a streaming response rather than rendered
disabled, and the settled turn above keeps its own actions. It asserted
the removed disabled-and-transparent behaviour and its opacity check only
held because the growing response shifted the row out from under the
pointer.

* style: Trim message edit chrome and stabilize the status row

The edit surface was a titled card sitting inside the conversation: a
bordered panel with an "Edit message" heading wrapping bordered fields,
which read as a settings dialog rather than an inline editor. Drop the
card background, border and heading, and take the footer buttons down to
the small size so the editor reads as a field in the message flow. The
captured row goes from 253px to 187px.

Move "Unsaved changes" into the footer and merge the rerun hint into the
same slot. Both previously added their own row, so typing pushed the rest
of the conversation down. The slot is clamped to two lines, which stays
under the 36px button row, so the footer height holds at 36px regardless
of which message is showing.

* test: Cover message edit layout stability

Add a mock e2e spec that measures the edit footer and section boxes and
asserts they hold steady as the status text appears, for both the
single-part user editor and a multi-part response.

The multi-part case needs an assistant message with two editable parts,
so add an E2E_THINK_REPLY marker to the fake model. Its think tags are
parsed downstream by the agents stream pipeline, which yields a reasoning
part followed by a text part.

* fix: Read the fork popover open state from its store

Fork mirrored the popover state into its own useState and reset it from an
onClose prop. Ariakit 0.4 has no onClose, and React's DOM types accept the
name on any element, so it type-checked, landed on a div and never fired.
Closing by Escape or an outside click therefore left the button reading as
active until the trigger was clicked again.

Read the state from the store instead so every close path clears it.

* fix: Keep the whole toolbar visible while an action is open

Only the triggered button escaped the hover fade, so opening the editor or
the fork popover left the row as a single floating button once the pointer
moved away. Mark the active button and have every action in the toolbar key
off it, so the group stays opaque for as long as a surface is open.

The marker is a dedicated class rather than the existing `active`, which
HoverButtons pins to the edit button of every assistant message and would
hold those toolbars open permanently.

The existing guard pressed Escape to close the editor while focus sat on the
body, so the editor never closed and its assertion only held because the
sibling faded regardless. Close the editor through its own control, and drop
focus before measuring the fade now that Escape returns it to the trigger.

* fix: Withhold copy while a response is still streaming

Text-to-speech, fork and feedback were all withheld from a message that is
still generating, but copy was rendered throughout, so the button offered to
put half a sentence on the clipboard. Gate it on the same condition.

That empties the toolbar for the duration, and SubRow collapses an empty row,
so a streaming response now carries no actions at all until it settles. Both
guards encoded the old contract: the unit test asserted copy was present and
counted a single button, and the browser guard used copy as its proof that the
toolbar had mounted. The settled turn above takes over that role.

* fix: Move retry navigation to the outer edge of a user turn

A user turn is right-aligned, but its sibling navigation rendered ahead of the
actions, so the retry counter sat inboard of the icons instead of under the
edge of the bubble it belongs to. Order it last on user turns.

* fix: Ride the stream instead of chasing it

Following a generating answer went through a helper throttled at 145ms, so the
thread caught up in visible jerks rather than flowing. It now writes the scroll
position directly on each frame, which is what an answer arriving a few pixels
at a time actually needs, and glides only for the one long trip a turn makes,
when sending has to travel from wherever the reader was down to the newest
word.

Whether to follow at all is now answered by where the reader is and which way
they were going, rather than by the abort flag. `useMessageProcess` raises that
flag on any wheel at all, downward ones included, through a throttle whose
trailing call lands after the gesture has ended, so nothing timed to the
gesture could outlive it. Scrolling down to the newest word could therefore
never resume the ride, while the scroll-to-bottom button, which touches no
wheel, always could.

Arrival is judged on the scroll it produces rather than the wheel tick that
started it, because wheel scrolling is animated and at tick time the thread is
still far short of where the tick is taking it. Arriving also counts from
further out than leaving does: while an answer streams the end recedes between
the last tick and the frame that measures it, so judging arrival as tightly as
departure leaves a reader unable to catch it at all.

* fix: Reveal retry navigation on hover while an answer generates

Copy, edit, fork and read-aloud are all withheld from a response that is still
generating, which left the retry counter as the only thing rendering under a
half-written answer. It now reveals on hover there, like the actions it sits
with, and stays put on a settled turn.

* fix: Keep a refused rerun from discarding the edit

While a response is streaming, the edit action stays available on every earlier
row, and those editors see a per-message submitting flag that is false, so
Update and rerun is enabled. The send itself is still refused: ask() returns
false for the duration of the active submission. Both editors ignored that and
closed anyway, so the draft went with them and no rerun ever started.

Both rerun paths now check the result and leave the editor untouched when the
send is refused, so the work survives until the thread is free.

* fix: Let an upward gesture beat the pending send glide

Sending arms a smooth glide down to the newest word, and the landing re-pins the
thread to the bottom. The landing was scheduled two ways, on scrollend and on a
700ms fallback, and neither was ever cancelled. A reader who changed their mind
and headed up mid-flight was pinned again regardless, then dragged back by the
next streaming resize. The fallback fires for the whole window, so this held even
after the glide had visibly settled.

The gesture now marks the glide interrupted, wherever it lets go of the bottom,
and the landing stands down when it sees that. A glide the reader leaves alone
still re-affirms the ride.

* fix: Fade retry navigation on every streaming response format

Every other action is withheld from the row that is still generating, so the
retry counter is the only thing left under a half-written answer. The plain text
row already faded it to hover-only there; the structured rows did not, and left
it sitting on its own.

Both structured paths now apply the same condition, and the class string the
three of them share moves next to the hover action styles it belongs with.

* i18n: Correct the copy the edit surface rewrite left behind

The multi-part hint told the reader to save first and then rerun, but a save
closes the editor and reopening seeds the drafts from what was just saved, so
there is nothing left to rerun and the button stays disabled. Rerunning carries a
single edited section by design, so the hint now states that limit rather than
pointing at a step that is not there.

Drop com_ui_save_submit as well: the per-part editor that used it is gone.

* test: Make the message visual baselines opt-in

The suite asserts sixteen screenshots and the repository tracks none, so
Playwright's default treats every one as a miss and the mock e2e job fails on
Linux. Baselines only compare cleanly against the machine that produced them, and
nothing here can generate ones that match the runner image.

The flows keep running and asserting their structure, which is where their value
was; only the pixel comparison is now gated behind E2E_VISUAL_SNAPSHOTS.

* style: Restore import order in the reworked message files

The repository sorter and CI disagreed with what these files were left holding
after the edit surface rework. No behavior change.

* test: Follow the reworded rerun hint in the edit layout spec

The multi-part hint was restated in the previous commit; this assertion still
expected the old wording and would have failed the mock e2e suite.

* fix: Leave the send glide alone while the answer streams in

Every delta of an answer reruns the scroll effect, and the plain follow writes
scrollTop outright, which cancels an animation on its first frame. So the glide a
send starts was killed by the first token to arrive and the reader was snapped
down instead of carried.

The follow now stands down while a glide is travelling, which is what the hook
already documented but only enforced on the resize path.

* fix: Write a saved edit onto the thread as it stands

An earlier turn stays editable while the newest answer streams, and the save
captured the thread before the request but wrote it back after. Every delta that
landed during the round trip was overwritten. Most of the time the next delta
re-merged and the damage showed as a one-frame truncation, but a save that
resolved after the stream's final write left the cache wrong for the rest of the
session.

The thread is now read once the request has resolved, which is what the content
part editor already did.

The editor actions in this file also wrap again rather than hold one unbreakable
row, for the reason given in the following commit.

* fix: Let the editor actions wrap on a narrow row

At 320px an assistant turn gives the editor about 252px once page padding, the
identity column and the row gap are taken out, and Cancel, Save and Update &
rerun need more than that in English alone. The group was pinned with shrink-0,
so it ran past the edge of the row instead of wrapping. A longer translated label
makes it worse, and the user turn had no margin left either.

Both editors wrap again, which is what the footer did before the status row was
folded into it.

* fix: Catch up to the new bottom when the glide lands

Following stands down for the length of the glide, so an answer that arrives
while it travels moves the bottom past the target the glide aimed at. A short
response that finished before the glide reported landing left the thread a few
lines short of its own end, with nothing left to correct it.

Landing now closes whatever gap opened, unless the reader took over on the way.

* test: Follow the renamed rerun button in the edit flow specs

The button became 'Update & rerun' when the edit surfaces were unified, but two
edit-flow specs still located 'Save & Submit' and would have waited for it until
they timed out. A type comment named the old button too.

* fix: Judge the first thread scroll against a real position

The direction check seeded its last-position ref at 0, so the first scroll
event on an opened thread, which arrives carrying a large positive
scrollTop, read as a jump downward. Near the end that cleared the abort
flag and re-pinned a reader to the stream they were scrolling away from.

Take the first event as a baseline and judge direction from the next.

* fix: Hold the content part editor to what it replaced

EditContentParts took over from EditTextPart and left two of its behaviors
behind.

An emptied box now blocks Save and rerun instead of persisting a blank
part. EditTextPart refused the same edit through its form's required rule
and the sibling EditMessage still does, so both editors hold one line. The
keyboard shortcuts reach the save paths directly, so they are guarded
there too, and the footer says why the buttons are down.

The editor also follows the chat direction again, taking dir and text
alignment from the same setting EditMessage reads.

* fix: Hold the footer height while a response streams

Every action is withheld from the row that is still generating, and a lone
sibling counter renders nothing, so the footer measured zero until the answer
landed and then sprang to the height of the buttons. The transcript stepped
upward under the reader at the moment a response completed.

The placeholder that used to reserve this space went when the footer became
unconditional, so hold the height on the row itself instead.

* fix: Remember where the thread was put before judging a gesture

Direction is judged against the last sample, and the thread is placed at its
end without the reader touching it. With no record of where it was put, their
first gesture was spent taking the baseline instead of being obeyed: a single
PageUp cleared no flag of its own, so the next streamed resize rode the reader
straight back to the end they were leaving.

Every programmatic move now records the position it left the thread at, so the
sentinel stands only until something has actually placed it.

* fix: Spend the start of a turn only once it can be honored

A reader who scrolls away during one answer leaves the abort flag raised, and
nothing lowers it until the next connection opens, which is after this effect
has already seen the send. Marking the turn as started on that first pass spent
it against a closed gate: by the time the flag cleared there was no start left
to honor, the reader was still detached, and the answer they had just asked for
streamed on offscreen.

Record the turn as started only on the pass that acts on it.

* fix: Show the part edits that survived a refused save

The editor saves every changed part through one button, but the endpoint
takes a single part per call and nothing rolls a write back. A part the
server refused therefore left the earlier ones stored while the editor
reported that the message could not be saved, so cancelling from there
walked away from edits that were already live.

Record the writes that landed and reconcile the transcript with them
whichever way the save ended. The refused parts are the only ones left
holding a draft, so a retry no longer rewrites what already arrived.

* fix: Stop a shared transcript from calling the sharer the reader

The share row reused the chat view's user label, which reads "You". It is
the screen-reader heading for the user turn, so anyone opening a share
link heard every prompt the sharer wrote credited to themselves.

Use the neutral "User" label on this surface. It keeps the localization
the row gained, unlike the untranslated string it replaced.

* fix: Let go of the stream when an interaction settles over several resizes

Expanding a tool result mid-answer renders the container first and fills it
once its contents arrive, so one gesture produces more than one resize. Only
the first was credited to the interaction. The second read the reader as still
riding the stream and put them back on the bottom they had just left.

The suppressed resize now settles the ride as well as the near-bottom measure,
using the position the interaction actually left the reader at, so an
interaction that kept them on the end still streams.

* fix: Edit inside a structured text part instead of flattening it

A text content part holds either a string or a { value, annotations } object.
The Assistants thread sync persists the structured form with its file
citations intact, and the editor reads the part through the same union, so
saving an edit wrote a bare string over the whole object and took every
citation with it.

The same object was handed to the tokenizer, which measures length, so a part
that had been edited this way also stored a NaN token count. Write the edit
into value, keep the rest of the part, and count the text itself.

* fix: Keep a saved part's citations in the transcript it is written back to

A text or think part holds either a bare string or a { value, annotations }
object, and the editor already read both through getPartText. Writing the
draft back into the local message cache put the string over the whole value,
so a response carrying file citations lost them the moment it was edited and
did not get them back until a refetch.

Reading and writing now go through the same accessor, so an edit lands in the
shape it was read from and the rest of the part survives.

* fix: Let the message editor follow the chosen font size

Editing a message dropped the draft to a fixed 14px regardless of the
Font Size setting. On dev the textarea carried the markdown class, so it
read --markdown-font-size like the rendered message does; restyling it
into a bordered box replaced that with text-sm, and the new per-part
editor was written the same way. Anyone on Extra Small, Large or Extra
Large saw the text jump the moment they entered edit mode.

Share the .message-content typography with the editors through a
message-editor-text class so a draft is sized like the message it
replaces and keeps tracking the setting.
2026-08-13 19:30:39 -04:00
Danny Avila
78eb0c98ce
🧭 refactor: Resolve Activity Phase Position Once Per Boundary (#14807)
* refactor(api): make anchor construction exhaustive

Bounded anchors were built in two places, each spreading one side and
hand-picking the rest, so any field nobody named was dropped silently and
nothing failed until a boundary landed badly. That already cost agentId
and then unresolvedToolStartIndex in consecutive review rounds, and
mergedAgentIds was never added to the demotion path at all.

Both constructors now assign an AnchorFields literal, mapped over
keyof Required<TrackedActivity>, so adding a field to TrackedActivity is
a type error at both sites until its anchor semantics are decided.

No behavior change: every field resolves to what the hand-picked versions
already produced. Adds a folding property over count, failure count,
agent attribution, and ordering.

* refactor(api): resolve activity position once per boundary

Positional fields on TrackedActivity are captured at different times
against an array that keeps moving, so each was a cache that could go
stale, collide, or be truncated — and closesBeforeBoundary read three of
them directly. Roughly two thirds of the review findings on #14785 were
that pattern: a proxy outranking, outliving, or standing in for the
rendered position.

resolvePosition now folds tool indices, the prior partition floor, the
unmaterialized fallback, and reasoning anchors into one value, and
closesBeforeBoundary takes only that value. A caller cannot reach past
it to a raw field, and the four branches the predicate used to carry
collapse into one comparison: an activity closes early exactly when
nothing locates it beyond the boundary.

Resolution also decides when the saved fallback has gone stale, so the
stripping that kept it out of snapshots is now a property of the resolved
value rather than a separate step.

No behavior change; the existing boundary and straddle regressions cover
both directions.
2026-08-13 19:29:42 -04:00
Danny Avila
619ed2f1fb
🧱 refactor: Make Activity Phase Anchor Construction Exhaustive (#14805)
Bounded anchors were built in two places, each spreading one side and
hand-picking the rest, so any field nobody named was dropped silently and
nothing failed until a boundary landed badly. That already cost agentId
and then unresolvedToolStartIndex in consecutive review rounds, and
mergedAgentIds was never added to the demotion path at all.

Both constructors now assign an AnchorFields literal, mapped over
keyof Required<TrackedActivity>, so adding a field to TrackedActivity is
a type error at both sites until its anchor semantics are decided.

No behavior change: every field resolves to what the hand-picked versions
already produced. Adds a folding property over count, failure count,
agent attribution, and ordering.
2026-08-13 19:29:18 -04:00
Danny Avila
05ed7ad8c0
🔖 fix: Split Activity Phases at Substantial Text (#14785)
* fix(api): split phases at substantial text

* tune(api): split phases after 200 text chars

* fix(api): reanchor substantial text boundaries

* test(api): type multi-phase payload captures

* fix(api): preserve activity phase boundaries

* fix(api): anchor retained phase partitions

* fix(api): persist phase partition anchors

* fix(api): harden activity phase boundaries

* fix(api): preserve bounded phase partitions

* fix: preserve final and delayed phase content

* refactor(api): partition phase state at one boundary

Boundary closure split fifteen separately-maintained fields by hand, and
each fix partitioned one more while the next stayed unguarded. Fold the
overflow bookkeeping into the tracked activity list so every counted
activity carries a position, and route the split through a single
partitionAt that returns both sides.

Counts are now summed from the partition instead of reconstructed by
subtraction, so a run past the anchor budget reports every activity it
performed rather than the truncated window. Snapshots move to version 3;
the reader still accepts versions 1 and 2 and rebuilds their unpositioned
remainder as a bounded anchor, dropping it when its evidence is stale.

Adds a boundary-conservation property covering every split point.

* fix(client): drop empty phase content segments

Late-child recovery can strip every index from a segment it already
claimed, leaving a content segment with no parts. Each one still mounts a
nested ContentParts that renders nothing, and it broke the exact-segment
expectation in the late-child regression from 831a00353.

Route the four content pushes through one guard that skips index-less
segments, matching the existing splice of fully recovered segments.

* fix(api): keep the run's answer outside the collapsed phase

The substantial-text boundary replaced completion's final-text boundary
outright, so a short reply from a provider that emits no phase metadata
was folded into the collapsed parent. That is the deterministic e2e
failure at activity-phases.spec.ts:182 and codex's short-final-answer
findings; 831a00353 fixed only the path where the provider labels the
step final_answer.

Restore the completion boundary at the last materialized visible text
whatever its length. Length now decides only whether intermediate text
earns a boundary, and semantic commentary still stays inside. The
"later work" check shares one predicate with partitionAt so the two
cannot drift.

Also splits a legacy v1/v2 remainder across the positions its saved tool
anchors still materialize at, each carrying its own id so it can be
located, and merges over-cap anchors by closest pair into the earlier
position instead of folding the oldest forward.

* fix(api): clear resolved anchors and keep folded agents

Two findings from the latest review:

A resumed activity whose tool was missing at construction kept its high
fallback anchor after that tool materialized at a lower index, so the
partition rejected it at any boundary below the stale value and pushed
pre-boundary work into the following phase. Drop the anchor once every
tracked call has materialized.

Folding anchors past the cap spread only the surviving side, silently
dropping the other's agent. close() now derives both marker attribution
and the summarizer payload from the partitioned activities, so a merged
anchor carries the union instead.

Both regressions are mutation-checked against their own fix.

* fix(api): anchor live batches awaiting materialization

A batch tracked after its child-label slot is reserved but before its
tool call reaches the shared array had no materialized position, and the
partition read "nothing materialized" as "happened earlier". A boundary
between the two then counted the batch in the earlier phase while ending
before its eventual tool call, stranding the tool outside its parent.

Record the tracked start as an unresolved anchor in that window so the
existing retain branch keeps the batch on its own side. Using the plain
fallback index instead is wrong: restored evidence-less activities carry
a synthesized index, not a position.

Regression is mutation-checked against its own fix.

* perf(api): partition in one pass and reanchor filtered batches

Dropping a batch's already-covered calls leaves a different activity
behind, but the batch position was still the covered call's index. The
survivor therefore inherited a position inside an emitted phase and was
consumed by it instead of being held for its own. Re-derive the start
from the retained ids, which also restores the unresolved-anchor signal
when none of them have materialized.

The boundary partition also classified every activity twice and rescanned
retained ones to reanchor, walking the shared content array several times
per activity per boundary. Build both sides in one pass with the
materialized tool indices computed once and threaded into the predicate.

Regression is mutation-checked; an earlier version of it was vacuous
because the tool materialized before completion, converging both paths.

* perf(api): carry tool indices through boundary resolution

The previous pass cached the materialized indices only in the partition
loop, so resolution still scanned for the batch start and again for the
fully-materialized check, and an empty result triggered a third scan
inside the boundary predicate.

Walk the shared content array once per activity and carry the indices
through resolution, classification, and reanchoring. findTrackedToolStart
becomes its own first element and is dropped.

* fix: trust rendered position over registration order

Three findings from the latest review:

Context partitioning only consulted the rendered index when the activity
position tied the closing count, so a parallel lane registered before the
closing tool hooks was assigned to the earlier phase despite rendering
after the boundary. An activity position is registration order; a
materialized index is proof, and now wins whenever it has one.

Snapshot restore bound pending reasoning to the first part sharing its
80-character anchor, which could replay a still-pending lane on the
earlier side of a boundary and delete it. An ambiguous anchor is treated
as unresolved.

Recovering the only filled child label out of a phase segment left its
hasContent flag set, rendering an expandable card with an empty body.

The context regression is mutation-checked against its own fix.

* fix(api): decide context by proven position, both directions

The previous change let a rendered index override registration order only
when it proved the text was after the boundary, and trusted that index
even when it was not provably this entry's.

Both gaps were reachable. Context registered after work that already
rendered ahead of it was retained despite rendering before the boundary,
and a restored entry with no step id whose excerpt repeats after the
boundary matched the later occurrence and moved to the wrong phase.

Locating now reports whether the position is authoritative — anchored by
a step index or a unique text match — and only then decides, in both
directions. Otherwise the saved activity position stands.

Each regression is mutation-checked against its own direction.

* fix(api): carry unresolved positions through folded anchors

Folding two anchors spread only the surviving side, dropping the later
one's unresolved fallback. A boundary between them then saw just the
earlier materialized tool index and closed the whole merged count,
counting work whose tool call had not appeared and leaving that call
outside its parent.

Carry the later fallback into the merged anchor; resolution already
clears it once every retained id materializes.

Regression is mutation-checked against its own fix.

* fix(client): keep phase headers recovery did not empty

A completed phase can carry no children after compaction — its summary
header is the whole segment. Recovery spliced any segment left with no
retained indices, so a later marker deleted that header even though it
recovered nothing from it.

Only drop a segment recovery actually emptied, not one that arrived
empty. Regression is mutation-checked against its own fix.
2026-08-13 16:16:09 -04:00
Danny Avila
bcbe26ab4c
🪑 fix: Rebase Activity Phase Bounds Onto Compacted Content and Unskip the MCPManager Suite (#14782)
* 🧭 fix: Rebase Activity Phase Bounds Onto Compacted Content

`filterMalformedContentParts` compacts the aggregator's content array —
`Array.prototype.filter` skips holes and drops malformed tool calls — but a
parent phase marker's `activity_start_index`/`activity_end_index` still address
the pre-filter positions. The array is routinely sparse: the aggregator writes
parts at provider-source indexes, so a model turn that emits no text before its
tool calls leaves an empty slot.

Every part after a hole therefore shifts left on persistence while the bounds
stay put, so the stored phase claims the wrong range — the final answer is
swallowed into the parent card and the marker's own slot is counted as a child.
The in-run analogue (`rebaseActivityPhaseBounds`) already rebases after
completion-time reshaping; the final compaction had no such step.

Rebase the bounds as part of the compaction, mapping each bound to the number
of retained parts ahead of it. The mapping is monotonic, so `start <= end <=
markerIndex` survives, and an identity mapping leaves untouched arrays — and
their marker objects — exactly as they were. Markers are copied rather than
mutated so the caller's array keeps its own coordinates, which the live stream
and the resume snapshot still address.

Fixes the `activity-phases` e2e failure on dev and the same defect on the two
resume persistence paths.

* 🔌 fix: Stop Replacing the Env Module in the MCPManager Suite

`MCPManager.test.ts` mocked `~/utils/env` with a factory that replaced the whole
module. #14780 then made `~/mcp/utils` read `ALLOWED_BODY_FIELDS` from that
module at module scope, so importing `~/mcp/oauth` -> `handler.ts` ->
`~/mcp/utils` evaluated `undefined.map(...)` and the suite died at import time.
All 111 of its tests have been silently skipped since; the shard has been red on
dev, on this PR, and on release-v0.8.8-rc1.

Spread the real module and keep only the mock that earns its place.
`processMCPEnv` stays a seam: fifteen cases drive it with `mockReturnValue` /
`mockImplementation` to hand the manager a specific processed config, and one
asserts its call count, so making it real would couple these tests to
env-substitution logic. `isPluginSourced` and `MCP_PLUGIN_SOURCE` were dropped —
the factory restated the real implementations verbatim and no test referenced
either, so they were duplication, not a seam.

111 tests now run and pass.

* 🧪 test: Stop Replacing the Env Module in Three More Suites

Same latent trap as the MCPManager suite: a `jest.mock('~/utils/env', ...)`
factory that replaces the whole module. These three pass today only because
their import graphs never reach `~/mcp/utils`, which reads `ALLOWED_BODY_FIELDS`
from that module at module scope — the next module-scope constant added to
`env.ts` would break all three the same silent way.

Each mock is kept only where it earns its place:

- `activityLabels/host.spec.ts` — dropped. `createSafeUser` was never referenced
  and the stub returned `undefined` where the real function returns `{}`, so the
  mock was strictly less faithful than the real, pure implementation.
- `run-codeTools.test.ts` — dropped. Neither `resolveHeaders` nor
  `createSafeUser` was referenced by any case.
- `run-summarization.test.ts` — `resolveHeaders` is now a spy wrapping the real
  implementation rather than an identity stub. One case asserts templated header
  values go through it, which only means something if the real substitution
  actually runs. `createSafeUser` dropped as unreferenced.

103 suites / 2708 tests green across `src/agents`, `src/utils`, and the
MCPManager suite.

* 📝 docs: Describe the Full Contract of filterMalformedContentParts

Per Copilot's review: the public JSDoc still described the function as only
dropping malformed tool calls, while the implementation also compacts empty
slots and rebases parent activity-phase bounds. The detail lived on the private
helper, so callers reading intellisense saw a stale contract.

State what it actually produces, note that compaction is inherent rather than
incidental (the aggregator writes at provider-source indexes, so the array is
frequently sparse), and add an example of a hole moving a phase bound. The
example was verified against the built runtime, not written from memory.
2026-08-13 07:52:12 -04:00
Danny Avila
6755544cee
🌍 ci: Harden Locize Translation Sync (#14784) 2026-08-13 07:29:29 -04:00
Danny Avila
155f71f81a
📱 fix: Show Quote Popup for Block Selections and on Touch Devices (#14777)
* 📱 fix: Show Quote Popup for Block Selections and on Touch Devices

The "Add to chat" popup never appeared for two whole classes of selection.

Block-granularity gestures (triple-click, double-click then word-drag) park
the selection's far boundary at the start of the next block. For a message's
closing block that boundary sits outside `.message-render` — on the composer
wrapper or the following message row — while selecting no text there, so the
anchor/focus equality check suppressed the popup. Triple-clicking any earlier
paragraph worked, which is what made this look like an edge case. The range is
now clamped to the message before the check, and selections that really do
carry visible text from another message are still refused.

Touch platforms could not reach the feature at all. A long-press, and every
drag of the native selection handles, emits no mouse event whatsoever — only
`selectionchange` — while the popup was shown exclusively from mouseup,
dblclick and keyup. Showing now also hangs off a settle-debounced
`selectionchange`, gated so an in-progress mouse drag still cannot flicker it.
Accepting was broken independently: the tap is also the gesture that dismisses
the selection, unmounting the button before `click` could land, so touch
commits on `pointerdown` instead. The desktop mousedown path is deliberately
unchanged, since preventDefault on `pointerdown` can suppress the
compatibility mousedown that click depends on.

Two UX consequences of the same code: scrolling re-anchors the popup rather
than dismissing it on the first event (the chat auto-scrolls constantly while
streaming, and a mobile URL bar collapsing fires resize), and touch selections
place the button below the text, clear of the OS Copy/Share callout, with a
44px tap target.

Covered by six e2e tests — three desktop, three on an emulated Pixel 5 with a
real touchscreen — each verified to fail against the pre-fix build.

* 🩹 fix: Address Review Findings and Repair the Scroll Specs

The two failing e2e shards were a defect in the specs, not the component.
`scrollMessages` reached for `.scrollbar-gutter-stable` with a document-wide
query, but the nav and side panels carry that class too, so it could grab a
sidebar list that never scrolls — 0px moved, and only in CI, where the nav
renders differently. The scroller is now reached from the message itself, the
way `MessageNav` does it. The specs also centre the selection first and nudge
by a quarter of the visible height, so the gesture cannot scroll the selection
clean out of view and then blame the popup for going with it.

Review findings, all in `QuoteButton`:

Visibility was tested against the window, but the list scrolls inside a bounded
container, so text can sit clipped under the header or the composer while its
un-clipped rect is still inside the window — leaving the popup floating over
unrelated UI. It is now clipped to the nearest scrollable ancestor.

Touch committed on the press, so starting a scroll on the button, or touching
it and thinking better of it, still added the quote. The excerpt is captured on
the press and committed on the release, and only when that release lands on the
button, restoring the cancellation every button is expected to have. Commit on
press existed because the tap dismisses the selection before `click` fires;
capturing the text up front keeps that safe, and an in-flight press is no
longer allowed to unmount its own target.

A visible popup also described the previous selection for up to the settle
window, so a tap while dragging a native selection handle queued the stale
excerpt. It is dropped as soon as a differing selection starts settling.

Finally, `viaTouch` survived from the last press into keyboard-driven
selections on hybrid devices, which could flip the popup into the touch layout;
keydown clears it.

The cancel path is covered by a new touch spec, verified to fail against a
commit-on-press build.

* 🧵 fix: Reconcile Cancelled Presses, Widen Clipping, Steady the Scroll Specs

Second review round, with one finding taken on trust and flagged rather than
claimed as proven.

A cancelled touch press could leave the popup backed by a selection that no
longer existed. A press deliberately keeps the button alive through a
collapsing selection so the release has a target to be judged against, but a
cancel then dropped the press without ever honouring the collapse it had
masked, so a later tap could add a dead excerpt. Ending a press without
committing now rechecks the live selection and dismisses if it went away.

Visibility now intersects every clipping ancestor of the message rather than
stopping at the nearest. This one is precautionary, not a proven fix: the
review that prompted it describes scroll containers *inside* a message (a wide
table, a code block) shadowing the outer chat scroller, but the walk starts
from the message element, so those are descendants and were never in the chain.
Behaviour is unchanged in the current layout — a spec covering a table-cell
selection passes identically with and without it — and it is kept only because
intersecting the whole chain stays correct if the list is ever nested inside a
further-clipped panel. The comment says exactly this.

The scroll specs were the real instability. They now move the selection between
two positions that are both on screen instead of nudging by a pixel count:
blind nudges kept pushing it under the composer, where the popup correctly
hides, and the chat's own auto-scroll made the landing spot unpredictable. They
also target the opening paragraph, since the closing one is the last content in
the conversation and cannot be carried upward from a list already at maximum
scroll.

The reply fixture gained a table so a selection inside a nested scroll container
is exercised, and a spec covers the cancelled press.

15/15 pass locally.

* 🪟 fix: Judge Quote-Popup Visibility From the Selection, on Both Axes

Third review round. All three findings held up, and each now has a spec that
fails without its fix.

Clipping is now measured from the selection rather than from the message, and
on both axes. A wide table or a long code line scrolls inside its own container
— and `overflow-x: auto` makes the computed `overflow-y` auto, so it clips
vertically too — which means scrolling it sideways carries the selected text out
of view while the message never moves. Walking up from the message could not see
those containers at all, and a vertical-only test could not see that motion.
This supersedes the previous round's precautionary widening, which was kept
without evidence; the evidence is now a spec that scrolls a table past its own
selection.

Publishing a settled selection also checks visibility. Nothing is tracked during
the 300ms settle interval, so a scroll inside that window never reached the
re-anchoring path, and the reading was published off-screen and then clamped
into view — stranding the popup over unrelated UI.

The cancelled-press spec now reproduces the ordering it describes. Collapsing
the selection and cancelling in one synchronous block let the asynchronous
`selectionchange` arrive after the press had ended, which is the ordinary path
and passes either way; it now waits for delivery in between, so the collapse
lands while the press is still masking it. Two other specs needed the same
scrutiny: `toBeHidden` is satisfied by an element that does not exist yet, so
the settle spec sits out the interval before asserting, and it scrolls just past
the container edge rather than to the end of the conversation, because a violent
scroll re-renders the messages and drops the selection for unrelated reasons.

The reply fixture's table is now wide enough to overflow sideways.

17/17 pass, and each new spec was re-run against a build with its own fix
reverted to confirm it fails there.
2026-08-13 00:36:45 -04:00
Danny Avila
df6e15a0de
🔖 feat: Bound Parent Activity Phases With an Exclusive End Index (#14768)
* 🧭 fix: Finalize Parent Activity Phases at Run Completion

* 🧭 fix: Preserve Activity Phase Boundaries

* 🎨 fix: Format Activity Phase Boundary Check

* 🧭 fix: Ignore Late Label Artifacts at Phase Completion

* 🧭 fix: Preserve Logical Activity Phase Membership

* 🩹 fix: Narrow Optional Activity Phase Marker

* fix activity phase tail boundaries

* fix activity phase test lint

* fix straddling activity phase batches

* preserve activity phase boundaries at scale

* fix persisted activity phase final boundary

* fix resumed activity phase edge cases

* fix sparse activity phase grouping

* fix sparse activity phase tail scan

* fix resumed activity phase text fallback

* fix sparse activity phase completion scans

* avoid sparse activity phase runtime scans

* stabilize sparse activity phase resumes

* support activity phases on current ts target

* preserve sparse phase reservations

* finalize activity phase boundary handling

* avoid sparse phase start scans

* fix activity phase final text bounds

* tighten activity phase summary boundaries

* format activity phase boundary checks

* leave final commentary outside activity phases

* recognize lane-tagged final activity text

* rebase retained activity boundaries on resume

* bound activity phase collection work

* correct resumed phase activity count

* resolve late reasoning before phase completion

* preserve lane-tagged final answers

* assert durable activity phase bounds in e2e

* preserve empty finalized activity phases

* ignore empty reasoning at phase completion

* format phase completion guard

* fix(api): retain overflow reasoning anchors

* perf(api): index overflow reasoning anchors

* perf(api): skip empty reasoning index scans

* fix(api): reconcile completion boundaries efficiently
2026-08-12 23:43:35 -04:00
Danny Avila
1a3e2aebcb
🛰️ fix: Attach Request-Scoped MCP Servers (#14780)
* fix: attach request-scoped MCP servers

* fix: satisfy MCP static checks

* fix: format MCP runtime hint
2026-08-12 23:43:02 -04:00
Danny Avila
c44d11ebf4
🧾 test: Pin the Transactions Config Wiring on the Fallback Path (#14779)
Follow-up to #14774. Its tests cover `AgentClient.recordTokenUsage` in
isolation, so the `BaseClient` half of the fix was unpinned: deleting the
`transactions` property from the call site restored the bug with the suite
still green.

These cases drive `sendMessage` with a real app config on `req` and assert the
resolved value reaches `recordTokenUsage` — disabled, the default when no
config is present, and the balance-enabled override that force-enables it.
Each fails if either half of #14774 is reverted.

They also isolate `options.endpoint` for the block. `options` is shared across
this file, and an endpoint left behind by an earlier case routes the
balance-enabled arrangement into `checkBalance`.
2026-08-12 23:42:23 -04:00
Danny Avila
298a3d9ee9
📦 chore: Update @librechat/agents to v3.4.6 (#14781) 2026-08-12 23:42:13 -04:00
Danny Avila
3a3a8dcad0
🖍️ refactor: Typed Console Colors and Hardened Script Helpers (#14778)
* 🛠️ refactor: Enhance console color handling and improve deleteNodeModules function

* refactor: Use coloredConsole for consistent console output in invite-user.js

* refactor: Convert year to string format in invite user payload
2026-08-12 22:48:36 -04:00
Danny Avila
8f1f961212
🧱 refactor: Require Broad Config Management for Base Field Mutations (#14775) 2026-08-12 22:32:38 -04:00
James Todaro
e696b07619
🧾 fix: Honor Disabled Transactions on the Token-Count Fallback Path (#14774)
`AgentClient.recordTokenUsage` had no `transactions` parameter, so the setting
never reached `createTransaction`, whose guard reads it from the object it is
handed. `transactions?.enabled === false` saw `undefined` and the write went
ahead.

This path is reached only from `BaseClient`'s fallback branch, when the provider
returns no usable stream usage, so the bulk path masked it wherever usage is
reported. Where it is not, the setting had no effect at all.
2026-08-12 22:31:13 -04:00
Danny Avila
dccef82254
🪶 chore: Aggregate Empty MCP Tool Logs (#14767)
* fix: aggregate empty MCP tool logs

* fix: retain server names in MCP tool logs
2026-08-12 22:29:49 -04:00