* 🤖 feat: make event actor HITL durable
* 🤖 fix: break event actor outcome cycle
* 🤖 fix: close durable actor terminal races
* fix: harden durable event actor recovery proofs
* fix: close event actor resume publication races
* feat: wire Keenable web-search provider into config, schema, and UI
Keenable landed as a search provider in @librechat/agents (#285, shipped in
3.2.58+), but LibreChat did not yet expose it. This adds the config/schema/UI
glue so it can be selected, mirroring the existing Tavily provider.
- data-provider: add `keenable` to SearchProvider type + SearchProviders enum,
keenableApiKey/keenableApiUrl schema fields, and a keenableSearchOptions block
(maxResults, site, attributionTitle, timeout).
- data-schemas: register keenable in webSearchAuth.providers and default the
key/URL placeholders in loadWebSearchConfig.
- api/web: pass keenableSearchOptions through to the provider and handle
Keenable's keyless model. Unlike other providers it authenticates with no key
(the public endpoint), picking up an optional key/URL when set; the URL
override is SSRF-preflighted like other user-provided URLs.
- client: add Keenable to the provider dropdown with an optional API-key input.
- docs: document KEENABLE_API_KEY/KEENABLE_API_URL in .env.example and a
webSearch example in librechat.example.yaml.
- tests: keyless + keyed auth resolution, config defaults, and schema parsing.
* fix: ESLint no-unused-vars and clarify Keenable yaml example
- Remove the now-unused RerankerTypes import in data-schemas web.ts (the lint
job runs with --max-warnings 0 on changed files, so this latent warning failed
CI once the file was touched).
- Note in the librechat.example.yaml Keenable stanza that a scraper (and
reranker) is still required for web search to load, and include a Firecrawl
scraper in the example.
* chore: fix import order drift (sort-imports)
* feat: add Keenable as a keyless scraper and select it without a pinned provider
The Keenable scraper landed in @librechat/agents#337, so wire the scraper
category the same way the search provider already is: `scraperProvider:
keenable` reads pages through Keenable's public fetch endpoint with no key
(a key only lifts rate limits, and the endpoint is overridden with
KEENABLE_FETCH_URL). Paired with `rerankerType: none` this makes a fully
keyless web-search stack possible for the first time.
Also closes the Codex finding on this PR: because none of Keenable's auth
fields are required, the generic auth loop skips it whenever it isn't pinned,
so a key submitted through the API-key dialog (which cannot pin a provider)
left the providers category unauthenticated. Keenable is now selected in that
case, gated on one of its values actually being present so installs that
configured nothing keep their current behavior. The scraper gets the same
fallback, additionally gated on Keenable being the resolved search provider,
so it never silently scrapes for another provider.
* fix: select the Keenable scraper from a supplied key, not only for Keenable search
The API-key dialog submits credentials and cannot pin a provider, so choosing
Keenable as the scraper while search stays on Serper/SearXNG/Tavily had no
effect: the unpinned-scraper fallback required Keenable to also be the resolved
search provider.
A supplied Keenable value now triggers it as well, which is the only signal the
dialog can send. The fallback still runs only when no keyed scraper
authenticated, and with neither trigger the category stays unauthenticated, so
a deployment that never configured Keenable is unaffected.
Note the fully keyless choice still cannot be expressed through the dialog:
Keenable's key is optional, so picking it with no key submits nothing at all.
librechat.example.yaml now documents pinning scraperProvider: keenable for that
case.
* style: Sort Keenable imports
* fix: Harden Keenable auth resolution
* fix: Preserve Keenable selection intent
* fix: Fail closed on invalid web search auth
* fix: close keenable auth gaps
* style: sort web auth imports
* fix: preserve web search selection integrity
* fix: isolate web search auth ordering
* fix: silence expected credential misses
* chore: bump agents sdk
* fix: preserve web search preference ownership
* fix: forward cleared Keenable endpoint
* style: sort web search hook imports
* fix: require intent for credential clears
---------
Co-authored-by: Ilya Bogin <ilya.bogin@keenable.ai>
* fix: upgrade redis dependencies and code to avoid elasticache bigint bug
* fix: preserve tls uri behavior with the node-redis v5 changes
* fix: satisfy node-redis v5 socket typings and clear lint in touched specs
The TLS spec passed `socket: { ca }` without `tls: true`, which node-redis
v5 accepts at runtime (the rediss:// scheme sets the flag) but its typings
reject, failing the type check. Assert the resolved socket options instead,
which covers scheme inference in both directions rather than only that the
constructor does not throw.
The benchmark spec carried two lint warnings that predate this branch and
only surface because CI lints changed files with --max-warnings=0: an unused
cache binding and a test with no assertions. Drop the binding and assert the
SCAN actually yielded keys, which is the behavior the page flattening
changed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014sYJcABr6NEFmVvPxsWhfy
---------
Co-authored-by: Arnau Berenguer Jiménez <arnau.berenguer@vista.com>
Co-authored-by: NoOPeEKS <arnauapps@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
* ⬆️ chore: Bump `@librechat/agents` to v3.6.0
Bumps the pin in `api` and `packages/api` from `^3.5.1` to `^3.6.0`. The
caret on `^3.5.1` cannot cross the minor, so both manifests and the lockfile
need the explicit bump.
v3.6.0 contains three changes over v3.5.1, all additive:
- `fix: Close Subagent Child-Graph Run Steps` — subagent child graphs run via
`workflow.invoke()` outside `Run.processStream`, so the terminal sweep never
reached their steps. They now close on both the success and error paths,
which is what makes `on_run_step_closed` reliable for subagent tool cards.
- `fix: Restore Run Steps Across Process Resumes` — open run-step lifecycle
state is now persisted in LangGraph checkpoints, so a step opened by one
process closes correctly after a resume on another.
- `feat: route code execution per agent profile` — new optional
`codeSessionKey` partition for code-session ids and file refs.
No breaking changes: every new field on the public type surface is optional,
and the package's own dependency set is unchanged between the two versions
(verified against the registry), so the lockfile diff is limited to the
`@librechat/agents` entry itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5
* 🔒 chore: Sync `bun.lock` with the agents v3.6.0 bump
`bun.lock` still recorded both workspace requirements and the resolved
package as `@librechat/agents@3.5.1`, which no longer satisfies `^3.6.0`, so
`bun install --frozen-lockfile` would reject the committed state.
`bun install --lockfile-only` cannot run in this environment: bun stores no
integrity for the `xlsx` URL dependency and therefore re-fetches
`cdn.sheetjs.com`, which the sandbox network policy denies (403 on CONNECT).
The entry was updated directly instead, which is exact here because the
package's dependency graph does not move between the two versions: its
`dependencies`, `peerDependencies` and `optionalPeers` at 3.6.0 are identical
to 3.5.1 (checked against the registry), so only the version, the resolution
id and the integrity hash change. The integrity matches the one npm resolved
into `package-lock.json`, and the two existing
`@librechat/agents/*` hoisting overrides stay valid because the dependency
set they resolve is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014vLhxCFMYkCaTsoFTiAjJ5
---------
Co-authored-by: Claude <noreply@anthropic.com>
`@modelcontextprotocol/sdk@1.30.0` is a small maintenance release on the 1.x line
(upstream's active line is now the 2.0.0 scoped packages). The range was already
`^1.29.0`, so only the lockfile pinned the old version; the manifests move too so
the floor matches what we test against.
Nothing in it is breaking. The four changed type declarations are additive —
optional `maxBufferSize` on `StdioServerParameters`, an optional third
constructor argument on `StdioServerTransport`, optional options on `ReadBuffer`,
optional `keepAliveMs` on the server transport — and the only manifest change is
`@hono/node-server` widening to `^1.19.9 || ^2.0.5`. No new dependencies.
Two behavior changes are worth knowing about even though neither is an API break.
`ReadBuffer` now caps a single stdio message at 10 MB (previously unbounded) and
errors the transport instead of growing, which is reachable through
`StdioClientTransport` if a stdio server returns a very large single result; it
takes `maxBufferSize` if that ever needs raising. And Content-Type handling
switched from substring search to parsed media types, client and server.
Most of the release is Streamable HTTP server hardening we do not run — a 15s SSE
keep-alive, `X-Accel-Buffering: no` on SSE responses, guards so a stale stream's
cancel cannot tear down its successor, and `_closed` checks so a transport closing
mid-request stops registering streams into swept maps. None of it changes how we
behave as a client. In particular it does not address the stale-stream 409 in
#14816: that keep-alive runs in whichever server we connect to, not here.
The same substring-vs-parse mistake the SDK corrected exists in our streamable
HTTP response guard, which classified a response as SSE with
`contentType.includes('text/event-stream')`. A `Content-Type` naming the SSE type
in a parameter — `text/plain; boundary=text/event-stream` — is not an event
stream, but matched. The guard then took `canEmitFallbackSSEError`, so an
oversized body was answered with a synthetic SSE error frame the caller reads as
a well-formed response body, rather than the throw a non-SSE response gets. The
check now compares the parsed media type, via a `mediaTypeEssence` helper added
to the header utils where `mergeHeaders` already lives.
Verified against 1.30.0 rather than assuming: the package was staged into the
worktree's own `node_modules` so it shadowed the shared install, and
`packages/api` `src/mcp` ran green on it — same four pre-existing red suites as
on 1.29.0 (`MCPReinitRecovery` plus three Redis `cache_integration` suites that
need a live Redis), no new failures.
* ⚡ feat: Add Gemini 3.7 Flash Support
Adds first-class support for Google's Gemini 3.7 Flash (`gemini-3.7-flash`)
for both the Gemini API (AI Studio) and Google Cloud Gemini Enterprise Agent
Platform, following the Gemini 3.6 Flash integration (#14369).
- Context window (1,048,576) in googleModels; API + cache pricing in tx.ts.
- Model dropdown (config.ts) and GOOGLE_MODELS examples for both integrations.
- Register the model in the Flash-family handler so it inherits the existing
strip of deprecated sampling params (temperature/topP/topK), rejected
penalty params, and thinkingBudget, and defaults to `medium` thinking.
- Generalize that handler's enumerated table from a [id, level] tuple to a
rule object, so a model can also declare thinking levels it rejects. Gemini
3.7 Flash errors on `minimal` (which the Google endpoint offers in its
thinkingLevel slider), so an explicit `minimal` is substituted with the
nearest supported level, `low`. Explicit low/medium/high pass through
unchanged.
- Apply Google's introductory pricing ($0.75 in / $3.75 out / $0.075 cached,
per 1M) to Gemini 3.7 Flash and correct Gemini 3.6 Flash to the same rates.
Both revert to $1.50 / $7.50 / $0.15 on 2027-01-01; noted at both call sites.
Resolves#14802
Ref: https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash
Ref: https://ai.google.dev/gemini-api/docs/pricing
* 📝 docs: Match the House Style for Promotional Rate Comments
Align the Gemini 3.6/3.7 Flash introductory-pricing notes with the existing
Sonnet 5 convention in the same file: one comment per group, naming the models
and the exact values to restore, so the manual follow-up is unambiguous.
No rate changes.
* ⬆️ chore: Bump `@librechat/agents` to 3.4.7 for Gemini 3.7 Flash Prefill
Unblocks this PR. `NO_PREFILL_GEMINI_MODELS` is model-enumerated in the agents
SDK, so 3.4.6 does not know `gemini-3.7-flash` forbids a trailing `model`-role
turn — editing an assistant reply and resubmitting would reach Google as a
prefill and return HTTP 400 on a model this PR adds to the default list.
3.4.7 (danny-avila/agents#412, released via #413) adds it. Verified the
published tarball: `3.4.6...3.4.7` touches only
`dist/{cjs,esm}/llm/google/utils/common.*` — the prefill array and its comment.
`dist/types` is byte-identical, so there is no API surface change.
Raises the declared range in both workspaces alongside the lock. `^3.4.6`
already permitted 3.4.7, but the fix is required rather than merely compatible,
so the floor should say so.
* 📦 chore: bump `@librechat/agents` to version 3.4.2
* 📦 chore: bump `mermaid` to version 11.16.1 and update related dependencies
* 📦 chore: bump `js-yaml` to version 4.3.1 in package-lock and data-provider
* 📦 chore: bump `nanoid` to version 3.3.18 in package.json and package-lock.json across multiple packages
* 🔧 fix: Remove stray `api/tsconfig.json` breaking e2e `~` alias
An empty `api/tsconfig.json` was accidentally committed with the agents bump.
Playwright's require hook resolves path aliases from the nearest path-config,
checking `tsconfig.json` before `jsconfig.json` in each folder, so the empty
file shadowed `api/jsconfig.json` — the only place `"~/*": ["./*"]` is defined.
Every e2e spec that calls `cleanupUser` then failed on
`Cannot find module '~/cache/getLogStores'` from `api/models/index.js`.
- delete the stray file and gitignore it so tooling can't re-commit it
- register `module-alias` in `cleanupUser` so backend requires resolve
regardless of which path-config Playwright happens to find
* 📦 chore: bump `@librechat/agents` to version 3.4.3 in package.json and package-lock.json
* test: cover streamed subagent results end to end
* test: assert real e2e conversation id
* test: harden streamed subagent e2e
* test: stop incompatible subagent fixtures
* chore: update @librechat/agents to version 3.4.0 in package.json and package-lock.json
* fix(langfuse): disable central fanout media uploads
* test(langfuse): cover fanout media policy in run config
* chore(deps): bump agents for Langfuse media policy
* chore(deps): bump agents to 3.3.13
* fix(langfuse): gate central fanout media uploads
* 🛡️ fix: Run message-filter PII patterns on a linear-time regex engine
The messageFilter.pii middleware compiled admin-configured customPatterns with the native RegExp engine and ran them synchronously against every message on the shared event loop, so a catastrophic-backtracking pattern such as (a+)+$ could stall the entire process (native RegExp takes tens of seconds at roughly 32 characters) and take the instance down for every user.
Compile these patterns with RE2JS, a linear-time RE2 port with no native addon, so catastrophic backtracking is impossible regardless of the pattern rather than something the code tries to detect. Patterns using features RE2 does not support, such as backreferences, fail to compile and are dropped and logged exactly as an invalid pattern already is. The filter only tests for a match, so this is a drop-in engine swap with no behavior change for valid patterns.
* 🛡️ fix: Reject RE2-incompatible messageFilter patterns at config load
The customPatterns regex was validated with native RegExp at config load, but the runtime now compiles it with a linear-time engine (RE2) that does not support backreferences or lookaround. Such a pattern passed validation, then failed to compile and was silently dropped at request time, quietly removing PII protection after upgrade.
Reject backreferences and lookaround during config validation with an explicit message, and document RE2 syntax in the example config instead of "JavaScript-flavor". The runtime engine remains the authoritative boundary and still drops-and-logs anything this load-time check misses.
* 🧹 test: Use direct MessageFilterPiiConfig annotations in the PII specs
The added ReDoS cases satisfy the exported MessageFilterPiiConfig type directly, so the `as unknown as` assertions were unnecessary. Annotate the config objects directly, matching the repo's type-safety guidance.
* 🧹 fix: Reject named backreferences in messageFilter patterns at config load
Extend the config-load check to also reject named backreferences (\k<name>), which are valid JavaScript regex but unsupported by the linear-time runtime engine, so they surface at load rather than being dropped at request time. Together with the existing numeric-backreference and lookaround checks this covers the RE2-incompatible construct set; the runtime engine remains authoritative.
* 🛡️ fix: Preserve Unicode whitespace matching in messageFilter starter patterns
RE2's \s is ASCII-only, so after the engine swap the built-in api-key and Bearer starters no
longer matched a secret separated by non-ASCII whitespace (e.g. a non-breaking space), which
native RegExp did match. Broaden the whitespace classes to [\s\p{Zs}] so those patterns keep
their original coverage, and add a regression test for a non-breaking-space separator.
* 🛡️ fix: Validate messageFilter patterns with the RE2 engine at config load
Replace the syntax blacklist (numeric/named backreferences, lookaround) with authoritative
validation: config load now compiles each custom pattern with the same linear-time engine the
runtime uses, so any RE2-incompatible construct (including control escapes like \cA) is rejected
at load with a clear error instead of being silently dropped at request time.
The validator is swappable and defaults to native RegExp so browser builds add no engine; the
server wires the RE2-backed check at startup via configureMessageFilterRegexValidator in both
entry points.
* 🛡️ fix: Match the full whitespace set in messageFilter starter patterns
RE2's `\s` omits the vertical tab and `\p{Zs}` omits U+2028, U+2029, and
U+FEFF, so a separator built from one of those characters slipped past the
`api-key` and `Bearer` starter patterns and reached the model. Broaden the
starter whitespace class to the full JavaScript whitespace set so those
separators are covered again.
* fix: fail closed when messageFilter.pii compiles to zero patterns
DB and admin config overrides bypass the RE2 schema validation (it only
runs at YAML load), so an override whose only pattern is RE2-incompatible
was dropped at compile time, left zero patterns, and let the request
through. compile() now returns a failClosed flag when a config declared
patterns but every one failed to compile; the middleware returns 400 and
findPiiMatchInMessages returns a distinct misconfigured match that the
OpenAI and Responses controllers surface with an admin-facing message.
* 🛡️ fix: Fail closed when any messageFilter.pii custom pattern drops
compile() previously set failClosed only when every pattern dropped (patterns.length === 0 && dropped > 0). With the default starters present, a single RE2-incompatible custom override incremented dropped but left patterns.length > 0, so the filter silently enforced only the surviving subset and text matching only the dropped rule passed.
failClosed now keys off dropped > 0, so any dropped custom pattern blocks with the misconfigured 400. YAML patterns are RE2-validated at load, so dropped stays 0 for valid configs and only unvalidated DB or admin overrides can trip it. Reframed the two keeps-others-active specs to assert fail-closed and added a default-starters partial-drop regression.
* 🧹 fix: Correct the misconfigured JSDoc and drop redundant casts in the PII specs
The misconfigured flag now means any configured custom pattern failed to compile, not that every pattern failed, so its JSDoc on PiiMatch is updated to match. The partial-drop regressions now use direct MessageFilterPiiConfig annotations instead of as-unknown-as casts, keeping the specs type-checked, consistent with the rest of the suite.
* 🚦 feat: Configurable Circuit Breakers for Runaway Streamed Tool Args
* docs: forewarn create_file about the streamed tool-argument limit
The breaker failing a near-limit write should not be the model's first
exposure to the bound. Both create_file variants now state the default
64 KB per-call limit and the incremental pattern (create the first
section, extend with edit_file) in the tool description and the content
parameter description.
* fix: keep skill create_file description under the provider advisory cap
The limit-guidance paragraph pushed the skill-aware description to 1169
chars, past the 1024-char advisory bound where providers may truncate.
The skill variant now carries the guidance only in its content parameter
description, which sits closest to the generated payload and is not at
truncation risk; the shorter code-sandbox variant keeps the full
paragraph.
* 🚦 feat: per-tool streamed-arg limits with a create_file default
Thirty days of production data show create_file is the only tool class
with legitimate near-limit arguments (p99 80.6 KiB; every other tool
p99 under 10 KiB). Rather than loosening the global 64 KiB cap for all
tools, the yaml gains maxToolCallArgBytesByTool (per-tool overrides,
keyed by model-facing tool name, 0 disables that tool's guard) and
LibreChat ships { create_file: 131072 } by default; yaml entries merge
over and can replace it. Pairs with maxToolCallArgBytesByTool support
in the agents SDK and stays inert until the dependency bump.
* test: pass per-tool spec configs as plain Partial literals
The as-TAgentsEndpoint casts fail TS2352 for object-valued fields:
comparability does not grant nested literals the implicit index
signature that plain assignability does, so casts carrying
maxToolCallArgBytesByTool never sufficiently overlap. The mapper
already accepts Partial<TAgentsEndpoint>, so the new cases pass
uncast literals instead.
* chore(deps): bump @librechat/agents to 3.3.12
Activates activity-label continuity end to end. The host side landed with
the activity-groups feature (#14391) — the per-run accumulator that reads
committed headers at request-build time, the `previousLabels` payload
field, resume seeding, and the bridge passthrough — but the SDK had no
field to receive them, so the traced generation path ignored the context
and only the direct fallback rendered it. v3.3.8 carries
danny-avila/agents#356, which adds `previousLabels` to
`RunActivityLabelOptions` and renders it as the label prompt's first
section (capped at 3, each entry whitespace-collapsed and clipped at 200
chars so one malformed header cannot forge prompt sections or inflate
every later request in the run).
Effect: consecutive same-activity batches now extend the run's story
instead of restating a line already on screen, and setup batches stop
being labeled with conclusions their tools had not yet established.
Also included between v3.3.7 and v3.3.8: danny-avila/agents#354, which
anchors summary coverage to a source message id.
Lockfile carries no transitive churn — 3.3.8's dependency tree is
identical to 3.3.7's.
Verified against the installed package: 73 activityLabels tests and 218
api agents-controller tests pass, `tsc --noEmit` clean on packages/api,
and the published build renders the capped, sanitized header section
(oversized labels clipped, embedded newlines flattened to inert text).
* 🎯 feat: Tool Intent Label Capability (tool_intents)
Adds the fourth member of the per-tool capability family (defer_loading,
allowed_callers, run_in_background): an admin capability
AgentCapabilities.tool_intents plus a per-tool
tool_options[name].describe_intent flag. Opted-in tools get an optional
intent string injected as the FIRST property of their schema — one
model-authored sentence per call, streamed to the client as the call's
live status label (args already reach the client verbatim, so no new
event plumbing). Native host tools (web_search, create_file/edit_file,
set_memory/delete_memory, ask_user_question) default on while the
capability is enabled; explicit false opts out. SDK-native intent
schemas (@librechat/agents coding suite) are recognized and left alone.
- packages/api/src/agents/intent.ts: structural sibling of
background.ts — first-key non-mutating injection with registry
parity (covers deferred/tool_search discovery), eligibility and
PTC-only skips, arg read/strip helpers, self-spawn strip for defs and
registry, ephemeral/model-spec synthesis with a tool_options merge so
the background and intent toggles compose.
- handlers.ts: intent runs BEFORE background injection so the label
stays the first streamed key when a tool carries both (pinned by
test); the arg is stripped before invocation unless the tool's own
schema declares it, on both the foreground and background-dispatch
paths; PTC target schemas are sanitized like background's.
- Capability plumbing through all four routes (endpoint initialize,
openai + responses controllers, the exported OpenAI-compatible
service) plus handoff discovery and added-convo agents, and the
intentToolNames execution channel via configurable.
- describe_intent on toolOptionsSchema (all three written-out Zod
annotations), ToolOptions, TEphemeralAgent, TModelSpec (+ zod), and
data-schemas doc comments (tool_options is Mixed — no migration).
- intent.spec.ts: 28 tests cloned from background.spec.ts structure,
including the intent+background key-order composition.
* 🧯 fix: Codex Review — Opt-Out Strips SDK-Native Intent, Skip mcp_all Placeholders
- An explicit describe_intent: false now REMOVES an SDK-native intent
property from the definition and registry entry, so the per-tool
opt-out actually disables the arg's token cost for tools like
web_search that carry the schema natively (SDK bodies tolerate its
absence). Previously the early return left the property in place.
- synthesizeIntentToolOptions skips lazily-expanded mcp_all
placeholders instead of recording options under names that
applyIntentLabels' exact-name matching can never match, and documents
the limitation (parity with synthesizeBackgroundToolOptions).
The P1 about the client not rendering the label is the documented
slicing: the UI streaming-label PR follows once #14391's ToolCallGroup
changes merge — args already reach the client, so that slice is purely
rendering.
* 🧯 fix: Codex Re-Review — Label Marker Guard, Capability Kill Switch, Late Defs, Service Threading
- removeIntentParam is now marker-guarded (the label contract's opening
instruction discriminates it), so an MCP/action tool's own business
`intent` parameter is never stripped by an opt-out or the disabled
path — previously an explicit false could remove a real, possibly
required argument.
- New sanitizeIntentLabels pass runs AFTER every registration step
(the skill catalog appends its SDK definition post-injection): with
tool_intents disabled it strips SDK-native intent labels from all
definitions and registry entries, making the capability a real kill
switch over their token cost; with it enabled it enforces explicit
per-tool opt-outs on late-registered definitions.
- ask_user_question removed from the native default-on set: its graph
tool is rebuilt in run.ts from its own Zod schema (also the HITL
card's wire shape), so definition-level injection never reached the
model. Its intent support lands with the HITL slice, which threads
the label into the interrupt payload deliberately.
- The exported OpenAI-compatible service now threads intentToolNames
into the run configurable, so the executor's PTC path can strip
host-injected intent schemas on that route like the in-repo
controllers do.
* 🧯 fix: Codex Round 2 — Post-Skill Injection, PTC Native Strip, Service Boundary, Honest Docs
- Intent injection now runs LAST in initializeAgent, after the skill
catalog — which both appends its own definition and REPLACES upgraded
ones (skill-aware read_file), clobbering an earlier injection while
intentToolNames still listed the tool. Injection PREPENDS while
background APPENDS, so intent stays the first schema property under
the new ordering (pinned by a reverse-order composition test).
- The PTC target-schema strip is now marker-guarded strip-ALL: SDK-
native intent labels (which are deliberately never in intentToolNames)
are removed from sandbox-advertised schemas alongside host-injected
ones; business intent params survive.
- toolIntentsAvailable on the exported service documents the loader
boundary: a custom LoadToolsFn returning only structured instances
bypasses definition/registry injection and sanitize by construction.
- librechat.example.yaml describes tool_intents as backend groundwork
with UI rendering in an upcoming release rather than promising a live
label today.
* 📦 chore: bump `@librechat/agents` to v3.3.6
Brings in the SDK half of tool intent labels (danny-avila/agents#347,
#349): intent-first schemas on the coding suite across all three
engines, plus web_search / subagent / skill / tool_search, and the
outcome / outcome_patch result channel.
Activates three host paths that were inert while no SDK tool shipped an
`intent` property — verified against the real 3.3.6 schemas:
- capability OFF now strips SDK-native labels (a real admin kill switch)
- explicit `describe_intent: false` removes them per tool
- host injection stays idempotent against an SDK schema, keeping
`intent` first and never double-injecting
* 🔬 test: Real-Provider Verification for Tool Intent Labels
Adds the live check the unit tests structurally cannot perform: whether a
real model actually authors the injected arg, places it FIRST, and gives
sibling calls to one tool distinct labels. Reuses the existing
real-provider harness (in-memory Mongo, seeded user, credential
neutralizer) and the existing stdio MCP fixture as a genuine tool, so no
external service is involved.
- e2e/config/librechat.real.yaml: adds the e2e-memory MCP server and the
tool_intents capability, giving the real model something to call. The
sibling spec asserts only relative token growth, so the extra schemas
do not perturb it.
- e2e/playwright.config.real.ts: optional Langfuse passthrough. The
LANGFUSE_* keys match the credential-neutralizer pattern and were being
blanked before the server booted; they are preserved explicitly, read
from the invoking environment only, and never written to the generated
config.
- e2e/specs/real/tool-intents.spec.ts: two facts stored in one turn, both
through the same tool, asserting intent is the first key of each call
and that the two labels differ. Args are read from persistence rather
than the DOM deliberately — no UI renders the label yet, and
persistence is what a reloaded conversation and the trace both read.
First run against claude-haiku-4-5 produced 'Recording the location of
the OAuth callback router' and 'Recording the location of the MCP
connection pool configuration' — distinct, first-position, no tool name.
Also updates tool-intent-spec.md: records the 3.3.7 removal of the tense
verb map with the evidence that motivated it, the trimmed description and
the marker's role as an API, and a new mandatory requirement that
client-side label rendering be gated on a server-sent signal rather than
the presence of an intent key (a tool's own business 'intent' parameter
would otherwise render as a status label).
* 📦 chore: bump `@librechat/agents` to v3.3.7 and dedupe the intent contract
Picks up danny-avila/agents#353: the tense verb map is gone (a bare
intent now displays unchanged, with completion carried by UI state), the
model-facing description is trimmed 502 → 289 chars, and both the marker
and the description are exported.
Stops redeclaring the SDK contract here:
- INTENT_LABEL_MARKER is imported instead of duplicated as a string
literal. Every removal path in this module keys on it, and a local copy
that drifted from the SDK's would make them all stop recognizing
SDK-native labels — failing OPEN, with labels left in schemas and
per-tool opt-outs silently inert.
- INTENT_DESCRIPTION is imported too, so host-injected tools and
SDK-native tools present the model with one identical instruction.
Keeping the old local copy would also have meant host-injected tools
still paying ~126 tokens per schema while SDK tools paid ~72.
Verified live against real Anthropic after the trim: two sibling calls to
one MCP tool produced 'Storing the OAuth callback router file location'
and 'Storing the MCP connection pool configuration file location' —
first-position and distinct, so the shorter description holds compliance.
Closes GHSA-664h-wqgq-64gw (CVSS 6.5, CWE-1321), a prototype pollution
in update casting via a __proto__-prefixed dotted path. Affected range
is >=8.0.0 <8.24.1, so 8.23.1 was flagged by npm audit.
* ⚡ feat: Add Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Support
Adds first-class support for Google's Gemini 3.6 Flash (`gemini-3.6-flash`)
and Gemini 3.5 Flash-Lite (`gemini-3.5-flash-lite`) for both the Gemini API
(AI Studio) and Google Cloud/Vertex integrations.
- Context window (1M) in googleModels; API + cache pricing in tx.ts.
- Model dropdown (config.ts) and GOOGLE_MODELS examples for both integrations.
- Generalize the Gemini 3.5 Flash overrides into a flash-family handler that
strips deprecated temperature/topP/topK and applies each model's default
thinking level (3.6 Flash: medium, 3.5 Flash-Lite: minimal), with
longest-prefix resolution so flash-lite does not collide with flash.
Ref: https://ai.google.dev/gemini-api/docs/latest-model#api-changes-and-parameter-updates
* 🩹 fix: Strip unsupported penalty params for Gemini Flash family
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash reject presencePenalty/
frequencyPenalty with HTTP 400 ("Penalty is not enabled for this model",
verified live). These pass through llmConfig via knownGoogleParams, so add
them to the flash-family strip list alongside the deprecated sampling params.
* 🩹 fix: Strip Flash-blocked params on custom Google endpoint path
For custom OpenAI-compatible endpoints with defaultParamsEndpoint=google,
getOpenAIConfig strips Flash-blocked params via getGoogleConfig but then
transformToOpenAIConfig re-applies raw addParams, undoing the strip. Filter
addParams through stripGeminiFlashBlockedParams before the transform so the
deprecated sampling / rejected penalty params cannot reach the provider.
* 🔧 chore: Update sharp package to version 0.35.3 in package-lock.json, api/package.json, and packages/api/package.json
* 🔧 chore: Update dependencies in package-lock.json to latest versions for @google/genai (2.13.0), @hono/node-server (1.19.14), fast-uri (3.1.4), hono (4.12.31), and svgo (2.8.3)
* 🔧 chore: Update dependencies in package.json and package-lock.json for @librechat/agents (3.2.67), @opentelemetry/sdk-node (0.221.0), and add new dependencies for @opentelemetry/propagator-jaeger (2.10.0) and protobufjs (7.6.5). Update monaco-editor version in client package.json to 0.56.0.
* 🔧 chore: Upgrade turbo package to version 2.10.5 in package.json and package-lock.json, and update schema reference in turbo.json
* 🩹 fix: Resolve CI breakage from bundled dependency bumps
Not related to the Gemini models — both are fallout from the dep bumps on
this branch:
- monaco-editor 0.56 changed IEditorHoverOptions.enabled from boolean to
'on' | 'off' | 'onKeyboardModifier'; update ArtifactCodeEditor to match
(mirrors the sibling occurrencesHighlight/matchBrackets pattern).
- sharp 0.35.3 fails resize+encode on a degenerate 1x1 PNG (vipspng: libpng
read error); the provider-file e2e fixture was 1x1, so use a 16x16 PNG.
Normal images are unaffected (verified 64x64 resize/encode/jpeg all OK).
* 📝 docs: Correct e2e image-fixture comment (bad IDAT CRC, not a sharp bug)
Root cause was the old 1x1 fixture's corrupt IDAT CRC (verified: IHDR/IEND
CRC OK, IDAT CRC BAD), which sharp 0.35.3's stricter libpng correctly rejects.
Not a dimension/resize edge case and not a sharp bug; comment now reflects that.
* 🧭 feat: Mid-Run Steering and Queued Messages for Agent Runs
Steering: submit a message while a run is generating; the server queues
it in the job store (cross-instance) and a run-scoped PostToolBatch hook
injects it into graph state at the next tool-batch boundary, records an
inline 'steer' content part on the response (replayed as a user message
on later turns), and streams on_steer_applied to the client.
Queuing: messages composed during a run auto-send as normal follow-up
turns after clean completion (one per final event, FIFO); user aborts
leave them as chips unless armed by interrupt-and-send.
Requires hook injectedMessages support in @librechat/agents
(danny-avila/agents#299); hard-gated via a capability probe so older
SDKs 501 the steer route instead of draining and dropping messages.
* 🧵 fix: Harden Steering Against Finalization Races and Route Guard Gaps
Addresses local Codex review findings on the steering feature:
- Close-and-drain the steer queue atomically at finalization (final event,
abort) so a steer POST racing teardown is rejected instead of 202-ACKed
and then silently cleared; the closed flag lives on the job hash and is
reset when a replacement job reuses the stream id.
- Clear inherited steer queues on createJob — a job replacement must not
drain the replaced run's messages.
- Keep steers queued across a HITL pause instead of draining them into
ephemeral client state: resumeState re-seeds chips on reload and the
resumed run injects them at its first tool boundary (steers key TTL now
extends to the approval window; on_steers_pending event removed).
- Queue the NO_ACTIVE_RUN steer fallback while the final SSE is still
settling — a direct send would be dropped by ask()'s in-flight guard.
- Reconcile the 202 ACK against on_steer_applied events that beat it over
the SSE, so a chip can't be re-minted after its removal event passed.
- Allow the per-send Steer override when the default action is queue.
- Apply the configured message rate limiters and the PII filter to
POST /chat/steer — a steer is model-bound user text.
* ✅ ci: Assert Steering Capability Probe Against the Installed SDK
CI installs the published @librechat/agents pin (pre-injectedMessages),
where isSteeringSupported() is legitimately false — the probe test now
asserts it mirrors the installed SDK's capability flag instead of
hardcoding the capability-bearing build's value. Verified against both
the published 3.2.61 dist and the agents#299 build.
* 🛟 fix: Preserve Steer Text Across Run-End, Error, and Abort Races
Codex round 2 (4 P2s):
- Applied-steer-id set survives run end (capped at 100) and converted
ids join it, so a 202 ACK that lands after final/abort drops its chip
instead of re-minting a stranded pending one.
- Failed runs no longer strand acknowledged chips: both error paths
convert local pending chips to queued follow-ups (chip text is
client-local), and the server closes the steer queue before emitting
the error so a racing steer POST gets 404 fallback instead of a 202
whose payload dies with the job.
- sendQueuedNow keys on steer availability, not the default action —
send-now on a queued chip is an explicit override for queue-preferring
users.
- Stop path consumes pendingSteers from the abort HTTP response as a
fallback for the SSE final event it may close before processing;
conversion is deduped so double delivery is a no-op (shared
useSteerConvert hook).
* 📎 feat: Carry Attachments Through During-Run Queued Messages
Steering stays text-only (SDK injection, inline STEER part, and replay
are all text), so a during-run submit with media now queues the whole
message as one unit instead of silently stranding the files:
- QueuedMessage gains `files`; composer attachments are consumed into
the queued item at queue time (steerFromComposer / queueFromComposer /
interruptAndSend), fixing the latent hazard where lingering composer
files glued onto whatever `ask` vacuumed up next.
- Enter-steer with attachments degrades to queue with an explanatory
toast; the per-send menu routes through the same composer-aware
wrappers.
- The drain and sendQueuedNow pass the item's files as `overrideFiles`;
media items never steer (send as a normal turn when idle, re-front
otherwise). ask() no longer clears composer state for caller-supplied
overrideFiles — only regenerate keeps that behavior.
- During-run submits hold while uploads are in flight, mirroring the
send button's filesLoading gate; queued chips show a paperclip count.
* 🎛️ feat: Rework During-Run Chips into Action Rows
Full-width rows above the composer (reference-UI parity): each queued
message shows a primary Steer/Send-now action, delete, and a "…" menu
with Edit message (restores text + attachments into the composer) and a
Turn on queueing/steering toggle that flips the Enter default. Steer
rows share the layout with status text; failed steers keep retry /
edit / queue-convert. The per-send menu gains the same default toggle.
Queued file refs now retain filename + bytes so edit-restore rebuilds
real composer entries (draft-recovery shape).
* 🖇️ feat: Steer With Attachments (Multimodal Mid-Run Injection)
Steering now carries media end-to-end instead of degrading to queue:
- The steer POST accepts sanitized attachment refs (cap 10; only
file_id is trusted — the drain re-fetches owner-scoped and re-derives
everything else). SteerQueueItem/TPendingSteer/SteerContentPart carry
`files` refs; encoded data is never persisted or queued.
- New api/server/services/Files/steering.js decouples attachment
building from the request path: encodeSteerContent reuses the exact
per-turn pipeline (addFileContextToMessage + processAttachments'
single-pass categorize/encode, SDK formatMessage assembly,
prependFileContext for extracted text) with zero new encoding code.
buildSteerMedia feeds the drain hook's new buildMedia seam (any
failure degrades that steer to text-only — words always land);
stampSteerPartMedia re-encodes past steer parts per turn with ONE
batched owner-scoped fetch and stamps a transient `media` array,
replaced immutably so it can never leak into a save. Replay honors
resendFiles like regular message media.
- The SDK's formatAgentMessages (the formatter agents actually use)
gained the steer replay branch on the PR branch; the local
formatMessages.js branch now mirrors the media preference.
- Client: steerFromComposer consumes composer files into the POST,
chips/seeding/conversions carry files everywhere (retry, queue
convert, abort/error recovery), queued media items steer for real,
and SteerBubble renders the steered attachments inline.
* 🧵 fix: Harden Steer Recovery Races and Drain Isolation
Codex round 3 (7 fixes):
- A 202 ACK landing after the run ended converts straight to a queued
follow-up (server queue is gone; no event will ever resolve a pending
chip for a finished run). Covers stream errors with in-flight POSTs.
- A Stop that lands pre-completion can arrive as a final with
unfinished:true and no aborted flag — runEnd now treats it as aborted
so queued messages are not auto-sent against the user's Stop.
- Leftover-steer conversion merges chronologically by createdAt instead
of appending, preserving the order the user composed.
- Auto-drained queued messages pass explicit (possibly empty)
overrideFiles/overrideQuotes/overrideManualSkills: a drain can no
longer vacuum up files, quotes, or skill picks staged in the composer
for the user's NEXT message (ask() treats overrideFiles != null as
authoritative).
- Failed-steer Retry and resume-on-load chip restoration keep the
steer's attachments.
- The job-replacement guard moved INSIDE the store's atomic
drain/close-and-drain (Lua createdAt compare; in-memory equivalent):
a stale run's hook or finalization can neither consume, close, nor
steal a replacement job's steer queue, and the drain hook drops its
separate check-then-drain round trip.
* 🧰 refactor: Typed Steer Controller, Single-Query Media Pass, Round-4 Fixes
Codex round 4 + efficiency tightening in one pass:
- Moved the steer guard ladder (validation, file sanitization via a
shared toSteerFileRef picker, ownership/tenant checks, status-guarded
enqueue) into packages/api as handleSteerRequest; api/steer.js is now
a thin wrapper. Ladder covered against the REAL in-memory job manager
in request.spec.ts; the api spec pins only the wrapper contract.
- Folded the steer replay stamp into the turn's ONE historical-files
query: collectHistoricalFileRefs also gathers steer-part refs, the
owner-scoped doc map rides client state, and stampSteerPartMedia
consumes it (no second round trip) while encoding parts in parallel.
- Stamped steer media now counts against the run budget (existing
multimodal counter over the non-text parts, folded into
indexTokenCountMap/promptTokens after the stamp).
- Steer route runs the PII filter BEFORE moderateText, matching chat.js
so blocked sensitive text never reaches the external moderation API.
- Interrupt & send survives the abort-response-beats-SSE-final race:
stopGenerating writes the run-end signal itself when the one-shot
interrupt flag is armed and no signal landed (double-fire safe).
- Resume reconciles chips against the server's still-queued list even
when EMPTY, clearing chips for steers applied while disconnected.
- The local formatter's steer flush preserves non-text assistant parts
(array-content AIMessage) instead of folding to text.
* 🔒 fix: Replay-Aware Capability Gate and Round-5 Race Closures
- isSteeringSupported now requires BOTH halves of the SDK contract:
injection (HOOK_INJECTED_MESSAGES_CAPABLE) AND replay
(ContentTypes.STEER, shipped in the same SDK commit as the
formatAgentMessages steer branch). An SDK that can inject but not
replay 501s the steer route — no release window can create steer
parts that would leak into provider-facing assistant content.
- The local formatter mirrors the SDK's anchor reset: a post-steer
tool_call mints a fresh AIMessage instead of attaching to the
pre-steer anchor (invalid provider ordering).
- Queued-chip send-now and the NO_ACTIVE_RUN fallback pass explicit
(possibly empty) overrideFiles so an idle send can't vacuum composer
files staged for a different draft.
- Redis createJob deletes the stale steer list BEFORE the replacement
hash is written as running — a steer 202-accepted against the new job
can never be wiped by the reset.
- Resumed-turn finalization mirrors the normal path's terminal drain:
createdAt-guarded close-and-drain, leftovers ride the resumed final
event as pendingSteers instead of being cleared by completeJob.
- buildSteerMedia restores composer order over the $in result so
multi-attachment steers reach the model in the order the user saw.
* ⚛️ fix: Atomic Job Replacement and Boundary-Clean Steering Module
Codex round 6 (5 fixed, 1 standing deferral):
- createJob resets the steer queue and writes the job hash in ONE
same-slot Lua script (JOB_CREATE_LUA): a steer POST can no longer
interleave between them on cluster, so a steer accepted against one
run can never be drained into another. Redis-validated.
- The steering media pipeline moved to packages/api
(agents/steering/media.ts) with injected getFiles and a structural
client interface — /api keeps zero steering logic; specs ported to
the DI seam.
- handleSteerRequest checks the job BEFORE the capability gate: a steer
racing completion on an unsupported SDK gets 404 (send-now) instead
of a 501 queue with no run-end signal left to drain it.
- useQueueDrain binds to the active conversation: navigating away
between the final SSE and the drain effect leaves the signal
unconsumed instead of submitting A's follow-up into B; the drain
fires on return.
- abortJob closes and drains the steer queue BEFORE the content
snapshot, so a drain-hook apply that lands pre-drain is captured
inline rather than lost between the snapshot and the terminal drain.
* 🚦 fix: Parked Run-End Signals, Interrupt Priority, Settled-Run Fallbacks
Codex round 7 (5 fixes):
- Run-end signals for a non-active conversation are PARKED per
conversation instead of squatting the shared index slot: a later run
finishing on the same pane can no longer overwrite them, and the
parked drain fires when the user returns.
- "Interrupt & send" front-inserts carry a priority flag that outranks
createdAt when abort leftovers merge back chronologically — the
urgent redirect drains first, not the oldest steer.
- STEER_UNSUPPORTED/RUN_PAUSED/QUEUE_FULL rejections landing after the
run settled mirror the NO_ACTIVE_RUN fallback and send immediately
(queueing would strand the text with no run-end signal left); on the
pinned SDK this is the common Enter-near-run-end path.
- A failed abort (e.g. 404 when the run completed first) still signals
the interrupt drain, so the queued interrupt message can't strand and
the armed flag can't leak onto a later run.
- Steered-image fallback alt text is localized (com_ui_attached_image).
* 📌 chore: Adopt Published @librechat/agents Types Post-Bump
dev's pin bump to ^3.2.62 (the release carrying injection + steer
replay) landed via merge; the steering runtime now uses the SDK's real
InjectedMessage/hook-output types instead of the local structural
mirrors that bridged the pre-publish window. The two-half capability
probe stays as the defensive gate for mismatched deployments — and the
capability spec now exercises its TRUE path against the published
package in CI.
* 🛅 feat: Park-and-Claim Steer Recovery + Host-View Content Reads
Codex round 8 (6 fixed incl. both P1s, 1 push-back):
- The long-deferred no-subscriber gap is closed: every terminal drain
(final, aborted-final, error, abortJob, resumed finalize) PARKS
acknowledged leftovers on the job hash (unrecoveredSteers), and the
status route claims them exactly once for inactive jobs — a client
that closed/reloaded past the transient final event restores its
steers as queued chips within the post-terminal TTL. A replacement
run clears the parked copy (a live client started it).
- Same-instance content reads are steer-complete: RedisJobStore now
caches the HOST content array (WeakRef) via setContentParts and
prefers it over the SDK graph cache, whose view never contains
host-authored steer parts; the graph fallback splice-INSERTS steer
chunks at their recorded host-view indices (the graph array is
unshifted, so assignment would overwrite SDK parts).
- Replay token accounting now counts prepended file-context text: full
stamped content minus the steer body (already counted), so large
steered documents hit the budget instead of bypassing pruning.
- The queue drain restores an item when ask() refuses without sending
(history not yet in cache after navigating back) — text is never
silently dropped.
- The armed interrupt flag travels WITH a parked run-end signal, so
another run on the same pane can neither consume nor clear it.
- parseTextParts extracts steer text (search indexing / audio).
* 🎛️ refactor: Single Send Slot + In-Thread Steer Messages
- Merge the during-run send affordance into the send/stop button slot:
with composer text the send button replaces Stop (Enter = default
action), hover reveals Steer/Queue/Interrupt rows with shortcuts;
drop the separate DuringRunActionsMenu chevron
- Add during-run keyboard chords: Cmd/Ctrl+Enter = non-default action,
Alt+Enter = interrupt & send (plain-Enter submitters only)
- Render steers as standard user messages in the thread: SteerPart
(icon + author header + user text presentation) replaces the
SteerBubble, and submitted steers appear immediately at the projected
injection point via the PendingSteers slot on the streaming message
- Keep composer rows only for recoverable states: failed steers
(retry/edit/queue) and queued follow-ups
* 🩹 fix: Keep the Replacement Submission Alive Across Abort Settlement
The aborted run's final SSE event fires before the abort HTTP response
resolves, so an armed interrupt & send drains and starts the NEXT
submission while the abort POST is still in flight. The response
handler's unconditional clearAllSubmissions() then reset the new
submission, aborting its stream attach before the subscribe — the
follow-up ran and persisted server-side but the live placeholder
finalized empty (content appeared only after reload).
useAbortCleanup captures the submission before the abort round-trip
and both settlement paths (success and 404-catch) clear only when the
captured submission is still current; a replacement stays untouched.
Plain Stop behavior is unchanged.
* 🧭 test: Playwright E2E for Mid-Run Steering and Queuing
- Add e2e/specs/mock/steering.spec.ts: steer mid-run (202 + immediate
in-thread pending part + real MCP tool boundary + words survive run
end), Cmd/Ctrl+Enter queue with auto-send after clean completion,
and Alt+Enter interrupt & send with the follow-up streaming into the
live view
- Add the E2E_STEER_TOOL_REPLY fake-model marker: slow preamble, a
real remember_fact MCP tool call (PostToolBatch boundary), then a
final turn
- Test 1 pins the run-end degradation contract while the SDK's
top-level agentId stamping bug blocks live injection; its header
documents the assertions to flip once the fixed SDK is pinned
* 🧷 fix: Job-Independent Steer Recovery + Expiry and Resume-Gap Parking
Codex round 10: the park-and-claim recovery had lifecycle holes.
- Move parked steers off the job hash onto their own bounded-TTL store
key (JOB_CREATE_LUA resets it; deleteJob leaves it alone): the default
completeJob path deletes the job record immediately, and the Redis
read path never deserialized the old hash field — recovery previously
worked only with STREAM_KEEP_COMPLETED_JOBS on the in-memory store
- Carry the owner identity inside the parked payload and authorize the
claim against it, so the status route recovers steers on its jobless
branch too (the common reload-after-terminal case); a non-owner claim
returns nothing and re-parks the payload
- Park queued steers on approval expiry: snapshot the frozen queue
before the requires_action→aborted CAS (whose terminal cleanup drops
the steers key) and park only when the CAS wins
- Mirror the terminal drain/park block in resume.js's failure path,
which previously let completeJob's backstop clear 202-accepted steers
- Close the Redis snapshot→subscribe resume gap: re-peek the queue
after attaching and re-surface missed on_steer_applied events from
the durable content view (synthesizeAppliedSteerEvents), updating
resumeState.pendingSteers to the live queue
* 📌 chore: Require @librechat/agents 3.2.63 + Applied-Steer E2E Contract
- Bump the @librechat/agents pin to ^3.2.63 in api/ and packages/api/:
it scopes the hook agentId marker to subagent child graphs, so the
steering drain hook fires at top-level tool-batch boundaries and
mid-run injection is active (danny-avila/agents PR 307)
- Flip e2e steering test 1 from the documented degradation contract to
the applied-steer contract: the optimistic in-thread part transitions
to the persisted part at the tool boundary and survives inside the
response after run end, with no queued follow-up turn
* 🎗️ feat: Steered Messages Join the Message-Nav Ribs
Steers are user messages, so they get their own clickable rib on the
navigation rail, interleaved at their in-thread position inside the
response that absorbed them (one DOM query in document order). SteerPart
anchors itself as #steer-<id> with a steer-render marker — both the
optimistic pending entry and the persisted part — and the rib carries
the user role label with a preview drawn from the steer's text body,
skipping the author header.
* ❎ feat: Cancel a Queued Steer Before Injection + True User-Message Alignment
- Add POST /chat/steer/cancel: removes ONE still-queued steer by id via
an atomic list rebuild (Redis Lua preserves order and TTL), authorized
against the job owner; removed:false is advisory — the cancel lost its
race to the drain or the run end, never an error
- Surface an × on the in-thread pending steer (server-acknowledged
entries only): optimistic removal, restored if the POST fails since
the server would still inject the words
- Outdent SteerPart past the response's icon column so steers sit flush
with top-level message rows, reading as regular user messages
* 🧯 fix: Round-11 Recovery Hardening + Provider-Free Pending Slot
- Reconcile the resume steer gap by steerId SETS, not queue length — a
steer added in the gap (or an equal-length drain+enqueue swap) now
refreshes resumeState.pendingSteers and still synthesizes the missed
on_steer_applied events
- Make completeJob's terminal backstop park: direct error-path callers
without the controllers' close-and-park no longer silently clear
202-accepted steers (createdAt-guarded closeAndDrain + owner park
before the terminal write)
- Persist the steer part BEFORE media encoding in the drain hook: an
abort inside the encode window can no longer lose a file-steer (the
part refs come from the enqueue-sanitized item; replay re-encodes
per turn unchanged)
- Move the parked-claim owner check INSIDE the atomic store claim
(substring gate in the Lua / in-memory equivalent): a non-owner probe
can no longer transiently delete the recovery payload; the app-side
parse stays authoritative
- Park queued steers in BOTH stores' own requires_action expiry
cleanup, which bypassed the manager-level sweep
- Sweep expired parked steers from the in-memory store's periodic
cleanup; restore a queued chip when send-now's submit is refused;
upsert steer ACKs so an SSE reconnect reseed cannot duplicate chips
- Mount the cancel mutation per steer item so the pending slot needs no
QueryClient on ordinary streaming renders (fixes the CI failure in
ContentParts.integration.test)
- Skipped delivery-gated parking (finding 8): transport receiver counts
cannot prove browser delivery, and gating the only durable copy on
them trades cosmetic chip resurrection for real text loss; the window
is already bounded by claim-on-read, createJob reset, and the TTL
* 🩺 fix: Annotate PARKED_STEERS_TTL_MS for isolatedDeclarations
tsdown's d.ts generation requires explicit types on exported consts
with computed initializers; tsc --noEmit does not run that check, so
the round-11 export slipped past local verification and broke Build
packages (and every downstream CI job that consumes the built dist).
* 🛟 fix: Round-12 Terminal-Path Recovery + Durable Steer Events
- Park queued steers before the stale-running reap deletes a crashed or
hung job in BOTH stores — the one terminal path with no controller
finalization; requires_action expiry parking refactored onto the same
snapshot/park helpers
- Enqueue instead of dropping when a steer fallback send is refused:
both the NO_ACTIVE_RUN branch and the settled-run rejection branch
now observe sendNow's false return
- Recover on the SSE reconnect-404 terminal path: convert local pending
steers to queued, claim parked steers via /chat/status, and write a
non-completed run-end signal so interrupt flags release without
auto-sending an unknown outcome
- Fall back to a positive parked-recovery TTL when completedTtl is 0
(SET EX 0 is invalid and silently killed recovery)
- Make on_steer_applied durable before publish: emitChunk gains a
durable option that awaits the chunk-log append (best-effort) ahead
of the transport publish; the default delta path stays fire-and-forget
* 🔐 fix: Round-13 Steer Authorization + Trusted File Refs
- Resolve client-supplied steer file refs against the DB owner-scoped
at enqueue and queue only DB-derived shapes (same filter as the
injection fetch, shared via refs.ts); any unresolved id fails loud
with 400 — spoofed type/filepath metadata can no longer be persisted
into assistant content or rendered in chat/share views
- Enforce agent authorization on /chat/steer against the ORIGINATING
run's job identity: the chat path's role gate (AGENTS:USE, with the
same non-agents-endpoint skip) plus the per-agent ACL check with the
capability bypass — revoked access mid-run can no longer inject;
cancel stays ownership-only (nothing model-bound)
- Mark steered uploads used after a successful enqueue (owner-scoped,
best-effort) so the upload-window TTL cannot reap a file the
persisted steer part references
- Consume the parked recovery copy after live delivery: converting
final/abort/error pendingSteers fires one owner-gated claim-on-read,
so dismissed chips can no longer resurrect on a later reload
* 🎙️ fix: Round-14 Composer-Context Fidelity + TTS and Queue-State Gaps
- Keep steer text out of generic assistant text extraction:
parseTextParts excludes STEER parts by default with an includeSteer
opt-in for the full-record surfaces (Meili indexing, aborted-response
persistence) — TTS callers no longer speak the user's own mid-run
words
- Mark queued uploads used at enqueue time via a minimal owner-scoped
POST /files/usage (fail-closed without a user; upload limiters do not
apply to a metadata touch), fired once wherever composer files enter
the queued state — the upload-window TTL can no longer reap a file
waiting out a long run or approval pause
- Carry quote chips and manual skill picks on queued items: captured
and consumed from the composer at queue/interrupt time exactly like
files, threaded through the drain and send-now overrides, and
restored by the queued row's Edit message
- Key an early-aborted FIRST turn's run-end signal to NEW_CONVO
(resolveRunEndTarget) so queued follow-ups stay visible on the
restored new-chat composer instead of parking under an optimistic
stream id the user never sees again
* 🧿 fix: Round-15 Gap Coverage + Consolidated Sweep (Share Leak, Abort Ids, Chip Hygiene)
- Run the resume steer-gap check for every still-active job: an empty
snapshot no longer skips the re-peek, and synthesis now keys on the
FRESH content view so an applied-in-gap steer that was never
snapshotted still re-surfaces (over-emission is benign — applied-id
dedupe, index-stable parts)
- Thread queued context through steer degradation: sendQueuedNow passes
the item's quotes/skills into submitSteer, and every fallback
(requeue or settled send) restores them instead of dropping to
text+files
- Stop shared links from leaking steer attachment refs: the share
snapshot now walks content — files-excluded shares strip steer-part
files entirely; files-included shares sanitize and share-route them
like top-level files (copy-on-write, non-steer content by reference)
- Seed pending-steer chips unconditionally on load/return so a steer
applied while away cannot linger as a stale chip beside its part
- Use the abort response's resolved job id: chips/drain-signal land
where the user actually is (NEW_CONVO for a new-held first turn,
consistent with resolveRunEndTarget) while the parked-copy claim hits
the resolved id instead of a no-op /chat/status/new
- Open steered documents like normal message files (FilePreviewDialog)
- Cap the applied-steer id set on the live path via a shared helper;
kept surviving run end deliberately (late-ACK race depends on it) and
fixed the atom comment that claimed otherwise
* 💡 fix: Un-light Steer Ribs When Their Node Is Replaced
Two stacked gaps kept a steer rib lit after scrolling away: the
pending→applied swap replaces the DOM node under the same id, which
produces no IntersectionObserver exit and — because the entry list
dedupes on (id, preview) — no entries change either, so the observer
kept watching a detached node; and the rail's mutation filter only
reacted to .message-render nodes, so steer-node swaps and removals
never triggered a refresh at all.
- reconcileObservedElements re-points the observer at replaced nodes
from the mutation-driven refresh regardless of entries identity,
dropping stale visibility until the fresh node reports (the observer
fires its initial intersection immediately, so a truly visible part
re-lights within a frame)
- The mutation filter now recognizes steer-render nodes alongside
message rows
* 🪪 fix: Round-16 Recovery Owner Fields + Context Stickiness + Share Labels
- Park resumed-run leftovers with the manager facade's metadata owner
fields: a bare job.userId is undefined on that shape, which made
every parked payload from a resumed HITL run unclaimable
- Keep a queued item's quotes/skills sticky through a successful steer
ACK: the pending chip carries them (client-only), reseeds preserve
them across reconnects, and every terminal conversion — local or
server-list, merged by steerId — restores them onto the queued item
- Convert resumeState.pendingSteers on the inactive status branch
(deduped against unrecoveredSteers) so steers observed in the
expired-pause-before-sweeper window convert instead of vanishing
until a later reload
- Label shared steer parts share-safely via the existing ShareContext:
a viewer's own name no longer appears on the sharer's steered
messages
* ✂️ fix: Carry Steer Context Through the Failed-Chip Edit Action
Retry and convert-to-queue already preserve a failed steer's carried
quotes/skills; Edit message dropped them on the way back to the
composer. It now restores them through the same context path.