mirror of
https://github.com/danny-avila/LibreChat.git
synced 2026-08-04 14:57:42 +00:00
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* 🧵 feat: Background Tool Calls for Agents & Model Specs Opt-in, poll-based background tool execution. The model marks an eligible tool call with `run_in_background: true`; the host executor registers a task, returns a handle immediately (so the graph turn resolves), runs the tool as a detached promise, and the model retrieves the result via a new `check_background_task` poll tool. Host-side only — no `@librechat/agents` change. - Opt-in mirrors `deferred_tools`: admin capability `run_in_background` (off by default) + per-tool `tool_options.run_in_background`. - Model specs / ephemeral agents: `TModelSpec.runInBackground` / `TEphemeralAgent.run_in_background` synthesize per-tool options; both paths converge at `initializeAgent`. - In-process task registry: scoped per user+conversation, idempotent by toolCallId (safe across resume/replay), capped, TTL-swept. - Excludes direct-path / host-special / code-session tools. Subagents and push notifications are deferred follow-ups. * 🩹 fix: Harden background tool calls (Codex review) - Reliable per-agent execution gate: thread the injected `run_in_background` tool names from `initializeAgent` through `configurable.backgroundToolNames` (`toolRegistry` only reaches the executor for PTC/tool_search), fixing the silent no-op + unstripped-arg leak for ordinary event-driven tools. - Enforce the per-tool opt-in at execution (`backgroundToolSet.has(name)`) so a non-opted-in tool can't be backgrounded via an extra arg. - Gate the `check_background_task` interception on the run actually enabling background, so a user tool sharing that name still executes. - Forward `backgroundToolsAvailable` to added-convo (multi-convo) agents. - Exclude `web_search`/`file_search` from eligibility — their results are turned into user-visible attachments/citations only by the foreground toolEndCallback. * 🩹 fix: Address Codex round 2 on background tool calls - Idempotency scoped to run+turn: provider tool-call ids repeat across turns (e.g. `call_0`), so key the dedupe map by `runId::toolCallId` and sweep orphaned mappings — a later turn no longer collides with a retained task. - Artifacts preserved: a backgrounded tool's artifact is processed through the same `toolEndCallback` as the foreground path (images/files/citations no longer silently dropped), best-effort/guarded. - Forward the `run_in_background` capability to connected-agent discovery and subagent `processAgent` init, so a child agent's own event-driven tools work the same as when it runs as primary. - Strip the injected flag on foreground calls of background-capable tools (the model may emit it as `false`) so strict MCP/action schemas don't reject. - `check_background_task` list path returns metadata only (result_available / result_chars), never full results — prevents context overflow; the full result is returned only when a specific id is requested. * 🩹 fix: Address Codex round 3 on background tool calls - Exclude background-capable tools from eager execution (run.ts): a speculative eager dispatch of a `run_in_background` call could launch the detached task with partial/stale args, and that side effect can't be canceled. - Reserve the `check_background_task` name: overwrite a colliding user/MCP tool with the host poll schema (with a warning) so the advertised schema matches the executor's interception instead of hijacking a mismatched tool. - Don't inject background schemas into pure subagents (spawn-tool child graphs) whose tools don't reach the host interceptor; keep it for primary/added/ connected agents. Subagent background is the durable follow-up. - Thread `backgroundToolsAvailable` + `backgroundToolNames` through the OpenAI-compatible and Responses agent routes (was chat-only), so the same agent/model spec behaves consistently across surfaces. - Exclude image-generation built-ins (dalle/flux/gemini_image_gen/image_gen_oai/ image_edit_oai) — artifact-first tools whose files can't reliably attach to an already-saved turn when backgrounded. * 🩹 fix: Address Codex round 4 on background tool calls - Sanitize self-spawn subagent inputs: strip `run_in_background` + the `check_background_task` def from the parent AgentInputs reused for self-spawn, so the isolated child (direct/child-graph path) doesn't advertise a background schema it can't honor. The SDK resolver keeps a provided `agentInputs` even with `self: true`. - Exclude `check_background_task` from PTC (`run_tools_with_code`) tool definitions — it's host-only and not callable from generated code. - Parse stringified JSON args before deciding background dispatch and before stripping the flag, so string-delivered `run_in_background` is honored and never leaks to strict object-schema tools. - Skip injection for tools that already declare their own `run_in_background` param (would otherwise hijack/strip it), and for non-object (string-input) schemas (would otherwise rewrite the input contract). * 🩹 fix: Address Codex round 5 on background tool calls - check_background_task now parses stringified JSON args, so providers that deliver args as a string can retrieve a specific task by id (not just list). - Include agentId in the background dedupe key (`agentId::runId::toolCallId`): two agents in the same run emitting the same provider id (e.g. `call_0`) now launch independent tasks instead of colliding. - Self-spawn sanitization also strips the background entries from the reused toolRegistry (not just toolDefinitions), so a child using tool_search/deferred loading can't rediscover the host-only run_in_background / check_background_task. * 🩹 fix: Strip run_in_background from PTC target tool schemas (Codex round 6) The PTC path already filtered out the host-only check_background_task poll tool but still exposed target tool schemas with the injected `run_in_background` param (the shared toolRegistry entries were mutated by applyBackgroundToolCalls). PTC codegen doesn't go through the host background interceptor, so it could pass the flag to an MCP/action tool (strict-schema rejection or silent foreground with no poll). Sanitize the PTC toolDefs like the self-spawn path does. * 🩹 fix: Sanitize background from explicit subagent inputs (Codex round 7) A child agent reachable as a top-level/handoff agent is initialized WITH the background capability, then reused as an explicit subagent via buildSubagentConfigs. Round 4 only sanitized the self-spawn case; this now applies the same stripBackgroundFromToolDefinitions/Registry to explicit child agentInputs when `child.backgroundToolNames` is non-empty, so an isolated child graph doesn't advertise a run_in_background / check_background_task contract it can't honor. * 🩹 fix: Reap stuck/expired background tasks (Codex round 8) - get() now sweeps before returning, so repeatedly polling a known background_task_id can't keep an expired completed task (and its retained result, up to 100k chars) alive past the one-hour completed TTL. - sweep() now reaps `running` tasks older than a 30-min running TTL, marking them errored. Previously a detached call that never settled (hung network / lost MCP connection) held a running slot forever, exhausting the per-conversation cap and rejecting every later dispatch. * 🩹 fix: Evict oldest settled tasks instead of blocking at the cap (Codex round 9) Only the running-task cap gates dispatch now. The total-tasks cap (MAX_TASKS_PER_BUCKET) bounds memory but no longer rejects new background calls: when full, it evicts the oldest settled (completed/error) tasks to make room. Previously 200 quick background calls in one conversation would block all new dispatches for up to the completed-task TTL, since polling doesn't remove settled tasks. Running is already capped, so room always frees. * 📝 docs: Frame background tool calls as within-turn (Codex P1 contract) Codex escalated the request-lifecycle findings to P1 on the grounds that the advertised "poll later" contract can't be honored for genuinely long-running calls (request-scoped MCP connections + the run abort signal are torn down at turn end). Align the model-facing contract with what the same-run implementation actually delivers: the run_in_background param, check_background_task, and the dispatch handle now instruct the model to collect the result WITHIN THE SAME TURN (backgrounded work isn't guaranteed to survive past the turn). This is within-turn parallelism; cross-turn survival of long-running calls remains the deliberate durable subagent follow-up. Copy/comment-only; no behavior change. * ♻️ refactor: Cross-turn background tool calls, leak-free Extend background tool calls from within-turn to cross-turn on a single process, since the mechanism already supports it: the run's abort signal never reaches the detached invoke (the graph forwards only configurable/ metadata to the tool-execute handler), so the floating promise keeps running past turn completion and its result stays in the in-process registry for a later turn to poll (get/list key only on user::conversation + id, never the dispatch run/turn). Guarantee no connection leak: ephemeral request-scoped MCP tools (runtime {{LIBRECHAT_BODY_*}} placeholders) capture their request-scoped store at creation and fall back to it, so config manipulation can't redirect them; their connection is torn down at request end. Tag such tools in createToolInstance and run them in the foreground instead of backgrounding them. Pooled/app-level MCP and structured tools are unaffected and survive cross-turn via their managed pools. Reword the model-facing contract (run_in_background, check_background_task, handle message, fileoverview) from within-turn to cross-turn on this server (not across restart/replica, which stays the durable follow-up). Tests: cross-turn poll retrieval; ephemeral MCP tool runs foreground. * 🐛 fix: Guard ephemeral MCP tag against a null server config createToolInstance can be reached with a null/stale capturedServerConfig (cached availableTools + getServerConfig returns null, as several MCP unit tests construct tools). The new unconditional requiresEphemeralUserConnection call then dereferenced config.source and threw during tool construction (CI: Tests api shard 2/3). Guard with the same serverConfig ? ... : false pattern the other callers use; a missing config is not request-scoped. * 🎨 fix: Deliver backgrounded tool artifacts on the poll turn A slow backgrounded MCP/action tool resolves after its dispatch turn is finalized: createToolEndCallback only appends to that turn's artifactPromises (already awaited) and writes to a closed stream, so the artifact (file/citation/ UI resource) was silently dropped — check_background_task recorded only the hasArtifact boolean. The cross-turn contract made this the common case. Hold the artifact on the task and deliver it through the LIVE poll turn's toolEndCallback the first time check_background_task collects that id (once, then cleared to free memory), attributed to the original tool. Same-turn and cross-turn now share this path since the model must poll to collect any result. Tests: registry claim-once; artifact delivered on poll not dispatch, idempotent. * ✨ feat: Agent-builder toggle for background tool calls + cap tool descriptions Add a per-MCP-tool "run in background" toggle in the agent builder, mirroring the programmatic/deferred pattern: gated on the admin `run_in_background` capability via useAgentCapabilities, read/written on tool_options[id] .run_in_background through useMCPToolOptions (per-tool + bulk mark-all), and rendered as a Zap toggle in MCPToolItem and McpSection with new locale keys. Also cap the section tool/server descriptions (McpSection, ToolSection, SkillSection) with max-h-40 overflow-y-auto so a long description scrolls instead of overflowing the dialog, matching MCPToolItem's existing cap. Tests: MCPToolItem renders/toggles the background button only when enabled. * 🧪 fix: Mock new background hook functions in McpSection spec * 🎨 fix: Restore background artifact when poll-turn delivery fails * 🛡️ fix: Harden background tool call edges from review findings - Error immediately (matching foreground) when a background-requested tool failed to load, instead of returning a success handle for a dead task - Exclude ephemeral request-scoped MCP tools at injection time so the model never sees a run_in_background param the executor would silently downgrade; flip the execute-time tag to fail closed on a missing server config - Source image-tool background exclusions from the shared imageGenTools set (adds missing stable-diffusion, an artifact-first live tool) instead of a hand-copied list - Add check_background_task to the eager-execution exclusion list: artifact collection is a one-shot claim that must not fire from a speculative snapshot the SDK may discard - Strip an imitated run_in_background arg on tools the executing agent never opted in (multi-agent history bleed), unless the tool's own schema declares the parameter - Truncate oversized stored results with an explicit marker via the shared truncateMiddle (moved to utils/text) instead of a silent slice - Document the at-most-once artifact delivery semantics honestly (the callback's downstream persistence is fire-and-forget, as in foreground) * ♻️ refactor: Deduplicate background tool-call plumbing and tighten types - Use the SDK's JsonSchemaType instead of a local duplicate; drop all as-unknown casts and type the poll-tool serializer explicitly - Drop derivable BackgroundTask state (progress, hasArtifact) and the dead `enabled` param/return on applyBackgroundToolCalls (guarded at the call site), which also skips the defs pass when nothing opted in - Fold the enable expression into synthesizeBackgroundToolOptions so the three load/added call sites can't drift - Throttle the registry's all-buckets sweep and always sweep the accessed bucket, so a hot poll loop is no longer O(total tasks server-wide); bound retained artifact memory with a size cap - Single-pass stripBackgroundFromToolDefinitions; pass metadata through to the poll-turn callback instead of a no-op reconstruction - Collapse the client's copy-pasted boolean option families into a keyed factory (also removes the shared-object mutation in the bulk toggles) and the six toggle-button copies into one OptionToggle component * 🧪 test: e2e coverage for cross-turn background tool calls Proves the full contract through the real pipeline (mock harness): an agent opts an MCP tool in via tool_options.run_in_background, the model dispatches it detached and receives the synthetic handle while the tool is still running (status=running in the rendered ack — the non-blocking guarantee without timing assertions), the tool completes after its turn finalized, and a later user turn recovers the task id from replayed history, polls check_background_task, and renders the collected result. - fake-mcp-server: slow_echo fixture tool (delayed echo) - fake-model: E2E_BACKGROUND_DISPATCH / E2E_BACKGROUND_COLLECT markers - e2e yaml: agents capabilities = defaults + run_in_background * 🔧 fix: Close two background capability gaps from review - Thread backgroundToolsAvailable through the OpenAI-compatible service (derived from app capabilities like codeEnvAvailable/statefulSessions), so agents with tool_options.run_in_background keep the feature on that route; fold the three capability derivations into one helper - Index ephemeral MCP servers by normalizeServerName when excluding tools from background injection: tool names embed the normalized server name while mcpConfig keys the original, so exotic server names previously escaped the injection-time exclusion * 🛂 fix: Fall back to configurable user identity for background task scoping The in-repo routes merge req into the tool-execute configurable, but external hosts of the exported OpenAI-compatible service inject their own loadTools and may not — tasks would then register under an empty user id, collapsing registry isolation to conversationId alone. Resolve the scoping id from req.user.id, then configurable.user_id / user, and cover the isolation with a foreign-user not_found test. * 🧹 chore: Apply repo import sorter to PR-touched files
1250 lines
40 KiB
JavaScript
1250 lines
40 KiB
JavaScript
const { nanoid } = require('nanoid');
|
|
const { v4: uuidv4 } = require('uuid');
|
|
const { logger } = require('@librechat/data-schemas');
|
|
const { Callback, ToolEndHandler, formatAgentMessages } = require('@librechat/agents');
|
|
const {
|
|
EModelEndpoint,
|
|
ResourceType,
|
|
PermissionBits,
|
|
hasPermissions,
|
|
AgentCapabilities,
|
|
} = require('librechat-data-provider');
|
|
const {
|
|
createRun,
|
|
applyContextToAgent,
|
|
buildToolSet,
|
|
buildAgentScopedContext,
|
|
buildAgentContextAttachmentsByAgentId,
|
|
createSafeUser,
|
|
initializeAgent,
|
|
loadSkillStates,
|
|
getBalanceConfig,
|
|
injectSkillPrimes,
|
|
extractManualSkills,
|
|
recordCollectedUsage,
|
|
createSubagentUsageSink,
|
|
getTransactionsConfig,
|
|
findPiiMatchInMessages,
|
|
discoverConnectedAgents,
|
|
createToolExecuteHandler,
|
|
getRemoteAgentPermissions,
|
|
resolveAgentScopedSkillIds,
|
|
// Responses API
|
|
writeDone,
|
|
buildResponse,
|
|
generateResponseId,
|
|
isValidationFailure,
|
|
emitResponseCreated,
|
|
createResponseContext,
|
|
createResponseTracker,
|
|
setupStreamingResponse,
|
|
emitResponseInProgress,
|
|
convertInputToMessages,
|
|
validateResponseRequest,
|
|
buildAggregatedResponse,
|
|
createResponseAggregator,
|
|
sendResponsesErrorResponse,
|
|
createResponsesEventHandlers,
|
|
createAggregatorEventHandlers,
|
|
} = require('@librechat/api');
|
|
const {
|
|
createResponsesToolEndCallback,
|
|
buildSummarizationHandlers,
|
|
markSummarizationUsage,
|
|
createToolEndCallback,
|
|
agentLogHandlerObj,
|
|
} = require('~/server/controllers/agents/callbacks');
|
|
const { loadAgentTools, loadToolsForExecution } = require('~/server/services/ToolService');
|
|
const {
|
|
findAccessibleResources,
|
|
getEffectivePermissions,
|
|
} = require('~/server/services/PermissionService');
|
|
const {
|
|
getSkillToolDeps,
|
|
getSkillDbMethods,
|
|
canAuthorSkillFiles,
|
|
withDeploymentSkillIds,
|
|
buildAgentToolContext,
|
|
enrichLoadedToolsWithAgentContext,
|
|
} = require('~/server/services/Endpoints/agents/skillDeps');
|
|
const { getModelsConfig } = require('~/server/controllers/ModelController');
|
|
const { resolveConfigServers } = require('~/server/services/MCP');
|
|
const { getMCPManager } = require('~/config');
|
|
const { logViolation } = require('~/cache');
|
|
const db = require('~/models');
|
|
|
|
/**
|
|
* Creates a tool loader function for the agent.
|
|
* @param {AbortSignal} signal - The abort signal
|
|
* @param {boolean} [definitionsOnly=true] - When true, returns only serializable
|
|
* tool definitions without creating full tool instances (for event-driven mode)
|
|
*/
|
|
function createToolLoader(signal, definitionsOnly = true) {
|
|
return async function loadTools({
|
|
req,
|
|
res,
|
|
tools,
|
|
model,
|
|
agentId,
|
|
provider,
|
|
tool_options,
|
|
tool_resources,
|
|
}) {
|
|
const agent = { id: agentId, tools, provider, model, tool_options };
|
|
try {
|
|
return await loadAgentTools({
|
|
req,
|
|
res,
|
|
agent,
|
|
signal,
|
|
tool_resources,
|
|
definitionsOnly,
|
|
streamId: null,
|
|
});
|
|
} catch (error) {
|
|
logger.error('Error loading tools for agent ' + agentId, error);
|
|
}
|
|
};
|
|
}
|
|
|
|
/**
|
|
* Convert Open Responses input items to internal messages
|
|
* @param {import('@librechat/api').InputItem[]} input
|
|
* @returns {Array} Internal messages
|
|
*/
|
|
function convertToInternalMessages(input) {
|
|
return convertInputToMessages(input);
|
|
}
|
|
|
|
/**
|
|
* Load messages from a previous response/conversation
|
|
* @param {string} conversationId - The conversation/response ID
|
|
* @param {string} userId - The user ID
|
|
* @returns {Promise<Array>} Messages from the conversation
|
|
*/
|
|
async function loadPreviousMessages(conversationId, userId) {
|
|
try {
|
|
const messages = await db.getMessages({ conversationId, user: userId });
|
|
if (!messages || messages.length === 0) {
|
|
return [];
|
|
}
|
|
|
|
// Convert stored messages to internal format
|
|
return messages.map((msg) => {
|
|
const internalMsg = {
|
|
role: msg.isCreatedByUser ? 'user' : 'assistant',
|
|
content: '',
|
|
messageId: msg.messageId,
|
|
};
|
|
|
|
// Handle content - could be string or array
|
|
if (typeof msg.text === 'string') {
|
|
internalMsg.content = msg.text;
|
|
} else if (Array.isArray(msg.content)) {
|
|
// Handle content parts
|
|
internalMsg.content = msg.content;
|
|
} else if (msg.text) {
|
|
internalMsg.content = String(msg.text);
|
|
}
|
|
|
|
return internalMsg;
|
|
});
|
|
} catch (error) {
|
|
logger.error('[Responses API] Error loading previous messages:', error);
|
|
return [];
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Save input messages to database
|
|
* @param {import('express').Request} req
|
|
* @param {string} conversationId
|
|
* @param {Array} inputMessages - Internal format messages
|
|
* @param {string} agentId
|
|
* @returns {Promise<void>}
|
|
*/
|
|
async function saveInputMessages(req, conversationId, inputMessages, agentId) {
|
|
for (const msg of inputMessages) {
|
|
if (msg.role === 'user') {
|
|
await db.saveMessage(
|
|
req,
|
|
{
|
|
messageId: msg.messageId || nanoid(),
|
|
conversationId,
|
|
parentMessageId: null,
|
|
isCreatedByUser: true,
|
|
text: typeof msg.content === 'string' ? msg.content : JSON.stringify(msg.content),
|
|
sender: 'User',
|
|
endpoint: EModelEndpoint.agents,
|
|
model: agentId,
|
|
},
|
|
{ context: 'Responses API - save user input' },
|
|
);
|
|
}
|
|
}
|
|
}
|
|
|
|
/**
|
|
* Save response output to database
|
|
* @param {import('express').Request} req
|
|
* @param {string} conversationId
|
|
* @param {string} responseId
|
|
* @param {import('@librechat/api').Response} response
|
|
* @param {string} agentId
|
|
* @returns {Promise<void>}
|
|
*/
|
|
async function saveResponseOutput(req, conversationId, responseId, response, agentId) {
|
|
// Extract text content from output items
|
|
let responseText = '';
|
|
for (const item of response.output) {
|
|
if (item.type === 'message' && item.content) {
|
|
for (const part of item.content) {
|
|
if (part.type === 'output_text' && part.text) {
|
|
responseText += part.text;
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
// Save the assistant message
|
|
await db.saveMessage(
|
|
req,
|
|
{
|
|
messageId: responseId,
|
|
conversationId,
|
|
parentMessageId: null,
|
|
isCreatedByUser: false,
|
|
text: responseText,
|
|
sender: 'Agent',
|
|
endpoint: EModelEndpoint.agents,
|
|
model: agentId,
|
|
finish_reason: response.status === 'completed' ? 'stop' : response.status,
|
|
tokenCount: response.usage?.output_tokens,
|
|
},
|
|
{ context: 'Responses API - save assistant response' },
|
|
);
|
|
}
|
|
|
|
/**
|
|
* Save or update conversation
|
|
* @param {import('express').Request} req
|
|
* @param {string} conversationId
|
|
* @param {string} agentId
|
|
* @param {object} agent
|
|
* @returns {Promise<void>}
|
|
*/
|
|
async function saveConversation(req, conversationId, agentId, agent) {
|
|
await db.saveConvo(
|
|
{
|
|
userId: req?.user?.id,
|
|
isTemporary: req?.body?.isTemporary,
|
|
interfaceConfig: req?.config?.interfaceConfig,
|
|
},
|
|
{
|
|
conversationId,
|
|
endpoint: EModelEndpoint.agents,
|
|
agentId,
|
|
title: agent?.name || 'Open Responses Conversation',
|
|
model: agent?.model,
|
|
},
|
|
{ context: 'Responses API - save conversation' },
|
|
);
|
|
}
|
|
|
|
/**
|
|
* Convert stored messages to Open Responses output format
|
|
* @param {Array} messages - Stored messages
|
|
* @returns {Array} Output items
|
|
*/
|
|
function convertMessagesToOutputItems(messages) {
|
|
const output = [];
|
|
|
|
for (const msg of messages) {
|
|
if (!msg.isCreatedByUser) {
|
|
output.push({
|
|
type: 'message',
|
|
id: msg.messageId,
|
|
role: 'assistant',
|
|
status: 'completed',
|
|
content: [
|
|
{
|
|
type: 'output_text',
|
|
text: msg.text || '',
|
|
annotations: [],
|
|
},
|
|
],
|
|
});
|
|
}
|
|
}
|
|
|
|
return output;
|
|
}
|
|
|
|
/**
|
|
* Create Response - POST /v1/responses
|
|
*
|
|
* Creates a model response following the Open Responses API specification.
|
|
* Supports both streaming and non-streaming responses.
|
|
*
|
|
* @param {import('express').Request} req
|
|
* @param {import('express').Response} res
|
|
*/
|
|
const createResponse = async (req, res) => {
|
|
const appConfig = req.config;
|
|
const requestStartTime = Date.now();
|
|
|
|
// Validate request
|
|
const validation = validateResponseRequest(req.body);
|
|
if (isValidationFailure(validation)) {
|
|
return sendResponsesErrorResponse(res, 400, validation.error);
|
|
}
|
|
|
|
const request = validation.request;
|
|
const agentId = request.model;
|
|
const isStreaming = request.stream === true;
|
|
const summarizationConfig = appConfig?.summarization;
|
|
|
|
// Look up the agent
|
|
const agent = await db.getAgent({ id: agentId });
|
|
if (!agent) {
|
|
return sendResponsesErrorResponse(
|
|
res,
|
|
404,
|
|
`Agent not found: ${agentId}`,
|
|
'not_found',
|
|
'model_not_found',
|
|
);
|
|
}
|
|
|
|
// Generate IDs
|
|
const responseId = generateResponseId();
|
|
const context = createResponseContext(request, responseId);
|
|
|
|
logger.debug(
|
|
`[Responses API] Request ${responseId} started for agent ${agentId}, stream: ${isStreaming}`,
|
|
);
|
|
|
|
// Set up abort controller
|
|
const abortController = new AbortController();
|
|
|
|
// Handle client disconnect
|
|
req.on('close', () => {
|
|
if (!abortController.signal.aborted) {
|
|
abortController.abort();
|
|
logger.debug('[Responses API] Client disconnected, aborting');
|
|
}
|
|
});
|
|
|
|
try {
|
|
if (request.previous_response_id != null) {
|
|
if (typeof request.previous_response_id !== 'string') {
|
|
return sendResponsesErrorResponse(
|
|
res,
|
|
400,
|
|
'previous_response_id must be a string',
|
|
'invalid_request',
|
|
);
|
|
}
|
|
if (!(await db.getConvo(req.user?.id, request.previous_response_id))) {
|
|
return sendResponsesErrorResponse(res, 404, 'Conversation not found', 'not_found');
|
|
}
|
|
}
|
|
|
|
const conversationId = request.previous_response_id ?? uuidv4();
|
|
const parentMessageId = null;
|
|
|
|
// Build allowed providers set
|
|
const allowedProviders = new Set(
|
|
appConfig?.endpoints?.[EModelEndpoint.agents]?.allowedProviders,
|
|
);
|
|
|
|
// Create tool loader
|
|
const loadTools = createToolLoader(abortController.signal);
|
|
const skillDbMethods = getSkillDbMethods();
|
|
|
|
// Initialize the agent first to check for disableStreaming
|
|
const endpointOption = {
|
|
endpoint: agent.provider,
|
|
model_parameters: agent.model_parameters ?? {},
|
|
};
|
|
|
|
// `filterFilesByAgentAccess` is intentionally omitted: it calls
|
|
// `checkPermission` with `resourceType: AGENT`, but this route
|
|
// authorizes callers through `REMOTE_AGENT` (via
|
|
// `getRemoteAgentPermissions`), so including it would silently drop
|
|
// owner-attached context files for any remote user who has
|
|
// `REMOTE_AGENT_VIEWER` but not direct `AGENT_VIEW`.
|
|
const dbMethods = {
|
|
getConvoFiles: db.getConvoFiles,
|
|
getFiles: db.getFiles,
|
|
getUserKey: db.getUserKey,
|
|
getMessages: db.getMessages,
|
|
updateFilesUsage: db.updateFilesUsage,
|
|
getUserKeyValues: db.getUserKeyValues,
|
|
getUserCodeFiles: db.getUserCodeFiles,
|
|
getToolFilesByIds: db.getToolFilesByIds,
|
|
getCodeGeneratedFiles: db.getCodeGeneratedFiles,
|
|
listSkillsByAccess: skillDbMethods.listSkillsByAccess,
|
|
listAlwaysApplySkills: skillDbMethods.listAlwaysApplySkills,
|
|
getSkillByName: skillDbMethods.getSkillByName,
|
|
};
|
|
|
|
const enabledCapabilities = new Set(
|
|
appConfig?.endpoints?.[EModelEndpoint.agents]?.capabilities,
|
|
);
|
|
const skillsCapabilityEnabled = enabledCapabilities.has(AgentCapabilities.skills);
|
|
const ephemeralSkillsToggle = req.body?.ephemeralAgent?.skills === true;
|
|
const accessibleSkillIds = skillsCapabilityEnabled
|
|
? withDeploymentSkillIds(
|
|
await findAccessibleResources({
|
|
userId: req.user.id,
|
|
role: req.user.role,
|
|
resourceType: ResourceType.SKILL,
|
|
requiredPermissions: PermissionBits.VIEW,
|
|
}),
|
|
)
|
|
: [];
|
|
const editableSkillIds = skillsCapabilityEnabled
|
|
? await findAccessibleResources({
|
|
userId: req.user.id,
|
|
role: req.user.role,
|
|
resourceType: ResourceType.SKILL,
|
|
requiredPermissions: PermissionBits.EDIT,
|
|
})
|
|
: [];
|
|
const skillCreateAllowed = skillsCapabilityEnabled
|
|
? await getSkillToolDeps().canCreateSkill({ req })
|
|
: false;
|
|
|
|
const { skillStates, defaultActiveOnShare } = await loadSkillStates({
|
|
userId: req.user.id,
|
|
appConfig,
|
|
getUserById: db.getUserById,
|
|
accessibleSkillIds,
|
|
});
|
|
|
|
const manualSkills = extractManualSkills(req.body);
|
|
|
|
const primaryScopedSkillIds = resolveAgentScopedSkillIds({
|
|
agent,
|
|
accessibleSkillIds,
|
|
skillsCapabilityEnabled,
|
|
ephemeralSkillsToggle,
|
|
});
|
|
const primaryScopedEditableSkillIds = resolveAgentScopedSkillIds({
|
|
agent,
|
|
accessibleSkillIds: editableSkillIds,
|
|
skillsCapabilityEnabled,
|
|
ephemeralSkillsToggle,
|
|
});
|
|
|
|
const primaryConfig = await initializeAgent(
|
|
{
|
|
req,
|
|
res,
|
|
loadTools,
|
|
requestFiles: [],
|
|
conversationId,
|
|
parentMessageId,
|
|
agent,
|
|
endpointOption,
|
|
allowedProviders,
|
|
isInitialAgent: true,
|
|
accessibleSkillIds: primaryScopedSkillIds,
|
|
skillAuthoringAvailable: canAuthorSkillFiles({
|
|
agent,
|
|
scopedEditableSkillIds: primaryScopedEditableSkillIds,
|
|
skillCreateAllowed,
|
|
skillsCapabilityEnabled,
|
|
ephemeralSkillsToggle,
|
|
}),
|
|
codeEnvAvailable: enabledCapabilities.has(AgentCapabilities.execute_code),
|
|
backgroundToolsAvailable: enabledCapabilities.has(AgentCapabilities.run_in_background),
|
|
statefulSessionsAvailable: enabledCapabilities.has(
|
|
AgentCapabilities.stateful_code_sessions,
|
|
),
|
|
skillStates,
|
|
defaultActiveOnShare,
|
|
manualSkills,
|
|
},
|
|
dbMethods,
|
|
);
|
|
|
|
/**
|
|
* Per-agent tool-execution context map, keyed by agentId. Ensures the
|
|
* ON_TOOL_EXECUTE callback routes each sub-agent's tool calls to the
|
|
* correct toolRegistry / userMCPAuthMap / tool_resources.
|
|
* @type {Map<string, {
|
|
* agent: object,
|
|
* toolRegistry?: import('@librechat/agents').LCToolRegistry,
|
|
* requestScopedConnections?: import('@librechat/api').RequestScopedMCPConnectionStore,
|
|
* userMCPAuthMap?: Record<string, Record<string, string>>,
|
|
* tool_resources?: object,
|
|
* actionsEnabled?: boolean,
|
|
* }>}
|
|
*/
|
|
const agentToolContexts = new Map();
|
|
agentToolContexts.set(
|
|
primaryConfig.id,
|
|
buildAgentToolContext({ agent, config: primaryConfig }),
|
|
);
|
|
|
|
// Only run BFS discovery (and pay `getModelsConfig` upfront) when the
|
|
// primary has edges to follow — the common API case is single-agent.
|
|
let handoffAgentConfigs = new Map();
|
|
let discoveredEdges = [];
|
|
let discoveredMCPAuthMap;
|
|
if (primaryConfig.edges?.length) {
|
|
const modelsConfig = await getModelsConfig(req);
|
|
({
|
|
agentConfigs: handoffAgentConfigs,
|
|
edges: discoveredEdges,
|
|
userMCPAuthMap: discoveredMCPAuthMap,
|
|
} = await discoverConnectedAgents(
|
|
{
|
|
req,
|
|
res,
|
|
primaryConfig,
|
|
endpointOption,
|
|
allowedProviders,
|
|
modelsConfig,
|
|
loadTools,
|
|
requestFiles: [],
|
|
conversationId,
|
|
parentMessageId,
|
|
// The route enforces REMOTE_AGENT on the primary; every discovered
|
|
// sub-agent must clear the same sharing boundary, not the looser
|
|
// in-app AGENT one.
|
|
resourceType: ResourceType.REMOTE_AGENT,
|
|
computeAccessibleSkillIds: (handoffAgent) =>
|
|
resolveAgentScopedSkillIds({
|
|
agent: handoffAgent,
|
|
accessibleSkillIds,
|
|
skillsCapabilityEnabled,
|
|
ephemeralSkillsToggle,
|
|
}),
|
|
computeSkillAuthoringAvailable: (handoffAgent) =>
|
|
canAuthorSkillFiles({
|
|
agent: handoffAgent,
|
|
scopedEditableSkillIds: resolveAgentScopedSkillIds({
|
|
agent: handoffAgent,
|
|
accessibleSkillIds: editableSkillIds,
|
|
skillsCapabilityEnabled,
|
|
ephemeralSkillsToggle,
|
|
}),
|
|
skillCreateAllowed,
|
|
skillsCapabilityEnabled,
|
|
ephemeralSkillsToggle,
|
|
}),
|
|
skillStates,
|
|
defaultActiveOnShare,
|
|
/** @see DiscoverConnectedAgentsParams.codeEnvAvailable */
|
|
codeEnvAvailable: enabledCapabilities.has(AgentCapabilities.execute_code),
|
|
backgroundToolsAvailable: enabledCapabilities.has(AgentCapabilities.run_in_background),
|
|
statefulSessionsAvailable: enabledCapabilities.has(
|
|
AgentCapabilities.stateful_code_sessions,
|
|
),
|
|
},
|
|
{
|
|
getAgent: db.getAgent,
|
|
// Use `getRemoteAgentPermissions` so sub-agent authorization
|
|
// matches what the route's `createCheckRemoteAgentAccess`
|
|
// middleware does for the primary: AGENT owners with the SHARE
|
|
// bit are treated as remotely authorized even without an
|
|
// explicit REMOTE_AGENT grant.
|
|
checkPermission: async ({ userId, role, resourceId, requiredPermission }) => {
|
|
const permissions = await getRemoteAgentPermissions(
|
|
{ getEffectivePermissions },
|
|
userId,
|
|
role,
|
|
resourceId,
|
|
);
|
|
return hasPermissions(permissions, requiredPermission);
|
|
},
|
|
logViolation,
|
|
db: dbMethods,
|
|
onAgentInitialized: (agentId, handoffAgent, config) => {
|
|
agentToolContexts.set(agentId, buildAgentToolContext({ agent: handoffAgent, config }));
|
|
},
|
|
initializeAgent,
|
|
},
|
|
));
|
|
}
|
|
|
|
primaryConfig.edges = discoveredEdges;
|
|
const runAgents = [primaryConfig, ...handoffAgentConfigs.values()];
|
|
const mergedMCPAuthMap = discoveredMCPAuthMap ?? primaryConfig.userMCPAuthMap;
|
|
|
|
const agentContextAttachmentsByAgentId = buildAgentContextAttachmentsByAgentId(runAgents);
|
|
const agentScopedContext = await buildAgentScopedContext({
|
|
agentIds: runAgents.map(({ id }) => id),
|
|
attachmentsByAgentId: agentContextAttachmentsByAgentId,
|
|
req,
|
|
});
|
|
|
|
const mcpManager = getMCPManager();
|
|
const configServers = await resolveConfigServers(req);
|
|
|
|
await Promise.all(
|
|
runAgents.map((runAgent) =>
|
|
applyContextToAgent({
|
|
agent: runAgent,
|
|
agentId: runAgent.id,
|
|
logger,
|
|
mcpManager,
|
|
configServers,
|
|
sharedRunContext: agentScopedContext.get(runAgent.id) ?? '',
|
|
}),
|
|
),
|
|
);
|
|
|
|
// Determine if streaming is enabled (check both request and agent config)
|
|
const streamingDisabled = !!primaryConfig.model_parameters?.disableStreaming;
|
|
const actuallyStreaming = isStreaming && !streamingDisabled;
|
|
|
|
// Load previous messages if previous_response_id is provided
|
|
let previousMessages = [];
|
|
if (request.previous_response_id) {
|
|
const userId = req.user?.id ?? 'api-user';
|
|
previousMessages = await loadPreviousMessages(request.previous_response_id, userId);
|
|
}
|
|
|
|
// Convert input to internal messages
|
|
const inputMessages = convertToInternalMessages(
|
|
typeof request.input === 'string' ? request.input : request.input,
|
|
);
|
|
|
|
const piiHit = findPiiMatchInMessages(inputMessages, appConfig?.messageFilter?.pii);
|
|
if (piiHit != null) {
|
|
return sendResponsesErrorResponse(
|
|
res,
|
|
400,
|
|
`Message contains a ${piiHit.label}. Remove it and try again.`,
|
|
'invalid_request',
|
|
'message_filter_pii_block',
|
|
);
|
|
}
|
|
|
|
// Merge previous messages with new input
|
|
const allMessages = [...previousMessages, ...inputMessages];
|
|
|
|
const toolSet = buildToolSet(primaryConfig);
|
|
const formatted = formatAgentMessages(allMessages, {}, toolSet);
|
|
const formattedMessages = formatted.messages;
|
|
const initialSummary = formatted.summary;
|
|
let indexTokenCountMap = formatted.indexTokenCountMap;
|
|
|
|
/**
|
|
* Inject manual + always-apply skill primes so the model sees SKILL.md
|
|
* bodies for this turn — parity with AgentClient's chat path. The
|
|
* Responses API uses its own response-builder shape, so LibreChat-
|
|
* style card SSE events don't apply; only the message-context part
|
|
* carries over.
|
|
*/
|
|
const manualSkillPrimes = primaryConfig.manualSkillPrimes;
|
|
const alwaysApplySkillPrimes = primaryConfig.alwaysApplySkillPrimes;
|
|
if (
|
|
(manualSkillPrimes && manualSkillPrimes.length > 0) ||
|
|
(alwaysApplySkillPrimes && alwaysApplySkillPrimes.length > 0)
|
|
) {
|
|
const primeResult = injectSkillPrimes({
|
|
initialMessages: formattedMessages,
|
|
indexTokenCountMap,
|
|
manualSkillPrimes,
|
|
alwaysApplySkillPrimes,
|
|
});
|
|
indexTokenCountMap = primeResult.indexTokenCountMap;
|
|
/* Surface the cap-driven always-apply truncation at the controller
|
|
layer too — `injectSkillPrimes` already logs internally, but the
|
|
controller-level warn includes endpoint context so operators can
|
|
tell at a glance which path hit the cap. Mirrors AgentClient's
|
|
warn in `client.js`. */
|
|
if (primeResult.alwaysApplyDropped > 0) {
|
|
logger.warn(
|
|
`[Responses API] Dropped ${primeResult.alwaysApplyDropped} always-apply prime(s) to stay within MAX_PRIMED_SKILLS_PER_TURN.`,
|
|
);
|
|
}
|
|
}
|
|
|
|
/* Stable for the turn: the primary prime list is fixed once
|
|
`initializeAgent` resolves and is used as the fallback when a
|
|
specific agent context is unavailable. `codeEnvAvailable` is read
|
|
per-agent from the stored tool context (admin cap AND that
|
|
agent's `tools` list includes `execute_code`) — a skills-only
|
|
agent never gains sandbox access even if the admin enabled the
|
|
capability globally. */
|
|
// Create tracker for streaming or aggregator for non-streaming
|
|
const tracker = actuallyStreaming ? createResponseTracker() : null;
|
|
const aggregator = actuallyStreaming ? null : createResponseAggregator();
|
|
|
|
// Set up response for streaming
|
|
if (actuallyStreaming) {
|
|
setupStreamingResponse(res);
|
|
|
|
// Create handler config
|
|
const handlerConfig = {
|
|
res,
|
|
context,
|
|
tracker,
|
|
};
|
|
|
|
// Emit response.created then response.in_progress per Open Responses spec
|
|
emitResponseCreated(handlerConfig);
|
|
emitResponseInProgress(handlerConfig);
|
|
|
|
// Create event handlers
|
|
const { handlers: responsesHandlers, finalizeStream } =
|
|
createResponsesEventHandlers(handlerConfig);
|
|
|
|
// Collect usage for balance tracking
|
|
const collectedUsage = [];
|
|
|
|
// Artifact promises for processing tool outputs
|
|
/** @type {Promise<import('librechat-data-provider').TAttachment | null>[]} */
|
|
const artifactPromises = [];
|
|
// Use Responses API-specific callback that emits librechat:attachment events
|
|
const toolEndCallback = createResponsesToolEndCallback({
|
|
req,
|
|
res,
|
|
tracker,
|
|
artifactPromises,
|
|
});
|
|
|
|
// Create tool execute options for event-driven tool execution
|
|
const toolExecuteOptions = {
|
|
loadTools: async (toolNames, agentId) => {
|
|
const ctx =
|
|
agentToolContexts.get(agentId) ?? agentToolContexts.get(primaryConfig.id) ?? {};
|
|
const result = await loadToolsForExecution({
|
|
req,
|
|
res,
|
|
toolNames,
|
|
agent: ctx.agent ?? agent,
|
|
signal: abortController.signal,
|
|
toolRegistry: ctx.toolRegistry,
|
|
backgroundToolNames: ctx.backgroundToolNames,
|
|
mcpAvailableTools: ctx.mcpAvailableTools,
|
|
requestScopedConnections: ctx.requestScopedConnections,
|
|
userMCPAuthMap: ctx.userMCPAuthMap,
|
|
tool_resources: ctx.tool_resources,
|
|
actionsEnabled: ctx.actionsEnabled,
|
|
});
|
|
return enrichLoadedToolsWithAgentContext({
|
|
result,
|
|
req,
|
|
ctx,
|
|
});
|
|
},
|
|
toolEndCallback,
|
|
...getSkillToolDeps(),
|
|
};
|
|
|
|
// Combine handlers
|
|
const handlers = {
|
|
on_message_delta: responsesHandlers.on_message_delta,
|
|
on_reasoning_delta: responsesHandlers.on_reasoning_delta,
|
|
on_run_step: responsesHandlers.on_run_step,
|
|
on_run_step_delta: responsesHandlers.on_run_step_delta,
|
|
on_chat_model_end: {
|
|
handle: (event, data, metadata) => {
|
|
responsesHandlers.on_chat_model_end.handle(event, data);
|
|
const usage = data?.output?.usage_metadata;
|
|
if (usage) {
|
|
const taggedUsage = markSummarizationUsage(usage, metadata);
|
|
collectedUsage.push(taggedUsage);
|
|
}
|
|
},
|
|
},
|
|
on_tool_end: new ToolEndHandler(toolEndCallback, logger),
|
|
on_run_step_completed: { handle: () => {} },
|
|
on_chain_stream: { handle: () => {} },
|
|
on_chain_end: { handle: () => {} },
|
|
on_agent_update: { handle: () => {} },
|
|
on_custom_event: { handle: () => {} },
|
|
on_tool_execute: createToolExecuteHandler(toolExecuteOptions),
|
|
on_agent_log: agentLogHandlerObj,
|
|
...(summarizationConfig?.enabled !== false
|
|
? buildSummarizationHandlers({ isStreaming: actuallyStreaming, res })
|
|
: {}),
|
|
};
|
|
|
|
// Create and run the agent
|
|
const userId = req.user?.id ?? 'api-user';
|
|
const userMCPAuthMap = mergedMCPAuthMap;
|
|
|
|
const run = await createRun({
|
|
agents: runAgents,
|
|
messages: formattedMessages,
|
|
indexTokenCountMap,
|
|
initialSummary,
|
|
runId: responseId,
|
|
summarizationConfig,
|
|
appConfig,
|
|
signal: abortController.signal,
|
|
customHandlers: handlers,
|
|
requestBody: {
|
|
messageId: responseId,
|
|
conversationId,
|
|
},
|
|
user: { id: userId },
|
|
tenantId: req.user?.tenantId,
|
|
/** Bills subagent child-run model calls (reported outside the
|
|
* streamEvents loop) into the same collectedUsage array. */
|
|
subagentUsageSink: createSubagentUsageSink(collectedUsage),
|
|
});
|
|
|
|
if (!run) {
|
|
throw new Error('Failed to create agent run');
|
|
}
|
|
|
|
// Process the stream
|
|
const config = {
|
|
runName: 'AgentRun',
|
|
configurable: {
|
|
thread_id: conversationId,
|
|
user_id: userId,
|
|
user: createSafeUser(req.user),
|
|
requestBody: {
|
|
messageId: responseId,
|
|
conversationId,
|
|
},
|
|
...(userMCPAuthMap != null && { userMCPAuthMap }),
|
|
},
|
|
signal: abortController.signal,
|
|
streamMode: 'values',
|
|
version: 'v2',
|
|
};
|
|
|
|
await run.processStream({ messages: formattedMessages }, config, {
|
|
callbacks: {
|
|
[Callback.TOOL_ERROR]: (graph, error, toolId) => {
|
|
logger.error(`[Responses API] Tool Error "${toolId}"`, error);
|
|
},
|
|
},
|
|
});
|
|
|
|
// Record token usage against balance
|
|
const balanceConfig = getBalanceConfig(appConfig);
|
|
const transactionsConfig = getTransactionsConfig(appConfig);
|
|
recordCollectedUsage(
|
|
{
|
|
spendTokens: db.spendTokens,
|
|
spendStructuredTokens: db.spendStructuredTokens,
|
|
pricing: { getMultiplier: db.getMultiplier, getCacheMultiplier: db.getCacheMultiplier },
|
|
bulkWriteOps: { insertMany: db.bulkInsertTransactions, updateBalance: db.updateBalance },
|
|
},
|
|
{
|
|
user: userId,
|
|
conversationId,
|
|
collectedUsage,
|
|
context: 'message',
|
|
messageId: responseId,
|
|
balance: balanceConfig,
|
|
transactions: transactionsConfig,
|
|
model: primaryConfig.model || agent.model_parameters?.model,
|
|
},
|
|
).catch((err) => {
|
|
logger.error('[Responses API] Error recording usage:', err);
|
|
});
|
|
|
|
// Finalize the stream
|
|
finalizeStream();
|
|
res.end();
|
|
|
|
const duration = Date.now() - requestStartTime;
|
|
logger.debug(`[Responses API] Request ${responseId} completed in ${duration}ms (streaming)`);
|
|
|
|
// Save to database if store: true
|
|
if (request.store === true) {
|
|
try {
|
|
// Save conversation
|
|
await saveConversation(req, conversationId, agentId, agent);
|
|
|
|
// Save input messages
|
|
await saveInputMessages(req, conversationId, inputMessages, agentId);
|
|
|
|
// Build response for saving (use tracker with buildResponse for streaming)
|
|
const finalResponse = buildResponse(context, tracker, 'completed');
|
|
await saveResponseOutput(req, conversationId, responseId, finalResponse, agentId);
|
|
|
|
logger.debug(
|
|
`[Responses API] Stored response ${responseId} in conversation ${conversationId}`,
|
|
);
|
|
} catch (saveError) {
|
|
logger.error('[Responses API] Error saving response:', saveError);
|
|
// Don't fail the request if saving fails
|
|
}
|
|
}
|
|
|
|
// Wait for artifact processing after response ends (non-blocking)
|
|
if (artifactPromises.length > 0) {
|
|
Promise.all(artifactPromises).catch((artifactError) => {
|
|
logger.warn('[Responses API] Error processing artifacts:', artifactError);
|
|
});
|
|
}
|
|
} else {
|
|
const aggregatorHandlers = createAggregatorEventHandlers(aggregator);
|
|
|
|
// Collect usage for balance tracking
|
|
const collectedUsage = [];
|
|
|
|
/** @type {Promise<import('librechat-data-provider').TAttachment | null>[]} */
|
|
const artifactPromises = [];
|
|
const toolEndCallback = createToolEndCallback({ req, res, artifactPromises, streamId: null });
|
|
|
|
const toolExecuteOptions = {
|
|
loadTools: async (toolNames, agentId) => {
|
|
const ctx =
|
|
agentToolContexts.get(agentId) ?? agentToolContexts.get(primaryConfig.id) ?? {};
|
|
const result = await loadToolsForExecution({
|
|
req,
|
|
res,
|
|
toolNames,
|
|
agent: ctx.agent ?? agent,
|
|
signal: abortController.signal,
|
|
toolRegistry: ctx.toolRegistry,
|
|
backgroundToolNames: ctx.backgroundToolNames,
|
|
mcpAvailableTools: ctx.mcpAvailableTools,
|
|
requestScopedConnections: ctx.requestScopedConnections,
|
|
userMCPAuthMap: ctx.userMCPAuthMap,
|
|
tool_resources: ctx.tool_resources,
|
|
actionsEnabled: ctx.actionsEnabled,
|
|
});
|
|
return enrichLoadedToolsWithAgentContext({
|
|
result,
|
|
req,
|
|
ctx,
|
|
});
|
|
},
|
|
toolEndCallback,
|
|
...getSkillToolDeps(),
|
|
};
|
|
|
|
const handlers = {
|
|
on_message_delta: aggregatorHandlers.on_message_delta,
|
|
on_reasoning_delta: aggregatorHandlers.on_reasoning_delta,
|
|
on_run_step: aggregatorHandlers.on_run_step,
|
|
on_run_step_delta: aggregatorHandlers.on_run_step_delta,
|
|
on_chat_model_end: {
|
|
handle: (event, data, metadata) => {
|
|
aggregatorHandlers.on_chat_model_end.handle(event, data);
|
|
const usage = data?.output?.usage_metadata;
|
|
if (usage) {
|
|
const taggedUsage = markSummarizationUsage(usage, metadata);
|
|
collectedUsage.push(taggedUsage);
|
|
}
|
|
},
|
|
},
|
|
on_tool_end: new ToolEndHandler(toolEndCallback, logger),
|
|
on_run_step_completed: { handle: () => {} },
|
|
on_chain_stream: { handle: () => {} },
|
|
on_chain_end: { handle: () => {} },
|
|
on_agent_update: { handle: () => {} },
|
|
on_custom_event: { handle: () => {} },
|
|
on_tool_execute: createToolExecuteHandler(toolExecuteOptions),
|
|
on_agent_log: agentLogHandlerObj,
|
|
...(summarizationConfig?.enabled !== false
|
|
? buildSummarizationHandlers({ isStreaming: false, res })
|
|
: {}),
|
|
};
|
|
|
|
const userId = req.user?.id ?? 'api-user';
|
|
const userMCPAuthMap = mergedMCPAuthMap;
|
|
|
|
const run = await createRun({
|
|
agents: runAgents,
|
|
messages: formattedMessages,
|
|
indexTokenCountMap,
|
|
initialSummary,
|
|
runId: responseId,
|
|
summarizationConfig,
|
|
appConfig,
|
|
signal: abortController.signal,
|
|
customHandlers: handlers,
|
|
requestBody: {
|
|
messageId: responseId,
|
|
conversationId,
|
|
},
|
|
user: { id: userId },
|
|
tenantId: req.user?.tenantId,
|
|
/** Bills subagent child-run model calls (reported outside the
|
|
* streamEvents loop) into the same collectedUsage array. */
|
|
subagentUsageSink: createSubagentUsageSink(collectedUsage),
|
|
});
|
|
|
|
if (!run) {
|
|
throw new Error('Failed to create agent run');
|
|
}
|
|
|
|
const config = {
|
|
runName: 'AgentRun',
|
|
configurable: {
|
|
thread_id: conversationId,
|
|
user_id: userId,
|
|
user: createSafeUser(req.user),
|
|
requestBody: {
|
|
messageId: responseId,
|
|
conversationId,
|
|
},
|
|
...(userMCPAuthMap != null && { userMCPAuthMap }),
|
|
},
|
|
signal: abortController.signal,
|
|
streamMode: 'values',
|
|
version: 'v2',
|
|
};
|
|
|
|
await run.processStream({ messages: formattedMessages }, config, {
|
|
callbacks: {
|
|
[Callback.TOOL_ERROR]: (graph, error, toolId) => {
|
|
logger.error(`[Responses API] Tool Error "${toolId}"`, error);
|
|
},
|
|
},
|
|
});
|
|
|
|
// Record token usage against balance
|
|
const balanceConfig = getBalanceConfig(appConfig);
|
|
const transactionsConfig = getTransactionsConfig(appConfig);
|
|
recordCollectedUsage(
|
|
{
|
|
spendTokens: db.spendTokens,
|
|
spendStructuredTokens: db.spendStructuredTokens,
|
|
pricing: { getMultiplier: db.getMultiplier, getCacheMultiplier: db.getCacheMultiplier },
|
|
bulkWriteOps: { insertMany: db.bulkInsertTransactions, updateBalance: db.updateBalance },
|
|
},
|
|
{
|
|
user: userId,
|
|
conversationId,
|
|
collectedUsage,
|
|
context: 'message',
|
|
messageId: responseId,
|
|
balance: balanceConfig,
|
|
transactions: transactionsConfig,
|
|
model: primaryConfig.model || agent.model_parameters?.model,
|
|
},
|
|
).catch((err) => {
|
|
logger.error('[Responses API] Error recording usage:', err);
|
|
});
|
|
|
|
if (artifactPromises.length > 0) {
|
|
try {
|
|
await Promise.all(artifactPromises);
|
|
} catch (artifactError) {
|
|
logger.warn('[Responses API] Error processing artifacts:', artifactError);
|
|
}
|
|
}
|
|
|
|
const response = buildAggregatedResponse(context, aggregator);
|
|
|
|
if (request.store === true) {
|
|
try {
|
|
await saveConversation(req, conversationId, agentId, agent);
|
|
|
|
await saveInputMessages(req, conversationId, inputMessages, agentId);
|
|
|
|
await saveResponseOutput(req, conversationId, responseId, response, agentId);
|
|
|
|
logger.debug(
|
|
`[Responses API] Stored response ${responseId} in conversation ${conversationId}`,
|
|
);
|
|
} catch (saveError) {
|
|
logger.error('[Responses API] Error saving response:', saveError);
|
|
// Don't fail the request if saving fails
|
|
}
|
|
}
|
|
|
|
res.json(response);
|
|
|
|
const duration = Date.now() - requestStartTime;
|
|
logger.debug(
|
|
`[Responses API] Request ${responseId} completed in ${duration}ms (non-streaming)`,
|
|
);
|
|
}
|
|
} catch (error) {
|
|
const errorMessage = error instanceof Error ? error.message : 'An error occurred';
|
|
logger.error('[Responses API] Error:', error);
|
|
|
|
// Check if we already started streaming (headers sent)
|
|
if (res.headersSent) {
|
|
// Headers already sent, write error event and close
|
|
writeDone(res);
|
|
res.end();
|
|
} else {
|
|
// Forward upstream provider status codes (e.g., Anthropic 400s) instead of masking as 500
|
|
const statusCode =
|
|
typeof error?.status === 'number' && error.status >= 400 && error.status < 600
|
|
? error.status
|
|
: 500;
|
|
const errorType = statusCode >= 400 && statusCode < 500 ? 'invalid_request' : 'server_error';
|
|
sendResponsesErrorResponse(res, statusCode, errorMessage, errorType);
|
|
}
|
|
}
|
|
};
|
|
|
|
/**
|
|
* List available agents as models - GET /v1/models (also works with /v1/responses/models)
|
|
*
|
|
* Returns a list of available agents the user has remote access to.
|
|
*
|
|
* @param {import('express').Request} req
|
|
* @param {import('express').Response} res
|
|
*/
|
|
const listModels = async (req, res) => {
|
|
try {
|
|
const userId = req.user?.id;
|
|
const userRole = req.user?.role;
|
|
|
|
if (!userId) {
|
|
return sendResponsesErrorResponse(res, 401, 'Authentication required', 'auth_error');
|
|
}
|
|
|
|
// Find agents the user has remote access to (VIEW permission on REMOTE_AGENT)
|
|
const accessibleAgentIds = await findAccessibleResources({
|
|
userId,
|
|
role: userRole,
|
|
resourceType: ResourceType.REMOTE_AGENT,
|
|
requiredPermissions: PermissionBits.VIEW,
|
|
});
|
|
|
|
// Get the accessible agents
|
|
let agents = [];
|
|
if (accessibleAgentIds.length > 0) {
|
|
agents = await db.getAgents({ _id: { $in: accessibleAgentIds } });
|
|
}
|
|
|
|
// Convert to models format
|
|
const models = agents.map((agent) => ({
|
|
id: agent.id,
|
|
object: 'model',
|
|
created: Math.floor(new Date(agent.createdAt).getTime() / 1000),
|
|
owned_by: agent.author ?? 'librechat',
|
|
// Additional metadata
|
|
name: agent.name,
|
|
description: agent.description,
|
|
provider: agent.provider,
|
|
}));
|
|
|
|
res.json({
|
|
object: 'list',
|
|
data: models,
|
|
});
|
|
} catch (error) {
|
|
logger.error('[Responses API] Error listing models:', error);
|
|
sendResponsesErrorResponse(
|
|
res,
|
|
500,
|
|
error instanceof Error ? error.message : 'Failed to list models',
|
|
'server_error',
|
|
);
|
|
}
|
|
};
|
|
|
|
/**
|
|
* Get Response - GET /v1/responses/:id
|
|
*
|
|
* Retrieves a stored response by its ID.
|
|
* The response ID maps to a conversationId in LibreChat's storage.
|
|
*
|
|
* @param {import('express').Request} req
|
|
* @param {import('express').Response} res
|
|
*/
|
|
const getResponse = async (req, res) => {
|
|
try {
|
|
const responseId = req.params.id;
|
|
const userId = req.user?.id;
|
|
|
|
if (!responseId) {
|
|
return sendResponsesErrorResponse(res, 400, 'Response ID is required');
|
|
}
|
|
|
|
// The responseId could be either the response ID or the conversation ID
|
|
// Try to find a conversation with this ID
|
|
const conversation = await db.getConvo(userId, responseId);
|
|
|
|
if (!conversation) {
|
|
return sendResponsesErrorResponse(
|
|
res,
|
|
404,
|
|
`Response not found: ${responseId}`,
|
|
'not_found',
|
|
'response_not_found',
|
|
);
|
|
}
|
|
|
|
// Load messages for this conversation
|
|
const messages = await db.getMessages({ conversationId: responseId, user: userId });
|
|
|
|
if (!messages || messages.length === 0) {
|
|
return sendResponsesErrorResponse(
|
|
res,
|
|
404,
|
|
`No messages found for response: ${responseId}`,
|
|
'not_found',
|
|
'response_not_found',
|
|
);
|
|
}
|
|
|
|
// Convert messages to Open Responses output format
|
|
const output = convertMessagesToOutputItems(messages);
|
|
|
|
// Find the last assistant message for usage info
|
|
const lastAssistantMessage = messages.filter((m) => !m.isCreatedByUser).pop();
|
|
|
|
// Build the response object
|
|
const response = {
|
|
id: responseId,
|
|
object: 'response',
|
|
created_at: Math.floor(new Date(conversation.createdAt || Date.now()).getTime() / 1000),
|
|
completed_at: Math.floor(new Date(conversation.updatedAt || Date.now()).getTime() / 1000),
|
|
status: 'completed',
|
|
incomplete_details: null,
|
|
model: conversation.agentId || conversation.model || 'unknown',
|
|
previous_response_id: null,
|
|
instructions: null,
|
|
output,
|
|
error: null,
|
|
tools: [],
|
|
tool_choice: 'auto',
|
|
truncation: 'disabled',
|
|
parallel_tool_calls: true,
|
|
text: { format: { type: 'text' } },
|
|
temperature: 1,
|
|
top_p: 1,
|
|
presence_penalty: 0,
|
|
frequency_penalty: 0,
|
|
top_logprobs: null,
|
|
reasoning: null,
|
|
user: userId,
|
|
usage: lastAssistantMessage?.tokenCount
|
|
? {
|
|
input_tokens: 0,
|
|
output_tokens: lastAssistantMessage.tokenCount,
|
|
total_tokens: lastAssistantMessage.tokenCount,
|
|
}
|
|
: null,
|
|
max_output_tokens: null,
|
|
max_tool_calls: null,
|
|
store: true,
|
|
background: false,
|
|
service_tier: 'default',
|
|
metadata: {},
|
|
safety_identifier: null,
|
|
prompt_cache_key: null,
|
|
};
|
|
|
|
res.json(response);
|
|
} catch (error) {
|
|
logger.error('[Responses API] Error getting response:', error);
|
|
sendResponsesErrorResponse(
|
|
res,
|
|
500,
|
|
error instanceof Error ? error.message : 'Failed to get response',
|
|
'server_error',
|
|
);
|
|
}
|
|
};
|
|
|
|
module.exports = {
|
|
createResponse,
|
|
getResponse,
|
|
listModels,
|
|
};
|