From fdc7e64bb78b326d94bec6b67370a32d02a3a713 Mon Sep 17 00:00:00 2001 From: Danny Avila Date: Tue, 16 Jun 2026 17:54:13 -0400 Subject: [PATCH] =?UTF-8?q?=F0=9F=AA=99=20feat:=20SDK-Aligned=20Context-Us?= =?UTF-8?q?age=20Projection=20(gauge=20for=20window-switch=20&=20snapshot-?= =?UTF-8?q?less=20branches)=20(#13801)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * πŸͺ™ feat: Context-usage projection β€” data-provider + client wiring Consumer side of the SDK-aligned context projection (agents `projectAgentContextUsage`). Adds the `/api/endpoints/context-projection` data-provider plumbing (endpoint, service, query key, `TContextProjectionRequest`) and a `useContextProjectionQuery` gated to fire only when no fresh snapshot covers the viewed branch. Wires `useTokenUsage` precedence to: live snapshot β†’ fresh persisted snapshot (window matches the resolved one) β†’ server projection β†’ per-message estimate. A model/window switch marks the baked snapshot stale (its `maxContextTokens` no longer matches) and falls to the projection β€” closing the gauge's window-switch (G1) and snapshot-less-branch (G2) gaps. Snapshot and projection share the render-relevant fields, so they render uniformly. Backend endpoint + agents version bump land in follow-up commits. Includes the design spec (CONTEXT_PROJECTION_SPEC.md). * πŸͺ™ feat: Context-projection backend endpoint POST /api/endpoints/context-projection β†’ resolveContextProjection (packages/api): reconstructs the viewed branch (parent-chain walk from messageId), resolves the agent config (instructions/provider/model/maxContextTokens), reuses LibreChat's stored per-message tokenCounts as the index map (no re-tokenizing), and calls the agents SDK projectAgentContextUsage β€” no model call. Thin controller injects db.getMessages/db.getAgent; route mirrors /token-config. First cut targets message-windowing accuracy; tool-schema tokens are deferred to a follow-up that reuses the full initializeAgent path. * 🩹 fix: Codex review on context projection (G1 guard, IDOR, recount, summary) - Guard `currentActive` against a stale window: a model/window switch on the current branch left the live snapshot outranking the projection (G1 didn't fire). Now defers to the projection unless streaming or the window matches. - Scope branch lookups to the authenticated user (`getMessages` filter + injected `userId`) β€” was loading any conversation by id (IDOR). - Recount messages with no stored `tokenCount` via the tokenizer instead of charging 0, so snapshot-less/imported histories don't under-report. - Fall back (null) for already-summarized branches rather than projecting from the full raw parent chain (the next call would send summary + tail); the client's summary-baseline-aware estimate handles them until a follow-up replays the summary boundary. * 🩹 fix: Codex round 2 β€” drop agent load, summary marker, edit-invalidation - Stop loading agent/model-spec config server-side (closes the agent-access IDOR and the spec-prompt special-casing). Provider/model/window now come from the client-resolved request (`limits.endpoint`/model β€” the agent's real provider, not the `agents` endpoint, so the tokenizer is right). Agent/spec/ promptPrefix instructions are uniformly deferred to the full-fidelity follow-up. - Detect summarized branches via the live path's `metadata.summaryUsedTokens` marker (was the wrong `summaryTokenCount` field) and fall back to the summary-aware estimate. - Invalidate the projection query on in-place message edits via a branch content `revision` in the cache key (the tail id is unchanged on edit). Deferred (valid, not a regression): same-window endpoint/model switch keeps a window-matched snapshot β€” needs endpoint/model persisted on the snapshot, which lands with the fidelity follow-up. Smoke-tested: fits / prunes / summarizedβ†’null / no-windowβ†’null. * πŸ›‘οΈ fix: make context projection strictly additive (no-regression) Revert the G1 window-match guard on the live/branch snapshot. When no explicit maxContextTokens is set (the common default), the SDK's snapshot window is reserve-derived (~0.9Β·(modelContext βˆ’ maxOutputTokens)) while useTokenLimits resolves the raw model context β€” so `snapshot.maxContextTokens === resolvedMax` is false for the SAME model, and the guard would wrongly drop a valid current-branch snapshot to projection/estimate post-stream (a regression in the default case, per initialize.ts:1240-1243). The projection now activates ONLY for snapshot-less branches (G2): the precedence is live snapshot β†’ persisted branch snapshot β†’ projection β†’ estimate, where the first two are byte-for-byte the prior behavior and the projection just slots ahead of the estimate. Window/model-switch (G1) detection needs the snapshot to carry its model/window and defers to the fidelity follow-up. * 🩹 fix: surface projections as estimates, not authoritative snapshots A first-cut projection carries the SDK's windowing but omits instruction/tool overhead, so rendering it as `isEstimate: false` showed a confident under-count for snapshot-less branches. Mark projection-sourced views `isEstimate: true` + `snapshotActive: false` (and drop the snapshot field) so they present as a better estimate than sumBranch β€” improved used/window number, estimate framing, no misleading granular breakdown with ~0 tools. Real snapshots stay authoritative. (Codex round 3, projection.ts:139.) * 🧹 chore: drop CONTEXT_PROJECTION_SPEC.md from the PR * 🎨 style: fix import-sort order in projection.ts (CI sort-imports check) * πŸ”§ chore: update @librechat/agents dependency to version 3.2.36 in package-lock.json and related package.json files * chore: npm audit fix * 🎨 style: fix import-sort order in data-service.ts (CI sort-imports check) * 🩹 fix: drop dead calibrationRatio in projectionParams (tsc never error) Inside the ternary, branchSnapshot is narrowed to null (the gate is ), so accessed a property on (frontend typecheck failure). It was also dead β€” there is never a snapshot to seed from in this branch β€” so just remove it. * Revert "chore: npm audit fix" This reverts commit 4cdb862d0c72f4656e5fcb85dcea4f6643e528d9. --- api/package.json | 2 +- .../ContextProjectionController.js | 31 ++++ api/server/routes/endpoints.js | 2 + client/src/data-provider/Endpoints/queries.ts | 39 +++++ client/src/hooks/Chat/useTokenUsage.ts | 119 ++++++++++---- package-lock.json | 10 +- packages/api/package.json | 2 +- packages/api/src/endpoints/index.ts | 1 + packages/api/src/endpoints/projection.ts | 145 ++++++++++++++++++ packages/data-provider/src/api-endpoints.ts | 2 + packages/data-provider/src/data-service.ts | 7 + packages/data-provider/src/keys.ts | 1 + packages/data-provider/src/types/runs.ts | 23 +++ 13 files changed, 348 insertions(+), 36 deletions(-) create mode 100644 api/server/controllers/ContextProjectionController.js create mode 100644 packages/api/src/endpoints/projection.ts diff --git a/api/package.json b/api/package.json index 13cd274a5a..dc8eb96686 100644 --- a/api/package.json +++ b/api/package.json @@ -46,7 +46,7 @@ "@azure/storage-blob": "^12.30.0", "@google/genai": "^2.8.0", "@keyv/redis": "^4.3.3", - "@librechat/agents": "^3.2.35", + "@librechat/agents": "^3.2.36", "@librechat/api": "*", "@librechat/data-schemas": "*", "@microsoft/microsoft-graph-client": "^3.0.7", diff --git a/api/server/controllers/ContextProjectionController.js b/api/server/controllers/ContextProjectionController.js new file mode 100644 index 0000000000..eaf9592e73 --- /dev/null +++ b/api/server/controllers/ContextProjectionController.js @@ -0,0 +1,31 @@ +const { logger } = require('@librechat/data-schemas'); +const { resolveContextProjection } = require('@librechat/api'); +const db = require('~/models'); + +/** + * Returns a server-side context-usage projection for the viewed branch + config + * (agents SDK, no model call) β€” powers the gauge for snapshot-less branches and + * after a model/window switch. Resolution lives in `@librechat/api`; this + * controller only injects request-scoped model accessors. + * @param {ServerRequest} req + * @param {ServerResponse} res + */ +async function contextProjectionController(req, res) { + try { + const params = req.body ?? {}; + if (!params.conversationId || !params.messageId) { + res.json(null); + return; + } + const projection = await resolveContextProjection( + { userId: req.user?.id, getMessages: db.getMessages }, + params, + ); + res.json(projection ?? null); + } catch (error) { + logger.error('[contextProjectionController]', error); + res.status(500).json({ error: 'Failed to resolve context projection' }); + } +} + +module.exports = contextProjectionController; diff --git a/api/server/routes/endpoints.js b/api/server/routes/endpoints.js index 8b1fceccc4..b11de153df 100644 --- a/api/server/routes/endpoints.js +++ b/api/server/routes/endpoints.js @@ -3,10 +3,12 @@ const requireJwtAuth = require('~/server/middleware/requireJwtAuth'); const configMiddleware = require('~/server/middleware/config/app'); const endpointController = require('~/server/controllers/EndpointController'); const tokenConfigController = require('~/server/controllers/TokenConfigController'); +const contextProjectionController = require('~/server/controllers/ContextProjectionController'); const router = express.Router(); /** Auth required for role/tenant-scoped endpoint config resolution. */ router.get('/', requireJwtAuth, endpointController); router.get('/token-config', requireJwtAuth, configMiddleware, tokenConfigController); +router.post('/context-projection', requireJwtAuth, configMiddleware, contextProjectionController); module.exports = router; diff --git a/client/src/data-provider/Endpoints/queries.ts b/client/src/data-provider/Endpoints/queries.ts index 294f81597d..1ea8b3c0aa 100644 --- a/client/src/data-provider/Endpoints/queries.ts +++ b/client/src/data-provider/Endpoints/queries.ts @@ -41,6 +41,45 @@ export const useTokenConfigQuery = ( }); }; +/** + * Server-side context-usage projection for the viewed branch + resolved config + * (agents SDK, no model call). Keyed on the fields that change what the next + * call would send β€” branch tail, endpoint/model/agent, window β€” so a branch or + * model/window switch refetches. Disabled until a branch tail is known. + */ +export const useContextProjectionQuery = ( + params: t.TContextProjectionRequest | null, + config?: UseQueryOptions, +): QueryObserverResult => { + const queriesEnabled = useRecoilValue(store.queriesEnabled); + return useQuery( + [ + QueryKeys.contextProjection, + params?.conversationId, + params?.messageId, + params?.endpoint, + params?.model, + params?.agentId, + params?.maxContextTokens, + params?.revision, + ], + () => dataService.getContextProjection(params as t.TContextProjectionRequest), + { + staleTime: Infinity, + refetchOnWindowFocus: false, + refetchOnReconnect: false, + refetchOnMount: false, + ...config, + enabled: + (config?.enabled ?? true) === true && + queriesEnabled && + params != null && + (params.conversationId?.length ?? 0) > 0 && + (params.messageId?.length ?? 0) > 0, + }, + ); +}; + /** * Auth-aware query key so unauthenticated (login page) and authenticated * (chat page) configs are cached independently, preventing stale diff --git a/client/src/hooks/Chat/useTokenUsage.ts b/client/src/hooks/Chat/useTokenUsage.ts index a241c33455..6a6d4ff845 100644 --- a/client/src/hooks/Chat/useTokenUsage.ts +++ b/client/src/hooks/Chat/useTokenUsage.ts @@ -2,7 +2,13 @@ import { useEffect, useMemo, useRef } from 'react'; import { useAtomValue, useSetAtom } from 'jotai'; import { useQueryClient } from '@tanstack/react-query'; import { Constants, QueryKeys } from 'librechat-data-provider'; -import type { TMessage, TConversation, TModelTokenomics } from 'librechat-data-provider'; +import type { + TMessage, + TConversation, + TModelTokenomics, + TContextUsageEvent, + TContextProjectionRequest, +} from 'librechat-data-provider'; import type { BranchTotals, BranchUsage } from '~/utils/tokens'; import type { ContextSnapshot } from '~/store/usage'; import { @@ -24,6 +30,7 @@ import { findBranchSnapshotAnchor, } from '~/utils'; import { useLatestMessageId } from '~/hooks/Messages/useLatestMessage'; +import { useContextProjectionQuery } from '~/data-provider'; import useTokenLimits from './useTokenLimits'; export interface TokenUsageParams { @@ -79,6 +86,55 @@ export default function useTokenUsage({ const setTotalUsage = useSetAtom(totalUsageFamily(conversationKey)); const limits = useTokenLimits(conversation); + /** Deepest persisted/live snapshot on the viewed branch (present only for + * turns generated with the feature on). Gates the projection fetch and is a + * render source. */ + const branchSnapshot = useMemo(() => { + if (snapshotsByAnchor.size === 0) { + return null; + } + const anchor = findBranchSnapshotAnchor( + conversationKey, + branchTotals.tailId, + snapshotsByAnchor, + ); + return anchor != null ? (snapshotsByAnchor.get(anchor) ?? null) : null; + }, [conversationKey, branchTotals.tailId, snapshotsByAnchor]); + + const resolvedMax = limits.maxContextTokens; + + /** Project the branch (agents SDK, no model call) ONLY when no persisted/live + * snapshot covers it β€” snapshot-less branches (G2: pre-feature history, + * imports, never-generated branches). A present snapshot stays authoritative; + * reliable window-switch (G1) detection needs the snapshot to carry its + * model/window (deferred to the fidelity follow-up), and the SDK window + * (reserve-derived) doesn't equal the client-resolved raw window, so we must + * NOT mis-flag a valid snapshot as stale here. Cached + refetched by branch/ + * endpoint/model/window/revision. */ + const projectionParams: TContextProjectionRequest | null = + !isSubmitting && + branchSnapshot == null && + conversation?.conversationId != null && + conversation.conversationId !== Constants.NEW_CONVO && + branchTotals.tailId != null && + conversation.endpoint != null + ? { + conversationId: conversation.conversationId, + messageId: branchTotals.tailId, + /** Resolved provider/model (e.g. an agent's actual provider, not the + * `agents` endpoint) so the server picks the right tokenizer. */ + endpoint: limits.endpoint || conversation.endpoint, + model: limits.model || conversation.model || undefined, + agentId: conversation.agent_id ?? undefined, + spec: conversation.spec ?? undefined, + maxContextTokens: resolvedMax, + /** Content revision so an in-place message edit (same tail id) refetches. */ + revision: branchTotals.input + branchTotals.output, + } + : null; + const { data: projectionData } = useContextProjectionQuery(projectionParams); + const projection = projectionData ?? null; + /** Branch/total provider usage is index-derived; the in-flight response is * the only live add (the pending holder), counted into both β€” it sits on the * active branch tail and inside the conversation. The backend prices each @@ -186,41 +242,46 @@ export default function useTokenUsage({ snapshot != null && (isSubmitting || (snapshot.anchorMessageId != null && branchTotals.containsAnchor)); - /** When the live snapshot belongs to another branch, recover this branch's - * own finalized snapshot (if it was generated this session) by walking the - * branch for its deepest stored anchor β€” keeps the granular rows on switch - * instead of dropping to coarse totals. */ - let activeSnapshot: ContextSnapshot | null = currentActive ? snapshot : null; - if (activeSnapshot == null && !isSubmitting && snapshotsByAnchor.size > 0) { - const anchor = findBranchSnapshotAnchor( - conversationKey, - branchTotals.tailId, - snapshotsByAnchor, - ); - activeSnapshot = anchor != null ? (snapshotsByAnchor.get(anchor) ?? null) : null; + /** Precedence: live/active snapshot β†’ persisted branch snapshot β†’ server + * projection (snapshot-less branches, G2) β†’ per-message estimate. The first + * two preserve the pre-projection behavior exactly; the projection only + * slots in ahead of the estimate when no snapshot exists. Snapshot and + * projection share the render-relevant fields, so they render uniformly. */ + let effective: ContextSnapshot | TContextUsageEvent | null = null; + /** A server projection is the SDK's windowing but, in this first cut, omits + * instruction/tool overhead β€” so it's surfaced as an ESTIMATE (a better one + * than sumBranch), never a false-authoritative number. Real snapshots stay + * authoritative. */ + let projected = false; + if (currentActive) { + effective = snapshot; + } else if (branchSnapshot != null) { + effective = branchSnapshot; + } else if (projection != null) { + effective = projection; + projected = true; } - if (activeSnapshot != null) { - const breakdown = activeSnapshot.breakdown; - const maxTokens = activeSnapshot.contextBudget ?? breakdown.maxContextTokens; - const instructionTokens = - activeSnapshot.effectiveInstructionTokens ?? breakdown.instructionTokens; + if (effective != null) { + const breakdown = effective.breakdown; + const maxTokens = effective.contextBudget ?? breakdown.maxContextTokens; + const instructionTokens = effective.effectiveInstructionTokens ?? breakdown.instructionTokens; const baseUsed = - activeSnapshot.remainingContextTokens != null - ? maxTokens - activeSnapshot.remainingContextTokens + effective.remainingContextTokens != null + ? maxTokens - effective.remainingContextTokens : instructionTokens + breakdown.messageTokens; - /** The snapshot is pre-invoke: in-flight output rides on `liveTokens` - * (0 unless streaming this branch), the last call's finalized output on - * `completedOutputTokens`. */ + /** The snapshot/projection is pre-invoke: in-flight output rides on + * `liveTokens` (0 unless streaming this branch), the last call's finalized + * output on `completedOutputTokens` (absent on a projection β†’ 0). */ const usedTokens = - Math.max(0, baseUsed) + liveTokens + (activeSnapshot.completedOutputTokens ?? 0); + Math.max(0, baseUsed) + liveTokens + (effective.completedOutputTokens ?? 0); return { usedTokens, maxTokens, percent: maxTokens > 0 ? Math.min((usedTokens / maxTokens) * 100, 100) : 0, - isEstimate: false, - snapshot: activeSnapshot, - snapshotActive: true, + isEstimate: projected, + snapshot: projected ? null : (effective as ContextSnapshot), + snapshotActive: !projected, branchTotals, branchUsage, totalUsage, @@ -266,7 +327,7 @@ export default function useTokenUsage({ hasUsage, liveTokens, limits, - snapshotsByAnchor, - conversationKey, + branchSnapshot, + projection, ]); } diff --git a/package-lock.json b/package-lock.json index 0557171349..66cf865512 100644 --- a/package-lock.json +++ b/package-lock.json @@ -61,7 +61,7 @@ "@azure/storage-blob": "^12.30.0", "@google/genai": "^2.8.0", "@keyv/redis": "^4.3.3", - "@librechat/agents": "^3.2.35", + "@librechat/agents": "^3.2.36", "@librechat/api": "*", "@librechat/data-schemas": "*", "@microsoft/microsoft-graph-client": "^3.0.7", @@ -11685,9 +11685,9 @@ } }, "node_modules/@librechat/agents": { - "version": "3.2.35", - "resolved": "https://registry.npmjs.org/@librechat/agents/-/agents-3.2.35.tgz", - "integrity": "sha512-k/6e3WOkYSgqZVFtgw5XeNpuA+1+y0vX4u4pqGeX3MM3ERDk729OQLdO1qFbOCxkcP2hCZppMzZ5JExSyKdplw==", + "version": "3.2.36", + "resolved": "https://registry.npmjs.org/@librechat/agents/-/agents-3.2.36.tgz", + "integrity": "sha512-y9Bm5CMACNZh6OfDCUcABK84WUYndAfgO36Bu5EEyLqDm1S6TvoYUwQ/6P3jdsMy+Vj1ESrxAfCdQgy7fxtcjg==", "license": "MIT", "dependencies": { "@anthropic-ai/sdk": "^0.92.0", @@ -44290,7 +44290,7 @@ "@azure/storage-blob": "^12.30.0", "@google/genai": "^2.8.0", "@keyv/redis": "^4.3.3", - "@librechat/agents": "^3.2.35", + "@librechat/agents": "^3.2.36", "@librechat/data-schemas": "*", "@modelcontextprotocol/sdk": "^1.29.0", "@opentelemetry/api": "^1.9.0", diff --git a/packages/api/package.json b/packages/api/package.json index a4713d9144..82e04dae28 100644 --- a/packages/api/package.json +++ b/packages/api/package.json @@ -113,7 +113,7 @@ "@azure/storage-blob": "^12.30.0", "@google/genai": "^2.8.0", "@keyv/redis": "^4.3.3", - "@librechat/agents": "^3.2.35", + "@librechat/agents": "^3.2.36", "@librechat/data-schemas": "*", "@modelcontextprotocol/sdk": "^1.29.0", "@opentelemetry/api": "^1.9.0", diff --git a/packages/api/src/endpoints/index.ts b/packages/api/src/endpoints/index.ts index 9e6e9dbac0..4be03df1e3 100644 --- a/packages/api/src/endpoints/index.ts +++ b/packages/api/src/endpoints/index.ts @@ -6,4 +6,5 @@ export * from './google'; export * from './models'; export * from './openai'; export * from './pricing'; +export * from './projection'; export * from './tokenConfig'; diff --git a/packages/api/src/endpoints/projection.ts b/packages/api/src/endpoints/projection.ts new file mode 100644 index 0000000000..62bc40dfed --- /dev/null +++ b/packages/api/src/endpoints/projection.ts @@ -0,0 +1,145 @@ +import { HumanMessage, AIMessage } from '@langchain/core/messages'; +import { Providers, createTokenCounter, projectAgentContextUsage } from '@librechat/agents'; +import type { TContextProjectionRequest, TContextUsageEvent } from 'librechat-data-provider'; +import type { BaseMessage } from '@langchain/core/messages'; + +interface ProjectionMessage { + messageId: string; + parentMessageId?: string | null; + tokenCount?: number; + isCreatedByUser?: boolean; + text?: string; + /** Compaction marker written by the live path (`agents/usage.ts`); its + * presence means the next call sends the summary + tail, not this raw chain. */ + metadata?: { summaryUsedTokens?: number }; +} + +export interface ContextProjectionDeps { + /** Authenticated requester β€” branch lookups are scoped to this user. */ + userId?: string; + getMessages: ( + filter: { conversationId: string; user?: string }, + select?: string, + ) => Promise; +} + +/** + * Walks the parent chain from `tailId` to root and returns the branch messages + * oldestβ†’newest. The visited set guards against cycles / self-referential links. + */ +function resolveBranch(messages: ProjectionMessage[], tailId: string): ProjectionMessage[] { + const byId = new Map(); + for (const message of messages) { + byId.set(message.messageId, message); + } + const branch: ProjectionMessage[] = []; + const seen = new Set(); + let currentId: string | null | undefined = tailId; + while (currentId != null && !seen.has(currentId)) { + const message = byId.get(currentId); + if (message == null) { + break; + } + seen.add(currentId); + branch.push(message); + currentId = message.parentMessageId; + } + return branch.reverse(); +} + +/** Maps an endpoint/provider string to the agents `Providers` enum. */ +function resolveProvider(value?: string): Providers { + if (value == null || value === '') { + return Providers.OPENAI; + } + const lower = value.toLowerCase(); + for (const provider of Object.values(Providers)) { + if (provider.toLowerCase() === lower) { + return provider; + } + } + if (lower.includes('anthropic') || lower.includes('claude')) { + return Providers.ANTHROPIC; + } + if (lower.includes('google') || lower.includes('gemini') || lower.includes('vertex')) { + return Providers.GOOGLE; + } + if (lower.includes('bedrock')) { + return Providers.BEDROCK; + } + return Providers.OPENAI; +} + +/** + * Server-side context-usage projection: reconstructs the viewed branch and asks + * the agents SDK what the next call's context would be, WITHOUT invoking the + * model. Provider/model/window come from the (client-resolved) request β€” no + * agent or model-spec config is loaded here, so there is no cross-user config + * exposure. Reuses LibreChat's already-calibrated per-message `tokenCount`s (no + * re-tokenizing). Returns null when there is no resolvable context window. + * NOTE: this first cut targets message-windowing accuracy β€” instruction and + * tool-schema tokens (agent instructions, `promptPrefix`, model-spec presets, + * tool schemas) are NOT yet included; a follow-up will reuse the full + * `initializeAgent`/send path for exact overhead and proper access control. + */ +export async function resolveContextProjection( + deps: ContextProjectionDeps, + params: TContextProjectionRequest, +): Promise { + const maxContextTokens = params.maxContextTokens; + if (maxContextTokens == null || maxContextTokens <= 0) { + return null; + } + + const stored = await deps.getMessages( + { conversationId: params.conversationId, user: deps.userId }, + 'messageId parentMessageId tokenCount isCreatedByUser text metadata', + ); + const branch = resolveBranch(stored, params.messageId); + if (branch.length === 0) { + return null; + } + + /** A summarized/compacted branch's next call sends the saved summary + the + * post-summary tail, NOT this raw parent chain β€” projecting from the full + * history would prune/count the wrong context and omit the summary. Detect it + * via the live path's `metadata.summaryUsedTokens` marker and fall back (null) + * so the client's summary-baseline-aware estimate handles these branches until + * a follow-up replays the summary boundary. */ + if (branch.some((message) => (message.metadata?.summaryUsedTokens ?? 0) > 0)) { + return null; + } + + const model = params.model; + const encoding = (model ?? '').toLowerCase().includes('claude') ? 'claude' : 'o200k_base'; + const tokenCounter = await createTokenCounter(encoding); + + const messages: BaseMessage[] = []; + const indexTokenCountMap: Record = {}; + for (let i = 0; i < branch.length; i++) { + const message = branch[i]; + const text = message.text ?? ''; + const lcMessage = + message.isCreatedByUser === true ? new HumanMessage(text) : new AIMessage(text); + messages.push(lcMessage); + /** Recount messages with no stored count (imported / pre-feature) rather + * than charging 0 β€” a real 0 and "unknown" must not collapse, or the + * snapshot-less histories this endpoint targets would under-report. */ + indexTokenCountMap[String(i)] = + message.tokenCount != null && message.tokenCount > 0 + ? message.tokenCount + : tokenCounter(lcMessage); + } + + return projectAgentContextUsage({ + agent: { + agentId: params.agentId ?? 'projection', + provider: resolveProvider(params.endpoint), + maxContextTokens, + }, + messages, + tokenCounter, + indexTokenCountMap, + calibrationRatio: params.calibrationRatio, + }); +} diff --git a/packages/data-provider/src/api-endpoints.ts b/packages/data-provider/src/api-endpoints.ts index 09864bea43..73f45096bc 100644 --- a/packages/data-provider/src/api-endpoints.ts +++ b/packages/data-provider/src/api-endpoints.ts @@ -150,6 +150,8 @@ export const aiEndpoints = () => `${BASE_URL}/api/endpoints`; export const tokenConfig = () => `${BASE_URL}/api/endpoints/token-config`; +export const contextProjection = () => `${BASE_URL}/api/endpoints/context-projection`; + export const models = () => `${BASE_URL}/api/models`; export const tokenizer = () => `${BASE_URL}/api/tokenizer`; diff --git a/packages/data-provider/src/data-service.ts b/packages/data-provider/src/data-service.ts index e008e8c514..d9fa692cec 100644 --- a/packages/data-provider/src/data-service.ts +++ b/packages/data-provider/src/data-service.ts @@ -1,4 +1,5 @@ import type { AxiosResponse } from 'axios'; +import type { TContextProjectionRequest, TContextUsageEvent } from './types/runs'; import type { TFileConfig } from './file-config'; import type * as t from './types'; import * as permissions from './accessPermissions'; @@ -249,6 +250,12 @@ export const getTokenConfig = (): Promise => { return request.get(endpoints.tokenConfig()); }; +export const getContextProjection = ( + payload: TContextProjectionRequest, +): Promise => { + return request.post(endpoints.contextProjection(), payload); +}; + export const getModels = async (): Promise => { return request.get(endpoints.models()); }; diff --git a/packages/data-provider/src/keys.ts b/packages/data-provider/src/keys.ts index 04a65beb51..0f15a7d05f 100644 --- a/packages/data-provider/src/keys.ts +++ b/packages/data-provider/src/keys.ts @@ -13,6 +13,7 @@ export enum QueryKeys { balance = 'balance', endpoints = 'endpoints', tokenConfig = 'tokenConfig', + contextProjection = 'contextProjection', presets = 'presets', searchResults = 'searchResults', tokenCount = 'tokenCount', diff --git a/packages/data-provider/src/types/runs.ts b/packages/data-provider/src/types/runs.ts index 1fdc07cd81..8e1033c07f 100644 --- a/packages/data-provider/src/types/runs.ts +++ b/packages/data-provider/src/types/runs.ts @@ -88,6 +88,29 @@ export type TContextUsageEvent = { completedOutputTokens?: number; }; +/** + * Request payload for a server-side context-usage projection: "what context + * would the next call send for this branch under this config", computed by the + * agents SDK without invoking the model. Powers the gauge in states the live + * snapshot can't cover (page load of a snapshot-less branch, window/model + * switch). `messageId` is the viewed branch's tail; the server walks its parent + * chain. + */ +export type TContextProjectionRequest = { + conversationId: string; + messageId: string; + endpoint: string; + model?: string; + agentId?: string; + spec?: string; + maxContextTokens?: number; + /** Provider-calibrated ratio from a prior snapshot, applied as a static seed. */ + calibrationRatio?: number; + /** Client-only cache-bust: a branch content revision so a message edit + * (which keeps the same tail id) refetches. The server ignores it. */ + revision?: number; +}; + /** * Per-response usage rollup persisted on `responseMessage.metadata.usage`, in * display units (input excludes cache; output includes repaired completion).