mirror of
https://github.com/danny-avila/LibreChat.git
synced 2026-08-27 04:07:05 +00:00
🪙 feat: SDK-Aligned Context-Usage Projection (gauge for window-switch & snapshot-less branches) (#13801)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions
* 🪙 feat: Context-usage projection — data-provider + client wiring
Consumer side of the SDK-aligned context projection (agents
`projectAgentContextUsage`). Adds the `/api/endpoints/context-projection`
data-provider plumbing (endpoint, service, query key, `TContextProjectionRequest`)
and a `useContextProjectionQuery` gated to fire only when no fresh snapshot
covers the viewed branch.
Wires `useTokenUsage` precedence to: live snapshot → fresh persisted snapshot
(window matches the resolved one) → server projection → per-message estimate.
A model/window switch marks the baked snapshot stale (its `maxContextTokens`
no longer matches) and falls to the projection — closing the gauge's
window-switch (G1) and snapshot-less-branch (G2) gaps. Snapshot and projection
share the render-relevant fields, so they render uniformly.
Backend endpoint + agents version bump land in follow-up commits. Includes the
design spec (CONTEXT_PROJECTION_SPEC.md).
* 🪙 feat: Context-projection backend endpoint
POST /api/endpoints/context-projection → resolveContextProjection (packages/api):
reconstructs the viewed branch (parent-chain walk from messageId), resolves the
agent config (instructions/provider/model/maxContextTokens), reuses LibreChat's
stored per-message tokenCounts as the index map (no re-tokenizing), and calls
the agents SDK projectAgentContextUsage — no model call. Thin controller injects
db.getMessages/db.getAgent; route mirrors /token-config.
First cut targets message-windowing accuracy; tool-schema tokens are deferred to
a follow-up that reuses the full initializeAgent path.
* 🩹 fix: Codex review on context projection (G1 guard, IDOR, recount, summary)
- Guard `currentActive` against a stale window: a model/window switch on the
current branch left the live snapshot outranking the projection (G1 didn't
fire). Now defers to the projection unless streaming or the window matches.
- Scope branch lookups to the authenticated user (`getMessages` filter +
injected `userId`) — was loading any conversation by id (IDOR).
- Recount messages with no stored `tokenCount` via the tokenizer instead of
charging 0, so snapshot-less/imported histories don't under-report.
- Fall back (null) for already-summarized branches rather than projecting from
the full raw parent chain (the next call would send summary + tail); the
client's summary-baseline-aware estimate handles them until a follow-up
replays the summary boundary.
* 🩹 fix: Codex round 2 — drop agent load, summary marker, edit-invalidation
- Stop loading agent/model-spec config server-side (closes the agent-access
IDOR and the spec-prompt special-casing). Provider/model/window now come from
the client-resolved request (`limits.endpoint`/model — the agent's real
provider, not the `agents` endpoint, so the tokenizer is right). Agent/spec/
promptPrefix instructions are uniformly deferred to the full-fidelity follow-up.
- Detect summarized branches via the live path's `metadata.summaryUsedTokens`
marker (was the wrong `summaryTokenCount` field) and fall back to the
summary-aware estimate.
- Invalidate the projection query on in-place message edits via a branch
content `revision` in the cache key (the tail id is unchanged on edit).
Deferred (valid, not a regression): same-window endpoint/model switch keeps a
window-matched snapshot — needs endpoint/model persisted on the snapshot, which
lands with the fidelity follow-up. Smoke-tested: fits / prunes / summarized→null
/ no-window→null.
* 🛡️ fix: make context projection strictly additive (no-regression)
Revert the G1 window-match guard on the live/branch snapshot. When no explicit
maxContextTokens is set (the common default), the SDK's snapshot window is
reserve-derived (~0.9·(modelContext − maxOutputTokens)) while useTokenLimits
resolves the raw model context — so `snapshot.maxContextTokens === resolvedMax`
is false for the SAME model, and the guard would wrongly drop a valid
current-branch snapshot to projection/estimate post-stream (a regression in the
default case, per initialize.ts:1240-1243).
The projection now activates ONLY for snapshot-less branches (G2): the
precedence is live snapshot → persisted branch snapshot → projection → estimate,
where the first two are byte-for-byte the prior behavior and the projection just
slots ahead of the estimate. Window/model-switch (G1) detection needs the
snapshot to carry its model/window and defers to the fidelity follow-up.
* 🩹 fix: surface projections as estimates, not authoritative snapshots
A first-cut projection carries the SDK's windowing but omits instruction/tool
overhead, so rendering it as `isEstimate: false` showed a confident under-count
for snapshot-less branches. Mark projection-sourced views `isEstimate: true` +
`snapshotActive: false` (and drop the snapshot field) so they present as a
better estimate than sumBranch — improved used/window number, estimate framing,
no misleading granular breakdown with ~0 tools. Real snapshots stay
authoritative. (Codex round 3, projection.ts:139.)
* 🧹 chore: drop CONTEXT_PROJECTION_SPEC.md from the PR
* 🎨 style: fix import-sort order in projection.ts (CI sort-imports check)
* 🔧 chore: update @librechat/agents dependency to version 3.2.36 in package-lock.json and related package.json files
* chore: npm audit fix
* 🎨 style: fix import-sort order in data-service.ts (CI sort-imports check)
* 🩹 fix: drop dead calibrationRatio in projectionParams (tsc never error)
Inside the ternary, branchSnapshot is narrowed to null (the gate is
), so accessed a
property on (frontend typecheck failure). It was also dead — there is
never a snapshot to seed from in this branch — so just remove it.
* Revert "chore: npm audit fix"
This reverts commit 4cdb862d0c.
This commit is contained in:
parent
c820dfb9a0
commit
fdc7e64bb7
13 changed files with 348 additions and 36 deletions
|
|
@ -41,6 +41,45 @@ export const useTokenConfigQuery = (
|
|||
});
|
||||
};
|
||||
|
||||
/**
|
||||
* Server-side context-usage projection for the viewed branch + resolved config
|
||||
* (agents SDK, no model call). Keyed on the fields that change what the next
|
||||
* call would send — branch tail, endpoint/model/agent, window — so a branch or
|
||||
* model/window switch refetches. Disabled until a branch tail is known.
|
||||
*/
|
||||
export const useContextProjectionQuery = (
|
||||
params: t.TContextProjectionRequest | null,
|
||||
config?: UseQueryOptions<t.TContextUsageEvent | null>,
|
||||
): QueryObserverResult<t.TContextUsageEvent | null> => {
|
||||
const queriesEnabled = useRecoilValue<boolean>(store.queriesEnabled);
|
||||
return useQuery<t.TContextUsageEvent | null>(
|
||||
[
|
||||
QueryKeys.contextProjection,
|
||||
params?.conversationId,
|
||||
params?.messageId,
|
||||
params?.endpoint,
|
||||
params?.model,
|
||||
params?.agentId,
|
||||
params?.maxContextTokens,
|
||||
params?.revision,
|
||||
],
|
||||
() => dataService.getContextProjection(params as t.TContextProjectionRequest),
|
||||
{
|
||||
staleTime: Infinity,
|
||||
refetchOnWindowFocus: false,
|
||||
refetchOnReconnect: false,
|
||||
refetchOnMount: false,
|
||||
...config,
|
||||
enabled:
|
||||
(config?.enabled ?? true) === true &&
|
||||
queriesEnabled &&
|
||||
params != null &&
|
||||
(params.conversationId?.length ?? 0) > 0 &&
|
||||
(params.messageId?.length ?? 0) > 0,
|
||||
},
|
||||
);
|
||||
};
|
||||
|
||||
/**
|
||||
* Auth-aware query key so unauthenticated (login page) and authenticated
|
||||
* (chat page) configs are cached independently, preventing stale
|
||||
|
|
|
|||
|
|
@ -2,7 +2,13 @@ import { useEffect, useMemo, useRef } from 'react';
|
|||
import { useAtomValue, useSetAtom } from 'jotai';
|
||||
import { useQueryClient } from '@tanstack/react-query';
|
||||
import { Constants, QueryKeys } from 'librechat-data-provider';
|
||||
import type { TMessage, TConversation, TModelTokenomics } from 'librechat-data-provider';
|
||||
import type {
|
||||
TMessage,
|
||||
TConversation,
|
||||
TModelTokenomics,
|
||||
TContextUsageEvent,
|
||||
TContextProjectionRequest,
|
||||
} from 'librechat-data-provider';
|
||||
import type { BranchTotals, BranchUsage } from '~/utils/tokens';
|
||||
import type { ContextSnapshot } from '~/store/usage';
|
||||
import {
|
||||
|
|
@ -24,6 +30,7 @@ import {
|
|||
findBranchSnapshotAnchor,
|
||||
} from '~/utils';
|
||||
import { useLatestMessageId } from '~/hooks/Messages/useLatestMessage';
|
||||
import { useContextProjectionQuery } from '~/data-provider';
|
||||
import useTokenLimits from './useTokenLimits';
|
||||
|
||||
export interface TokenUsageParams {
|
||||
|
|
@ -79,6 +86,55 @@ export default function useTokenUsage({
|
|||
const setTotalUsage = useSetAtom(totalUsageFamily(conversationKey));
|
||||
const limits = useTokenLimits(conversation);
|
||||
|
||||
/** Deepest persisted/live snapshot on the viewed branch (present only for
|
||||
* turns generated with the feature on). Gates the projection fetch and is a
|
||||
* render source. */
|
||||
const branchSnapshot = useMemo(() => {
|
||||
if (snapshotsByAnchor.size === 0) {
|
||||
return null;
|
||||
}
|
||||
const anchor = findBranchSnapshotAnchor(
|
||||
conversationKey,
|
||||
branchTotals.tailId,
|
||||
snapshotsByAnchor,
|
||||
);
|
||||
return anchor != null ? (snapshotsByAnchor.get(anchor) ?? null) : null;
|
||||
}, [conversationKey, branchTotals.tailId, snapshotsByAnchor]);
|
||||
|
||||
const resolvedMax = limits.maxContextTokens;
|
||||
|
||||
/** Project the branch (agents SDK, no model call) ONLY when no persisted/live
|
||||
* snapshot covers it — snapshot-less branches (G2: pre-feature history,
|
||||
* imports, never-generated branches). A present snapshot stays authoritative;
|
||||
* reliable window-switch (G1) detection needs the snapshot to carry its
|
||||
* model/window (deferred to the fidelity follow-up), and the SDK window
|
||||
* (reserve-derived) doesn't equal the client-resolved raw window, so we must
|
||||
* NOT mis-flag a valid snapshot as stale here. Cached + refetched by branch/
|
||||
* endpoint/model/window/revision. */
|
||||
const projectionParams: TContextProjectionRequest | null =
|
||||
!isSubmitting &&
|
||||
branchSnapshot == null &&
|
||||
conversation?.conversationId != null &&
|
||||
conversation.conversationId !== Constants.NEW_CONVO &&
|
||||
branchTotals.tailId != null &&
|
||||
conversation.endpoint != null
|
||||
? {
|
||||
conversationId: conversation.conversationId,
|
||||
messageId: branchTotals.tailId,
|
||||
/** Resolved provider/model (e.g. an agent's actual provider, not the
|
||||
* `agents` endpoint) so the server picks the right tokenizer. */
|
||||
endpoint: limits.endpoint || conversation.endpoint,
|
||||
model: limits.model || conversation.model || undefined,
|
||||
agentId: conversation.agent_id ?? undefined,
|
||||
spec: conversation.spec ?? undefined,
|
||||
maxContextTokens: resolvedMax,
|
||||
/** Content revision so an in-place message edit (same tail id) refetches. */
|
||||
revision: branchTotals.input + branchTotals.output,
|
||||
}
|
||||
: null;
|
||||
const { data: projectionData } = useContextProjectionQuery(projectionParams);
|
||||
const projection = projectionData ?? null;
|
||||
|
||||
/** Branch/total provider usage is index-derived; the in-flight response is
|
||||
* the only live add (the pending holder), counted into both — it sits on the
|
||||
* active branch tail and inside the conversation. The backend prices each
|
||||
|
|
@ -186,41 +242,46 @@ export default function useTokenUsage({
|
|||
snapshot != null &&
|
||||
(isSubmitting || (snapshot.anchorMessageId != null && branchTotals.containsAnchor));
|
||||
|
||||
/** When the live snapshot belongs to another branch, recover this branch's
|
||||
* own finalized snapshot (if it was generated this session) by walking the
|
||||
* branch for its deepest stored anchor — keeps the granular rows on switch
|
||||
* instead of dropping to coarse totals. */
|
||||
let activeSnapshot: ContextSnapshot | null = currentActive ? snapshot : null;
|
||||
if (activeSnapshot == null && !isSubmitting && snapshotsByAnchor.size > 0) {
|
||||
const anchor = findBranchSnapshotAnchor(
|
||||
conversationKey,
|
||||
branchTotals.tailId,
|
||||
snapshotsByAnchor,
|
||||
);
|
||||
activeSnapshot = anchor != null ? (snapshotsByAnchor.get(anchor) ?? null) : null;
|
||||
/** Precedence: live/active snapshot → persisted branch snapshot → server
|
||||
* projection (snapshot-less branches, G2) → per-message estimate. The first
|
||||
* two preserve the pre-projection behavior exactly; the projection only
|
||||
* slots in ahead of the estimate when no snapshot exists. Snapshot and
|
||||
* projection share the render-relevant fields, so they render uniformly. */
|
||||
let effective: ContextSnapshot | TContextUsageEvent | null = null;
|
||||
/** A server projection is the SDK's windowing but, in this first cut, omits
|
||||
* instruction/tool overhead — so it's surfaced as an ESTIMATE (a better one
|
||||
* than sumBranch), never a false-authoritative number. Real snapshots stay
|
||||
* authoritative. */
|
||||
let projected = false;
|
||||
if (currentActive) {
|
||||
effective = snapshot;
|
||||
} else if (branchSnapshot != null) {
|
||||
effective = branchSnapshot;
|
||||
} else if (projection != null) {
|
||||
effective = projection;
|
||||
projected = true;
|
||||
}
|
||||
|
||||
if (activeSnapshot != null) {
|
||||
const breakdown = activeSnapshot.breakdown;
|
||||
const maxTokens = activeSnapshot.contextBudget ?? breakdown.maxContextTokens;
|
||||
const instructionTokens =
|
||||
activeSnapshot.effectiveInstructionTokens ?? breakdown.instructionTokens;
|
||||
if (effective != null) {
|
||||
const breakdown = effective.breakdown;
|
||||
const maxTokens = effective.contextBudget ?? breakdown.maxContextTokens;
|
||||
const instructionTokens = effective.effectiveInstructionTokens ?? breakdown.instructionTokens;
|
||||
const baseUsed =
|
||||
activeSnapshot.remainingContextTokens != null
|
||||
? maxTokens - activeSnapshot.remainingContextTokens
|
||||
effective.remainingContextTokens != null
|
||||
? maxTokens - effective.remainingContextTokens
|
||||
: instructionTokens + breakdown.messageTokens;
|
||||
/** The snapshot is pre-invoke: in-flight output rides on `liveTokens`
|
||||
* (0 unless streaming this branch), the last call's finalized output on
|
||||
* `completedOutputTokens`. */
|
||||
/** The snapshot/projection is pre-invoke: in-flight output rides on
|
||||
* `liveTokens` (0 unless streaming this branch), the last call's finalized
|
||||
* output on `completedOutputTokens` (absent on a projection → 0). */
|
||||
const usedTokens =
|
||||
Math.max(0, baseUsed) + liveTokens + (activeSnapshot.completedOutputTokens ?? 0);
|
||||
Math.max(0, baseUsed) + liveTokens + (effective.completedOutputTokens ?? 0);
|
||||
return {
|
||||
usedTokens,
|
||||
maxTokens,
|
||||
percent: maxTokens > 0 ? Math.min((usedTokens / maxTokens) * 100, 100) : 0,
|
||||
isEstimate: false,
|
||||
snapshot: activeSnapshot,
|
||||
snapshotActive: true,
|
||||
isEstimate: projected,
|
||||
snapshot: projected ? null : (effective as ContextSnapshot),
|
||||
snapshotActive: !projected,
|
||||
branchTotals,
|
||||
branchUsage,
|
||||
totalUsage,
|
||||
|
|
@ -266,7 +327,7 @@ export default function useTokenUsage({
|
|||
hasUsage,
|
||||
liveTokens,
|
||||
limits,
|
||||
snapshotsByAnchor,
|
||||
conversationKey,
|
||||
branchSnapshot,
|
||||
projection,
|
||||
]);
|
||||
}
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue