🪙 feat: SDK-Aligned Context-Usage Projection (gauge for window-switch & snapshot-less branches) (#13801)
Some checks are pending
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Waiting to run
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Waiting to run
GitNexus Index / index (push) Waiting to run
GitNexus Index / post-index (push) Blocked by required conditions

* 🪙 feat: Context-usage projection — data-provider + client wiring

Consumer side of the SDK-aligned context projection (agents
`projectAgentContextUsage`). Adds the `/api/endpoints/context-projection`
data-provider plumbing (endpoint, service, query key, `TContextProjectionRequest`)
and a `useContextProjectionQuery` gated to fire only when no fresh snapshot
covers the viewed branch.

Wires `useTokenUsage` precedence to: live snapshot → fresh persisted snapshot
(window matches the resolved one) → server projection → per-message estimate.
A model/window switch marks the baked snapshot stale (its `maxContextTokens`
no longer matches) and falls to the projection — closing the gauge's
window-switch (G1) and snapshot-less-branch (G2) gaps. Snapshot and projection
share the render-relevant fields, so they render uniformly.

Backend endpoint + agents version bump land in follow-up commits. Includes the
design spec (CONTEXT_PROJECTION_SPEC.md).

* 🪙 feat: Context-projection backend endpoint

POST /api/endpoints/context-projection → resolveContextProjection (packages/api):
reconstructs the viewed branch (parent-chain walk from messageId), resolves the
agent config (instructions/provider/model/maxContextTokens), reuses LibreChat's
stored per-message tokenCounts as the index map (no re-tokenizing), and calls
the agents SDK projectAgentContextUsage — no model call. Thin controller injects
db.getMessages/db.getAgent; route mirrors /token-config.

First cut targets message-windowing accuracy; tool-schema tokens are deferred to
a follow-up that reuses the full initializeAgent path.

* 🩹 fix: Codex review on context projection (G1 guard, IDOR, recount, summary)

- Guard `currentActive` against a stale window: a model/window switch on the
  current branch left the live snapshot outranking the projection (G1 didn't
  fire). Now defers to the projection unless streaming or the window matches.
- Scope branch lookups to the authenticated user (`getMessages` filter +
  injected `userId`) — was loading any conversation by id (IDOR).
- Recount messages with no stored `tokenCount` via the tokenizer instead of
  charging 0, so snapshot-less/imported histories don't under-report.
- Fall back (null) for already-summarized branches rather than projecting from
  the full raw parent chain (the next call would send summary + tail); the
  client's summary-baseline-aware estimate handles them until a follow-up
  replays the summary boundary.

* 🩹 fix: Codex round 2 — drop agent load, summary marker, edit-invalidation

- Stop loading agent/model-spec config server-side (closes the agent-access
  IDOR and the spec-prompt special-casing). Provider/model/window now come from
  the client-resolved request (`limits.endpoint`/model — the agent's real
  provider, not the `agents` endpoint, so the tokenizer is right). Agent/spec/
  promptPrefix instructions are uniformly deferred to the full-fidelity follow-up.
- Detect summarized branches via the live path's `metadata.summaryUsedTokens`
  marker (was the wrong `summaryTokenCount` field) and fall back to the
  summary-aware estimate.
- Invalidate the projection query on in-place message edits via a branch
  content `revision` in the cache key (the tail id is unchanged on edit).

Deferred (valid, not a regression): same-window endpoint/model switch keeps a
window-matched snapshot — needs endpoint/model persisted on the snapshot, which
lands with the fidelity follow-up. Smoke-tested: fits / prunes / summarized→null
/ no-window→null.

* 🛡️ fix: make context projection strictly additive (no-regression)

Revert the G1 window-match guard on the live/branch snapshot. When no explicit
maxContextTokens is set (the common default), the SDK's snapshot window is
reserve-derived (~0.9·(modelContext − maxOutputTokens)) while useTokenLimits
resolves the raw model context — so `snapshot.maxContextTokens === resolvedMax`
is false for the SAME model, and the guard would wrongly drop a valid
current-branch snapshot to projection/estimate post-stream (a regression in the
default case, per initialize.ts:1240-1243).

The projection now activates ONLY for snapshot-less branches (G2): the
precedence is live snapshot → persisted branch snapshot → projection → estimate,
where the first two are byte-for-byte the prior behavior and the projection just
slots ahead of the estimate. Window/model-switch (G1) detection needs the
snapshot to carry its model/window and defers to the fidelity follow-up.

* 🩹 fix: surface projections as estimates, not authoritative snapshots

A first-cut projection carries the SDK's windowing but omits instruction/tool
overhead, so rendering it as `isEstimate: false` showed a confident under-count
for snapshot-less branches. Mark projection-sourced views `isEstimate: true` +
`snapshotActive: false` (and drop the snapshot field) so they present as a
better estimate than sumBranch — improved used/window number, estimate framing,
no misleading granular breakdown with ~0 tools. Real snapshots stay
authoritative. (Codex round 3, projection.ts:139.)

* 🧹 chore: drop CONTEXT_PROJECTION_SPEC.md from the PR

* 🎨 style: fix import-sort order in projection.ts (CI sort-imports check)

* 🔧 chore: update @librechat/agents dependency to version 3.2.36 in package-lock.json and related package.json files

* chore: npm audit fix

* 🎨 style: fix import-sort order in data-service.ts (CI sort-imports check)

* 🩹 fix: drop dead calibrationRatio in projectionParams (tsc never error)

Inside the ternary, branchSnapshot is narrowed to null (the gate is
), so  accessed a
property on  (frontend typecheck failure). It was also dead — there is
never a snapshot to seed from in this branch — so just remove it.

* Revert "chore: npm audit fix"

This reverts commit 4cdb862d0c.
This commit is contained in:
Danny Avila 2026-06-16 17:54:13 -04:00 committed by GitHub
parent c820dfb9a0
commit fdc7e64bb7
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
13 changed files with 348 additions and 36 deletions

View file

@ -46,7 +46,7 @@
"@azure/storage-blob": "^12.30.0",
"@google/genai": "^2.8.0",
"@keyv/redis": "^4.3.3",
"@librechat/agents": "^3.2.35",
"@librechat/agents": "^3.2.36",
"@librechat/api": "*",
"@librechat/data-schemas": "*",
"@microsoft/microsoft-graph-client": "^3.0.7",

View file

@ -0,0 +1,31 @@
const { logger } = require('@librechat/data-schemas');
const { resolveContextProjection } = require('@librechat/api');
const db = require('~/models');
/**
* Returns a server-side context-usage projection for the viewed branch + config
* (agents SDK, no model call) powers the gauge for snapshot-less branches and
* after a model/window switch. Resolution lives in `@librechat/api`; this
* controller only injects request-scoped model accessors.
* @param {ServerRequest} req
* @param {ServerResponse} res
*/
async function contextProjectionController(req, res) {
try {
const params = req.body ?? {};
if (!params.conversationId || !params.messageId) {
res.json(null);
return;
}
const projection = await resolveContextProjection(
{ userId: req.user?.id, getMessages: db.getMessages },
params,
);
res.json(projection ?? null);
} catch (error) {
logger.error('[contextProjectionController]', error);
res.status(500).json({ error: 'Failed to resolve context projection' });
}
}
module.exports = contextProjectionController;

View file

@ -3,10 +3,12 @@ const requireJwtAuth = require('~/server/middleware/requireJwtAuth');
const configMiddleware = require('~/server/middleware/config/app');
const endpointController = require('~/server/controllers/EndpointController');
const tokenConfigController = require('~/server/controllers/TokenConfigController');
const contextProjectionController = require('~/server/controllers/ContextProjectionController');
const router = express.Router();
/** Auth required for role/tenant-scoped endpoint config resolution. */
router.get('/', requireJwtAuth, endpointController);
router.get('/token-config', requireJwtAuth, configMiddleware, tokenConfigController);
router.post('/context-projection', requireJwtAuth, configMiddleware, contextProjectionController);
module.exports = router;

View file

@ -41,6 +41,45 @@ export const useTokenConfigQuery = (
});
};
/**
* Server-side context-usage projection for the viewed branch + resolved config
* (agents SDK, no model call). Keyed on the fields that change what the next
* call would send branch tail, endpoint/model/agent, window so a branch or
* model/window switch refetches. Disabled until a branch tail is known.
*/
export const useContextProjectionQuery = (
params: t.TContextProjectionRequest | null,
config?: UseQueryOptions<t.TContextUsageEvent | null>,
): QueryObserverResult<t.TContextUsageEvent | null> => {
const queriesEnabled = useRecoilValue<boolean>(store.queriesEnabled);
return useQuery<t.TContextUsageEvent | null>(
[
QueryKeys.contextProjection,
params?.conversationId,
params?.messageId,
params?.endpoint,
params?.model,
params?.agentId,
params?.maxContextTokens,
params?.revision,
],
() => dataService.getContextProjection(params as t.TContextProjectionRequest),
{
staleTime: Infinity,
refetchOnWindowFocus: false,
refetchOnReconnect: false,
refetchOnMount: false,
...config,
enabled:
(config?.enabled ?? true) === true &&
queriesEnabled &&
params != null &&
(params.conversationId?.length ?? 0) > 0 &&
(params.messageId?.length ?? 0) > 0,
},
);
};
/**
* Auth-aware query key so unauthenticated (login page) and authenticated
* (chat page) configs are cached independently, preventing stale

View file

@ -2,7 +2,13 @@ import { useEffect, useMemo, useRef } from 'react';
import { useAtomValue, useSetAtom } from 'jotai';
import { useQueryClient } from '@tanstack/react-query';
import { Constants, QueryKeys } from 'librechat-data-provider';
import type { TMessage, TConversation, TModelTokenomics } from 'librechat-data-provider';
import type {
TMessage,
TConversation,
TModelTokenomics,
TContextUsageEvent,
TContextProjectionRequest,
} from 'librechat-data-provider';
import type { BranchTotals, BranchUsage } from '~/utils/tokens';
import type { ContextSnapshot } from '~/store/usage';
import {
@ -24,6 +30,7 @@ import {
findBranchSnapshotAnchor,
} from '~/utils';
import { useLatestMessageId } from '~/hooks/Messages/useLatestMessage';
import { useContextProjectionQuery } from '~/data-provider';
import useTokenLimits from './useTokenLimits';
export interface TokenUsageParams {
@ -79,6 +86,55 @@ export default function useTokenUsage({
const setTotalUsage = useSetAtom(totalUsageFamily(conversationKey));
const limits = useTokenLimits(conversation);
/** Deepest persisted/live snapshot on the viewed branch (present only for
* turns generated with the feature on). Gates the projection fetch and is a
* render source. */
const branchSnapshot = useMemo(() => {
if (snapshotsByAnchor.size === 0) {
return null;
}
const anchor = findBranchSnapshotAnchor(
conversationKey,
branchTotals.tailId,
snapshotsByAnchor,
);
return anchor != null ? (snapshotsByAnchor.get(anchor) ?? null) : null;
}, [conversationKey, branchTotals.tailId, snapshotsByAnchor]);
const resolvedMax = limits.maxContextTokens;
/** Project the branch (agents SDK, no model call) ONLY when no persisted/live
* snapshot covers it snapshot-less branches (G2: pre-feature history,
* imports, never-generated branches). A present snapshot stays authoritative;
* reliable window-switch (G1) detection needs the snapshot to carry its
* model/window (deferred to the fidelity follow-up), and the SDK window
* (reserve-derived) doesn't equal the client-resolved raw window, so we must
* NOT mis-flag a valid snapshot as stale here. Cached + refetched by branch/
* endpoint/model/window/revision. */
const projectionParams: TContextProjectionRequest | null =
!isSubmitting &&
branchSnapshot == null &&
conversation?.conversationId != null &&
conversation.conversationId !== Constants.NEW_CONVO &&
branchTotals.tailId != null &&
conversation.endpoint != null
? {
conversationId: conversation.conversationId,
messageId: branchTotals.tailId,
/** Resolved provider/model (e.g. an agent's actual provider, not the
* `agents` endpoint) so the server picks the right tokenizer. */
endpoint: limits.endpoint || conversation.endpoint,
model: limits.model || conversation.model || undefined,
agentId: conversation.agent_id ?? undefined,
spec: conversation.spec ?? undefined,
maxContextTokens: resolvedMax,
/** Content revision so an in-place message edit (same tail id) refetches. */
revision: branchTotals.input + branchTotals.output,
}
: null;
const { data: projectionData } = useContextProjectionQuery(projectionParams);
const projection = projectionData ?? null;
/** Branch/total provider usage is index-derived; the in-flight response is
* the only live add (the pending holder), counted into both it sits on the
* active branch tail and inside the conversation. The backend prices each
@ -186,41 +242,46 @@ export default function useTokenUsage({
snapshot != null &&
(isSubmitting || (snapshot.anchorMessageId != null && branchTotals.containsAnchor));
/** When the live snapshot belongs to another branch, recover this branch's
* own finalized snapshot (if it was generated this session) by walking the
* branch for its deepest stored anchor keeps the granular rows on switch
* instead of dropping to coarse totals. */
let activeSnapshot: ContextSnapshot | null = currentActive ? snapshot : null;
if (activeSnapshot == null && !isSubmitting && snapshotsByAnchor.size > 0) {
const anchor = findBranchSnapshotAnchor(
conversationKey,
branchTotals.tailId,
snapshotsByAnchor,
);
activeSnapshot = anchor != null ? (snapshotsByAnchor.get(anchor) ?? null) : null;
/** Precedence: live/active snapshot persisted branch snapshot server
* projection (snapshot-less branches, G2) per-message estimate. The first
* two preserve the pre-projection behavior exactly; the projection only
* slots in ahead of the estimate when no snapshot exists. Snapshot and
* projection share the render-relevant fields, so they render uniformly. */
let effective: ContextSnapshot | TContextUsageEvent | null = null;
/** A server projection is the SDK's windowing but, in this first cut, omits
* instruction/tool overhead so it's surfaced as an ESTIMATE (a better one
* than sumBranch), never a false-authoritative number. Real snapshots stay
* authoritative. */
let projected = false;
if (currentActive) {
effective = snapshot;
} else if (branchSnapshot != null) {
effective = branchSnapshot;
} else if (projection != null) {
effective = projection;
projected = true;
}
if (activeSnapshot != null) {
const breakdown = activeSnapshot.breakdown;
const maxTokens = activeSnapshot.contextBudget ?? breakdown.maxContextTokens;
const instructionTokens =
activeSnapshot.effectiveInstructionTokens ?? breakdown.instructionTokens;
if (effective != null) {
const breakdown = effective.breakdown;
const maxTokens = effective.contextBudget ?? breakdown.maxContextTokens;
const instructionTokens = effective.effectiveInstructionTokens ?? breakdown.instructionTokens;
const baseUsed =
activeSnapshot.remainingContextTokens != null
? maxTokens - activeSnapshot.remainingContextTokens
effective.remainingContextTokens != null
? maxTokens - effective.remainingContextTokens
: instructionTokens + breakdown.messageTokens;
/** The snapshot is pre-invoke: in-flight output rides on `liveTokens`
* (0 unless streaming this branch), the last call's finalized output on
* `completedOutputTokens`. */
/** The snapshot/projection is pre-invoke: in-flight output rides on
* `liveTokens` (0 unless streaming this branch), the last call's finalized
* output on `completedOutputTokens` (absent on a projection 0). */
const usedTokens =
Math.max(0, baseUsed) + liveTokens + (activeSnapshot.completedOutputTokens ?? 0);
Math.max(0, baseUsed) + liveTokens + (effective.completedOutputTokens ?? 0);
return {
usedTokens,
maxTokens,
percent: maxTokens > 0 ? Math.min((usedTokens / maxTokens) * 100, 100) : 0,
isEstimate: false,
snapshot: activeSnapshot,
snapshotActive: true,
isEstimate: projected,
snapshot: projected ? null : (effective as ContextSnapshot),
snapshotActive: !projected,
branchTotals,
branchUsage,
totalUsage,
@ -266,7 +327,7 @@ export default function useTokenUsage({
hasUsage,
liveTokens,
limits,
snapshotsByAnchor,
conversationKey,
branchSnapshot,
projection,
]);
}

10
package-lock.json generated
View file

@ -61,7 +61,7 @@
"@azure/storage-blob": "^12.30.0",
"@google/genai": "^2.8.0",
"@keyv/redis": "^4.3.3",
"@librechat/agents": "^3.2.35",
"@librechat/agents": "^3.2.36",
"@librechat/api": "*",
"@librechat/data-schemas": "*",
"@microsoft/microsoft-graph-client": "^3.0.7",
@ -11685,9 +11685,9 @@
}
},
"node_modules/@librechat/agents": {
"version": "3.2.35",
"resolved": "https://registry.npmjs.org/@librechat/agents/-/agents-3.2.35.tgz",
"integrity": "sha512-k/6e3WOkYSgqZVFtgw5XeNpuA+1+y0vX4u4pqGeX3MM3ERDk729OQLdO1qFbOCxkcP2hCZppMzZ5JExSyKdplw==",
"version": "3.2.36",
"resolved": "https://registry.npmjs.org/@librechat/agents/-/agents-3.2.36.tgz",
"integrity": "sha512-y9Bm5CMACNZh6OfDCUcABK84WUYndAfgO36Bu5EEyLqDm1S6TvoYUwQ/6P3jdsMy+Vj1ESrxAfCdQgy7fxtcjg==",
"license": "MIT",
"dependencies": {
"@anthropic-ai/sdk": "^0.92.0",
@ -44290,7 +44290,7 @@
"@azure/storage-blob": "^12.30.0",
"@google/genai": "^2.8.0",
"@keyv/redis": "^4.3.3",
"@librechat/agents": "^3.2.35",
"@librechat/agents": "^3.2.36",
"@librechat/data-schemas": "*",
"@modelcontextprotocol/sdk": "^1.29.0",
"@opentelemetry/api": "^1.9.0",

View file

@ -113,7 +113,7 @@
"@azure/storage-blob": "^12.30.0",
"@google/genai": "^2.8.0",
"@keyv/redis": "^4.3.3",
"@librechat/agents": "^3.2.35",
"@librechat/agents": "^3.2.36",
"@librechat/data-schemas": "*",
"@modelcontextprotocol/sdk": "^1.29.0",
"@opentelemetry/api": "^1.9.0",

View file

@ -6,4 +6,5 @@ export * from './google';
export * from './models';
export * from './openai';
export * from './pricing';
export * from './projection';
export * from './tokenConfig';

View file

@ -0,0 +1,145 @@
import { HumanMessage, AIMessage } from '@langchain/core/messages';
import { Providers, createTokenCounter, projectAgentContextUsage } from '@librechat/agents';
import type { TContextProjectionRequest, TContextUsageEvent } from 'librechat-data-provider';
import type { BaseMessage } from '@langchain/core/messages';
interface ProjectionMessage {
messageId: string;
parentMessageId?: string | null;
tokenCount?: number;
isCreatedByUser?: boolean;
text?: string;
/** Compaction marker written by the live path (`agents/usage.ts`); its
* presence means the next call sends the summary + tail, not this raw chain. */
metadata?: { summaryUsedTokens?: number };
}
export interface ContextProjectionDeps {
/** Authenticated requester — branch lookups are scoped to this user. */
userId?: string;
getMessages: (
filter: { conversationId: string; user?: string },
select?: string,
) => Promise<ProjectionMessage[]>;
}
/**
* Walks the parent chain from `tailId` to root and returns the branch messages
* oldestnewest. The visited set guards against cycles / self-referential links.
*/
function resolveBranch(messages: ProjectionMessage[], tailId: string): ProjectionMessage[] {
const byId = new Map<string, ProjectionMessage>();
for (const message of messages) {
byId.set(message.messageId, message);
}
const branch: ProjectionMessage[] = [];
const seen = new Set<string>();
let currentId: string | null | undefined = tailId;
while (currentId != null && !seen.has(currentId)) {
const message = byId.get(currentId);
if (message == null) {
break;
}
seen.add(currentId);
branch.push(message);
currentId = message.parentMessageId;
}
return branch.reverse();
}
/** Maps an endpoint/provider string to the agents `Providers` enum. */
function resolveProvider(value?: string): Providers {
if (value == null || value === '') {
return Providers.OPENAI;
}
const lower = value.toLowerCase();
for (const provider of Object.values(Providers)) {
if (provider.toLowerCase() === lower) {
return provider;
}
}
if (lower.includes('anthropic') || lower.includes('claude')) {
return Providers.ANTHROPIC;
}
if (lower.includes('google') || lower.includes('gemini') || lower.includes('vertex')) {
return Providers.GOOGLE;
}
if (lower.includes('bedrock')) {
return Providers.BEDROCK;
}
return Providers.OPENAI;
}
/**
* Server-side context-usage projection: reconstructs the viewed branch and asks
* the agents SDK what the next call's context would be, WITHOUT invoking the
* model. Provider/model/window come from the (client-resolved) request no
* agent or model-spec config is loaded here, so there is no cross-user config
* exposure. Reuses LibreChat's already-calibrated per-message `tokenCount`s (no
* re-tokenizing). Returns null when there is no resolvable context window.
* NOTE: this first cut targets message-windowing accuracy instruction and
* tool-schema tokens (agent instructions, `promptPrefix`, model-spec presets,
* tool schemas) are NOT yet included; a follow-up will reuse the full
* `initializeAgent`/send path for exact overhead and proper access control.
*/
export async function resolveContextProjection(
deps: ContextProjectionDeps,
params: TContextProjectionRequest,
): Promise<TContextUsageEvent | null> {
const maxContextTokens = params.maxContextTokens;
if (maxContextTokens == null || maxContextTokens <= 0) {
return null;
}
const stored = await deps.getMessages(
{ conversationId: params.conversationId, user: deps.userId },
'messageId parentMessageId tokenCount isCreatedByUser text metadata',
);
const branch = resolveBranch(stored, params.messageId);
if (branch.length === 0) {
return null;
}
/** A summarized/compacted branch's next call sends the saved summary + the
* post-summary tail, NOT this raw parent chain projecting from the full
* history would prune/count the wrong context and omit the summary. Detect it
* via the live path's `metadata.summaryUsedTokens` marker and fall back (null)
* so the client's summary-baseline-aware estimate handles these branches until
* a follow-up replays the summary boundary. */
if (branch.some((message) => (message.metadata?.summaryUsedTokens ?? 0) > 0)) {
return null;
}
const model = params.model;
const encoding = (model ?? '').toLowerCase().includes('claude') ? 'claude' : 'o200k_base';
const tokenCounter = await createTokenCounter(encoding);
const messages: BaseMessage[] = [];
const indexTokenCountMap: Record<string, number> = {};
for (let i = 0; i < branch.length; i++) {
const message = branch[i];
const text = message.text ?? '';
const lcMessage =
message.isCreatedByUser === true ? new HumanMessage(text) : new AIMessage(text);
messages.push(lcMessage);
/** Recount messages with no stored count (imported / pre-feature) rather
* than charging 0 a real 0 and "unknown" must not collapse, or the
* snapshot-less histories this endpoint targets would under-report. */
indexTokenCountMap[String(i)] =
message.tokenCount != null && message.tokenCount > 0
? message.tokenCount
: tokenCounter(lcMessage);
}
return projectAgentContextUsage({
agent: {
agentId: params.agentId ?? 'projection',
provider: resolveProvider(params.endpoint),
maxContextTokens,
},
messages,
tokenCounter,
indexTokenCountMap,
calibrationRatio: params.calibrationRatio,
});
}

View file

@ -150,6 +150,8 @@ export const aiEndpoints = () => `${BASE_URL}/api/endpoints`;
export const tokenConfig = () => `${BASE_URL}/api/endpoints/token-config`;
export const contextProjection = () => `${BASE_URL}/api/endpoints/context-projection`;
export const models = () => `${BASE_URL}/api/models`;
export const tokenizer = () => `${BASE_URL}/api/tokenizer`;

View file

@ -1,4 +1,5 @@
import type { AxiosResponse } from 'axios';
import type { TContextProjectionRequest, TContextUsageEvent } from './types/runs';
import type { TFileConfig } from './file-config';
import type * as t from './types';
import * as permissions from './accessPermissions';
@ -249,6 +250,12 @@ export const getTokenConfig = (): Promise<t.TTokenConfigMap> => {
return request.get(endpoints.tokenConfig());
};
export const getContextProjection = (
payload: TContextProjectionRequest,
): Promise<TContextUsageEvent | null> => {
return request.post(endpoints.contextProjection(), payload);
};
export const getModels = async (): Promise<t.TModelsConfig> => {
return request.get(endpoints.models());
};

View file

@ -13,6 +13,7 @@ export enum QueryKeys {
balance = 'balance',
endpoints = 'endpoints',
tokenConfig = 'tokenConfig',
contextProjection = 'contextProjection',
presets = 'presets',
searchResults = 'searchResults',
tokenCount = 'tokenCount',

View file

@ -88,6 +88,29 @@ export type TContextUsageEvent = {
completedOutputTokens?: number;
};
/**
* Request payload for a server-side context-usage projection: "what context
* would the next call send for this branch under this config", computed by the
* agents SDK without invoking the model. Powers the gauge in states the live
* snapshot can't cover (page load of a snapshot-less branch, window/model
* switch). `messageId` is the viewed branch's tail; the server walks its parent
* chain.
*/
export type TContextProjectionRequest = {
conversationId: string;
messageId: string;
endpoint: string;
model?: string;
agentId?: string;
spec?: string;
maxContextTokens?: number;
/** Provider-calibrated ratio from a prior snapshot, applied as a static seed. */
calibrationRatio?: number;
/** Client-only cache-bust: a branch content revision so a message edit
* (which keeps the same tail id) refetches. The server ignores it. */
revision?: number;
};
/**
* Per-response usage rollup persisted on `responseMessage.metadata.usage`, in
* display units (input excludes cache; output includes repaired completion).