LibreChat/e2e/benchmarks-reasoning/payload.ts
Danny Avila 4f5808d9ae
🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan (#14494)
* 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan

Adds a Playwright benchmark that streams one long, unsplit <think> block
(18k chars — 4x the legacy SplitStreamHandler blockThreshold) plus 6k chars
of markdown through the real mock-model agents pipeline, with react-scan
injected to tally per-component renders. It verifies the legacy content-part
splitting (removed in #10533) is not needed for rendering performance:

- The whole reasoning section lands in ONE think part (a single Thoughts
  toggle) — nothing re-splits it anywhere in the pipeline.
- rAF coalescing bounds the think box to ~1 render per 43 streamed chunks
  (122 renders / 5,290 chunks).
- MarkdownBlock renders stay O(blocks + flushes) (153 renders / 2,092 text
  chunks), not O(blocks x tokens).
- Long tasks during the 13.4s stream: one 96ms task; total render time 885ms.
- Typing after the long transcript leaves transcript components quiet
  (<=2 renders across 40 keystrokes).

Runs against the vite dev server (prod minification strips displayName
assignments, which react-scan needs for naming). react-scan itself is not a
repo dependency: install with `npm i --no-save react-scan` or point
REACT_SCAN_PATH at its auto.global.js bundle.

Also fixes the mock e2e stack for local runs: a developer .env with
CHECK_BALANCE=true leaked through neutralizeCredentialEnv (not
credential-shaped) and made every streaming mock spec fail with a
token_balance violation, since the fresh e2e user has no balance record.
vanillaOverrides now pins CHECK_BALANCE=false.

* 🩹 fix: Address Codex Review — Payload Integrity, Frame Bounds, Proxy Port

- Assert the full 18k-char reasoning payload survives the pipeline: expand
  the Thoughts toggle and compare rendered think text against the source
  (whitespace-normalized), instead of only counting toggles.
- Derive render bounds from elapsed frames (60fps + headroom) rather than
  chunk counts, so the coalescing assertion stays meaningful regardless of
  how many chunks stream before resetPerf; apply the same bound to
  MarkdownBlock.
- Tighten main-thread budgets: worst long task < 250ms and long-task total
  < 10% of stream wall time (baseline: one 51-96ms task per run).
- Pass BACKEND_PORT derived from the configured E2E base URL to the vite dev
  server so its /api proxy follows a non-default app-server port.

* 🧭 fix: Address Codex Round 2 — Typed Global, Drained Observer, Full-Payload Checks

- Declare window.__PERF__ via global Window augmentation; drop the
  as-unknown-as double casts from both perf helpers.
- Retain the longtask PerformanceObserver and drain takeRecords() before
  every snapshot/reset so stalls landing near the final render are counted.
- Start the wall clock at the same instant as the tally reset so frame
  bounds and long-task percentages divide by exactly the measured interval.
- Verify the complete markdown body: every generated section heading
  (exact-match), the exact list-item and table counts, and the generated
  code block — END_MARKER alone only proved the suffix rendered.
- Require positive ThinkingContent/MarkdownBlock render counts so a renamed
  component or dropped instrumentation cannot void the upper bounds.
- Cap cumulative render time at 25% of stream wall time to catch sustained
  sub-50ms work that never surfaces as a long task.

* 🧷 fix: Address Codex Round 3 — Page Clock, Completion Wait, Exact Payload Checks

- Measure the stream interval on the page's own clock: reset stamps the
  start, the snapshot evaluation reads the end, so bounds divide by exactly
  the tallied window including work between marker paint and snapshot.
- Wait for the Stop generating button to hide before snapshotting, so
  generation finalization (usage chunk, terminal events, save re-render) is
  inside the measured interval.
- Compare the rendered think text exactly (edges trimmed only) — internal
  paragraph breaks are user-visible under whitespace-pre-wrap and must
  survive verbatim.
- Verify the markdown prose, not just structure: per-section doubled-sentence
  paragraph and both list-item texts, exact table count with cell values, and
  both generated code lines.
- Derive the vite proxy port via getE2EServerAddress() so implicit ports in
  E2E_BASE_URL (default 80/443) agree between the app server and the proxy.
- Pin react-scan@0.5.7 in the README — thresholds are calibrated against its
  instrumentation semantics.

* 🪛 fix: Address Codex Round 4 — Pre-Send Reset, Count Every Code Block

- Reset the tally immediately BEFORE triggering the send: with a 1ms chunk
  delay, the earliest deltas can render between the response headers
  resolving and a post-send evaluation, which the old order erased from the
  measurement.
- Assert Math.floor(sectionCount / 3) occurrences of both generated code
  lines via code-element locators instead of .first(), so dropped later
  code blocks can no longer pass the payload check.

* 🎛️ fix: Address Codex Round 5 — First-Render Clock, Expanded Box, Typing Budget

- Stamp the wall clock at the FIRST render after each reset (inside
  onRender) so idle request-setup time between reset and stream start never
  pads the frame, long-task, or render-time denominators.
- Seed showThinking=true so the reasoning box streams EXPANDED — the heavier
  live-layout path — and drop the post-hoc expand click.
- Bound the typing phase itself: worst long task < 150ms and cumulative
  render time < 25% of the typed interval, so input lag without transcript
  re-renders still fails.
- Derive the vite dev server host from getE2EServerAddress() alongside the
  port, so a non-localhost E2E base URL keeps the app server, listen host,
  and /api proxy in agreement.

* 🧿 fix: Address Codex Round 6 — Stream-Anchored Clock, IPv6 Proxy, Rate-Free Bounds

- Anchor the stream clock to the first ThinkingContent render — the payload
  opens with reasoning, so that is the first assistant-content paint —
  keeping composer renders and idle request setup out of the denominators.
- Bracket IPv6 HOST values when building the vite /api proxy target in
  client/vite.config.ts; unbracketed ::1 produced an unparseable URL.
- Add an absolute cumulative long-task budget (<300ms) to the typing phase
  so repeated sub-threshold stalls cannot evade the worst-case check or
  dilute the ratio via inflated elapsed time.
- Add chunk-relative companion bounds (renders < chunks/4) for both
  ThinkingContent and MarkdownBlock, and hard-pin MOCK_LLM_CHUNK_DELAY_MS=1,
  so a slower stream can no longer loosen the coalescing assertions.
2026-07-28 22:18:24 -04:00

57 lines
1.9 KiB
TypeScript

/**
* Deterministic long-form reply payload for the reasoning-stream perf benchmark.
*
* The reasoning section intentionally far exceeds the legacy 4500-char
* `blockThreshold` the old SplitStreamHandler used, so streaming it as ONE
* contiguous think part exercises exactly the case the legacy splitting
* existed to protect against.
*/
export const SENTENCE =
'The quarterly analytics review shows sustained growth across every referral channel, with notable spikes on launch days. ';
export const THINK_TARGET_CHARS = 18000;
export const TEXT_TARGET_CHARS = 6000;
export const END_MARKER = 'END-OF-BENCH-STREAM';
export function buildThinkSection(): string {
let think = '';
let i = 0;
while (think.length < THINK_TARGET_CHARS) {
i += 1;
think += `Step ${i}: ${SENTENCE}`;
if (i % 6 === 0) {
think += '\n\n';
}
}
return think;
}
export function buildTextSection(): string {
const codeBlock =
'```ts\nexport function estimate(total: number, rate: number): number {\n return Math.round(total * rate);\n}\n```\n\n';
const table =
'| Month | Visitors | Growth |\n| --- | --- | --- |\n| Jan | 120000 | 4% |\n| Feb | 135500 | 12% |\n| Mar | 151200 | 11% |\n\n';
let text = '# Milestone Report\n\n';
let j = 0;
while (text.length < TEXT_TARGET_CHARS) {
j += 1;
text += `## Section ${j}\n\n${SENTENCE}${SENTENCE}\n\n- Point one for section ${j}\n- Point two for section ${j}\n\n`;
if (j % 3 === 0) {
text += codeBlock;
}
if (j % 4 === 0) {
text += table;
}
}
text += `\n\n${END_MARKER}\n`;
return text;
}
export function buildReasoningPayload(): string {
return `<think>${buildThinkSection()}</think>\n\n${buildTextSection()}`;
}
/** Mirrors the mock FakeChatModel's default whitespace split strategy. */
export function countModelChunks(text: string): number {
return text.split(/(?<=\s+)|(?=\s+)/).length;
}