mirror of
https://github.com/danny-avila/LibreChat.git
synced 2026-08-27 12:13:30 +00:00
* 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan Adds a Playwright benchmark that streams one long, unsplit <think> block (18k chars — 4x the legacy SplitStreamHandler blockThreshold) plus 6k chars of markdown through the real mock-model agents pipeline, with react-scan injected to tally per-component renders. It verifies the legacy content-part splitting (removed in #10533) is not needed for rendering performance: - The whole reasoning section lands in ONE think part (a single Thoughts toggle) — nothing re-splits it anywhere in the pipeline. - rAF coalescing bounds the think box to ~1 render per 43 streamed chunks (122 renders / 5,290 chunks). - MarkdownBlock renders stay O(blocks + flushes) (153 renders / 2,092 text chunks), not O(blocks x tokens). - Long tasks during the 13.4s stream: one 96ms task; total render time 885ms. - Typing after the long transcript leaves transcript components quiet (<=2 renders across 40 keystrokes). Runs against the vite dev server (prod minification strips displayName assignments, which react-scan needs for naming). react-scan itself is not a repo dependency: install with `npm i --no-save react-scan` or point REACT_SCAN_PATH at its auto.global.js bundle. Also fixes the mock e2e stack for local runs: a developer .env with CHECK_BALANCE=true leaked through neutralizeCredentialEnv (not credential-shaped) and made every streaming mock spec fail with a token_balance violation, since the fresh e2e user has no balance record. vanillaOverrides now pins CHECK_BALANCE=false. * 🩹 fix: Address Codex Review — Payload Integrity, Frame Bounds, Proxy Port - Assert the full 18k-char reasoning payload survives the pipeline: expand the Thoughts toggle and compare rendered think text against the source (whitespace-normalized), instead of only counting toggles. - Derive render bounds from elapsed frames (60fps + headroom) rather than chunk counts, so the coalescing assertion stays meaningful regardless of how many chunks stream before resetPerf; apply the same bound to MarkdownBlock. - Tighten main-thread budgets: worst long task < 250ms and long-task total < 10% of stream wall time (baseline: one 51-96ms task per run). - Pass BACKEND_PORT derived from the configured E2E base URL to the vite dev server so its /api proxy follows a non-default app-server port. * 🧭 fix: Address Codex Round 2 — Typed Global, Drained Observer, Full-Payload Checks - Declare window.__PERF__ via global Window augmentation; drop the as-unknown-as double casts from both perf helpers. - Retain the longtask PerformanceObserver and drain takeRecords() before every snapshot/reset so stalls landing near the final render are counted. - Start the wall clock at the same instant as the tally reset so frame bounds and long-task percentages divide by exactly the measured interval. - Verify the complete markdown body: every generated section heading (exact-match), the exact list-item and table counts, and the generated code block — END_MARKER alone only proved the suffix rendered. - Require positive ThinkingContent/MarkdownBlock render counts so a renamed component or dropped instrumentation cannot void the upper bounds. - Cap cumulative render time at 25% of stream wall time to catch sustained sub-50ms work that never surfaces as a long task. * 🧷 fix: Address Codex Round 3 — Page Clock, Completion Wait, Exact Payload Checks - Measure the stream interval on the page's own clock: reset stamps the start, the snapshot evaluation reads the end, so bounds divide by exactly the tallied window including work between marker paint and snapshot. - Wait for the Stop generating button to hide before snapshotting, so generation finalization (usage chunk, terminal events, save re-render) is inside the measured interval. - Compare the rendered think text exactly (edges trimmed only) — internal paragraph breaks are user-visible under whitespace-pre-wrap and must survive verbatim. - Verify the markdown prose, not just structure: per-section doubled-sentence paragraph and both list-item texts, exact table count with cell values, and both generated code lines. - Derive the vite proxy port via getE2EServerAddress() so implicit ports in E2E_BASE_URL (default 80/443) agree between the app server and the proxy. - Pin react-scan@0.5.7 in the README — thresholds are calibrated against its instrumentation semantics. * 🪛 fix: Address Codex Round 4 — Pre-Send Reset, Count Every Code Block - Reset the tally immediately BEFORE triggering the send: with a 1ms chunk delay, the earliest deltas can render between the response headers resolving and a post-send evaluation, which the old order erased from the measurement. - Assert Math.floor(sectionCount / 3) occurrences of both generated code lines via code-element locators instead of .first(), so dropped later code blocks can no longer pass the payload check. * 🎛️ fix: Address Codex Round 5 — First-Render Clock, Expanded Box, Typing Budget - Stamp the wall clock at the FIRST render after each reset (inside onRender) so idle request-setup time between reset and stream start never pads the frame, long-task, or render-time denominators. - Seed showThinking=true so the reasoning box streams EXPANDED — the heavier live-layout path — and drop the post-hoc expand click. - Bound the typing phase itself: worst long task < 150ms and cumulative render time < 25% of the typed interval, so input lag without transcript re-renders still fails. - Derive the vite dev server host from getE2EServerAddress() alongside the port, so a non-localhost E2E base URL keeps the app server, listen host, and /api proxy in agreement. * 🧿 fix: Address Codex Round 6 — Stream-Anchored Clock, IPv6 Proxy, Rate-Free Bounds - Anchor the stream clock to the first ThinkingContent render — the payload opens with reasoning, so that is the first assistant-content paint — keeping composer renders and idle request setup out of the denominators. - Bracket IPv6 HOST values when building the vite /api proxy target in client/vite.config.ts; unbracketed ::1 produced an unparseable URL. - Add an absolute cumulative long-task budget (<300ms) to the typing phase so repeated sub-threshold stalls cannot evade the worst-case check or dilute the ratio via inflated elapsed time. - Add chunk-relative companion bounds (renders < chunks/4) for both ThinkingContent and MarkdownBlock, and hard-pin MOCK_LLM_CHUNK_DELAY_MS=1, so a slower stream can no longer loosen the coalescing assertions.
62 lines
2.3 KiB
TypeScript
62 lines
2.3 KiB
TypeScript
import { defineConfig } from '@playwright/test';
|
|
import path from 'node:path';
|
|
import mockConfig from './playwright.config.mock';
|
|
import { buildReasoningPayload } from './benchmarks-reasoning/payload';
|
|
import { getE2EServerAddress } from './setup/env';
|
|
|
|
/**
|
|
* Reasoning-stream perf benchmark config.
|
|
*
|
|
* Streams one long `<think>…</think>` + markdown reply through the real
|
|
* mock-model agents pipeline so react-scan can verify that rendering a single
|
|
* large, unsplit reasoning part (plus long markdown text) stays
|
|
* render-bounded — confirming the legacy content-part splitting is not needed.
|
|
*
|
|
* Tests run against the vite dev server (port 3090, proxying /api to the mock
|
|
* backend on 3080): the dev build keeps component names, which react-scan
|
|
* needs for per-component tallies — the production minifier (oxc) strips
|
|
* `displayName` assignments.
|
|
*/
|
|
process.env.MOCK_LLM_REPLY = buildReasoningPayload();
|
|
/** Pinned, not defaulted: the render-count thresholds are calibrated against
|
|
* this delivery rate, and a slower stream would loosen the frame-derived
|
|
* bounds. */
|
|
process.env.MOCK_LLM_CHUNK_DELAY_MS = '1';
|
|
|
|
const rootPath = path.resolve(__dirname, '..');
|
|
const { host: backendHost, port: backendPort } = getE2EServerAddress();
|
|
const devHost = backendHost.includes(':') ? `[${backendHost}]` : backendHost;
|
|
const DEV_SERVER_URL = `http://${devHost}:3090`;
|
|
|
|
const appServer = Array.isArray(mockConfig.webServer)
|
|
? mockConfig.webServer[0]
|
|
: mockConfig.webServer;
|
|
|
|
export default defineConfig({
|
|
...mockConfig,
|
|
testDir: 'benchmarks-reasoning',
|
|
outputDir: 'benchmarks-reasoning/.test-results',
|
|
timeout: 10 * 60 * 1000,
|
|
retries: 0,
|
|
reporter: [['line']],
|
|
use: {
|
|
...mockConfig.use,
|
|
baseURL: DEV_SERVER_URL,
|
|
},
|
|
webServer: [
|
|
...(appServer ? [appServer] : []),
|
|
{
|
|
command: 'npm run frontend:dev',
|
|
cwd: rootPath,
|
|
/** The mock env exports PORT for the backend; vite reads PORT for its
|
|
* own listen port, so pin the dev server back to 3090 and point its
|
|
* /api proxy (HOST + BACKEND_PORT in client/vite.config.ts) at the
|
|
* host/port the app server actually binds per the E2E base URL. */
|
|
env: { ...process.env, PORT: '3090', HOST: backendHost, BACKEND_PORT: backendPort },
|
|
url: DEV_SERVER_URL,
|
|
stdout: 'pipe',
|
|
timeout: 180_000,
|
|
reuseExistingServer: false,
|
|
},
|
|
],
|
|
});
|