mirror of
https://github.com/danny-avila/LibreChat.git
synced 2026-08-04 14:57:42 +00:00
* 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan Adds a Playwright benchmark that streams one long, unsplit <think> block (18k chars — 4x the legacy SplitStreamHandler blockThreshold) plus 6k chars of markdown through the real mock-model agents pipeline, with react-scan injected to tally per-component renders. It verifies the legacy content-part splitting (removed in #10533) is not needed for rendering performance: - The whole reasoning section lands in ONE think part (a single Thoughts toggle) — nothing re-splits it anywhere in the pipeline. - rAF coalescing bounds the think box to ~1 render per 43 streamed chunks (122 renders / 5,290 chunks). - MarkdownBlock renders stay O(blocks + flushes) (153 renders / 2,092 text chunks), not O(blocks x tokens). - Long tasks during the 13.4s stream: one 96ms task; total render time 885ms. - Typing after the long transcript leaves transcript components quiet (<=2 renders across 40 keystrokes). Runs against the vite dev server (prod minification strips displayName assignments, which react-scan needs for naming). react-scan itself is not a repo dependency: install with `npm i --no-save react-scan` or point REACT_SCAN_PATH at its auto.global.js bundle. Also fixes the mock e2e stack for local runs: a developer .env with CHECK_BALANCE=true leaked through neutralizeCredentialEnv (not credential-shaped) and made every streaming mock spec fail with a token_balance violation, since the fresh e2e user has no balance record. vanillaOverrides now pins CHECK_BALANCE=false. * 🩹 fix: Address Codex Review — Payload Integrity, Frame Bounds, Proxy Port - Assert the full 18k-char reasoning payload survives the pipeline: expand the Thoughts toggle and compare rendered think text against the source (whitespace-normalized), instead of only counting toggles. - Derive render bounds from elapsed frames (60fps + headroom) rather than chunk counts, so the coalescing assertion stays meaningful regardless of how many chunks stream before resetPerf; apply the same bound to MarkdownBlock. - Tighten main-thread budgets: worst long task < 250ms and long-task total < 10% of stream wall time (baseline: one 51-96ms task per run). - Pass BACKEND_PORT derived from the configured E2E base URL to the vite dev server so its /api proxy follows a non-default app-server port. * 🧭 fix: Address Codex Round 2 — Typed Global, Drained Observer, Full-Payload Checks - Declare window.__PERF__ via global Window augmentation; drop the as-unknown-as double casts from both perf helpers. - Retain the longtask PerformanceObserver and drain takeRecords() before every snapshot/reset so stalls landing near the final render are counted. - Start the wall clock at the same instant as the tally reset so frame bounds and long-task percentages divide by exactly the measured interval. - Verify the complete markdown body: every generated section heading (exact-match), the exact list-item and table counts, and the generated code block — END_MARKER alone only proved the suffix rendered. - Require positive ThinkingContent/MarkdownBlock render counts so a renamed component or dropped instrumentation cannot void the upper bounds. - Cap cumulative render time at 25% of stream wall time to catch sustained sub-50ms work that never surfaces as a long task. * 🧷 fix: Address Codex Round 3 — Page Clock, Completion Wait, Exact Payload Checks - Measure the stream interval on the page's own clock: reset stamps the start, the snapshot evaluation reads the end, so bounds divide by exactly the tallied window including work between marker paint and snapshot. - Wait for the Stop generating button to hide before snapshotting, so generation finalization (usage chunk, terminal events, save re-render) is inside the measured interval. - Compare the rendered think text exactly (edges trimmed only) — internal paragraph breaks are user-visible under whitespace-pre-wrap and must survive verbatim. - Verify the markdown prose, not just structure: per-section doubled-sentence paragraph and both list-item texts, exact table count with cell values, and both generated code lines. - Derive the vite proxy port via getE2EServerAddress() so implicit ports in E2E_BASE_URL (default 80/443) agree between the app server and the proxy. - Pin react-scan@0.5.7 in the README — thresholds are calibrated against its instrumentation semantics. * 🪛 fix: Address Codex Round 4 — Pre-Send Reset, Count Every Code Block - Reset the tally immediately BEFORE triggering the send: with a 1ms chunk delay, the earliest deltas can render between the response headers resolving and a post-send evaluation, which the old order erased from the measurement. - Assert Math.floor(sectionCount / 3) occurrences of both generated code lines via code-element locators instead of .first(), so dropped later code blocks can no longer pass the payload check. * 🎛️ fix: Address Codex Round 5 — First-Render Clock, Expanded Box, Typing Budget - Stamp the wall clock at the FIRST render after each reset (inside onRender) so idle request-setup time between reset and stream start never pads the frame, long-task, or render-time denominators. - Seed showThinking=true so the reasoning box streams EXPANDED — the heavier live-layout path — and drop the post-hoc expand click. - Bound the typing phase itself: worst long task < 150ms and cumulative render time < 25% of the typed interval, so input lag without transcript re-renders still fails. - Derive the vite dev server host from getE2EServerAddress() alongside the port, so a non-localhost E2E base URL keeps the app server, listen host, and /api proxy in agreement. * 🧿 fix: Address Codex Round 6 — Stream-Anchored Clock, IPv6 Proxy, Rate-Free Bounds - Anchor the stream clock to the first ThinkingContent render — the payload opens with reasoning, so that is the first assistant-content paint — keeping composer renders and idle request setup out of the denominators. - Bracket IPv6 HOST values when building the vite /api proxy target in client/vite.config.ts; unbracketed ::1 produced an unparseable URL. - Add an absolute cumulative long-task budget (<300ms) to the typing phase so repeated sub-threshold stalls cannot evade the worst-case check or dilute the ratio via inflated elapsed time. - Add chunk-relative companion bounds (renders < chunks/4) for both ThinkingContent and MarkdownBlock, and hard-pin MOCK_LLM_CHUNK_DELAY_MS=1, so a slower stream can no longer loosen the coalescing assertions. |
||
|---|---|---|
| .. | ||
| payload.ts | ||
| README.md | ||
| reasoning-stream.perf.spec.ts | ||
Reasoning-Stream Perf Benchmark (react-scan)
Verifies that streaming one long, unsplit reasoning block (plus long markdown
text) through the real mock-model agents pipeline stays render-bounded — i.e.
the legacy content-part splitting (SplitStreamHandler / blockThreshold,
removed in #10533) is not needed for rendering performance.
What it measures, via react-scan injected into the page:
- Per-component render counts and render time while a ~18k-char
<think>block and ~6k-char markdown reply stream token by token. - That the whole reasoning section lands in one think part (a single "Thoughts" toggle) — no re-splitting anywhere in the pipeline.
- rAF coalescing: the think box re-renders far fewer times than there are streamed chunks.
- Markdown block memoization:
MarkdownBlockrenders stay ~O(tokens + blocks), not O(tokens × blocks). - Main-thread health: long-task totals bounded relative to stream wall time.
- Typing latency after the long transcript: transcript components must not re-render per keystroke.
Run
react-scan is not a repo dependency; provide the bundle path. The recorded
baselines and thresholds were measured with react-scan 0.5.7 — instrumentation
overhead and onRender semantics are version-dependent, so keep it pinned:
npm i --no-save react-scan@0.5.7
npx playwright test --config=e2e/playwright.config.reasoning-perf.ts
or point REACT_SCAN_PATH at an existing
react-scan@0.5.7/dist/auto.global.js.
Requires a built client (client/dist) like the other mock e2e configs.