LibreChat/e2e/benchmarks-reasoning
Danny Avila 4f5808d9ae
🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan (#14494)
* 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan

Adds a Playwright benchmark that streams one long, unsplit <think> block
(18k chars — 4x the legacy SplitStreamHandler blockThreshold) plus 6k chars
of markdown through the real mock-model agents pipeline, with react-scan
injected to tally per-component renders. It verifies the legacy content-part
splitting (removed in #10533) is not needed for rendering performance:

- The whole reasoning section lands in ONE think part (a single Thoughts
  toggle) — nothing re-splits it anywhere in the pipeline.
- rAF coalescing bounds the think box to ~1 render per 43 streamed chunks
  (122 renders / 5,290 chunks).
- MarkdownBlock renders stay O(blocks + flushes) (153 renders / 2,092 text
  chunks), not O(blocks x tokens).
- Long tasks during the 13.4s stream: one 96ms task; total render time 885ms.
- Typing after the long transcript leaves transcript components quiet
  (<=2 renders across 40 keystrokes).

Runs against the vite dev server (prod minification strips displayName
assignments, which react-scan needs for naming). react-scan itself is not a
repo dependency: install with `npm i --no-save react-scan` or point
REACT_SCAN_PATH at its auto.global.js bundle.

Also fixes the mock e2e stack for local runs: a developer .env with
CHECK_BALANCE=true leaked through neutralizeCredentialEnv (not
credential-shaped) and made every streaming mock spec fail with a
token_balance violation, since the fresh e2e user has no balance record.
vanillaOverrides now pins CHECK_BALANCE=false.

* 🩹 fix: Address Codex Review — Payload Integrity, Frame Bounds, Proxy Port

- Assert the full 18k-char reasoning payload survives the pipeline: expand
  the Thoughts toggle and compare rendered think text against the source
  (whitespace-normalized), instead of only counting toggles.
- Derive render bounds from elapsed frames (60fps + headroom) rather than
  chunk counts, so the coalescing assertion stays meaningful regardless of
  how many chunks stream before resetPerf; apply the same bound to
  MarkdownBlock.
- Tighten main-thread budgets: worst long task < 250ms and long-task total
  < 10% of stream wall time (baseline: one 51-96ms task per run).
- Pass BACKEND_PORT derived from the configured E2E base URL to the vite dev
  server so its /api proxy follows a non-default app-server port.

* 🧭 fix: Address Codex Round 2 — Typed Global, Drained Observer, Full-Payload Checks

- Declare window.__PERF__ via global Window augmentation; drop the
  as-unknown-as double casts from both perf helpers.
- Retain the longtask PerformanceObserver and drain takeRecords() before
  every snapshot/reset so stalls landing near the final render are counted.
- Start the wall clock at the same instant as the tally reset so frame
  bounds and long-task percentages divide by exactly the measured interval.
- Verify the complete markdown body: every generated section heading
  (exact-match), the exact list-item and table counts, and the generated
  code block — END_MARKER alone only proved the suffix rendered.
- Require positive ThinkingContent/MarkdownBlock render counts so a renamed
  component or dropped instrumentation cannot void the upper bounds.
- Cap cumulative render time at 25% of stream wall time to catch sustained
  sub-50ms work that never surfaces as a long task.

* 🧷 fix: Address Codex Round 3 — Page Clock, Completion Wait, Exact Payload Checks

- Measure the stream interval on the page's own clock: reset stamps the
  start, the snapshot evaluation reads the end, so bounds divide by exactly
  the tallied window including work between marker paint and snapshot.
- Wait for the Stop generating button to hide before snapshotting, so
  generation finalization (usage chunk, terminal events, save re-render) is
  inside the measured interval.
- Compare the rendered think text exactly (edges trimmed only) — internal
  paragraph breaks are user-visible under whitespace-pre-wrap and must
  survive verbatim.
- Verify the markdown prose, not just structure: per-section doubled-sentence
  paragraph and both list-item texts, exact table count with cell values, and
  both generated code lines.
- Derive the vite proxy port via getE2EServerAddress() so implicit ports in
  E2E_BASE_URL (default 80/443) agree between the app server and the proxy.
- Pin react-scan@0.5.7 in the README — thresholds are calibrated against its
  instrumentation semantics.

* 🪛 fix: Address Codex Round 4 — Pre-Send Reset, Count Every Code Block

- Reset the tally immediately BEFORE triggering the send: with a 1ms chunk
  delay, the earliest deltas can render between the response headers
  resolving and a post-send evaluation, which the old order erased from the
  measurement.
- Assert Math.floor(sectionCount / 3) occurrences of both generated code
  lines via code-element locators instead of .first(), so dropped later
  code blocks can no longer pass the payload check.

* 🎛️ fix: Address Codex Round 5 — First-Render Clock, Expanded Box, Typing Budget

- Stamp the wall clock at the FIRST render after each reset (inside
  onRender) so idle request-setup time between reset and stream start never
  pads the frame, long-task, or render-time denominators.
- Seed showThinking=true so the reasoning box streams EXPANDED — the heavier
  live-layout path — and drop the post-hoc expand click.
- Bound the typing phase itself: worst long task < 150ms and cumulative
  render time < 25% of the typed interval, so input lag without transcript
  re-renders still fails.
- Derive the vite dev server host from getE2EServerAddress() alongside the
  port, so a non-localhost E2E base URL keeps the app server, listen host,
  and /api proxy in agreement.

* 🧿 fix: Address Codex Round 6 — Stream-Anchored Clock, IPv6 Proxy, Rate-Free Bounds

- Anchor the stream clock to the first ThinkingContent render — the payload
  opens with reasoning, so that is the first assistant-content paint —
  keeping composer renders and idle request setup out of the denominators.
- Bracket IPv6 HOST values when building the vite /api proxy target in
  client/vite.config.ts; unbracketed ::1 produced an unparseable URL.
- Add an absolute cumulative long-task budget (<300ms) to the typing phase
  so repeated sub-threshold stalls cannot evade the worst-case check or
  dilute the ratio via inflated elapsed time.
- Add chunk-relative companion bounds (renders < chunks/4) for both
  ThinkingContent and MarkdownBlock, and hard-pin MOCK_LLM_CHUNK_DELAY_MS=1,
  so a slower stream can no longer loosen the coalescing assertions.
2026-07-28 22:18:24 -04:00
..
payload.ts 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan (#14494) 2026-07-28 22:18:24 -04:00
README.md 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan (#14494) 2026-07-28 22:18:24 -04:00
reasoning-stream.perf.spec.ts 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan (#14494) 2026-07-28 22:18:24 -04:00

Reasoning-Stream Perf Benchmark (react-scan)

Verifies that streaming one long, unsplit reasoning block (plus long markdown text) through the real mock-model agents pipeline stays render-bounded — i.e. the legacy content-part splitting (SplitStreamHandler / blockThreshold, removed in #10533) is not needed for rendering performance.

What it measures, via react-scan injected into the page:

  • Per-component render counts and render time while a ~18k-char <think> block and ~6k-char markdown reply stream token by token.
  • That the whole reasoning section lands in one think part (a single "Thoughts" toggle) — no re-splitting anywhere in the pipeline.
  • rAF coalescing: the think box re-renders far fewer times than there are streamed chunks.
  • Markdown block memoization: MarkdownBlock renders stay ~O(tokens + blocks), not O(tokens × blocks).
  • Main-thread health: long-task totals bounded relative to stream wall time.
  • Typing latency after the long transcript: transcript components must not re-render per keystroke.

Run

react-scan is not a repo dependency; provide the bundle path. The recorded baselines and thresholds were measured with react-scan 0.5.7 — instrumentation overhead and onRender semantics are version-dependent, so keep it pinned:

npm i --no-save react-scan@0.5.7
npx playwright test --config=e2e/playwright.config.reasoning-perf.ts

or point REACT_SCAN_PATH at an existing react-scan@0.5.7/dist/auto.global.js.

Requires a built client (client/dist) like the other mock e2e configs.