* 🧪 test: Reasoning-Stream Render Perf Benchmark via react-scan Adds a Playwright benchmark that streams one long, unsplit <think> block (18k chars — 4x the legacy SplitStreamHandler blockThreshold) plus 6k chars of markdown through the real mock-model agents pipeline, with react-scan injected to tally per-component renders. It verifies the legacy content-part splitting (removed in #10533) is not needed for rendering performance: - The whole reasoning section lands in ONE think part (a single Thoughts toggle) — nothing re-splits it anywhere in the pipeline. - rAF coalescing bounds the think box to ~1 render per 43 streamed chunks (122 renders / 5,290 chunks). - MarkdownBlock renders stay O(blocks + flushes) (153 renders / 2,092 text chunks), not O(blocks x tokens). - Long tasks during the 13.4s stream: one 96ms task; total render time 885ms. - Typing after the long transcript leaves transcript components quiet (<=2 renders across 40 keystrokes). Runs against the vite dev server (prod minification strips displayName assignments, which react-scan needs for naming). react-scan itself is not a repo dependency: install with `npm i --no-save react-scan` or point REACT_SCAN_PATH at its auto.global.js bundle. Also fixes the mock e2e stack for local runs: a developer .env with CHECK_BALANCE=true leaked through neutralizeCredentialEnv (not credential-shaped) and made every streaming mock spec fail with a token_balance violation, since the fresh e2e user has no balance record. vanillaOverrides now pins CHECK_BALANCE=false. * 🩹 fix: Address Codex Review — Payload Integrity, Frame Bounds, Proxy Port - Assert the full 18k-char reasoning payload survives the pipeline: expand the Thoughts toggle and compare rendered think text against the source (whitespace-normalized), instead of only counting toggles. - Derive render bounds from elapsed frames (60fps + headroom) rather than chunk counts, so the coalescing assertion stays meaningful regardless of how many chunks stream before resetPerf; apply the same bound to MarkdownBlock. - Tighten main-thread budgets: worst long task < 250ms and long-task total < 10% of stream wall time (baseline: one 51-96ms task per run). - Pass BACKEND_PORT derived from the configured E2E base URL to the vite dev server so its /api proxy follows a non-default app-server port. * 🧭 fix: Address Codex Round 2 — Typed Global, Drained Observer, Full-Payload Checks - Declare window.__PERF__ via global Window augmentation; drop the as-unknown-as double casts from both perf helpers. - Retain the longtask PerformanceObserver and drain takeRecords() before every snapshot/reset so stalls landing near the final render are counted. - Start the wall clock at the same instant as the tally reset so frame bounds and long-task percentages divide by exactly the measured interval. - Verify the complete markdown body: every generated section heading (exact-match), the exact list-item and table counts, and the generated code block — END_MARKER alone only proved the suffix rendered. - Require positive ThinkingContent/MarkdownBlock render counts so a renamed component or dropped instrumentation cannot void the upper bounds. - Cap cumulative render time at 25% of stream wall time to catch sustained sub-50ms work that never surfaces as a long task. * 🧷 fix: Address Codex Round 3 — Page Clock, Completion Wait, Exact Payload Checks - Measure the stream interval on the page's own clock: reset stamps the start, the snapshot evaluation reads the end, so bounds divide by exactly the tallied window including work between marker paint and snapshot. - Wait for the Stop generating button to hide before snapshotting, so generation finalization (usage chunk, terminal events, save re-render) is inside the measured interval. - Compare the rendered think text exactly (edges trimmed only) — internal paragraph breaks are user-visible under whitespace-pre-wrap and must survive verbatim. - Verify the markdown prose, not just structure: per-section doubled-sentence paragraph and both list-item texts, exact table count with cell values, and both generated code lines. - Derive the vite proxy port via getE2EServerAddress() so implicit ports in E2E_BASE_URL (default 80/443) agree between the app server and the proxy. - Pin react-scan@0.5.7 in the README — thresholds are calibrated against its instrumentation semantics. * 🪛 fix: Address Codex Round 4 — Pre-Send Reset, Count Every Code Block - Reset the tally immediately BEFORE triggering the send: with a 1ms chunk delay, the earliest deltas can render between the response headers resolving and a post-send evaluation, which the old order erased from the measurement. - Assert Math.floor(sectionCount / 3) occurrences of both generated code lines via code-element locators instead of .first(), so dropped later code blocks can no longer pass the payload check. * 🎛️ fix: Address Codex Round 5 — First-Render Clock, Expanded Box, Typing Budget - Stamp the wall clock at the FIRST render after each reset (inside onRender) so idle request-setup time between reset and stream start never pads the frame, long-task, or render-time denominators. - Seed showThinking=true so the reasoning box streams EXPANDED — the heavier live-layout path — and drop the post-hoc expand click. - Bound the typing phase itself: worst long task < 150ms and cumulative render time < 25% of the typed interval, so input lag without transcript re-renders still fails. - Derive the vite dev server host from getE2EServerAddress() alongside the port, so a non-localhost E2E base URL keeps the app server, listen host, and /api proxy in agreement. * 🧿 fix: Address Codex Round 6 — Stream-Anchored Clock, IPv6 Proxy, Rate-Free Bounds - Anchor the stream clock to the first ThinkingContent render — the payload opens with reasoning, so that is the first assistant-content paint — keeping composer renders and idle request setup out of the denominators. - Bracket IPv6 HOST values when building the vite /api proxy target in client/vite.config.ts; unbracketed ::1 produced an unparseable URL. - Add an absolute cumulative long-task budget (<300ms) to the typing phase so repeated sub-threshold stalls cannot evade the worst-case check or dilute the ratio via inflated elapsed time. - Add chunk-relative companion bounds (renders < chunks/4) for both ThinkingContent and MarkdownBlock, and hard-pin MOCK_LLM_CHUNK_DELAY_MS=1, so a slower stream can no longer loosen the coalescing assertions. |
||
|---|---|---|
| .devcontainer | ||
| .do/gitnexus | ||
| .github | ||
| .husky | ||
| .vscode | ||
| api | ||
| client | ||
| config | ||
| e2e | ||
| helm | ||
| otel/langfuse-fanout | ||
| packages | ||
| redis-config | ||
| scripts | ||
| skill | ||
| src/tests | ||
| utils | ||
| .dockerignore | ||
| .env.example | ||
| .gitattributes | ||
| .gitignore | ||
| .nvmrc | ||
| .prettierrc | ||
| AGENTS.md | ||
| bun.lock | ||
| CLAUDE.md | ||
| deploy-compose.langfuse-fanout.yml | ||
| deploy-compose.yml | ||
| docker-compose.langfuse-fanout.yml | ||
| docker-compose.override.yml.example | ||
| docker-compose.yml | ||
| Dockerfile | ||
| Dockerfile.multi | ||
| eslint.config.mjs | ||
| librechat.example.yaml | ||
| LICENSE | ||
| package-lock.json | ||
| package.json | ||
| rag.yml | ||
| README.md | ||
| README.zh.md | ||
| turbo.json | ||
LibreChat
English · 中文
✨ Features
-
🖥️ UI & Experience inspired by ChatGPT with enhanced design and features
-
🤖 AI Model Selection:
- Anthropic (Claude), AWS Bedrock, OpenAI, Azure OpenAI, Google, Vertex AI, OpenAI Responses API (incl. Azure)
- Custom Endpoints: Use any OpenAI-compatible API with LibreChat, no proxy required
- Compatible with Local & Remote AI Providers:
- Ollama, groq, Cohere, Mistral AI, Apple MLX, koboldcpp, together.ai,
- OpenRouter, Helicone, Perplexity, ShuttleAI, Deepseek, Qwen, and more
-
- Secure, Sandboxed Execution in Python, Node.js (JS/TS), Go, C/C++, Java, PHP, Rust, and Fortran
- Seamless File Handling: Upload, process, and download files directly
- No Privacy Concerns: Fully isolated and secure execution
- Open-Source & Self-Hostable: powered by ClickHouse/code-interpreter
-
🔦 Agents & Tools Integration:
- LibreChat Agents:
- No-Code Custom Assistants: Build specialized, AI-driven helpers
- Agent Marketplace: Discover and deploy community-built agents
- Collaborative Sharing: Share agents with specific users and groups
- Flexible & Extensible: Use MCP Servers, tools, file search, code execution, and more
- Skills: Create reusable
SKILL.mdinstruction bundles for manual, automatic, or always-on agent workflows - Subagents: Delegate focused work to isolated child agent runs with their own context windows
- Compatible with Custom Endpoints, OpenAI, Azure, Anthropic, AWS Bedrock, Google, Vertex AI, Responses API, and more
- Model Context Protocol (MCP) Support for Tools
- LibreChat Agents:
-
🔍 Web Search:
- Search the internet and retrieve relevant information to enhance your AI context
- Combines search providers, content scrapers, and result rerankers for optimal results
- Customizable Jina Reranking: Configure custom Jina API URLs for reranking services
- Learn More →
-
🪄 Generative UI with Code Artifacts:
- Code Artifacts allow creation of React, HTML, and Mermaid diagrams directly in chat
-
🎨 Image Generation & Editing
- Text-to-image and image-to-image with GPT-Image-1
- Text-to-image with DALL-E (3/2), Stable Diffusion, Flux, or any MCP server
- Produce stunning visuals from prompts or refine existing images with a single instruction
-
💾 Presets & Context Management:
- Create, Save, & Share Custom Presets
- Switch between AI Endpoints and Presets mid-chat
- Edit, Resubmit, and Continue Messages with Conversation branching
- Create and share prompts with specific users and groups
- Fork Messages & Conversations for Advanced Context control
-
💬 Multimodal & File Interactions:
- Upload and analyze images with Claude 3, GPT-4.5, GPT-4o, o1, Llama-Vision, and Gemini 📸
- Chat with Files using Custom Endpoints, OpenAI, Azure, Anthropic, AWS Bedrock, & Google 🗃️
-
🌎 Multilingual UI:
- English, 中文 (简体), 中文 (繁體), العربية, Deutsch, Español, Français, Italiano
- Polski, Português (PT), Português (BR), Русский, 日本語, Svenska, 한국어, Tiếng Việt
- Türkçe, Nederlands, עברית, Català, Čeština, Dansk, Eesti, فارسی
- Suomi, Magyar, Հայերեն, Bahasa Indonesia, ქართული, Latviešu, ไทย, ئۇيغۇرچە
-
🧠 Reasoning UI:
- Dynamic Reasoning UI for Chain-of-Thought/Reasoning AI models like DeepSeek-R1
-
🎨 Customizable Interface:
- Customizable Dropdown & Interface that adapts to both power users and newcomers
-
- Never lose a response: AI responses automatically reconnect and resume if your connection drops
- Multi-Tab & Multi-Device Sync: Open the same chat in multiple tabs or pick up on another device
- Production-Ready: Works from single-server setups to horizontally scaled deployments with Redis
-
🗣️ Speech & Audio:
- Chat hands-free with Speech-to-Text and Text-to-Speech
- Automatically send and play Audio
- Supports OpenAI, Azure OpenAI, and Elevenlabs
-
📥 Import & Export Conversations:
- Import Conversations from LibreChat, ChatGPT, Chatbot UI
- Export conversations as screenshots, markdown, text, json
-
🔍 Search & Discovery:
- Search all messages/conversations
-
👥 Multi-User & Secure Access:
- Multi-User, Secure Authentication with OAuth2, LDAP, & Email Login Support
- Built-in Moderation, and Token spend tools
-
🎛️ Admin Panel:
- Browser-based UI to manage users, groups, roles, and configuration overrides
- Edit settings and per-role/group permissions live, without redeploying
- Bundled with the Docker Compose stacks for one-command setup
-
⚙️ Configuration & Deployment:
- Configure Proxy, Reverse Proxy, Docker, & many Deployment options
- Use S3 with CloudFront for stable media links, edge delivery, signed cookies, and secured downloads
- Use completely local or deploy on the cloud
-
📖 Open-Source & Community:
- Completely Open-Source & Built in Public
- Community-driven development, support, and feedback
For a thorough review of our features, see our docs here 📚
🪶 All-In-One AI Conversations with LibreChat
LibreChat is a self-hosted AI chat platform that unifies all major AI providers in a single, privacy-focused interface.
Beyond chat, LibreChat provides AI Agents, Model Context Protocol (MCP) support, Artifacts, Code Interpreter, custom actions, conversation search, and enterprise-ready multi-user authentication.
Open source, actively developed, and built for anyone who values control over their AI infrastructure.
🌐 Resources
GitHub Repo:
- RAG API: github.com/danny-avila/rag_api
- Website: github.com/LibreChat-AI/librechat.ai
Other:
- Website: librechat.ai
- Documentation: librechat.ai/docs
- Blog: librechat.ai/blog
📝 Changelog
Keep up with the latest updates by visiting the releases page and notes:
⚠️ Please consult the changelog for breaking changes before updating.
⭐ Star History
✨ Contributions
Contributions, suggestions, bug reports and fixes are welcome!
For new features, components, or extensions, please open an issue and discuss before sending a PR.
If you'd like to help translate LibreChat into your language, we'd love your contribution! Improving our translations not only makes LibreChat more accessible to users around the world but also enhances the overall user experience. Please check out our Translation Guide.
💖 This project exists in its current state thanks to all the people who contribute
🎉 Special Thanks
We thank Locize for their translation management tools that support multiple languages in LibreChat.