mirror of
https://github.com/danny-avila/LibreChat.git
synced 2026-08-04 14:57:42 +00:00
* 🐛 fix: Dedup retried start-generation requests to prevent duplicate billing A lost or reset start-generation response makes the client re-POST the identical payload (up to 3x on network errors). The resumable-stream controller had no idempotency: createJob unconditionally overwrote the running job without aborting the prior one, so both requests ran full LLM completions and both billed while the UI showed only one (#14339). Add a stable per-submission clientRequestId (uuid, fresh per ask() so a regenerate differs, reused across the start-generation retries) and an atomic claim on the job store keyed by userId:clientRequestId. The first request wins and generates; a retried POST loses the claim and receives the original stream, which the client subscribes to and replays - no second billed generation. - IJobStore.claimIdempotencyKey/releaseIdempotencyKey (in-memory Map+TTL, Redis single-key SET NX PX + GET Lua, cluster-safe) - GenerationJobManager.claimGeneration/releaseGeneration (20m TTL) - Controller claims before the concurrency check, dedups with a resumed response, releases on start-failure/429 - clientRequestId threaded through TSubmission/TPayload/createPayload * 🐛 fix: Harden start-generation dedup (Codex review) Address three P2 findings on the idempotency path: - Resume replay: a deduped retry now subscribes with resume=true so the client replays prior content and any pending-action from the running stream instead of only live events (cross-replica / HITL correctness). startGeneration returns { streamId, resumed } and the response's status:'resumed' drives the subscribe mode. - Wait for the job record: a duplicate that loses the claim now waits briefly for the winner to create the job before returning the stream (a stream with no job 404s terminally). If the winner has not materialized, return 503 SERVER_NOT_READY so the client retries via the existing readiness path instead of attaching to a dead stream. - Release only owned claims: track whether the request actually won the claim; the 429 and init-error paths no longer release a claim owned by another in-flight generation (fail-open path could erase it and re-enable double billing). Adds controller tests covering dedup, the 503 race fallback, win-then- create, and claim-release ownership on 429 / fail-open. * 🐛 fix: Don't trap deduped retries on missing job records (Codex review) The previous round returned 503 SERVER_NOT_READY when a deduped retry's job record was absent. But a missing job usually means the original generation already completed and was cleaned up (cleanupOnComplete) — the correct recovery is to return the stream and let the client's subscribe 404 handler refetch the persisted messages. The 503 instead trapped the send in a readiness-retry loop until the client's window expired. Keep the bounded wait (it still covers the job-about-to-be-created race) but always return the resumed stream afterward; a gone/never-created job recovers via the client's existing 404 path instead of being treated as indefinitely starting. Updated the controller test accordingly. * 🐛 fix: Gate deduped resume on claim age, not just job presence (Codex review) Removing the 503 entirely (previous round) reintroduced the inverse race: if the winning request stalls between claimGeneration and createJob, a losing duplicate saw no job, returned status:'resumed' anyway, and the client subscribed to a stream that did not exist yet — the 404 handler tore the turn down while the winner went on to generate and bill with no UI attached. Distinguish the two missing-job cases by claim age (claimedAt now travels on the claim value): - fresh claim, no job yet → winner is still starting → 503 SERVER_NOT_READY so the client retries via the readiness path (bounded, not indefinite). - old claim, no job → the original already completed and was cleaned up (or the winner died) → attach; the client's 404 handler refetches. Tests cover both age branches. * 🐛 fix: Scope dedup fail-open + keep resumed convos on 404 (Codex review) - Fail-open only on claim acquisition: a store error while checking an already-confirmed existing claim no longer falls through to createJob (which would start a second billed generation during a Redis hiccup). Once claim.existing is known, a job-lookup error returns 503 retry. - Don't drop a resumed convo on 404: the optimistic-conversation cleanup in useResumableSSE now runs only for fresh (non-resume) subscribes. A deduped resume whose original completed and was cleaned up 404s, but its conversation is persisted and must stay in the sidebar. Adds a controller test for the job-lookup-error path (503, no createJob). * 🐛 fix: Reconcile resumed convos on 404 instead of guessing (Codex review) Round-4's !isResume guard fixed the completed-and-cleaned case (don't drop a persisted convo) but left the inverse: a new-conversation retry deduped to a claim whose original worker died before persisting still resumes, 404s, and — with removal skipped — leaves a phantom /c/<streamId> sidebar entry. Stop guessing keep-vs-remove on a resume 404. Reconcile against the server: invalidate the conversations list so a real (persisted) convo stays and a phantom is dropped. Fresh (non-resume) optimistic streams still prune immediately. Adds a client test for the resume path. * 🐛 fix: Finalize failed job before releasing its claim (Codex review) In the initialization-error catch, the idempotency claim was released before completeJob(streamId). A racing retry could win the released key and createJob() the same streamId while this catch was still running, and completeJob() (not guarded by the original createdAt) would then abort the replacement. Finalize the failed job first, then release the claim. Adds a controller test asserting completeJob precedes releaseGeneration. * 🐛 fix: Clear claims on destroy + survive completeJob failure (Codex review) - InMemoryJobStore.destroy() now clears the idempotencyClaims map, so a reused/reconfigured store instance doesn't dedup a fresh start against a torn-down job's stale claim. - Init-error cleanup: completeJob() is swallowed so a store-hiccup rejection can no longer skip the idempotency-key release and the pending-request decrement (which would wedge the retry behind the claim and leak the concurrency slot). A failed completeJob finalized nothing, so releasing afterward still can't abort a later replacement. Tests: claims cleared on destroy; release + pending decrement still run when completeJob rejects. |
||
|---|---|---|
| .. | ||
| app | ||
| cache | ||
| config | ||
| db | ||
| models | ||
| server | ||
| strategies | ||
| test | ||
| utils | ||
| jest.config.js | ||
| jsconfig.json | ||
| package.json | ||
| typedefs.js | ||