Commit graph

8 commits

Author SHA1 Message Date
Danny Avila
d369c649ed
🪥 chore: Run CI's Static Checks on Each Commit's Diff (#15303)
* 🪝 chore: Run Static Checks on Every Commit

Adds `scripts/static-checks.mts`, a local port of the Static Checks CI job
(.github/workflows/static-checks.yml) scoped to the files in a diff. It
resolves the changed-file list, applies the same `dorny/paths-filter` groups
the job uses, and runs whichever checks those paths activate — ESLint,
Prettier, import order, ESLint config validation, package.json validation, and
(behind `--full`) config migration tests, unused i18n keys and depcheck. Like
the job, every selected check runs even after one fails and the failures are
summarized at the end.

The pre-commit hook keeps lint-staged for the per-file layer, which verifies
the exact staged content of partially staged files, then runs the script for
everything lint-staged cannot cover. lint-staged now uses the job's ESLint
invocation, so warnings fail locally the way they fail CI.

The slow gates stay opt-in (`npm run static-checks:full`, or
`STATIC_CHECKS_FULL=1`) to keep commit latency unchanged.

Hooks were never installed: `config/prepare.js` existed but no `prepare` script
called it, so the hook only ran where `core.hooksPath` had been set by hand.
Replaces it with an inline `prepare` — both Dockerfiles run `npm ci` after
copying only the manifests, so a `node config/prepare.js` step would fail the
image build, and husky is absent from `--omit=dev` installs.

The i18n scan is a single pass over the source identifiers rather than one grep
per key, verified to flag exactly the same keys as the CI loop (including the
substring and dynamic-key cases) in 0.5s instead of 14s.

* 🩹 fix: Address Codex Round 1 on the Static Checks Runner

Activate gates from the unfiltered changed-path list. `dorny/paths-filter`
matches deleted paths too, so gating on the `--diff-filter=ACMRTUXB` list the
per-file steps use let a delete-only commit — the last reference to a
translation key, say — slip past the i18n and depcheck gates. The two lists are
now derived separately, the way CI derives them.

Pass `-m` to `git diff-tree` in `--commit` mode. Without it a merge commit
emits no paths at all, so `--commit <merge-sha>` reported "Nothing to check";
a real merge in this repo's history goes from 0 to 97 files.

Build the workspaces the config suite imports instead of skipping when `dist`
is absent. The three `dist` directories are gitignored and `npm ci` does not
produce them, so a fresh checkout reported a pass for a gate that never ran —
and an existing `dist` could be stale. Each is a sub-second tsdown build.

Resolve a global depcheck through a shell on Windows, where npm exposes it as
`depcheck.cmd` and `spawnSync` cannot see the shim.

Skip dot directories when walking for imports. `.claude/worktrees/` can hold a
full checkout per branch — 117 on this machine — and the root-wide scan behind
the depcheck gate walked every one of them.

Records the remaining boundary in the header: the per-file checks see exact
staged content via lint-staged, while the tree-wide gates read the working
tree, as running them by hand would.

* 🩹 fix: Address Codex Round 2 on the Static Checks Runner

Treat an unresolvable checker as a failure. ESLint and Prettier missing meant
the runner printed "All affected static checks passed" without having linted
anything; only depcheck, which CI installs globally and this documents as
optional, may still skip.

Catch per-check exceptions. The runner promises that every selected check runs
even after a failure, but a throw — a malformed translation JSON, say —
escaped and cancelled the checks after it. Each is now recorded as that
check's failure; verified that depcheck still runs after i18n throws.

Restrict `--commit` to the checked-out commit. Paths came from the named
commit while contents came from the working tree, so an older revision was
scored against the wrong file contents: a file added then deleted vanished,
and one modified since was read at its newer contents. It now fails with a
pointer to `--against`.

Reject unknown options. `--ful` silently ran the fast tier and `--commmit HEAD`
treated `HEAD` as a file path, both exiting 0 and implying gates had run.

Cover the runner in CI. `scripts/**` was absent from the workflow's trigger
paths, so a PR touching only the script the pre-commit hook now depends on got
no Static Checks run — and ESLint has no flat-config match for
`scripts/**/*.mts`, so nothing else loads it either. Adds the trigger path, a
`runner` filter group and a step that runs the script against the PR's own diff.

* 🔗 feat: Add Circular Dependency and TypeScript Gates

Both already run in CI as jobs of the Backend Unit Tests workflow; this brings
them to the local runner so they land before a push rather than after.

Circular dependencies (`node config/circular-deps.mjs`) is fast enough at 0.9s
to sit in the per-commit tier, gated on the same paths that trigger the CI job.

TypeScript stays opt-in behind `--full`: the five projects cost between 2.3s
and 20.9s each, which is too much per commit. Each project declares the paths
that can affect it — its own sources plus its upstream packages — so an edit to
data-provider still typechecks data-schemas, api, packages/client and client,
while an `api/**`-only change runs none of them, since no typechecked project
includes that directory. The builds a project's imports resolve through are
made first, mirroring the CI jobs' dependency on the build artifacts, through a
helper the config suite now shares.

Also addresses codex round 3:

Reject conflicting target selectors. `--against origin/dev package.json`
silently checked only the file, and `--against <bad-ref> --commit HEAD` never
resolved the bad base, so a caller could believe a range had been checked.

Require a clean worktree in `--commit` mode. The HEAD-only restriction was not
enough: contents still come from the working tree, so an uncommitted edit was
scored against the named commit — an invalid uncommitted package.json failing a
valid HEAD, or an uncommitted fix masking a defect in it.

The summary now names how many checks were skipped rather than reporting a
bare pass, and a typecheck failure carries the stale-workspace-build hint —
inside a git worktree `librechat-data-provider` resolves to the main checkout,
whose dist can predate the branch and shows up as missing properties.

* 🩹 fix: Address Codex Round 4 on the Static Checks Runner

Diff `--against` from the merge base. A two-dot diff reports the base branch's
own commits in reverse once it advances, so `--against origin/dev` scored 64
files for a branch that changed 7, activating gates for files the branch never
touched. Three dots makes the documented PR-style command mean what it says.

Typecheck on root manifest changes. Both review workflows trigger their
TypeScript jobs on package.json and package-lock.json, because a dependency or
@types bump breaks compilation on its own; the local filter ignored them, so
`static-checks:full` passed where CI would fail.

Validate every workspace manifest. The list mirrored the four the CI step
happens to name, so a malformed packages/api, data-provider or data-schemas
manifest passed validation in the revision modes, which have no lint-staged
pass behind them. Both lists now cover all seven.

Make the runner smoke execute a check. `--list` never runs one, and a
script-only PR activates no group, so the CI coverage added for exactly that
case could pass with the execution path untouched. It now runs against an
explicit target.

* 🩹 fix: Address Codex Round 5 on the Static Checks Runner

Activate the JSON gate for every manifest it validates. Round 4 added the four
workspace manifests to the validation list but not to the filter that turns the
gate on, so a malformed packages/data-provider or data-schemas manifest still
passed when it was the only changed file — the list grew and the trigger did
not. The same two entries also feed the unused-package calculation, reached
through api/package.json's @librechat/data-schemas dependency.

Include the owning workflows in the imported gates' filters. Circular
dependencies and TypeScript come from the review workflows, both of which list
their own YAML in `on.paths` and therefore rerun those jobs when the workflow
changes; locally the gates stayed inactive, so a change to how they are built
or invoked could bypass the local equivalent. Added to the group filters and to
the per-project predicates, since a workflow-only change would otherwise
activate the group and then select no project.

Bound command batches by characters rather than file count. Windows caps a
command line at 32767 characters, far below POSIX ARG_MAX, and a count does not
bound that: 400 of this repository's longer paths already come to 30176
characters before the executable and fixed arguments. Verified that a list
spanning several batches still reports a defect in its final file.
2026-08-28 10:08:37 -04:00
Danny Avila
c77e6a5ad2
🌉 fix: Preserve Missing Locize Translations (#14917) 2026-08-17 02:20:25 -04:00
Danny Avila
5ca667c258
🛡️ fix: Validate Translation Contracts Against English (#14914) 2026-08-16 22:55:05 -04:00
Danny Avila
6755544cee
🌍 ci: Harden Locize Translation Sync (#14784) 2026-08-13 07:29:29 -04:00
Danny Avila
b807292997
📉 perf: Bound Early Event Buffering for Detached Generations (#14612)
Some checks failed
Docker Dev Branch Images Build / build (Dockerfile, lc-dev, node) (push) Has been cancelled
Docker Dev Branch Images Build / build (Dockerfile.multi, lc-dev-api, api-build) (push) Has been cancelled
GitNexus Index / index (push) Has been cancelled
GitNexus Index / post-index (push) Has been cancelled
* 📉 perf: Bound Early Event Buffering for Detached Generations

A generation streaming with no attached subscriber re-entered buffering
mode on every disconnect and retained each emitted event in
earlyEventBuffer for its remaining duration. A single 26-minute detached
run (~58,800 tool-argument deltas) grew the heap past 2 GiB with GC cost
climbing alongside it, while client reconnects always resume from
durable state and discard that local buffer anyway.

- Close the early buffer after the first attachment drains it in Redis
  mode; the durable chunk log and pub/sub own recovery from then on,
  matching how cross-replica subscribers already attach.
- Enforce hard bounds (5,000 events / 8 MB estimated) in both modes; on
  overflow the buffer is discarded and closed, with recovery falling back
  to the durable chunk log (Redis) or resume snapshot (in-memory).
- Add a generation_stream_early_buffer_overflows_total counter and
  earlyBufferedEvents/Bytes gauges on getRuntimeStats() for visibility.
- Add incident-shaped regression tests and update specs that pinned the
  old post-disconnect re-buffering contract.

* fix: redirect post-overflow first attachments to resume recovery

A buffer discarded by the overflow guard left the initial non-resume
SSE attachment with nothing to replay, silently omitting pre-attach
output until the final event. Track the overflow on the runtime and
close such attachments with the existing reconnect signal instead: the
client already re-attaches with resume=true on transport failure and
its sync frame reconstructs the discarded output from durable/snapshot
state. Adds no per-event work; the check is one boolean per attachment.

* fix: enforce buffer bounds when restoring canceled resume captures

Captured emissions restored by a resume canceled before activation
bypassed the early-buffer hard cap, so one oversized restoration could
persist past the limits with no later emission to trip the guard.
Restoration now applies the same overflow-and-close behavior through a
shared helper, and the restore-cap spec fails before this change
(5 events / ~10MB retained) and passes after.

* chore: add Redis management scripts and update package.json for Redis commands
2026-08-03 19:03:37 -04:00
Danny Avila
ad4ed67070
🟦 chore: Convert Activity-Label Eval Harness to TypeScript (#14530) 2026-07-31 09:58:55 -04:00
Danny Avila
a07c0e4ae8
🧪 chore: Add the Activity-Label Prose Eval Harness (#14527)
Grades fast-model activity-label headers against a fixed corpus so
instruction changes are measured rather than eyeballed on one
conversation. This existed untracked while the continuity work was
developed; committing it because it is the only reproducible record of
WHY `ACTIVITY_INSTRUCTION` is ordered and capped the way it is.

- captured.json: 9 real production payloads pulled verbatim from Langfuse
  with the headers that shipped. Irreplaceable — traces age out.
- corpus.js: 17 cases / 28 steps. The captured run replays as one
  sequence, plus synthetic cases for the modes it never exercised
  (all-failed, partial, parallel batches, rapid near-duplicates, entry
  overflow, truncated output, error-shaped success). Multi-step cases
  chain each generated label into the next step's context, which is what
  makes cross-batch redundancy measurable at all.
- prompt.js: faithful port of the SDK's buildActivityLabelPrompt so
  synthetic cases render the bytes production sends, plus a
  previousLabelCap knob for continuity-window experiments.
- variants.js: single-factor instruction variants. The baseline is read
  from the BUILT package (workspace resolution, then dist, then
  LABEL_EVAL_DIST) so a variant can never be graded against a stale copy
  of the shipped instruction.
- checks.js: length/punctuation/markdown/tool-echo/count-echo, plus
  overlap split into `restate` (adds nothing over an earlier header) vs
  `template` (same frame, new payload — often fine).
- run.js / rescore.js: live runner on the production wire shape
  (max_tokens 256) and an offline re-grader, so metric fixes never
  require re-spending on the API.

Results are gitignored — regenerable, and 292K of the 364K. A full sweep
is ~$0.03 per variant and ~45s.

Findings are recorded in the README, two of them counter-intuitive:
enumerating acceptable opening verbs ANCHORED the model rather than
diversifying it (Confirmed 18→23, opener diversity halved), and diverse
examples alone changed nothing. Sentence order is load-bearing, so a
tidying reshuffle of ACTIVITY_INSTRUCTION regresses real output.
2026-07-30 09:22:34 -04:00
Danny Avila
bfb6b224d2
🔧 chore: Update ESLint config, Import Sorting script, Test Sharding, Bump @librechat/agents (#13552)
* 🔧 chore: Update ESLint config, add import sorting script, Test Sharding, Bump `@librechat/agents`

* Change 'no-nested-ternary' rule from 'warn' to 'error' in ESLint config
* Add new scripts for sorting imports in the project
* Update lint-staged configuration to include import sorting
* Modify GitHub Actions workflows to support sharding for unit tests

* chore: remove nested ternary expressions

* refactor: Extract scale multiplier logic into a separate function in CircleRender component
* refactor: Simplify auto-refill rendering logic in Balance component for better readability
* refactor: Improve width style handling in DataTable components for clarity and maintainability

* chore: remove CircleRender component

* delete: Remove CircleRender component as it is no longer needed in the project

* chore: Bump @librechat/agents to version 3.2.31 and update Node.js engine requirement

* Update @librechat/agents dependency from 3.2.2 to 3.2.31 in package-lock.json, api/package.json, and packages/api/package.json
* Change Node.js engine requirement from >=20.0.0 to >=24.0.0 in @librechat/agents

* chore: Add import sorting check to ESLint CI workflow

* Implement a new job in the GitHub Actions workflow to verify import ordering on changed files.
* The job checks for changes in specific file types and reports any import order drift, providing instructions for local fixes.
2026-06-06 12:31:55 -04:00