LibreChat/packages
Danny Avila c570517768
🐘 test: refresh FerretDB harness, registry-derived models, bulkWrite coverage (#14679)
* fix(data-schemas): refresh FerretDB harness model coverage, fix compile errors, add bulkWrite differentials

Track 2 of the search-stack plan (PLAN.md "FerretDB track"):

- Replace the three hand-rolled 29-model MODEL_SCHEMAS maps in
  multiTenancy/sharding/orgOperations.ferretdb.spec.ts with a shared
  getModelSchemas(mongoose) helper (misc/ferretdb/schemas.ts) derived from
  the live createModels() registry, so coverage tracks all 37 current
  models automatically instead of drifting. Matches the reference pattern
  in misc/documentdb/compat.documentdb.spec.ts.

- Fix the 3 compile-broken specs this uncovered: all three imported a
  `projectSchema` from '~/schema' that no longer exists (superseded by
  `chatProjectSchema`), which `tsc --noEmit` flags as TS2724 but the
  babel-based jest transform silently let through as `undefined`. Removing
  the hand-rolled maps removes the bad import as a side effect; verified
  clean with tsc across misc/ferretdb and misc/documentdb.

- Add misc/ferretdb/bulkWrite.ferretdb.spec.ts: differential specs for the
  five bulkWrite flows the plan names as actually at risk (import via
  bulkSaveConvos/bulkSaveMessages, bulkWriteAclEntries,
  bulkIncrementTagCounts, Transaction.insertMany, file-TTL bulkWrite via
  extendFilesTTL). Each flow runs identical operations against a real
  mongodb-memory-server (always) and, when FERRETDB_URI is set, against
  FerretDB, asserting normalized result equality. Multi-document
  transactions already degrade via the existing supportsTransactions
  probe — not duplicated here.

- Land the Spike A BSON-legibility findings (bson-legibility.md,
  bson-inventory.txt) from the bson-legibility-spike-6e38de worktree so
  decision 2's evidence is in-repo.

Verified against a real FerretDB 2.7.0 + postgres-documentdb 17 stack
(docker compose -f misc/ferretdb/docker-compose.ferretdb.yml): all 10
harness spec files pass individually, including all 10 bulkWrite.ferretdb
tests (5 mongodb-memory-server baselines + 5 FerretDB differentials). Full
packages/data-schemas src/ suite (1907 tests) unaffected.

* 📝 docs: make the BSON projection findings self-contained

The doc was written for readers who already knew the internal shorthand — it
opened on "Spike A executed, Spike B scoped" and referred to Options 1/2/3 and
"the handoff" without ever defining them, so a reader arriving from the repo
could not follow the argument or act on the recommendation.

Reframed around what the document actually investigates: the question is stated
up front, the three candidate mechanisms are named in a table before they are
compared, and the recommendation refers to them by name. No findings, numbers,
or SQL changed.

* 📝 docs: drop internal planning references from spec header

* fix(data-schemas): keep FerretDB harness schema derivation side-effect free

`getModelSchemas()` derived its map by calling `createModels(mongoose)`,
which carried three consequences the harness did not want:

- Model creation applies the tenant-isolation plugin to the module-level
  schema singletons, so every harness read and write inherited middleware
  that throws under `TENANT_ISOLATION_STRICT=true`.
- Registering on the default connection meant the benchmark's own
  `mongoose.connect()` auto-created 37 collections in the URI's base
  database, adding a database and dozens of collections to the very
  catalog metrics it measures.
- The unfiltered registry provisioned app-wide control-plane models
  (`SystemGrant`, `AuditLog`, `SkillSyncCredential`, `SkillSyncStatus`)
  into every org database.

The helper now builds the registry on a throwaway Mongoose instance,
returns schemas rebuilt from their own definition, options, and declared
indexes, skips the four app-wide models (validated against the registry so
a rename fails loudly), and memoizes the result.

Also in this pass:

- The "adds a new collection" migration test used `AuditLog`, which
  provisioning had already created, so it silently reused the production
  model and ignored its proposed schema. It now uses a fixture model absent
  from the registry and asserts the collection is missing beforehand and
  carries the proposed compound index afterwards.
- `bulkWrite` flows run inside `runAsSystem()`; they drive production
  methods unscoped, as a cross-tenant maintenance job does, and otherwise
  fail closed under strict tenant isolation.
- Phase 2's sparse-index assertion pinned a count the User schema no longer
  declares; it now checks that each index type round-trips.
2026-08-07 10:45:58 -04:00
..
api 🫆 chore: Remove Published Credential Defaults (#14680) 2026-08-07 07:25:05 -04:00
client fix: Restore WCAG AA Contrast for Text Tokens & Hide Edit Action While Streaming (#14677) 2026-08-07 00:52:07 -04:00
data-provider 🛟 fix: Stop Agents When Code Resources Cannot Recover (#14651) 2026-08-07 00:33:28 -04:00
data-schemas 🐘 test: refresh FerretDB harness, registry-derived models, bulkWrite coverage (#14679) 2026-08-07 10:45:58 -04:00