LibreChat/packages/data-schemas
Danny Avila c570517768
🐘 test: refresh FerretDB harness, registry-derived models, bulkWrite coverage (#14679)
* fix(data-schemas): refresh FerretDB harness model coverage, fix compile errors, add bulkWrite differentials

Track 2 of the search-stack plan (PLAN.md "FerretDB track"):

- Replace the three hand-rolled 29-model MODEL_SCHEMAS maps in
  multiTenancy/sharding/orgOperations.ferretdb.spec.ts with a shared
  getModelSchemas(mongoose) helper (misc/ferretdb/schemas.ts) derived from
  the live createModels() registry, so coverage tracks all 37 current
  models automatically instead of drifting. Matches the reference pattern
  in misc/documentdb/compat.documentdb.spec.ts.

- Fix the 3 compile-broken specs this uncovered: all three imported a
  `projectSchema` from '~/schema' that no longer exists (superseded by
  `chatProjectSchema`), which `tsc --noEmit` flags as TS2724 but the
  babel-based jest transform silently let through as `undefined`. Removing
  the hand-rolled maps removes the bad import as a side effect; verified
  clean with tsc across misc/ferretdb and misc/documentdb.

- Add misc/ferretdb/bulkWrite.ferretdb.spec.ts: differential specs for the
  five bulkWrite flows the plan names as actually at risk (import via
  bulkSaveConvos/bulkSaveMessages, bulkWriteAclEntries,
  bulkIncrementTagCounts, Transaction.insertMany, file-TTL bulkWrite via
  extendFilesTTL). Each flow runs identical operations against a real
  mongodb-memory-server (always) and, when FERRETDB_URI is set, against
  FerretDB, asserting normalized result equality. Multi-document
  transactions already degrade via the existing supportsTransactions
  probe — not duplicated here.

- Land the Spike A BSON-legibility findings (bson-legibility.md,
  bson-inventory.txt) from the bson-legibility-spike-6e38de worktree so
  decision 2's evidence is in-repo.

Verified against a real FerretDB 2.7.0 + postgres-documentdb 17 stack
(docker compose -f misc/ferretdb/docker-compose.ferretdb.yml): all 10
harness spec files pass individually, including all 10 bulkWrite.ferretdb
tests (5 mongodb-memory-server baselines + 5 FerretDB differentials). Full
packages/data-schemas src/ suite (1907 tests) unaffected.

* 📝 docs: make the BSON projection findings self-contained

The doc was written for readers who already knew the internal shorthand — it
opened on "Spike A executed, Spike B scoped" and referred to Options 1/2/3 and
"the handoff" without ever defining them, so a reader arriving from the repo
could not follow the argument or act on the recommendation.

Reframed around what the document actually investigates: the question is stated
up front, the three candidate mechanisms are named in a table before they are
compared, and the recommendation refers to them by name. No findings, numbers,
or SQL changed.

* 📝 docs: drop internal planning references from spec header

* fix(data-schemas): keep FerretDB harness schema derivation side-effect free

`getModelSchemas()` derived its map by calling `createModels(mongoose)`,
which carried three consequences the harness did not want:

- Model creation applies the tenant-isolation plugin to the module-level
  schema singletons, so every harness read and write inherited middleware
  that throws under `TENANT_ISOLATION_STRICT=true`.
- Registering on the default connection meant the benchmark's own
  `mongoose.connect()` auto-created 37 collections in the URI's base
  database, adding a database and dozens of collections to the very
  catalog metrics it measures.
- The unfiltered registry provisioned app-wide control-plane models
  (`SystemGrant`, `AuditLog`, `SkillSyncCredential`, `SkillSyncStatus`)
  into every org database.

The helper now builds the registry on a throwaway Mongoose instance,
returns schemas rebuilt from their own definition, options, and declared
indexes, skips the four app-wide models (validated against the registry so
a rename fails loudly), and memoizes the result.

Also in this pass:

- The "adds a new collection" migration test used `AuditLog`, which
  provisioning had already created, so it silently reused the production
  model and ignored its proposed schema. It now uses a fixture model absent
  from the registry and asserts the collection is missing beforehand and
  carries the proposed compound index afterwards.
- `bulkWrite` flows run inside `runAsSystem()`; they drive production
  methods unscoped, as a cross-tenant maintenance job does, and otherwise
  fail closed under strict tenant isolation.
- Phase 2's sparse-index assertion pinned a count the User schema no longer
  declares; it now checks that each index type round-trips.
2026-08-07 10:45:58 -04:00
..
misc 🐘 test: refresh FerretDB harness, registry-derived models, bulkWrite coverage (#14679) 2026-08-07 10:45:58 -04:00
src 🔗 fix: Prevent duplicate share rows from clobbering pending role edits (dedupe adds by stable id) (#14655) 2026-08-06 12:41:01 -04:00
.gitignore
babel.config.cjs
jest.config.mjs 🐘 feat: FerretDB Compatibility (#11769) 2026-03-21 14:28:49 -04:00
LICENSE
package.json 🌯 chore: Retire Rollup-Era devDependencies After tsdown Migration (#14496) 2026-07-28 22:08:52 -04:00
README.md
tsconfig.build.json 📦 refactor: Consolidate DB models, encapsulating Mongoose usage in data-schemas (#11830) 2026-03-21 14:28:53 -04:00
tsconfig.json 🔧 chore: Enforce isolatedDeclarations in data-schemas tsconfig (#13593) 2026-06-08 09:55:20 -04:00
tsconfig.spec.json 📦 chore: Update TypeScript Config for TS v7 (#12794) 2026-04-23 12:51:03 -04:00
tsdown.config.mjs 🌀 ci: Deterministic Circular Dependency Checks (#14579) 2026-08-01 14:43:26 -04:00

LibreChat Data Schemas Package

This package provides the database schemas, models, types, and methods for LibreChat using Mongoose ODM.

📁 Package Structure

packages/data-schemas/
├── src/
│   ├── schema/         # Mongoose schema definitions
│   ├── models/         # Model factory functions
│   ├── types/          # TypeScript type definitions
│   ├── methods/        # Database operation methods
│   ├── common/         # Shared constants and enums
│   ├── config/         # Configuration files (winston, etc.)
│   └── index.ts        # Main package exports

🏗️ Architecture Patterns

1. Schema Files (src/schema/)

Schema files define the Mongoose schema structure. They follow these conventions:

  • Naming: Use lowercase filenames (e.g., user.ts, accessRole.ts)
  • Imports: Import types from ~/types for TypeScript support
  • Exports: Export only the schema as default

Example:

import { Schema } from 'mongoose';
import type { IUser } from '~/types';

const userSchema = new Schema<IUser>(
  {
    name: { type: String },
    email: { type: String, required: true },
    // ... other fields
  },
  { timestamps: true }
);

export default userSchema;

2. Type Definitions (src/types/)

Type files define TypeScript interfaces and types. They follow these conventions:

  • Base Type: Define a plain type without Mongoose Document properties
  • Document Interface: Extend the base type with Document and _id
  • Enums/Constants: Place related enums in the type file or common/ if shared

Example:

import type { Document, Types } from 'mongoose';

export type User = {
  name?: string;
  email: string;
  // ... other fields
};

export type IUser = User &
  Document & {
    _id: Types.ObjectId;
  };

3. Model Factory Functions (src/models/)

Model files create Mongoose models using factory functions. They follow these conventions:

  • Function Name: create[EntityName]Model
  • Singleton Pattern: Check if model exists before creating
  • Type Safety: Use the corresponding interface from types

Example:

import userSchema from '~/schema/user';
import type * as t from '~/types';

export function createUserModel(mongoose: typeof import('mongoose')) {
  return mongoose.models.User || mongoose.model<t.IUser>('User', userSchema);
}

4. Database Methods (src/methods/)

Method files contain database operations for each entity. They follow these conventions:

  • Function Name: create[EntityName]Methods
  • Return Type: Export a type for the methods object
  • Operations: Include CRUD operations and entity-specific queries

Example:

import type { Model } from 'mongoose';
import type { IUser } from '~/types';

export function createUserMethods(mongoose: typeof import('mongoose')) {
  async function findUserById(userId: string): Promise<IUser | null> {
    const User = mongoose.models.User as Model<IUser>;
    return await User.findById(userId).lean();
  }

  async function createUser(userData: Partial<IUser>): Promise<IUser> {
    const User = mongoose.models.User as Model<IUser>;
    return await User.create(userData);
  }

  return {
    findUserById,
    createUser,
    // ... other methods
  };
}

export type UserMethods = ReturnType<typeof createUserMethods>;

5. Main Exports (src/index.ts)

The main index file exports:

  • createModels() - Factory function for all models
  • createMethods() - Factory function for all methods
  • Type exports from ~/types
  • Shared utilities and constants

🚀 Adding a New Entity

To add a new entity to the data-schemas package, follow these steps:

Step 1: Create the Type Definition

Create src/types/[entityName].ts:

import type { Document, Types } from 'mongoose';

export type EntityName = {
  /** Field description */
  fieldName: string;
  // ... other fields
};

export type IEntityName = EntityName &
  Document & {
    _id: Types.ObjectId;
  };

Step 2: Update Types Index

Add to src/types/index.ts:

export * from './entityName';

Step 3: Create the Schema

Create src/schema/[entityName].ts:

import { Schema } from 'mongoose';
import type { IEntityName } from '~/types';

const entityNameSchema = new Schema<IEntityName>(
  {
    fieldName: { type: String, required: true },
    // ... other fields
  },
  { timestamps: true }
);

export default entityNameSchema;

Step 4: Create the Model Factory

Create src/models/[entityName].ts:

import entityNameSchema from '~/schema/entityName';
import type * as t from '~/types';

export function createEntityNameModel(mongoose: typeof import('mongoose')) {
  return (
    mongoose.models.EntityName || 
    mongoose.model<t.IEntityName>('EntityName', entityNameSchema)
  );
}

Step 5: Update Models Index

Add to src/models/index.ts:

  1. Import the factory function:
import { createEntityNameModel } from './entityName';
  1. Add to the return object in createModels():
EntityName: createEntityNameModel(mongoose),

Step 6: Create Database Methods

Create src/methods/[entityName].ts:

import type { Model, Types } from 'mongoose';
import type { IEntityName } from '~/types';

export function createEntityNameMethods(mongoose: typeof import('mongoose')) {
  async function findEntityById(id: string | Types.ObjectId): Promise<IEntityName | null> {
    const EntityName = mongoose.models.EntityName as Model<IEntityName>;
    return await EntityName.findById(id).lean();
  }

  // ... other methods

  return {
    findEntityById,
    // ... other methods
  };
}

export type EntityNameMethods = ReturnType<typeof createEntityNameMethods>;

Step 7: Update Methods Index

Add to src/methods/index.ts:

  1. Import the methods:
import { createEntityNameMethods, type EntityNameMethods } from './entityName';
  1. Add to the return object in createMethods():
...createEntityNameMethods(mongoose),
  1. Add to the AllMethods type:
export type AllMethods = UserMethods &
  // ... other methods
  EntityNameMethods;

📝 Best Practices

  1. Consistent Naming: Use lowercase for filenames, PascalCase for types/interfaces
  2. Type Safety: Always use TypeScript types, avoid any
  3. JSDoc Comments: Document complex fields and methods
  4. Indexes: Define database indexes in schema files for query performance
  5. Validation: Use Mongoose schema validation for data integrity
  6. Lean Queries: Use .lean() for read operations when you don't need Mongoose document methods

🔧 Common Patterns

Enums and Constants

Place shared enums in src/common/:

// src/common/permissions.ts
export enum PermissionBits {
  VIEW = 1,
  EDIT = 2,
  DELETE = 4,
  SHARE = 8,
}

Compound Indexes

For complex queries, add compound indexes:

schema.index({ field1: 1, field2: 1 });
schema.index(
  { uniqueField: 1 },
  { 
    unique: true, 
    partialFilterExpression: { uniqueField: { $exists: true } }
  }
);

Virtual Properties

Add computed properties using virtuals:

schema.virtual('fullName').get(function() {
  return `${this.firstName} ${this.lastName}`;
});

🧪 Testing

When adding new entities, ensure:

  • Types compile without errors
  • Models can be created successfully
  • Methods handle edge cases (null checks, validation)
  • Indexes are properly defined for query patterns

📚 Resources