Schema Design Patterns

Intermediate
13 min

Schema Design Patterns

Embedding versus referencing is the first decision; beyond it, MongoDB teams rely on a catalogue of named patterns for recurring problems: documents that grow forever, joins on every page view, totals recomputed on each read, and fields whose names differ per document. After this lesson you will be able to recognise those problems, apply the matching pattern, and avoid the anti-patterns that most often cause production incidents.

The Principle Behind Every Pattern

Model for the workload, not for the entities: ask what the application reads together, writes together, and how often. Data that is read together should be stored together; data that grows without limit should not live inside another document. Every pattern below is a controlled trade: a little duplication or write-time work in exchange for cheap, predictable reads.

| Pattern | Problem it solves | Trade-off | |---|---|---| | Extended Reference | join on every read | duplicated fields must be kept fresh | | Subset | large arrays where only a few items are shown | second collection for the full data | | Computed | totals recomputed on every read | extra write per source change | | Bucket | one document per event floods the collection | slightly more complex writes | | Attribute | many optional, differently named fields | key/value array is less readable | | Outlier | a few documents break the normal design | code must handle two shapes | | Schema Versioning | changing shape across millions of documents | readers handle several versions |

Extended Reference

Instead of a bare customerId, embed the handful of customer fields the order screen displays. Choose fields that rarely change (name, email) and keep the reference id so you can still look up the full record:

javascript
// When the customer renames themselves, fan the change out db.orders.updateMany({ "customer._id": 42 }, { $set: { "customer.name": "Ada King" } })

Subset and Computed

Subset keeps the most relevant part of a large relation inside the parent — the ten latest reviews on a product — and stores the complete set in a reviews collection. The product page needs one read; the "all reviews" page pages through the second collection.

Computed stores derived values at write time:

javascript
// On each new review, update the aggregate in the same write path db.products.updateOne( { _id: productId }, { $inc: { "stats.reviewCount": 1, "stats.ratingSum": 4 } } ) // Average on read: stats.ratingSum / stats.reviewCount

Bucket

Time-series data (sensor readings, page views, price ticks) produces millions of tiny documents. The bucket pattern groups them — one document per sensor per hour — with pre-computed count and sum fields, as in the sample code. Reads scan far fewer documents, and the count: { $lt: 360 } filter keeps each bucket bounded: once full, the upsert creates a new one.

MongoDB 5.0+ ships this idea as a collection type: db.createCollection("readings", { timeseries: { timeField: "t", metaField: "sensorId", granularity: "minutes" } }) stores buckets internally while you insert and query plain documents.

Irregular Data: Attribute, Outlier, Polymorphic and Versioning

Attribute: when products have unpredictable specifications, { specs: [{ k: "color", v: "red" }, { k: "weightKg", v: 4 }] } lets one compound index on { "specs.k": 1, "specs.v": 1 } serve every attribute query, instead of an index per field.

Outlier: design for the common case and mark the exceptions. A book embeds its buyer ids until a bestseller has millions; that document gets hasExtras: true and the overflow moves to a separate collection read only by the outlier code path.

Polymorphic: related types share a collection with a type discriminator, so one query returns all vehicles while type-specific fields stay optional.

Schema Versioning: add a schemaVersion field. New code writes the latest version, reads every version, and migrates a document lazily when it is next saved — no downtime, no giant migration script:

javascript
const user = db.users.findOne({ _id: id }); if ((user.schemaVersion ?? 1) < 2) { // v2 split "name" into first/last const [first, ...rest] = user.name.split(" "); db.users.updateOne({ _id: id }, { $set: { first, last: rest.join(" "), schemaVersion: 2 }, $unset: { name: "" } }); }

Anti-Patterns to Avoid

  • Unbounded arrays. Comments on a viral post will hit the 16 MB limit; bound them or move them out.
  • Bloated documents. Loading a 2 MB document to show a title wastes RAM; keep working-set fields together and cold data elsewhere.
  • Unnecessary indexes that slow every write without serving a query.
  • Separating data that is always accessed together, recreating relational joins for no benefit.
Quick Quiz
Question 1 of 3

Which pattern replaces a `$lookup` on every order page with a few duplicated customer fields inside the order?

Key Takeaways

  • Design for the workload: store together what is read together and keep growth bounded.
  • Extended Reference and Subset trade small, controlled duplication for single-read pages.
  • Computed and Bucket move work to write time, which pays off when reads dominate.
  • Attribute, Outlier and Polymorphic patterns handle irregular data without one index or collection per variant.
  • Schema Versioning allows shape changes without downtime; avoid unbounded arrays, bloated documents and needless indexes.

Next lesson: Schema Validation with $jsonSchema — enforce the structure you designed at the database level.

Schema Design Patterns - MongoDB | CodeYourCraft | CodeYourCraft