DataLoader: Batching and Caching to Solve N+1

Advanced
13 min

DataLoader: Batching and Caching to Solve N+1

The previous chapter showed why a list of 50 posts with an author field triggers 51 database queries. This chapter fixes it properly with DataLoader, the small utility Facebook released alongside GraphQL for exactly this purpose. After this lesson you will be able to write batch functions, wire loaders into the request context, and avoid the mistakes that silently disable batching.

How DataLoader Batches

A DataLoader wraps a batch function that takes an array of keys and returns a promise of an array of values. Each loader.load(key) call returns a promise immediately but does not run the batch function; DataLoader waits until the current event-loop tick finishes, collects every key requested in that window, and calls the batch function once. Because GraphQL resolves sibling fields concurrently, all 50 author resolvers call load() in the same tick and the batch function runs one WHERE id IN (...) query.

bash
npm install dataloader
javascript
import DataLoader from "dataloader"; export function createLoaders(db) { return { userById: new DataLoader(async (ids) => { const rows = await db.users.findByIds([...ids]); const byId = new Map(rows.map((u) => [u.id, u])); return ids.map((id) => byId.get(id) ?? null); }), }; }

The batch function has two strict contracts: the returned array must have the same length as ids, and the value at each index must match the key at that index. Databases do not guarantee the order of IN results, so the code maps rows by id and rebuilds the array in key order. Missing keys become null (or an Error instance to make that key fail).

Wiring Loaders into Context

Loaders belong in the per-request context so each request gets fresh instances:

javascript
const server = new ApolloServer({ typeDefs, resolvers }); const { url } = await startStandaloneServer(server, { context: async () => ({ db, loaders: createLoaders(db) }), }); const resolvers = { Query: { posts: (_p, _a, { db }) => db.posts.findAll() }, Post: { author: (post, _args, { loaders }) => loaders.userById.load(post.authorId), }, };

With this in place, { posts { title author { name } } } executes two queries regardless of how many posts are returned.

Per-Request Caching

DataLoader also memoises: calling load("1") twice in one request returns the same promise, so a user who wrote ten posts is fetched once. This cache lives inside the loader instance, which is exactly why loaders must be created in the context function rather than at module level. A module-level singleton would still batch, but it would serve stale data to every later request and leak memory.

After a mutation updates a record, clear the entry so later fields in the same operation see the new value:

javascript
updateUser: async (_p, { id, input }, { db, loaders }) => { const user = await db.users.update(id, input); loaders.userById.clear(id).prime(id, user); return user; },

clear(key) removes a cached value and prime(key, value) seeds the cache without a database call. Use loadMany([...ids]) when a resolver needs several keys at once.

Loading One-to-Many Relations

Batching works for lists too. The batch function returns an array of arrays, one per key:

javascript
postsByAuthorId: new DataLoader(async (authorIds) => { const rows = await db.posts.findByAuthorIds([...authorIds]); return authorIds.map((id) => rows.filter((p) => p.authorId === id)); }),

User.posts then resolves with loaders.postsByAuthorId.load(user.id) and the whole { users { posts { title } } } tree costs two queries.

Common Mistakes

  • Returning rows in database order instead of key order, which assigns the wrong author to posts.
  • Awaiting inside a loop before calling load(), which splits the calls across ticks and defeats batching.
  • Passing object keys; DataLoader compares keys by identity, so pass primitives or set cacheKeyFn.
  • Creating one loader per resolver call instead of per request, so there is never anything to batch.
Quick Quiz
Question 1 of 2

Why must loaders be created inside the context function rather than once at module load?

Key Takeaways

  • DataLoader collects every load(key) call made in one event-loop tick and issues a single batched request.
  • The batch function must return an array with the same length and order as its keys.
  • Create loaders per request in the context function; the built-in cache is meant to last one request only.
  • Use clear() and prime() after mutations, and loadMany() for several keys.
  • One-to-many relations batch the same way by returning an array of arrays.

Next lesson: Cursor Pagination and Relay Connections — implement stable, efficient pagination with edges, nodes and pageInfo.

DataLoader: Batching and Caching to Solve N+1 - GraphQL | CodeYourCraft | CodeYourCraft