Text Search: $regex, Text Indexes and Atlas Search

Intermediate
13 min

Text Search: $regex, Text Indexes and Atlas Search

"Find documents containing this word" sounds simple, but the right tool depends on whether you need pattern matching, word-level relevance, or a search box that tolerates typos. MongoDB offers three levels: the $regex operator, text indexes with $text, and Atlas Search. After this lesson you will be able to write each of them, know which one uses an index, and pick the right level for a feature.

$regex: Pattern Matching

$regex evaluates a Perl-compatible regular expression against a string field. Both syntaxes below are equivalent:

javascript
db.articles.find({ title: /^mongo/i }) db.articles.find({ title: { $regex: "^mongo", $options: "i" } }) db.users.find({ email: { $regex: "@example\\.com$" } }) // suffix match db.articles.find({ title: { $not: /draft/i } }) // negation

Options: i (case-insensitive), m (multi-line anchors), s (dot matches newline), x (ignore whitespace in the pattern).

Performance is the catch. A case-sensitive prefix pattern such as ^Mongo turns into an index range scan on an indexed field. Anything else — a leading wildcard, an unanchored pattern, or the i option — has to test every index key or every document. $regex is fine for admin filters and small collections; it is not a search engine.

Text Indexes and $text

A text index tokenises string fields into words, lowercases them, removes stop words and stems them (running becomes run) for the configured language. $text then searches that index:

javascript
db.articles.createIndex( { title: "text", body: "text" }, { weights: { title: 5, body: 1 }, default_language: "english" } ) db.articles.find( { $text: { $search: "index performance -sharding" } }, { title: 1, score: { $meta: "textScore" } } ).sort({ score: { $meta: "textScore" } })

Search-string rules: terms are combined with OR, "quoted phrases" must appear verbatim, and a leading - excludes a term. Weights scale the relevance score per field, and $meta: "textScore" exposes it for projection and sorting. Options $caseSensitive, $diacriticSensitive and $language override the defaults per query.

Constraints to plan around:

  • A collection can have one text index, though it may cover several fields (or every string field with { "$**": "text" }).
  • $text must sit at the top level of the query, and in an aggregation it must be inside the first $match.
  • No fuzzy matching, no autocomplete, no "did you mean"; results are word-level only.

Atlas Search

Atlas Search is a Lucene-based engine that runs beside your Atlas cluster (it is available on every tier, including the free M0). You define a search index — in the Atlas UI, the Atlas CLI or the Admin API — and query it with the $search aggregation stage, which must be the first stage of the pipeline.

json
{ "mappings": { "dynamic": false, "fields": { "title": [{ "type": "string" }, { "type": "autocomplete" }], "body": { "type": "string" }, "category": { "type": "token" } } } }
javascript
db.articles.aggregate([ { $search: { index: "default", compound: { must: [{ text: { query: "index performance", path: ["title", "body"], fuzzy: { maxEdits: 1 } } }], filter: [{ equals: { path: "category", value: "database" } }] } } }, { $project: { title: 1, score: { $meta: "searchScore" } } }, { $limit: 10 } ])

Operators include text, phrase, autocomplete (search-as-you-type), range, wildcard, equals and compound (must, should, mustNot, filter). Add highlight to return matched snippets, use $searchMeta for facet counts, and use the sibling $vectorSearch stage for embedding-based semantic search. The index updates asynchronously, so a document inserted a moment ago may take a second to become searchable.

Choosing the Right Level

| Need | Use | Index-backed? | |---|---|---| | exact prefix or simple pattern, small data | $regex | only case-sensitive ^prefix | | word search with relevance, self-hosted | text index + $text | yes | | typo tolerance, autocomplete, facets, highlighting | Atlas Search $search | yes (Lucene) | | "similar meaning" search over embeddings | $vectorSearch | yes (vector index) |

Common Mistakes

  • Using /.*term.*/i as site search. It scans everything and cannot rank results.
  • Creating a second text index. The call fails; extend the existing one instead.
  • Placing $search after $match. It must be the first stage; put filters inside compound.filter.
  • Forgetting to sort by score. $text returns matches in arbitrary order unless you sort on textScore.
Quick Quiz
Question 1 of 3

Which `$regex` query can use a regular index on `title` efficiently?

Key Takeaways

  • $regex is pattern matching; only a case-sensitive ^prefix benefits from a normal index.
  • A text index tokenises and stems words; $text searches it, ranks with textScore and supports phrases and negation.
  • One text index per collection, $text at the top level of the query, no fuzzy matching.
  • Atlas Search adds Lucene features — fuzzy text, autocomplete, facets, highlighting — through the $search stage, which must come first.
  • Match the tool to the feature: admin filters, word search, or a user-facing search box.

Next lesson: Reading explain() Plans and the Database Profiler — see exactly how MongoDB executes a query and find the slow ones in production.

Text Search: $regex, Text Indexes and Atlas Search - MongoDB | CodeYourCraft | CodeYourCraft