MongoDB interview questions & answers
Document databases: schema design and embedding vs referencing, indexes, the aggregation pipeline, replication, sharding and when to choose NoSQL.
Official reference: MongoDB manual
Top 59 MongoDB interview questions most asked first
1.What is MongoDB, and when would you choose it over a relational database?easy
MongoDB is a document database. It stores records as BSON documents (JSON-like, with nested objects and arrays) grouped into collections, with a flexible schema, a rich query language, secondary indexes, an aggregation framework, built-in replication through replica sets and horizontal scaling through sharding.
I'd pick it when:
- The data is naturally hierarchical or varies per record: product catalogs, user profiles, content, events, IoT readings
- The app mostly reads and writes whole aggregates, so one document can hold what would be several joined tables
- The schema evolves quickly, or the dataset will outgrow one server
I'd lean relational when the domain is dominated by many-to-many relationships and ad hoc joins, or when most operations touch many entities transactionally. MongoDB does have ACID multi-document transactions and schema validation, so the real decision is about data shape and access patterns, not "no transactions".
What interviewers listen for- Document database storing BSON documents in collections
- Flexible schema, rich queries, indexes, aggregation
- Replica sets for HA, sharding for scale-out
- Fits aggregates read together and evolving schemas
- Relational fits join-heavy, highly relational domains
Likely follow-up: What would make you move a MongoDB workload back to SQL?
2.How does MongoDB compare with a SQL database? Map the main concepts and explain the key differences.easy
The terminology maps fairly directly:
- table → collection, row → document, column → field
- primary key → the
_idfield - join → embedding related data, or
$lookupin an aggregation GROUP BY→ the$groupstage
The bigger differences are in philosophy. SQL databases enforce a fixed schema and normalize data by entity, then join at query time. MongoDB lets documents in a collection differ (you can add
$jsonSchemavalidation when you want rules), and you model around access patterns, storing data that's read together in one document.MongoDB doesn't enforce foreign keys, so referential integrity is the application's job. On scaling, MongoDB has sharding built in, whereas many relational setups scale up first. Both offer ACID transactions; in MongoDB a single-document write is always atomic, and multi-document transactions exist but cost more.
What interviewers listen for- Collection, document, field,
_idmap to table, row, column, PK - Joins become embedding or
$lookup - Flexible schema with optional validation
- Model by access pattern, not by entity
- No enforced foreign keys; sharding built in
Likely follow-up: What are the main categories of NoSQL databases?
3.What is BSON, and why does MongoDB store BSON instead of plain JSON?easy
BSON is a binary-encoded serialization of JSON-like documents. MongoDB uses it on disk and on the wire for two reasons.
First, richer types. JSON only has strings, numbers, booleans, null, objects and arrays. BSON adds
ObjectId,Date, 32- and 64-bit integers, doubles,Decimal128for exact money values, binary data, regular expressions and more, so a date stays a date and you can range-query and index it correctly.Second, efficient traversal. Elements are type-tagged and length-prefixed, so the server can skip over fields without parsing everything.
A document is an ordered set of field/value pairs, up to 16 MiB and 100 levels of nesting. Documents live in collections, which live in databases. Both are created implicitly on first insert. When mongosh prints BSON it uses helpers like
ObjectId(...)andISODate(...); drivers exchange it as Extended JSON when text is needed.db.orders.insertOne({ customerId: ObjectId("64f1c2a9e4b0a1b2c3d4e5f6"), total: Decimal128("59.90"), qty: NumberInt(3), placedAt: new Date(), items: [{ sku: "A1", price: Decimal128("19.96") }] });What interviewers listen for- Binary, type-tagged, length-prefixed JSON-like format
- Extra types: ObjectId, Date, Int32/Int64, Decimal128, binary
- Faster to scan than parsing text JSON
- Documents max 16 MiB, 100 nesting levels
- Collections and databases are created implicitly
Likely follow-up: Why use
Decimal128rather than a double for prices?4.What is the
_idfield, and what is inside an ObjectId?easyEvery document in a standard collection needs an
_idthat acts as its primary key. It must be unique within the collection, it's immutable once set, and it can be any BSON type except an array, a regex or undefined. MongoDB automatically creates a unique index on_idfor every collection. If you insert a document without one, the driver generates an ObjectId.An ObjectId is 12 bytes:
- a 4-byte timestamp: seconds since the Unix epoch
- a 5-byte random value generated once per client process
- a 3-byte incrementing counter, initialized to a random value
So IDs are generated client-side with no coordination, and
getTimestamp()recovers the creation time. Because they roughly increase over time, sorting by_idapproximates insertion order, but it isn't strictly monotonic across different clients. That same "always increasing" property makes a plain ranged ObjectId a poor shard key.What interviewers listen for- Primary key: unique, immutable, automatically indexed
- Any BSON type except array, regex or undefined
- 4-byte timestamp, 5-byte random, 3-byte counter
- Generated by the driver without coordination
- Roughly time-ordered; bad as a ranged shard key
Likely follow-up: When would you use a natural key instead of an ObjectId for
_id?5.How do you decide whether to embed related data in a document or reference it from another collection?easy
The guiding principle is data that's accessed together should be stored together, and I start from embedding unless there's a reason not to.
Embed when:
- The child is owned by the parent and read with it, like an order's line items or a user's addresses
- The relationship is one-to-few and the array is bounded
- You want the parent and children updated atomically in a single-document write
Reference when:
- The "many" side is large or grows without limit, so the document would bloat toward the 16 MiB cap
- The child needs to be queried or updated on its own
- The data is shared by many parents and changes often, so duplicating it would be expensive to keep in sync
- It's many-to-many
Often the answer is a hybrid: reference the full record but embed a few frequently read fields (the extended reference pattern). Always design from the queries the application actually runs.
What interviewers listen for- Data accessed together is stored together
- Embed bounded, owned, read-together children
- Reference unbounded, shared or independently queried data
- Embedding gives single-document atomicity
- Hybrid: reference plus a few duplicated fields
Likely follow-up: How would you model blog posts and their comments?
6.Why do we need indexes in MongoDB, and which index types does it support?easy
Without a suitable index a query has to do a collection scan (
COLLSCAN), reading every document. An index is an ordered B-tree of field values pointing to documents, so the server can jump to matching entries and even return results in index order without sorting.Index types:
- Single field and compound (several fields, order matters)
- Multikey: created automatically when an indexed field holds an array
- Text: keyword search with stemming; one per collection
- Geospatial:
2dsphereand2d - Hashed: used mainly for hashed sharding
- Wildcard:
{ "$**": 1 }for unpredictable field names
Index properties include
unique,partialFilterExpression,sparse, TTL (expireAfterSeconds),hiddenand a collation for case-insensitive matching. Every collection gets a unique_idindex.The trade-off: each index costs RAM and disk, and slows every insert, update and delete that touches its fields.
What interviewers listen for- Avoids COLLSCAN; B-tree lookups and index-ordered sorts
- Single, compound, multikey, text, geo, hashed, wildcard
- Properties: unique, partial, sparse, TTL, hidden
- Automatic unique index on
_id - Each index costs memory and write speed
Likely follow-up: How would you find indexes that are never used?
7.What is the aggregation pipeline? Walk me through what this query does.easy
The aggregation pipeline is MongoDB's framework for transforming and analyzing data on the server. Documents flow through an ordered list of stages, and each stage's output is the next stage's input, much like a Unix pipe.
This one returns the top five customers by paid revenue this year:
$matchfilters to paid orders since January. Placed first, it can use an index onstatusandplacedAt$groupmakes one document percustomerId, summingamountand counting orders with$sum: 1$sortand$limitkeep the five biggest totals; the optimizer merges them into a top-k sort$projectreshapes the output, renaming_idtocustomerId
Other common stages are
$lookupfor joins,$unwindto flatten arrays,$addFields/$set,$facet,$count, and$outor$mergeto write results to a collection.db.orders.aggregate([ { $match: { status: "paid", placedAt: { $gte: ISODate("2026-01-01") } } }, { $group: { _id: "$customerId", total: { $sum: "$amount" }, orders: { $sum: 1 } } }, { $sort: { total: -1 } }, { $limit: 5 }, { $project: { _id: 0, customerId: "$_id", total: 1, orders: 1 } } ]);What interviewers listen for- Ordered stages; output of one feeds the next
$matchearly so it can use indexes$groupwith accumulators like$sum,$avg$sortplus$limitbecomes a top-k sort$projectreshapes;$out/$mergepersist results
Likely follow-up: What is the memory limit for a pipeline stage, and what happens when you exceed it?
8.What is a replica set, and what happens when the primary goes down?mid
A replica set is a group of
mongodprocesses holding the same data. One member is the primary and accepts all writes; it records them in its oplog, a capped collection that secondaries copy and replay asynchronously. Members exchange heartbeats every two seconds.If secondaries can't reach the primary for longer than
electionTimeoutMillis(10 seconds by default), an eligible secondary calls an election, and the candidate that wins votes from a majority of voting members becomes primary. The manual says a new primary is typically elected in about 12 seconds or less with default settings. During the election no writes are accepted, but secondaries can keep serving reads if your read preference allows it. Drivers enable retryable writes by default, so many writes are retried once automatically.When the old primary comes back it rejoins as a secondary. Any of its writes that never reached a majority are rolled back, which is why
w: "majority"matters. A set can have up to 50 members, but at most 7 voting ones; use an odd number of voters.What interviewers listen for- One primary takes writes; secondaries replay its oplog
- Election after 10 s without primary contact (default)
- New primary needs a majority of voting members
- Retryable writes smooth over failover
- Non-majority writes on the old primary roll back
Likely follow-up: What is an arbiter, and why is it often discouraged?
9.What is sharding in MongoDB, and what are the components of a sharded cluster?mid
Sharding is horizontal partitioning: a collection's documents are spread across several shards so data size and write throughput can exceed what one server handles.
A sharded cluster has three parts:
- Shards: each is a replica set holding a subset of the data
mongosrouters: the app connects to these; they look up metadata and route each operation to the right shards- Config servers: a replica set storing cluster metadata, such as which ranges live where
The shard key, one or more fields chosen per collection, decides where each document goes. Data is split into chunks (contiguous shard key ranges, 128 MB by default), and the balancer migrates chunks to keep data evenly spread. Queries that include the shard key are targeted to specific shards; queries without it become scatter-gather across all shards.
Sharding adds real operational complexity, so I'd first make sure indexes, schema and vertical scaling on a replica set are exhausted.
What interviewers listen for- Horizontal partitioning of a collection across shards
- Shards are replica sets; mongos routes; config servers store metadata
- Shard key decides document placement
- Chunks of 128 MB default, moved by the balancer
- Targeted vs scatter-gather queries
Likely follow-up: What happens to a query that does not include the shard key?
10.Does MongoDB support ACID transactions? Show how you would transfer money between two accounts.easy
Yes, at two levels.
A write to a single document is always atomic, including all its embedded documents and arrays. With good modeling, that covers most use cases.
For changes across documents or collections, MongoDB has multi-document ACID transactions: on replica sets since 4.0 and on sharded clusters since 4.2. They don't work on a standalone server. Inside a transaction, reads come from a consistent snapshot (on sharded clusters, ask for read concern
"snapshot"to guarantee it across shards), commit is all-or-nothing, and no other client sees the changes until commit.In mongosh,
session.withTransaction()runs the callback, commits, and retries the whole transaction or the commit on transient errors. Throwing inside the callback aborts and rolls everything back. That's why the snippet checksmodifiedCount: the filterbalance: { $gte: 100 }makes the debit conditional, and if it matched nothing we throw instead of committing half a transfer.Transactions cost more than single-document writes, so the manual advises against using them as a substitute for good schema design.
const session = db.getMongo().startSession(); try { session.withTransaction(() => { const accounts = session.getDatabase("bank").accounts; const debit = accounts.updateOne( { _id: "A", balance: { $gte: 100 } }, { $inc: { balance: -100 } } ); if (debit.modifiedCount !== 1) throw new Error("insufficient funds"); accounts.updateOne({ _id: "B" }, { $inc: { balance: 100 } }); }); } finally { session.endSession(); }What interviewers listen for- Single-document writes are always atomic
- Multi-document: replica sets 4.0+, sharded 4.2+
- Not available on standalone servers
withTransactionretries transient errors; throw to abort- More expensive; prefer modeling first
Likely follow-up: What is a
TransientTransactionError?11.How do you order the fields of a compound index? Explain the ESR rule with this query.mid
A compound index is sorted by its first field, then by the second within each value of the first, and so on. Two consequences:
- Prefix rule:
{ a: 1, b: 1, c: 1 }supports queries ona,a + banda + b + c, but not onbalone - Sort direction: it supports
sort({ a: 1, b: 1 })and the exact reverse, but not a mixed{ a: 1, b: -1 }
The ESR guideline says: Equality fields first, then Sort fields, then Range fields. Equality first means everything after it stays in sorted order. Sort next lets MongoDB read keys already in
placedAtorder and skip an in-memory sort. Range last, because a range ontotalplaced beforeplacedAtwould break that ordering.Watch out:
$ne,$ninand$regexcount as range operators. And if the range filter is very selective, putting it before the sort field (ERS) can examine fewer keys, so confirm withexplain().db.orders.find({ status: "shipped", total: { $gt: 100 } }) .sort({ placedAt: -1 }); // ESR: Equality, then Sort, then Range db.orders.createIndex({ status: 1, placedAt: -1, total: 1 });What interviewers listen for- Prefix rule: leftmost fields must be used
- Equality, then Sort, then Range
- Sort field before range avoids an in-memory sort
$ne,$nin,$regexbehave as range predicates- Verify the choice with
explain()
Likely follow-up: How does
$infit into ESR when the query also sorts?- Prefix rule:
12.How do you use
explain()to check whether a query is efficient? What is the difference between COLLSCAN and IXSCAN?midexplain()shows the plan the query optimizer chose. It has three verbosity modes:queryPlanner(the default, plan only),executionStats(runs the winning plan and reports counters) andallPlansExecution(also partial stats for rejected plans).In
winningPlan, read the stage tree from the leaf up:- COLLSCAN: a full collection scan, reading every document. Fine for tiny collections, a red flag otherwise
- IXSCAN: scanning an index range, usually followed by FETCH to load the documents
- A SORT stage means an in-memory sort; if it's absent, the index supplied the order
Then compare
nReturned,totalKeysExaminedandtotalDocsExaminedinexecutionStats. The ideal is close to 1:1:1. If you examine 50,000 keys to return 20 documents, the index is weakly selective or its field order is wrong.totalDocsExamined: 0means a covered query.rejectedPlansshows which other indexes the planner considered.db.orders.find({ customerId: 42, status: "shipped" }) .sort({ placedAt: -1 }) .explain("executionStats"); // For pipelines: db.orders.explain("executionStats").aggregate([{ $match: { customerId: 42 } }]);What interviewers listen for- Modes: queryPlanner, executionStats, allPlansExecution
- COLLSCAN reads every document; IXSCAN reads an index range
- A SORT stage means an in-memory sort
- Compare nReturned vs keys and docs examined
- Aim for roughly 1:1:1
Likely follow-up: How would you find slow queries in production in the first place?
13.How do you choose a good shard key, and what happens if you pick a bad one?hard
I evaluate candidates on four properties:
- High cardinality: each distinct key value can live in only one chunk, so a key like
continentcaps you at seven chunks, and adding more shards won't help - Low frequency: if a few values dominate, their chunks grow huge, can't be split, and become hot spots or jumbo chunks
- Not monotonic: with timestamps or ObjectIds, every insert lands in the chunk with the
maxKeyupper bound, so one shard takes all the writes - Query isolation: the key should appear in most queries so
mongoscan target one shard instead of scatter-gather
A compound key often satisfies all four, for example
{ customerId: 1, orderDate: 1 }: customer gives spread and targeting, date adds cardinality. Hashing a monotonic field fixes write hot spots but makes range queries broadcast.A bad key used to be permanent. Now you can refine it by adding suffix fields with
refineCollectionShardKey, or fully reshard withreshardCollectionsince 5.0, though that's a heavy operation. From 7.0,analyzeShardKeyreports cardinality, frequency and monotonicity before you commit.What interviewers listen for- High cardinality and low frequency
- Avoid monotonically increasing keys for ranged sharding
- Include the key in common queries to target shards
- Compound keys balance spread and targeting
- Refine or reshard (5.0+) if you got it wrong
Likely follow-up: Why is
{ createdAt: 1 }a poor shard key for an events collection?- High cardinality: each distinct key value can live in only one chunk, so a key like
14.Are updates to a single document atomic in MongoDB? How do you avoid lost updates under concurrency?easy
Yes. Any write to one document is atomic, even if it changes several fields, embedded documents and arrays at once. Other readers see either the old document or the new one, never half of it.
The first pattern is a classic lost update: two clients both read
stock: 5, both write3, and one sale vanishes. The fix is to let the server do the math with update operators like$inc,$set,$pushor$min, and to put the precondition in the filter. The second statement only matches if at least two units remain, so checkingmodifiedCounttells you whether the reservation succeeded.This atomicity is per document.
updateManyapplies atomically to each matched document but not to the set as a whole; other operations can interleave. If several documents must change together, either model them into one document or use a multi-document transaction.// Race-prone: read, compute in the app, write back const p = db.products.findOne({ _id: 1 }); db.products.updateOne({ _id: 1 }, { $set: { stock: p.stock - 2 } }); // Atomic: the check and the change are one document write db.products.updateOne( { _id: 1, stock: { $gte: 2 } }, { $inc: { stock: -2 }, $set: { updatedAt: new Date() } } );What interviewers listen for- Single-document writes are atomic, including nested data
- Use
$inc/$setinstead of read-modify-write - Put preconditions in the filter; check
modifiedCount updateManyis atomic per document, not overall- Cross-document atomicity needs a transaction
Likely follow-up: How would you implement optimistic locking with a version field?
15.How do you write queries with
find()? Explain the common query operators used here.easyfind(filter, projection)returns a cursor over matching documents. Separate conditions in the filter are implicitly ANDed.The main operator groups:
- Comparison:
$eq,$ne,$gt,$gte,$lt,$lte,$in,$nin - Logical:
$and,$or,$nor,$not - Element:
$existschecks presence,$typechecks the BSON type - Evaluation:
$regex, and$exprto use aggregation expressions, such as comparing two fields of the same document - Array:
$all(contains every value),$size,$elemMatch
So this finds working-age users in India or the US, on the pro plan or with more than 100 credits, not soft-deleted, tagged both beta and mobile, who have spent more than their budget.
For performance: equality and ranges use indexes well;
$ne,$ninand unanchored regexes are rarely selective; a case-sensitive prefix regex like/^abc/can use an index range.db.users.find({ age: { $gte: 18, $lt: 65 }, country: { $in: ["IN", "US"] }, $or: [{ plan: "pro" }, { credits: { $gt: 100 } }], deletedAt: { $exists: false }, tags: { $all: ["beta", "mobile"] }, $expr: { $gt: ["$spent", "$budget"] } });What interviewers listen for- Filter fields are implicitly ANDed
- Comparison, logical, element, evaluation and array operators
$exprcompares fields within a document$allneeds every value;$inneeds any- Anchored prefix regex can use an index
Likely follow-up: What is the difference between
$inand$all?- Comparison:
16.MongoDB has no JOIN keyword. How do you combine data from two collections?mid
There are three options.
First, avoid the join by embedding the data, or by copying a few fields you always need (the extended reference pattern).
Second,
$lookupin an aggregation, which performs a left outer join. For each input document it finds matches in thefromcollection and puts them in an array field named byas. An order with no customer gets an empty array.$unwindthen turns the one-element array into an embedded object, but it drops documents whose array is empty unless you setpreserveNullAndEmptyArrays: true. The pipeline form, withletandpipeline, supports extra conditions and sub-queries on the foreign side.Third, application-level joins: query orders, collect the IDs, then one
$inquery for customers. That's what Mongoose'spopulate()does.Whichever you choose, index the foreign field. Without one, each lookup can scan the whole foreign collection. If nearly every read needs a join, that's a sign to revisit the schema.
db.orders.aggregate([ { $match: { status: "paid" } }, { $lookup: { from: "customers", localField: "customerId", foreignField: "_id", as: "customer" } }, { $unwind: "$customer" }, { $project: { total: 1, "customer.name": 1, "customer.email": 1 } } ]);What interviewers listen for- Embed or duplicate fields to avoid joins
$lookupis a left outer join into an array$unwinddrops non-matches unless preserveNullAndEmptyArrays- Index the
foreignField - App-level joins with
$inalso work
Likely follow-up: Can the
fromcollection of a$lookupbe sharded?17.What is write concern? Explain
w,jandwtimeout, and what the default is.midWrite concern is the level of acknowledgment you ask for before a write is reported as successful. It trades latency for durability.
w: 0: fire and forget, no acknowledgmentw: 1: the primary applied it. It can be rolled back if the primary fails before replicatingw: "majority": a majority of data-bearing voting members have it. It survives failoverw: <n>or a custom tag name for specific member counts or data centersj: true: wait until the write is in the on-disk journalwtimeout: how long to wait for the concern, in milliseconds
Since MongoDB 5.0 the implicit default is
w: "majority". The exception: with arbiters, if data-bearing voters don't exceed the voting majority, for example primary-secondary-arbiter, the default isw: 1.Gotcha: a
wtimeouterror doesn't undo anything. The write may already be applied and may still replicate; you only learn the guarantee wasn't confirmed in time, so the retry logic must be idempotent.What interviewers listen for- Acknowledgment level: latency vs durability
w: 1can roll back; majority survives failoverj: truewaits for the on-disk journal- Default
w: "majority"since 5.0, except some arbiter setups wtimeouterrors do not undo the write
Likely follow-up: Why does a primary-secondary-arbiter set default to
w: 1?18.What is the difference between read concern and read preference?mid
They answer different questions.
Read preference decides which member serves the read:
primary(default),primaryPreferred,secondary,secondaryPreferred,nearest- optionally refined with tag sets (for example, "same region") and
maxStalenessSecondsto skip lagging secondaries
Read concern decides which version of the data is returned, meaning its consistency and durability:
"local"(default): the member's latest data, which might later roll back"available": similar, lowest latency, with sharding caveats"majority": only data acknowledged by a majority, so it won't be rolled back"linearizable": reflects every majority write completed before the read began"snapshot": a consistent point-in-time view, used in transactions
Replication is asynchronous, so reading from secondaries can return stale data even with
"majority". Use secondaries for analytics or geo-local reads that tolerate lag, not for read-your-own-writes. Transactions that read must useprimary.What interviewers listen for- Read preference: which member serves the read
- Read concern: which data version, consistency and durability
- Defaults:
primaryand"local" - Majority reads never see rolled-back data
- Secondary reads can be stale
Likely follow-up: How can a client read its own writes when reading from secondaries?
19.What MongoDB schema design patterns do you know, and when would you use each?mid
MongoDB documents a catalog of reusable patterns. The ones I reach for most:
- Bucket: group many small items, like sensor readings per device per hour, into one document with an array. Fewer documents and smaller indexes
- Subset: keep only the hot part in the main document, such as the 10 latest reviews, and the rest in another collection
- Extended reference: copy a few frequently read fields of a referenced document, like a customer's name on each order, to avoid joins
- Computed: precompute values such as totals or averages on write instead of on every read
- Outlier: design for the typical case and flag the rare huge document, then overflow its extras elsewhere
- Attribute: turn many similar fields into an array of key/value pairs so one index covers them
- Polymorphic: different shapes in one collection, sharing common fields
- Schema versioning: a
schemaVersionfield lets old and new shapes coexist during a migration
Others include approximation, tree, pre-allocation and document versioning. Each trades some write complexity or duplication for faster reads.
What interviewers listen for- Bucket, subset, extended reference, computed, outlier
- Attribute, polymorphic, schema versioning
- Patterns are driven by access patterns
- Trade duplication or write work for read speed
- Patterns can be combined
Likely follow-up: How would you model product reviews for a site where a few products have 100,000 reviews?
20.How do you model one-to-few, one-to-many and one-to-squillions relationships?mid
The size of the "N" drives the choice, along with whether the children stand alone.
- One-to-few: embed the children as an array of subdocuments. A person's handful of addresses lives inside the person document, read and written together
- One-to-many, up to a few hundred or thousand: store an array of references in the parent, such as a product holding the ObjectIds of its parts. Children are their own documents, and the same approach gives you many-to-many without a join table
- One-to-squillions: put a reference to the parent in each child, like log messages each carrying
hostId. An array would grow without bound and hit 16 MiB
The rules of thumb behind this: favor embedding unless there's a compelling reason not to; needing to access a child on its own is such a reason; arrays should not grow without bound; don't be afraid of application-level joins if you index properly; and weigh read-to-write ratio before denormalizing a field. Everything depends on the access patterns.
What interviewers listen for- One-to-few: embed subdocuments
- One-to-many: array of child references in the parent
- One-to-squillions: parent reference in each child
- Arrays must not grow without bound
- Denormalize fields that are read often, updated rarely
Likely follow-up: How would you get the 20 most recent log messages for a host efficiently?
21.What is the maximum document size in MongoDB, and what do you do if your data needs more?easy
A BSON document can be at most 16 MiB. The limit exists so a single document can't consume an excessive amount of RAM, or bandwidth when it's sent over the network.
In practice, a document getting anywhere near that size is a modeling smell, usually an array that keeps growing: comments, events, log lines. Every read and update of that document drags the whole thing through the cache, so performance suffers long before the hard limit. Fixes:
- Reference the growing children from their own collection instead of embedding them
- Use the bucket pattern to cap items per document
- Use the subset or outlier pattern to keep only the hot part inline
For genuinely large binary content like videos or PDFs, use GridFS, which splits a file into chunks (255 kB by default) stored in
fs.chunkswith metadata infs.files. Many teams instead put files in object storage and keep only the URL and metadata in MongoDB.What interviewers listen for- 16 MiB max BSON document size
- Near-limit documents signal unbounded arrays
- Reference, bucket, subset or outlier patterns
- GridFS splits large files into chunks
- Object storage plus metadata is a common alternative
Likely follow-up: When would you choose GridFS over object storage?
22.What is a covered query, and how do you confirm one?mid
A covered query is answered entirely from an index, without reading any documents. That's fast because index keys are small and usually in RAM.
For a query to be covered:
- every field in the filter is in the same index
- every field returned is in that index
- no filter field is compared to
null
The classic trap is
_id: it's returned by default, so unless it's part of the index, the projection must say_id: 0, as above. Other limits: a multikey index can't cover a query on its array field, and on a sharded collection the index must contain the shard key for a query throughmongosto be covered.To confirm, run
explain("executionStats"). A covered plan has no FETCH stage, often showsPROJECTION_COVERED, and reportstotalDocsExamined: 0whiletotalKeysExaminedis greater than zero.db.users.createIndex({ email: 1, name: 1 }); db.users.find( { email: "ana@example.com" }, { _id: 0, email: 1, name: 1 } ).explain("executionStats"); // PROJECTION_COVERED over IXSCAN, totalDocsExamined: 0What interviewers listen for- Filter and returned fields all in one index
- Exclude
_idunless it is indexed - No FETCH stage; totalDocsExamined is 0
- Multikey indexes cannot cover their array fields
- Sharded: index must include the shard key
Likely follow-up: Why might adding one field to a projection make a fast query slow?
23.What is an upsert? What document does this create when nothing matches, and how do you make it safe under concurrency?easy
An upsert is "update if it exists, insert if it doesn't", enabled with
{ upsert: true }onupdateOne,updateMany,replaceOneorfindOneAndUpdate.If a document matches, it just gets
viewsincremented. If none matches, MongoDB builds a new document from the equality conditions in the filter (pageandday), applies the update operators, soviewsbecomes 1, and applies$setOnInsert, which only runs on insert. Range conditions like$gtare not copied into the new document, and an_idis generated if none is given. The result reportsupsertedId.The concurrency trap: two clients can both find no match and both insert, creating duplicates. The fix is a unique index on the filter fields. Then exactly one insert wins. When the filter is equality on exactly the unique index's fields, the server retries the losers as updates; otherwise they get a duplicate key error that the app should retry.
db.pageViews.createIndex({ page: 1, day: 1 }, { unique: true }); db.pageViews.updateOne( { page: "/pricing", day: "2026-09-28" }, { $inc: { views: 1 }, $setOnInsert: { firstSeenAt: new Date() } }, { upsert: true } );What interviewers listen for- Update if found, otherwise insert
- New doc built from filter equality fields plus operators
$setOnInsertapplies only on insert- Concurrent upserts can duplicate without a unique index
- Unique index on filter fields prevents duplicates
Likely follow-up: What does
upsertedIdcontain when the upsert updated an existing document?24.Explain
$push,$addToSetand$pull. How would you keep an array capped at the 50 newest items?easyAll three change arrays atomically inside one document, without rewriting the array from the app.
$pushappends a value, creating the array if the field is missing. Pushing an array appends it as one element; use$eachto add several$addToSetadds a value only if it isn't already present, giving set semantics. For subdocuments, the comparison is on the whole value, including field order$pullremoves every element equal to a value or matching a condition, such as all comments byspammer
$pushalso takes modifiers that all require$each:$positionto insert at an index,$sortto order the array, and$sliceto trim it. They run in a fixed order: insert, then sort, then slice. The feed update above keeps only the 50 newest items, a simple way to stop an array from growing without bound.$popremoves just the first or last element.db.posts.updateOne({ _id: 1 }, { $push: { comments: { user: "ana", text: "Nice" } } }); db.posts.updateOne({ _id: 1 }, { $addToSet: { tags: { $each: ["mongodb", "nosql"] } } }); db.posts.updateOne({ _id: 1 }, { $pull: { tags: "draft" } }); db.posts.updateOne({ _id: 1 }, { $pull: { comments: { user: "spammer" } } }); db.feeds.updateOne( { _id: 7 }, { $push: { items: { $each: [{ at: new Date() }], $sort: { at: -1 }, $slice: 50 } } } );What interviewers listen for$pushappends;$eachadds several values$addToSetskips values already present$pullremoves all matching elements$slicewith$eachcaps array length- Modifiers apply: position, sort, then slice
Likely follow-up: How do you update one specific element inside an array?
25.What is a projection in MongoDB, and what rules apply when writing one?easy
A projection is the second argument to
find()(or a$projectstage) that chooses which fields come back.db.users.find({}, { name: 1, email: 1 })returns onlyname,emailand_id.The rules:
_idis included by default; add_id: 0to drop it- You can't mix inclusion (
1) and exclusion (0) in one projection, except for_id - Use dot notation for nested fields, like
"address.city": 1 - Arrays have projection operators:
$slicefor the first or last N elements,$elemMatchfor the first matching element, and positional$for the element matched by the filter
Projections cut network transfer and client memory, and they're required for covered queries. One nuance: unless the query is covered, the server still loads the full document into its cache. Projection shrinks the response, not the disk read, so it isn't a fix for bloated documents.
What interviewers listen for- Selects which fields are returned
_idincluded unless_id: 0- No mixing include and exclude, except
_id $slice,$elemMatch,$project array elements- Reduces transfer; enables covered queries
Likely follow-up: How would you return only the last 5 comments of a post?
26.Why do these two queries return different results for the same document? When do you need
$elemMatch?midWhen you query array fields with separate conditions, each condition can be satisfied by a different element. In the first query, the element with
product: "xyz"satisfies the first condition and the element withscore: 10satisfies the second, so the document matches, even though no single result is "xyz with a score of at least 8".$elemMatchrequires one element to satisfy all conditions together. The xyz element has a score of 5, so the second query correctly finds nothing.The same thing happens with scalar arrays:
{ scores: { $gt: 80, $lt: 85 } }matches[90, 70], because 90 is above 80 and 70 is below 85, while$elemMatch: { $gt: 80, $lt: 85 }needs a value actually between them.So I use
$elemMatchwhenever I have two or more conditions on the same array element. For one condition, plain dot notation is enough. There's also a projection$elemMatchthat returns only the first matching element.// { _id: 1, results: [ { product: "abc", score: 10 }, { product: "xyz", score: 5 } ] } db.survey.find({ "results.product": "xyz", "results.score": { $gte: 8 } }); // matches _id 1 db.survey.find({ results: { $elemMatch: { product: "xyz", score: { $gte: 8 } } } }); // no matchWhat interviewers listen for- Separate conditions can match different elements
$elemMatchneeds one element to match all- Applies to subdocument and scalar arrays
- Use it for two or more conditions on one element
- Projection
$elemMatchreturns first matching element
Likely follow-up: How does a multikey index handle an
$elemMatchquery with a range?27.What is a TTL index, and how precise is its expiry?easy
A TTL index is a single-field index with
expireAfterSeconds. A background task deletes documents once that many seconds have passed since the date in the indexed field. It's ideal for sessions, tokens, caches and temporary logs.Two styles: a fixed lifetime (
lastSeenplus 30 minutes), orexpireAfterSeconds: 0with a per-documentexpireAtdate, so each document expires at its own time.Things to know:
- It's not precise. The TTL monitor runs every 60 seconds, and under heavy load deletes can lag further, so filter on the date too if stale data mustn't be seen
- The field must be a BSON Date or an array of dates; with an array, the earliest date counts. Documents where it's missing or not a date never expire
- It must be a single-field index: compound indexes ignore
expireAfterSeconds, and_idcan't be a TTL index - On a replica set only the primary deletes; the deletes replicate like any other write
// Delete sessions 30 minutes after lastSeen db.sessions.createIndex({ lastSeen: 1 }, { expireAfterSeconds: 1800 }); // Or expire each document at its own time db.invites.createIndex({ expireAt: 1 }, { expireAfterSeconds: 0 }); db.invites.insertOne({ code: "X1", expireAt: new Date("2026-10-01T00:00:00Z") });What interviewers listen for- Single-field index with
expireAfterSeconds - Field must be a Date or array of dates
- Monitor runs every 60 s; expiry is approximate
expireAfterSeconds: 0for per-document expiry times- Missing or non-date fields never expire
Likely follow-up: How do you change the TTL of an existing index?
28.How do you paginate results in MongoDB? Why is
skip()a problem on large collections?midThe simple approach is
skip((page - 1) * size).limit(size). It works for small data, but the server still has to walk past every skipped entry, so page 5,000 costs far more than page 1. Results can also shift or repeat while users page, because inserts and deletes move the offsets.For large or infinite-scroll lists I use range-based (keyset or cursor) pagination. Sort on an indexed key, remember the last item returned, and ask for items after it. Each page is an index seek plus
limit, so it costs the same at any depth, and concurrent inserts don't cause duplicates.The sort must be unique to be stable. MongoDB's sort isn't stable for equal keys, so add
_idas a tie-breaker, as in the snippet, backed by an index like{ status: 1, createdAt: -1, _id: -1 }. The trade-off: you can't jump straight to page 37, which most UIs don't need anyway. The client gets an opaque cursor encodingcreatedAtand_id.// First page db.posts.find({ status: "published" }) .sort({ createdAt: -1, _id: -1 }).limit(20); // Next page: continue after the last post you returned db.posts.find({ status: "published", $or: [ { createdAt: { $lt: last.createdAt } }, { createdAt: last.createdAt, _id: { $lt: last._id } } ] }).sort({ createdAt: -1, _id: -1 }).limit(20);What interviewers listen forskipwalks past skipped entries: cost grows with depth- Offsets drift as data changes
- Range pagination: filter after the last seen key
- Unique sort with an
_idtie-breaker - Back it with a matching compound index
Likely follow-up: How would you show a total page count without an expensive count on every request?
29.Why does the second insert fail here, and how would a partial index fix it?mid
A unique index stores a null key for documents where the field is missing, and uniqueness applies to that null like any other value. The first user without an email takes the null slot, and the second collides with it.
A partial index only indexes documents that match
partialFilterExpression. Combined withunique: true, uniqueness is enforced only among documents that have an email, so any number of users can omit it. The filter supports equality,$exists: true,$gt/$lt-style comparisons,$type,$and,$orand$in.Partial indexes also shrink index size, for example indexing only
{ status: "active" }orders. The catch: the planner uses one only when the query's filter implies the partial filter. A query onemailalone may not qualify; include the condition, such asemail: { $exists: true }.Sparse indexes do something similar, skipping documents missing the field, but the manual recommends partial indexes because they're more expressive. Note that creating a unique index fails if existing data already has duplicates.
db.users.createIndex({ email: 1 }, { unique: true }); db.users.insertOne({ name: "A" }); // ok db.users.insertOne({ name: "B" }); // E11000 duplicate key error // Better: drop that index and enforce uniqueness only where email exists db.users.createIndex( { email: 1 }, { unique: true, partialFilterExpression: { email: { $exists: true } } } );What interviewers listen for- Missing field is indexed as null
- Unique index allows only one null
- Partial index indexes only matching documents
- Query must imply the partial filter to use it
- Prefer partial over sparse indexes
Likely follow-up: How would you enforce case-insensitive uniqueness of emails?
30.How do
$unwindand$groupwork together? Explain this best-sellers report.mid$unwinddeconstructs an array: an order with three items becomes three documents, each holding one item initems. Documents where the array is missing or empty are dropped unless you passpreserveNullAndEmptyArrays: true.$groupthen collapses documents that share a group key, given as_id. Here each SKU becomes one output document, using accumulators:$sumof quantities,$sumof an expression for revenue, and$avg. Others include$min,$max,$first,$last,$pushand$addToSet. Grouping by_id: nullaggregates the whole input into one document.Things interviewers look for:
$groupoutput order is not guaranteed, hence the explicit$sort$groupis blocking and each stage is limited to 100 MB of RAM; beyond that it spills to disk when disk use is allowed, which is the default since 6.0, or errors if it isn't$unwindmultiplies document counts, so$matchfirst to shrink the input
// orders: { _id, status, items: [ { sku, qty, price } ] } db.orders.aggregate([ { $match: { status: "paid" } }, { $unwind: "$items" }, { $group: { _id: "$items.sku", unitsSold: { $sum: "$items.qty" }, revenue: { $sum: { $multiply: ["$items.qty", "$items.price"] } }, avgQty: { $avg: "$items.qty" } } }, { $sort: { revenue: -1 } } ]);What interviewers listen for$unwindemits one document per array element- Drops empty arrays unless preserveNullAndEmptyArrays
$groupkey is_id;nullgroups everything- Accumulators:
$sum,$avg,$push,$addToSet - Group output is unordered; 100 MB stage limit
Likely follow-up: How could you get the top-selling SKU per category in one pipeline?
31.What is Mongoose, and what are schemas and models?easy
Mongoose is an ODM (object data modeling library) for Node.js that sits on top of the official MongoDB driver. It adds structure to MongoDB's flexible documents from the application side.
- A schema declares fields, types, defaults and validators, plus options like
timestamps, which addscreatedAtandupdatedAt - A model is compiled from a schema with
mongoose.model(). It's the class you query with (User.find,User.create), and it maps to a collection, by default the lowercased plural of the name, sousers - A document is an instance of a model, with change tracking and
save()
Mongoose casts values (the string
"42"becomes a number), validates before saving, and adds middleware, virtuals andpopulate().Gotchas:
unique: trueis not a validator; it only asks Mongoose to build a unique index. Also, update validators don't run onupdateOneand similar unless you passrunValidators: true. Schemas are app-side only; other clients can still write anything unless you also add server-side$jsonSchemavalidation.const mongoose = require("mongoose"); const userSchema = new mongoose.Schema({ email: { type: String, required: true, unique: true, lowercase: true }, name: String, age: { type: Number, min: 0 }, roles: { type: [String], default: ["user"] } }, { timestamps: true }); const User = mongoose.model("User", userSchema); // collection "users" await mongoose.connect(process.env.MONGODB_URI); await User.create({ email: "Ana@Example.com", name: "Ana" });What interviewers listen for- ODM for Node.js on top of the driver
- Schema defines types, defaults, validators
- Model maps to a collection and runs queries
- Casting, validation, middleware, populate
uniquebuilds an index; it is not a validator
Likely follow-up: When would you skip Mongoose and use the native driver?
- A schema declares fields, types, defaults and validators, plus options like
32.How does Mongoose
populate()work, and how is it different from$lookup? When would you uselean()?midYou store a reference as
{ type: Schema.Types.ObjectId, ref: "User" }, thenPost.find().populate("author", "name")replaces eachauthorID with the referenced user document.Under the hood, populate is not a server-side join. Mongoose runs the main query, collects the referenced IDs, and runs a separate query per populated path, essentially
User.find({ _id: { $in: ids } }), then stitches the results together in Node. That's one extra round trip per path, not one per document, so it avoids N+1. But deep or nested populates add up, and you can't filter the parent by a populated field in the same query. For that, use an aggregation with$lookup, which does the join inside the database.Virtual populate covers the reverse side, like a user's posts, via
localFieldandforeignField, without storing an array of IDs on the parent.lean()returns plain JavaScript objects instead of full Mongoose documents: no change tracking, getters, virtuals orsave(), but much less memory and faster. I use it for read-only endpoints.What interviewers listen forrefplus ObjectId declares the relationship- Populate runs a separate
$inquery per path $lookupjoins inside the server- Virtual populate for the reverse relationship
lean()returns plain objects for fast reads
Likely follow-up: How would you sort posts by their author's name?
33.A MongoDB-backed API has become slow. How do you find and fix the problem?hard
I work from evidence, not guesses.
- Find the slow operations. The slow query log records operations over
slowms, 100 ms by default, and the profiler (db.setProfilingLevel(1)) stores them insystem.profile.db.currentOp()shows what's running now. On Atlas, the Query Profiler and Performance Advisor do this for you. - Explain them. Look for COLLSCAN, in-memory SORT stages and a poor ratio of keys examined to documents returned. Fix with compound indexes following ESR, and covered queries where possible.
- Check memory. If the working set (hot data plus indexes) doesn't fit in the WiredTiger cache, reads hit disk. Look at cache eviction and page faults.
- Check the schema. Bloated documents, unbounded arrays and
$lookupon every read are design problems that indexes can't fix. - Check the client. Use one pooled client per process, project only needed fields, and batch writes with
bulkWrite. - Check for redundant indexes with
$indexStats; they slow writes.
Only after that would I scale hardware or shard.
What interviewers listen for- Slow query log and profiler first
explain(): COLLSCAN, SORT, keys vs docs examined- Working set must fit in the WiredTiger cache
- Schema fixes beat adding hardware
- Pooled clients, projections, bulk writes
Likely follow-up: What does a high ratio of documents examined to returned tell you?
- Find the slow operations. The slow query log records operations over
34.What are the most common MongoDB schema and indexing anti-patterns?mid
The ones MongoDB itself warns about, and I see most:
- Unbounded arrays: embedding comments, events or followers that grow forever. Documents bloat toward 16 MiB, updates rewrite big documents, and multikey indexes explode
- Too many indexes: every index slows writes and competes for RAM. Drop unused or redundant ones;
{ a: 1 }is redundant if{ a: 1, b: 1 }exists - Bloated documents: storing rarely used large fields with hot data, so every read pulls them into cache. Split them out with the subset pattern
- Separating data that's accessed together: normalizing like SQL, then
$lookupon every request - Massive numbers of collections, such as one per user or per day. Each collection and index has overhead
- Case-insensitive regex queries like
/^ana$/i, which can't use an index efficiently. Use a collation-based case-insensitive index instead
Also watch for deep
skip()pagination and read-modify-write updates instead of atomic operators.What interviewers listen for- Unbounded arrays
- Unnecessary or redundant indexes
- Bloated documents and separated hot data
- Too many collections
- Case-insensitive regex without a collation index
Likely follow-up: How do you find indexes that are never used?
35.What are the costs and limits of multi-document transactions, and how do you use them safely in production?hard
Transactions work, but they aren't free:
- Time limit: a transaction must finish within 60 seconds by default (
transactionLifetimeLimitSeconds), or it's aborted - Lock waits: it waits only 5 ms by default to acquire a lock before aborting
- Write conflicts: if another operation modifies a document the transaction is writing, the transaction fails with a
TransientTransactionErrorlabel and must be retried from the start - Cache pressure: open transactions pin old snapshots in the WiredTiger cache, so long ones hurt everyone
- Sharded clusters: a transaction spanning shards uses a two-phase commit, which adds round trips and latency
Safe usage:
- Keep transactions short and small; do reads and computation before starting
- Use the callback API (
withTransaction) so transient errors andUnknownTransactionCommitResultare retried - Make sure queries inside are indexed
- Use
w: "majority"for commits you can't lose - Prefer a single-document design when one is possible
A standalone server doesn't support transactions, and they can't write to capped collections or the
admin,configandlocaldatabases.What interviewers listen for- 60 s default lifetime; 5 ms lock wait
- Write conflicts raise TransientTransactionError; retry
- Long transactions pressure the WiredTiger cache
- Cross-shard commits use two-phase commit
- Keep them short; use
withTransactionretries
Likely follow-up: What should you do when a commit fails with
UnknownTransactionCommitResult?- Time limit: a transaction must finish within 60 seconds by default (
36.Compare hashed and ranged sharding. Which would you use for an events collection keyed by timestamp?hard
Ranged sharding splits data into contiguous ranges of the shard key. Documents with nearby keys live together, so range queries are targeted to few shards, and zones can pin ranges to regions. The weakness is a monotonically increasing key like a timestamp or ObjectId: every new document falls in the top range, so one shard absorbs all inserts while the balancer tries to catch up.
Hashed sharding shards on a hash of one field. Nearby values scatter, so writes spread evenly even for monotonic keys, and equality queries are still targeted. The cost: range queries become broadcast to all shards, and hashed fields can't be arrays.
For events keyed by time, I'd avoid a plain ranged
{ ts: 1 }. Options:{ deviceId: 1, ts: 1 }if most queries are per device: spread plus targeted, time-ordered reads{ _id: "hashed" }if writes dominate and queries rarely filter by time range- A compound key with one hashed field, like
{ tenantId: 1, ts: "hashed" }, to keep tenant-targeted queries while spreading inserts
What interviewers listen for- Ranged: targeted range queries, monotonic hot spots
- Hashed: even writes, range queries broadcast
- Equality queries target a shard in both
- Compound shard keys can combine both properties
- Pick based on query and write patterns
Likely follow-up: How do zones work with ranged sharding?
37.MongoDB is "schemaless". How do you enforce a schema on the server anyway?mid
MongoDB is really flexible-schema: by default documents in a collection can differ, but you can attach a validator that the server checks on every insert and update, whichever client writes.
The usual tool is
$jsonSchema, based on JSON Schema draft 4 with MongoDB extensions.bsonTypechecks BSON types likeint,dateorobjectId;requiredlists mandatory fields;propertiessets per-field rules likeminimum,patternandenum. You can also use query operators in the validator. Add or change rules on an existing collection withcollMod.Two settings control strictness:
validationLevel:strict(default) checks all inserts and updates;moderatedoesn't enforce rules on updates to existing documents that were already invalid, which helps when retrofitting rules onto old datavalidationAction:error(default) rejects the write;warnallows it and logs the violation; recent versions adderrorAndLog
Existing documents aren't re-checked when you add a validator. Users with the right privilege can pass
bypassDocumentValidationfor migrations.db.createCollection("users", { validator: { $jsonSchema: { bsonType: "object", required: ["email", "createdAt"], properties: { email: { bsonType: "string", pattern: "^.+@.+$" }, age: { bsonType: "int", minimum: 0 }, createdAt: { bsonType: "date" } } } }, validationLevel: "moderate", validationAction: "error" });What interviewers listen for- Validator enforced by the server for all clients
$jsonSchemawithbsonType,required,properties- validationLevel strict (default) or moderate
- validationAction error (default) or warn
- Apply to existing collections with
collMod
Likely follow-up: How would you roll out a new required field without breaking old documents?
38.What are change streams, and how do you make a consumer survive restarts?mid
Change streams give applications a real-time feed of data changes, built on the oplog, without tailing it by hand. You call
watch()on a collection, a database or the whole deployment, and can pass an aggregation pipeline to filter or reshape events. They require a replica set or sharded cluster, and they only report changes committed to a majority, so you never see events that later roll back.Each event has an
operationType(insert,update,replace,delete,invalidateand others) anddocumentKey. Updates carry only the changed fields unless you setfullDocument: "updateLookup", which fetches the current version of the document; that may include later changes. Since 6.0 you can enable pre- and post-images on a collection instead.For restarts, every event's
_idis a resume token. Persist it after processing, and reopen withresumeAfter(orstartAfter, which also works after an invalidate). Resuming only works while that point is still in the oplog, so size the oplog for your longest expected outage.Typical uses: cache invalidation, syncing a search index, notifications and event-driven microservices.
const cs = db.orders.watch( [{ $match: { operationType: { $in: ["insert", "update"] } } }], { fullDocument: "updateLookup" } ); while (!cs.isClosed()) { if (cs.hasNext()) { const ev = cs.next(); printjson({ op: ev.operationType, id: ev.documentKey._id }); saveResumeToken(ev._id); // persist it with your own storage } }What interviewers listen for- Real-time change feed built on the oplog
- Needs a replica set or sharded cluster
- Only majority-committed changes are reported
- Resume with the saved token via
resumeAfter - Resumable only within the oplog window
Likely follow-up: How would you guarantee exactly-once processing of change events?
39.What is a capped collection, and when would you use one instead of a TTL index?easy
A capped collection is a fixed-size collection that behaves like a circular buffer. You create it with
db.createCollection("log", { capped: true, size: 100000 }), wheresizeis in bytes and an optionalmaxlimits the document count. When it's full, the oldest documents are removed automatically to make room.It keeps documents in insertion order, so reading the newest entries is cheap, and it supports tailable cursors that stay open and stream new documents as they arrive, like
tail -f. MongoDB's own replication oplog is a capped collection.Restrictions: capped collections can't be sharded, can't be written inside transactions, can't be the target of
$out, and shouldn't be updated in ways that grow documents. With concurrent writers, insertion order isn't guaranteed.These days the manual generally recommends TTL indexes instead. Capped collections serialize writes, so they perform worse under concurrency, and TTL lets you expire by age rather than by total size. I'd choose capped only when "keep the last N megabytes" is exactly the requirement.
What interviewers listen for- Fixed size; oldest documents removed automatically
- Insertion order and tailable cursors
- The oplog is a capped collection
- Cannot shard or write in transactions
- TTL indexes are usually the better choice
Likely follow-up: What is a tailable cursor?
40.How do you implement text search in MongoDB? What are the limits of a text index?mid
On self-managed MongoDB, create a text index, for example
createIndex({ title: "text", body: "text" }, { weights: { title: 5 } }), then query with{ $text: { $search: "coffee grinder" } }.The text index tokenizes strings, removes stop words and applies stemming for the chosen language, so "grinders" matches "grinder". The search string supports quoted phrases and
-termto exclude. You can project and sort by relevance with{ $meta: "textScore" }.Limits:
- Only one text index per collection, though it can cover many fields, or all strings with
"$**" - No fuzzy matching, typo tolerance, autocomplete or partial-word matching
- Relevance scoring is basic and hard to tune
- Unanchored
$regexisn't an alternative; it can't use an index efficiently
For real search features, MongoDB recommends MongoDB Search (Atlas Search, built on Apache Lucene) with the
$searchstage: fuzzy matching, autocomplete, facets and custom analyzers, plus Vector Search for semantic similarity.What interviewers listen for- Text index plus
$textwith$search - Stemming and stop words per language
- Sort by
$meta: "textScore" - One text index per collection; no fuzzy or autocomplete
- Atlas Search (Lucene) for real search features
Likely follow-up: How would you build search-as-you-type?
- Only one text index per collection, though it can cover many fields, or all strings with
41.How do multikey indexes work, and what restrictions and performance traps come with them?hard
When you index a field that holds an array in any document, MongoDB automatically makes the index multikey: it stores one index entry per array element, all pointing to the same document. That's what makes
{ tags: "mongodb" }fast.Restrictions:
- In a compound multikey index, each document can have at most one indexed field that is an array.
{ tags: 1, categories: 1 }is fine until one document has both as arrays, and then that insert fails - A multikey index can't be a shard key index, and hashed indexes can't be multikey
- It can't cover a query that returns the array field
Performance traps:
- Index size grows with array length. A document with 5,000 tags adds 5,000 keys, and every push updates the index. That's another reason to keep arrays bounded
- Bounds: for
{ scores: { $gt: 80, $lt: 85 } }without$elemMatch, different elements may satisfy each condition, so MongoDB can't simply intersect the bounds; it may scan a wider range. With$elemMatch, it can combine them into one tight range
What interviewers listen for- One index entry per array element
- Compound: at most one array field per document
- Cannot be a shard key; hashed cannot be multikey
- Large arrays inflate index size and write cost
$elemMatchallows tighter index bounds
Likely follow-up: How can you tell from
explain()that an index is multikey?- In a compound multikey index, each document can have at most one indexed field that is an array.
42.How do you optimize an aggregation pipeline? What does the optimizer already do for you?hard
The most important rule: only the start of a pipeline can use indexes, so put
$match, and$sortwhen possible, first, and filter as early as you can. On a sharded collection, a leading$matchon the shard key also targets fewer shards.The optimizer rewrites some things automatically:
- moves
$matchfilters ahead of$project,$addFields/$setor$sortwhen they don't depend on computed fields - coalesces
$sort+$limitinto a top-k sort, merges adjacent$match,$limitand$skipstages - folds
$unwind(and a following$match) into a preceding$lookup - analyzes field dependencies so only needed fields flow through
What I still do by hand:
- Avoid
$unwind+$groupjust to rebuild an array; use array operators like$filter,$mapand$reduce - Index the
foreignFieldof every$lookup - Keep an eye on the 100 MB per-stage memory limit; spilling to disk works but is slower
- Precompute heavy reports with
$mergeinto a summary collection - Check the result with
explain()on the aggregate
What interviewers listen for- Only leading stages can use indexes
$matchand$sortas early as possible- Optimizer reorders and coalesces stages
- Use array operators instead of unwind-and-regroup
- 100 MB stage limit;
$mergefor precomputed results
Likely follow-up: When can a
$sortfollowed by$groupwith$firstuse an index?- moves
43.What does
$facetdo? How would you return search results, a total count and filter counts in one query?mid$facetruns several sub-pipelines over the same input documents within a single stage. Each sub-pipeline's output becomes an array field in one result document. It's the natural fit for faceted search pages: one round trip returns the first page of results, the total count, counts per brand and price bands.Things to know:
- The sub-pipelines are independent; one can't use another's output. Add stages after
$facetto combine them - Sub-pipelines don't use indexes. Put a selective
$matchbefore$facet, which can use an index; a pipeline that starts with$facetdoes a collection scan - The combined output is a single document, so it must fit in 16 MiB, and each sub-pipeline stage has the usual 100 MB memory limit
- You can't nest
$facetor use stages like$out,$mergeor$geoNearinside it
For very large result sets, separate queries (or Atlas Search's
$searchMetafor facet counts) may be cheaper.db.products.aggregate([ { $match: { category: "laptops", price: { $lte: 2000 } } }, { $facet: { results: [{ $sort: { price: 1, _id: 1 } }, { $limit: 20 }], total: [{ $count: "count" }], byBrand: [{ $group: { _id: "$brand", n: { $sum: 1 } } }, { $sort: { n: -1 } }], priceBands: [{ $bucket: { groupBy: "$price", boundaries: [0, 500, 1000, 2001], default: "other" } }] } } ]);What interviewers listen for- Multiple sub-pipelines on the same input
- Returns one document with an array per facet
- Sub-pipelines cannot use indexes;
$matchfirst - Output limited to 16 MiB
- No nested
$facet,$outor$mergeinside
Likely follow-up: Why might a
$facetcount be slow on a large collection?- The sub-pipelines are independent; one can't use another's output. Add stages after
44.Explain replica set member configuration (priority, votes, hidden, delayed, arbiters) and when a rollback happens.hard
Member options shape elections and roles:
- priority: higher-priority secondaries call elections sooner and are more likely to win;
priority: 0means the member can never become primary - votes: 0 or 1; at most 7 members vote, out of up to 50 total. Non-voting members must have priority 0
- hidden: priority 0 and invisible to client reads; useful for backups or analytics
- delayed: a hidden member applying the oplog after a set delay, giving a window to recover from something like an accidental drop
- arbiter: votes but holds no data; it weakens durability and can make the default write concern
w: 1
Elections use a Raft-like protocol: a candidate needs votes from a majority of voting members, and a primary that loses contact with a majority steps down.
A rollback happens when the old primary rejoins holding writes that never replicated to the new primary's side. Those writes are undone and saved as BSON files under
dbpath/rollbackfor manual review. Writes acknowledged withw: "majority"are never rolled back.What interviewers listen for- Priority 0 members never become primary
- Max 7 voting members of up to 50
- Hidden and delayed members for backups and recovery
- Majority vote to elect; isolated primary steps down
- Rollback removes non-majority writes; majority prevents it
Likely follow-up: How would you lay out a replica set across three data centers?
- priority: higher-priority secondaries call elections sooner and are more likely to win;
45.What are chunks, how does the balancer move them, and what is a jumbo chunk?hard
A chunk is a contiguous range of shard key values owned by one shard. The config servers store the mapping, and
mongosuses it to route queries. The default range size is 128 MB.The balancer runs on the config server primary. In current versions it balances by data size per collection, not chunk count: a round starts when the gap between the shards with the most and least data for a collection reaches the migration threshold, about three times the range size (384 MB by default). Chunks are split when they need to be moved.
A migration copies the documents to the recipient, catches up on changes made meanwhile, commits the new ownership on the config servers, then deletes the range from the donor asynchronously. Until cleanup, the donor's copies are orphaned documents.
A jumbo chunk exceeds the range size but can't be split, because it holds a single shard-key value. It signals low cardinality or high frequency; refining the shard key is the real fix. A balancing window can restrict migrations to off-peak hours.
What interviewers listen for- Chunk: contiguous shard-key range owned by one shard
- Default range size 128 MB
- Balancer balances by data size in current versions
- Migrations leave orphans until the donor cleans up
- Jumbo chunk: single key value too big to split
Likely follow-up: Why should you stop the balancer during a filesystem snapshot backup?
46.How do you secure a MongoDB deployment?mid
I go through the manual's security checklist:
- Enable access control. Self-managed MongoDB doesn't enforce authentication by default; set
security.authorization: enabled. SCRAM is the default mechanism, with x.509 as an option, and LDAP and Kerberos in Enterprise - Least-privilege roles: create a user administrator first, then one user per app or person with only the roles it needs, such as
readorreadWriteon one database, or custom roles. No sharedrootaccounts - Limit network exposure:
mongodbinds to localhost by default; if you widenbindIp, restrict it with firewalls, security groups or private networking. Never expose port 27017 to the internet - Encrypt: TLS for all traffic, including between members; encryption at rest; Client-Side Field Level Encryption or Queryable Encryption for sensitive fields
- Auditing (Enterprise and Atlas), patching, and running
mongodas a dedicated OS user
In application code, guard against operator injection: if a login handler passes
req.body.passwordstraight into a filter, an attacker can send{ "$ne": null }. Cast inputs to expected types, or enable Mongoose'ssanitizeFilter.What interviewers listen for- Turn on authorization; SCRAM by default
- Least-privilege RBAC, one user per app
- Bind to private interfaces, firewall port 27017
- TLS, encryption at rest, field-level encryption
- Prevent operator injection from user input
Likely follow-up: What is the difference between CSFLE and Queryable Encryption?
- Enable access control. Self-managed MongoDB doesn't enforce authentication by default; set
47.What are the options for backing up MongoDB, and what are their trade-offs?mid
First, replication isn't a backup: an accidental
deleteManyreplicates to every secondary within seconds.The main approaches:
mongodump/mongorestore: a logical BSON export. Simple and portable, and the manual positions it for small deployments: it's slow on big datasets, indexes are rebuilt on restore, and it adds load. On a replica set,--oplogcaptures writes during the dump somongorestore --oplogReplayrestores a consistent point- Filesystem snapshots (LVM, EBS and similar): fast and suited to large data. Journaling must be on the same volume. For a sharded cluster you must stop the balancer and snapshot every shard and a config server at about the same moment
cp/rsyncof data files: only with writes stopped- Managed backups: Atlas Cloud Backups, or Ops Manager and Cloud Manager, which take snapshots and support point-in-time recovery, including for sharded clusters
Whatever you choose, define your RPO and RTO, and regularly test restores. A delayed hidden member can help you recover from human error, but it doesn't replace real backups.
What interviewers listen for- Replication is not a backup
mongodumpfor small deployments;--oplogfor consistency- Filesystem snapshots for large data; journal on same volume
- Sharded snapshots need the balancer stopped
- Atlas or Ops Manager for point-in-time restores
Likely follow-up: How would you restore a single collection that was dropped an hour ago?
48.What is MongoDB Atlas, and what does it give you over running MongoDB yourself?easy
Atlas is MongoDB's fully managed cloud database service, running on AWS, Google Cloud and Azure. You choose a cluster tier, cloud and region, and Atlas deploys a replica set or sharded cluster for you.
What you stop doing yourself:
- Operations: provisioning, patching, version upgrades and scaling, including storage and compute auto-scaling
- Backups: cloud snapshots with point-in-time restore
- Monitoring: metrics, alerts, a Query Profiler and a Performance Advisor that suggests indexes
- Security baseline: authentication required, TLS, encryption at rest and IP access lists or private networking
It also bundles platform features that don't exist in a plain
mongod: Atlas Search (Lucene-based full-text search), Vector Search, triggers, Charts, multi-region and multi-cloud clusters, and Online Archive for cold data.There's a free tier for learning (M0) and paid tiers from small shared clusters up to large dedicated ones. The trade-offs are cost at scale and less low-level control over the servers.
What interviewers listen for- Managed MongoDB on AWS, Google Cloud and Azure
- Automates patching, scaling and backups
- Monitoring, Query Profiler, Performance Advisor
- Secure defaults: auth, TLS, network access lists
- Adds Atlas Search and Vector Search
Likely follow-up: How does an application connect to Atlas securely?
49.What is Mongoose middleware? What is the bug in relying only on this
pre("save")hook to hash passwords?midMiddleware (hooks) are functions Mongoose runs before (
pre) or after (post) an operation. There are four kinds:- Document middleware:
validate,save,initand others, wherethisis the document - Query middleware:
find,findOne,findOneAndUpdate,updateOne,deleteManyand so on, wherethisis the query, not a document - Aggregate middleware for
aggregate() - Model middleware such as
insertManyandbulkWrite
With an async function you don't need to call
next(). Hooks must be registered beforemongoose.model()compiles the schema.The bug:
savehooks don't run forupdateOne,findOneAndUpdateand other query-based updates. That update goes straight to the database, so the password is stored in plain text. Fixes: load the document, set the field and callsave(); or add a matchingpre("updateOne")orpre("findOneAndUpdate")hook that hashes the value in the update; or route all password changes through one service method.userSchema.pre("save", async function () { if (!this.isModified("password")) return; this.password = await bcrypt.hash(this.password, 12); }); // Elsewhere in the codebase: await User.updateOne({ _id: id }, { password: req.body.password });What interviewers listen for- Pre and post hooks around operations
- Document, query, aggregate and model middleware
thisis the query in query middleware- Save hooks do not run on
updateOneorfindOneAndUpdate - Register hooks before compiling the model
Likely follow-up: How would you implement soft delete with query middleware?
- Document middleware:
50.Explain the bucket pattern. How does this update implement it for sensor readings?hard
The bucket pattern groups many small, related records into one document per bucket, usually per entity per time window, instead of one document per event. A sensor writing every second produces 86,400 tiny documents a day; bucketing turns that into a few hundred documents, with far fewer index entries, better compression and faster range reads.
This update appends a reading to the sensor's current bucket for the day, but only while
countis under 200. When the bucket is full, the filter matches nothing and the upsert creates a new bucket. ItssensorIdanddaycome from the filter's equality fields, while the$ltcondition isn't copied. The bucket also keeps precomputed summary fields (count,sumTemp,first,last), so averages and ranges don't need to unwind the array. That's the computed pattern layered on top.The count cap keeps arrays bounded. For time-series workloads, time series collections, available since 5.0, apply this bucketing internally, so today I'd try one of those first.
db.readings.updateOne( { sensorId: "s-17", day: ISODate("2026-09-28"), count: { $lt: 200 } }, { $push: { samples: { t: new Date(), temp: 21.4 } }, $inc: { count: 1, sumTemp: 21.4 }, $min: { first: new Date() }, $max: { last: new Date() } }, { upsert: true } );What interviewers listen for- Group many small records into bounded buckets
- Fewer documents and index entries, faster range reads
- Upsert with a count cap starts a new bucket
- Precomputed summaries per bucket
- Time series collections bucket automatically (5.0+)
Likely follow-up: How would you query the average temperature for one sensor over a week?
51.Explain the subset and extended reference patterns with an example of each.hard
Both reduce what a read has to fetch, at the cost of some duplication.
Subset pattern. A product page shows the 10 newest reviews, but popular products have thousands. Embedding them all makes the product document huge, and every product read drags them into RAM. Instead, keep all reviews in a
reviewscollection and embed only the latest 10 in the product. Adding a review is two writes: insert intoreviews, then$pushonto the product with$each,$sortand$slice: 10. The working set shrinks; "see all reviews" is a second query.Extended reference pattern. An order references its customer by
_id, but every order view needs the name and shipping address. Rather than a$lookupon each read, copy just those fields into the order:customer: { _id, name, shippingAddress }.The trade-off is keeping duplicated data in sync. Pick fields that rarely change, or where a historical snapshot is actually correct: the address an order shipped to shouldn't change later. If updates must propagate, use a background job or a change stream.
What interviewers listen for- Subset: embed only the hot part, rest elsewhere
- Shrinks documents and the working set
- Extended reference: copy a few fields of a reference
- Avoids
$lookupon frequent reads - Duplicate only stable or snapshot-worthy fields
Likely follow-up: How would you propagate a customer name change to existing orders?
52.What are the computed and outlier patterns, and what problems do they solve?hard
Computed pattern: when the same value is calculated over and over on reads, store the result instead. A movie page shows the average rating and number of ratings; rather than aggregating millions of rating documents each time, every new rating also runs
$inc: { ratingCount: 1, ratingSum: 5 }on the movie, and the average is derived from those two fields. For expensive but less urgent figures, a scheduled pipeline with$mergecan refresh a summary collection. It suits read-heavy workloads, and the cost is extra write work plus deciding how fresh the numbers need to be.Outlier pattern: design for the typical document and handle the rare extreme one separately. Most books have a few hundred buyers in their embedded array, but a bestseller has millions, which would break the 16 MiB limit. Keep the embedded array for everyone, cap it, and set a flag such as
hasExtras: trueon outliers, with overflow in separate documents. The app checks the flag and fetches the extras only when needed. The common path stays fast; the price is extra application logic.What interviewers listen for- Computed: store derived values instead of recomputing
- Update on write with
$inc, or batch with$merge - Outlier: optimize for the typical document
- Flag outliers and overflow their extra data
- Both trade write or app complexity for read speed
Likely follow-up: How would you keep a computed average correct if ratings can be edited?
53.How do you update specific elements inside an array? Explain
$,$[]and$[<identifier>].midMongoDB has three positional operators for updating array elements in place:
$: refers to the first element that matched the array condition in the query filter. The array field must appear in the filter, asitems.skudoes here, so this increments B's quantity$[]: the all-positional operator, which updates every element of the array$[<identifier>]: the filtered positional operator, which updates every element matching a condition you give in thearrayFiltersoption. The identifier must start with a lowercase letter, and each identifier needs exactly one array filter
$[<identifier>]also nests, for example"grades.$[g].scores.$[s]", for arrays inside arrays, which plain$can't handle.All of these are atomic single-document updates, so there's no need to read the array, change it in the app and write it back. If you're matching elements by a unique key like
sku,$is simplest; usearrayFilterswhen several elements should change.// { _id: 1, items: [ { sku: "A", qty: 1 }, { sku: "B", qty: 2 } ] } // $ : the first element matched by the filter db.carts.updateOne({ _id: 1, "items.sku": "B" }, { $inc: { "items.$.qty": 1 } }); // $[] : every element db.carts.updateOne({ _id: 1 }, { $set: { "items.$[].reserved": false } }); // $[<identifier>] : every element matching arrayFilters db.carts.updateOne( { _id: 1 }, { $set: { "items.$[low].flag": "restock" } }, { arrayFilters: [{ "low.qty": { $lt: 2 } }] } );What interviewers listen for$: first element matched by the query filter$[]: all elements$[id]plusarrayFilters: all matching elements- Filtered positional works in nested arrays
- In-place, atomic; no read-modify-write
Likely follow-up: What happens with
$if the filter matches two array elements?54.How would you build a job queue where workers never grab the same job, and how do you do optimistic locking in MongoDB?hard
Both rely on the fact that a single-document update is atomic, and that the filter is evaluated as part of that atomic write.
For the queue,
findOneAndUpdatefinds, modifies and returns one document in one operation. Two workers can race for the same pending job, but only one update can flip it frompendingtorunning; the other matches the next job instead.sortpicks the oldest one, and an index on{ status: 1, createdAt: 1 }keeps it fast. By default it returns the document before the update;returnDocument: "after"(orreturnNewDocument: truein mongosh) returns the new version. Add a lease timeout so jobs from crashed workers can be reclaimed.For optimistic locking, keep a
versionfield. Include the version you read in the filter and increment it in the update. If someone else saved first,matchedCountis 0, and you reload and retry. No locks are held between read and write. Mongoose offers the same behavior with itsoptimisticConcurrencyschema option.// Claim the oldest pending job atomically const job = db.jobs.findOneAndUpdate( { status: "pending" }, { $set: { status: "running", workerId: "w-3", startedAt: new Date() } }, { sort: { createdAt: 1 }, returnDocument: "after" } ); // Optimistic locking: write only if nobody changed it since we read it const res = db.docs.updateOne( { _id: doc._id, version: doc.version }, { $set: { body: newBody }, $inc: { version: 1 } } ); if (res.matchedCount === 0) { /* conflict: reload and retry */ }What interviewers listen for- Filter plus update is one atomic operation
findOneAndUpdateclaims and returns a document- Returns the pre-update doc unless told otherwise
- Version field in filter for optimistic locking
matchedCountof 0 means a conflict
Likely follow-up: How would you generate a sequential invoice number safely?
55.What is WiredTiger, and how do its cache, compression and journaling affect you?mid
WiredTiger is MongoDB's default storage engine, the layer that actually stores documents and indexes on disk.
- Concurrency: it uses document-level concurrency control, so different clients can write different documents in the same collection simultaneously. Conflicting writes to the same document are detected and retried transparently. It also provides MVCC snapshots, which transactions build on
- Cache: the internal cache defaults to the larger of 50% of (RAM − 1 GB) or 256 MB. MongoDB also benefits from the OS file cache. Performance depends on the working set, meaning hot documents plus indexes, fitting in memory
- Compression: collections use snappy block compression by default (zlib and zstd are options), and indexes use prefix compression, which often shrinks disk use a lot
- Durability: it writes checkpoints every 60 seconds, and a write-ahead journal records changes in between, so after a crash MongoDB recovers from the last checkpoint plus the journal
When sizing a server, I size RAM so indexes and hot data fit in the cache.
What interviewers listen for- Default storage engine
- Document-level concurrency, MVCC snapshots
- Cache: larger of 50% of (RAM − 1 GB) or 256 MB
- Snappy for data, prefix compression for indexes
- Checkpoints every 60 s plus a journal
Likely follow-up: What happens to performance when the working set exceeds the cache?
56.What is the difference between
countDocuments()andestimatedDocumentCount()?easyThey trade accuracy for speed.
countDocuments(filter)takes a query filter and actually counts matching documents, running an aggregation under the hood. It's accurate and can use indexes, but on a big result it still has to walk every matching index entry or document, so counting millions of rows isn't free.estimatedDocumentCount()takes no filter and reads the collection's metadata, so it returns almost instantly even for huge collections. It can be off after an unclean shutdown, and on a sharded cluster it doesn't filter out orphaned documents.The older
count()helper is best avoided: without a filter it also relies on metadata and can be approximate, which confused many people. Drivers replaced it with these two explicit methods.In practice I use
estimatedDocumentCount()for dashboards like "about 12 million users", andcountDocuments()with an indexed filter when the number must be exact. For paginated UIs, I avoid exact totals on every request, caching them or showing "more results" instead.What interviewers listen forcountDocumentsis exact and accepts a filterestimatedDocumentCountuses metadata, no filter, fast- Estimates can drift after unclean shutdowns or with orphans
- Avoid the legacy
count()helper - Exact counts on huge sets are expensive
Likely follow-up: How would you show the total number of results for a search page cheaply?
57.What happens when one document fails in
insertMany()? How do ordered and unordered bulk writes differ?easyinsertMany()is not atomic across documents; each insert succeeds or fails on its own.- With the default
ordered: true, the server inserts in order and stops at the first error, say a duplicate key on document 50. Documents 1 to 49 stay inserted, and nothing after 50 is attempted - With
ordered: false, it keeps going past errors and reports all failures at the end. It can also be faster, especially on sharded clusters, because operations don't wait for the previous one
bulkWrite()works the same way but mixes operation types in one call:insertOne,updateOne,updateMany,replaceOne,deleteOneanddeleteMany. That's much faster than individual calls because it saves round trips. The driver splits very large requests into batches (the server's maximum is 100,000 operations per batch).On failure you get a bulk write error listing each failed operation's index and error, plus counts of what succeeded. If the whole set must be all-or-nothing, run it inside a transaction. Otherwise make the operations idempotent, for example upserts, so a retry is safe.
What interviewers listen for- Not atomic: earlier successes stay inserted
- Ordered (default) stops at the first error
- Unordered continues and reports all errors
bulkWritemixes operation types, fewer round trips- Use a transaction for all-or-nothing
Likely follow-up: Why is
ordered: falseoften faster on a sharded cluster?- With the default
58.What is a wildcard index, and when would you use one instead of regular indexes?hard
A wildcard index indexes fields whose names you don't know in advance.
{ "$**": 1 }indexes every field in the document, and{ "attributes.$**": 1 }indexes every subfield underattributes. You can include or exclude specific paths withwildcardProjectionwhen using$**.The typical use case is user-defined or highly variable data: product attributes that differ by category, custom fields in a SaaS app, or event payloads. One wildcard index serves queries like
{ "attributes.color": "red" }and{ "attributes.voltage": 220 }without creating one index per attribute. Since 7.0, compound wildcard indexes such as{ tenantId: 1, "attributes.$**": 1 }combine a known field with the wildcard part.Limitations:
- It can't be
uniqueor TTL, and it can't be a shard key - It's sparse, so it can't support
$exists: false - It supports only one field predicate per query, and can't match whole embedded documents or arrays
It isn't a replacement for designed indexes on known query patterns. The attribute pattern, an array of key/value pairs with a compound index, is the older alternative.
What interviewers listen for- Indexes unknown or variable field names
$**for all fields or under a path- Compound wildcard indexes since 7.0
- Not unique, TTL or shard key
- Attribute pattern is the alternative
Likely follow-up: How does the attribute pattern compare to a wildcard index?
- It can't be
59.A user updates their profile, then immediately sees the old version because reads go to secondaries. How do you fix this?hard
Replication is asynchronous, so a read routed to a secondary can arrive before that secondary has applied the user's write. That breaks read-your-own-writes.
Options, simplest first:
- Read from the primary for that flow. It's the default read preference, and many apps only use secondaries for analytics anyway
- Use a causally consistent session. Within a session, if reads use
"majority"read concern and writes use"majority"write concern, MongoDB guarantees read-your-writes, monotonic reads, monotonic writes and writes-follow-reads, even when operations hit different members. The driver tracks the session's latest operation time and sends it with the next read, so the secondary waits until it has caught up to that point before answering. Drivers enable causal consistency on explicit sessions by default; the app must pass the session to every related operation - Return the updated document from the write, for example with
findOneAndUpdate, and show that, instead of re-reading
The guarantees apply only within one session, so across different devices or services you still need one of the other approaches.
What interviewers listen for- Async replication makes secondary reads stale
- Read from primary for read-your-writes flows
- Causally consistent sessions with majority concerns
- Secondary waits for the session's operation time
- Guarantees only hold within one session
Likely follow-up: What does
maxStalenessSecondsdo, and why is it not enough here?
No questions match that filter.
Prefer multiple choice? All 22 MongoDB MCQs with answers →