Ch. 12

MongoDB interview questions & answers

Document databases: schema design and embedding vs referencing, indexes, the aggregation pipeline, replication, sharding and when to choose NoSQL.

59 interview questions22 quiz questions0 notes
your progress0%

Top 59 MongoDB interview questions most asked first

  1. 1.What is MongoDB, and when would you choose it over a relational database?easy

    MongoDB is a document database. It stores records as BSON documents (JSON-like, with nested objects and arrays) grouped into collections, with a flexible schema, a rich query language, secondary indexes, an aggregation framework, built-in replication through replica sets and horizontal scaling through sharding.

    I'd pick it when:

    • The data is naturally hierarchical or varies per record: product catalogs, user profiles, content, events, IoT readings
    • The app mostly reads and writes whole aggregates, so one document can hold what would be several joined tables
    • The schema evolves quickly, or the dataset will outgrow one server

    I'd lean relational when the domain is dominated by many-to-many relationships and ad hoc joins, or when most operations touch many entities transactionally. MongoDB does have ACID multi-document transactions and schema validation, so the real decision is about data shape and access patterns, not "no transactions".

    What interviewers listen for
    • Document database storing BSON documents in collections
    • Flexible schema, rich queries, indexes, aggregation
    • Replica sets for HA, sharding for scale-out
    • Fits aggregates read together and evolving schemas
    • Relational fits join-heavy, highly relational domains

    Likely follow-up: What would make you move a MongoDB workload back to SQL?

  2. 2.How does MongoDB compare with a SQL database? Map the main concepts and explain the key differences.easy

    The terminology maps fairly directly:

    • table → collection, row → document, column → field
    • primary key → the _id field
    • join → embedding related data, or $lookup in an aggregation
    • GROUP BY → the $group stage

    The bigger differences are in philosophy. SQL databases enforce a fixed schema and normalize data by entity, then join at query time. MongoDB lets documents in a collection differ (you can add $jsonSchema validation when you want rules), and you model around access patterns, storing data that's read together in one document.

    MongoDB doesn't enforce foreign keys, so referential integrity is the application's job. On scaling, MongoDB has sharding built in, whereas many relational setups scale up first. Both offer ACID transactions; in MongoDB a single-document write is always atomic, and multi-document transactions exist but cost more.

    What interviewers listen for
    • Collection, document, field, _id map to table, row, column, PK
    • Joins become embedding or $lookup
    • Flexible schema with optional validation
    • Model by access pattern, not by entity
    • No enforced foreign keys; sharding built in

    Likely follow-up: What are the main categories of NoSQL databases?

  3. 3.What is BSON, and why does MongoDB store BSON instead of plain JSON?easy

    BSON is a binary-encoded serialization of JSON-like documents. MongoDB uses it on disk and on the wire for two reasons.

    First, richer types. JSON only has strings, numbers, booleans, null, objects and arrays. BSON adds ObjectId, Date, 32- and 64-bit integers, doubles, Decimal128 for exact money values, binary data, regular expressions and more, so a date stays a date and you can range-query and index it correctly.

    Second, efficient traversal. Elements are type-tagged and length-prefixed, so the server can skip over fields without parsing everything.

    A document is an ordered set of field/value pairs, up to 16 MiB and 100 levels of nesting. Documents live in collections, which live in databases. Both are created implicitly on first insert. When mongosh prints BSON it uses helpers like ObjectId(...) and ISODate(...); drivers exchange it as Extended JSON when text is needed.

    db.orders.insertOne({
      customerId: ObjectId("64f1c2a9e4b0a1b2c3d4e5f6"),
      total: Decimal128("59.90"),
      qty: NumberInt(3),
      placedAt: new Date(),
      items: [{ sku: "A1", price: Decimal128("19.96") }]
    });
    What interviewers listen for
    • Binary, type-tagged, length-prefixed JSON-like format
    • Extra types: ObjectId, Date, Int32/Int64, Decimal128, binary
    • Faster to scan than parsing text JSON
    • Documents max 16 MiB, 100 nesting levels
    • Collections and databases are created implicitly

    Likely follow-up: Why use Decimal128 rather than a double for prices?

  4. 4.What is the _id field, and what is inside an ObjectId?easy

    Every document in a standard collection needs an _id that acts as its primary key. It must be unique within the collection, it's immutable once set, and it can be any BSON type except an array, a regex or undefined. MongoDB automatically creates a unique index on _id for every collection. If you insert a document without one, the driver generates an ObjectId.

    An ObjectId is 12 bytes:

    • a 4-byte timestamp: seconds since the Unix epoch
    • a 5-byte random value generated once per client process
    • a 3-byte incrementing counter, initialized to a random value

    So IDs are generated client-side with no coordination, and getTimestamp() recovers the creation time. Because they roughly increase over time, sorting by _id approximates insertion order, but it isn't strictly monotonic across different clients. That same "always increasing" property makes a plain ranged ObjectId a poor shard key.

    What interviewers listen for
    • Primary key: unique, immutable, automatically indexed
    • Any BSON type except array, regex or undefined
    • 4-byte timestamp, 5-byte random, 3-byte counter
    • Generated by the driver without coordination
    • Roughly time-ordered; bad as a ranged shard key

    Likely follow-up: When would you use a natural key instead of an ObjectId for _id?

  5. 5.How do you decide whether to embed related data in a document or reference it from another collection?easy

    The guiding principle is data that's accessed together should be stored together, and I start from embedding unless there's a reason not to.

    Embed when:

    • The child is owned by the parent and read with it, like an order's line items or a user's addresses
    • The relationship is one-to-few and the array is bounded
    • You want the parent and children updated atomically in a single-document write

    Reference when:

    • The "many" side is large or grows without limit, so the document would bloat toward the 16 MiB cap
    • The child needs to be queried or updated on its own
    • The data is shared by many parents and changes often, so duplicating it would be expensive to keep in sync
    • It's many-to-many

    Often the answer is a hybrid: reference the full record but embed a few frequently read fields (the extended reference pattern). Always design from the queries the application actually runs.

    What interviewers listen for
    • Data accessed together is stored together
    • Embed bounded, owned, read-together children
    • Reference unbounded, shared or independently queried data
    • Embedding gives single-document atomicity
    • Hybrid: reference plus a few duplicated fields

    Likely follow-up: How would you model blog posts and their comments?

  6. 6.Why do we need indexes in MongoDB, and which index types does it support?easy

    Without a suitable index a query has to do a collection scan (COLLSCAN), reading every document. An index is an ordered B-tree of field values pointing to documents, so the server can jump to matching entries and even return results in index order without sorting.

    Index types:

    • Single field and compound (several fields, order matters)
    • Multikey: created automatically when an indexed field holds an array
    • Text: keyword search with stemming; one per collection
    • Geospatial: 2dsphere and 2d
    • Hashed: used mainly for hashed sharding
    • Wildcard: { "$**": 1 } for unpredictable field names

    Index properties include unique, partialFilterExpression, sparse, TTL (expireAfterSeconds), hidden and a collation for case-insensitive matching. Every collection gets a unique _id index.

    The trade-off: each index costs RAM and disk, and slows every insert, update and delete that touches its fields.

    What interviewers listen for
    • Avoids COLLSCAN; B-tree lookups and index-ordered sorts
    • Single, compound, multikey, text, geo, hashed, wildcard
    • Properties: unique, partial, sparse, TTL, hidden
    • Automatic unique index on _id
    • Each index costs memory and write speed

    Likely follow-up: How would you find indexes that are never used?

  7. 7.What is the aggregation pipeline? Walk me through what this query does.easy

    The aggregation pipeline is MongoDB's framework for transforming and analyzing data on the server. Documents flow through an ordered list of stages, and each stage's output is the next stage's input, much like a Unix pipe.

    This one returns the top five customers by paid revenue this year:

    • $match filters to paid orders since January. Placed first, it can use an index on status and placedAt
    • $group makes one document per customerId, summing amount and counting orders with $sum: 1
    • $sort and $limit keep the five biggest totals; the optimizer merges them into a top-k sort
    • $project reshapes the output, renaming _id to customerId

    Other common stages are $lookup for joins, $unwind to flatten arrays, $addFields/$set, $facet, $count, and $out or $merge to write results to a collection.

    db.orders.aggregate([
      { $match: { status: "paid", placedAt: { $gte: ISODate("2026-01-01") } } },
      { $group: { _id: "$customerId", total: { $sum: "$amount" }, orders: { $sum: 1 } } },
      { $sort: { total: -1 } },
      { $limit: 5 },
      { $project: { _id: 0, customerId: "$_id", total: 1, orders: 1 } }
    ]);
    What interviewers listen for
    • Ordered stages; output of one feeds the next
    • $match early so it can use indexes
    • $group with accumulators like $sum, $avg
    • $sort plus $limit becomes a top-k sort
    • $project reshapes; $out/$merge persist results

    Likely follow-up: What is the memory limit for a pipeline stage, and what happens when you exceed it?

  8. 8.What is a replica set, and what happens when the primary goes down?mid

    A replica set is a group of mongod processes holding the same data. One member is the primary and accepts all writes; it records them in its oplog, a capped collection that secondaries copy and replay asynchronously. Members exchange heartbeats every two seconds.

    If secondaries can't reach the primary for longer than electionTimeoutMillis (10 seconds by default), an eligible secondary calls an election, and the candidate that wins votes from a majority of voting members becomes primary. The manual says a new primary is typically elected in about 12 seconds or less with default settings. During the election no writes are accepted, but secondaries can keep serving reads if your read preference allows it. Drivers enable retryable writes by default, so many writes are retried once automatically.

    When the old primary comes back it rejoins as a secondary. Any of its writes that never reached a majority are rolled back, which is why w: "majority" matters. A set can have up to 50 members, but at most 7 voting ones; use an odd number of voters.

    What interviewers listen for
    • One primary takes writes; secondaries replay its oplog
    • Election after 10 s without primary contact (default)
    • New primary needs a majority of voting members
    • Retryable writes smooth over failover
    • Non-majority writes on the old primary roll back

    Likely follow-up: What is an arbiter, and why is it often discouraged?

  9. 9.What is sharding in MongoDB, and what are the components of a sharded cluster?mid

    Sharding is horizontal partitioning: a collection's documents are spread across several shards so data size and write throughput can exceed what one server handles.

    A sharded cluster has three parts:

    • Shards: each is a replica set holding a subset of the data
    • mongos routers: the app connects to these; they look up metadata and route each operation to the right shards
    • Config servers: a replica set storing cluster metadata, such as which ranges live where

    The shard key, one or more fields chosen per collection, decides where each document goes. Data is split into chunks (contiguous shard key ranges, 128 MB by default), and the balancer migrates chunks to keep data evenly spread. Queries that include the shard key are targeted to specific shards; queries without it become scatter-gather across all shards.

    Sharding adds real operational complexity, so I'd first make sure indexes, schema and vertical scaling on a replica set are exhausted.

    What interviewers listen for
    • Horizontal partitioning of a collection across shards
    • Shards are replica sets; mongos routes; config servers store metadata
    • Shard key decides document placement
    • Chunks of 128 MB default, moved by the balancer
    • Targeted vs scatter-gather queries

    Likely follow-up: What happens to a query that does not include the shard key?

  10. 10.Does MongoDB support ACID transactions? Show how you would transfer money between two accounts.easy

    Yes, at two levels.

    A write to a single document is always atomic, including all its embedded documents and arrays. With good modeling, that covers most use cases.

    For changes across documents or collections, MongoDB has multi-document ACID transactions: on replica sets since 4.0 and on sharded clusters since 4.2. They don't work on a standalone server. Inside a transaction, reads come from a consistent snapshot (on sharded clusters, ask for read concern "snapshot" to guarantee it across shards), commit is all-or-nothing, and no other client sees the changes until commit.

    In mongosh, session.withTransaction() runs the callback, commits, and retries the whole transaction or the commit on transient errors. Throwing inside the callback aborts and rolls everything back. That's why the snippet checks modifiedCount: the filter balance: { $gte: 100 } makes the debit conditional, and if it matched nothing we throw instead of committing half a transfer.

    Transactions cost more than single-document writes, so the manual advises against using them as a substitute for good schema design.

    const session = db.getMongo().startSession();
    try {
      session.withTransaction(() => {
        const accounts = session.getDatabase("bank").accounts;
        const debit = accounts.updateOne(
          { _id: "A", balance: { $gte: 100 } },
          { $inc: { balance: -100 } }
        );
        if (debit.modifiedCount !== 1) throw new Error("insufficient funds");
        accounts.updateOne({ _id: "B" }, { $inc: { balance: 100 } });
      });
    } finally {
      session.endSession();
    }
    What interviewers listen for
    • Single-document writes are always atomic
    • Multi-document: replica sets 4.0+, sharded 4.2+
    • Not available on standalone servers
    • withTransaction retries transient errors; throw to abort
    • More expensive; prefer modeling first

    Likely follow-up: What is a TransientTransactionError?

  11. 11.How do you order the fields of a compound index? Explain the ESR rule with this query.mid

    A compound index is sorted by its first field, then by the second within each value of the first, and so on. Two consequences:

    • Prefix rule: { a: 1, b: 1, c: 1 } supports queries on a, a + b and a + b + c, but not on b alone
    • Sort direction: it supports sort({ a: 1, b: 1 }) and the exact reverse, but not a mixed { a: 1, b: -1 }

    The ESR guideline says: Equality fields first, then Sort fields, then Range fields. Equality first means everything after it stays in sorted order. Sort next lets MongoDB read keys already in placedAt order and skip an in-memory sort. Range last, because a range on total placed before placedAt would break that ordering.

    Watch out: $ne, $nin and $regex count as range operators. And if the range filter is very selective, putting it before the sort field (ERS) can examine fewer keys, so confirm with explain().

    db.orders.find({ status: "shipped", total: { $gt: 100 } })
      .sort({ placedAt: -1 });
    
    // ESR: Equality, then Sort, then Range
    db.orders.createIndex({ status: 1, placedAt: -1, total: 1 });
    What interviewers listen for
    • Prefix rule: leftmost fields must be used
    • Equality, then Sort, then Range
    • Sort field before range avoids an in-memory sort
    • $ne, $nin, $regex behave as range predicates
    • Verify the choice with explain()

    Likely follow-up: How does $in fit into ESR when the query also sorts?

  12. 12.How do you use explain() to check whether a query is efficient? What is the difference between COLLSCAN and IXSCAN?mid

    explain() shows the plan the query optimizer chose. It has three verbosity modes: queryPlanner (the default, plan only), executionStats (runs the winning plan and reports counters) and allPlansExecution (also partial stats for rejected plans).

    In winningPlan, read the stage tree from the leaf up:

    • COLLSCAN: a full collection scan, reading every document. Fine for tiny collections, a red flag otherwise
    • IXSCAN: scanning an index range, usually followed by FETCH to load the documents
    • A SORT stage means an in-memory sort; if it's absent, the index supplied the order

    Then compare nReturned, totalKeysExamined and totalDocsExamined in executionStats. The ideal is close to 1:1:1. If you examine 50,000 keys to return 20 documents, the index is weakly selective or its field order is wrong. totalDocsExamined: 0 means a covered query. rejectedPlans shows which other indexes the planner considered.

    db.orders.find({ customerId: 42, status: "shipped" })
      .sort({ placedAt: -1 })
      .explain("executionStats");
    
    // For pipelines:
    db.orders.explain("executionStats").aggregate([{ $match: { customerId: 42 } }]);
    What interviewers listen for
    • Modes: queryPlanner, executionStats, allPlansExecution
    • COLLSCAN reads every document; IXSCAN reads an index range
    • A SORT stage means an in-memory sort
    • Compare nReturned vs keys and docs examined
    • Aim for roughly 1:1:1

    Likely follow-up: How would you find slow queries in production in the first place?

  13. 13.How do you choose a good shard key, and what happens if you pick a bad one?hard

    I evaluate candidates on four properties:

    • High cardinality: each distinct key value can live in only one chunk, so a key like continent caps you at seven chunks, and adding more shards won't help
    • Low frequency: if a few values dominate, their chunks grow huge, can't be split, and become hot spots or jumbo chunks
    • Not monotonic: with timestamps or ObjectIds, every insert lands in the chunk with the maxKey upper bound, so one shard takes all the writes
    • Query isolation: the key should appear in most queries so mongos can target one shard instead of scatter-gather

    A compound key often satisfies all four, for example { customerId: 1, orderDate: 1 }: customer gives spread and targeting, date adds cardinality. Hashing a monotonic field fixes write hot spots but makes range queries broadcast.

    A bad key used to be permanent. Now you can refine it by adding suffix fields with refineCollectionShardKey, or fully reshard with reshardCollection since 5.0, though that's a heavy operation. From 7.0, analyzeShardKey reports cardinality, frequency and monotonicity before you commit.

    What interviewers listen for
    • High cardinality and low frequency
    • Avoid monotonically increasing keys for ranged sharding
    • Include the key in common queries to target shards
    • Compound keys balance spread and targeting
    • Refine or reshard (5.0+) if you got it wrong

    Likely follow-up: Why is { createdAt: 1 } a poor shard key for an events collection?

  14. 14.Are updates to a single document atomic in MongoDB? How do you avoid lost updates under concurrency?easy

    Yes. Any write to one document is atomic, even if it changes several fields, embedded documents and arrays at once. Other readers see either the old document or the new one, never half of it.

    The first pattern is a classic lost update: two clients both read stock: 5, both write 3, and one sale vanishes. The fix is to let the server do the math with update operators like $inc, $set, $push or $min, and to put the precondition in the filter. The second statement only matches if at least two units remain, so checking modifiedCount tells you whether the reservation succeeded.

    This atomicity is per document. updateMany applies atomically to each matched document but not to the set as a whole; other operations can interleave. If several documents must change together, either model them into one document or use a multi-document transaction.

    // Race-prone: read, compute in the app, write back
    const p = db.products.findOne({ _id: 1 });
    db.products.updateOne({ _id: 1 }, { $set: { stock: p.stock - 2 } });
    
    // Atomic: the check and the change are one document write
    db.products.updateOne(
      { _id: 1, stock: { $gte: 2 } },
      { $inc: { stock: -2 }, $set: { updatedAt: new Date() } }
    );
    What interviewers listen for
    • Single-document writes are atomic, including nested data
    • Use $inc/$set instead of read-modify-write
    • Put preconditions in the filter; check modifiedCount
    • updateMany is atomic per document, not overall
    • Cross-document atomicity needs a transaction

    Likely follow-up: How would you implement optimistic locking with a version field?

  15. 15.How do you write queries with find()? Explain the common query operators used here.easy

    find(filter, projection) returns a cursor over matching documents. Separate conditions in the filter are implicitly ANDed.

    The main operator groups:

    • Comparison: $eq, $ne, $gt, $gte, $lt, $lte, $in, $nin
    • Logical: $and, $or, $nor, $not
    • Element: $exists checks presence, $type checks the BSON type
    • Evaluation: $regex, and $expr to use aggregation expressions, such as comparing two fields of the same document
    • Array: $all (contains every value), $size, $elemMatch

    So this finds working-age users in India or the US, on the pro plan or with more than 100 credits, not soft-deleted, tagged both beta and mobile, who have spent more than their budget.

    For performance: equality and ranges use indexes well; $ne, $nin and unanchored regexes are rarely selective; a case-sensitive prefix regex like /^abc/ can use an index range.

    db.users.find({
      age: { $gte: 18, $lt: 65 },
      country: { $in: ["IN", "US"] },
      $or: [{ plan: "pro" }, { credits: { $gt: 100 } }],
      deletedAt: { $exists: false },
      tags: { $all: ["beta", "mobile"] },
      $expr: { $gt: ["$spent", "$budget"] }
    });
    What interviewers listen for
    • Filter fields are implicitly ANDed
    • Comparison, logical, element, evaluation and array operators
    • $expr compares fields within a document
    • $all needs every value; $in needs any
    • Anchored prefix regex can use an index

    Likely follow-up: What is the difference between $in and $all?

  16. 16.MongoDB has no JOIN keyword. How do you combine data from two collections?mid

    There are three options.

    First, avoid the join by embedding the data, or by copying a few fields you always need (the extended reference pattern).

    Second, $lookup in an aggregation, which performs a left outer join. For each input document it finds matches in the from collection and puts them in an array field named by as. An order with no customer gets an empty array. $unwind then turns the one-element array into an embedded object, but it drops documents whose array is empty unless you set preserveNullAndEmptyArrays: true. The pipeline form, with let and pipeline, supports extra conditions and sub-queries on the foreign side.

    Third, application-level joins: query orders, collect the IDs, then one $in query for customers. That's what Mongoose's populate() does.

    Whichever you choose, index the foreign field. Without one, each lookup can scan the whole foreign collection. If nearly every read needs a join, that's a sign to revisit the schema.

    db.orders.aggregate([
      { $match: { status: "paid" } },
      { $lookup: {
          from: "customers",
          localField: "customerId",
          foreignField: "_id",
          as: "customer"
      } },
      { $unwind: "$customer" },
      { $project: { total: 1, "customer.name": 1, "customer.email": 1 } }
    ]);
    What interviewers listen for
    • Embed or duplicate fields to avoid joins
    • $lookup is a left outer join into an array
    • $unwind drops non-matches unless preserveNullAndEmptyArrays
    • Index the foreignField
    • App-level joins with $in also work

    Likely follow-up: Can the from collection of a $lookup be sharded?

  17. 17.What is write concern? Explain w, j and wtimeout, and what the default is.mid

    Write concern is the level of acknowledgment you ask for before a write is reported as successful. It trades latency for durability.

    • w: 0: fire and forget, no acknowledgment
    • w: 1: the primary applied it. It can be rolled back if the primary fails before replicating
    • w: "majority": a majority of data-bearing voting members have it. It survives failover
    • w: <n> or a custom tag name for specific member counts or data centers
    • j: true: wait until the write is in the on-disk journal
    • wtimeout: how long to wait for the concern, in milliseconds

    Since MongoDB 5.0 the implicit default is w: "majority". The exception: with arbiters, if data-bearing voters don't exceed the voting majority, for example primary-secondary-arbiter, the default is w: 1.

    Gotcha: a wtimeout error doesn't undo anything. The write may already be applied and may still replicate; you only learn the guarantee wasn't confirmed in time, so the retry logic must be idempotent.

    What interviewers listen for
    • Acknowledgment level: latency vs durability
    • w: 1 can roll back; majority survives failover
    • j: true waits for the on-disk journal
    • Default w: "majority" since 5.0, except some arbiter setups
    • wtimeout errors do not undo the write

    Likely follow-up: Why does a primary-secondary-arbiter set default to w: 1?

  18. 18.What is the difference between read concern and read preference?mid

    They answer different questions.

    Read preference decides which member serves the read:

    • primary (default), primaryPreferred, secondary, secondaryPreferred, nearest
    • optionally refined with tag sets (for example, "same region") and maxStalenessSeconds to skip lagging secondaries

    Read concern decides which version of the data is returned, meaning its consistency and durability:

    • "local" (default): the member's latest data, which might later roll back
    • "available": similar, lowest latency, with sharding caveats
    • "majority": only data acknowledged by a majority, so it won't be rolled back
    • "linearizable": reflects every majority write completed before the read began
    • "snapshot": a consistent point-in-time view, used in transactions

    Replication is asynchronous, so reading from secondaries can return stale data even with "majority". Use secondaries for analytics or geo-local reads that tolerate lag, not for read-your-own-writes. Transactions that read must use primary.

    What interviewers listen for
    • Read preference: which member serves the read
    • Read concern: which data version, consistency and durability
    • Defaults: primary and "local"
    • Majority reads never see rolled-back data
    • Secondary reads can be stale

    Likely follow-up: How can a client read its own writes when reading from secondaries?

  19. 19.What MongoDB schema design patterns do you know, and when would you use each?mid

    MongoDB documents a catalog of reusable patterns. The ones I reach for most:

    • Bucket: group many small items, like sensor readings per device per hour, into one document with an array. Fewer documents and smaller indexes
    • Subset: keep only the hot part in the main document, such as the 10 latest reviews, and the rest in another collection
    • Extended reference: copy a few frequently read fields of a referenced document, like a customer's name on each order, to avoid joins
    • Computed: precompute values such as totals or averages on write instead of on every read
    • Outlier: design for the typical case and flag the rare huge document, then overflow its extras elsewhere
    • Attribute: turn many similar fields into an array of key/value pairs so one index covers them
    • Polymorphic: different shapes in one collection, sharing common fields
    • Schema versioning: a schemaVersion field lets old and new shapes coexist during a migration

    Others include approximation, tree, pre-allocation and document versioning. Each trades some write complexity or duplication for faster reads.

    What interviewers listen for
    • Bucket, subset, extended reference, computed, outlier
    • Attribute, polymorphic, schema versioning
    • Patterns are driven by access patterns
    • Trade duplication or write work for read speed
    • Patterns can be combined

    Likely follow-up: How would you model product reviews for a site where a few products have 100,000 reviews?

  20. 20.How do you model one-to-few, one-to-many and one-to-squillions relationships?mid

    The size of the "N" drives the choice, along with whether the children stand alone.

    • One-to-few: embed the children as an array of subdocuments. A person's handful of addresses lives inside the person document, read and written together
    • One-to-many, up to a few hundred or thousand: store an array of references in the parent, such as a product holding the ObjectIds of its parts. Children are their own documents, and the same approach gives you many-to-many without a join table
    • One-to-squillions: put a reference to the parent in each child, like log messages each carrying hostId. An array would grow without bound and hit 16 MiB

    The rules of thumb behind this: favor embedding unless there's a compelling reason not to; needing to access a child on its own is such a reason; arrays should not grow without bound; don't be afraid of application-level joins if you index properly; and weigh read-to-write ratio before denormalizing a field. Everything depends on the access patterns.

    What interviewers listen for
    • One-to-few: embed subdocuments
    • One-to-many: array of child references in the parent
    • One-to-squillions: parent reference in each child
    • Arrays must not grow without bound
    • Denormalize fields that are read often, updated rarely

    Likely follow-up: How would you get the 20 most recent log messages for a host efficiently?

  21. 21.What is the maximum document size in MongoDB, and what do you do if your data needs more?easy

    A BSON document can be at most 16 MiB. The limit exists so a single document can't consume an excessive amount of RAM, or bandwidth when it's sent over the network.

    In practice, a document getting anywhere near that size is a modeling smell, usually an array that keeps growing: comments, events, log lines. Every read and update of that document drags the whole thing through the cache, so performance suffers long before the hard limit. Fixes:

    • Reference the growing children from their own collection instead of embedding them
    • Use the bucket pattern to cap items per document
    • Use the subset or outlier pattern to keep only the hot part inline

    For genuinely large binary content like videos or PDFs, use GridFS, which splits a file into chunks (255 kB by default) stored in fs.chunks with metadata in fs.files. Many teams instead put files in object storage and keep only the URL and metadata in MongoDB.

    What interviewers listen for
    • 16 MiB max BSON document size
    • Near-limit documents signal unbounded arrays
    • Reference, bucket, subset or outlier patterns
    • GridFS splits large files into chunks
    • Object storage plus metadata is a common alternative

    Likely follow-up: When would you choose GridFS over object storage?

  22. 22.What is a covered query, and how do you confirm one?mid

    A covered query is answered entirely from an index, without reading any documents. That's fast because index keys are small and usually in RAM.

    For a query to be covered:

    • every field in the filter is in the same index
    • every field returned is in that index
    • no filter field is compared to null

    The classic trap is _id: it's returned by default, so unless it's part of the index, the projection must say _id: 0, as above. Other limits: a multikey index can't cover a query on its array field, and on a sharded collection the index must contain the shard key for a query through mongos to be covered.

    To confirm, run explain("executionStats"). A covered plan has no FETCH stage, often shows PROJECTION_COVERED, and reports totalDocsExamined: 0 while totalKeysExamined is greater than zero.

    db.users.createIndex({ email: 1, name: 1 });
    
    db.users.find(
      { email: "ana@example.com" },
      { _id: 0, email: 1, name: 1 }
    ).explain("executionStats");
    // PROJECTION_COVERED over IXSCAN, totalDocsExamined: 0
    What interviewers listen for
    • Filter and returned fields all in one index
    • Exclude _id unless it is indexed
    • No FETCH stage; totalDocsExamined is 0
    • Multikey indexes cannot cover their array fields
    • Sharded: index must include the shard key

    Likely follow-up: Why might adding one field to a projection make a fast query slow?

  23. 23.What is an upsert? What document does this create when nothing matches, and how do you make it safe under concurrency?easy

    An upsert is "update if it exists, insert if it doesn't", enabled with { upsert: true } on updateOne, updateMany, replaceOne or findOneAndUpdate.

    If a document matches, it just gets views incremented. If none matches, MongoDB builds a new document from the equality conditions in the filter (page and day), applies the update operators, so views becomes 1, and applies $setOnInsert, which only runs on insert. Range conditions like $gt are not copied into the new document, and an _id is generated if none is given. The result reports upsertedId.

    The concurrency trap: two clients can both find no match and both insert, creating duplicates. The fix is a unique index on the filter fields. Then exactly one insert wins. When the filter is equality on exactly the unique index's fields, the server retries the losers as updates; otherwise they get a duplicate key error that the app should retry.

    db.pageViews.createIndex({ page: 1, day: 1 }, { unique: true });
    
    db.pageViews.updateOne(
      { page: "/pricing", day: "2026-09-28" },
      {
        $inc: { views: 1 },
        $setOnInsert: { firstSeenAt: new Date() }
      },
      { upsert: true }
    );
    What interviewers listen for
    • Update if found, otherwise insert
    • New doc built from filter equality fields plus operators
    • $setOnInsert applies only on insert
    • Concurrent upserts can duplicate without a unique index
    • Unique index on filter fields prevents duplicates

    Likely follow-up: What does upsertedId contain when the upsert updated an existing document?

  24. 24.Explain $push, $addToSet and $pull. How would you keep an array capped at the 50 newest items?easy

    All three change arrays atomically inside one document, without rewriting the array from the app.

    • $push appends a value, creating the array if the field is missing. Pushing an array appends it as one element; use $each to add several
    • $addToSet adds a value only if it isn't already present, giving set semantics. For subdocuments, the comparison is on the whole value, including field order
    • $pull removes every element equal to a value or matching a condition, such as all comments by spammer

    $push also takes modifiers that all require $each: $position to insert at an index, $sort to order the array, and $slice to trim it. They run in a fixed order: insert, then sort, then slice. The feed update above keeps only the 50 newest items, a simple way to stop an array from growing without bound. $pop removes just the first or last element.

    db.posts.updateOne({ _id: 1 }, { $push: { comments: { user: "ana", text: "Nice" } } });
    db.posts.updateOne({ _id: 1 }, { $addToSet: { tags: { $each: ["mongodb", "nosql"] } } });
    db.posts.updateOne({ _id: 1 }, { $pull: { tags: "draft" } });
    db.posts.updateOne({ _id: 1 }, { $pull: { comments: { user: "spammer" } } });
    
    db.feeds.updateOne(
      { _id: 7 },
      { $push: { items: { $each: [{ at: new Date() }], $sort: { at: -1 }, $slice: 50 } } }
    );
    What interviewers listen for
    • $push appends; $each adds several values
    • $addToSet skips values already present
    • $pull removes all matching elements
    • $slice with $each caps array length
    • Modifiers apply: position, sort, then slice

    Likely follow-up: How do you update one specific element inside an array?

  25. 25.What is a projection in MongoDB, and what rules apply when writing one?easy

    A projection is the second argument to find() (or a $project stage) that chooses which fields come back. db.users.find({}, { name: 1, email: 1 }) returns only name, email and _id.

    The rules:

    • _id is included by default; add _id: 0 to drop it
    • You can't mix inclusion (1) and exclusion (0) in one projection, except for _id
    • Use dot notation for nested fields, like "address.city": 1
    • Arrays have projection operators: $slice for the first or last N elements, $elemMatch for the first matching element, and positional $ for the element matched by the filter

    Projections cut network transfer and client memory, and they're required for covered queries. One nuance: unless the query is covered, the server still loads the full document into its cache. Projection shrinks the response, not the disk read, so it isn't a fix for bloated documents.

    What interviewers listen for
    • Selects which fields are returned
    • _id included unless _id: 0
    • No mixing include and exclude, except _id
    • $slice, $elemMatch, $ project array elements
    • Reduces transfer; enables covered queries

    Likely follow-up: How would you return only the last 5 comments of a post?

  26. 26.Why do these two queries return different results for the same document? When do you need $elemMatch?mid

    When you query array fields with separate conditions, each condition can be satisfied by a different element. In the first query, the element with product: "xyz" satisfies the first condition and the element with score: 10 satisfies the second, so the document matches, even though no single result is "xyz with a score of at least 8".

    $elemMatch requires one element to satisfy all conditions together. The xyz element has a score of 5, so the second query correctly finds nothing.

    The same thing happens with scalar arrays: { scores: { $gt: 80, $lt: 85 } } matches [90, 70], because 90 is above 80 and 70 is below 85, while $elemMatch: { $gt: 80, $lt: 85 } needs a value actually between them.

    So I use $elemMatch whenever I have two or more conditions on the same array element. For one condition, plain dot notation is enough. There's also a projection $elemMatch that returns only the first matching element.

    // { _id: 1, results: [ { product: "abc", score: 10 }, { product: "xyz", score: 5 } ] }
    
    db.survey.find({ "results.product": "xyz", "results.score": { $gte: 8 } });
    // matches _id 1
    
    db.survey.find({ results: { $elemMatch: { product: "xyz", score: { $gte: 8 } } } });
    // no match
    What interviewers listen for
    • Separate conditions can match different elements
    • $elemMatch needs one element to match all
    • Applies to subdocument and scalar arrays
    • Use it for two or more conditions on one element
    • Projection $elemMatch returns first matching element

    Likely follow-up: How does a multikey index handle an $elemMatch query with a range?

  27. 27.What is a TTL index, and how precise is its expiry?easy

    A TTL index is a single-field index with expireAfterSeconds. A background task deletes documents once that many seconds have passed since the date in the indexed field. It's ideal for sessions, tokens, caches and temporary logs.

    Two styles: a fixed lifetime (lastSeen plus 30 minutes), or expireAfterSeconds: 0 with a per-document expireAt date, so each document expires at its own time.

    Things to know:

    • It's not precise. The TTL monitor runs every 60 seconds, and under heavy load deletes can lag further, so filter on the date too if stale data mustn't be seen
    • The field must be a BSON Date or an array of dates; with an array, the earliest date counts. Documents where it's missing or not a date never expire
    • It must be a single-field index: compound indexes ignore expireAfterSeconds, and _id can't be a TTL index
    • On a replica set only the primary deletes; the deletes replicate like any other write
    // Delete sessions 30 minutes after lastSeen
    db.sessions.createIndex({ lastSeen: 1 }, { expireAfterSeconds: 1800 });
    
    // Or expire each document at its own time
    db.invites.createIndex({ expireAt: 1 }, { expireAfterSeconds: 0 });
    db.invites.insertOne({ code: "X1", expireAt: new Date("2026-10-01T00:00:00Z") });
    What interviewers listen for
    • Single-field index with expireAfterSeconds
    • Field must be a Date or array of dates
    • Monitor runs every 60 s; expiry is approximate
    • expireAfterSeconds: 0 for per-document expiry times
    • Missing or non-date fields never expire

    Likely follow-up: How do you change the TTL of an existing index?

  28. 28.How do you paginate results in MongoDB? Why is skip() a problem on large collections?mid

    The simple approach is skip((page - 1) * size).limit(size). It works for small data, but the server still has to walk past every skipped entry, so page 5,000 costs far more than page 1. Results can also shift or repeat while users page, because inserts and deletes move the offsets.

    For large or infinite-scroll lists I use range-based (keyset or cursor) pagination. Sort on an indexed key, remember the last item returned, and ask for items after it. Each page is an index seek plus limit, so it costs the same at any depth, and concurrent inserts don't cause duplicates.

    The sort must be unique to be stable. MongoDB's sort isn't stable for equal keys, so add _id as a tie-breaker, as in the snippet, backed by an index like { status: 1, createdAt: -1, _id: -1 }. The trade-off: you can't jump straight to page 37, which most UIs don't need anyway. The client gets an opaque cursor encoding createdAt and _id.

    // First page
    db.posts.find({ status: "published" })
      .sort({ createdAt: -1, _id: -1 }).limit(20);
    
    // Next page: continue after the last post you returned
    db.posts.find({
      status: "published",
      $or: [
        { createdAt: { $lt: last.createdAt } },
        { createdAt: last.createdAt, _id: { $lt: last._id } }
      ]
    }).sort({ createdAt: -1, _id: -1 }).limit(20);
    What interviewers listen for
    • skip walks past skipped entries: cost grows with depth
    • Offsets drift as data changes
    • Range pagination: filter after the last seen key
    • Unique sort with an _id tie-breaker
    • Back it with a matching compound index

    Likely follow-up: How would you show a total page count without an expensive count on every request?

  29. 29.Why does the second insert fail here, and how would a partial index fix it?mid

    A unique index stores a null key for documents where the field is missing, and uniqueness applies to that null like any other value. The first user without an email takes the null slot, and the second collides with it.

    A partial index only indexes documents that match partialFilterExpression. Combined with unique: true, uniqueness is enforced only among documents that have an email, so any number of users can omit it. The filter supports equality, $exists: true, $gt/$lt-style comparisons, $type, $and, $or and $in.

    Partial indexes also shrink index size, for example indexing only { status: "active" } orders. The catch: the planner uses one only when the query's filter implies the partial filter. A query on email alone may not qualify; include the condition, such as email: { $exists: true }.

    Sparse indexes do something similar, skipping documents missing the field, but the manual recommends partial indexes because they're more expressive. Note that creating a unique index fails if existing data already has duplicates.

    db.users.createIndex({ email: 1 }, { unique: true });
    db.users.insertOne({ name: "A" });  // ok
    db.users.insertOne({ name: "B" });  // E11000 duplicate key error
    
    // Better: drop that index and enforce uniqueness only where email exists
    db.users.createIndex(
      { email: 1 },
      { unique: true, partialFilterExpression: { email: { $exists: true } } }
    );
    What interviewers listen for
    • Missing field is indexed as null
    • Unique index allows only one null
    • Partial index indexes only matching documents
    • Query must imply the partial filter to use it
    • Prefer partial over sparse indexes

    Likely follow-up: How would you enforce case-insensitive uniqueness of emails?

  30. 30.How do $unwind and $group work together? Explain this best-sellers report.mid

    $unwind deconstructs an array: an order with three items becomes three documents, each holding one item in items. Documents where the array is missing or empty are dropped unless you pass preserveNullAndEmptyArrays: true.

    $group then collapses documents that share a group key, given as _id. Here each SKU becomes one output document, using accumulators: $sum of quantities, $sum of an expression for revenue, and $avg. Others include $min, $max, $first, $last, $push and $addToSet. Grouping by _id: null aggregates the whole input into one document.

    Things interviewers look for:

    • $group output order is not guaranteed, hence the explicit $sort
    • $group is blocking and each stage is limited to 100 MB of RAM; beyond that it spills to disk when disk use is allowed, which is the default since 6.0, or errors if it isn't
    • $unwind multiplies document counts, so $match first to shrink the input
    // orders: { _id, status, items: [ { sku, qty, price } ] }
    db.orders.aggregate([
      { $match: { status: "paid" } },
      { $unwind: "$items" },
      { $group: {
          _id: "$items.sku",
          unitsSold: { $sum: "$items.qty" },
          revenue: { $sum: { $multiply: ["$items.qty", "$items.price"] } },
          avgQty: { $avg: "$items.qty" }
      } },
      { $sort: { revenue: -1 } }
    ]);
    What interviewers listen for
    • $unwind emits one document per array element
    • Drops empty arrays unless preserveNullAndEmptyArrays
    • $group key is _id; null groups everything
    • Accumulators: $sum, $avg, $push, $addToSet
    • Group output is unordered; 100 MB stage limit

    Likely follow-up: How could you get the top-selling SKU per category in one pipeline?

  31. 31.What is Mongoose, and what are schemas and models?easy

    Mongoose is an ODM (object data modeling library) for Node.js that sits on top of the official MongoDB driver. It adds structure to MongoDB's flexible documents from the application side.

    • A schema declares fields, types, defaults and validators, plus options like timestamps, which adds createdAt and updatedAt
    • A model is compiled from a schema with mongoose.model(). It's the class you query with (User.find, User.create), and it maps to a collection, by default the lowercased plural of the name, so users
    • A document is an instance of a model, with change tracking and save()

    Mongoose casts values (the string "42" becomes a number), validates before saving, and adds middleware, virtuals and populate().

    Gotchas: unique: true is not a validator; it only asks Mongoose to build a unique index. Also, update validators don't run on updateOne and similar unless you pass runValidators: true. Schemas are app-side only; other clients can still write anything unless you also add server-side $jsonSchema validation.

    const mongoose = require("mongoose");
    
    const userSchema = new mongoose.Schema({
      email: { type: String, required: true, unique: true, lowercase: true },
      name: String,
      age: { type: Number, min: 0 },
      roles: { type: [String], default: ["user"] }
    }, { timestamps: true });
    
    const User = mongoose.model("User", userSchema); // collection "users"
    
    await mongoose.connect(process.env.MONGODB_URI);
    await User.create({ email: "Ana@Example.com", name: "Ana" });
    What interviewers listen for
    • ODM for Node.js on top of the driver
    • Schema defines types, defaults, validators
    • Model maps to a collection and runs queries
    • Casting, validation, middleware, populate
    • unique builds an index; it is not a validator

    Likely follow-up: When would you skip Mongoose and use the native driver?

  32. 32.How does Mongoose populate() work, and how is it different from $lookup? When would you use lean()?mid

    You store a reference as { type: Schema.Types.ObjectId, ref: "User" }, then Post.find().populate("author", "name") replaces each author ID with the referenced user document.

    Under the hood, populate is not a server-side join. Mongoose runs the main query, collects the referenced IDs, and runs a separate query per populated path, essentially User.find({ _id: { $in: ids } }), then stitches the results together in Node. That's one extra round trip per path, not one per document, so it avoids N+1. But deep or nested populates add up, and you can't filter the parent by a populated field in the same query. For that, use an aggregation with $lookup, which does the join inside the database.

    Virtual populate covers the reverse side, like a user's posts, via localField and foreignField, without storing an array of IDs on the parent.

    lean() returns plain JavaScript objects instead of full Mongoose documents: no change tracking, getters, virtuals or save(), but much less memory and faster. I use it for read-only endpoints.

    What interviewers listen for
    • ref plus ObjectId declares the relationship
    • Populate runs a separate $in query per path
    • $lookup joins inside the server
    • Virtual populate for the reverse relationship
    • lean() returns plain objects for fast reads

    Likely follow-up: How would you sort posts by their author's name?

  33. 33.A MongoDB-backed API has become slow. How do you find and fix the problem?hard

    I work from evidence, not guesses.

    • Find the slow operations. The slow query log records operations over slowms, 100 ms by default, and the profiler (db.setProfilingLevel(1)) stores them in system.profile. db.currentOp() shows what's running now. On Atlas, the Query Profiler and Performance Advisor do this for you.
    • Explain them. Look for COLLSCAN, in-memory SORT stages and a poor ratio of keys examined to documents returned. Fix with compound indexes following ESR, and covered queries where possible.
    • Check memory. If the working set (hot data plus indexes) doesn't fit in the WiredTiger cache, reads hit disk. Look at cache eviction and page faults.
    • Check the schema. Bloated documents, unbounded arrays and $lookup on every read are design problems that indexes can't fix.
    • Check the client. Use one pooled client per process, project only needed fields, and batch writes with bulkWrite.
    • Check for redundant indexes with $indexStats; they slow writes.

    Only after that would I scale hardware or shard.

    What interviewers listen for
    • Slow query log and profiler first
    • explain(): COLLSCAN, SORT, keys vs docs examined
    • Working set must fit in the WiredTiger cache
    • Schema fixes beat adding hardware
    • Pooled clients, projections, bulk writes

    Likely follow-up: What does a high ratio of documents examined to returned tell you?

  34. 34.What are the most common MongoDB schema and indexing anti-patterns?mid

    The ones MongoDB itself warns about, and I see most:

    • Unbounded arrays: embedding comments, events or followers that grow forever. Documents bloat toward 16 MiB, updates rewrite big documents, and multikey indexes explode
    • Too many indexes: every index slows writes and competes for RAM. Drop unused or redundant ones; { a: 1 } is redundant if { a: 1, b: 1 } exists
    • Bloated documents: storing rarely used large fields with hot data, so every read pulls them into cache. Split them out with the subset pattern
    • Separating data that's accessed together: normalizing like SQL, then $lookup on every request
    • Massive numbers of collections, such as one per user or per day. Each collection and index has overhead
    • Case-insensitive regex queries like /^ana$/i, which can't use an index efficiently. Use a collation-based case-insensitive index instead

    Also watch for deep skip() pagination and read-modify-write updates instead of atomic operators.

    What interviewers listen for
    • Unbounded arrays
    • Unnecessary or redundant indexes
    • Bloated documents and separated hot data
    • Too many collections
    • Case-insensitive regex without a collation index

    Likely follow-up: How do you find indexes that are never used?

  35. 35.What are the costs and limits of multi-document transactions, and how do you use them safely in production?hard

    Transactions work, but they aren't free:

    • Time limit: a transaction must finish within 60 seconds by default (transactionLifetimeLimitSeconds), or it's aborted
    • Lock waits: it waits only 5 ms by default to acquire a lock before aborting
    • Write conflicts: if another operation modifies a document the transaction is writing, the transaction fails with a TransientTransactionError label and must be retried from the start
    • Cache pressure: open transactions pin old snapshots in the WiredTiger cache, so long ones hurt everyone
    • Sharded clusters: a transaction spanning shards uses a two-phase commit, which adds round trips and latency

    Safe usage:

    • Keep transactions short and small; do reads and computation before starting
    • Use the callback API (withTransaction) so transient errors and UnknownTransactionCommitResult are retried
    • Make sure queries inside are indexed
    • Use w: "majority" for commits you can't lose
    • Prefer a single-document design when one is possible

    A standalone server doesn't support transactions, and they can't write to capped collections or the admin, config and local databases.

    What interviewers listen for
    • 60 s default lifetime; 5 ms lock wait
    • Write conflicts raise TransientTransactionError; retry
    • Long transactions pressure the WiredTiger cache
    • Cross-shard commits use two-phase commit
    • Keep them short; use withTransaction retries

    Likely follow-up: What should you do when a commit fails with UnknownTransactionCommitResult?

  36. 36.Compare hashed and ranged sharding. Which would you use for an events collection keyed by timestamp?hard

    Ranged sharding splits data into contiguous ranges of the shard key. Documents with nearby keys live together, so range queries are targeted to few shards, and zones can pin ranges to regions. The weakness is a monotonically increasing key like a timestamp or ObjectId: every new document falls in the top range, so one shard absorbs all inserts while the balancer tries to catch up.

    Hashed sharding shards on a hash of one field. Nearby values scatter, so writes spread evenly even for monotonic keys, and equality queries are still targeted. The cost: range queries become broadcast to all shards, and hashed fields can't be arrays.

    For events keyed by time, I'd avoid a plain ranged { ts: 1 }. Options:

    • { deviceId: 1, ts: 1 } if most queries are per device: spread plus targeted, time-ordered reads
    • { _id: "hashed" } if writes dominate and queries rarely filter by time range
    • A compound key with one hashed field, like { tenantId: 1, ts: "hashed" }, to keep tenant-targeted queries while spreading inserts
    What interviewers listen for
    • Ranged: targeted range queries, monotonic hot spots
    • Hashed: even writes, range queries broadcast
    • Equality queries target a shard in both
    • Compound shard keys can combine both properties
    • Pick based on query and write patterns

    Likely follow-up: How do zones work with ranged sharding?

  37. 37.MongoDB is "schemaless". How do you enforce a schema on the server anyway?mid

    MongoDB is really flexible-schema: by default documents in a collection can differ, but you can attach a validator that the server checks on every insert and update, whichever client writes.

    The usual tool is $jsonSchema, based on JSON Schema draft 4 with MongoDB extensions. bsonType checks BSON types like int, date or objectId; required lists mandatory fields; properties sets per-field rules like minimum, pattern and enum. You can also use query operators in the validator. Add or change rules on an existing collection with collMod.

    Two settings control strictness:

    • validationLevel: strict (default) checks all inserts and updates; moderate doesn't enforce rules on updates to existing documents that were already invalid, which helps when retrofitting rules onto old data
    • validationAction: error (default) rejects the write; warn allows it and logs the violation; recent versions add errorAndLog

    Existing documents aren't re-checked when you add a validator. Users with the right privilege can pass bypassDocumentValidation for migrations.

    db.createCollection("users", {
      validator: { $jsonSchema: {
        bsonType: "object",
        required: ["email", "createdAt"],
        properties: {
          email: { bsonType: "string", pattern: "^.+@.+$" },
          age: { bsonType: "int", minimum: 0 },
          createdAt: { bsonType: "date" }
        }
      } },
      validationLevel: "moderate",
      validationAction: "error"
    });
    What interviewers listen for
    • Validator enforced by the server for all clients
    • $jsonSchema with bsonType, required, properties
    • validationLevel strict (default) or moderate
    • validationAction error (default) or warn
    • Apply to existing collections with collMod

    Likely follow-up: How would you roll out a new required field without breaking old documents?

  38. 38.What are change streams, and how do you make a consumer survive restarts?mid

    Change streams give applications a real-time feed of data changes, built on the oplog, without tailing it by hand. You call watch() on a collection, a database or the whole deployment, and can pass an aggregation pipeline to filter or reshape events. They require a replica set or sharded cluster, and they only report changes committed to a majority, so you never see events that later roll back.

    Each event has an operationType (insert, update, replace, delete, invalidate and others) and documentKey. Updates carry only the changed fields unless you set fullDocument: "updateLookup", which fetches the current version of the document; that may include later changes. Since 6.0 you can enable pre- and post-images on a collection instead.

    For restarts, every event's _id is a resume token. Persist it after processing, and reopen with resumeAfter (or startAfter, which also works after an invalidate). Resuming only works while that point is still in the oplog, so size the oplog for your longest expected outage.

    Typical uses: cache invalidation, syncing a search index, notifications and event-driven microservices.

    const cs = db.orders.watch(
      [{ $match: { operationType: { $in: ["insert", "update"] } } }],
      { fullDocument: "updateLookup" }
    );
    
    while (!cs.isClosed()) {
      if (cs.hasNext()) {
        const ev = cs.next();
        printjson({ op: ev.operationType, id: ev.documentKey._id });
        saveResumeToken(ev._id); // persist it with your own storage
      }
    }
    What interviewers listen for
    • Real-time change feed built on the oplog
    • Needs a replica set or sharded cluster
    • Only majority-committed changes are reported
    • Resume with the saved token via resumeAfter
    • Resumable only within the oplog window

    Likely follow-up: How would you guarantee exactly-once processing of change events?

  39. 39.What is a capped collection, and when would you use one instead of a TTL index?easy

    A capped collection is a fixed-size collection that behaves like a circular buffer. You create it with db.createCollection("log", { capped: true, size: 100000 }), where size is in bytes and an optional max limits the document count. When it's full, the oldest documents are removed automatically to make room.

    It keeps documents in insertion order, so reading the newest entries is cheap, and it supports tailable cursors that stay open and stream new documents as they arrive, like tail -f. MongoDB's own replication oplog is a capped collection.

    Restrictions: capped collections can't be sharded, can't be written inside transactions, can't be the target of $out, and shouldn't be updated in ways that grow documents. With concurrent writers, insertion order isn't guaranteed.

    These days the manual generally recommends TTL indexes instead. Capped collections serialize writes, so they perform worse under concurrency, and TTL lets you expire by age rather than by total size. I'd choose capped only when "keep the last N megabytes" is exactly the requirement.

    What interviewers listen for
    • Fixed size; oldest documents removed automatically
    • Insertion order and tailable cursors
    • The oplog is a capped collection
    • Cannot shard or write in transactions
    • TTL indexes are usually the better choice

    Likely follow-up: What is a tailable cursor?

  40. 41.How do multikey indexes work, and what restrictions and performance traps come with them?hard

    When you index a field that holds an array in any document, MongoDB automatically makes the index multikey: it stores one index entry per array element, all pointing to the same document. That's what makes { tags: "mongodb" } fast.

    Restrictions:

    • In a compound multikey index, each document can have at most one indexed field that is an array. { tags: 1, categories: 1 } is fine until one document has both as arrays, and then that insert fails
    • A multikey index can't be a shard key index, and hashed indexes can't be multikey
    • It can't cover a query that returns the array field

    Performance traps:

    • Index size grows with array length. A document with 5,000 tags adds 5,000 keys, and every push updates the index. That's another reason to keep arrays bounded
    • Bounds: for { scores: { $gt: 80, $lt: 85 } } without $elemMatch, different elements may satisfy each condition, so MongoDB can't simply intersect the bounds; it may scan a wider range. With $elemMatch, it can combine them into one tight range
    What interviewers listen for
    • One index entry per array element
    • Compound: at most one array field per document
    • Cannot be a shard key; hashed cannot be multikey
    • Large arrays inflate index size and write cost
    • $elemMatch allows tighter index bounds

    Likely follow-up: How can you tell from explain() that an index is multikey?

  41. 42.How do you optimize an aggregation pipeline? What does the optimizer already do for you?hard

    The most important rule: only the start of a pipeline can use indexes, so put $match, and $sort when possible, first, and filter as early as you can. On a sharded collection, a leading $match on the shard key also targets fewer shards.

    The optimizer rewrites some things automatically:

    • moves $match filters ahead of $project, $addFields/$set or $sort when they don't depend on computed fields
    • coalesces $sort + $limit into a top-k sort, merges adjacent $match, $limit and $skip stages
    • folds $unwind (and a following $match) into a preceding $lookup
    • analyzes field dependencies so only needed fields flow through

    What I still do by hand:

    • Avoid $unwind + $group just to rebuild an array; use array operators like $filter, $map and $reduce
    • Index the foreignField of every $lookup
    • Keep an eye on the 100 MB per-stage memory limit; spilling to disk works but is slower
    • Precompute heavy reports with $merge into a summary collection
    • Check the result with explain() on the aggregate
    What interviewers listen for
    • Only leading stages can use indexes
    • $match and $sort as early as possible
    • Optimizer reorders and coalesces stages
    • Use array operators instead of unwind-and-regroup
    • 100 MB stage limit; $merge for precomputed results

    Likely follow-up: When can a $sort followed by $group with $first use an index?

  42. 43.What does $facet do? How would you return search results, a total count and filter counts in one query?mid

    $facet runs several sub-pipelines over the same input documents within a single stage. Each sub-pipeline's output becomes an array field in one result document. It's the natural fit for faceted search pages: one round trip returns the first page of results, the total count, counts per brand and price bands.

    Things to know:

    • The sub-pipelines are independent; one can't use another's output. Add stages after $facet to combine them
    • Sub-pipelines don't use indexes. Put a selective $match before $facet, which can use an index; a pipeline that starts with $facet does a collection scan
    • The combined output is a single document, so it must fit in 16 MiB, and each sub-pipeline stage has the usual 100 MB memory limit
    • You can't nest $facet or use stages like $out, $merge or $geoNear inside it

    For very large result sets, separate queries (or Atlas Search's $searchMeta for facet counts) may be cheaper.

    db.products.aggregate([
      { $match: { category: "laptops", price: { $lte: 2000 } } },
      { $facet: {
          results: [{ $sort: { price: 1, _id: 1 } }, { $limit: 20 }],
          total: [{ $count: "count" }],
          byBrand: [{ $group: { _id: "$brand", n: { $sum: 1 } } }, { $sort: { n: -1 } }],
          priceBands: [{ $bucket: {
            groupBy: "$price", boundaries: [0, 500, 1000, 2001], default: "other"
          } }]
      } }
    ]);
    What interviewers listen for
    • Multiple sub-pipelines on the same input
    • Returns one document with an array per facet
    • Sub-pipelines cannot use indexes; $match first
    • Output limited to 16 MiB
    • No nested $facet, $out or $merge inside

    Likely follow-up: Why might a $facet count be slow on a large collection?

  43. 44.Explain replica set member configuration (priority, votes, hidden, delayed, arbiters) and when a rollback happens.hard

    Member options shape elections and roles:

    • priority: higher-priority secondaries call elections sooner and are more likely to win; priority: 0 means the member can never become primary
    • votes: 0 or 1; at most 7 members vote, out of up to 50 total. Non-voting members must have priority 0
    • hidden: priority 0 and invisible to client reads; useful for backups or analytics
    • delayed: a hidden member applying the oplog after a set delay, giving a window to recover from something like an accidental drop
    • arbiter: votes but holds no data; it weakens durability and can make the default write concern w: 1

    Elections use a Raft-like protocol: a candidate needs votes from a majority of voting members, and a primary that loses contact with a majority steps down.

    A rollback happens when the old primary rejoins holding writes that never replicated to the new primary's side. Those writes are undone and saved as BSON files under dbpath/rollback for manual review. Writes acknowledged with w: "majority" are never rolled back.

    What interviewers listen for
    • Priority 0 members never become primary
    • Max 7 voting members of up to 50
    • Hidden and delayed members for backups and recovery
    • Majority vote to elect; isolated primary steps down
    • Rollback removes non-majority writes; majority prevents it

    Likely follow-up: How would you lay out a replica set across three data centers?

  44. 45.What are chunks, how does the balancer move them, and what is a jumbo chunk?hard

    A chunk is a contiguous range of shard key values owned by one shard. The config servers store the mapping, and mongos uses it to route queries. The default range size is 128 MB.

    The balancer runs on the config server primary. In current versions it balances by data size per collection, not chunk count: a round starts when the gap between the shards with the most and least data for a collection reaches the migration threshold, about three times the range size (384 MB by default). Chunks are split when they need to be moved.

    A migration copies the documents to the recipient, catches up on changes made meanwhile, commits the new ownership on the config servers, then deletes the range from the donor asynchronously. Until cleanup, the donor's copies are orphaned documents.

    A jumbo chunk exceeds the range size but can't be split, because it holds a single shard-key value. It signals low cardinality or high frequency; refining the shard key is the real fix. A balancing window can restrict migrations to off-peak hours.

    What interviewers listen for
    • Chunk: contiguous shard-key range owned by one shard
    • Default range size 128 MB
    • Balancer balances by data size in current versions
    • Migrations leave orphans until the donor cleans up
    • Jumbo chunk: single key value too big to split

    Likely follow-up: Why should you stop the balancer during a filesystem snapshot backup?

  45. 46.How do you secure a MongoDB deployment?mid

    I go through the manual's security checklist:

    • Enable access control. Self-managed MongoDB doesn't enforce authentication by default; set security.authorization: enabled. SCRAM is the default mechanism, with x.509 as an option, and LDAP and Kerberos in Enterprise
    • Least-privilege roles: create a user administrator first, then one user per app or person with only the roles it needs, such as read or readWrite on one database, or custom roles. No shared root accounts
    • Limit network exposure: mongod binds to localhost by default; if you widen bindIp, restrict it with firewalls, security groups or private networking. Never expose port 27017 to the internet
    • Encrypt: TLS for all traffic, including between members; encryption at rest; Client-Side Field Level Encryption or Queryable Encryption for sensitive fields
    • Auditing (Enterprise and Atlas), patching, and running mongod as a dedicated OS user

    In application code, guard against operator injection: if a login handler passes req.body.password straight into a filter, an attacker can send { "$ne": null }. Cast inputs to expected types, or enable Mongoose's sanitizeFilter.

    What interviewers listen for
    • Turn on authorization; SCRAM by default
    • Least-privilege RBAC, one user per app
    • Bind to private interfaces, firewall port 27017
    • TLS, encryption at rest, field-level encryption
    • Prevent operator injection from user input

    Likely follow-up: What is the difference between CSFLE and Queryable Encryption?

  46. 47.What are the options for backing up MongoDB, and what are their trade-offs?mid

    First, replication isn't a backup: an accidental deleteMany replicates to every secondary within seconds.

    The main approaches:

    • mongodump / mongorestore: a logical BSON export. Simple and portable, and the manual positions it for small deployments: it's slow on big datasets, indexes are rebuilt on restore, and it adds load. On a replica set, --oplog captures writes during the dump so mongorestore --oplogReplay restores a consistent point
    • Filesystem snapshots (LVM, EBS and similar): fast and suited to large data. Journaling must be on the same volume. For a sharded cluster you must stop the balancer and snapshot every shard and a config server at about the same moment
    • cp / rsync of data files: only with writes stopped
    • Managed backups: Atlas Cloud Backups, or Ops Manager and Cloud Manager, which take snapshots and support point-in-time recovery, including for sharded clusters

    Whatever you choose, define your RPO and RTO, and regularly test restores. A delayed hidden member can help you recover from human error, but it doesn't replace real backups.

    What interviewers listen for
    • Replication is not a backup
    • mongodump for small deployments; --oplog for consistency
    • Filesystem snapshots for large data; journal on same volume
    • Sharded snapshots need the balancer stopped
    • Atlas or Ops Manager for point-in-time restores

    Likely follow-up: How would you restore a single collection that was dropped an hour ago?

  47. 48.What is MongoDB Atlas, and what does it give you over running MongoDB yourself?easy

    Atlas is MongoDB's fully managed cloud database service, running on AWS, Google Cloud and Azure. You choose a cluster tier, cloud and region, and Atlas deploys a replica set or sharded cluster for you.

    What you stop doing yourself:

    • Operations: provisioning, patching, version upgrades and scaling, including storage and compute auto-scaling
    • Backups: cloud snapshots with point-in-time restore
    • Monitoring: metrics, alerts, a Query Profiler and a Performance Advisor that suggests indexes
    • Security baseline: authentication required, TLS, encryption at rest and IP access lists or private networking

    It also bundles platform features that don't exist in a plain mongod: Atlas Search (Lucene-based full-text search), Vector Search, triggers, Charts, multi-region and multi-cloud clusters, and Online Archive for cold data.

    There's a free tier for learning (M0) and paid tiers from small shared clusters up to large dedicated ones. The trade-offs are cost at scale and less low-level control over the servers.

    What interviewers listen for
    • Managed MongoDB on AWS, Google Cloud and Azure
    • Automates patching, scaling and backups
    • Monitoring, Query Profiler, Performance Advisor
    • Secure defaults: auth, TLS, network access lists
    • Adds Atlas Search and Vector Search

    Likely follow-up: How does an application connect to Atlas securely?

  48. 49.What is Mongoose middleware? What is the bug in relying only on this pre("save") hook to hash passwords?mid

    Middleware (hooks) are functions Mongoose runs before (pre) or after (post) an operation. There are four kinds:

    • Document middleware: validate, save, init and others, where this is the document
    • Query middleware: find, findOne, findOneAndUpdate, updateOne, deleteMany and so on, where this is the query, not a document
    • Aggregate middleware for aggregate()
    • Model middleware such as insertMany and bulkWrite

    With an async function you don't need to call next(). Hooks must be registered before mongoose.model() compiles the schema.

    The bug: save hooks don't run for updateOne, findOneAndUpdate and other query-based updates. That update goes straight to the database, so the password is stored in plain text. Fixes: load the document, set the field and call save(); or add a matching pre("updateOne") or pre("findOneAndUpdate") hook that hashes the value in the update; or route all password changes through one service method.

    userSchema.pre("save", async function () {
      if (!this.isModified("password")) return;
      this.password = await bcrypt.hash(this.password, 12);
    });
    
    // Elsewhere in the codebase:
    await User.updateOne({ _id: id }, { password: req.body.password });
    What interviewers listen for
    • Pre and post hooks around operations
    • Document, query, aggregate and model middleware
    • this is the query in query middleware
    • Save hooks do not run on updateOne or findOneAndUpdate
    • Register hooks before compiling the model

    Likely follow-up: How would you implement soft delete with query middleware?

  49. 50.Explain the bucket pattern. How does this update implement it for sensor readings?hard

    The bucket pattern groups many small, related records into one document per bucket, usually per entity per time window, instead of one document per event. A sensor writing every second produces 86,400 tiny documents a day; bucketing turns that into a few hundred documents, with far fewer index entries, better compression and faster range reads.

    This update appends a reading to the sensor's current bucket for the day, but only while count is under 200. When the bucket is full, the filter matches nothing and the upsert creates a new bucket. Its sensorId and day come from the filter's equality fields, while the $lt condition isn't copied. The bucket also keeps precomputed summary fields (count, sumTemp, first, last), so averages and ranges don't need to unwind the array. That's the computed pattern layered on top.

    The count cap keeps arrays bounded. For time-series workloads, time series collections, available since 5.0, apply this bucketing internally, so today I'd try one of those first.

    db.readings.updateOne(
      { sensorId: "s-17", day: ISODate("2026-09-28"), count: { $lt: 200 } },
      {
        $push: { samples: { t: new Date(), temp: 21.4 } },
        $inc: { count: 1, sumTemp: 21.4 },
        $min: { first: new Date() },
        $max: { last: new Date() }
      },
      { upsert: true }
    );
    What interviewers listen for
    • Group many small records into bounded buckets
    • Fewer documents and index entries, faster range reads
    • Upsert with a count cap starts a new bucket
    • Precomputed summaries per bucket
    • Time series collections bucket automatically (5.0+)

    Likely follow-up: How would you query the average temperature for one sensor over a week?

  50. 51.Explain the subset and extended reference patterns with an example of each.hard

    Both reduce what a read has to fetch, at the cost of some duplication.

    Subset pattern. A product page shows the 10 newest reviews, but popular products have thousands. Embedding them all makes the product document huge, and every product read drags them into RAM. Instead, keep all reviews in a reviews collection and embed only the latest 10 in the product. Adding a review is two writes: insert into reviews, then $push onto the product with $each, $sort and $slice: 10. The working set shrinks; "see all reviews" is a second query.

    Extended reference pattern. An order references its customer by _id, but every order view needs the name and shipping address. Rather than a $lookup on each read, copy just those fields into the order: customer: { _id, name, shippingAddress }.

    The trade-off is keeping duplicated data in sync. Pick fields that rarely change, or where a historical snapshot is actually correct: the address an order shipped to shouldn't change later. If updates must propagate, use a background job or a change stream.

    What interviewers listen for
    • Subset: embed only the hot part, rest elsewhere
    • Shrinks documents and the working set
    • Extended reference: copy a few fields of a reference
    • Avoids $lookup on frequent reads
    • Duplicate only stable or snapshot-worthy fields

    Likely follow-up: How would you propagate a customer name change to existing orders?

  51. 52.What are the computed and outlier patterns, and what problems do they solve?hard

    Computed pattern: when the same value is calculated over and over on reads, store the result instead. A movie page shows the average rating and number of ratings; rather than aggregating millions of rating documents each time, every new rating also runs $inc: { ratingCount: 1, ratingSum: 5 } on the movie, and the average is derived from those two fields. For expensive but less urgent figures, a scheduled pipeline with $merge can refresh a summary collection. It suits read-heavy workloads, and the cost is extra write work plus deciding how fresh the numbers need to be.

    Outlier pattern: design for the typical document and handle the rare extreme one separately. Most books have a few hundred buyers in their embedded array, but a bestseller has millions, which would break the 16 MiB limit. Keep the embedded array for everyone, cap it, and set a flag such as hasExtras: true on outliers, with overflow in separate documents. The app checks the flag and fetches the extras only when needed. The common path stays fast; the price is extra application logic.

    What interviewers listen for
    • Computed: store derived values instead of recomputing
    • Update on write with $inc, or batch with $merge
    • Outlier: optimize for the typical document
    • Flag outliers and overflow their extra data
    • Both trade write or app complexity for read speed

    Likely follow-up: How would you keep a computed average correct if ratings can be edited?

  52. 53.How do you update specific elements inside an array? Explain $, $[] and $[<identifier>].mid

    MongoDB has three positional operators for updating array elements in place:

    • $: refers to the first element that matched the array condition in the query filter. The array field must appear in the filter, as items.sku does here, so this increments B's quantity
    • $[]: the all-positional operator, which updates every element of the array
    • $[<identifier>]: the filtered positional operator, which updates every element matching a condition you give in the arrayFilters option. The identifier must start with a lowercase letter, and each identifier needs exactly one array filter

    $[<identifier>] also nests, for example "grades.$[g].scores.$[s]", for arrays inside arrays, which plain $ can't handle.

    All of these are atomic single-document updates, so there's no need to read the array, change it in the app and write it back. If you're matching elements by a unique key like sku, $ is simplest; use arrayFilters when several elements should change.

    // { _id: 1, items: [ { sku: "A", qty: 1 }, { sku: "B", qty: 2 } ] }
    
    // $ : the first element matched by the filter
    db.carts.updateOne({ _id: 1, "items.sku": "B" }, { $inc: { "items.$.qty": 1 } });
    
    // $[] : every element
    db.carts.updateOne({ _id: 1 }, { $set: { "items.$[].reserved": false } });
    
    // $[<identifier>] : every element matching arrayFilters
    db.carts.updateOne(
      { _id: 1 },
      { $set: { "items.$[low].flag": "restock" } },
      { arrayFilters: [{ "low.qty": { $lt: 2 } }] }
    );
    What interviewers listen for
    • $: first element matched by the query filter
    • $[]: all elements
    • $[id] plus arrayFilters: all matching elements
    • Filtered positional works in nested arrays
    • In-place, atomic; no read-modify-write

    Likely follow-up: What happens with $ if the filter matches two array elements?

  53. 54.How would you build a job queue where workers never grab the same job, and how do you do optimistic locking in MongoDB?hard

    Both rely on the fact that a single-document update is atomic, and that the filter is evaluated as part of that atomic write.

    For the queue, findOneAndUpdate finds, modifies and returns one document in one operation. Two workers can race for the same pending job, but only one update can flip it from pending to running; the other matches the next job instead. sort picks the oldest one, and an index on { status: 1, createdAt: 1 } keeps it fast. By default it returns the document before the update; returnDocument: "after" (or returnNewDocument: true in mongosh) returns the new version. Add a lease timeout so jobs from crashed workers can be reclaimed.

    For optimistic locking, keep a version field. Include the version you read in the filter and increment it in the update. If someone else saved first, matchedCount is 0, and you reload and retry. No locks are held between read and write. Mongoose offers the same behavior with its optimisticConcurrency schema option.

    // Claim the oldest pending job atomically
    const job = db.jobs.findOneAndUpdate(
      { status: "pending" },
      { $set: { status: "running", workerId: "w-3", startedAt: new Date() } },
      { sort: { createdAt: 1 }, returnDocument: "after" }
    );
    
    // Optimistic locking: write only if nobody changed it since we read it
    const res = db.docs.updateOne(
      { _id: doc._id, version: doc.version },
      { $set: { body: newBody }, $inc: { version: 1 } }
    );
    if (res.matchedCount === 0) { /* conflict: reload and retry */ }
    What interviewers listen for
    • Filter plus update is one atomic operation
    • findOneAndUpdate claims and returns a document
    • Returns the pre-update doc unless told otherwise
    • Version field in filter for optimistic locking
    • matchedCount of 0 means a conflict

    Likely follow-up: How would you generate a sequential invoice number safely?

  54. 55.What is WiredTiger, and how do its cache, compression and journaling affect you?mid

    WiredTiger is MongoDB's default storage engine, the layer that actually stores documents and indexes on disk.

    • Concurrency: it uses document-level concurrency control, so different clients can write different documents in the same collection simultaneously. Conflicting writes to the same document are detected and retried transparently. It also provides MVCC snapshots, which transactions build on
    • Cache: the internal cache defaults to the larger of 50% of (RAM − 1 GB) or 256 MB. MongoDB also benefits from the OS file cache. Performance depends on the working set, meaning hot documents plus indexes, fitting in memory
    • Compression: collections use snappy block compression by default (zlib and zstd are options), and indexes use prefix compression, which often shrinks disk use a lot
    • Durability: it writes checkpoints every 60 seconds, and a write-ahead journal records changes in between, so after a crash MongoDB recovers from the last checkpoint plus the journal

    When sizing a server, I size RAM so indexes and hot data fit in the cache.

    What interviewers listen for
    • Default storage engine
    • Document-level concurrency, MVCC snapshots
    • Cache: larger of 50% of (RAM − 1 GB) or 256 MB
    • Snappy for data, prefix compression for indexes
    • Checkpoints every 60 s plus a journal

    Likely follow-up: What happens to performance when the working set exceeds the cache?

  55. 56.What is the difference between countDocuments() and estimatedDocumentCount()?easy

    They trade accuracy for speed.

    countDocuments(filter) takes a query filter and actually counts matching documents, running an aggregation under the hood. It's accurate and can use indexes, but on a big result it still has to walk every matching index entry or document, so counting millions of rows isn't free.

    estimatedDocumentCount() takes no filter and reads the collection's metadata, so it returns almost instantly even for huge collections. It can be off after an unclean shutdown, and on a sharded cluster it doesn't filter out orphaned documents.

    The older count() helper is best avoided: without a filter it also relies on metadata and can be approximate, which confused many people. Drivers replaced it with these two explicit methods.

    In practice I use estimatedDocumentCount() for dashboards like "about 12 million users", and countDocuments() with an indexed filter when the number must be exact. For paginated UIs, I avoid exact totals on every request, caching them or showing "more results" instead.

    What interviewers listen for
    • countDocuments is exact and accepts a filter
    • estimatedDocumentCount uses metadata, no filter, fast
    • Estimates can drift after unclean shutdowns or with orphans
    • Avoid the legacy count() helper
    • Exact counts on huge sets are expensive

    Likely follow-up: How would you show the total number of results for a search page cheaply?

  56. 57.What happens when one document fails in insertMany()? How do ordered and unordered bulk writes differ?easy

    insertMany() is not atomic across documents; each insert succeeds or fails on its own.

    • With the default ordered: true, the server inserts in order and stops at the first error, say a duplicate key on document 50. Documents 1 to 49 stay inserted, and nothing after 50 is attempted
    • With ordered: false, it keeps going past errors and reports all failures at the end. It can also be faster, especially on sharded clusters, because operations don't wait for the previous one

    bulkWrite() works the same way but mixes operation types in one call: insertOne, updateOne, updateMany, replaceOne, deleteOne and deleteMany. That's much faster than individual calls because it saves round trips. The driver splits very large requests into batches (the server's maximum is 100,000 operations per batch).

    On failure you get a bulk write error listing each failed operation's index and error, plus counts of what succeeded. If the whole set must be all-or-nothing, run it inside a transaction. Otherwise make the operations idempotent, for example upserts, so a retry is safe.

    What interviewers listen for
    • Not atomic: earlier successes stay inserted
    • Ordered (default) stops at the first error
    • Unordered continues and reports all errors
    • bulkWrite mixes operation types, fewer round trips
    • Use a transaction for all-or-nothing

    Likely follow-up: Why is ordered: false often faster on a sharded cluster?

  57. 58.What is a wildcard index, and when would you use one instead of regular indexes?hard

    A wildcard index indexes fields whose names you don't know in advance. { "$**": 1 } indexes every field in the document, and { "attributes.$**": 1 } indexes every subfield under attributes. You can include or exclude specific paths with wildcardProjection when using $**.

    The typical use case is user-defined or highly variable data: product attributes that differ by category, custom fields in a SaaS app, or event payloads. One wildcard index serves queries like { "attributes.color": "red" } and { "attributes.voltage": 220 } without creating one index per attribute. Since 7.0, compound wildcard indexes such as { tenantId: 1, "attributes.$**": 1 } combine a known field with the wildcard part.

    Limitations:

    • It can't be unique or TTL, and it can't be a shard key
    • It's sparse, so it can't support $exists: false
    • It supports only one field predicate per query, and can't match whole embedded documents or arrays

    It isn't a replacement for designed indexes on known query patterns. The attribute pattern, an array of key/value pairs with a compound index, is the older alternative.

    What interviewers listen for
    • Indexes unknown or variable field names
    • $** for all fields or under a path
    • Compound wildcard indexes since 7.0
    • Not unique, TTL or shard key
    • Attribute pattern is the alternative

    Likely follow-up: How does the attribute pattern compare to a wildcard index?

  58. 59.A user updates their profile, then immediately sees the old version because reads go to secondaries. How do you fix this?hard

    Replication is asynchronous, so a read routed to a secondary can arrive before that secondary has applied the user's write. That breaks read-your-own-writes.

    Options, simplest first:

    • Read from the primary for that flow. It's the default read preference, and many apps only use secondaries for analytics anyway
    • Use a causally consistent session. Within a session, if reads use "majority" read concern and writes use "majority" write concern, MongoDB guarantees read-your-writes, monotonic reads, monotonic writes and writes-follow-reads, even when operations hit different members. The driver tracks the session's latest operation time and sends it with the next read, so the secondary waits until it has caught up to that point before answering. Drivers enable causal consistency on explicit sessions by default; the app must pass the session to every related operation
    • Return the updated document from the write, for example with findOneAndUpdate, and show that, instead of re-reading

    The guarantees apply only within one session, so across different devices or services you still need one of the other approaches.

    What interviewers listen for
    • Async replication makes secondary reads stale
    • Read from primary for read-your-writes flows
    • Causally consistent sessions with majority concerns
    • Secondary waits for the session's operation time
    • Guarantees only hold within one session

    Likely follow-up: What does maxStalenessSeconds do, and why is it not enough here?

Prefer multiple choice? All 22 MongoDB MCQs with answers →

esc