Ch. 7 · Java

java.lang.OutOfMemoryError: Java Heap Space: Diagnosis and Fixes

Fix java.lang.OutOfMemoryError: Java heap space by checking -Xmx, capturing a heap dump and finding what retains memory in Eclipse MAT.

~8 min readadvancedupdated Oct 4, 2026

Your application died, or started failing requests, with this:

Exception in thread "main" java.lang.OutOfMemoryError: Java heap space
	at java.base/java.util.Arrays.copyOf(Arrays.java:3512)
	at java.base/java.util.ArrayList.grow(ArrayList.java:237)
	at java.base/java.util.ArrayList.add(ArrayList.java:483)
	at com.example.report.ExportJob.loadRows(ExportJob.java:41)
Text

It means the JVM needed space for a new object, ran a full garbage collection, and still could not find enough free memory within the maximum heap size. Either the heap is too small for the work you are doing, or something is holding on to objects that are no longer needed (a leak). The frames show where the final allocation failed, which is often an innocent bystander; the real question is what filled the heap before that moment.

Quick fix checklist

  • Confirm the message: Java heap space is about the object heap. Other variants (below) need different fixes.
  • Check the real maximum heap with java -XX:+PrintFlagsFinal -version | grep MaxHeapSize or jcmd <pid> VM.flags.
  • In containers, set -XX:MaxRAMPercentage (for example 75) instead of relying on the default of 25 percent of the memory limit.
  • Add -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/some/dir so the next failure leaves evidence.
  • Open the dump in Eclipse MAT and read the dominator tree to find what retains the most memory.
  • Fix leaks by bounding caches and removing listeners; fix large workloads by streaming or paging data.

Before you start

You need shell access where the JVM runs, a full JDK there if you want jcmd, disk space for a heap dump about the size of the used heap, and Eclipse Memory Analyzer (MAT) on your workstation. Remember the GC basics: an object stays alive while a chain of references leads to it from a GC root (a thread’s stack, a static field).

Recognise the other variants so you do not chase the wrong problem:

  • GC overhead limit exceeded: the Parallel collector spent almost all its time collecting and recovered almost nothing. Treat it like heap space.
  • Metaspace: too many classes loaded, often class loaders leaked by redeploys or generated proxies. More heap will not help.
  • unable to create native thread: the operating system refused another thread. Look at thread counts and process limits, not the heap.

Why it happens

The JVM reserves a heap with a fixed maximum (-Xmx). Without that flag, modern JDKs default to a quarter of the physical memory or of the container’s memory limit. A JVM in a 2 GB container therefore gets roughly a 512 MB heap, which surprises teams who sized the container generously.

When the heap fills, the collector frees unreachable objects. If, after a full collection, the live data plus the new allocation still does not fit, the JVM throws OutOfMemoryError. Two very different situations produce that:

  1. A genuine workload that needs more memory than you gave it: a big batch, a large in-memory index, more concurrent requests.
  2. A leak: objects that are logically dead but still reachable, so the collector must keep them. Usually a collection that only grows (a static map, a cache without eviction, a listener list that is never cleaned up) or a ThreadLocal on a pooled thread.

The distinction matters because more heap fixes the first and only delays the second.

Step-by-step walkthrough

Step 1: Confirm the heap you actually have

# Effective defaults for this JVM on this machine or container
java -XX:+PrintFlagsFinal -version | grep -E 'MaxHeapSize|MaxRAMPercentage'

# Flags of a running process
jcmd <pid> VM.flags
Terminal

MaxHeapSize is in bytes. If it is far below what you expected, set it explicitly. On a VM or bare metal, -Xmx4g is fine. In Kubernetes or Docker, prefer a percentage so the heap follows the container limit:

java -XX:MaxRAMPercentage=75.0 -jar app.jar
Terminal

Leave headroom for thread stacks, metaspace and direct buffers, which live outside the heap. A container that exceeds its limit is killed by the kernel (exit code 137, OOMKilled) with no Java exception at all; that is a different failure.

Step 2: Capture a heap dump

Add these flags permanently in staging and production. They cost nothing until an error occurs:

java -Xmx2g \
     -XX:+HeapDumpOnOutOfMemoryError \
     -XX:HeapDumpPath=/var/dumps \
     -Xlog:gc*:file=/var/log/app/gc.log:time,uptime \
     -jar app.jar
Terminal

If the process is still alive but memory keeps climbing, take a dump on demand. Use an absolute path, because a relative one is resolved against the target process’s working directory:

jcmd <pid> GC.heap_dump /var/dumps/app-before-oom.hprof
jcmd <pid> GC.class_histogram | head -20
Terminal

The class histogram is a quick first look: millions of instances of one domain class is already a lead. A heap dump pauses the application and contains whatever data was in memory, including tokens and personal records, so handle it like a database backup.

Step 3: Read the dominator tree in Eclipse MAT

Open the .hprof file in MAT and let it run the Leak Suspects report. Then open the dominator tree. An object dominates another if every path from a GC root to the second passes through the first, so the dominator tree groups memory by “what would be freed if this object went away”. Sort by retained heap.

Look for one entry that retains a large share of the heap, such as a single HashMap holding 70 percent. Expand it to see its entries, then use “Path to GC Roots” (excluding weak and soft references) to find the static field or thread keeping it alive; that name usually points straight at the responsible code.

If no single object dominates and memory is spread across many request-scoped objects, you are probably looking at a genuinely large workload or too much concurrency rather than a leak.

Step 4: Apply the matching fix and watch the trend

For a leak, bound or release what is held. For a large workload, stop materialising everything at once; the worked scenario below shows both. Afterwards, watch old-generation occupancy after collections in the GC log or with jstat -gcutil <pid> 5000. A healthy service returns to a similar baseline; a leak shows a baseline that climbs for hours.

Worked scenario

A nightly export job built a CSV of all orders and crashed with the trace at the top of this article. The heap dump showed two things in the dominator tree: an ArrayList of 4 million OrderRow objects owned by ExportJob.loadRows, and a static HashMap in CustomerNames holding 900,000 entries.

The broken code:

class CustomerNames {
    private static final Map<Long, String> CACHE = new HashMap<>();

    static String nameFor(long id, CustomerRepository repo) {
        return CACHE.computeIfAbsent(id, repo::findName);
    }
}

class ExportJob {
    List<OrderRow> loadRows(Connection conn) throws SQLException {
        List<OrderRow> rows = new ArrayList<>();
        try (PreparedStatement ps = conn.prepareStatement("SELECT id, customer_id, total FROM orders");
             ResultSet rs = ps.executeQuery()) {
            while (rs.next()) {
                rows.add(new OrderRow(rs.getLong(1), rs.getLong(2), rs.getBigDecimal(3)));
            }
        }
        return rows;
    }
}
java

Diagnosis: the list is a workload problem (the table grows every day, so eventually any heap is too small). The static cache is a leak: it never evicts, and it lives as long as the class loader, so every customer ever exported stays in memory across runs.

The fix streams rows straight to the output and bounds the cache:

class CustomerNames {
    private static final int MAX_ENTRIES = 10_000;

    private static final Map<Long, String> CACHE = Collections.synchronizedMap(
        new LinkedHashMap<>(16, 0.75f, true) {
            @Override
            protected boolean removeEldestEntry(Map.Entry<Long, String> eldest) {
                return size() > MAX_ENTRIES;
            }
        });

    static String nameFor(long id, CustomerRepository repo) {
        return CACHE.computeIfAbsent(id, repo::findName);
    }
}

class ExportJob {
    void export(Connection conn, Writer out, CustomerRepository repo) throws SQLException, IOException {
        conn.setAutoCommit(false); // some drivers (PostgreSQL) only stream with autocommit off
        try (PreparedStatement ps = conn.prepareStatement("SELECT id, customer_id, total FROM orders")) {
            ps.setFetchSize(1_000);
            try (ResultSet rs = ps.executeQuery()) {
                while (rs.next()) {
                    long customerId = rs.getLong(2);
                    out.write(rs.getLong(1) + "," + CustomerNames.nameFor(customerId, repo)
                        + "," + rs.getBigDecimal(3) + "\n");
                }
            }
        }
    }
}
java

Memory use is now proportional to the fetch size and the cache bound, not the table size. A real service might use Caffeine with size and time limits; the point is that every cache needs eviction. Files work the same way: Files.readAllLines loads everything, while a BufferedReader processes one line at a time.

Common mistake

The reflex fix is to double -Xmx and redeploy. For a workload that has genuinely grown, that can be correct. For a leak it just moves the crash from Tuesday to Thursday, and a bigger heap also means longer pauses. Only raise the heap after a dump shows the retained memory is legitimate.

Other mistakes:

  • Calling System.gc() in code. The JVM already collects fully before throwing; an explicit call cannot free reachable objects.
  • Catching OutOfMemoryError and continuing. The error can strike any thread mid-operation and leave data structures half updated. Log, dump and restart (-XX:+ExitOnOutOfMemoryError makes the restart reliable under an orchestrator).
  • Reading the top stack frame as the culprit. The thread that failed to allocate 40 bytes is not necessarily the one that consumed 2 GB.

Verify the behavior

Run the export against a staging copy of the data with a deliberately small heap (-Xmx256m) before and after the change: the old code should throw, the new code should finish, and the post-collection baseline in the GC log should stay flat.

For the cache bound, add a plain unit test so nobody removes the eviction later:

import static org.junit.jupiter.api.Assertions.assertTrue;

import org.junit.jupiter.api.Test;

class CustomerNamesTest {
    @Test
    void cacheNeverExceedsItsBound() {
        CustomerRepository repo = id -> "customer-" + id;
        for (long id = 0; id < 50_000; id++) {
            CustomerNames.nameFor(id, repo);
        }
        assertTrue(CustomerNames.size() <= 10_000);
    }
}
java

This assumes a small package-private static int size() accessor on CustomerNames that returns CACHE.size(), and that CustomerRepository is a functional interface with a single String findName(long id) method.

Interview exercise

Your service runs in a container with a 4 GB memory limit and no JVM flags. It throws OutOfMemoryError: Java heap space under load, but the container’s memory graph shows only about 1.3 GB used. What is going on, and what do you do?

Answer and reasoning

With no -Xmx, the JVM sets the maximum heap to 25 percent of the container limit, so about 1 GB. The heap is full long before the container is, which is why the graph looks fine. First set -XX:MaxRAMPercentage=75.0 (or an explicit -Xmx), leaving the rest for metaspace, thread stacks and native memory. Then make sure the extra heap is needed rather than masking a leak: enable HeapDumpOnOutOfMemoryError, read the dominator tree, and check that the post-GC baseline is stable under steady load. If it keeps rising, the fix is in the code, not the flags.

Continue learning

More in Java

esc