Ch. 25 · Go

Go Worker Pools and Goroutine Leaks Explained

How to bound concurrency in Go with a worker pool or errgroup, why early returns leak goroutines, and how to find leaks with pprof.

~7 min readadvancedupdated Oct 6, 2026

“Write a worker pool in Go” is one of the most common live-coding prompts in backend interviews, often phrased as “fetch 10,000 URLs with at most 20 requests in flight”. The second half of the question is the interesting one: what happens to your goroutines when the caller gives up early, and how would you prove in production that they are not leaking?

Interviewers ask because goroutine leaks are the Go equivalent of a memory leak. They pass every unit test, then show up weeks later as a slow climb in memory and goroutine count until the pod is killed. A correct pool is a lesson in channel ownership, and this guide builds one step by step with Go 1.23.

Before you start

You should know how buffered and unbuffered channels block, that only the sender should close a channel, and how sync.WaitGroup and context.Context work. Familiarity with net/http/pprof is useful for the debugging section.

The short answer

A worker pool has three parts: a producer that sends jobs on a channel and closes it, a fixed number of workers that range over the jobs channel, and a coordinator that waits for every worker with a WaitGroup and then closes the results channel. The concurrency limit is the number of workers. To avoid leaks, every goroutine needs an exit path: pass a context, cancel it when the consumer stops early, and make every blocking send select on ctx.Done(). errgroup.Group with SetLimit packages the same idea when tasks return errors.

How it works

The design rests on one rule: each channel has exactly one owner that closes it, and the owner is whoever knows no more values are coming.

  • The producer owns jobs. It closes jobs after sending the last item, which ends every worker’s range loop.
  • The workers share results as senders, so none of them can close it. A small coordinator goroutine waits for all workers, then closes results.
  • The consumer ranges over results until it is closed.
func Process(ctx context.Context, urls []string, workers int) ([]Result, error) {
	ctx, cancel := context.WithCancel(ctx)
	defer cancel() // runs on every return, releasing all goroutines below

	jobs := make(chan string)
	results := make(chan Result)

	go func() { // producer
		defer close(jobs)
		for _, u := range urls {
			select {
			case jobs <- u:
			case <-ctx.Done():
				return
			}
		}
	}()

	var wg sync.WaitGroup
	for range workers {
		wg.Add(1)
		go func() { // worker
			defer wg.Done()
			for u := range jobs {
				r := fetch(ctx, u)
				select {
				case results <- r:
				case <-ctx.Done():
					return
				}
			}
		}()
	}

	go func() { // coordinator
		wg.Wait()
		close(results)
	}()

	var out []Result
	for r := range results {
		if r.Err != nil {
			return nil, r.Err // deferred cancel unblocks producer and workers
		}
		out = append(out, r)
	}
	return out, nil
}
go

Trace the early return. cancel closes ctx.Done(). The producer is either blocked on jobs <- u or finished; the select lets it exit and close jobs. Workers blocked on results <- r take the ctx.Done() case and return; workers waiting on jobs see it closed and return. When the last worker calls Done, the coordinator closes results. Every goroutine ends.

Step-by-step walkthrough

Step 1: Decide the limit and why

The number of workers is the concurrency limit, so choose it from the bottleneck. For CPU-bound work, runtime.GOMAXPROCS(0) is a sensible start. For I/O-bound work, the limit is usually set by the downstream: a database pool of 20 connections or a partner API that allows 50 requests in parallel. More workers than the downstream can serve only moves the queue somewhere less visible.

Step 2: Make the producer cancellable

Without the select, a producer blocked on jobs <- u never notices cancellation, because no worker will receive again. The producer is the goroutine most often forgotten in leak analysis, since it looks like plain sequential code.

Step 3: Protect every send in the workers

The worker’s select on the results send is the heart of leak prevention. If a worker could block on results <- r with no ctx.Done() alternative, any consumer that stops reading early would strand it. Pass ctx into fetch as well, so in-flight HTTP requests abort promptly instead of running to their own timeout.

Step 4: Use errgroup when tasks return errors

func FetchAll(ctx context.Context, urls []string) ([]Result, error) {
	g, ctx := errgroup.WithContext(ctx)
	g.SetLimit(20) // at most 20 goroutines run at once; g.Go blocks when full

	results := make([]Result, len(urls))
	for i, u := range urls {
		g.Go(func() error {
			r, err := fetchURL(ctx, u)
			if err != nil {
				return fmt.Errorf("fetch %s: %w", u, err)
			}
			results[i] = r // each goroutine writes its own index: no race
			return nil
		})
	}
	if err := g.Wait(); err != nil {
		return nil, err // the first error; ctx was cancelled for the rest
	}
	return results, nil
}
go

errgroup lives in golang.org/x/sync/errgroup. The first non-nil error cancels the derived context, and Wait returns that error after every goroutine finishes. Writing into a preallocated slice by index keeps input order without a results channel.

Step 5: Recover panics per job when one bad input must not crash the pool

func safeFetch(ctx context.Context, u string) (r Result) {
	defer func() {
		if p := recover(); p != nil {
			r = Result{URL: u, Err: fmt.Errorf("panic: %v", p)}
		}
	}()
	return fetch(ctx, u)
}
go

An unrecovered panic in any worker terminates the whole process, and recover only works in the goroutine that panicked. Convert panics to errors inside the worker if inputs are untrusted.

Worked scenario

An image service validates uploaded batches by checking each file against a scanner API. After a week in production, memory grows from 200 MB to 3 GB and the pod is OOM-killed every few days.

// Broken
func Validate(files []string) error {
	jobs := make(chan string, len(files))
	errs := make(chan error) // unbuffered, no context
	for _, f := range files {
		jobs <- f
	}
	close(jobs)

	for range 8 {
		go func() {
			for f := range jobs {
				errs <- scan(f) // blocks forever once Validate has returned
			}
		}()
	}
	for range files {
		if err := <-errs; err != nil {
			return err // early return: up to 8 workers stranded
		}
	}
	return nil
}
go

The goroutine profile tells the story:

curl -s 'http://localhost:6060/debug/pprof/goroutine?debug=1' | head -5
# goroutine profile: total 48213
# 48160 @ 0x... 0x... 0x...
# #	0x...	main.Validate.func1+0x...	/src/validate.go:15
Terminal

Forty-eight thousand goroutines share one stack, parked at line 15, the errs <- scan(f) send. Every batch with an invalid file returned early and stranded up to eight workers, each holding a scanner response and its buffers.

The fix uses errgroup, which cancels the remaining scans and waits for them:

func Validate(ctx context.Context, files []string) error {
	g, ctx := errgroup.WithContext(ctx)
	g.SetLimit(8)
	for _, f := range files {
		g.Go(func() error { return scan(ctx, f) })
	}
	return g.Wait()
}
go

Validate now returns only after every goroutine it started has finished, which is the property that makes leaks impossible by construction. After deploying, the goroutine count graph flattens.

Common mistake

  • Closing the results channel from a worker. The first worker to finish closes it, and the others panic on send.
  • Calling wg.Add inside the worker, which races with wg.Wait in the coordinator.
  • Waiting for workers before reading results on an unbuffered channel. The workers block on send, Wait never returns: a deadlock.
  • A goroutine per item with no limit over a large input, which overwhelms memory or the downstream service.
  • Returning early without cancelling, the root cause of most leaks.

Verify the behavior

func TestProcessReleasesGoroutinesOnError(t *testing.T) {
	defer goleak.VerifyNone(t) // go.uber.org/goleak

	urls := make([]string, 100)
	for i := range urls {
		urls[i] = fmt.Sprintf("u%d", i)
	}
	urls[3] = "bad" // fetch returns an error for this one

	if _, err := Process(context.Background(), urls, 8); err == nil {
		t.Fatal("expected an error")
	}
}
go

goleak.VerifyNone fails the test if any unexpected goroutine is still running after Process returns, printing its stack. Remove the ctx.Done() cases from the pool and this test fails with workers parked on the results send. Run it with go test -race -run TestProcessReleasesGoroutinesOnError -v. In production, export runtime.NumGoroutine() as a metric and alert on sustained growth.

Follow-up questions

How do you add rate limiting on top of a concurrency limit? Use golang.org/x/time/rate: call limiter.Wait(ctx) before each request. Concurrency caps requests in flight; a rate limiter caps requests per second.

How do you keep results in input order with a channel-based pool? Send the index with each job and write results into a preallocated slice, or reorder in the consumer.

When is a buffered results channel enough to prevent leaks? When its capacity is at least the number of sends that can happen after the consumer stops, such as one slot per task for a fixed-size batch.

Why not start one goroutine per item? For hundreds of items it is fine. For millions, memory and downstream pressure make a bound essential.

Interview exercise

A function starts 4 workers that range over a jobs channel pre-filled with 10 items and closed, sends results on an unbuffered results channel, and runs a coordinator that closes results after wg.Wait(). The caller reads 3 results and returns. There is no context. How many goroutines are leaked, and where is each blocked?

Answer and reasoning

Five goroutines leak. Exactly 3 sends complete, and after that every worker ends up blocked on results <- r because nobody will receive again: one holding a result it produced before the caller stopped, the others after taking one more job each. In total 7 jobs are taken (3 delivered, 4 stuck in blocked sends) and 3 stay in the buffer, unprocessed. The coordinator is blocked in wg.Wait(), since no worker will ever call Done. The producer is not leaked because the jobs were buffered and the channel closed before the workers started. Fixing it requires a context cancelled when the caller returns, with each worker’s send in a select on ctx.Done(); then the workers return, the coordinator closes results, and nothing is left. Counting blocked goroutines by where they wait is exactly how you read a goroutine profile.

Continue learning

Practise with the Go interview questions and the Go MCQs. Related notes: Go context cancellation, goroutines and the Go scheduler and backpressure in system design. Primary sources: the Go blog post on pipelines and cancellation, the errgroup package documentation and the net/http/pprof documentation.

More in Go

esc