“Write a worker pool in Go” is one of the most common live-coding prompts in backend interviews, often phrased as “fetch 10,000 URLs with at most 20 requests in flight”. The second half of the question is the interesting one: what happens to your goroutines when the caller gives up early, and how would you prove in production that they are not leaking?
Interviewers ask because goroutine leaks are the Go equivalent of a memory leak. They pass every unit test, then show up weeks later as a slow climb in memory and goroutine count until the pod is killed. A correct pool is a lesson in channel ownership, and this guide builds one step by step with Go 1.23.
Before you start
You should know how buffered and unbuffered channels block, that only the sender should close a channel, and how sync.WaitGroup and context.Context work. Familiarity with net/http/pprof is useful for the debugging section.
The short answer
A worker pool has three parts: a producer that sends jobs on a channel and closes it, a fixed number of workers that range over the jobs channel, and a coordinator that waits for every worker with a WaitGroup and then closes the results channel. The concurrency limit is the number of workers. To avoid leaks, every goroutine needs an exit path: pass a context, cancel it when the consumer stops early, and make every blocking send select on ctx.Done(). errgroup.Group with SetLimit packages the same idea when tasks return errors.
How it works
The design rests on one rule: each channel has exactly one owner that closes it, and the owner is whoever knows no more values are coming.
- The producer owns
jobs. It closesjobsafter sending the last item, which ends every worker’srangeloop. - The workers share
resultsas senders, so none of them can close it. A small coordinator goroutine waits for all workers, then closesresults. - The consumer ranges over
resultsuntil it is closed.
func Process(ctx context.Context, urls []string, workers int) ([]Result, error) {
ctx, cancel := context.WithCancel(ctx)
defer cancel() // runs on every return, releasing all goroutines below
jobs := make(chan string)
results := make(chan Result)
go func() { // producer
defer close(jobs)
for _, u := range urls {
select {
case jobs <- u:
case <-ctx.Done():
return
}
}
}()
var wg sync.WaitGroup
for range workers {
wg.Add(1)
go func() { // worker
defer wg.Done()
for u := range jobs {
r := fetch(ctx, u)
select {
case results <- r:
case <-ctx.Done():
return
}
}
}()
}
go func() { // coordinator
wg.Wait()
close(results)
}()
var out []Result
for r := range results {
if r.Err != nil {
return nil, r.Err // deferred cancel unblocks producer and workers
}
out = append(out, r)
}
return out, nil
}Trace the early return. cancel closes ctx.Done(). The producer is either blocked on jobs <- u or finished; the select lets it exit and close jobs. Workers blocked on results <- r take the ctx.Done() case and return; workers waiting on jobs see it closed and return. When the last worker calls Done, the coordinator closes results. Every goroutine ends.
Step-by-step walkthrough
Step 1: Decide the limit and why
The number of workers is the concurrency limit, so choose it from the bottleneck. For CPU-bound work, runtime.GOMAXPROCS(0) is a sensible start. For I/O-bound work, the limit is usually set by the downstream: a database pool of 20 connections or a partner API that allows 50 requests in parallel. More workers than the downstream can serve only moves the queue somewhere less visible.
Step 2: Make the producer cancellable
Without the select, a producer blocked on jobs <- u never notices cancellation, because no worker will receive again. The producer is the goroutine most often forgotten in leak analysis, since it looks like plain sequential code.
Step 3: Protect every send in the workers
The worker’s select on the results send is the heart of leak prevention. If a worker could block on results <- r with no ctx.Done() alternative, any consumer that stops reading early would strand it. Pass ctx into fetch as well, so in-flight HTTP requests abort promptly instead of running to their own timeout.
Step 4: Use errgroup when tasks return errors
func FetchAll(ctx context.Context, urls []string) ([]Result, error) {
g, ctx := errgroup.WithContext(ctx)
g.SetLimit(20) // at most 20 goroutines run at once; g.Go blocks when full
results := make([]Result, len(urls))
for i, u := range urls {
g.Go(func() error {
r, err := fetchURL(ctx, u)
if err != nil {
return fmt.Errorf("fetch %s: %w", u, err)
}
results[i] = r // each goroutine writes its own index: no race
return nil
})
}
if err := g.Wait(); err != nil {
return nil, err // the first error; ctx was cancelled for the rest
}
return results, nil
}errgroup lives in golang.org/x/sync/errgroup. The first non-nil error cancels the derived context, and Wait returns that error after every goroutine finishes. Writing into a preallocated slice by index keeps input order without a results channel.
Step 5: Recover panics per job when one bad input must not crash the pool
func safeFetch(ctx context.Context, u string) (r Result) {
defer func() {
if p := recover(); p != nil {
r = Result{URL: u, Err: fmt.Errorf("panic: %v", p)}
}
}()
return fetch(ctx, u)
}An unrecovered panic in any worker terminates the whole process, and recover only works in the goroutine that panicked. Convert panics to errors inside the worker if inputs are untrusted.
Worked scenario
An image service validates uploaded batches by checking each file against a scanner API. After a week in production, memory grows from 200 MB to 3 GB and the pod is OOM-killed every few days.
// Broken
func Validate(files []string) error {
jobs := make(chan string, len(files))
errs := make(chan error) // unbuffered, no context
for _, f := range files {
jobs <- f
}
close(jobs)
for range 8 {
go func() {
for f := range jobs {
errs <- scan(f) // blocks forever once Validate has returned
}
}()
}
for range files {
if err := <-errs; err != nil {
return err // early return: up to 8 workers stranded
}
}
return nil
}The goroutine profile tells the story:
curl -s 'http://localhost:6060/debug/pprof/goroutine?debug=1' | head -5
# goroutine profile: total 48213
# 48160 @ 0x... 0x... 0x...
# # 0x... main.Validate.func1+0x... /src/validate.go:15Forty-eight thousand goroutines share one stack, parked at line 15, the errs <- scan(f) send. Every batch with an invalid file returned early and stranded up to eight workers, each holding a scanner response and its buffers.
The fix uses errgroup, which cancels the remaining scans and waits for them:
func Validate(ctx context.Context, files []string) error {
g, ctx := errgroup.WithContext(ctx)
g.SetLimit(8)
for _, f := range files {
g.Go(func() error { return scan(ctx, f) })
}
return g.Wait()
}Validate now returns only after every goroutine it started has finished, which is the property that makes leaks impossible by construction. After deploying, the goroutine count graph flattens.
Common mistake
- Closing the results channel from a worker. The first worker to finish closes it, and the others panic on send.
- Calling
wg.Addinside the worker, which races withwg.Waitin the coordinator. - Waiting for workers before reading results on an unbuffered channel. The workers block on send,
Waitnever returns: a deadlock. - A goroutine per item with no limit over a large input, which overwhelms memory or the downstream service.
- Returning early without cancelling, the root cause of most leaks.
Verify the behavior
func TestProcessReleasesGoroutinesOnError(t *testing.T) {
defer goleak.VerifyNone(t) // go.uber.org/goleak
urls := make([]string, 100)
for i := range urls {
urls[i] = fmt.Sprintf("u%d", i)
}
urls[3] = "bad" // fetch returns an error for this one
if _, err := Process(context.Background(), urls, 8); err == nil {
t.Fatal("expected an error")
}
}goleak.VerifyNone fails the test if any unexpected goroutine is still running after Process returns, printing its stack. Remove the ctx.Done() cases from the pool and this test fails with workers parked on the results send. Run it with go test -race -run TestProcessReleasesGoroutinesOnError -v. In production, export runtime.NumGoroutine() as a metric and alert on sustained growth.
Follow-up questions
How do you add rate limiting on top of a concurrency limit? Use golang.org/x/time/rate: call limiter.Wait(ctx) before each request. Concurrency caps requests in flight; a rate limiter caps requests per second.
How do you keep results in input order with a channel-based pool? Send the index with each job and write results into a preallocated slice, or reorder in the consumer.
When is a buffered results channel enough to prevent leaks? When its capacity is at least the number of sends that can happen after the consumer stops, such as one slot per task for a fixed-size batch.
Why not start one goroutine per item? For hundreds of items it is fine. For millions, memory and downstream pressure make a bound essential.
Interview exercise
A function starts 4 workers that range over a jobs channel pre-filled with 10 items and closed, sends results on an unbuffered results channel, and runs a coordinator that closes results after wg.Wait(). The caller reads 3 results and returns. There is no context. How many goroutines are leaked, and where is each blocked?
Answer and reasoning
Five goroutines leak. Exactly 3 sends complete, and after that every worker ends up blocked on results <- r because nobody will receive again: one holding a result it produced before the caller stopped, the others after taking one more job each. In total 7 jobs are taken (3 delivered, 4 stuck in blocked sends) and 3 stay in the buffer, unprocessed. The coordinator is blocked in wg.Wait(), since no worker will ever call Done. The producer is not leaked because the jobs were buffered and the channel closed before the workers started. Fixing it requires a context cancelled when the caller returns, with each worker’s send in a select on ctx.Done(); then the workers return, the coordinator closes results, and nothing is left. Counting blocked goroutines by where they wait is exactly how you read a goroutine profile.
Continue learning
Practise with the Go interview questions and the Go MCQs. Related notes: Go context cancellation, goroutines and the Go scheduler and backpressure in system design. Primary sources: the Go blog post on pipelines and cancellation, the errgroup package documentation and the net/http/pprof documentation.