“What is a data race, and how would you find one in Go?” is a staple of Go interviews, usually followed by a snippet where a thousand goroutines increment a counter and the candidate must say what it prints. Senior rounds add a twist: the service crashed with fatal error: concurrent map writes, so what happened and why did recover not help?
Interviewers ask because races are the most common correctness bug in concurrent Go code and because Go ships an unusually good tool for them. Knowing what the race detector can and cannot prove separates “I added a mutex until it stopped failing” from a reasoned fix. Examples target Go 1.23, with Go 1.25’s WaitGroup.Go noted.
Before you start
You should know how to start goroutines and wait for them with sync.WaitGroup, and the basics of methods with pointer receivers. It helps to know that compilers and CPUs reorder memory operations when they can prove a single thread will not notice, which is exactly what makes unsynchronised sharing unpredictable.
The short answer
A data race happens when two goroutines access the same variable concurrently, at least one access is a write, and no synchronisation orders them. The Go memory model gives no useful guarantee for such a program: updates can be lost, values can be stale or even half-written. You fix races by synchronising with a sync.Mutex, an atomic from sync/atomic, or a channel that hands ownership from one goroutine to another, and you find them with the race detector, go test -race, which reports conflicting accesses that actually happen while the tests run.
How it works
Consider the classic counter:
var count int
var wg sync.WaitGroup
for range 1000 {
wg.Add(1)
go func() {
defer wg.Done()
count++ // load count, add 1, store count
}()
}
wg.Wait()
fmt.Println(count) // usually less than 1000, and different on each runcount++ is three steps. Two goroutines can both load 41, both add one and both store 42, losing an increment. Even with GOMAXPROCS=1 the program is still racy by the language’s rules; it may just fail less often.
The memory model defines which writes a read is guaranteed to see. The guarantee comes only from synchronized-before edges: an Unlock before the next Lock of the same mutex, a channel send before the matching receive completes, an atomic store before an atomic load that observes it, WaitGroup.Done before the Wait it releases. Without such an edge, there is no “before”, so the compiler may keep a value in a register or reorder stores.
The race detector is built on ThreadSanitizer. When you compile with -race, every memory access is instrumented, and the runtime tracks the happens-before relation between goroutines using vector clocks. When two accesses to the same address are not ordered and one is a write, it prints a report with both stacks and where each goroutine was created. Instrumentation typically raises memory use five to ten times and slows execution two to twenty times, so it belongs in tests, CI and canaries rather than all of production. A program that found races exits with status 66 by default.
Step-by-step walkthrough
Step 1: Let the detector show you the race
go run -race ./cmd/counter==================
WARNING: DATA RACE
Read at 0x00c000012118 by goroutine 8:
main.main.func1()
/src/cmd/counter/main.go:14 +0x...
Previous write at 0x00c000012118 by goroutine 7:
main.main.func1()
/src/cmd/counter/main.go:14 +0x...
==================
Found 1 data race(s)
exit status 66The real report also shows where each goroutine was created; it is shortened here. Two goroutines touch the same address on line 14, one reading and one writing. The report names the accesses, not the fix; deciding how to synchronise is your job.
Step 2: Guard shared state with a mutex inside a type
type Counter struct {
mu sync.Mutex
n map[string]int
}
func (c *Counter) Inc(key string) {
c.mu.Lock()
defer c.mu.Unlock()
c.n[key]++
}
func (c *Counter) Get(key string) int {
c.mu.Lock()
defer c.mu.Unlock()
return c.n[key]
}Keep the mutex next to the fields it protects, unexported, and lock in every method that reads or writes them. Pointer receivers are essential: a value receiver would copy the struct and its mutex, so each call would lock a different lock. go vet reports copied locks.
Step 3: Use atomics for single values
var count atomic.Int64 // typed atomics since Go 1.19
for range 1000 {
wg.Add(1)
go func() {
defer wg.Done()
count.Add(1)
}()
}
wg.Wait()
fmt.Println(count.Load()) // always 1000Atomics suit independent counters, flags (atomic.Bool) and pointer swaps (atomic.Pointer[T]). As soon as two values must change together, such as a balance and a version, use a mutex; two separate atomics do not form a transaction.
Step 4: Get WaitGroup ordering right
var wg sync.WaitGroup
for _, job := range jobs {
wg.Add(1) // before starting the goroutine
go func() {
defer wg.Done()
run(job)
}()
}
wg.Wait()
// Go 1.25 and later: Add and Done handled for you
// wg.Go(func() { run(job) })Calling Add inside the goroutine races with Wait, which may see a zero counter and return before the work starts. Pass a *sync.WaitGroup, never a copy.
Step 5: Prefer simple locks, then measure
sync.RWMutex lets many readers hold the lock at once and is worth it only when reads dominate and the critical section is long enough for the extra bookkeeping to pay off. Mutexes are not reentrant: calling Lock twice on the same mutex from one goroutine deadlocks. Never hold a lock while doing network I/O or calling a user callback, because every other goroutine then waits on a remote system.
Worked scenario
A pricing service keeps exchange rates in a map. Handlers read it; a background goroutine refreshes it every minute. It works for weeks, then crashes several pods at once with:
fatal error: concurrent map read and map write// Broken
type Rates struct{ m map[string]float64 }
func (r *Rates) Get(cur string) float64 { return r.m[cur] } // handlers
func (r *Rates) refresh(latest map[string]float64) { // background goroutine
for k, v := range latest {
r.m[k] = v
}
}The runtime detects some concurrent map misuse and aborts the process. This is a fatal error, not a panic, so recover in the HTTP server’s handler wrapper cannot catch it, and every in-flight request on the pod dies. The crash appeared only when a refresh happened to coincide with heavy read traffic, which is why it took weeks.
Two good fixes. The direct one locks every access:
type Rates struct {
mu sync.RWMutex
m map[string]float64
}
func (r *Rates) Get(cur string) float64 {
r.mu.RLock()
defer r.mu.RUnlock()
return r.m[cur]
}with refresh taking r.mu.Lock(). Because refreshes replace the whole data set, an even simpler design builds a new map and swaps a pointer, so readers never lock:
type Rates struct{ m atomic.Pointer[map[string]float64] }
func (r *Rates) Get(cur string) float64 {
if m := r.m.Load(); m != nil {
return (*m)[cur]
}
return 0
}
func (r *Rates) Replace(latest map[string]float64) {
r.m.Store(&latest) // latest must not be modified after this call
}Readers see either the old map or the new one, never a half-updated one, and the old map is garbage-collected once no reader holds it.
Common mistake
- “The race detector proves my code is race-free.” It reports only races that execute during the run. Untested paths stay unchecked.
- “Writing an int is atomic on x86, so it is fine.” The compiler can cache, reorder or eliminate unsynchronised accesses regardless of hardware.
- Locking writes but not reads. A read concurrent with a write is still a race.
- Value receivers on types with a mutex, which copy the lock.
- Trying to
recoverfrom concurrent map writes. It is a fatal runtime error that ends the process.
Verify the behavior
func TestRatesConcurrentAccess(t *testing.T) {
var r Rates
r.Replace(map[string]float64{"EUR": 1.1})
var wg sync.WaitGroup
for i := range 50 {
wg.Add(2)
go func() { defer wg.Done(); _ = r.Get("EUR") }()
go func() {
defer wg.Done()
r.Replace(map[string]float64{"EUR": 1.0 + float64(i)/100})
}()
}
wg.Wait()
}Run it with go test -race -count=20 -run TestRatesConcurrentAccess. If Replace writes into the shared map without a lock, as refresh did, the detector reports the conflicting read and write, usually on the first run. Against either fixed version, it stays quiet across repetitions. Make -race the default in CI for packages with concurrency.
Follow-up questions
When is sync.Map the right choice? When keys are written once and read many times, or goroutines work on disjoint key sets. For general use, a map plus a mutex is clearer and typed.
Is Go’s mutex fair? Mostly. Since Go 1.9, a waiter that fails to acquire the lock for over 1 ms switches it into starvation mode, where ownership passes directly to the waiting goroutines in order.
What does TryLock do? Added in Go 1.18, it acquires the lock only if it is free. Correct uses are rare; it often hides a design problem.
Can a program with no races still deadlock? Yes. Lock ordering problems and forgotten unlocks cause deadlocks; the race detector does not look for them.
Interview exercise
In the counter program from “How it works”, what values can be printed? Would it be correct with GOMAXPROCS=1? Name two minimal fixes and say which you would choose.
Answer and reasoning
In principle any value up to 1000 can be printed, most often something slightly below 1000, and it varies between runs because concurrent increments overwrite each other. With GOMAXPROCS=1 it may print 1000 more often, since goroutines rarely interleave inside count++, but the program is still a data race under the memory model and the detector still reports it. The two minimal fixes are a sync.Mutex around the increment or an atomic.Int64 with Add(1). For a single counter the atomic is the better choice: it is simpler, cannot be forgotten in a code path, and avoids lock contention. If the counter later had to change together with other fields, a mutex would be the right tool.
Continue learning
Practise with the Go interview questions and the Go MCQs. Related notes: race conditions and atomic state transitions, mutexes versus semaphores and Java thread visibility. Primary sources: the Go memory model, the data race detector guide and the sync package documentation.