“What is the difference between a goroutine and a thread?” opens almost every Go interview, and the better interviewers follow it with “So how does the scheduler actually run a million of them?” A one-line answer (“goroutines are lightweight threads”) gets you past the first question and stuck on the second.
Interviewers ask because the scheduler explains real production behaviour: CPU throttling in containers, processes with more threads than CPUs, and why a goroutine per request is fine while a goroutine per row of a huge table is not. This guide targets Go 1.23 and notes where Go 1.25 changed the defaults.
Before you start
You should know what a process and an OS thread are, and that the kernel switches threads on cores with a relatively expensive context switch. Be clear on concurrency (independently progressing tasks) versus parallelism (tasks running at the same instant on different cores), and be able to start a goroutine with go and wait for it with a sync.WaitGroup.
The short answer
A goroutine is a function executing concurrently under the Go runtime’s control, not the operating system’s. It starts with a stack of about 2 KB that grows on demand, so it is far cheaper to create than a thread with a megabyte-scale stack. The runtime multiplexes many goroutines onto a small number of OS threads using an M:N scheduler built from three parts: G (goroutine), M (OS thread) and P (processor, a scheduling context). GOMAXPROCS sets the number of Ps and therefore how many goroutines execute Go code in parallel.
How it works
The scheduler’s three entities fit together like this:
- G is a goroutine: its stack, its instruction pointer and its status (runnable, running, waiting).
- M is an OS thread. Ms are created as needed and can be idle.
- P is a logical processor holding a local run queue of runnable Gs. An M must own a P to execute Go code.
There are exactly GOMAXPROCS Ps. Each M that holds a P repeatedly picks the next G from its P’s local queue and runs it until the G blocks, finishes or is preempted. If the local queue is empty, the M checks the global run queue and the netpoller, then tries to steal half of another P’s queue. It also checks the global queue periodically so goroutines there are not starved.
Blocking is where goroutines earn their reputation:
- Channel, mutex or
time.Sleepwaits park the G in user space. The M immediately runs another G. No thread blocks. - Network I/O goes through the netpoller (epoll on Linux, kqueue on macOS). A goroutine reading an empty socket is parked and made runnable when data arrives, so ten thousand idle connections block zero threads.
- Blocking system calls (many file operations, cgo calls) block the M itself. The runtime detaches the P from that M, using a background monitor thread called
sysmonif the call takes long, and hands the P to another M so the remaining goroutines keep running. That is why a Go process can have far more threads thanGOMAXPROCS.
Stacks grow by copying: when a function prologue finds too little space, the runtime allocates a bigger stack and moves the frames. The default maximum is 1 GB on 64-bit platforms, so infinite recursion ends in goroutine stack exceeds 1000000000-byte limit.
You can inspect the basic numbers from inside a program:
package main
import (
"fmt"
"runtime"
"time"
)
func main() {
fmt.Println("CPUs:", runtime.NumCPU())
fmt.Println("GOMAXPROCS:", runtime.GOMAXPROCS(0)) // 0 means "query, don't change"
for i := 0; i < 10_000; i++ {
go func() { time.Sleep(time.Minute) }()
}
fmt.Println("goroutines:", runtime.NumGoroutine()) // 10001: the sleepers plus main
}The sleepers use roughly 20 MB of stack and create no extra threads: each is a parked G waiting on a runtime timer.
Step-by-step walkthrough
Step 1: Start goroutines and actually wait for them
The first behaviour to state precisely: when main returns, the program exits, and running goroutines are simply abandoned. The runtime does not wait for them.
func main() {
var wg sync.WaitGroup
for i := range 3 {
wg.Add(1) // before the go statement, never inside the goroutine
go func() {
defer wg.Done()
fmt.Println("worker", i)
}()
}
wg.Wait()
}The print order is not deterministic. Because the module targets Go 1.22 or later, each iteration has its own i, so each worker prints a distinct number.
Step 2: See what GOMAXPROCS controls
GOMAXPROCS limits parallel execution of Go code, not the number of goroutines and not the number of threads. A CPU-bound job shows the difference clearly:
func work(n int) int {
sum := 0
for i := range n {
sum += i % 7
}
return sum
}
func run(procs int) time.Duration {
runtime.GOMAXPROCS(procs)
start := time.Now()
var wg sync.WaitGroup
for range 8 {
wg.Add(1)
go func() { defer wg.Done(); work(200_000_000) }()
}
wg.Wait()
return time.Since(start)
}On an 8-core machine, run(1) takes roughly eight times as long as run(8): with one P, the eight goroutines take turns on one thread. They are concurrent in both cases but parallel only in the second. I/O-bound services show a much smaller difference, because most goroutines are parked.
Step 3: Distinguish the three kinds of blocking
When asked “what happens when a goroutine blocks?”, answer by category: a channel or mutex wait parks the G, a socket read parks it on the netpoller, and a blocking syscall blocks the M while the P moves to another M. GODEBUG=schedtrace=1000 prints scheduler state once a second:
GODEBUG=schedtrace=1000 ./server
# SCHED 2004ms: gomaxprocs=8 idleprocs=6 threads=14 spinningthreads=1 ... runqueue=0 [0 1 0 0 0 0 0 0]threads=14 with gomaxprocs=8 is normal: some threads are in syscalls, one runs sysmon, some are idle. The brackets show each P’s local queue length. The runtime aborts if it needs more than 10,000 threads (debug.SetMaxThreads), so goroutines piling up in blocking syscalls can exhaust OS resources.
Step 4: Rely on preemption, but do not abuse it
Before Go 1.14, a goroutine could only be switched out at function calls and other safe points, so a tight loop with no calls could monopolise its P. Go 1.14 added asynchronous preemption: sysmon notices a goroutine that has run for more than about 10 ms and sends its thread a signal (SIGURG on Unix) that forces a switch.
func main() {
runtime.GOMAXPROCS(1)
go func() {
for { // no function calls: pre-1.14 this starved main forever
}
}()
time.Sleep(50 * time.Millisecond)
fmt.Println("main still runs") // printed on Go 1.14+
}runtime.Gosched() is therefore almost never needed. Preemption keeps the program responsive, but the busy loop still burns a full core.
Worked scenario
A payments API runs in Kubernetes with a CPU limit of 2 on nodes with 64 cores. It is built with Go 1.23. Under moderate load, p99 latency jumps from 15 ms to over 200 ms, while average CPU usage looks low.
The cause is the default GOMAXPROCS. Before Go 1.25 it equals the CPUs the process can see (64), not the cgroup quota (2). Up to 64 threads run Go code, including parallel GC workers, so the container spends its 200 ms of quota per 100 ms CFS period in a few milliseconds and is throttled for the rest. Requests arriving meanwhile wait, which shows up as tail latency.
# Broken: 64 Ps competing for a 2-CPU quota
resources:
limits:
cpu: "2"There are three fixes, from simplest to most robust:
# Fix A: tell the runtime explicitly (any Go version)
env:
- name: GOMAXPROCS
value: "2"// Fix B: on Go 1.24 or older, derive it from the cgroup at startup
import _ "go.uber.org/automaxprocs"Fix C is to build with Go 1.25 or later (with the module’s go line at 1.25 or higher, since the new default is tied to it), whose default GOMAXPROCS on Linux takes the cgroup CPU limit into account and is re-checked periodically if the limit changes. Setting GOMAXPROCS in the environment or calling runtime.GOMAXPROCS disables it, so remove stale overrides when you upgrade. Afterwards container_cpu_cfs_throttled_periods_total should drop to near zero.
Common mistake
- “Goroutines are green threads, so Go is single-threaded.” Goroutines run in parallel on up to
GOMAXPROCSthreads at once. - “GOMAXPROCS limits how many goroutines or threads can exist.” It limits only the Ps. A program can have millions of goroutines and many more threads than Ps.
- “Goroutines are free.” Each costs kilobytes and holds whatever it references; unbounded fan-out can exhaust memory or flood a downstream service.
- “Use runtime.Gosched() to be fair.” Since Go 1.14 the scheduler preempts on its own.
Verify the behavior
A small test makes the “main does not wait” and “goroutines are cheap” claims concrete:
func TestManyGoroutinesAreCheap(t *testing.T) {
before := runtime.NumGoroutine()
stop := make(chan struct{})
var started sync.WaitGroup
for range 100_000 {
started.Add(1)
go func() { started.Done(); <-stop }()
}
started.Wait()
if got := runtime.NumGoroutine() - before; got < 100_000 {
t.Fatalf("expected 100000 extra goroutines, got %d", got)
}
close(stop) // release them so the test does not leak
}Run it with go test -run TestManyGoroutinesAreCheap -v and watch top: memory rises by a few hundred megabytes at most and the thread count stays small. For behaviour under load, go test -trace trace.out then go tool trace trace.out shows goroutines moving between Ps.
Follow-up questions
Why can a Go process have more threads than GOMAXPROCS? Threads blocked in syscalls or cgo calls release their P but still exist, and the runtime keeps idle threads and a sysmon thread around.
How does this compare to Java virtual threads? Both are M:N user-space scheduling; Go has had it from the start, so every library assumes it.
How big can a goroutine’s stack get? Up to 1 GB on 64-bit systems by default, configurable with debug.SetMaxStack. Deep recursion beyond that is a fatal error.
Interview exercise
A service sets GOMAXPROCS=4. It receives 2,000 concurrent requests; each one makes an outbound HTTPS call taking 300 ms and then spends 2 ms of CPU building a response. How many OS threads do you expect, and will the requests be served in parallel?
Answer and reasoning
Expect a small number of threads, perhaps a dozen, not 2,000. While waiting 300 ms for the network, each request’s goroutine is parked on the netpoller and holds no thread. Only the 2 ms CPU phases need a P, and at most four run at once. Since 2,000 requests × 2 ms is 4 seconds of CPU work spread over 4 Ps, the CPU part alone needs about one second, so the requests overlap heavily and finish in roughly 1.3 seconds of wall time rather than 2,000 × 302 ms. The reasoning shows you understand that concurrency comes from parking waiting goroutines cheaply, while parallelism is capped by the number of Ps.
Continue learning
Practise with the Go interview questions and the Go MCQs. Related notes: Go worker pools and goroutine leaks, processes versus threads and context switching overhead. For primary sources, read the Go FAQ on goroutines, the runtime package documentation and the Go 1.25 release notes for container-aware GOMAXPROCS.