Compute, storage, databases, VPC networking, IAM, messaging, observability, IaC, security, DR and a 3-tier HA design: the AWS facts interviewers probe.
AMCompiled by Aditya Mishra · Technical Lead, BNP Paribas
The AWS services, limits and trade-offs that come up in cloud and system design interviews, compressed into one sheet.
Global infrastructure & shared model
Region: an isolated geographic area with multiple Availability Zones (AZs): one or more data centers each, with independent power and networking and low-latency links between them.
Edge locations serve CloudFront, Route 53 and Global Accelerator; Local Zones and Outposts bring AWS closer to users or on premises.
Pick a Region for data residency, latency, service availability and price.
Most services are regional; IAM, Route 53, CloudFront and Organizations are global.
Shared responsibility: AWS secures of the cloud (facilities, hardware, network, hypervisor); you secure in the cloud (data, IAM, encryption, network rules, guest OS and app patching).
The line moves with abstraction: on EC2 you patch the OS, on RDS AWS patches the engine, on Lambda you own only code and permissions.
steady usage; Compute SP also covers Fargate and Lambda
Reserved Instances
1 or 3 years, fixed instance attributes
up to 72%
steady workloads; zonal RIs also reserve capacity
Spot
none; reclaimed with a 2-minute warning
up to 90%
fault-tolerant batch, CI, stateless workers
Storage & S3
Option
Type
Scope
Notes
S3
object store
Region (≥3 AZs)
unlimited objects, up to 50 TB each
EBS
block volume for one instance
one AZ
persists after stop; snapshots to move it; gp3 is the default SSD
EFS
managed NFS file system
Region (multi-AZ)
shared by many Linux instances and containers; elastic
Instance store
local disk on the host
the instance
fastest, but data is lost on stop, hibernate or terminate
S3 class
Min duration
Access
Use
Standard
none
milliseconds
hot data
Intelligent-Tiering
none
milliseconds (optional archive tiers)
unknown or changing patterns; no retrieval fee
Standard-IA / One Zone-IA
30 days
milliseconds, retrieval fee
backups; One Zone for re-creatable data
Glacier Instant Retrieval
90 days
milliseconds
archives read about quarterly
Glacier Flexible Retrieval
90 days
minutes to hours (restore first)
archives read about yearly
Glacier Deep Archive
180 days
hours (restore first)
compliance archives, cheapest
Durability is designed for 11 nines (99.999999999%) across classes; IA classes bill a 128 KB minimum object size.
Strong read-after-write consistency for all PUTs, DELETEs and LISTs.
Single PUT up to 5 GB; use multipart upload from about 100 MB (required above 5 GB), up to 50 TB.
Performance: at least 3,500 writes and 5,500 reads per second per prefix; spread keys across prefixes.
New buckets block public access, disable ACLs and encrypt with SSE-S3 by default; use SSE-KMS for key control and audit.
Also: versioning (once on, it can only be suspended), Object Lock (WORM), lifecycle rules, replication (needs versioning on both sides), presigned URLs, event notifications.
relational OLTP; Multi-AZ for HA, read replicas for reads
Aurora
MySQL/PostgreSQL-compatible, storage kept as 6 copies across 3 AZs
higher throughput, fast failover, up to 15 replicas, Global Database, Serverless v2
DynamoDB
serverless key-value/document NoSQL
single-digit-ms at any scale with known access patterns
ElastiCache
in-memory Valkey, Redis OSS, Memcached
caching, sessions, leaderboards, rate limits
Redshift
columnar MPP data warehouse
OLAP over large data; Spectrum queries S3
RDS Multi-AZ = synchronous standby plus automatic DNS failover (HA). Read replica = asynchronous copy for scaling reads (can be promoted, may lag).
DynamoDB primary key: partition key, or partition + sort key. The partition key is hashed, so pick a high-cardinality one to avoid hot partitions.
Query needs the partition key (plus optional sort-key conditions such as begins_with, between); Scan reads the whole table. Both return at most 1 MB per call, so paginate with LastEvaluatedKey.
GSI
LSI
Keys
any partition + sort key
same partition key, different sort key
Created
any time
only at table creation
Consistency
eventually consistent only
strong or eventual
Capacity
its own
shares the table’s
Limit
20 per table (default quota)
5 per table; 10 GB per partition-key value
Item max 400 KB. 1 RCU = one strongly consistent read/s of up to 4 KB (or two eventually consistent); 1 WCU = one write/s of up to 1 KB.
Also: on-demand vs provisioned capacity, Streams, TTL, transactions, DAX cache, global tables (multi-Region, last writer wins).
Networking
VPC: regional, IPv4 CIDR from /16 to /28; a subnet lives in one AZ, and AWS reserves 5 addresses in each.
Public subnet = its route table sends 0.0.0.0/0 to the internet gateway (one per VPC, managed, highly available). Private subnets reach out through a NAT gateway.
A zonal NAT gateway sits in one AZ’s public subnet with an Elastic IP: run one per AZ. Newer regional NAT gateways expand across AZs automatically.
default NACL allows all; a new custom NACL denies all
Sources
CIDRs or other security groups
CIDRs only
VPC peering: 1:1, non-transitive, no overlapping CIDRs, cross-account and cross-Region. Transit Gateway: hub-and-spoke with transitive routing across many VPCs, VPNs and Direct Connect, segmented with route tables.
VPC endpoints: gateway endpoints (S3 and DynamoDB only, a route-table entry, no charge) vs interface endpoints (PrivateLink ENIs with private IPs, most services, hourly and per-GB charges).
Route 53 routing policies: simple, weighted, latency, failover, geolocation, geoproximity, IP-based, multivalue answer. Health checks drive failover; alias records point the zone apex at ELB, CloudFront or S3.
CloudFront: CDN at edge locations in front of S3 (locked down with Origin Access Control), ALB or custom origins; CloudFront Functions and Lambda@Edge run code at the edge.
IAM
Identities: the root user (MFA, then lock away), users (avoid long-lived keys), groups, roles (temporary credentials), and federation through IAM Identity Center, SAML or OIDC.
Policy types: identity-based, resource-based (bucket, KMS key, SQS, Lambda policies), permission boundaries, Organizations SCPs/RCPs and session policies. Boundaries, SCPs/RCPs and session policies only limit; they never grant.
Evaluation: everything starts as an implicit deny → any explicit Deny wins → otherwise it needs an Allow (identity or resource policy in the same account) that every applicable limit also permits. Cross-account access needs an allow on both sides.
Roles & STS: a trust policy says who may assume the role, a permissions policy what it can do. AssumeRole credentials last 1 hour by default (15 minutes up to the role’s 1–12 hour maximum); role chaining caps sessions at 1 hour.
Roles everywhere: EC2 instance profiles, Lambda execution roles, ECS task roles, EKS Pod Identity, cross-account roles with an ExternalId for third parties (confused deputy).
Least privilege: specific actions on specific ARNs, conditions (aws:PrincipalOrgID, aws:MultiFactorAuthPresent), tags for ABAC, IAM Access Analyzer to find unused or external access.
fan-out to SQS, Lambda, HTTP, email, SMS; FIFO topics keep order
none (FIFO topics can archive)
notifying many subscribers
EventBridge
event bus with content-based rules
AWS, SaaS and custom events
optional archive and replay
event-driven integration
Kinesis Data Streams
sharded, replayable log
ordered per shard (partition key); many consumers
24 h default, up to 365 days
real-time streams and analytics
SQS: visibility timeout default 30 s (max 12 h), long polling up to 20 s, delay up to 15 min, batches of 10. Add a DLQ with maxReceiveCount.
Payload caps: SQS and EventBridge accept 1 MiB (older material says 256 KB); SNS topics default to 256 KiB. For bigger data put it in S3 and send a pointer.
Observability & IaC
CloudWatch: metrics (EC2 basic every 5 min, detailed every 1 min; memory and disk need the agent), logs and Logs Insights, alarms, dashboards.
CloudTrail: API audit trail (who did what, when, from where). 90 days of management events in Event history; create a trail to keep logs in S3; data events (S3 object, Lambda invoke) cost extra.
X-Ray: distributed tracing and service maps (instrument with OpenTelemetry). AWS Config: resource configuration history and compliance rules.
CloudFormation
CDK
Terraform
Language
YAML/JSON templates
TypeScript, Python, Java, C#, Go
HCL
State
managed by AWS (stacks)
synthesizes CloudFormation
state file you store (e.g. S3 with locking)
Scope
AWS only
AWS only
multi-cloud and SaaS providers
Strengths
change sets, drift detection, StackSets
loops, types, reusable constructs
plan/apply, huge provider ecosystem
Security services
KMS: HSM-backed keys governed by key policies; envelope encryption (GenerateDataKey: encrypt locally with a data key, store that key encrypted under the KMS key) because Encrypt takes only 4 KB; every use is logged in CloudTrail.
Secrets Manager
Parameter Store
Built for
secrets (DB passwords, API keys)
config and secrets (SecureString via KMS)
Rotation
built in (Lambda), native for RDS/Aurora
none built in
Size
up to 64 KB
4 KB standard, 8 KB advanced
Cost
per secret per month plus API calls
standard tier free, advanced charged
Extras
cross-Region replication
hierarchies (/app/prod/db), versions
WAF: layer-7 web ACLs (managed rules, rate limits, IP sets, SQLi/XSS) on CloudFront, ALB, API Gateway and more.
RPO = how much data you can lose (time since the last good copy). RTO = how long recovery may take.
Strategy
RPO / RTO
DR Region runs
Cost
Backup & restore
hours
backups only; redeploy with IaC
$
Pilot light
tens of minutes
replicated data and core infra; app servers switched off
$$
Warm standby
minutes
a scaled-down but fully working copy
$$$
Multi-site active/active
near zero
full production, serving traffic
$$$$
Replication doesn’t protect against corruption or deletion: keep point-in-time backups too (AWS Backup, versioning), and test restores.
3-tier HA architecture
Edge: Route 53 alias → CloudFront with WAF and Shield; static assets from a private S3 bucket through OAC.
VPC across 2–3 AZs, each with a public subnet (ALB nodes, NAT gateway), a private app subnet and a private data subnet.
App tier: internet-facing ALB → Auto Scaling group (or ECS on Fargate) in every AZ’s app subnet, ELB health checks, stateless servers (sessions in ElastiCache or DynamoDB).
Data tier: Aurora or RDS Multi-AZ in the data subnets, read replicas for read traffic, ElastiCache in front for hot reads.
Chained security groups: ALB allows 443 from anywhere; app allows only the ALB’s SG; the database allows its port only from the app SG.
Egress via NAT gateways, S3/DynamoDB via gateway endpoints, secrets in Secrets Manager, KMS at rest, IAM roles for compute.
Operations: CloudWatch alarms, CloudTrail, AWS Backup, everything in IaC, a DR Region matched to RPO/RTO.
Interview tip
In design rounds, say the failure you are protecting against: an instance (Auto Scaling group), an AZ (multi-AZ subnets, Multi-AZ database) or a Region (DR strategy). Then map each layer to it.
Quick answers
Region vs AZ vs edge location? Geographic area; isolated data center(s) inside it; CDN/DNS point of presence close to users.
Security group vs NACL? Stateful, allow-only, per instance; stateless, allow and deny, per subnet.
How should EC2 access S3? An IAM role via an instance profile (temporary credentials), never stored access keys.
SQS vs SNS? SQS is a pull queue where each message is processed by one consumer; SNS pushes each message to every subscriber.
Multi-AZ vs read replica? Multi-AZ is for availability (synchronous standby, automatic failover); read replicas scale reads asynchronously.
ECS vs EKS? ECS is simpler and AWS-native; EKS gives you Kubernetes APIs and portability. Both run on EC2 or Fargate.
Lambda vs containers? Lambda for short, spiky, event-driven work; containers for long-running or steady services.
CloudTrail vs CloudWatch? CloudTrail records API calls (audit); CloudWatch collects metrics and logs and alarms on them (operations).
Secrets Manager vs Parameter Store? Rotation and cross-Region replication vs cheap hierarchical config (with SecureString for simple secrets).
What makes an app highly available on AWS? Multiple AZs at every tier, a load balancer with health checks, Auto Scaling, a Multi-AZ database and stateless servers.
Gotchas & traps
s3:ListBucket applies to the bucket ARN (arn:aws:s3:::b); object actions apply to arn:aws:s3:::b/*. Mixing them up is the most common policy bug.
In a versioned bucket a delete only adds a delete marker: old versions are still billed until lifecycle rules expire them.
Security groups can’t deny: block an IP with a NACL or WAF.
NACLs are stateless: forget the ephemeral return ports and responses silently drop.
A Lambda function attached to a VPC has no internet access unless its subnets route through a NAT gateway (or use VPC endpoints).
NAT gateway per-GB processing and cross-AZ data transfer add up; send S3/DynamoDB traffic through gateway endpoints.
SQS Standard can deliver duplicates and out of order, so make consumers idempotent; set the visibility timeout longer than processing time.
An ACM certificate for CloudFront must be requested in us-east-1, whatever Region the origin is in.