AWS · cheat sheet

AWS

Compute, storage, databases, VPC networking, IAM, messaging, observability, IaC, security, DR and a 3-tier HA design: the AWS facts interviewers probe.

The AWS services, limits and trade-offs that come up in cloud and system design interviews, compressed into one sheet.

Global infrastructure & shared model

  • Region: an isolated geographic area with multiple Availability Zones (AZs): one or more data centers each, with independent power and networking and low-latency links between them.
  • Edge locations serve CloudFront, Route 53 and Global Accelerator; Local Zones and Outposts bring AWS closer to users or on premises.
  • Pick a Region for data residency, latency, service availability and price.
  • Most services are regional; IAM, Route 53, CloudFront and Organizations are global.
  • Shared responsibility: AWS secures of the cloud (facilities, hardware, network, hypervisor); you secure in the cloud (data, IAM, encryption, network rules, guest OS and app patching).
  • The line moves with abstraction: on EC2 you patch the OS, on RDS AWS patches the engine, on Lambda you own only code and permissions.

Compute

Service What Use when
EC2 virtual machines, full control custom OS/agents, steady long-running work, lift-and-shift
Lambda functions per event, pay per request and duration spiky, event-driven work under 15 min
ECS AWS-native container orchestrator containers without Kubernetes complexity
EKS managed Kubernetes control plane Kubernetes portability and ecosystem
Fargate serverless compute for ECS tasks and EKS pods containers without managing nodes
Elastic Beanstalk PaaS: provisions EC2, ASG, ELB for your code quick deploys of standard web apps
  • Lambda limits: 900 s timeout, 128–10,240 MB memory (CPU scales with it), /tmp 512–10,240 MB, 6 MB synchronous and 1 MB asynchronous payload, default 1,000 concurrent executions per Region (a soft quota).
  • Cold starts: provisioned concurrency, smaller packages, SnapStart where supported.
  • An Auto Scaling group keeps N healthy instances across AZs; target tracking (e.g. 50% CPU) is the usual policy.
  • ALB: layer 7, host/path routing, gRPC, WebSockets. NLB: layer 4 TCP/UDP, static IPs, extreme throughput. GWLB: inline appliances.

EC2 pricing

Model Commitment Discount vs On-Demand Best for
On-Demand none none spiky, short or unknown workloads
Savings Plans $/hour for 1 or 3 years up to 72% steady usage; Compute SP also covers Fargate and Lambda
Reserved Instances 1 or 3 years, fixed instance attributes up to 72% steady workloads; zonal RIs also reserve capacity
Spot none; reclaimed with a 2-minute warning up to 90% fault-tolerant batch, CI, stateless workers

Storage & S3

Option Type Scope Notes
S3 object store Region (≥3 AZs) unlimited objects, up to 50 TB each
EBS block volume for one instance one AZ persists after stop; snapshots to move it; gp3 is the default SSD
EFS managed NFS file system Region (multi-AZ) shared by many Linux instances and containers; elastic
Instance store local disk on the host the instance fastest, but data is lost on stop, hibernate or terminate
S3 class Min duration Access Use
Standard none milliseconds hot data
Intelligent-Tiering none milliseconds (optional archive tiers) unknown or changing patterns; no retrieval fee
Standard-IA / One Zone-IA 30 days milliseconds, retrieval fee backups; One Zone for re-creatable data
Glacier Instant Retrieval 90 days milliseconds archives read about quarterly
Glacier Flexible Retrieval 90 days minutes to hours (restore first) archives read about yearly
Glacier Deep Archive 180 days hours (restore first) compliance archives, cheapest
  • Durability is designed for 11 nines (99.999999999%) across classes; IA classes bill a 128 KB minimum object size.
  • Strong read-after-write consistency for all PUTs, DELETEs and LISTs.
  • Single PUT up to 5 GB; use multipart upload from about 100 MB (required above 5 GB), up to 50 TB.
  • Performance: at least 3,500 writes and 5,500 reads per second per prefix; spread keys across prefixes.
  • New buckets block public access, disable ACLs and encrypt with SSE-S3 by default; use SSE-KMS for key control and audit.
  • Also: versioning (once on, it can only be suspended), Object Lock (WORM), lifecycle rules, replication (needs versioning on both sides), presigned URLs, event notifications.

Databases & DynamoDB

Service Type Pick it for
RDS managed PostgreSQL, MySQL, MariaDB, Oracle, SQL Server, Db2 relational OLTP; Multi-AZ for HA, read replicas for reads
Aurora MySQL/PostgreSQL-compatible, storage kept as 6 copies across 3 AZs higher throughput, fast failover, up to 15 replicas, Global Database, Serverless v2
DynamoDB serverless key-value/document NoSQL single-digit-ms at any scale with known access patterns
ElastiCache in-memory Valkey, Redis OSS, Memcached caching, sessions, leaderboards, rate limits
Redshift columnar MPP data warehouse OLAP over large data; Spectrum queries S3
  • RDS Multi-AZ = synchronous standby plus automatic DNS failover (HA). Read replica = asynchronous copy for scaling reads (can be promoted, may lag).
  • DynamoDB primary key: partition key, or partition + sort key. The partition key is hashed, so pick a high-cardinality one to avoid hot partitions.
  • Query needs the partition key (plus optional sort-key conditions such as begins_with, between); Scan reads the whole table. Both return at most 1 MB per call, so paginate with LastEvaluatedKey.
GSI LSI
Keys any partition + sort key same partition key, different sort key
Created any time only at table creation
Consistency eventually consistent only strong or eventual
Capacity its own shares the table’s
Limit 20 per table (default quota) 5 per table; 10 GB per partition-key value
  • Item max 400 KB. 1 RCU = one strongly consistent read/s of up to 4 KB (or two eventually consistent); 1 WCU = one write/s of up to 1 KB.
  • Also: on-demand vs provisioned capacity, Streams, TTL, transactions, DAX cache, global tables (multi-Region, last writer wins).

Networking

  • VPC: regional, IPv4 CIDR from /16 to /28; a subnet lives in one AZ, and AWS reserves 5 addresses in each.
  • Public subnet = its route table sends 0.0.0.0/0 to the internet gateway (one per VPC, managed, highly available). Private subnets reach out through a NAT gateway.
  • A zonal NAT gateway sits in one AZ’s public subnet with an Elastic IP: run one per AZ. Newer regional NAT gateways expand across AZs automatically.
Security group Network ACL
Attached to ENI / instance subnet
State stateful: replies allowed automatically stateless: allow return (ephemeral) ports explicitly
Rules allow only allow and deny
Evaluation all rules numbered, lowest first, first match wins
Defaults new SG: no inbound, all outbound default NACL allows all; a new custom NACL denies all
Sources CIDRs or other security groups CIDRs only
  • VPC peering: 1:1, non-transitive, no overlapping CIDRs, cross-account and cross-Region. Transit Gateway: hub-and-spoke with transitive routing across many VPCs, VPNs and Direct Connect, segmented with route tables.
  • VPC endpoints: gateway endpoints (S3 and DynamoDB only, a route-table entry, no charge) vs interface endpoints (PrivateLink ENIs with private IPs, most services, hourly and per-GB charges).
  • Route 53 routing policies: simple, weighted, latency, failover, geolocation, geoproximity, IP-based, multivalue answer. Health checks drive failover; alias records point the zone apex at ELB, CloudFront or S3.
  • CloudFront: CDN at edge locations in front of S3 (locked down with Origin Access Control), ALB or custom origins; CloudFront Functions and Lambda@Edge run code at the edge.

IAM

  • Identities: the root user (MFA, then lock away), users (avoid long-lived keys), groups, roles (temporary credentials), and federation through IAM Identity Center, SAML or OIDC.
  • Policy types: identity-based, resource-based (bucket, KMS key, SQS, Lambda policies), permission boundaries, Organizations SCPs/RCPs and session policies. Boundaries, SCPs/RCPs and session policies only limit; they never grant.
  • Evaluation: everything starts as an implicit deny → any explicit Deny wins → otherwise it needs an Allow (identity or resource policy in the same account) that every applicable limit also permits. Cross-account access needs an allow on both sides.
  • Roles & STS: a trust policy says who may assume the role, a permissions policy what it can do. AssumeRole credentials last 1 hour by default (15 minutes up to the role’s 1–12 hour maximum); role chaining caps sessions at 1 hour.
  • Roles everywhere: EC2 instance profiles, Lambda execution roles, ECS task roles, EKS Pod Identity, cross-account roles with an ExternalId for third parties (confused deputy).
  • Least privilege: specific actions on specific ARNs, conditions (aws:PrincipalOrgID, aws:MultiFactorAuthPresent), tags for ABAC, IAM Access Analyzer to find unused or external access.
{
  "Version": "2012-10-17",
  "Statement": [
    { "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject"],
      "Resource": "arn:aws:s3:::app-uploads/tenant-a/*" },
    { "Effect": "Allow", "Action": "s3:ListBucket",
      "Resource": "arn:aws:s3:::app-uploads",
      "Condition": { "StringLike": { "s3:prefix": "tenant-a/*" } } },
    { "Effect": "Deny", "Action": "s3:*", "Resource": "arn:aws:s3:::app-uploads/*",
      "Condition": { "Bool": { "aws:SecureTransport": "false" } } }
  ]
}
JSON

Messaging & streaming

Service Model Ordering & delivery Retention Use for
SQS queue; consumers poll Standard: at-least-once, best-effort order. FIFO: ordered per message group, exactly-once processing 1 min–14 days (default 4) decoupling, buffering
SNS pub/sub push fan-out to SQS, Lambda, HTTP, email, SMS; FIFO topics keep order none (FIFO topics can archive) notifying many subscribers
EventBridge event bus with content-based rules AWS, SaaS and custom events optional archive and replay event-driven integration
Kinesis Data Streams sharded, replayable log ordered per shard (partition key); many consumers 24 h default, up to 365 days real-time streams and analytics
  • SQS: visibility timeout default 30 s (max 12 h), long polling up to 20 s, delay up to 15 min, batches of 10. Add a DLQ with maxReceiveCount.
  • Payload caps: SQS and EventBridge accept 1 MiB (older material says 256 KB); SNS topics default to 256 KiB. For bigger data put it in S3 and send a pointer.

Observability & IaC

  • CloudWatch: metrics (EC2 basic every 5 min, detailed every 1 min; memory and disk need the agent), logs and Logs Insights, alarms, dashboards.
  • CloudTrail: API audit trail (who did what, when, from where). 90 days of management events in Event history; create a trail to keep logs in S3; data events (S3 object, Lambda invoke) cost extra.
  • X-Ray: distributed tracing and service maps (instrument with OpenTelemetry). AWS Config: resource configuration history and compliance rules.
CloudFormation CDK Terraform
Language YAML/JSON templates TypeScript, Python, Java, C#, Go HCL
State managed by AWS (stacks) synthesizes CloudFormation state file you store (e.g. S3 with locking)
Scope AWS only AWS only multi-cloud and SaaS providers
Strengths change sets, drift detection, StackSets loops, types, reusable constructs plan/apply, huge provider ecosystem

Security services

  • KMS: HSM-backed keys governed by key policies; envelope encryption (GenerateDataKey: encrypt locally with a data key, store that key encrypted under the KMS key) because Encrypt takes only 4 KB; every use is logged in CloudTrail.
Secrets Manager Parameter Store
Built for secrets (DB passwords, API keys) config and secrets (SecureString via KMS)
Rotation built in (Lambda), native for RDS/Aurora none built in
Size up to 64 KB 4 KB standard, 8 KB advanced
Cost per secret per month plus API calls standard tier free, advanced charged
Extras cross-Region replication hierarchies (/app/prod/db), versions
  • WAF: layer-7 web ACLs (managed rules, rate limits, IP sets, SQLi/XSS) on CloudFront, ALB, API Gateway and more.
  • Shield Standard: free, automatic layer 3/4 DDoS protection. Shield Advanced: paid, adds a response team, advanced detection and DDoS cost protection.
  • GuardDuty: threat detection from CloudTrail, VPC Flow Logs and DNS logs.
  • Also: Inspector (vulnerability scans), Macie (sensitive data in S3), Security Hub (all findings in one place).

Well-Architected & DR

  • Six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, Sustainability.
  • RPO = how much data you can lose (time since the last good copy). RTO = how long recovery may take.
Strategy RPO / RTO DR Region runs Cost
Backup & restore hours backups only; redeploy with IaC $
Pilot light tens of minutes replicated data and core infra; app servers switched off $$
Warm standby minutes a scaled-down but fully working copy $$$
Multi-site active/active near zero full production, serving traffic $$$$
  • Replication doesn’t protect against corruption or deletion: keep point-in-time backups too (AWS Backup, versioning), and test restores.

3-tier HA architecture

  1. Edge: Route 53 alias → CloudFront with WAF and Shield; static assets from a private S3 bucket through OAC.
  2. VPC across 2–3 AZs, each with a public subnet (ALB nodes, NAT gateway), a private app subnet and a private data subnet.
  3. App tier: internet-facing ALB → Auto Scaling group (or ECS on Fargate) in every AZ’s app subnet, ELB health checks, stateless servers (sessions in ElastiCache or DynamoDB).
  4. Data tier: Aurora or RDS Multi-AZ in the data subnets, read replicas for read traffic, ElastiCache in front for hot reads.
  5. Chained security groups: ALB allows 443 from anywhere; app allows only the ALB’s SG; the database allows its port only from the app SG.
  6. Egress via NAT gateways, S3/DynamoDB via gateway endpoints, secrets in Secrets Manager, KMS at rest, IAM roles for compute.
  7. Operations: CloudWatch alarms, CloudTrail, AWS Backup, everything in IaC, a DR Region matched to RPO/RTO.

Interview tip

In design rounds, say the failure you are protecting against: an instance (Auto Scaling group), an AZ (multi-AZ subnets, Multi-AZ database) or a Region (DR strategy). Then map each layer to it.

Quick answers

  • Region vs AZ vs edge location? Geographic area; isolated data center(s) inside it; CDN/DNS point of presence close to users.
  • Security group vs NACL? Stateful, allow-only, per instance; stateless, allow and deny, per subnet.
  • How should EC2 access S3? An IAM role via an instance profile (temporary credentials), never stored access keys.
  • SQS vs SNS? SQS is a pull queue where each message is processed by one consumer; SNS pushes each message to every subscriber.
  • Multi-AZ vs read replica? Multi-AZ is for availability (synchronous standby, automatic failover); read replicas scale reads asynchronously.
  • ECS vs EKS? ECS is simpler and AWS-native; EKS gives you Kubernetes APIs and portability. Both run on EC2 or Fargate.
  • Lambda vs containers? Lambda for short, spiky, event-driven work; containers for long-running or steady services.
  • CloudTrail vs CloudWatch? CloudTrail records API calls (audit); CloudWatch collects metrics and logs and alarms on them (operations).
  • Secrets Manager vs Parameter Store? Rotation and cross-Region replication vs cheap hierarchical config (with SecureString for simple secrets).
  • What makes an app highly available on AWS? Multiple AZs at every tier, a load balancer with health checks, Auto Scaling, a Multi-AZ database and stateless servers.

Gotchas & traps

  • s3:ListBucket applies to the bucket ARN (arn:aws:s3:::b); object actions apply to arn:aws:s3:::b/*. Mixing them up is the most common policy bug.
  • In a versioned bucket a delete only adds a delete marker: old versions are still billed until lifecycle rules expire them.
  • Security groups can’t deny: block an IP with a NACL or WAF.
  • NACLs are stateless: forget the ephemeral return ports and responses silently drop.
  • A Lambda function attached to a VPC has no internet access unless its subnets route through a NAT gateway (or use VPC endpoints).
  • NAT gateway per-GB processing and cross-AZ data transfer add up; send S3/DynamoDB traffic through gateway endpoints.
  • SQS Standard can deliver duplicates and out of order, so make consumers idempotent; set the visibility timeout longer than processing time.
  • An ACM certificate for CloudFront must be requested in us-east-1, whatever Region the origin is in.
esc