Ch. 14

AWS interview questions & answers

The AWS services interviewers ask about: compute, storage, IAM, networking, databases, serverless, messaging and well-architected design.

59 interview questions21 quiz questions0 notes
your progress0%

Top 59 AWS interview questions most asked first

  1. 1.Explain the difference between an AWS Region, an Availability Zone and an edge location.easy

    A Region is a separate geographic area, like us-east-1 or eu-west-1. Regions are isolated from each other, most services are Regional, and your data stays in a Region unless you move or replicate it, which matters for latency, compliance and data residency.

    An Availability Zone is one or more discrete data centers inside a Region, with independent power, cooling and networking, linked to the other AZs by low-latency connections. Every Region has several AZs, and spreading resources across at least two of them is the basic high-availability pattern: if one AZ fails, the others keep serving.

    Edge locations are points of presence in many more cities than there are Regions. CloudFront, Route 53 and AWS Global Accelerator use them to cache content, answer DNS queries and terminate connections close to users.

    I'd also mention Local Zones and Wavelength Zones, which push compute closer to specific metro areas or 5G networks.

    What interviewers listen for
    • Region: isolated geographic area; most services are Regional
    • AZ: independent data centers within a Region
    • Deploy across multiple AZs for high availability
    • Edge locations serve CloudFront, Route 53, Global Accelerator
    • Region choice drives latency, compliance, cost, service availability

    Likely follow-up: How do you choose which Region to deploy a new workload to? · Is us-east-1a the same physical AZ in every AWS account?

  2. 2.What is the AWS shared responsibility model? Give examples of what AWS is responsible for and what you are responsible for.easy

    AWS is responsible for security of the cloud; the customer is responsible for security in the cloud.

    • AWS: the physical data centers, hardware, global network and virtualization layer, plus the software that runs its managed services.
    • Customer: your data, IAM users, roles and permissions, encryption choices, network rules such as security groups and NACLs, and how you configure every service you use.

    The line moves with the type of service. On EC2 (infrastructure as a service) you patch the guest OS, the runtime and your application. On RDS, AWS patches the database engine and OS, but you still own database users, network access, encryption and backup settings. On Lambda or S3 AWS manages even more, yet the code, the data and the access policies are still yours.

    A classic example: a publicly readable S3 bucket, or a security group that opens SSH to 0.0.0.0/0, is the customer's problem, not AWS's.

    What interviewers listen for
    • AWS: security of the cloud (facilities, hardware, hypervisor)
    • Customer: security in the cloud (data, IAM, configuration)
    • Boundary shifts from IaaS to managed to serverless
    • Patching the EC2 guest OS is the customer's job
    • Misconfigurations like public buckets are on the customer

    Likely follow-up: Who patches the operating system on an RDS instance compared with an EC2 instance?

  3. 3.Compare the EC2 pricing models: On-Demand, Reserved Instances, Savings Plans and Spot. When would you use each?easy
    • On-Demand: pay for compute by the second or hour with no commitment. Best for spiky, short-lived or unpredictable workloads, and for anything you haven't sized yet.
    • Savings Plans: commit to a consistent spend per hour for 1 or 3 years in exchange for a discount. Compute Savings Plans apply across instance families and Regions, and also to Fargate and Lambda; EC2 Instance Savings Plans give a deeper discount but are tied to one instance family in one Region.
    • Reserved Instances: the older commitment model, 1 or 3 years for specific instance attributes. Standard RIs have the biggest discount; Convertible RIs can be exchanged. Zonal RIs also reserve capacity.
    • Spot: spare capacity at a steep discount, but AWS can reclaim it with a two-minute warning. Use it for fault-tolerant, stateless or flexible work such as batch, CI, big data and containers.

    A common strategy: a Savings Plan for the steady baseline, On-Demand for the variable part, and Spot for interruptible capacity.

    What interviewers listen for
    • On-Demand: no commitment, highest unit price
    • Savings Plans: 1- or 3-year spend commitment, flexible
    • Reserved Instances: commitment to specific instance attributes
    • Spot: big discount, two-minute interruption notice
    • Commit for the baseline; On-Demand and Spot for the rest

    Likely follow-up: What kinds of workloads should never run on Spot? · How do you get guaranteed capacity without a 1- or 3-year term?

  4. 4.What are IAM users, groups, roles and policies, and how do they relate to each other?easy
    • Users are identities for a person or workload with long-term credentials: a console password and/or access keys.
    • Groups are collections of users. Attaching policies to a group is how you manage permissions for many users at once. Groups can't be nested and can't be used as a principal in a policy.
    • Roles are identities without long-term credentials. A trusted principal (an EC2 instance, a Lambda function, a user in another account, a federated user) assumes the role and gets temporary credentials from STS.
    • Policies are JSON documents made of statements with Effect, Action, Resource and an optional Condition. Identity-based policies attach to users, groups or roles; resource-based policies, like bucket policies, attach to a resource and name a Principal.

    Current best practice: people sign in through IAM Identity Center or federation, workloads use roles, IAM users with access keys are the exception, and the root user is locked away with MFA.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Action": ["s3:GetObject", "s3:PutObject"],
          "Resource": "arn:aws:s3:::app-uploads/*"
        }
      ]
    }
    What interviewers listen for
    • Users: long-term credentials for a person or app
    • Groups: attach permissions to many users at once
    • Roles: assumed, temporary credentials via STS
    • Policies: JSON with Effect, Action, Resource, Condition
    • Prefer roles and federation over IAM user access keys

    Likely follow-up: What is the difference between an AWS managed policy, a customer managed policy and an inline policy?

  5. 5.What is the difference between a security group and a network ACL?easy
    • Scope: a security group is attached to network interfaces, so effectively to instances, load balancers or Lambda functions in a VPC. A network ACL applies to a whole subnet.
    • State: security groups are stateful: if a request is allowed in, the response is allowed out automatically. NACLs are stateless: you must allow the return traffic explicitly, usually the ephemeral port range.
    • Rules: security groups have only allow rules, all evaluated together. NACLs have numbered allow and deny rules, evaluated from the lowest number, and the first match wins.
    • Defaults: a new security group allows all outbound and no inbound traffic. The default NACL allows everything; a new custom NACL denies everything until you add rules.

    Security groups can use another security group as the source, which is how you say "only the load balancer may reach the app tier". I use security groups as the main control and NACLs as a coarse subnet guardrail, for example to block a known bad CIDR range.

    # App tier accepts port 8080 only from the ALB's security group
    aws ec2 authorize-security-group-ingress \
      --group-id sg-0app1234example \
      --protocol tcp --port 8080 \
      --source-group sg-0alb5678example
    What interviewers listen for
    • Security group: interface level; NACL: subnet level
    • Security groups stateful; NACLs stateless (return ports)
    • SG allow-only; NACL allow and deny, ordered rules
    • Security groups can reference other security groups
    • Default SG blocks inbound; default NACL allows all

    Likely follow-up: A NACL allows inbound 443 and outbound 443 only. Why do HTTPS clients time out?

  6. 6.Walk me through the main S3 storage classes and when you would use each.easy
    • S3 Standard: frequently accessed data with millisecond access. The default.
    • S3 Intelligent-Tiering: unknown or changing access patterns. S3 moves objects between access tiers automatically for a small per-object monitoring fee, with no retrieval charges.
    • S3 Standard-IA and S3 One Zone-IA: infrequently accessed data that still needs millisecond access. Cheaper storage, but per-GB retrieval fees and a minimum storage duration. One Zone-IA keeps data in a single AZ, so use it only for data you can re-create.
    • S3 Glacier Instant Retrieval: archives read rarely, say once a quarter, but needed in milliseconds.
    • S3 Glacier Flexible Retrieval: archives restored in minutes to hours.
    • S3 Glacier Deep Archive: the cheapest class, for long-term retention such as compliance data, restored within hours.
    • S3 Express One Zone: single-AZ, single-digit-millisecond latency for performance-critical workloads, using directory buckets.

    Apart from the One Zone classes, data is stored across multiple AZs. Lifecycle rules move objects down the tiers as they age.

    What interviewers listen for
    • Standard for hot data; Intelligent-Tiering for unknown patterns
    • IA classes: cheaper storage, retrieval fees, minimum duration
    • One Zone classes lose data if that AZ is lost
    • Glacier classes trade retrieval time for price
    • Lifecycle rules automate transitions between classes

    Likely follow-up: When is Intelligent-Tiering not a good fit?

  7. 7.What makes a subnet public or private in a VPC, and how do instances in a private subnet reach the internet?easy

    A subnet is public when its route table sends 0.0.0.0/0 (and ::/0 for IPv6) to an internet gateway. Instances there also need a public IPv4 or Elastic IP address to be reachable. A private subnet has no route to the internet gateway.

    Private instances that need outbound access, for patches or third-party APIs, get a route to a NAT gateway. The NAT gateway translates their private addresses to its own public IP and only allows connections initiated from inside, so nothing on the internet can open a connection to them. For IPv6 the equivalent is an egress-only internet gateway.

    A classic (zonal) NAT gateway sits in a public subnet of one AZ, so for high availability you deploy one per AZ and point each private route table at the NAT gateway in its own AZ. AWS has also added a regional NAT gateway mode that spans AZs automatically. If the traffic is only to services like S3 or DynamoDB, VPC endpoints avoid NAT charges entirely.

    PublicRoute:
      Type: AWS::EC2::Route
      Properties:
        RouteTableId: !Ref PublicRouteTable
        DestinationCidrBlock: 0.0.0.0/0
        GatewayId: !Ref InternetGateway
    PrivateRouteA:
      Type: AWS::EC2::Route
      Properties:
        RouteTableId: !Ref PrivateRouteTableA
        DestinationCidrBlock: 0.0.0.0/0
        NatGatewayId: !Ref NatGatewayA
    What interviewers listen for
    • Public means a route to an internet gateway
    • Private subnets use a NAT gateway for outbound only
    • One zonal NAT gateway per AZ for high availability
    • Egress-only internet gateway for IPv6
    • VPC endpoints avoid NAT for AWS service traffic

    Likely follow-up: Why is a single NAT gateway a risk for a multi-AZ application? · What is the difference between a NAT gateway and a NAT instance?

  8. 8.When would you choose an Application Load Balancer, a Network Load Balancer or a Gateway Load Balancer?mid
    • ALB works at layer 7 (HTTP, HTTPS, gRPC, WebSockets). It routes by host, path, headers, query string or source IP to target groups of instances, IP addresses, containers or Lambda functions. It terminates TLS, integrates with AWS WAF and can authenticate users through Cognito or any OIDC provider. It's the default for web apps and microservices.
    • NLB works at layer 4 (TCP, UDP, TLS). It handles very high throughput at low latency, gives you a static IP per AZ (you can bring Elastic IPs) and can preserve the client's source IP. Use it for non-HTTP protocols, extreme performance, fixed IPs that partners allowlist, or in front of a PrivateLink endpoint service.
    • GWLB works at layer 3 and inserts fleets of third-party virtual appliances, like firewalls or intrusion detection, transparently into the traffic path. It talks to the appliances with GENEVE encapsulation and is reached through Gateway Load Balancer endpoints.

    The Classic Load Balancer is legacy; I wouldn't use it for new designs.

    What interviewers listen for
    • ALB: layer 7, content-based routing, WAF, TLS termination
    • NLB: layer 4, static IPs, low latency, source IP preserved
    • NLB fronts PrivateLink endpoint services
    • GWLB: layer 3, transparent third-party appliances via GENEVE
    • Classic Load Balancer is legacy

    Likely follow-up: A partner needs to allowlist fixed IP addresses for your HTTP API behind an ALB. What do you do?

  9. 9.What is the difference between RDS Multi-AZ and read replicas?mid

    They solve different problems: Multi-AZ is for availability, read replicas are for read scaling.

    • Multi-AZ DB instance deployment: RDS keeps a synchronous standby in another AZ that serves no traffic. If the primary or its AZ fails, RDS fails over automatically by pointing the same DNS endpoint at the standby, typically within a minute or two, without losing committed transactions.
    • Multi-AZ DB cluster deployment (RDS for MySQL and PostgreSQL): a writer plus two readable standbys in three AZs, so you get failover and some read capacity.
    • Read replicas use asynchronous replication, so they can lag behind. Each has its own endpoint, the application must send reads to it, and it can live in another Region. Promoting a replica to a standalone database is a manual step, which makes a cross-Region replica a handy DR building block.

    Many production setups use both: Multi-AZ for the primary, plus read replicas for reporting or read-heavy traffic.

    What interviewers listen for
    • Multi-AZ: synchronous standby, automatic failover
    • Multi-AZ instance standby serves no reads
    • Read replicas: asynchronous, scale reads, can lag
    • Replicas can be cross-Region; promotion is manual
    • Multi-AZ DB cluster has two readable standbys

    Likely follow-up: How does the application find the new primary after a failover? · How does this differ in Aurora?

  10. 10.Design a highly available three-tier web application on AWS.hard

    I'd spread every tier across at least two, ideally three, Availability Zones in one Region.

    • Edge: Route 53 for DNS, CloudFront for static assets and caching, AWS WAF on CloudFront or the load balancer.
    • Web and app tier: an internet-facing ALB in public subnets forwarding to an Auto Scaling group (or ECS services) in private subnets. ELB health checks replace bad instances, target tracking scales on CPU or request count, and instances are stateless: sessions in ElastiCache or DynamoDB, uploads in S3.
    • Data tier: RDS Multi-AZ or Aurora in private database subnets, with read replicas or ElastiCache for read-heavy traffic, and automated backups with point-in-time recovery.
    • Network: security groups chained ALB to app to database, a NAT gateway per AZ, VPC endpoints for S3.
    • Operations: everything in infrastructure as code, CloudWatch alarms, CloudTrail, secrets in Secrets Manager.

    Then I'd say what it survives (an instance or AZ failure) and what it doesn't (a Regional outage), and propose a DR strategy if the RTO requires one.

    What interviewers listen for
    • Every tier spread across multiple AZs
    • ALB plus Auto Scaling group of stateless instances
    • RDS Multi-AZ or Aurora in private subnets
    • Security groups chained tier to tier
    • State the failure modes; DR for Regional outages

    Likely follow-up: How would you deploy new versions of this app with zero downtime? · What would you change to survive the loss of a whole Region?

  11. 11.How does an EC2 Auto Scaling group work, and what is a launch template?easy

    A launch template is a versioned definition of how to launch an instance: AMI, instance type, key pair, security groups, IAM instance profile, user data, volumes and tags. It replaces the older launch configurations, which AWS has deprecated.

    An Auto Scaling group uses a launch template to keep a fleet between a minimum and maximum size, aiming at a desired capacity. It spreads instances across the subnets (AZs) you give it, replaces unhealthy instances and registers new ones with load balancer target groups.

    Scaling options:

    • Target tracking: keep a metric such as average CPU at 50%. Usually the best default.
    • Step and simple scaling driven by CloudWatch alarms.
    • Scheduled scaling for known peaks and predictive scaling from historical patterns.

    Use the ELB health check type so instances that are running but failing the app's health check get replaced. Also worth knowing: lifecycle hooks, instance refresh for rolling AMI updates, warm pools, and mixed instances policies that combine On-Demand and Spot.

    aws autoscaling create-auto-scaling-group \
      --auto-scaling-group-name web-asg \
      --launch-template 'LaunchTemplateName=web-lt,Version=$Latest' \
      --min-size 2 --max-size 6 --desired-capacity 2 \
      --vpc-zone-identifier "subnet-0aaa1111example,subnet-0bbb2222example" \
      --target-group-arns arn:aws:elasticloadbalancing:us-east-1:111122223333:targetgroup/web/0123456789abcdef \
      --health-check-type ELB --health-check-grace-period 300
    What interviewers listen for
    • Launch template: versioned instance configuration
    • ASG keeps min, max and desired capacity across AZs
    • Target tracking as the default scaling policy
    • ELB health checks replace instances failing the app check
    • Instance refresh, lifecycle hooks, mixed On-Demand and Spot

    Likely follow-up: How do you let in-flight requests finish before an instance is terminated during scale-in?

  12. 12.What is the difference between EBS volumes and instance store volumes?easy

    EBS is network-attached block storage. A volume lives in one AZ, persists independently of the instance, can be detached and reattached, and can be backed up with snapshots, which are incremental, stored by AWS in S3 and copyable across Regions. Volume types include general purpose SSD (gp3, gp2), provisioned IOPS SSD (io2, io1) and HDD types for throughput or cold data (st1, sc1). With gp3 you set IOPS and throughput independently of the volume size.

    Instance store is disk physically attached to the host. It's very fast and included in the instance price, but it's ephemeral: data survives a reboot but is lost when the instance stops, hibernates or terminates, or if the underlying hardware fails. Only some instance types have it.

    So EBS for root volumes, self-managed databases and anything that must persist; instance store for caches, scratch space, buffers, or data that is already replicated elsewhere, such as nodes of a distributed database.

    What interviewers listen for
    • EBS: network-attached, persistent, scoped to one AZ
    • EBS snapshots are incremental and copyable across Regions
    • Instance store: physically attached, very fast, ephemeral
    • Instance store data lost on stop, terminate, hardware failure
    • gp3 decouples IOPS and throughput from size

    Likely follow-up: How would you move an EBS volume to a different Availability Zone?

  13. 13.Compare S3, EBS and EFS. Which would you use for what?easy
    • S3 is object storage: objects in buckets, accessed over an HTTPS API. It scales virtually without limit, stores data across multiple AZs by default, and suits static assets, backups, data lakes, logs and media. It isn't a file system you mount for low-latency random writes.
    • EBS is block storage attached to a single EC2 instance (except Multi-Attach on some provisioned IOPS volumes). It's scoped to one AZ, low latency, and what you use for boot volumes and databases you run yourself.
    • EFS is a managed NFS file system: many instances, containers and Lambda functions across AZs can mount it at the same time, it grows and shrinks automatically, and you pay for the storage you use. Good for shared content, home directories and lift-and-shift apps that expect a POSIX file system.

    For Windows file shares, high-performance computing or NetApp compatibility there's the FSx family, such as FSx for Windows File Server, FSx for Lustre and FSx for NetApp ONTAP.

    What interviewers listen for
    • S3: object storage over an API, virtually unlimited
    • EBS: block storage for one instance, scoped to an AZ
    • EFS: shared NFS file system mounted across AZs
    • EFS grows automatically; pay for what you store
    • FSx for Windows shares, Lustre and ONTAP
  14. 14.Compare SQS, SNS, EventBridge and Kinesis. How do you decide which one to use?mid
    • SQS is a queue: producers send messages, consumers poll and delete them. It decouples services and buffers load, and each message is handled by one consumer. It comes as standard or FIFO.
    • SNS is push-based pub/sub: a message published to a topic fans out to many subscribers, such as SQS queues, Lambda functions, HTTP endpoints, email or SMS. SNS to several SQS queues is the classic fan-out, giving each consumer its own durable copy.
    • EventBridge is an event bus with content-based routing: rules match fields in the event JSON and send matches to many kinds of targets. It receives events from AWS services and SaaS partners and supports archive and replay. Best for event-driven integration.
    • Kinesis Data Streams is a stream: records are ordered per shard by partition key and retained for a set period, so several consumers read independently and can replay. Use it for high-throughput real-time data like clickstreams, logs and IoT.

    Rule of thumb: work queue, SQS; fan-out, SNS; routing events between services, EventBridge; ordered, replayable streams, Kinesis.

    What interviewers listen for
    • SQS: pull-based queue, one consumer per message
    • SNS: push fan-out to many subscribers
    • EventBridge: rules route events by content
    • Kinesis: ordered shards, retention, replay, many consumers
    • SNS-to-SQS fan-out pattern

    Likely follow-up: Where does Amazon MSK fit into this comparison? · How would you guarantee ordering end to end?

  15. 15.How do you choose between RDS, Aurora and DynamoDB?mid
    • RDS runs managed relational engines: MySQL, PostgreSQL, MariaDB, Oracle, SQL Server and Db2. Pick it for a standard or commercial engine, SQL joins and transactions, when the load fits one primary plus replicas.
    • Aurora is AWS's MySQL- and PostgreSQL-compatible engine with a distributed storage layer: data is copied across three AZs, storage grows automatically, replicas share the same volume so lag is low, and failover is fast. It adds Global Database for cross-Region reads and DR, and Aurora Serverless v2 for variable load. Choose it for demanding relational workloads, usually at a higher price than RDS.
    • DynamoDB is a serverless key-value and document database with consistent single-digit-millisecond latency at virtually any scale. You design tables around known access patterns (partition and sort keys, GSIs); there are no joins and ad hoc queries are limited. Great for sessions, carts, user profiles, IoT and other high-scale, key-based access.

    So: relational data with flexible queries, RDS or Aurora; massive scale with well-known key-based access, DynamoDB.

    What interviewers listen for
    • RDS: managed standard engines, including commercial ones
    • Aurora: shared storage across three AZs, fast failover
    • Aurora Global Database and Serverless v2
    • DynamoDB: serverless NoSQL designed around access patterns
    • Decide by data model, query flexibility and scale

    Likely follow-up: What would make you migrate from RDS for PostgreSQL to Aurora PostgreSQL?

  16. 16.What is a Lambda cold start, and how do you reduce its impact?mid

    When no idle execution environment is available, Lambda must create one: load your code, start the runtime and run the initialization code outside your handler. That Init phase adds latency to that request, which is the cold start. The environment is then reused for later invocations until Lambda recycles it.

    Ways to reduce it:

    • Keep the package small and load only what you need; heavy frameworks and big dependency trees hurt most.
    • Do expensive setup once, outside the handler: SDK clients, connections, configuration.
    • Provisioned concurrency keeps a set number of environments initialized. It costs extra and suits latency-sensitive APIs.
    • SnapStart (recent Java, Python and .NET runtimes) resumes new environments from a snapshot of an initialized one. It can't be combined with provisioned concurrency.
    • More memory also means proportionally more CPU, which speeds up init.

    Cold starts usually hit a small share of requests, so I'd measure first, using the Init Duration in the logs or traces.

    import os
    import boto3
    
    # Init phase: runs once per execution environment, reused while warm
    table = boto3.resource("dynamodb").Table(os.environ["TABLE_NAME"])
    
    def handler(event, context):
        # Runs on every invocation
        item = table.get_item(Key={"pk": event["id"]}).get("Item")
        return {"statusCode": 200 if item else 404}
    What interviewers listen for
    • Init phase: create environment, start runtime, run init code
    • Create clients outside the handler for reuse
    • Provisioned concurrency keeps environments warm, at extra cost
    • SnapStart restores from a snapshot of initialized state
    • Smaller packages and more memory speed up init

    Likely follow-up: Why can SnapStart cause problems for code that generates unique IDs or random seeds during init?

  17. 17.What are the six pillars of the AWS Well-Architected Framework?easy
    • Operational Excellence: run and monitor systems and keep improving: infrastructure as code, small reversible changes, observability, learning from failures.
    • Security: protect data and systems: a strong identity foundation with least privilege, traceability, security at every layer, encryption in transit and at rest.
    • Reliability: recover from failures and meet demand: automatic recovery, horizontal scaling, tested recovery procedures, managed change.
    • Performance Efficiency: use resources efficiently as demand and technology change: pick the right services, use serverless, go global quickly, experiment.
    • Cost Optimization: avoid unnecessary spend: pay for consumption, right-size, measure efficiency, attribute costs to owners.
    • Sustainability: reduce environmental impact: maximize utilization, prefer managed services, shrink the resources your workload needs.

    The pillars trade off against each other; multi-Region reliability, for instance, costs more. The AWS Well-Architected Tool lets you review a workload against each pillar's questions, and lenses extend it for areas like serverless.

    What interviewers listen for
    • Operational Excellence, Security, Reliability
    • Performance Efficiency, Cost Optimization, Sustainability
    • Each pillar has design principles and best practices
    • Pillars involve trade-offs, e.g. reliability versus cost
    • Well-Architected Tool and lenses for workload reviews

    Likely follow-up: Which pillar would you prioritize for an early-stage startup, and why?

  18. 18.How do IAM roles and STS AssumeRole work?mid

    A role has two kinds of policy:

    • A trust policy, a resource-based policy on the role that says who may assume it: an AWS service such as lambda.amazonaws.com, an account, a specific role, or a federated identity provider.
    • Permissions policies that say what the role can do once assumed.

    A trusted principal calls sts:AssumeRole (for cross-account access its own identity policy must also allow that action) and gets temporary credentials: an access key ID, a secret access key and a session token. They last for the session duration, one hour by default and configurable up to the role's maximum, then expire on their own, so there's nothing long-lived to leak or rotate.

    Services use this behind the scenes: EC2 instance profiles, Lambda execution roles and ECS task roles all give the workload temporary credentials that the SDK refreshes automatically. Roles are also the standard mechanism for cross-account access, and AssumeRoleWithWebIdentity and AssumeRoleWithSAML cover federation.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Principal": { "Service": "lambda.amazonaws.com" },
          "Action": "sts:AssumeRole"
        }
      ]
    }
    What interviewers listen for
    • Trust policy: who can assume the role
    • Permissions policies: what the role can do
    • STS returns temporary credentials with a session token
    • Instance profiles, Lambda and ECS task roles use roles
    • Standard mechanism for cross-account access and federation

    Likely follow-up: What is role chaining, and how does it limit the session duration?

  19. 19.Explain the disaster recovery strategies on AWS and how they relate to RPO and RTO.hard

    RPO (recovery point objective) is how much data you can afford to lose, measured in time; RTO (recovery time objective) is how long you can be down. AWS's DR whitepaper describes four strategies, from cheapest to most expensive:

    • Backup and restore: back up data, AMIs and IaC to another Region; after a disaster, redeploy and restore. RPO and RTO in hours.
    • Pilot light: data is replicated continuously to live data stores in the DR Region and core infrastructure exists, but app servers are switched off until needed. Tens of minutes.
    • Warm standby: a scaled-down but fully working copy runs all the time and can take traffic immediately; you scale it up. Minutes.
    • Multi-site active/active: full capacity in several Regions, all serving traffic. Near-zero RTO, and near-zero RPO with asynchronous replication.

    Tools include AWS Backup, S3 replication, Aurora Global Database, DynamoDB global tables and Route 53 health checks. Two strong points: fail over with data plane operations, and remember replication also copies corruption, so keep point-in-time backups and test the plan.

    What interviewers listen for
    • RPO: tolerable data loss; RTO: tolerable downtime
    • Backup and restore: hours, cheapest
    • Pilot light: data live, compute off; tens of minutes
    • Warm standby: scaled-down running copy; minutes
    • Active/active: near zero, costliest; still keep backups

    Likely follow-up: How would you switch traffic to the DR Region in a pilot light setup? · Why does the whitepaper recommend relying on data plane rather than control plane operations during failover?

  20. 20.How do you control access to an S3 bucket? Compare IAM policies, bucket policies, ACLs and Block Public Access.mid
    • IAM identity policies grant your own users and roles access to buckets and objects. Best for managing what a principal can do across many resources.
    • Bucket policies are resource-based policies on the bucket. Use them for cross-account access, service principals like CloudFront, and bucket-wide guardrails such as denying non-TLS requests or requiring a specific VPC endpoint.
    • ACLs are the legacy, pre-IAM mechanism. New buckets default to the Object Ownership setting "bucket owner enforced", which disables ACLs, and AWS recommends keeping them off.
    • Block Public Access overrides policies and ACLs to stop public access. It's on by default for new buckets and can also be enforced for the whole account.

    Buckets are private by default. Within one account an allow in either the IAM policy or the bucket policy is enough, and an explicit deny anywhere wins. Access points help with shared datasets, and IAM Access Analyzer flags buckets shared publicly or with other accounts.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Sid": "DenyInsecureTransport",
          "Effect": "Deny",
          "Principal": "*",
          "Action": "s3:*",
          "Resource": ["arn:aws:s3:::app-data", "arn:aws:s3:::app-data/*"],
          "Condition": { "Bool": { "aws:SecureTransport": "false" } }
        }
      ]
    }
    What interviewers listen for
    • IAM policies for your own principals
    • Bucket policies for cross-account access and guardrails
    • ACLs are legacy; disabled by default via Object Ownership
    • Block Public Access overrides everything, on by default
    • Private by default; explicit deny always wins

    Likely follow-up: How would you let another account upload objects while your account keeps ownership of them?

  21. 21.How do S3 versioning and lifecycle rules work, and how would you use them together?easy

    Versioning keeps every version of an object. An overwrite creates a new version, and a delete without a version ID only adds a delete marker, so the previous version can be restored. It protects against accidental deletes and overwrites, and replication requires it. Once enabled it can be suspended but never fully turned off. MFA Delete and Object Lock add stronger protection.

    Lifecycle rules act on objects by age, scoped by prefix or tags:

    • Transition objects to cheaper classes, for example Standard to Standard-IA after 30 days, then to Glacier.
    • Expire current versions after a retention period.
    • Expire noncurrent versions so versioning doesn't grow your bill forever.
    • Abort incomplete multipart uploads and clean up expired delete markers.

    Together: version the bucket for protection, and let lifecycle rules age data into cheaper tiers and trim old versions.

    {
      "Rules": [
        {
          "ID": "logs-tiering",
          "Filter": { "Prefix": "logs/" }, "Status": "Enabled",
          "Transitions": [
            { "Days": 30, "StorageClass": "STANDARD_IA" },
            { "Days": 90, "StorageClass": "GLACIER" }
          ],
          "NoncurrentVersionExpiration": { "NoncurrentDays": 30 },
          "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
        }
      ]
    }
    What interviewers listen for
    • Versioning keeps all versions; deletes add a delete marker
    • Versioning can be suspended, not disabled
    • Lifecycle transitions move data to cheaper classes
    • Expire noncurrent versions to control cost
    • Abort incomplete multipart uploads

    Likely follow-up: A user deleted an object in a versioned bucket. How do you get it back?

  22. 22.What is the difference between CloudWatch, CloudTrail and X-Ray?easy
    • CloudWatch is monitoring: metrics from AWS services plus your own custom metrics, logs with retention settings, Logs Insights queries and metric filters, alarms that notify through SNS or trigger actions like Auto Scaling, and dashboards. It answers "how is my system behaving?".
    • CloudTrail is the audit log of API calls: who did what, when, from where, on which resource. Management events are recorded by default and visible in the event history for 90 days; a trail delivers events to S3 for long-term retention, and data events such as S3 object reads are opt-in. It answers "who changed this?".
    • X-Ray is distributed tracing: follow one request through API Gateway, Lambda, containers and downstream calls, with a service map showing where latency and errors come from. AWS now recommends OpenTelemetry-based instrumentation over the older X-Ray SDKs.

    A classic gotcha: EC2 doesn't report memory or disk usage out of the box; you need the CloudWatch agent.

    aws cloudwatch put-metric-alarm \
      --alarm-name web-high-cpu \
      --namespace AWS/EC2 --metric-name CPUUtilization \
      --dimensions Name=AutoScalingGroupName,Value=web-asg \
      --statistic Average --period 300 --evaluation-periods 2 \
      --threshold 80 --comparison-operator GreaterThanThreshold \
      --alarm-actions arn:aws:sns:us-east-1:111122223333:ops-alerts
    What interviewers listen for
    • CloudWatch: metrics, logs, alarms, dashboards
    • CloudTrail: audit trail of API calls
    • X-Ray: distributed tracing and service map
    • Memory and disk metrics need the CloudWatch agent
    • Trail to S3 for retention beyond 90 days

    Likely follow-up: Someone deleted a production security group rule. How do you find out who did it?

  23. 23.Explain DynamoDB partition keys and sort keys. How do you choose them?mid

    Every item is identified by its primary key:

    • Simple: a partition key only. DynamoDB hashes it to decide which physical partition stores the item, and GetItem needs the exact value.
    • Composite: partition key plus sort key. Items sharing a partition key form an item collection stored together and ordered by the sort key, so Query can fetch a range with begins_with, between or comparisons, in order.

    Choosing keys starts from access patterns, not entities. The partition key should have high cardinality and spread traffic evenly, like userId or orderId, not status or today's date. The sort key models the one-to-many relationship or ordering you need, for example ORDER#2024-05-01#123 so a customer's orders come back sorted by date.

    In single-table design, generic key names like PK and SK with prefixed values let one table hold several entity types. Patterns the table's key can't serve go to a GSI; a Scan should be the exception.

    What interviewers listen for
    • Partition key is hashed to choose the partition
    • Sort key orders items within a partition key
    • Query needs partition key equality; sort key allows ranges
    • High-cardinality, evenly accessed partition keys
    • Design from access patterns; GSIs for the rest

    Likely follow-up: How would you model customers and their orders so you can list a customer's recent orders?

  24. 24.What is the difference between SQS standard and FIFO queues?mid
    • Standard queues offer nearly unlimited throughput, at-least-once delivery (a message is occasionally delivered more than once) and best-effort ordering. Consumers must be idempotent.
    • FIFO queues give ordering within a message group and exactly-once processing. The queue name ends in .fifo, every message carries a MessageGroupId, and a message sent again within the 5-minute deduplication interval is dropped, based on an explicit MessageDeduplicationId or content-based deduplication (a SHA-256 hash of the body). Throughput is lower than standard, though batching and high-throughput mode raise it a lot.

    Ordering is per message group, not global, so the group ID is your parallelism knob: different groups, say different customers, are processed in parallel, while messages within one group are strictly sequential. While a message in a group is in flight, the next message in that group isn't delivered.

    Use FIFO when order or duplicates really matter, such as commands against one account; otherwise use standard with idempotent consumers.

    What interviewers listen for
    • Standard: at-least-once, best-effort order, huge throughput
    • FIFO: ordered per message group, exactly-once processing
    • Deduplication ID or content-based, 5-minute window
    • Message group ID sets ordering scope and parallelism
    • Standard queue consumers must be idempotent

    Likely follow-up: Does exactly-once processing in FIFO mean your consumer can skip idempotency?

  25. 25.What is the SQS visibility timeout, and how do dead-letter queues work?mid

    When a consumer receives a message, SQS doesn't remove it; it hides it for the visibility timeout (30 seconds by default, up to 12 hours). The consumer must call DeleteMessage before it expires. If the consumer crashes or is too slow, the message reappears and another consumer gets it, which is how SQS avoids losing messages and also why duplicates happen.

    So set the timeout above your processing time, extend it with ChangeMessageVisibility as a heartbeat for long jobs, and when Lambda consumes the queue set the visibility timeout to at least six times the function timeout, as AWS recommends.

    A dead-letter queue catches poison messages. The source queue's redrive policy sets maxReceiveCount: once a message has been received that many times without being deleted, SQS moves it to the DLQ. Alarm on the DLQ's depth, inspect the messages, fix the bug, then redrive them back to the source queue. A FIFO queue needs a FIFO DLQ.

    import boto3
    
    sqs = boto3.client("sqs")
    QUEUE_URL = "https://sqs.us-east-1.amazonaws.com/111122223333/orders"
    
    while True:
        # Long polling: wait up to 20 s for messages instead of returning empty
        resp = sqs.receive_message(QueueUrl=QUEUE_URL, MaxNumberOfMessages=10, WaitTimeSeconds=20)
        for msg in resp.get("Messages", []):
            process(msg["Body"])  # must finish within the visibility timeout
            sqs.delete_message(QueueUrl=QUEUE_URL, ReceiptHandle=msg["ReceiptHandle"])
    What interviewers listen for
    • Received messages are hidden, not deleted
    • Delete before the timeout or the message reappears
    • Extend with ChangeMessageVisibility for long jobs
    • DLQ via redrive policy and maxReceiveCount
    • Alarm on DLQ depth, fix, then redrive

    Likely follow-up: Your Lambda consumer times out after 60 seconds and messages are processed twice. What would you check?

  26. 26.How should an application running on EC2 get AWS credentials to call S3?mid

    Never put access keys in the AMI, the code or config files. Create an IAM role that trusts ec2.amazonaws.com, give it least-privilege permissions (say s3:GetObject on one bucket prefix), and attach it to the instance through an instance profile, the container that passes a role to EC2.

    The instance metadata service (IMDS) then serves temporary credentials for that role. The AWS SDKs and CLI find them automatically through the default credential provider chain and refresh them before they expire, so your code just creates an S3 client without any keys.

    Worth adding: require IMDSv2, which needs a session token obtained with a PUT request and helps protect against SSRF attacks that try to steal credentials from the metadata endpoint. The same idea applies elsewhere: Lambda execution roles, ECS task roles, and EKS Pod Identity or IAM roles for service accounts on Kubernetes.

    # IMDSv2: get a session token, then list the attached role's credentials
    TOKEN=$(curl -s -X PUT "http://169.254.169.254/latest/api/token" \
      -H "X-aws-ec2-metadata-token-ttl-seconds: 21600")
    curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
      http://169.254.169.254/latest/meta-data/iam/security-credentials/
    What interviewers listen for
    • IAM role attached through an instance profile
    • IMDS serves temporary, automatically rotated credentials
    • SDK default credential chain picks them up
    • Enforce IMDSv2 to mitigate SSRF credential theft
    • Never store long-term access keys on instances

    Likely follow-up: How does a pod running on EKS get its own AWS permissions?

  27. 27.How does AWS evaluate policies to decide whether a request is allowed?hard

    Everything starts as an implicit deny: nothing is allowed unless a policy allows it. Then:

    • Explicit deny wins. AWS collects every applicable policy: Organizations SCPs and RCPs, resource-based policies, identity-based policies, permissions boundaries and session policies. If any statement denies the request, the answer is deny.
    • Guardrails must allow. SCPs (on principals) and RCPs (on resources) never grant anything; they set the maximum. If they don't allow the action, it's denied.
    • Something must grant. Within one account, an allow in either the identity-based or the resource-based policy is generally enough. Role trust policies and KMS key policies are exceptions that must allow the principal explicitly.
    • Boundaries and session policies intersect. If present, they must allow the action too.

    For cross-account requests both sides must allow: the caller's identity policy and the resource's policy (or the trust policy of a role to assume).

    In the snippet, every object action on reports/* is allowed except DeleteObject, because the explicit deny wins.

    {
      "Version": "2012-10-17",
      "Statement": [
        { "Effect": "Allow", "Action": "s3:*", "Resource": "arn:aws:s3:::reports/*" },
        { "Effect": "Deny", "Action": "s3:DeleteObject", "Resource": "arn:aws:s3:::reports/*" }
      ]
    }
    What interviewers listen for
    • Default is implicit deny
    • An explicit deny in any policy always wins
    • SCPs, RCPs and boundaries limit, never grant
    • Same account: identity or resource policy can grant
    • Cross-account: both sides must allow

    Likely follow-up: Do SCPs restrict users and roles in the organization's management account?

  28. 28.What options do you have for encrypting data in S3?mid

    At rest, S3 encrypts every new object by default with SSE-S3, using keys S3 manages. The other options:

    • SSE-KMS: keys in AWS KMS, either the AWS managed key or your own customer managed key. You get key policies controlling who can decrypt, CloudTrail records of key use, and rotation. Enable S3 Bucket Keys to cut KMS request volume and cost.
    • DSSE-KMS: dual-layer encryption with KMS keys, for specific compliance requirements.
    • SSE-C: you send your own key with every request and S3 doesn't store it. It's now disabled by default on new buckets and rarely the right choice.
    • Client-side encryption: encrypt before uploading, so S3 never sees plaintext.

    In transit, use HTTPS and enforce it with a bucket policy that denies requests where aws:SecureTransport is false.

    SSE-KMS with a customer managed key is the usual pick for separation of duties: someone with s3:GetObject but no kms:Decrypt on the key still can't read the data.

    aws s3api put-bucket-encryption --bucket app-data \
      --server-side-encryption-configuration '{
        "Rules": [{
          "ApplyServerSideEncryptionByDefault": {
            "SSEAlgorithm": "aws:kms",
            "KMSMasterKeyID": "alias/app-data"
          },
          "BucketKeyEnabled": true
        }]
      }'
    What interviewers listen for
    • SSE-S3 applied to new objects by default
    • SSE-KMS adds key policies, audit trail, separation of duties
    • S3 Bucket Keys reduce KMS request costs
    • SSE-C and client-side encryption for customer-held keys
    • Enforce TLS with aws:SecureTransport in the bucket policy

    Likely follow-up: A role has s3:GetObject on a bucket but gets Access Denied on objects. What would you check?

  29. 29.What is an S3 presigned URL, and when would you use one?mid

    A presigned URL grants time-limited access to one object operation, such as a GET to download or a PUT to upload, to someone who has no AWS credentials. Your backend signs the URL with its own credentials, and S3 checks the signature, the expiry and the signer's permissions when the request arrives.

    Typical uses:

    • Direct uploads from browsers or mobile apps: your API authorizes the user and returns a presigned PUT URL, so the file goes straight to S3 instead of streaming through your servers or API Gateway.
    • Private downloads: invoices, reports or media only an authenticated user should get.

    Things to know: the URL can only do what the signer is allowed to do. With Signature Version 4 the maximum lifetime is 7 days, and a URL signed with temporary credentials, such as a role's, stops working when those credentials expire. Anyone holding the URL can use it until then, so keep expiries short. For content behind CloudFront, use CloudFront signed URLs or signed cookies instead.

    import boto3
    
    s3 = boto3.client("s3")
    url = s3.generate_presigned_url(
        "get_object",
        Params={"Bucket": "invoices-prod", "Key": "2024/inv-123.pdf"},
        ExpiresIn=900,  # seconds
    )
    What interviewers listen for
    • Time-limited access to one object operation
    • Uses the signer's credentials and permissions
    • Direct uploads from clients bypass your servers
    • Up to 7 days with SigV4; ends when temporary credentials expire
    • Bearer token: keep expiries short

    Likely follow-up: How would you restrict the maximum size of a file users upload directly to S3?

  30. 30.What is CloudFront, and how does it improve performance and security?easy

    CloudFront is AWS's content delivery network. Users are routed to a nearby edge location; if the object is cached there it's served immediately, otherwise CloudFront fetches it from the origin (an S3 bucket, an ALB, API Gateway or any HTTP server) over the AWS network, often through regional edge caches, and caches it.

    Performance: lower latency, less load on the origin, TLS terminated close to the user, compression, HTTP/2 and HTTP/3. Cache behaviors map path patterns to origins, each with a cache policy that sets TTLs and which headers, cookies and query strings form the cache key. Stale content is cleared with invalidations or, better, versioned file names.

    Security: AWS Shield Standard against common DDoS attacks, AWS WAF, HTTPS with ACM certificates, geographic restrictions, signed URLs or cookies for private content, and origin access control so an S3 origin only accepts requests from your distribution.

    For logic at the edge, CloudFront Functions handle lightweight header and URL rewrites, and Lambda@Edge handles heavier work.

    What interviewers listen for
    • CDN caching at edge locations close to users
    • Origins: S3, ALB, API Gateway, any HTTP server
    • Cache behaviors and cache policies define the cache key
    • Shield Standard, WAF, HTTPS, signed URLs and cookies
    • Origin access control locks S3 to the distribution

    Likely follow-up: How do you release a new version of a single-page app without users getting stale files?

  31. 31.What routing policies does Route 53 support, and when would you use each?mid
    • Simple: a single resource, such as one web server or load balancer.
    • Weighted: split traffic in proportions you choose, for example 90/10 for a canary or a gradual migration.
    • Latency-based: answer with the Region that gives the user the lowest latency; the usual choice for multi-Region active/active.
    • Failover: active-passive. Route 53 returns the primary while its health check passes, otherwise the secondary. The basis of DNS-level disaster recovery.
    • Geolocation: route by the user's continent, country or US state, for compliance, localized content or licensing. Always add a default record.
    • Geoproximity: route by the distance between users and your resources, with a bias to shift traffic toward or away from a location.
    • IP-based: route by the client's source IP ranges, for example a known ISP's CIDR blocks.
    • Multivalue answer: return up to eight healthy records at random.

    Most policies combine with health checks, and alias records point to AWS resources such as ALBs or CloudFront, even at the zone apex. Remember TTLs: failover only takes effect as clients' DNS caches expire.

    What interviewers listen for
    • Weighted for canaries and gradual cutovers
    • Latency-based for multi-Region active/active
    • Failover plus health checks for active-passive DR
    • Geolocation, geoproximity and IP-based differ in routing input
    • Alias records for AWS resources, including the zone apex

    Likely follow-up: Why might some users keep reaching the failed Region for a while after a DNS failover?

  32. 32.How do you connect multiple VPCs? Compare VPC peering and Transit Gateway.mid

    VPC peering is a direct private connection between two VPCs, in the same or different accounts and Regions. There's no extra hop or bandwidth bottleneck, and you pay only for data transfer. But it's not transitive: if A peers with B and B peers with C, A can't reach C through B. CIDR ranges can't overlap, and with many VPCs you end up managing a full mesh of connections and route entries.

    Transit Gateway is a Regional hub-and-spoke router. Each VPC, VPN connection or Direct Connect gateway attaches once, and transit gateway route tables decide who can reach whom, so you can segment networks: prod can't reach dev, but everyone reaches shared services. Transit gateways in different Regions can peer with each other. The trade-off is cost: hourly charges per attachment plus per-GB data processing.

    So peering for a handful of VPCs or heavy point-to-point traffic, Transit Gateway once you have many VPCs or hybrid connectivity. To expose just one service rather than a whole network, use PrivateLink.

    What interviewers listen for
    • Peering: direct, one-to-one, not transitive
    • No overlapping CIDRs; a full mesh grows quickly
    • Transit Gateway: Regional hub-and-spoke with route tables
    • Transit Gateway adds attachment and data processing costs
    • PrivateLink exposes a single service instead

    Likely follow-up: Two VPCs you need to connect have overlapping CIDR ranges. What are your options?

  33. 33.What are ECS, EKS and Fargate, and how would you choose between them?mid
    • ECS is AWS's own container orchestrator. You write task definitions (image, CPU, memory, ports, IAM task role) and run them as services that keep a desired count behind a load balancer. It's simpler than Kubernetes, tightly integrated with IAM, ALB and CloudWatch, and has no control plane fee.
    • EKS is managed Kubernetes: AWS runs the control plane and you get the standard Kubernetes API and ecosystem, like Helm and operators, plus portability to other clouds and on-premises. The cost is more complexity and a per-cluster fee.
    • Fargate isn't an orchestrator but a serverless compute engine for containers that works with both ECS and EKS. No EC2 instances to patch or scale; you pay for the vCPU and memory each task or pod requests.

    My default is ECS on Fargate for a team that wants containers on AWS with minimal operations; EKS when the team already runs Kubernetes or needs its ecosystem; EC2 capacity for either when you need GPUs, specific instance types, or the lowest cost at steady scale, often with Spot.

    What interviewers listen for
    • ECS: AWS-native orchestrator, task definitions and services
    • EKS: managed Kubernetes control plane, ecosystem, portability
    • Fargate: serverless compute for both ECS and EKS
    • EC2 capacity for GPUs, special instance types, cost
    • Choose by team skills and operational appetite

    Likely follow-up: What trade-offs come with running containers on Fargate instead of EC2?

  34. 34.What is the difference between a global secondary index and a local secondary index in DynamoDB?mid
    • Key: a GSI can use a completely different partition key and optional sort key. An LSI keeps the table's partition key and only changes the sort key.
    • Lifecycle: GSIs can be added or deleted at any time. LSIs must be defined when the table is created and can't be added or removed later.
    • Consistency: GSIs support only eventually consistent reads; LSIs also support strongly consistent reads.
    • Capacity: a GSI has its own throughput and is updated asynchronously; an LSI uses the table's capacity.
    • Size: with an LSI, all items for one partition key value (table plus indexes) must stay within 10 GB; GSIs have no such limit.
    • Attributes: a GSI query returns only projected attributes, while an LSI can fetch the rest from the table.

    The default quotas are 20 GSIs and 5 LSIs per table. In practice I almost always use GSIs, and pick an LSI only for a strongly consistent alternate sort order known up front. Watch out: in provisioned mode, an under-provisioned GSI can throttle writes to the base table.

    OrdersTable:
      Type: AWS::DynamoDB::Table
      Properties:
        BillingMode: PAY_PER_REQUEST
        AttributeDefinitions:
          - { AttributeName: PK, AttributeType: S }
          - { AttributeName: SK, AttributeType: S }
          - { AttributeName: email, AttributeType: S }
        KeySchema: [{ AttributeName: PK, KeyType: HASH }, { AttributeName: SK, KeyType: RANGE }]
        GlobalSecondaryIndexes:
          - IndexName: by-email
            KeySchema: [{ AttributeName: email, KeyType: HASH }]
            Projection: { ProjectionType: KEYS_ONLY }
    What interviewers listen for
    • GSI: any partition key; LSI: same partition key
    • GSIs added anytime; LSIs only at table creation
    • GSIs eventually consistent only; LSIs allow strong reads
    • GSI has its own capacity; LSI shares the table
    • LSI imposes a 10 GB limit per partition key value
  35. 35.What does least privilege mean in AWS, and how do you apply it in practice?easy

    Least privilege means each identity gets only the permissions it needs for its task, on only the resources it needs, for only as long as it needs them. It limits the blast radius of a bug, a leaked credential or a compromised workload.

    In practice:

    • Scope actions and resources: dynamodb:GetItem on one table ARN instead of dynamodb:* on *. Add conditions such as source VPC endpoint, tags or MFA.
    • Prefer roles with temporary credentials over IAM users with long-term access keys, one role per workload.
    • Tighten over time: AWS managed policies help you start, then IAM Access Analyzer can generate a policy from CloudTrail activity and flag unused or externally shared access, and last-accessed data shows permissions nobody uses.
    • Guardrails: SCPs in AWS Organizations and permissions boundaries cap what can ever be granted, for example when developers create their own roles.
    • People: federate through IAM Identity Center, require MFA, and don't use the root user day to day.

    Review regularly, because permissions tend to accumulate.

    What interviewers listen for
    • Only needed actions, on needed resources, for needed time
    • Specific ARNs and conditions instead of wildcards
    • Roles and temporary credentials over access keys
    • Access Analyzer and last-accessed data to trim permissions
    • SCPs and permissions boundaries as guardrails
  36. 36.When would you use Secrets Manager instead of Systems Manager Parameter Store?easy

    Both store values encrypted with KMS, are controlled with IAM, and can be read at runtime by Lambda, ECS, CloudFormation and your own code.

    Secrets Manager is built for secrets:

    • Automatic rotation on a schedule, with managed rotation for database credentials such as RDS and Lambda-based rotation for anything else.
    • Replication of secrets to other Regions, and resource-based policies for cross-account access.
    • Priced per secret per month plus API calls.

    Parameter Store is a general configuration store:

    • Hierarchical names like /app/prod/db-url, String, StringList and encrypted SecureString types, with version history.
    • The standard tier has no additional charge; the advanced tier allows larger values, more parameters and parameter policies such as expiration, for a fee.
    • No built-in rotation.

    My rule: database passwords, API keys and anything that should rotate go in Secrets Manager; endpoints, feature flags and other configuration go in Parameter Store. Either way, cache values in the application instead of calling the API on every request.

    What interviewers listen for
    • Both KMS-encrypted and controlled by IAM
    • Secrets Manager: rotation, replication, per-secret pricing
    • Parameter Store: hierarchical config, SecureString, free standard tier
    • Anything that must rotate goes in Secrets Manager
    • Cache values instead of fetching on every request

    Likely follow-up: How would you rotate a database password without breaking running applications?

  37. 37.Compare CloudFormation, the AWS CDK and Terraform for infrastructure as code.easy
    • CloudFormation is AWS's native IaC service. You write declarative JSON or YAML templates and deploy them as stacks; AWS keeps the state. You get change sets to preview updates, automatic rollback on failure, drift detection, and StackSets for many accounts and Regions. Downsides: verbose templates, and AWS only.
    • AWS CDK defines infrastructure in TypeScript, Python, Java, C# or Go. Constructs bundle sensible defaults or whole patterns, and you get loops, types and reuse. cdk synth produces a CloudFormation template, so deployment, state and rollback are still CloudFormation's.
    • Terraform uses HCL and providers for AWS and many other platforms, so one tool covers several clouds and SaaS products. terraform plan shows the diff before apply. You manage the state file yourself, typically in S3 with locking, and protect it because it can contain secrets. OpenTofu is an open-source fork.

    I'd pick CDK or CloudFormation for AWS-only teams that want managed state, and Terraform for multi-cloud setups or where it's already the organization's standard.

    AWSTemplateFormatVersion: '2010-09-09'
    Parameters:
      Env:
        Type: String
        AllowedValues: [dev, prod]
    Resources:
      AssetsBucket:
        Type: AWS::S3::Bucket
        Properties:
          BucketName: !Sub 'acme-assets-${Env}'
          VersioningConfiguration:
            Status: Enabled
    Outputs:
      BucketArn: { Value: !GetAtt AssetsBucket.Arn }
    What interviewers listen for
    • CloudFormation: native, declarative, AWS-managed state
    • Change sets, rollback, drift detection, StackSets
    • CDK: programming languages that synthesize CloudFormation
    • Terraform: HCL, multi-cloud providers, self-managed state
    • Choose by cloud scope and existing team standards

    Likely follow-up: Someone changed a resource by hand in the console. How do you detect and fix the drift?

  38. 38.What does API Gateway do, and how do REST APIs, HTTP APIs and WebSocket APIs differ?mid

    API Gateway is a managed front door for APIs: routing, authorization, throttling, CORS, custom domains, stages and logging, with integrations to Lambda, HTTP backends, other AWS services, and private resources in a VPC through VPC links.

    • REST APIs have the most features: API keys and usage plans for per-client throttling, request validation and transformation, caching, AWS WAF, resource policies, private endpoints and canary deployments.
    • HTTP APIs have fewer features but are cheaper and simpler, with native JWT authorizers for Cognito or any OIDC provider. A good default for plain Lambda or HTTP proxies.
    • WebSocket APIs keep persistent two-way connections for chat, notifications or live dashboards. Routes such as $connect and $disconnect invoke your backend, which pushes messages back through a connection callback URL.

    Authorization options include IAM, Lambda authorizers and Cognito. Two limits to know: REST API integrations time out after 29 seconds by default (it can be raised for Regional and private APIs), and payloads are capped at 10 MB, so large files go straight to S3 with presigned URLs.

    What interviewers listen for
    • Managed front door: routing, auth, throttling, stages
    • REST APIs: API keys, usage plans, caching, WAF, validation
    • HTTP APIs: cheaper and simpler, native JWT authorizers
    • WebSocket APIs for persistent two-way connections
    • Timeout and payload limits push big files to S3

    Likely follow-up: How would you protect a public API from one client consuming all the capacity?

  39. 39.Explain Lambda concurrency: the account quota, reserved concurrency and provisioned concurrency.hard

    Concurrency is the number of requests a function handles at the same moment, roughly requests per second times average duration. Each concurrent request needs its own execution environment.

    • Account quota: by default 1,000 concurrent executions per Region, shared by all functions and raisable. Beyond it, synchronous calls are throttled with a 429 error, while asynchronous and event source invocations are retried. Lambda also limits how fast each function can scale out.
    • Reserved concurrency sets aside a fixed amount for one function. It's both a guarantee (others can't starve it) and a cap (it can't go higher), which is handy to protect a downstream database. Setting it to 0 stops all invocations. No extra charge.
    • Provisioned concurrency keeps environments pre-initialized on a version or alias, removing cold starts for that capacity; beyond it the function scales on demand. It costs extra and can be scheduled or auto scaled.

    For SQS event sources, the event source mapping's maximum concurrency setting is a gentler way to cap consumers than reserved concurrency.

    # Guarantee and cap this function at 100 concurrent executions
    aws lambda put-function-concurrency --function-name orders-api \
      --reserved-concurrent-executions 100
    
    # Keep 20 environments initialized on the "live" alias
    aws lambda put-provisioned-concurrency-config --function-name orders-api \
      --qualifier live --provisioned-concurrent-executions 20
    What interviewers listen for
    • Concurrency is roughly requests per second times duration
    • Regional account quota shared by all functions
    • Reserved concurrency: guarantee plus cap, no charge
    • Provisioned concurrency: pre-initialized environments, extra cost
    • Synchronous invocations are throttled with 429 errors

    Likely follow-up: A Lambda function overwhelms an RDS database during traffic spikes. What would you change?

  40. 41.How does AWS KMS work, and what is envelope encryption?hard

    KMS manages encryption keys in hardware security modules, and a KMS key's material never leaves KMS unencrypted. Every key has a key policy; IAM policies and grants add to it, and every use is logged in CloudTrail. Keys are AWS owned (used inside services), AWS managed (created per service in your account, like aws/s3), or customer managed, where you control the policy, rotation and deletion.

    Encrypt only accepts small payloads (up to 4 KB), and sending all your data to KMS would be slow anyway. So services use envelope encryption:

    • GenerateDataKey returns a plaintext data key plus the same key encrypted under your KMS key.
    • Encrypt the data locally with the plaintext key, then discard it from memory.
    • Store the encrypted data key next to the ciphertext.
    • To decrypt, send the encrypted data key to KMS Decrypt and decrypt the data locally.

    S3, EBS and RDS encryption work this way. You get speed, no size limit, and access to the data hinges on permission to decrypt one small key.

    import boto3
    
    kms = boto3.client("kms")
    resp = kms.generate_data_key(KeyId="alias/app-data", KeySpec="AES_256")
    plaintext_key = resp["Plaintext"]       # use locally, never store it
    encrypted_key = resp["CiphertextBlob"]  # store next to the ciphertext
    
    # ... encrypt the payload locally with plaintext_key (for example AES-GCM) ...
    
    # Later: recover the data key in order to decrypt
    plaintext_key = kms.decrypt(CiphertextBlob=encrypted_key)["Plaintext"]
    What interviewers listen for
    • KMS key material never leaves KMS unencrypted
    • Key policies, IAM and grants control use; CloudTrail logs it
    • GenerateDataKey returns plaintext and encrypted data keys
    • Encrypt locally; store only the encrypted data key
    • AWS owned, AWS managed and customer managed keys

    Likely follow-up: What happens to data encrypted under a KMS key that is scheduled for deletion? · Does rotating a KMS key re-encrypt existing data?

  41. 42.The AWS bill has grown much faster than traffic. How would you approach cost optimization?mid

    Visibility first, then the biggest levers.

    • Visibility: Cost Explorer grouped by service, account and cost allocation tags; AWS Budgets with alerts; Cost Anomaly Detection. Each team should own its spend.
    • Right-size: use Compute Optimizer and CloudWatch utilization to shrink over-provisioned instances, databases and containers, and consider Graviton instances for better price-performance.
    • Stop paying for idle: schedule non-production environments off at night; delete unattached EBS volumes, old snapshots, idle load balancers and unused public IPs.
    • Pricing models: Savings Plans or Reserved Instances for the steady baseline, Spot for fault-tolerant work.
    • Storage: S3 lifecycle rules and Intelligent-Tiering, gp2 to gp3 migrations, sensible CloudWatch Logs retention.
    • Data transfer: NAT gateway processing and cross-AZ traffic often hide surprises. Use VPC endpoints for S3 and DynamoDB and CloudFront to reduce egress.
    • Architecture: serverless or managed services for spiky workloads, caching to shrink databases.

    Trusted Advisor gives quick wins. Then make it continuous with tagging policies, per-team budgets and regular reviews.

    What interviewers listen for
    • Visibility first: Cost Explorer, tags, Budgets, anomaly detection
    • Right-size with Compute Optimizer; consider Graviton
    • Remove idle resources; schedule non-production
    • Savings Plans for the baseline, Spot for flexible work
    • Watch data transfer, NAT gateways and storage tiers

    Likely follow-up: Which hidden data transfer costs have you seen surprise teams on AWS?

  42. 43.Design a serverless backend where users upload images that are resized and then served through an API.hard
    • Upload: the client calls an API (API Gateway HTTP API with a Cognito JWT authorizer) to get a presigned S3 PUT URL, then uploads straight to an uploads bucket. No file bytes pass through API Gateway or Lambda.
    • Trigger: an S3 event notification goes to an SQS queue, and a Lambda function consumes it in batches. The queue absorbs bursts and retries failures, and a DLQ catches poison messages.
    • Processing: Lambda, with enough memory for the image library, writes the resized variants to a separate processed bucket (never the source bucket, which would trigger itself in a loop) and records metadata and status in DynamoDB.
    • Serving: CloudFront with origin access control in front of the processed bucket; the API reads metadata from DynamoDB.
    • Operations: idempotent processing keyed by object key and version, alarms on DLQ depth and errors, a least-privilege role per function, and everything defined in SAM or CDK.

    If the pipeline grows into several steps, like moderation and notifications, I'd orchestrate them with Step Functions. The design scales to zero when idle.

    What interviewers listen for
    • Presigned URLs for direct uploads to S3
    • S3 event to SQS to Lambda: buffered, retried processing
    • Write outputs to another bucket to avoid event loops
    • DynamoDB for metadata; CloudFront with OAC for delivery
    • Idempotency, DLQ alarms, least-privilege roles

    Likely follow-up: How would you tell the client that processing has finished?

  43. 44.Compare DynamoDB on-demand and provisioned capacity modes.mid
    • On-demand: no capacity planning. DynamoDB scales automatically and you pay per read and write request. AWS now describes it as the default and recommended mode for most workloads. Ideal for new tables, unpredictable or spiky traffic, and tables that are often idle. You can set a maximum throughput per table to cap spend.
    • Provisioned: you set read and write capacity units per second and pay for them hourly whether or not you use them. Add auto scaling (target tracking on utilization) and, for steady long-term workloads, reserved capacity for a lower price. Best when traffic is predictable and you want tight cost control.

    Capacity units: one RCU is one strongly consistent read per second of an item up to 4 KB, or two eventually consistent reads; one WCU is one write per second of up to 1 KB. Transactional requests consume twice as much.

    Either mode can still throttle if traffic concentrates on one partition key, because each partition has its own throughput ceiling. Switching modes is allowed only a limited number of times.

    What interviewers listen for
    • On-demand: pay per request, no capacity planning
    • On-demand is AWS's recommended default for most tables
    • Provisioned: RCUs and WCUs, auto scaling, reserved capacity
    • RCU covers a 4 KB strong read; WCU a 1 KB write
    • Hot partitions can throttle in either mode
  44. 45.What is a hot partition in DynamoDB, and how do you fix one?hard

    DynamoDB spreads items across partitions by hashing the partition key, and each partition has a ceiling of 3,000 read units and 1,000 write units per second. A hot partition appears when a disproportionate share of traffic hits one key or a few: a celebrity's profile, today's date as the key, a global counter, a status attribute. You get throttled while the table has spare capacity.

    DynamoDB helps: burst capacity absorbs short spikes, and adaptive capacity shifts throughput to busy partitions and can split partitions that stay hot. But a single hot item can never exceed the per-partition limit.

    Fixes:

    • Choose a higher-cardinality key, like userId instead of country.
    • Write sharding: add a suffix, such as 2024-05-01#7 out of 10 shards, to spread writes, and query all shards in parallel when reading.
    • Cache hot reads with DAX or ElastiCache.
    • Aggregate counters instead of updating one item per event.
    • Check GSIs: a low-cardinality GSI key can be hot even when the table isn't.

    CloudWatch Contributor Insights shows the most accessed keys.

    What interviewers listen for
    • Per-partition ceiling: 3,000 RCU and 1,000 WCU
    • Caused by low-cardinality or skewed keys
    • Burst and adaptive capacity help, within limits
    • Write sharding with suffixes; parallel reads across shards
    • Cache hot reads; Contributor Insights finds hot keys

    Likely follow-up: How would you implement a like counter for a post that goes viral?

  45. 46.When would you use ElastiCache, and which caching strategies would you apply?mid

    ElastiCache is a managed in-memory data store for sub-millisecond reads. Use it to offload repeated database queries, store sessions, build leaderboards with sorted sets, rate limit, or do pub/sub. It supports Valkey, Redis OSS and Memcached, either as a serverless cache or as node-based clusters.

    • Valkey or Redis OSS: rich data structures, replicas with automatic failover across AZs, cluster mode for sharding, persistence options. The usual choice.
    • Memcached: a simple multi-threaded key-value cache without replication or persistence, fine for purely disposable data.

    Strategies:

    • Cache-aside (lazy loading): read the cache; on a miss, read the database and populate the cache. Only requested data is cached, but misses are slow and data can go stale.
    • Write-through: update the cache on every database write. Fresher data, but you also cache things nobody reads.
    • TTLs on keys to bound staleness, with some jitter so many keys don't expire together.

    Also plan invalidation on writes, protection against a stampede when a hot key expires, and the eviction policy for when memory fills up.

    What interviewers listen for
    • In-memory store: sub-millisecond reads, offloads databases
    • Valkey or Redis OSS: structures, replicas, failover
    • Memcached: simple multi-threaded cache, no replication
    • Cache-aside versus write-through trade-offs
    • TTLs with jitter; guard against cache stampedes
  46. 47.What is AWS Step Functions, and when would you choose Standard or Express workflows?mid

    Step Functions orchestrates workflows as state machines written in Amazon States Language. States include Task (call Lambda or an AWS service API directly), Choice, Parallel, Map, Wait, Succeed and Fail, and each state can declare Retry and Catch, so error handling lives in the workflow instead of your code. The .sync pattern waits for a job such as an ECS task to finish, and .waitForTaskToken pauses until a person or external system calls back.

    • Standard workflows run up to one year with exactly-once execution and a full, inspectable execution history, priced per state transition. Use them for long-running, auditable or non-idempotent processes like order fulfilment or payments.
    • Express workflows run up to five minutes and are priced by executions, duration and memory, for high-volume event processing. Asynchronous Express is at-least-once and synchronous Express at-most-once, so steps should be idempotent, and they don't support .sync or callbacks.

    I reach for Step Functions when a Lambda starts calling other Lambdas and managing its own retries.

    {
      "StartAt": "ChargeCard",
      "States": {
        "ChargeCard": {
          "Type": "Task",
          "Resource": "arn:aws:states:::lambda:invoke",
          "Parameters": { "FunctionName": "charge-card", "Payload.$": "$" },
          "Retry": [{ "ErrorEquals": ["States.TaskFailed"], "MaxAttempts": 3, "BackoffRate": 2 }],
          "Catch": [{ "ErrorEquals": ["States.ALL"], "Next": "PaymentFailed" }],
          "End": true
        },
        "PaymentFailed": { "Type": "Fail", "Error": "PaymentFailed" }
      }
    }
    What interviewers listen for
    • State machines in Amazon States Language
    • Retry and Catch per state for error handling
    • .sync and .waitForTaskToken integration patterns
    • Standard: up to a year, exactly-once, full history
    • Express: up to five minutes, high volume, idempotent steps
  47. 48.What are the 7 Rs of migrating workloads to AWS?mid

    AWS describes seven migration strategies, chosen per application after a portfolio assessment:

    • Retire: decommission applications nobody needs.
    • Retain: keep them where they are for now, because of compliance, dependencies, a recent upgrade or low value.
    • Rehost ("lift and shift"): move servers as they are, often automated with AWS Application Migration Service.
    • Relocate: move at the platform level without changing the application, for example a VMware estate to a VMware-based platform on AWS, or resources to another VPC, Region or account.
    • Repurchase ("drop and shop"): replace the application with a SaaS product or a newer version.
    • Replatform ("lift, tinker and shift"): targeted optimizations such as a self-managed database moving to RDS or an app moving into containers, without changing the core architecture.
    • Refactor or re-architect: redesign with cloud-native services such as serverless or microservices; the most effort and the most benefit.

    For large migrations AWS generally recommends rehost, replatform and relocate to move quickly, then modernizing once in the cloud. AWS DMS handles database migrations and DataSync moves file data.

    What interviewers listen for
    • Retire, Retain, Rehost, Relocate
    • Repurchase, Replatform, Refactor or re-architect
    • Chosen per application after a portfolio assessment
    • Move fast with rehost and replatform, modernize later
    • Refactor: most effort, most cloud-native benefit

    Likely follow-up: Which strategy would you choose for a legacy Oracle database with a tight data center exit deadline?

  48. 49.How do different event sources invoke Lambda, and how does error handling differ between them?hard

    There are three invocation models:

    • Synchronous: the caller waits for the result, as with API Gateway, an ALB or an SDK Invoke. Lambda doesn't retry; the caller sees the error and decides.
    • Asynchronous: Lambda queues the event and acknowledges immediately, as with S3 notifications, SNS and EventBridge. On failure Lambda retries twice by default, and you can send failed events to an on-failure destination or a DLQ.
    • Event source mappings: Lambda polls SQS, Kinesis, DynamoDB Streams, Kafka or Amazon MQ and invokes your function with batches.

    With polling sources, errors affect the whole batch. For SQS, failed messages reappear after the visibility timeout and eventually land in the queue's DLQ. For Kinesis and DynamoDB Streams, a failing batch blocks its shard until it succeeds or the records expire, so configure maximum retries, maximum record age, bisect batch on error and an on-failure destination.

    For SQS and streams, enable partial batch responses so only failed items are retried, and make handlers idempotent.

    def handler(event, context):
        failures = []
        for record in event["Records"]:
            try:
                process(record["body"])
            except Exception:
                failures.append({"itemIdentifier": record["messageId"]})
        # Requires ReportBatchItemFailures on the SQS event source mapping
        return {"batchItemFailures": failures}
    What interviewers listen for
    • Synchronous: caller handles errors, no Lambda retries
    • Asynchronous: two retries by default, destinations or DLQ
    • Event source mappings poll SQS, Kinesis, DynamoDB Streams
    • A failing stream batch blocks its shard
    • Partial batch responses and idempotent handlers
  49. 50.What is an AMI, and what is EC2 user data used for?easy

    An AMI (Amazon Machine Image) is the template an instance boots from: snapshots of the root volume and any other volumes, plus metadata like the architecture and block device mapping. AMIs are Regional; you can copy them to other Regions and share them with other accounts. You can start from AWS images like Amazon Linux, Marketplace images, or your own golden AMIs built with EC2 Image Builder or Packer.

    User data is a script or cloud-init configuration passed at launch. By default it runs once, as root, on the instance's first boot, and is typically used to install packages, fetch configuration or register the instance somewhere. It's limited to 16 KB and readable through the instance metadata service, so never put secrets in it.

    The trade-off: baking software into the AMI gives fast, consistent boots, which matters for Auto Scaling; heavy user data keeps images generic but slows every launch and can fail at boot. A common middle ground is a golden AMI with the OS and dependencies plus a short user data script for environment-specific settings.

    #!/bin/bash
    # Runs as root on first boot (Amazon Linux 2023)
    dnf install -y nginx
    systemctl enable --now nginx
    echo "Deployed at $(date)" > /usr/share/nginx/html/index.html
    What interviewers listen for
    • AMI: volume snapshots plus launch metadata
    • AMIs are Regional; copy or share them
    • User data runs once as root on first boot by default
    • Never put secrets in user data
    • Golden AMI plus light user data for fast scaling
  50. 51.How are EC2 instance types named, and which instance families would you pick for different workloads?easy

    A name like m7g.xlarge reads as family (m), generation (7), optional attributes (g for AWS Graviton, a for AMD, i for Intel, d for local NVMe instance storage, n for enhanced networking), then size (large, xlarge, 2xlarge and so on, with resources roughly doubling at each step).

    Main families:

    • General purpose (m, t): balanced CPU and memory for web and application servers. t instances are burstable: they earn CPU credits while idle and spend them in bursts, which suits low, spiky loads, but watch for credit exhaustion.
    • Compute optimized (c): more CPU per GB of memory, for batch jobs, encoding and busy web servers.
    • Memory optimized (r, x): in-memory databases, caches and large analytics.
    • Storage optimized (i, d): high local disk IOPS or throughput, for NoSQL databases and data warehousing.
    • Accelerated computing (p, g, inf, trn): GPUs or AWS Inferentia and Trainium chips for ML and graphics.

    Newer generations usually give better price-performance, and Graviton (Arm) instances often more so. I'd benchmark and check Compute Optimizer rather than guess.

    What interviewers listen for
    • Name: family, generation, attributes, size
    • m and t general purpose; t is burstable with credits
    • c compute, r and x memory, i and d storage
    • p, g, inf and trn for accelerated computing
    • Newer generations and Graviton improve price-performance
  51. 52.What consistency model does S3 provide today?mid

    S3 provides strong read-after-write consistency for object operations in all Regions, and has since late 2020. After a successful PUT of a new object, an overwrite or a DELETE, any subsequent GET, HEAD or LIST reflects the change. A job can write an object and immediately list the prefix and see it, with no extra cost or configuration.

    Nuances worth knowing:

    • Atomic per key only: a reader gets the old or the new object, never a partial one, but there are no transactions across keys.
    • Concurrent writers: if two clients write the same key at the same time, the latest timestamp wins. S3 doesn't lock objects, so coordinate in your application; conditional writes such as If-None-Match help avoid overwriting.
    • Bucket configuration is eventually consistent: a deleted bucket can briefly still appear in a listing, and after enabling versioning for the first time AWS recommends waiting before writing objects.

    Older articles that describe eventual consistency for overwrites and deletes are out of date.

    What interviewers listen for
    • Strong read-after-write for PUT, overwrite and DELETE
    • LIST reflects writes immediately
    • Atomic per key; no multi-key transactions
    • Concurrent writes: latest wins; use conditional writes
    • Bucket configuration changes are eventually consistent
  52. 53.What is IAM Identity Center, and why is it preferred over IAM users for people?mid

    IAM Identity Center, the successor to AWS Single Sign-On, is the recommended way to give people access across an AWS Organization. Users sign in once, through the access portal or the CLI with aws sso login, and pick an account and a role.

    • Identity source: its own directory, Active Directory, or an external identity provider such as Okta or Microsoft Entra ID over SAML 2.0, with SCIM to sync users and groups automatically.
    • Permission sets define what someone can do. You assign a group plus a permission set to one or more accounts, and Identity Center creates matching IAM roles in each account.
    • Temporary credentials: every session is a role session with short-lived credentials, so there are no long-term access keys to leak, and leavers lose access once they're disabled in the identity provider.

    Compared with IAM users in every account, you get one place to manage people, MFA and offboarding, consistent access across many accounts, and CloudTrail records of who assumed which role. IAM users remain for rare edge cases.

    What interviewers listen for
    • Central workforce access across Organizations accounts
    • Connects external identity providers via SAML and SCIM
    • Permission sets become IAM roles in each account
    • Short-lived credentials, no long-term access keys
    • One place for MFA and offboarding
  53. 54.What are Lambda layers, and when would you use them?easy

    A layer is a .zip archive of libraries, a custom runtime or other files that you publish separately and attach to functions. Lambda extracts layers into /opt in the execution environment, and the managed runtimes include paths under /opt in their search paths, so your code imports them normally.

    Why use them:

    • Share dependencies across many functions instead of bundling the same library into each package.
    • Smaller deployment packages, so code updates upload faster and the console editor stays usable.
    • Separate concerns: update function code and dependencies independently.
    • Distribute a custom runtime or tooling such as monitoring agents.

    Limits and caveats: up to five layers per function, and their unzipped size still counts toward the function's deployment size limit. Layer versions are immutable and referenced by exact ARN, so upgrades are explicit. Layers only work with .zip functions; with container images you build dependencies into the image. AWS also advises against layers for Go and Rust, where dependencies belong in the compiled binary.

    What interviewers listen for
    • Zip of dependencies extracted to /opt
    • Share libraries across functions; smaller packages
    • Up to five layers; size counts toward the limit
    • Immutable versions referenced by exact ARN
    • Not supported for container image functions
  54. 55.An EC2 instance in a private subnet cannot reach the internet to download packages. How do you troubleshoot it?hard

    I'd follow the packet's path from the instance outward:

    • Route table: is the private subnet associated with a route table that sends 0.0.0.0/0 to a NAT gateway? A common mistake is the subnet silently using the main route table.
    • NAT gateway: is it available, sitting in a public subnet whose route table sends 0.0.0.0/0 to the internet gateway, with an Elastic IP?
    • Security group: outbound rules must allow the traffic, such as HTTPS on 443.
    • Network ACLs: the NACLs on both subnets must allow the outbound traffic and the return traffic on ephemeral ports, because NACLs are stateless.
    • DNS: can the instance resolve names? Check the VPC's DNS settings and any custom DHCP options.
    • The host: proxy settings or a local firewall.

    Tools: VPC Reachability Analyzer to test the path, VPC Flow Logs to see rejected traffic, and Session Manager for a shell without SSH. If the instance only needs AWS APIs like S3, a VPC endpoint may be the better fix.

    What interviewers listen for
    • Private route table: default route to a NAT gateway
    • NAT gateway in a public subnet routed to the IGW
    • Security group outbound rules and NACL ephemeral ports
    • Check DNS resolution and host-level proxies
    • Reachability Analyzer and VPC Flow Logs
  55. 56.How would you give an application in account A access to resources in account B?hard

    There are two main patterns, and in both each account must allow the access.

    • Assume a role in account B, the most general option. B creates a role whose trust policy trusts A's specific role and whose permissions policy grants the needed access. A's role gets an identity policy allowing sts:AssumeRole on that role's ARN. The application calls AssumeRole, receives temporary credentials and acts as B's role. It works for any service and keeps B in control.
    • A resource-based policy in account B, for services that support one, such as S3 bucket policies, KMS key policies, SQS, SNS and Lambda. B names A's role as the Principal, and A's role also needs an identity policy allowing those actions. The caller keeps its own identity, with no role switching. For KMS-encrypted data, the key policy must allow A too.

    When the other party is a third party, add an sts:ExternalId condition, as in the snippet, to prevent the confused deputy problem. Within your own organization, conditions like aws:PrincipalOrgID keep access inside it.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Principal": { "AWS": "arn:aws:iam::111111111111:role/partner-reporting" },
          "Action": "sts:AssumeRole",
          "Condition": { "StringEquals": { "sts:ExternalId": "b7f3c2e9-example" } }
        }
      ]
    }
    What interviewers listen for
    • Both accounts must allow cross-account access
    • Pattern one: assume a role in the target account
    • Pattern two: resource policy naming the external principal
    • KMS key policies must also allow the caller
    • External ID for third parties; aws:PrincipalOrgID in-org
  56. 57.A Lambda-based API backed by RDS starts failing under load with "too many connections" errors. What is going on, and how do you fix it?hard

    Each Lambda execution environment opens its own database connection. When traffic spikes, Lambda scales out to hundreds or thousands of concurrent environments, each holding a connection, while a relational database supports a finite number of connections tied to its size. Idle environments also keep their connections open until they're recycled.

    Fixes, usually combined:

    • RDS Proxy: a managed connection pool between Lambda and RDS or Aurora. Functions connect to the proxy, which shares a smaller set of database connections among them. It also handles failovers more gracefully and can use IAM authentication with credentials kept in Secrets Manager.
    • Cap concurrency: reserved concurrency on the function, or maximum concurrency on an SQS event source, so the database never sees more load than it can take.
    • Reuse connections: open the connection outside the handler so warm invocations reuse it.
    • Change the access pattern: serve reads from a cache or read replicas, put writes on a queue, or move high-concurrency data to DynamoDB.

    I'd confirm the diagnosis first with the database's DatabaseConnections metric in CloudWatch.

    What interviewers listen for
    • Each concurrent Lambda environment holds its own connection
    • RDS Proxy pools and shares database connections
    • Reserved concurrency caps load on the database
    • Open connections outside the handler for reuse
    • Caches, queues or DynamoDB for very high concurrency
  57. 58.How would you host a static website or single-page app on S3 and CloudFront securely?hard
    • Private bucket: keep Block Public Access on and skip the S3 website endpoint; the bucket is only an origin.
    • CloudFront with origin access control (OAC): CloudFront signs its requests to S3, and the bucket policy allows s3:GetObject only to the cloudfront.amazonaws.com service principal for that distribution's ARN. OAC replaces the older origin access identity and also works with SSE-KMS objects, as long as the key policy lets CloudFront decrypt.
    • HTTPS and domain: an ACM certificate, which for CloudFront must be in us-east-1, a Route 53 alias record, and a viewer protocol policy that redirects HTTP to HTTPS.
    • SPA routing: index.html as the default root object, and custom error responses that map 403 and 404 to /index.html, or a CloudFront Function that rewrites paths.
    • Caching: long TTLs for fingerprinted assets, short TTLs or an invalidation for index.html on each deploy.
    • Extras: AWS WAF, security headers through a response headers policy, and access logging.

    The result: HTTPS, global caching, and a bucket nobody can read directly.

    {
      "Version": "2012-10-17",
      "Statement": [
        {
          "Effect": "Allow",
          "Principal": { "Service": "cloudfront.amazonaws.com" },
          "Action": "s3:GetObject",
          "Resource": "arn:aws:s3:::acme-web-assets/*",
          "Condition": { "StringEquals": { "AWS:SourceArn": "arn:aws:cloudfront::111122223333:distribution/EDFDVBD6EXAMPLE" } }
        }
      ]
    }
    What interviewers listen for
    • Private bucket with Block Public Access on
    • OAC plus a bucket policy scoped to the distribution
    • ACM certificate in us-east-1 for CloudFront
    • Map 403 and 404 to index.html for SPA routes
    • Long TTLs for hashed assets, short for index.html
  58. 59.How do you run production workloads on Spot Instances safely?hard

    Spot capacity can be reclaimed at any time with a two-minute interruption notice, so the workload has to tolerate losing instances.

    • Pick suitable workloads: stateless web tiers behind a load balancer, containers, CI, batch, big data, and ML training with checkpoints. Not a single-instance database.
    • Diversify: allow many instance types and sizes across all AZs, and use the price-capacity-optimized allocation strategy so the fleet draws from the deepest capacity pools. Attribute-based instance type selection makes this easier.
    • Mix with On-Demand: an Auto Scaling group with a mixed instances policy keeps an On-Demand base, covered by Savings Plans, with Spot above it.
    • Handle interruptions: catch the notice through EventBridge or instance metadata, then drain connections, deregister from the target group and checkpoint work. Capacity Rebalancing lets the group launch replacements early when an instance is at elevated risk.
    • Design for retries: idempotent jobs, work pulled from queues, and no important state on the instance.

    On containers, ECS capacity providers and Karpenter on EKS build these practices in.

    What interviewers listen for
    • Two-minute interruption notice; design for instance loss
    • Diversify instance types, sizes and AZs
    • Use the price-capacity-optimized allocation strategy
    • On-Demand base plus Spot via mixed instances policy
    • Drain and checkpoint on notice; Capacity Rebalancing

Prefer multiple choice? All 21 AWS MCQs with answers →

esc