SRE Interview Preparation
Ace your technical interviews with flashcards and detailed STAR-format incident talk-tracks.
Cloud Cost Optimization
70 cardsCloud Cost Optimizationjunior
If a log category is "excluded" from a storage bucket via a routing rule, does that mean the data is deleted?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What are the three EKS security scanning tools covered and what does each focus on?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What does it mean for a metric to be "unused" in a cost-optimization context, and how would you check for this?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is a cluster autoscaler, and why does a Kubernetes cluster need one?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is a GCP SKU, and why would you filter billing data by it?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is a Graviton instance and why would you use it?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is an EC2 Launch Template and why does it matter for ASG migrations?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is an S3 (or GCS) lifecycle policy, and what problem does it solve?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is AWS Cost Explorer and why is it the first tool you enable for cost optimization?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is AWS Trusted Advisor, and what categories of recommendations does it provide?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is gcloud auth login used for?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is the difference between kubectl drain and kubectl cordon?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What is the difference between Reserved Instances and Savings Plans?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What's a quick, low-risk way to reduce Prometheus-related cost without removing any metrics?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What's the difference between GCP's _Default and _Required log buckets?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What's the difference between GitHub authentication over HTTPS with a password versus SSH keys?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What's the difference between Intel, AMD, and ARM/Graviton instance families on AWS, from a cost perspective?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What's the difference between reducing log retention and excluding log categories, as two separate cost-optimization techniques?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
What tags should every AWS resource have, and why?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
Why is resource tagging described as foundational to cost optimization work?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
Why might a compliance/audit log bucket be configured to retain data for 400 days even when the regulatory minimum is only 365 days?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationjunior
Why might an organization end up running two overlapping observability stacks (e.g., Prometheus/Grafana and a cloud-native monitoring tool) at the same time?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
A client asks for a "rollback plan" before you delete/filter historical log data. What do you need to clarify with them before proceeding?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
A client asks you to confirm that reducing a log bucket's retention period from 30 to 7 days won't affect their compliance posture. How would you verify and communicate this?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
A client has a $98K/month AWS bill. Walk me through how you'd approach reducing it.
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
A client wants to know why their cloud bill increased significantly compared to last month. How would you approach answering this credibly?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Explain the difference between excluding a metric via a monitoring console's UI versus blocking it at scrape time, and why the distinction matters for cost.
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Explain the difference in migration complexity between a standalone EC2 instance, an instance in an Auto Scaling Group, and an EKS node group instance — and why does the difficulty ordering change depending on migration direction?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Explain the PRC framework for right-sizing decisions.
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Explain the trade-off between using a managed cloud service versus self-hosting the equivalent open-source tool (e.g., a managed message queue vs. self-hosted Kafka).
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
How do you migrate 170 EC2 instances from Intel to AMD with zero unplanned downtime?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
How would you determine which compliance frameworks (e.g., SOC 2, ISO, HIPAA, PCI-DSS) apply to your organization's infrastructure, if you're a DevOps engineer without direct visibility into that decision?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
How would you structure a cost optimization engagement for a client with no existing infrastructure documentation?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Walk through an ASG Intel→AMD migration with zero downtime.
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Walk through how you'd approach a situation where a documented cloud CLI command fails, and you're not sure why.
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Walk through the safe, step-by-step process for migrating a standalone EC2 instance from Intel to AMD in a production environment.
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
What are the mandatory Kubernetes safeguards before running Spot instances in production?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
What is MaxSessions / right-sizing, and how do you determine the correct instance size?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
What's the difference between a bastion host and a full PAM (Privileged Access Management) solution like CyberArk, and when would you use each?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
What's the practical difference between a Savings Plan and a Reserved Instance, and when would you choose one over the other?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
When is it appropriate to stop EC2 instances to save cost, and when is it not?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Why does EKS node group migration require ~2–3 minutes of downtime, while ASG migration has zero downtime?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Why does the recommended migration path go Intel → AMD → ARM instead of directly Intel → ARM?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Why is GCP's Active Assist described as a meaningful differentiator versus AWS Trusted Advisor?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Why might a security scanning tool like ScoutSuite be described as evaluating an environment "from an attacker's perspective" rather than being purely compliance-driven, and when would you choose it over a compliance-framework-oriented tool like Prowler?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
Why would an organization choose ARM-based (Graviton) instances over x86 (Intel/AMD), and what's the catch?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationmid
You're told a GCP account's logging costs are unexpectedly high. Walk through how you'd investigate and address it.
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
A client insists on continuing to run both a self-hosted Prometheus/Grafana stack and a cloud-native monitoring tool in parallel, citing team preference. How would you approach cost optimization given this constraint, rather than pushing for consolidation?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
A client refuses to let you make application-level logging changes, but wants logging costs reduced. What are your levers, and what are the trade-offs of each?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
A client wants a 50/50 multi-cloud split "for cost savings." How would you push back or reframe this conversation?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
A colleague proposes buying AWS's recommended Savings Plan directly from the console's default suggestion. What risks would you flag, and what would you do instead?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
A Kubescape scan finds 8 of 14 control checks failing for a pod. The most critical finding is "running as root." What's the complete fix, and why does it matter?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
A team wants to move their stateful MySQL pod to a Spot instance to save costs. What would you tell them?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
Compare CastAI, Karpenter, and a standard Kubernetes cluster-autoscaler. When would you recommend each?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
Design a sequencing strategy for a multi-phase cloud cost optimization engagement, using the principle demonstrated in this session (start with the lowest-risk changes).
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
Design a systematic inventory-building process for a GCP account with no existing documentation, using the gcloud CLI. What categories would you capture, and why does the organizational structure matter?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
Design a systematic process for auditing and reducing an organization's Cloud Monitoring/observability costs, using the approach demonstrated in this session as a starting point.
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
Explain the tradeoffs between Compute Savings Plans, EC2 Savings Plans, and Reserved Instances.
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How do you implement FinOps for a multi-account AWS organization with 20 accounts?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How does Karpenter differ from a standard EKS managed node group with Cluster Autoscaler?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How should FinOps governance be structured across a multi-account AWS Organization to prevent runaway costs, without making the FinOps team a bottleneck for every provisioning decision?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How would you decide whether a given security finding from a general-purpose scanning tool (like ScoutSuite) warrants immediate remediation versus being logged for a later, more formal audit phase?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How would you design a cost-optimization approach for a client's EC2 fleet where 116 out of 170 instances are candidates for architecture migration, while ensuring no production risk?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How would you explain to a non-technical stakeholder (e.g., a compliance officer) the difference between "we stopped storing these logs in our operational bucket" and "we deleted this data," in a way that would satisfy an audit conversation?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How would you reconcile a scenario where the aggregate billing figure for a service (e.g., "30,000 GB over 90 days") doesn't cleanly match a per-day calculation presented separately (e.g., "30 GB/day × 30 days")? What would you do before presenting a cost-savings number to a client?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
How would you scope a security audit for a client with 13+ separate cloud accounts and no architecture documentation, and what audit categories would you define?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
What are the risks of buying a Savings Plan without doing the mathematics first?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
What's the risk profile of making infrastructure changes to a system with no clear ownership or point of contact, and how do you mitigate it?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
Why can you NOT specify an IAM instance profile in the Launch Template when creating EKS node groups via CLI?
Tags: aws, cost-optimizationReveal Answer →
Core Syscall Knowledge
Cloud Cost Optimizationsenior
You've configured a log exclusion rule and want to verify it's actually working as intended. What would your verification process look like, and why is timing important?
Tags: gcp, cost-optimizationReveal Answer →
Core Syscall Knowledge
Kubernetes
46 cardsKubernetesjunior
A pod is in Pending state. Where do you look first?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
A pod is stuck in Pending with no node or IP assigned. Should you check CNI logs first?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
Does a Pod Disruption Budget prevent a pod from being scheduled?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
In what order does the kubelet evict pods under memory pressure?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What are the three Kubernetes QoS classes, and how are they assigned?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What do kubectl drain and kubectl cordon each do, and why are both needed when migrating a node group?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What does it mean when a Kubernetes pod is stuck in Pending state with no node or IP assigned?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What is a taint and how does it affect pod scheduling?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What is the difference between a pod being OOMKilled and a pod being Evicted?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What is the difference between kubectl cordon and kubectl taint?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What's the difference between modifying an ASG's Launch Template directly versus creating a new Launch Template version?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What's the difference between requiredDuringSchedulingIgnoredDuringExecution and preferredDuringSchedulingIgnoredDuringExecution in Kubernetes affinity rules?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What's the difference between the Kubernetes scheduler's "filter" phase and "score" phase?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
What's the first command you'd run to understand why a specific pod is stuck in Pending state?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesjunior
Why might a Compute Savings Plan be preferable to an EC2 Instance Savings Plan for an organization using multiple AWS compute services?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
A checkout-api pod is Pending. kubectl describe pod shows: "0 nodes available: 1 node had taint {workload=batch:NoSchedule}, 1 node was cordoned, 1 didn't match pod's node affinity." What are the three fixes and which should you apply in a P0?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
A node has a new taint added to it, but the pods that were already running on that node before the taint was added are still running fine. Why doesn't the taint affect them?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
A production pod shows status Evicted. What should you check first, and why?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesmid
A team wants to use spot instances for a production Kubernetes workload but is worried about interruption risk. What architectural safeguards would you recommend?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Explain podAntiAffinity with required vs. preferred and give a real-world scenario where using required would cause a production outage.
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Explain why an engineer might deliberately avoid right-sizing instances during a cloud-to-cloud migration, even though both activities individually seem like sound cost-optimization practice.
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Explain why converting a pod anti-affinity rule from "required" to "preferred" is a meaningful fix, but why it might not be sufficient on its own.
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Walk through why migrating an EKS node group's architecture requires downtime, while migrating a standalone EC2 instance or an ASG-managed instance typically doesn't.
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
What is a "retry storm," and how did it appear in this incident?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesmid
What is the difference between nodeSelector and nodeAffinity, and why should nodeSelector be avoided in autoscaling environments?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
What's the difference between stating a root cause and stating a symptom in an incident report, and why does this distinction matter?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetesmid
What should be included in a Kubernetes incident's RCA to make it useful to both on-call engineers and non-technical executives?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Why is "allocatable memory" different from a node's advertised memory capacity, and why does that matter?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Why might kubectl top nodes show "normal" CPU while the node is actually in trouble?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetesmid
Why might simply uncordoning a previously-cordoned node "fix" a stuck pod, but not actually be a correct or complete fix?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetesmid
You've fixed a pod's anti-affinity rule from required to preferred, but the pod is still stuck in Pending. What would you check next?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
A namespace has no ResourceQuota/LimitRange, and a BestEffort batch job is co-located with a Burstable production service on the same node group. What's the risk, and how would you redesign this?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetessenior
A team's PDB, anti-affinity rule, and node selector are each individually reasonable, but together they create a scheduling deadlock. How would you design a process to catch this class of compound misconfiguration before it reaches production?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetessenior
A team wants to enforce strict one-pod-per-node placement for a critical, high-availability service using required pod anti-affinity, but their cluster doesn't reliably have enough distinct, correctly-configured nodes available to satisfy this at all times (e.g., during a temporary node group scaling event). What are the trade-offs of different approaches to this tension?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
Describe the interaction between nodeSelector, podAntiAffinity, and tolerations as AND conditions in the Kubernetes scheduler. How can these compound to deadlock a pod?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
Design a diagnostic approach for a Kubernetes incident where a pod is stuck in Pending, using the frameworks discussed in this session.
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetessenior
Design a safe, staged process for migrating a stateless EKS workload's underlying node group from Intel to a different processor architecture, including how you'd verify success and how you'd roll back if needed.
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
Design a systematic diagnostic sequence for a pod stuck in Pending state in an EKS cluster where all standard component health checks (nodes, CNI, kube-proxy) report healthy.
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
Explain the full causal chain that turned an architectural memory-allocation assumption into a customer-facing checkout outage.
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetessenior
How do PriorityClass/preemption and QoS-based eviction interact — are they the same mechanism?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetessenior
How would you decide which specific EC2 instances in a large fleet are good candidates for Savings Plan coverage versus Spot versus on-demand, as part of a comprehensive cost-optimization plan?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
How would you evaluate whether an RCA is "production grade," using the specific criteria discussed in this session, if you were reviewing a colleague's RCA before it goes to a client?
Tags: kubernetes, rcaReveal Answer →
Core Syscall Knowledge
Kubernetessenior
How would you prevent the specific class of incident demonstrated in this session (a taint added to nodes without corresponding tolerations being added to existing deployments) from recurring in a real production environment?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
What's the strategic value of running kube-hunter, kubescape, and kube-bench together, rather than choosing just one, and how would you prioritize adoption if resource/time constraints only allowed a phased rollout?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Kubernetessenior
Why can a CoreDNS-related outage be one of the hardest failure modes to diagnose from logs alone?
Tags: kubernetes, dnsReveal Answer →
Core Syscall Knowledge
Kubernetessenior
You're in a P0. The checkout-api pod has been Pending for 8 hours. The business wants it fixed NOW, but the engineering lead says "we can't compromise our HA topology." How do you resolve the conflict?
Tags: kubernetes, schedulingReveal Answer →
Core Syscall Knowledge
Linux & Networking
30 cardsLinux & Networkingjunior
If a TCP handshake to port 22 succeeds via netcat, but ping to the same host fails completely, what does that tell you?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
List the steps of an SSH login at a high level.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
Name five "common fix" checks for an SSH outage.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
SSH times out connecting to an EC2 instance. Ping also fails. But nc -zv <ip> 22 succeeds. What does this tell you?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
Walk through the OSI model layers from L1 to L7 in the context of debugging an SSH connection failure.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
What are two alternative ways to access an EC2 instance if SSH access is completely broken?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
What is a bastion/jump server and why use one?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
What is ARP and why does an arp -n check matter during SSH troubleshooting?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
What's the canonical incident-response sequence?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingjunior
What's the difference between an SSH connection that times out versus one that hangs after authentication succeeds?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
A production incident shows a mix of failing and succeeding network-layer tests (e.g., ARP fails, but a TCP connection to a specific port succeeds). How would you characterize this kind of failure mode, and what should your next diagnostic steps be?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
Four ways to access an EC2 instance when SSH is broken. List them in order of preference and explain the tradeoff.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
How can DNS cause intermittent SSH login lag?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
SSH is hanging at login (after entering password/accepting key), but it eventually connects after 30 seconds. What are the likely causes?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
The affected users change every hour and the app dashboard is green. What does that tell you?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
Walk through why a disciplined engineer would continue investigating after finding 100% packet loss on a ping test, rather than immediately concluding the network is down.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
What is the "green dashboard" problem?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
What is the UseDNS setting in sshd_config, and how could it cause an SSH session to hang specifically after authentication succeeds?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
Why might MaxSessions and resource exhaustion be ruled out quickly here?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingmid
You run ssh -vvv and see "Authentication succeeded (publickey)" followed by "channel 0 opened, shell allocated" — and then the terminal freezes. What do you investigate?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
An incident involves a compound fault with 5 contributing factors. You fix one (UseDNS), verify SSH works, and close the incident. Three days later, the same symptom returns. What went wrong and how do you prevent it?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
Critique relying on "what changed?" as your first question.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
Design a troubleshooting approach for an incident where SSH access to a critical production server is itself broken, preventing you from directly investigating the server's own configuration. What's your overall strategy?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
Explain how GSSAPI/Kerberos can inject login latency, and a counter-argument.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
Explain the "networking stack alive for TCP but dead for ICMP/ARP" scenario. What causes it and what does it indicate?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
How would you decide whether to apply a general troubleshooting framework like OSI versus a system-specific framework (e.g., a Kubernetes-specific troubleshooting model) to a given production incident?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
How would you detect a single "naughty user" doing heavy transfer / port-forwarding via the bastion?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
This incident was deliberately engineered as a combination of multiple contributing factors (a narrow ephemeral port range, a UseDNS setting, and a partially degraded kernel networking stack) rather than a single root cause. What does this imply about how you should approach root-cause analysis and remediation for real production incidents?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
Walk the layered (OSI-adapted) framework for a cloud VM and why data-link folds into physical.
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Linux & Networkingsenior
You're in a P0 incident. SSH to production is broken. SSM agent is offline. The instance has no serial console password set. The application is serving traffic via a different path (not SSH-dependent). What do you do?
Tags: linux, sshReveal Answer →
Core Syscall Knowledge
Structured Debugging
26 cardsStructured Debuggingjunior
How would you determine whether a DNS resolution failure is caused by your local machine's resolver or by a broader network/DNS problem?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
Name the five categories that production outages typically fall into.
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
Should you run multiple independent microservices inside a single Kubernetes pod?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
What does a high or rising number of established connections (as shown by ss -s) typically indicate?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
What is a "retry storm," and why can it be worse than an external attack?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
What is the difference between a liveness probe and a readiness probe in Kubernetes?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
What's the difference between what ping and curl actually test, and why might one succeed while the other fails?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
What's the difference between what ping, curl, and nc each actually test, and why would you use all three?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingjunior
Why is DNS often described as "the root of most production outages"?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
A pod is showing as "Not Ready" but is not restarting. What's the most likely cause and how would you investigate?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
A production incident shows CPU, RAM, and disk all reporting as "healthy" in monitoring, yet a specific service is clearly struggling. What's a diagnostic angle that basic resource monitoring might miss?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
A service status check that normally responds instantly is now taking 8 seconds to respond. What would you conclude, and what would you do next?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
An organisation's production outage was traced to a legacy load balancer's DNS resolver going down. What category of outage is this, and what's the systemic fix (not just the immediate fix)?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
Describe a structured approach to debugging a production outage where logs and metrics show no obvious errors, but customers are reporting latency and intermittent errors.
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
Explain the cascading failure pattern seen in the Slack outage, step by step.
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
In a team-based incident response scenario, how would you avoid the kind of coordination failure seen in this session (where one responder's unannounced fix invalidated another responder's ongoing diagnosis)?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
Walk through a structured approach to diagnosing an ambiguous "intermittent connectivity issue" on a Linux server, from the outside in.
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingmid
Walk through the OSI-layer approach to troubleshooting a suspected network issue.
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
A junior engineer on your team says a system issue "just fixed itself" during a live troubleshooting session, and they can no longer reproduce a bug you were actively diagnosing together. What questions would you ask before concluding the issue is actually resolved?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
Design a lightweight incident-response process for a mid-size engineering organisation that doesn't yet have a formal RCA/knowledge-base practice. What are the minimum viable components, based on the principles discussed in this session?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
Explain why "assumption-driven" debugging is inefficient in production incidents, and what specifically makes "structured" debugging faster — not just theoretically, but in terms of what's actually different about the process.
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
How would you design monitoring/alerting to catch a retry-storm pattern (like the one discussed in this session) before it causes a full outage, rather than discovering it only during manual ss -s inspection?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
How would you structure a live war-room troubleshooting session to be maximally useful for a cohort of already-experienced engineers, based on the feedback given in this session?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
In the unresolved bastion/SSH scenario, several plausible-sounding hypotheses were ruled out one by one (firewall rules, FD limit misconfiguration, max session caps, user permissions, DNS, multi-region routing, NAT gateway). What does this progressive elimination process illustrate about real incident response, and why might it be valuable that the session didn't resolve the incident?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
Why is it important to test DNS resolution against multiple targets/methods (local resolver vs. direct upstream query) rather than a single dig command, especially in a production incident?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Structured Debuggingsenior
You're debugging an intermittent SSH access issue affecting a random subset of users (not tied to specific identities), where restarting the SSH daemon provides only temporary relief before the issue recurs. What does this pattern suggest about the nature of the underlying problem, and what would you check next?
Tags: debugging, sreReveal Answer →
Core Syscall Knowledge
Career & Interview Strategy
21 cardsCareer & Interview Strategyjunior
What core fundamentals should every DevOps engineer master first?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategyjunior
What is an ATS and why does it matter for your CV?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategyjunior
What's the difference between how reads and writes are handled in this system's multi-region design?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategyjunior
What's the recommended CV section order?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategyjunior
Why does this system avoid synchronous cross-region database writes?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategyjunior
Why might a large fintech system choose to decompose into 500 microservices rather than a smaller number of larger services?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategyjunior
Why use Gmail over Yahoo on a résumé?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
Explain "open endpoint vs. closed endpoint" with examples.
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
Explain the "home-region ownership" pattern and how it prevents a double-withdrawal scenario in a multi-region banking system.
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
How does Naukri's matching differ from LinkedIn's, and how do you optimize each?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
How does the CAP theorem relate to the multi-region architecture decisions described in this session?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
Walk through a STAR-format RCA for a CoreDNS outage.
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
What operational risk does over-decomposing an application into too many microservices introduce, according to the discussion in this session?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategymid
Why do experienced engineers' CVs fail to stand out, and what's the fix?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
A multi-region system's architecture team wants to justify why they've deliberately built in the assumption that any given region can fail at any time, rather than treating regional failure as a rare edge case. How would you frame this design philosophy to a skeptical stakeholder who sees it as over-engineering?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
Critique the advice to set salary fields to 0 and notice period to 30 days. What are the risks?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
Design a strategy for a fintech platform that wants lower write latency for non-local users than the home-region ownership pattern provides, without sacrificing correctness.
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
How portable is cloud expertise, and how do you demonstrate it?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
How would you evaluate whether an organization considering a shift from fine-grained microservices to a coarser "service-based" architecture (as one participant described) is making the right call?
Tags: career, programReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
How would you present knowledge you have but never ran in production, without misrepresenting yourself?
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Career & Interview Strategysenior
Make the case against this workshop's keyword/volume-maximization approach.
Tags: career, interviewReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscaling
17 cardsKubernetes Autoscalingjunior
How does Karpenter know which subnets and security groups to use when provisioning a new node?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingjunior
What does it mean that Karpenter has a "consolidation delay" when scaling down?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingjunior
What is CastAI and how is it different from Karpenter?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingjunior
What is the difference between Karpenter and the Kubernetes Cluster Autoscaler?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingjunior
What's the core functional difference between Karpenter and CastAI?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingjunior
Why might Karpenter choose a Spot instance for one workload and an on-demand instance for another?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingmid
A team says "Karpenter already optimizes our costs, so we don't need a separate cost-monitoring tool." How would you respond?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingmid
How would you decide whether a given batch workload is a good candidate for Spot instances?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingmid
Walk through the 6 steps to install Karpenter on an EKS cluster.
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingmid
What are the main risks or downsides of over-customizing Karpenter's NodePool configuration for every workload type upfront?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingmid
What does Karpenter's consolidation feature do and why does it save cost?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingmid
You're on-boarding Karpenter into a production EKS cluster that currently uses custom hardened AMIs and a blue-green patching strategy. How do you maintain this with Karpenter?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingsenior
A prospective client is based in the European Union and has strict data-residency requirements. You're evaluating CastAI (US-only SaaS) versus Karpenter (open-source, in-cluster) for their EKS autoscaling needs. Walk through your decision process.
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingsenior
Compare capacity planning approaches: traditional (static max) vs. Karpenter-managed. When is each appropriate?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingsenior
Design a node pool strategy for an EKS cluster that needs to support both a real-time customer-facing calculation service and a large nightly Spark analytics job, using the principles discussed in this session.
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingsenior
How would you use the "functional vs. non-functional requirement" distinction to structure a broader infrastructure tool evaluation (not just autoscaling), and why does treating security/compliance as a distinct, dedicated evaluation category matter?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
Kubernetes Autoscalingsenior
Karpenter provisioning is taking longer than expected during a traffic spike. What would you investigate?
Tags: kubernetes, autoscalingReveal Answer →
Core Syscall Knowledge
AI Systems Design
16 cardsAI Systems Designjunior
What are the four core reasons naive AI agents fail in production?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designjunior
What is the general rule of thumb for how response latency affects user experience?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designjunior
Why is storing conversation history in a pod's local RAM a problem in a Kubernetes environment?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
Describe the three-tier memory hierarchy used in a production agentic system and what belongs in each tier.
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
Explain the three-tier memory hierarchy for production AI agents.
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
How does a DAG orchestrator solve the tool cascading problem?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
How does converting a sequential agent workflow into a DAG (Directed Acyclic Graph) improve performance?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
What is speculative execution in the context of AI agents?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
What is "tool cascading" and why is it a problem?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
What is TTFT and how do you optimise it?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
What is wrong with a naive LLM agent and why can't it be used in production?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designmid
Why use a small language model (SLM) for intent classification instead of the same large LLM used for reasoning?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designsenior
Explain the difference between reducing TTFT via system-prompt caching versus KV caching.
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designsenior
In the layered security model described for agentic systems, why is separating the "prompt construction" layer from the "data access" layer an effective mitigation against prompt injection?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designsenior
Walk through the full optimized architecture for a multi-intent customer support query, from user input to final response.
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
AI Systems Designsenior
What is speculative execution in the context of AI agents, and what latency does it actually reduce?
Tags: ai-agents, llmReveal Answer →
Core Syscall Knowledge
CI/CD & Automation
14 cardsCI/CD & Automationjunior
What's the difference between in-place patching and immutable rotational patching?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationjunior
Why might checking only that a pod's status is "Running" and passing its readiness probe be insufficient to confirm a service is actually healthy after a deployment?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationjunior
Why should multi-region deployments or patching operations always be sequential rather than parallel?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
A self-service application allows developers to request cloud resource access, which is then approved by their manager and automatically provisioned. What are the key architectural components needed to build this safely?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
Explain the difference between in-place patching and immutable (rotational) patching. Which should you use and why?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
Explain the expand-migrate-contract pattern for database schema changes, and why the order matters.
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
How do you handle DB schema migrations during a canary deployment where two versions of the application run simultaneously?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
The 3 phases of a production CI/CD pipeline — what are they and what does each gate check?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
Why did this organization deliberately avoid including patching/security-hardening automation in the same self-service application used for routine access requests?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
Why is configuration automation kept separate from a self-service application?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationmid
Why is rollback done in parallel across all regions while forward deployment is sequential?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationsenior
A participant in a training session directly tells the instructor that the material has stayed too high-level and hasn't explained the system-design reasoning behind key architectural choices. How should this kind of feedback be handled, and what does the instructor's actual response in this session illustrate about good practice?
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationsenior
A team wants to run a canary deployment where a new application version requires a database column that doesn't exist yet, and the old version must keep running correctly during the rollout. Walk through exactly how you'd sequence this change to avoid an outage.
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
CI/CD & Automationsenior
Design a patching automation system for a 3-region, 500-microservice Kubernetes-based infrastructure that minimizes the risk of a patching-induced production outage.
Tags: cicd, automationReveal Answer →
Core Syscall Knowledge
Systems Design & Architecture
5 cardsSystems Design & Architecturemid
How does Kong rate limiting work and how is it different per endpoint?
Tags: architecture, microservicesReveal Answer →
Core Syscall Knowledge
Systems Design & Architecturemid
How does rollback work, and why not kubectl rollout undo?
Tags: architecture, microservicesReveal Answer →
Core Syscall Knowledge
Systems Design & Architecturemid
How does the access automation work for onboarding a new engineer?
Tags: architecture, microservicesReveal Answer →
Core Syscall Knowledge
Systems Design & Architecturemid
How does the canary deployment work and why is it done via Istio rather than a separate Deployment?
Tags: architecture, microservicesReveal Answer →
Core Syscall Knowledge
Systems Design & Architecturemid
Walk through how TitanGrid deploys a new version to production.
Tags: architecture, microservicesReveal Answer →
Core Syscall Knowledge
Linux & Scripting
2 cardsLinux & Scriptingjunior
What is the difference between #!/bin/bash and #!/bin/sh, and which should you use in production?
Tags: linux, shell-scriptingReveal Answer →
Core Syscall Knowledge
Linux & Scriptingsenior
Why are set -e, set -u, and set -o pipefail critical in production Shell Scripts?
Tags: linux, shell-scriptingReveal Answer →
Core Syscall Knowledge
Network Protocols
1 cardsNetwork Protocolsmid
What is the sequence of system calls executed when a Linux process establishes a TCP socket connection?
Tags: linux, syscallsReveal Answer →
Core Syscall Knowledge