Reference library

SRE Field Manuals

Exhaustive references on multi-region topology, cost consolidation, and scalable systems — written the way you'd want them at 3am.

24 manuals
16h of reading
17 topics covered
Topics #cost-optimization 22#ai-agents 22#aws 21#kubernetes 20#debugging 18#cloud-k8s 18#gcp 16#cicd 14#autoscaling 10#dns 9
senior 45m read

Production-Grade AI Agent Systems Design

A full treatment of production agent architecture — the failure modes of naive designs, the four structural fixes, and the security surface agentic systems introduce.

#kubernetes#aws
Read →
senior 45m read

AI Agent Systems Design: The Four Failure Modes

The four ways naive agent architectures fail in production — latency, state, tool cascading and observability — and the architectural answer to each.

#ai-agents#systems-design
Read →
senior 45m read

DevOps & SRE Career Optimization: Resume and Job-Platform Guidance

How to structure a DevOps/SRE resume that survives ATS filtering, and how to present the same experience across LinkedIn and job platforms.

#cloud-k8s#gcp
Read →
junior 20m read

Naukri & LinkedIn Profile Optimization for DevOps Engineers

Optimizing Naukri and LinkedIn profiles for DevOps and SRE roles: headline, skills, keyword targeting and recruiter visibility.

#kubernetes#aws
Read →
senior 45m read

Resume & CV Optimization for DevOps Engineers

A section-by-section rebuild of a DevOps/SRE CV — structure, wording, and how to frame achievements so they survive both ATS and a recruiter skim.

#kubernetes#cloud-k8s
Read →
mid 25m read

Debugging Memory Leaks in Go Runtime

A step-by-step SRE guide to diagnosing and mitigating high memory consumption inside compiled Go binaries using pprof heap visualization tools.

#go#profiling
Read →
mid 20m read

DevOps & SRE Career Optimization: The Complete Guide

Structuring a DevOps/SRE career narrative: an ATS-survivable resume, the STAR and RCA framing for experience, and profile optimisation across job platforms.

#kubernetes#cloud-k8s
Read →
senior 45m read

Karpenter & CastAI: Node Pool Design and Tool Selection

Designing Karpenter node pools for mixed real-time and batch workloads, and deciding when CastAI earns its place alongside them.

#kubernetes#cloud-k8s
Read →
senior 45m read

Structured Debugging in Practice: A Worked Walkthrough

Structured debugging applied end to end: forming a hypothesis, ruling out layers in order, and reaching a root cause without guessing.

#kubernetes#cloud-k8s
Read →
senior 35m read

Kubernetes Node Optimization: Karpenter & Autoscaling

A deep dive into Kubernetes autoscaling strategies, configuring Karpenter NodePools for cost-effective bin packing and rapid scheduling.

#kubernetes#aws
Read →
senior 45m read

EKS Cost Optimization with Karpenter and CastAI

Cutting EKS spend with Karpenter and CastAI — capacity planning, node pool design, migrating off Cluster Autoscaler, and where each tool pays off.

#kubernetes#cloud-k8s
Read →
senior 45m read

AWS Cost Optimization: A Healthcare Platform Case Study

A layered AWS cost optimization walkthrough on a healthcare platform: tagging, EC2 right-sizing, CPU architecture migration, storage and Savings Plans.

#kubernetes#cloud-k8s
Read →
senior 45m read

AWS Cost Optimization: Instance Migration, Spot and Savings Plans

Continuing an AWS cost programme: Intel to AMD to ARM migration, right-sizing, Spot adoption, Savings Plans strategy and EKS security scanning.

#kubernetes#cloud-k8s
Read →
senior 45m read

Titan Grid: Architecting a 500-Microservice Fintech Platform

Architecting a 500-microservice fintech platform from first principles — environment strategy, cluster topology and the trade-offs behind each decision.

#kubernetes#cloud-k8s
Read →
senior 45m read

Titan Grid: Configuration Automation and CI/CD Pipeline Design

Configuration automation, self-service tooling and CI/CD pipeline design for a large multi-environment microservice platform.

#kubernetes#cloud-k8s
Read →
senior 45m read

Masterclass: Designing Scalable, Low-Latency, Production-Grade AI Agent Systems

Designing agent systems that hold up under production load: DAG orchestration, a three-tier memory hierarchy, intent routing, speculative execution and TTFT optimisation.

#kubernetes#aws
Read →
senior 45m read

AWS Cost Optimization: Engagement Overview and Approach

How a real AWS cost optimization engagement is scoped and sequenced, from infrastructure inventory to a prioritised savings roadmap.

#kubernetes#cloud-k8s
Read →
senior 45m read

EC2 and EKS Migration with Security Scanning

Migrating EC2 and EKS node groups across CPU architectures, with Kubernetes security scanning wired in as a recurring job.

#kubernetes#cloud-k8s
Read →
senior 45m read

Titan Grid: Platform Architecture from First Principles

Designing a 500-microservice fintech platform from scratch: requirements, environment layout, cluster boundaries and the reasoning behind each choice.

#kubernetes#cloud-k8s
Read →
senior 45m read

Titan Grid: Configuration Automation, Self-Service and CI/CD Design

Configuration automation, self-service applications and CI/CD pipeline design for a platform spanning hundreds of services.

#kubernetes#cloud-k8s
Read →
senior 45m read

Client Engagement Walkthrough: Fintech GCP Infrastructure

A walkthrough of real fintech GCP infrastructure — cost exposure, security posture and high-availability gaps, and what to fix first.

#kubernetes#cloud-k8s
Read →
junior 20m read

GCP Cloud Monitoring Cost Optimization and Security Scanning

Reducing GCP Cloud Monitoring spend without losing signal, and layering security scanning onto the same infrastructure.

#kubernetes#cloud-k8s
Read →
senior 45m read

GCP Inventory Building and Cloud Logging Cost Optimization

Building a GCP resource inventory from nothing, then using it to find and cut Cloud Logging costs.

#cloud-k8s#aws
Read →
junior 20m read

GCP Cloud Logging Cost Optimization: An Implementation Walkthrough

Implementing GCP Cloud Logging cost reductions end to end: exclusion filters, sinks, retention tuning and verifying the savings.

#kubernetes#cloud-k8s
Read →