TechOps Examples
Hey — It's Govardhana MK 👋
Welcome to another technical edition.
Every Tuesday – You’ll receive a free edition with a byte-size use case, remote job opportunities, top news, tools, and articles.
Every Thursday and Saturday – You’ll receive a special edition with a deep dive use case, remote job opportunities and articles.
Top engineers at Anthropic and OpenAI say AI now writes 100% of their code.
If you're not using AI, you're spending 40 hours doing what they do in 4.
These 100+ Claude Code hacks fix that and help you ship 10x faster.
Sign up for The Code and get:
100+ Claude Code hacks used by top engineers — free
The Code newsletter — learn the latest AI tools, tips, and skills to code faster with AI in 5 minutes a day
If you’re not a subscriber, here’s what you missed last week.
To receive all the full articles and support TechOps Examples, consider subscribing:
One-time 25% OFF on all annual plans of memberships. Closes Soon.
🧠 USE CASE
Kubernetes Autoscaling - HPA, Cluster Autoscaler
Applications rarely receive a constant workload. One hour they handle thousands of requests, the next hour only a few. Kubernetes solves this by autoscaling, automatically adjusting compute resources based on demand.
Autoscaling in Kubernetes happens at multiple levels:
Pod level scaling: Horizontal Pod Autoscaler (HPA) and Vertical Pod Autoscaler (VPA)
Node level scaling: Cluster Autoscaler (CA)
Today we focus on the two most commonly used in production, HPA and Cluster Autoscaler, and how they work together.
Horizontal Pod Autoscaler (HPA)
HPA adjusts the number of Pod replicas in a deployment based on real time load metrics.

The Metrics Server collects resource usage from running Pods, such as CPU, memory, or custom metrics.
HPA continuously checks these values against defined thresholds.
When the average metric exceeds the target, HPA increases the replica count.
The Deployment controller creates new Pods to match the desired state.
When load decreases, HPA scales down by reducing replicas.
HPA ensures the application scales horizontally, creating more Pods for more work and fewer Pods when idle. It only manages Pods, not cluster capacity. If no space is left to schedule new Pods, they enter a Pending state.
Cluster Autoscaler (CA)
Cluster Autoscaler monitors the scheduling status of Pods. When Pods remain Pending due to insufficient cluster resources, CA requests new nodes from the infrastructure provider.

A Pod cannot be scheduled because there is not enough CPU or memory on any node.
CA detects this condition and increases the node count in the node group, such as AWS Auto Scaling Group, Azure VMSS, or GCP MIG.
The cloud provider provisions a new node.
Kubernetes schedules the Pending Pods on the new node.
When nodes stay underused for a defined period, CA can safely remove them.
Cluster Autoscaler ensures there is always enough capacity for Pods requested by workload controllers. Together, they keep workloads responsive and infrastructure cost efficient.

Production Notes
Metrics Server is mandatory for HPA. Without it, no scaling decisions are made.
HPA responds faster than CA. Node provisioning can take a few minutes depending on the cloud provider.
Configure resource requests and limits correctly so HPA decisions reflect actual workload pressure.
🔴 Get my DevOps & Kubernetes ebooks! (free for Premium Club and Personal Tier newsletter subscribers)
Looking to promote your company, product, service, or event to 50,000+ DevOps and Cloud Professionals? Let's work together.


