Reduce the AWS Bills
We have hearda lot of people complain about growing AWS bills that eat into their numbers. Some blame AWS billing policies, others blame their fate. The fact is that the bill is high because of some common problems that can be identified and fixed.
This blog gives a detailed account of how we saved a startup, pulling their bill down to 60%. It gives precise details and checklists for how to go about the task. It covers the technical as well as non-technical problems that lead to a high AWS bill.
Ofcourse, we changed all the names!
A Field Guide to AWS Cost Optimization
- Part 0 β The Bill That Broke the CFO](#part-0–the-bill-that-broke-the-cfo) (start here)
Book One β The Infrastructure (fix it this quarter)
- Part 1 β The Instance That Was Too Big to Fail (right-sizing, Graviton, Spot, scheduling)
- Part 2 β Commitment Issues (Savings Plans & Reserved Instances)
- Part 3 β Redundancy Theater (Multi-AZ/region overuse, over-replication, backups)
- Part 4 β The Case of the Mysterious NAT Gateway (networking & data transfer)
- Part 5 β The Bucket That Ate the Budget (S3 & storage lifecycle)
- Part 6 β Full Scan of Shame (databases & analytics)
- Part 7 β Serverless, Not Costless (Lambda & event patterns)
- Part 8 β Cache Rules Everything Around Me (caching & CDN)
- Part 9 β Everything Is a Kafka, Nothing Is a Kafka (over-engineering)
- Part 10 β The Boring Stuff That Saves Millions (governance, security, multi-account)
Book Two β The Architecture (refactor it over quarters)
- Part 11 β The Distributed Monolith and Other Horror Stories (anti-patterns I)
- Part 12 β It Compiles, Ship It (anti-patterns II)
Book Three β The Humans (the real reason)
- Part 13 β The Real Proble - Humans not Servers (the organizational blockers)
- Epilogue β Ninety Days Later (results, the two changes that stuck, the prioritization napkin)
Part 0 β The Bill That Broke the CFO
It was a Tuesday, which is when bad news at Nimbus traditionally arrives.
Priya Nair was three sips into her coffee when the all-hands calendar invite landed: “AWS Spend β Urgent β 30 min.” No agenda. Just the digital equivalent of a parent saying “we need to talk.” She knew, the way you know a pothole is coming, that this was going to be her problem soon.
Nimbus, for the uninitiated, is a Series-C consumer photo-sharing app. Forty million people use it to post pictures of their lunches, their dogs, and β increasingly β AI-generated portraits of their dogs eating lunch. It is a lovely business. It also, as of last month, spends $340,000 a month on AWS, a number that had grown so quietly and so consistently that nobody had thought to ask why until the CFO, doing the thing CFOs do, looked at a graph.
The graph went up and to the right. The revenue graph also went up and to the right, but noticeably less steeply. When two lines diverge like that, someone eventually gets a Tuesday meeting.
In that meeting, three things happened. First, Chad Ellison β Nimbus’s CTO, a man who owns more Patagonia vests than the average Patagonia store β explained that a big cloud bill is “a sign of scale” and that he’d literally bragged about it at a conference last month. (“We spend the GDP of a small island on AWS,” he’d said, to applause.) Second, Marcus Webb β Principal Architect, designer of the sacred Nimbus Platform 2.0 β insisted the architecture was correct and that any savings would compromise resilience “at scale.” Third, the CFO looked at Priya, the most senior engineer in the room who hadn’t yet said anything defensive, and said: “You. Figure out where the money’s going. You’re the FinOps lead now.”
Priya did not want to be the FinOps lead. Nobody wants to be the FinOps lead. It is a job title that sounds like a punishment and, for the first two weeks, feels like one. But she opened Cost Explorer, and what she found over the following ninety days is the reason this series exists.
What this series actually is
This is The Nimbus Chronicles β a field guide to AWS cost optimization told as the story of one company clawing back a runaway bill. It is meant to be funny, because the alternative to laughing at your cloud bill is crying at it, and it is meant to be accurate, because a joke that teaches you the wrong thing is just an expensive joke.
Every episode follows Priya through one category of waste: what it looked like at Nimbus, why it was happening, the real numbers, and exactly how she fixed it. The cast will become familiar. There’s Marcus, who can justify any expense with the phrase “we might need it at scale.” There’s Deepak, a genuinely brilliant engineer who has decided that the monolith β lovingly known as PhotoBlob β is finished and shall never be touched again. There’s Jenny, the junior engineer who actually reads the bill and keeps getting ignored. There’s Big Bertha, an m5.24xlarge running a single nightly cron job at 3% CPU. And there is the NAT gateway. Oh, we’ll get to the NAT gateway.
The map
The waste at Nimbus fell into three broad territories, and so does this series.
First, the infrastructure β the knobs you can turn this quarter. Oversized instances and the wrong chips (Part 1). Buying commitments like an adult (Part 2). Redundancy you’re paying for and will never need (Part 3). The invisible networking bill and its patron saint, the NAT gateway (Part 4). Storage that never forgets (Part 5). Databases behaving badly (Part 6). Serverless that turned out not to be costless (Part 7). Caching, or the art of not doing the same work twice (Part 8). Over-engineering, a.k.a. running a Kafka cluster for 500 messages a day (Part 9). And the boring governance plumbing that quietly saves millions (Part 10).
Second, the architecture β the technical debt you’ll refactor over quarters, not days. The distributed monolith and its friends (Part 11). And all the ways application code fights the cloud instead of using it (Part 12).
Third, the humans β because the real reason a bill stays huge is almost never technical (Part 13). Egos, incentives, vanity metrics, and the CTO who brags about the number that’s bleaking money.
We close with an epilogue: where Nimbus landed, what actually moved the needle, and the two changes that made everything else stick.
A note before we start: cost optimization is not about making everything as small and cheap as possible. A checkout service that falls over on Black Friday to save forty dollars is not a win. The goal is to stop paying for things that create no value β idle cores, duplicated data, redundancy nobody needs, work done twice β while keeping every dollar that actually buys you reliability, speed, or sleep. Priya’s north star, taped to her monitor by week three, read: “Cheap where it doesn’t matter, so we can spend where it does.”
Pour a coffee. Bertha’s been running all night for no reason. Let’s go find out why.
Part 1 β The Instance That Was Too Big to Fail
The email arrived at 8:47 a.m. Subject line: “Question about AWS.” No greeting. Just a screenshot of the Nimbus bill β $340,127.44 β and a single sentence from the CFO: “I was told this was a photo app, not a particle accelerator.”
Priya Nair read it twice. She was a Staff Engineer who, as of eleven minutes ago, was also the “FinOps lead” β a title she had not applied for and could not find in the org chart. She poured coffee, opened Cost Explorer, and filtered to EC2. Compute was 61% of everything. Somewhere in that number was a machine the team had affectionately named Big Bertha.
Big Bertha was an m5.24xlarge. Ninety-six vCPUs. 384 GiB of RAM. Roughly $4.60 an hour on-demand, which is about $3,350 a month, which is about $40,000 a year. Bertha ran exactly one thing: a nightly cron job that resized some thumbnails at 3 a.m. Priya pulled up CloudWatch. Bertha’s peak CPU utilization, ever, was 3%.
“We might need it at scale,” said Marcus Webb, appearing behind her like a compliance ghost. He had designed Bertha. He would defend Bertha.
Priya sipped her coffee. “Marcus. It’s running a cron job. At three percent. You could do this on a Raspberry Pi taped to the wall.”
The most expensive words in the cloud: “just in case”
Compute waste at Nimbus wasn’t one villain. It was a hundred small ones. Here’s the tour Priya took, clipboard in hand.
Instances that are simply too big
Bertha is the mascot, but she has cousins everywhere. Over-provisioned instances β sized for an imagined future β are the single most common form of cloud waste. The fix is boring and effective: right-sizing. AWS Compute Optimizer analyzes actual CPU, memory, and network history and tells you, per instance, “you could drop two sizes.” Dropping Bertha from a 24xlarge to, say, a 2xlarge isn’t a discount; it’s a 12x cut on that line. Trusted Advisor flags the low-utilization offenders too. Priya started a spreadsheet titled “Bertha and Friends.”
Instances that do nothing at all
Then there were the idle instances β dev boxes spun up for a demo in 2024, still humming, doing nothing but generating charges. Eight of them. t3.large, ~$60/month each. Nobody could say what they were for, which is the surest sign you can turn them off.
Old-generation instances
Nimbus was running a fleet of m4 instances. AWS’s newer generations β m5, m6 β are cheaper and faster per dollar. Same story for t2 β t3: t3 is a straight upgrade at lower cost. Moving an m4 to an m6i is often a ~10β20% price improvement for more performance. There is almost never a reason to stay on old gen; you’re paying more to run slower.
Intel/AMD vs. ARM Graviton
The bigger lever: Graviton. AWS’s ARM-based chips (the g suffix β m6g, c7g, r7g) run about 20% cheaper than the equivalent Intel/AMD instances, often with better performance-per-watt. If your workload compiles for ARM β and most modern runtimes (Java, Go, Node, Python) do β this is close to free money. Priya flagged the stateless web tier as candidate one.
The wrong instance family entirely
Instance families are shaped for different jobs: c (compute) is CPU-heavy and RAM-light; m is balanced; r is memory-heavy. Nimbus was running an in-memory cache on an m5 β starving for RAM, wasting vCPUs. Moving it to an r family gave more memory per dollar. Meanwhile a CPU-bound encoder sat on an r instance paying for RAM it never touched. They had it backwards. Matching family to workload shape is free performance.
Burstable T-family misuse
The T-family (t3, t4g) is burstable: it earns CPU credits when idle and spends them when busy. Brilliant for spiky, low-average workloads. But Nimbus had a steadily-busy API service on a t3.medium, blowing through credits and β because they’d left unlimited mode on β paying surcharge fees for sustained CPU. The T-family charges you both ways: idle-but-oversized wastes the baseline; busy-past-credits racks up overage. A steady workload belongs on an m or c, not a T.
The idle GPU
In us-west-2 sat a p3.2xlarge β a GPU instance at roughly $3/hour β spun up for an ML experiment that ended in Q1. Nobody killed it. That’s ~$2,200/month for a graphics card doing absolutely nothing. Idle GPUs are the most expensive idle you can own. Priya terminated it on the spot and felt something close to joy.
The storage nobody looks at
Compute has a shadow: the disks and images bolted to it.
- Over-provisioned EBS volumes: 1 TB volumes that were 6% full, sized “to be safe.” You pay for provisioned capacity, not usage.
- gp2 β gp3: The default
gp2volume type is superseded by gp3, which is about 20% cheaper per GB and lets you tune IOPS and throughput independently instead of buying a bigger disk just to get faster I/O. Migrating a fleet of gp2 volumes to gp3 is often a one-click, no-downtime ~20% storage cut. - Orphaned snapshots and unused AMIs: Thousands of EBS snapshots from long-deleted volumes, plus a graveyard of AMIs nobody had booted in two years β each quietly billing for storage. Priya found ~$1,900/month here alone.
Time is a resource, and you’re paying for all of it
Non-prod running 24⁄7
Nimbus’s dev, staging, and QA environments ran around the clock. Nobody works nights and weekends (allegedly). If you schedule non-prod off outside business hours β say 8 a.m.β8 p.m., weekdays only β you eliminate roughly 70% of those hours. Cut a $30,000/month non-prod footprint by 70% and you’ve found $21,000/month by doing nothing but turning things off when everyone’s asleep.
No autoscaling
The production web fleet was a static fleet sized for peak β enough servers for the busiest minute of the busiest day, running that many 24⁄7. That’s like renting a 40-seat bus to commute alone because once a year you carpool. Autoscaling adds and removes capacity as load changes, so you pay for the average, not the maximum.
Autoscaling, but misconfigured
Worse than no autoscaling is autoscaling that scales up fast and scales down never. Nimbus’s group had a timid scale-down policy β a 30-minute cooldown and a conservative threshold β so it ballooned during traffic spikes and then sat there, fat and happy, for hours. Tune the scale-down as deliberately as the scale-up, or you get all the cost of peak with none of the savings.
Not using Spot
For anything fault-tolerant, batch, or stateless β the thumbnail pipeline, the video transcoders, CI runners β Spot Instances offer up to 90% off on-demand. Yes, AWS can reclaim them with two minutes’ notice, which is exactly why you use them for interruptible work, not your database. Nimbus’s entire batch encoding fleet was on-demand. Moving it to Spot was a potential ~$18K/month swing.
Half-empty container nodes
Finally, the containers. Nimbus ran EKS, and Priya found nodes at 40% packed β big EC2 instances hosting a handful of small pods, the rest of the capacity billed and empty. This is a bin-packing problem. Tools like Karpenter or cluster-autoscaler consolidate pods onto fewer, right-sized nodes and drain the empties. For spiky container work, Fargate Spot runs tasks on spare capacity at a deep discount. Better bin-packing alone took the cluster from ~40% to ~80% utilization β halving the node count.
By lunch, Priya had a spreadsheet, a Graviton migration plan, and a very quiet GPU. Marcus was still muttering “at scale.” Big Bertha, mercifully, was scheduled for demolition.
Takeaways
- Right-size over-provisioned instances β use Compute Optimizer and Trusted Advisor to find machines running far below capacity, and drop them one or more sizes.
- Kill idle instances β if nobody can explain what a running box does, that’s your cue to turn it off.
- Upgrade old-generation instances β move m4βm5/m6 and t2βt3 for lower cost and better performance.
- Adopt Graviton (ARM) β the
g-suffix instances run ~20% cheaper than Intel/AMD for compatible workloads. - Match the instance family to the workload β c for compute-bound, m for balanced, r for memory-heavy; the wrong family wastes what you’re paying for.
- Stop misusing the burstable T-family β it charges both ways (wasted baseline when oversized, credit overage when sustained); steady workloads belong on m/c.
- Terminate idle GPU instances β an unused GPU box is the most expensive idle you can own.
- Right-size over-provisioned EBS volumes β you pay for provisioned capacity, not what you actually use.
- Migrate gp2 β gp3 β ~20% cheaper per GB with independently tunable IOPS/throughput, usually no downtime.
- Delete orphaned EBS snapshots and unused AMIs β they bill quietly forever.
- Schedule non-prod off nights and weekends β eliminate ~70% of those hours by turning dev/staging/QA off when nobody’s using them.
- Add autoscaling β never run a static fleet permanently sized for peak.
- Fix misconfigured autoscaling β tune scale-down as carefully as scale-up, or you keep peak cost forever.
- Use Spot Instances β up to 90% off for fault-tolerant, batch, and stateless workloads.
- Bin-pack container nodes β use Karpenter/cluster-autoscaler and Fargate Spot to eliminate half-empty ECS/EKS nodes.
Part 2 β Commitment Issues
Chad Ellison swept into the FinOps sync forty minutes late, Patagonia vest zipped to the chin, radiating the confidence of a man who had just told a conference audience that Nimbus spends “the GDP of a small island” on AWS. He said it like a boast. Priya had watched the recording. The audience had laughed. The CFO, she suspected, had not.
“I’ve been thinking,” Chad announced, which was never a phrase that saved money. “We should just buy three-year, all-upfront, everything. Lock in the discount. Maximum savings. Boom.” He mimed a small explosion with both hands.
Priya set down her coffee. “Chad. If we commit three years to our current fleet, and half that fleet is Big Bertha and eight idle dev boxes, we will have signed a thirty-six-month lease on machines we’re actively trying to delete.”
Chad’s hands lowered slowly. “…Go on.”
The problem with paying full price forever
Right now, Nimbus paid 100% on-demand for its entire compute footprint β including a rock-steady baseline that never, ever went away. That baseline was the mistake. On-demand is the flexible, pay-per-second, no-strings pricing. You pay a premium for the privilege of walking away at any moment. But Nimbus was never going to walk away from its core API servers. It was paying walk-away prices for machines that had been running continuously for two years.
That’s what Savings Plans and Reserved Instances are for. In exchange for a 1-year or 3-year commitment, AWS discounts your compute by up to ~72%. You promise to spend a certain amount (Savings Plans) or reserve specific capacity (RIs); AWS gives you a large discount on the part of your usage that was going to happen anyway.
Savings Plans vs. Reserved Instances
- Reserved Instances (RIs) are the older model. You reserve a specific instance type in a specific region (Standard RIs) for 1 or 3 years. Their one superpower: a Zonal RI can guarantee capacity reservation in a given Availability Zone β useful when you absolutely must be able to launch that instance during a regional crunch.
- Savings Plans are the newer, simpler model. You commit to a dollar-per-hour spend (e.g., “$20/hour of compute”) rather than to specific instances, and AWS applies the discount automatically to matching usage.
Both hit up to ~72% off on-demand at the 3-year, all-upfront end. But the flexibility differs enormously, and that’s where Chad’s plan went sideways.
The wrong commitment type
There are three flavors, and picking wrong is its own kind of waste:
- Compute Savings Plans β the most flexible. The discount follows your usage across instance family, size, region, OS, and even between EC2, Fargate, and Lambda. Migrate from m5 to Graviton m7g? Move regions? The Savings Plan just… keeps applying. Slightly smaller max discount, dramatically less risk.
- EC2 Instance Savings Plans β a deeper discount, but locked to a specific instance family in a specific region (e.g., “m-family in us-east-1”). You can change size and OS within that family, but stray outside it and you’re back to on-demand.
- Reserved Instances β deepest and most rigid; best reserved (pun intended) for stable workloads you’re certain won’t move, or when you specifically need that zonal capacity reservation.
Priya’s rule for Nimbus: default to Compute Savings Plans for flexibility, use EC2 Instance SPs or RIs only for the workloads so stable and so understood that the extra few points of discount are worth surrendering the ability to change your mind. Given that they were mid-migration to Graviton (see Part 1), locking into an EC2 Instance SP for the Intel m-family would have been paying extra to strand themselves on the chips they were trying to leave.
The one insight that matters: commit to the baseline, not the peak
Here is the whole game, on one napkin.
Plot Nimbus’s compute usage over a week. There’s a baseline β the floor the usage never drops below, maybe 50 vCPUs’ worth, humming 24⁄7. And there’s a peak β the Friday-night photo-upload surge that briefly triples it. The correct move is to commit to the stable baseline and leave the peak on on-demand (or Spot).
Say the baseline is a steady $12/hour of on-demand compute and the peaks spike to $30/hour for a few hours a day. If you buy a Savings Plan for $12/hour, you discount the part that’s always there β up to ~72% off that slice β and let the spiky top float on flexible pricing. Concretely: $12/hour on-demand is ~$8,760/month; at a ~50% effective Savings Plan discount that’s ~$4,380/month, saving roughly $4,380/month on the baseline alone**, with zero risk, because that usage is never going away.
Chad’s plan β commit to the peak β is the trap. If you buy a $30/hour commitment to cover the spikes, you’re now paying $30/hour all the time, including the 20 hours a day you only use $12. You’ve converted a discount into a penalty. Over-committing locks in unusable spend: you pay for reserved capacity that sits idle, and it’s often worse than the on-demand you were trying to avoid. Commit the floor. Never the ceiling.
The commitments you already have (and are wasting)
Nimbus wasn’t starting from zero. Somebody, at some point, had bought RIs. Priya opened the coverage and utilization reports in Cost Explorer and found the usual archaeology.
Unused and underutilized RIs and Savings Plans
Some RIs had low utilization β the team had migrated the underlying instances to a different family months ago, and the rigid Standard RIs were now attached to nothing, discounting usage that no longer existed. That’s pure waste: you prepaid for a discount and then deleted the thing it applied to. The utilization report shows this as a percentage; anything well under 100% is money evaporating. The coverage report shows the opposite gap β how much of your eligible on-demand usage isn’t covered by any commitment, i.e., where you’re still overpaying.
RIs expiring into coverage gaps
Worse, a batch of 1-year RIs was about to expire. When an RI or Savings Plan lapses and nobody renews it, the covered usage silently falls back to full on-demand β a quiet, automatic price increase that shows up as a nasty surprise on next month’s bill. Priya set calendar alerts 30 days before every expiration. Commitment coverage isn’t “set and forget”; it’s a rolling ladder you have to keep climbing.
Not sharing across the Organization
Nimbus had multiple AWS accounts under one AWS Organization β prod, staging, data, that one account nobody claims. By default, with RI/Savings Plan sharing enabled at the Organization level, a commitment bought in one account applies to matching usage across all linked accounts. Nimbus had sharing off, so an underused RI in the prod account sat idle while the data account paid full on-demand for the exact instance type that RI covered. Flipping on sharing let the discounts flow to wherever the usage actually was β no new purchase required, just a checkbox and a pooled commitment.
How Priya actually rolled it out
She didn’t buy three-year-all-upfront-everything. She did the responsible, slightly boring thing:
- Finish the right-sizing and Graviton migration first (Part 1), so the baseline she committed to was the optimized baseline, not the wasteful one.
- Buy a Compute Savings Plan covering ~80% of that stable baseline, on a 1-year term to start β enough discount to matter, short enough to adapt.
- Turn on Organization-wide sharing so every account drank from the same discounted well.
- Set expiration alerts and a monthly ritual of reading the coverage and utilization reports.
Chad, to his credit, came around. “So we’re still saving a ton,” he said, “we’re just… not signing our own hostage note.”
“Exactly,” said Priya. “Commit to what’s boring and permanent. Stay flexible on everything else.”
Takeaways
- Stop paying all-on-demand for a steady baseline β on-demand is walk-away pricing; you’re overpaying for usage that never walks away.
- Use Savings Plans / Reserved Instances β commit 1 or 3 years for up to ~72% off compute.
- Commit to the stable baseline, not the peak β discount the floor your usage never drops below; over-committing to the peak locks in unusable, idle spend.
- Pick the right commitment type β default to flexible Compute Savings Plans (portable across family/region/service); use EC2 Instance SPs or RIs only for truly stable workloads, and RIs specifically when you need zonal capacity reservation.
- Fix unused/underutilized RIs and Savings Plans β check the utilization report; anything under 100% is prepaid discount applied to nothing.
- Watch coverage gaps β the coverage report shows eligible usage still on full on-demand.
- Track RI/Savings Plan expirations β lapsed commitments silently fall back to on-demand; set alerts and renew on a rolling basis.
- Enable RI/Savings Plan sharing across your AWS Organization β let commitments in one account discount matching usage in all linked accounts.
Part 3 β Redundancy Theater
The disaster recovery drill was Marcus’s idea. We would simulate the loss of an entire AWS region and prove that Nimbus, home to 40 million people’s brunch photos, could shrug off a meteor strike on us-east-1.
Priya sat in the back of the war room, laptop open to Cost Explorer, watching Marcus click through his architecture diagram. Everything was blue. Everything was doubled. Everything, somewhere, had a twin quietly humming in another Availability Zone, another region, another account, drawing breath and dollars in equal measure.
“See,” Marcus said, gesturing at a constellation of boxes, “if us-east-1 goes down, we fail over to us-west-2 seamlessly. Full replication. Multi-AZ on everything. We’re bulletproof.”
“We’re paying for bulletproof glass on the garden shed,” Priya said, not quite under her breath. Chad, mid-sip of cold brew, laughed. Marcus did not. “We might need it at scale,” he said, and Priya wrote the phrase on a sticky note and pressed it to the side of her monitor like a specimen.
The Theater, and What the Audience Actually Paid For
Here is the thing nobody wants to say out loud in the DR meeting: redundancy is not free, and most redundancy at most companies is theater. It looks like resilience. It costs like resilience. But nobody has ever tested whether the understudy can actually go on.
Nimbus was spending an enormous amount on standbys, replicas, and copies that existed because someone once checked a box labeled “highly available” and never checked it again. Let’s walk the set.
Multi-AZ Everywhere, by Default
Multi-AZ is the reflex. You spin up an RDS instance, you tick “Multi-AZ deployment,” and AWS provisions a synchronous standby in a second AZ. Your write latency ticks up slightly, your failover becomes automatic β and your instance cost roughly doubles, because you are now paying for two instances instead of one.
That is a genuinely good trade for your production user database. It is a terrible trade for the seven non-production databases Nimbus had also set to Multi-AZ out of pure muscle memory.
Concrete math: a db.r6g.2xlarge in us-east-1 runs about $0.968/hour on-demand for a single instance β roughly $706/month. Flip on Multi-AZ and you are at about $1,412/month for that one database. Nimbus had this configuration on a dev database that three engineers touched during business hours and that would be gently mourned, then rebuilt from a snapshot, if it ever vanished. That’s about $700/month of standby guarding a database whose real recovery requirement was “recreate it Tuesday.”
The rule Priya wrote down: Multi-AZ is for genuinely critical, stateful production services where an AZ failure would be a real incident. Dev, test, and staging can run single-AZ. If they fall over, you relaunch them. That is what dev is for.
Multi-Region Replication of Things Nobody Would Miss
Marcus’s crown jewel was cross-region everything. And some of it was legitimate: the core user metadata, the billing ledger β yes, replicate those, because losing them is an extinction event.
But Nimbus was also replicating thumbnail caches, transcoded video derivatives, and analytics event logs across regions “for resilience.” These are all regenerable. The thumbnails can be re-rendered from originals. The analytics logs are already in the data warehouse. Paying to keep a hot copy of them 2,000 miles away protects against a disaster that, if it happened, you’d solve by pressing “rebuild.”
The discipline here is to start from your actual DR and RPO requirements β Recovery Point Objective, how much data you can afford to lose β and replicate only what those numbers demand. If your RPO for thumbnail caches is “who cares, regenerate them,” they do not belong in a cross-region replication policy.
Over-Replicated Storage: S3 CRR, Ghost Replicas, Idle Standbys
This is where the money hides, because storage redundancy is quiet.
S3 Cross-Region Replication (CRR): Nimbus had CRR turned on for a 400 TB bucket of user uploads. CRR means you pay for the destination storage again (another ~$0.023/GB-month for S3 Standard), plus inter-region data transfer at roughly $0.02/GB for every object replicated. 400 TB of duplicated storage is about $9,400/month just for the second copy, before a byte of transfer. Some of that bucket genuinely needed geographic redundancy. A large chunk of it was derived data that did not.
Read replicas nobody reads: Nimbus had three RDS read replicas on the PhotoBlob database. Flow logs and Performance Insights showed one of them served zero query traffic for the trailing 90 days. It was created for a reporting feature that shipped, got deprecated, and left its replica behind like a tenant who moved out but kept paying rent. That’s another full instance-cost β call it $700+/month β for a machine answering no questions.
Unused standbys: the DR region had warm standby EC2 fleets that had never once taken production traffic. Warm standby is a valid strategy, but a warm standby you never test is just an expensive standby. If you’re paying to keep it warm, you should be failing over to it on a schedule to prove it works β otherwise downgrade to pilot-light (minimal always-on footprint, scale up on disaster) and pocket the difference.
Duplicate Environments Breeding in the Dark
Priya inventoried the non-prod accounts and found: a staging stack, a qa stack, a staging-2 stack from a migration that finished in 2024, a demo environment for a sales team that now used a SaaS demo tool, and something called perf-test-DO-NOT-DELETE that no one could explain and everyone was afraid of.
Five overlapping environments, each with its own load balancers, databases, and half-scaled compute. Most were 80% idle. This is redundancy nobody even chose β it accreted. The fix is unglamorous: consolidate. One well-maintained staging environment that mirrors production is worth more than five stale ones, and it costs a fraction. Nimbus collapsed the five into two (staging + a shared ephemeral QA that spins up per-branch and tears down nightly).
Backups: Over-Retained and Over-Frequent
Finally, backups β where good intentions become a storage line item that only grows.
Nimbus was taking hourly automated snapshots of PhotoBlob and retaining them for two years. Snapshots are incremental, but two years of hourly snapshots across a large database is a genuinely large pile of EBS snapshot storage (~$0.05/GB-month), and it was climbing every hour of every day.
Nobody had ever restored a snapshot older than about a week. The real recovery need was “roll back a bad deploy” (hours) and “satisfy the 90-day compliance window” (quarters, not the whole database, and not hourly granularity). Priya matched retention to actual recovery needs: hourly snapshots kept for 7 days, daily for 30, monthly for a year for compliance. Same protection for every recovery scenario that had ever actually occurred β a fraction of the storage.
How the Drill Ended
They ran Marcus’s meteor-strike drill. Failover to us-west-2 worked for the core services β the ones that genuinely warranted the redundancy. And the exercise revealed, gently and expensively, that two-thirds of the doubled infrastructure was protecting things that either regenerated themselves or nobody would notice were gone.
Marcus kept his cross-region replication on the user ledger and the billing data, where it belonged. Priya turned off CRR on the thumbnail bucket, killed the ghost read replica, consolidated the environment zoo, and put the backups on a diet. The sticky note β “we might need it at scale” β stayed on her monitor, but she added one underneath it: “at which scale, for which failure, tested when?”
The redundancy that survived was the redundancy that answered those three questions. Everything else was theater, and the audience was the CFO.
Takeaways
- Don’t default everything to Multi-AZ. Multi-AZ roughly doubles instance cost β reserve it for genuinely critical stateful production services; run dev, test, and staging single-AZ and rebuild on failure.
- Replicate multi-region only what DR/RPO truly requires. Regenerable data (thumbnails, transcodes, analytics logs) doesn’t need cross-region copies; drive replication from real Recovery Point Objectives, not reflex.
- Hunt over-replicated storage. Audit S3 Cross-Region Replication (you pay for the second copy and inter-region transfer), delete read replicas serving zero traffic, and downgrade untested warm standbys to pilot-light.
- Consolidate duplicate environments. Overlapping staging/QA/demo stacks accrete and sit idle; collapse them into one maintained staging environment plus ephemeral per-branch QA.
- Match backup retention and frequency to real recovery needs. Hourly snapshots retained for years cost storage every hour for restores that never happen β tier retention (short-term hourly, longer-term daily/monthly) to the recovery scenarios that actually occur.
Part 4 β The Case of the Mysterious NAT Gateway
The line item just said “EC2-Other.” $31,247 of it, per month, and nobody could say what it was.
Jenny found it first, because Jenny reads Cost Explorer the way other people read the news. She wandered over to Priya’s desk holding a laptop like it was a small animal that had bitten her. “There’s a category called EC2-Other and it’s our fourth-biggest cost and I have no idea what’s in it.”
Priya knew, the way you know the sound of a specific creaking stair in a house you grew up in. “That’s the network,” she said. “That’s where the invisible bill lives.” And somewhere in that $31,247, purring contentedly, was the thing the team had named years ago in a postmortem and never stopped fearing: the NAT gateway from hell.
The Invisible Bill
Compute and storage have honest bills. An instance costs what it costs; a bucket bills by the gigabyte. Networking is different. Networking charges you for movement β bytes crossing invisible boundaries you didn’t know you’d drawn β and the charges only show up as cryptic usage types buried in “EC2-Other.” Nobody provisions a data-transfer charge. It just happens, quietly, a fraction of a cent at a time, forty million times an hour.
Let’s turn on the lights.
The NAT Gateway From Hell
A NAT Gateway lets resources in a private subnet reach the internet (to pull packages, call third-party APIs, hit public endpoints) without being publicly addressable themselves. Useful. But it bills two ways at once:
- A per-hour charge, ~$0.045/hour β about $32/month just for existing, per gateway, per AZ.
- A per-GB processing charge, ~$0.045/GB, on everything that passes through it.
That second number is the killer. Nimbus routed a firehose of traffic through NAT gateways β including, it turned out, traffic to S3. Every thumbnail write, every upload to the PhotoBlob bucket, was leaving the private subnet, passing through the NAT gateway (at $0.045/GB), and then hitting S3. At Nimbus’s volume β call it 300 TB/month through NAT β that’s roughly $13,500/month in NAT processing alone. For traffic to a service AWS will let you reach for free.
The Free Win Everyone Misses: Gateway VPC Endpoints
Here is the single highest-leverage fix in this entire series.
S3 and DynamoDB have Gateway VPC Endpoints, and they are free. No hourly charge, no per-GB charge. You add a route in your route table, and traffic to S3/DynamoDB goes over the AWS private network instead of out through the NAT gateway. You stop paying $0.045/GB. You stop paying for the internet round-trip you never needed.
Nimbus’s 300 TB/month of S3 traffic through NAT dropped to $0 in processing the moment a Gateway Endpoint was added. That one route-table change was the biggest single win Priya found all quarter. Deepak, who’d built the original VPC, went a bit quiet. “It was one checkbox,” he said. “It was always one checkbox.”
Interface Endpoints and PrivateLink for Everything Else
S3 and DynamoDB get the free Gateway endpoints. Other AWS services β ECR, Secrets Manager, SQS, Kinesis, CloudWatch, SSM, and friends β use Interface VPC Endpoints (PrivateLink) instead. These keep traffic to those services off the NAT gateway and on the private network.
The catch: Interface endpoints aren’t free. Each bills $0.01/hour per AZ ($7.20/month per AZ) plus ~$0.01/GB processed. Still far cheaper than NAT’s $0.045/GB for high-volume services β Nimbus’s container fleet pulling images from ECR all day paid down its NAT bill fast by adding an ECR interface endpoint.
But there’s a trap on the other side, which brings us to sprawl.
Interface Endpoint Sprawl
After the ECR win, someone got enthusiastic and started adding interface endpoints for every service, in every AZ, across every VPC. Because each is billed per-hour, per-AZ, this adds up. Fifteen interface endpoints across three AZs is 45 Γ $7.20 = ~$324/month in hourly charges before any traffic β and several were for services with trivial traffic that would’ve been cheaper left on NAT.
Interface endpoints are worth it for high-volume service traffic. For a service you call twice a day, the endpoint’s hourly cost exceeds the NAT processing you’d save. Right-size them; don’t collect them.
Cross-AZ Traffic: The Tax You Pay for Sloppy Placement
Traffic between Availability Zones costs ~$0.01/GB in each direction β so ~$0.02/GB round-trip. This is small per gigabyte and enormous in aggregate when your services are chatty and scattered carelessly across AZs.
Nimbus’s PhotoBlob monolith talked constantly to a caching tier that had been placed, for no reason anyone remembered, in a different AZ. Every cache read hopped an AZ boundary and back. VPC Flow Logs showed 80 TB/month of this cross-AZ chatter β roughly $1,600/month to move data between two machines that could have been AZ-aligned. Inter-service chattiness is fine within an AZ (free) and expensive across one. Co-locate the talkers.
Public Subnets, Wrong Paths, and Forced NAT
Several of the mysterious charges traced back to subnet design. Resources that should have been in private subnets with private routing were sitting in public subnets, or were forcing traffic out through NAT because no private path existed. A misplaced resource in the wrong subnet turns a free private hop into a NAT-processed, cross-AZ, sometimes-internet-egressing mess.
The broader principle: the data access path matters as much as the data. Reaching another AWS service can go over the public internet route (you pay egress), through a private endpoint/PrivateLink (cheap or free), or over Direct Connect (for on-prem). Same bytes, wildly different bills, decided entirely by which route table entry wins.
Internet Egress: The Priciest Bytes
Data leaving AWS for the public internet is the most expensive transfer of all β roughly $0.09/GB for the first tier out of us-east-1, tapering with volume. Nimbus serves photos to 40 million users, so egress is unavoidable β but how you serve it isn’t.
Serving images straight from S3 or EC2 to users pays full egress every time. Putting CloudFront in front changes the economics: CloudFront’s egress rates are lower, transfer from S3/EC2 to CloudFront is free, and caching means you serve most requests without touching the origin at all. Moving photo delivery behind CloudFront cut Nimbus’s origin egress dramatically β the CDN both lowered the per-GB rate and slashed how many GB left the origin in the first place.
Cross-Region and Sustained High-Volume Transfer
Cross-region traffic (~$0.02/GB between regions) had crept in via the replication we trimmed back in Part 3 β another reason to replicate only what DR truly requires.
And for sustained, high-volume transfer β think a steady pipe to an on-prem data center or partner β the internet route is the wrong tool. Direct Connect provides a dedicated private connection with lower per-GB rates for consistent heavy traffic; a VPN over the internet is cheaper to stand up but still pays internet egress and gives less predictable throughput. For a Nimbus analytics feed shipping terabytes to a partner monthly, Direct Connect beat both VPN and raw internet egress once volume stayed high month over month.
Idle Load Balancers and Unattached Elastic IPs
Finally, the loose change that bills by the hour whether you use it or not:
- Idle load balancers. Nimbus had four ALBs from decommissioned services still running. An ALB costs $0.0225/hour ($16/month) plus LCU charges β even with zero traffic, the hourly clock runs. Four dead ALBs β $65/month for nothing.
- Unattached Elastic IPs. An EIP is free while attached to a running instance. The moment it’s unattached (or attached to a stopped instance), AWS bills it $0.005/hour ($3.60/month) β deliberately, to discourage IP hoarding. Nimbus had 11 orphaned EIPs. Small money, but it’s money for literally nothing.
Cracking the Case
The tools that solved it: Cost Explorer with usage-type filters (filter on NatGateway-Bytes, DataTransfer-Regional-Bytes, DataTransfer-Out-Bytes to see exactly which movement is costing what) and VPC Flow Logs (to see which services are talking, across which boundaries, in what volume). Between them, “EC2-Other” stopped being a mystery and became a to-do list.
Jenny presented the findings. Chad, who bragged that Nimbus spent “the GDP of a small island” on AWS, was delighted to learn that a meaningful chunk of that GDP had been spent moving photos in a circle through a gateway they didn’t need. The NAT gateway from hell still runs β every VPC needs some egress β but it now handles only the traffic that genuinely has to reach the open internet. Everything else found a cheaper, quieter road.
Takeaways
- Add S3 and DynamoDB Gateway VPC Endpoints today β they’re free and stop you paying NAT’s ~$0.045/GB to reach services AWS lets you reach at no cost. This is usually the single biggest network win.
- Know that NAT Gateways bill twice: ~$0.045/hour and ~$0.045/GB processed. Route high-volume AWS traffic off NAT and onto endpoints.
- Use Interface/PrivateLink endpoints for other AWS services (ECR, Secrets Manager, SQS, etc.) when traffic is high-volume β but they cost ~$0.01/hr per AZ + ~$0.01/GB, so avoid interface-endpoint sprawl; skip them for low-traffic services where NAT is cheaper.
- Cut cross-AZ traffic (~$0.01/GB each way) by co-locating chatty services in the same AZ; use VPC Flow Logs to find inter-service chatter crossing AZ boundaries.
- Fix subnet design and data access paths: don’t leave resources in public subnets or force NAT where a private endpoint/PrivateLink route would work; pick the right path (internet vs private endpoint vs Direct Connect) deliberately.
- Reduce internet egress β the priciest transfer (~$0.09/GB) β by fronting user-facing delivery with CloudFront (free S3/EC2-to-CloudFront transfer, lower egress rates, caching cuts origin volume).
- Watch cross-region traffic (~$0.02/GB) and replicate only what DR requires.
- For sustained high-volume transfer, choose Direct Connect over VPN or internet egress β dedicated private throughput at lower per-GB rates beats paying repeated internet egress.
- Delete idle load balancers and release unattached Elastic IPs β both bill hourly (~$16/mo per idle ALB, ~$3.60/mo per orphaned EIP) whether used or not.
- Investigate “EC2-Other” with Cost Explorer usage-type filters and VPC Flow Logs to turn the invisible network bill into a concrete list of fixes.
Part 5 β The Bucket That Ate the Budget
Priya found the bucket at 11:40 on a Tuesday, which is the hour of day when nothing good is ever discovered.
It was called nimbus-photoblob-prod, and according to S3 Storage Lens it held 4.1 petabytes. She read the number three times. Then she opened Cost Explorer, filtered to S3, and watched a single line item bloom across the chart like a bruise: $71K a month. For one service. For a company that, as Chad enjoyed reminding the all-hands, “spends the GDP of a small island” on AWS.
“That can’t be right,” said Jenny, leaning over. “Forty million users, sure, but the photos aren’t that big.”
“They’re not,” said Priya. “The photos are fine. It’s everything around the photos.” She scrolled. Thumbnails. Re-encodes. A folder named temp/ last touched in 2024. And, beautifully, a top-level prefix called bertha-migration-backup/ that nobody could explain and everybody was afraid to delete. Big Bertha, reaching out from the grave to bill us.
The photos are fine. The plumbing is a crime scene.
S3 pricing rewards you for putting cold data in cold places and punishes you for hoarding hot copies of things nobody looks at. Nimbus was doing the opposite with enormous discipline.
Lifecycle policies, or the absence thereof
Every object in nimbus-photoblob-prod was sitting in S3 Standard, which in us-east-1 runs about $0.023 per GB-month. That’s the right tier for a photo uploaded ten minutes ago that the user is actively sharing. It is an insane tier for a photo from 2019 that will be opened, on average, never.
S3 offers a temperature ladder, and lifecycle policies let objects walk down it automatically based on age:
- Standard: ~ $0.023/GB-mo β hot, frequently accessed.
- Standard-Infrequent Access (IA): ~ $0.0125/GB-mo β cheaper storage, but you pay a per-GB retrieval fee.
- Glacier Flexible Retrieval: ~ $0.0036/GB-mo β minutes-to-hours to restore.
- Glacier Deep Archive: ~ $0.00099/GB-mo β cheapest on the menu, 12-hour restores.
Do the arithmetic on one petabyte of genuinely cold photos. In Standard that’s roughly $23,000/month. The same petabyte in Deep Archive is about $990/month. Nimbus had at least two petabytes that hadn’t been touched in over a year. A lifecycle rule β Standard for 30 days, IA at 30 days, Glacier at 90, Deep Archive at 365 β would have moved them without a human ever opening a ticket.
Priya wrote the rule that afternoon. Deepak asked whether restores would “make profile loads slow.” Priya explained, patiently, that you only transition data users aren’t loading, which is the entire point.
Intelligent-Tiering for the stuff you can’t predict
Lifecycle rules assume you know an object’s access pattern by its age. For a lot of Nimbus data β shared albums that go viral for a week two years after upload β you don’t. That’s what S3 Intelligent-Tiering is for: S3 monitors access per object and moves it between frequent and infrequent tiers automatically, for a small monitoring fee (about $0.0025 per 1,000 objects/month). No retrieval charges when data warms back up.
The catch is the monitoring fee: it’s death by a thousand cuts if your objects are tiny. Which brings us to Nimbus’s second crime.
A billion tiny objects
PhotoBlob stored every image as a family: original, plus thumbnails at six sizes, plus a JSON sidecar of EXIF data. Seven-plus objects per photo, most of them a few kilobytes. Across the fleet, roughly 1.5 billion objects.
Storage of tiny objects is cheap. Requests are not. S3 charges per request β about $0.0004 per 1,000 GET and $0.005 per 1,000 PUT in Standard. When your app renders a feed by issuing millions of GETs against millions of postage-stamp thumbnails, the request line becomes real money, and IA/Glacier tiers add per-object minimums and retrieval fees that make tiny cold objects worse, not better. The fix is architectural β sprite the thumbnails, cache aggressively at CloudFront, and stop treating S3 like a filesystem β but even naming the problem via Storage Lens’s object-count and request metrics was new information to the team.
The uploads that never finished
Then Priya found the quiet one. Multipart uploads let you send a big file in chunks. If a chunk fails and the client never calls Complete or Abort, those uploaded parts sit in the bucket billing you at full Standard rate, invisible in the normal object listing.
Nimbus’s mobile clients dropped connections constantly β subways, elevators, the void. Storage Lens reported incomplete multipart uploads totaling ~180 TB. That’s roughly $4,100/month for uploads that, by definition, no user ever completed.
The fix is one lifecycle rule: AbortIncompleteMultipartUpload after 7 days. It should be on literally every bucket you own. Priya added it everywhere and felt, briefly, like a person who had accomplished something.
Versioning without a curfew
nimbus-photoblob-prod had versioning enabled β sensible, protects against accidental deletes. But versioning without a non-current version expiration rule means every overwrite keeps the old copy forever, and every “delete” just drops a delete marker on top of a stack of retained versions you’re still paying to store.
Storage Lens showed non-current versions accounting for nearly 300 TB. Nobody could restore a two-year-old thumbnail version and nobody would want to. A lifecycle rule to expire non-current versions after 30 days and clean up expired delete markers reclaimed the lot β another ~$7,000/month of storage nobody knew existed.
Logs, temp, dupes, and abandoned buckets
The rest was housekeeping, but housekeeping at petabyte scale is money:
- Logs and temp data in Standard. ALB logs, CloudFront logs, and that
temp/folder were all sitting in Standard. Logs are write-once, read-during-an-incident. A lifecycle rule to IA at 30 days and Deep Archive (or outright expiration) at 90 fixed it. - Duplicate and stale data. The
bertha-migration-backup/prefix was a full copy of a migration that completed in 2024. 410 TB, pure duplicate. Priya verified the source was intact, then expired it. - Abandoned buckets. Storage Lens’s account-wide view surfaced eleven buckets from dead experiments, one still named
marcus-test-DO-NOT-DELETE(created by Marcus, who had left three teams ago). Together: a few more terabytes and a governance headache.
The block storage nobody looked at
S3 gets the headlines, but Priya ran the same audit on block storage:
- Over-provisioned EBS. Dozens of
gp2/gp3volumes provisioned at 500 GB and sitting 8% full, plus a graveyard of unattached volumes still billing at ~$0.08/GB-month. Right-sizing and deleting orphans clawed back a few thousand a month. - EFS in the wrong tier. A shared EFS filesystem held CI artifacts, all in EFS Standard (~$0.30/GB-mo). Enabling EFS Infrequent Access with a lifecycle policy dropped rarely-touched files to about $0.016/GB-mo β roughly a 90% cut on that slice.
The number at the bottom
By Friday, Priya’s spreadsheet said the storage changes β lifecycle transitions, aborted multiparts, expired versions, deleted duplicates, cold-tiered logs, right-sized EBS/EFS β would take the S3-and-friends line from ~$71K to roughly $28K a month. She didn’t delete a single photo a user cared about.
Chad, told the number, nodded gravely and said, “That’s the GDP of a smaller island.” Priya took the win.
Takeaways
- Add lifecycle transition rules to walk aging objects Standard β Infrequent Access β Glacier β Deep Archive; one cold petabyte drops from ~$23K/mo to ~$990/mo.
- Enable S3 Intelligent-Tiering for data with unknown or spiky access patterns so S3 moves objects between tiers automatically β but watch the per-object monitoring fee on tiny files.
- Add an AbortIncompleteMultipartUpload lifecycle rule (e.g., 7 days) to every bucket to stop paying for uploads that never completed.
- Expire non-current object versions and clean up delete markers β versioning without an expiration policy keeps every overwrite forever.
- Attack request (GET/PUT) costs from many tiny objects by spriting/consolidating, caching at CloudFront, and rethinking one-object-per-thumbnail designs.
- Move logs and temp data out of Standard into IA/Glacier or expire them outright; they’re write-once, read-rarely.
- Hunt duplicate/stale data and abandoned buckets with account-wide tooling and delete verified copies of completed migrations.
- Right-size EBS, delete orphaned volumes, and enable the EFS Infrequent Access tier to stop over-paying for block and file storage.
- Use S3 Storage Lens as the single pane of glass for object counts, incomplete multiparts, non-current versions, and per-bucket cost drivers.
Caution!
Such tiering often helps reducing the storage cost. But not always! When the files are very small (few KB), intelligent tiering can actually increase the cost - because S3 computes the storage cost on bulk units. So read the documents before you jump in and start following the popular trend.
Part 6 β Full Scan of Shame
The Slack message from Deepak was seven words long: “Profile loads are slow again, not my code.”
Priya opened Performance Insights for the primary RDS instance, sorted by load, and there it was β a single query eating 60% of database time, running thousands of times a minute: SELECT * FROM photos WHERE user_id = ?, with no index on user_id. Every profile load in PhotoBlob triggered a full table scan across 900 million rows. Deepak’s code. Deepak’s very own code, scanning the entire photos table to render one avatar.
“It’s a full scan of shame,” said Jenny, who had recently learned what a full table scan was and was enjoying it enormously.
Marcus wandered over. “We might need the flexibility of not indexing,” he offered, “at scale.” Priya wrote an index. It took ninety seconds. The query went from 4 seconds to 3 milliseconds, and the database’s CPU graph fell off a cliff. Nobody needed the flexibility.
That was the appetizer. The database bill was $96,000/month, and Priya had opinions about all of it.
The relational layer
Over-provisioned RDS and Aurora
The primary was a db.r5.8xlarge (32 vCPU, 256 GB RAM) running at 11% average CPU. It had been sized during the Big Bertha panic of 2025 and never revisited. Right-sizing to a db.r6g.2xlarge β a Graviton instance, which delivers roughly 20% better price-performance than the equivalent x86 β cut the instance cost by more than half. On-demand, an r5.8xlarge runs about $4.00/hour ($2,900/mo); an r6g.2xlarge is about $0.80/hour ($580/mo). Same workload, once the index existed.
Storage was on old-style gp2. Moving to gp3 decoupled IOPS from capacity and let Priya provision baseline 3,000 IOPS without over-buying storage, saving another ~20% on the storage line.
Idle databases and Aurora Serverless v2
There were fourteen RDS instances. Two were the real workload. The other twelve were per-team “staging” and “analytics” databases averaging under 5% CPU, several idle overnight and all weekend. Priya consolidated the stragglers onto shared instances, and for the genuinely bursty internal tools she moved them to Aurora Serverless v2, which scales capacity in fine-grained ACUs and can idle down when nobody’s querying β so the finance team’s once-a-day dashboard stops paying for a full instance 24⁄7.
Missing indexes, N+1, and the writer taking all the reads
The full-scan avatar query was the headline, but the slow-query log and Performance Insights surfaced two more patterns:
- N+1 queries. The feed renderer fetched a list of 50 photos, then issued 50 separate queries for each photo’s owner. Fifty-one round trips where a single JOIN or
WHERE id IN (...)would do. At feed-load volume this was millions of pointless queries an hour. Deepak grumbled, then batched them. - Reads hitting the writer. Nimbus had two read replicas. The application connected all traffic β reads and writes β to the primary endpoint anyway, because someone had hardcoded it in 2023. Pointing read queries at the reader endpoint offloaded ~70% of query volume from the writer, which is what made the right-sizing safe.
No connection pooling
Every application container opened its own pool, and at peak the database hit its max_connections ceiling and started rejecting connections β which the team had “fixed” by scaling up the instance for more RAM. The actual fix was RDS Proxy, which multiplexes thousands of application connections onto a small pool of database connections, smooths failover, and let Priya keep the smaller instance without connection storms.
DynamoDB: the two classic mistakes
Nimbus’s session and activity-feed data lived in DynamoDB, and it was making both textbook errors.
- Provisioned vs. on-demand mismatch. The sessions table was provisioned at 40,000 RCU/WCU to survive a traffic spike that happened twice a year, and sat at 8% utilization the rest of the time. For that spiky, unpredictable shape, on-demand mode (pay per request) was dramatically cheaper. Conversely, a steady high-throughput table should stay provisioned (ideally with a reserved capacity commitment) β the point is matching the mode to the traffic, and Nimbus had it backwards on both.
- Scan instead of Query. The activity feed was built with
Scanand a filter expression, reading the entire table and throwing away 99.9% of it before returning results. DynamoDB bills for data read, not data returned, so a filtered Scan pays for everything. The redesign added a GSI keyed onuser_id+ timestamp so the feed became aQuerytouching only that user’s items. Deepak’s PhotoBlob habit of scanning everything had a cousin over in NoSQL.
Backups, snapshots, and unused replicas
- One of the two read replicas served zero traffic (leftover from a migration) β deleted.
- Automated backup retention was set to 35 days on every instance, plus a cron job taking manual snapshots nightly and never deleting them. Hundreds of orphaned snapshots billed at ~$0.095/GB-month. Priya dialed retention to a sane 7 days for staging, kept 35 only where compliance required it, and expired the manual snapshot graveyard.
The analytics and big-data layer
The data platform was a second, quieter budget fire.
Athena scanning too much
Athena charges ~$5.00 per TB scanned. The analytics team’s dashboards queried raw JSON event logs with SELECT *, unpartitioned, and each dashboard refresh scanned ~2 TB β about $10 a click, thousands of times a day. Priya’s fixes were the canonical Athena playbook: partition the data by date, convert from JSON to Parquet (columnar, compressed), and select only the needed columns. The same dashboard query dropped from 2 TB scanned to under 40 GB β roughly $0.20 instead of $10.
Redshift running idle
A ra3.4xlarge Redshift cluster ran 24⁄7 for reports generated between 6 and 9 a.m. Options Priya weighed: pause the cluster outside business hours (RA3 supports pause/resume), or move to Redshift Serverless so it bills only while queries run. For a 3-hours-a-day workload, serverless cut the line item by ~80%.
EMR clusters left running
Two EMR clusters had been spun up for one-off backfills months ago and never terminated β a persistent bill for zero jobs. The pattern for batch analytics is transient clusters that launch, run the job, and self-destruct: enable auto-termination, and use Spot instances for task nodes (60β90% off on-demand) since task nodes are stateless and interruption-tolerant.
Glue over-provisioned
The nightly ETL Glue job was configured with 20 DPUs and finished in 6 minutes using a fraction of them. Right-sizing to the DPUs the job actually needed (and using auto-scaling where available) cut its cost proportionally β you pay per DPU-hour, so over-provisioning is pure waste.
Kinesis over-provisioned shards
The ingestion stream was provisioned at 50 shards sized for a launch that never fully materialized, billing per shard-hour regardless of throughput. Switching Kinesis Data Streams to on-demand mode let it scale with actual traffic instead of paying for permanent headroom.
OpenSearch oversized
The log-search OpenSearch domain ran a fat hot-tier cluster holding 18 months of logs nobody searched past week two. Priya moved older indices to UltraWarm (and cold storage beyond that), set an index lifecycle policy to delete stale indices after 90 days, and right-sized the hot nodes. Hot storage is expensive; UltraWarm is a fraction of it for data you rarely query.
The number at the bottom
Right-sizing, Graviton, gp3, killed idle instances, RDS Proxy, reader-endpoint routing, DynamoDB mode fixes and the Query redesign, plus the analytics cleanup, took the data layer from $96K to about $41K a month. The single largest contributor to the performance win was a ninety-second index on user_id.
Deepak, informed that his profile-load query was the most expensive line in the database, was quiet for a moment. Then he said, “So it was my code.” Personal growth.
Takeaways
- Right-size RDS/Aurora and move to Graviton DB instances and gp3 storage β an idle
r5.8xlargeat 11% CPU became anr6g.2xlargeat roughly half the cost. - Consolidate idle/underused databases and use Aurora Serverless v2 for bursty internal workloads so capacity idles down.
- Kill full table scans by adding the right indexes β use slow-query logs and Performance Insights to find them (Deepak’s avatar query went 4s β 3ms).
- Eliminate N+1 query patterns by batching with JOINs or
INlists instead of one query per row. - Add connection pooling with RDS Proxy instead of scaling the instance up to survive connection storms.
- Route read traffic to read replicas via the reader endpoint instead of hammering the writer.
- Match DynamoDB capacity mode to traffic β on-demand for spiky/unpredictable tables, provisioned (or reserved) for steady high throughput.
- Replace DynamoDB Scan with Query by redesigning keys/GSIs, since you pay for data scanned, not returned.
- Delete unused read replicas and trim over-retained backups/snapshots to sane retention windows.
- Cut Athena cost by partitioning, converting to Parquet/ORC, compressing, and selecting only needed columns β you’re charged ~$5/TB scanned.
- Pause idle Redshift or switch to Redshift Serverless / RA3 for workloads that only run a few hours a day.
- Use transient EMR clusters with auto-termination and Spot task nodes instead of leaving clusters running.
- Right-size Glue jobs to the DPUs they actually use.
- Switch over-provisioned Kinesis to on-demand shard mode.
- Right-size OpenSearch with UltraWarm/cold storage and lifecycle policies to delete stale indices.
Part 7 β Serverless, Not Costless
Priya found the recursive Lambda the way you find a gas leak: by noticing something was very wrong before you understood what.
The CloudWatch dashboard showed a function named thumbnail-fanout invoking 14 million times an hour. Nimbus has 40 million users. Nobody uploads 14 million photos an hour at 2 a.m. on a Tuesday. Priya traced it: the function wrote a resized image back to the same S3 prefix that triggered it. Every write fired a new event. Every event resized the image. Every resize wrote another object. Big Bertha, the ingest bucket, was cheerfully feeding a function that ate itself.
“It’s serverless,” Chad had told the CFO in March. “It scales to zero. It’s basically free.” Chad was wearing the Patagonia vest. The bill for Lambda alone that month was $31K, and roughly $9K of it was one function calling itself in a circle like a dog after its tail.
Serverless is wonderful. Serverless is not free. Let’s go through the ways Nimbus paid for the difference.
Memory is CPU, and CPU is money
The first thing Priya fixed was the most common Lambda mistake in the world: someone set memory to 128 MB “to save money.”
Lambda memory is not just memory. Memory is the dial that also sets vCPU β proportionally, up to roughly 6 vCPUs at 10,240 MB. At 128 MB you get a sliver of a core. Nimbus’s image-resize function at 128 MB took 8,000 ms to process a photo. Bumped to 1,024 MB, it took 900 ms.
You bill for GB-seconds. Do the math:
- 128 MB Γ 8.0 s = 0.125 GB Γ 8.0 = 1.0 GB-s
- 1,024 MB Γ 0.9 s = 1.0 GB Γ 0.9 = 0.9 GB-s
The “cheap” 128 MB setting was more expensive per invocation and nearly nine times slower. Too low starves the CPU and stretches billed duration. Too high β say pinning everything to 10 GB β burns money the workload never uses. The sweet spot is empirical, not vibes.
Priya ran AWS Lambda Power Tuning (the open-source Step Functions state machine) across each function. It sweeps memory settings, plots cost against speed, and hands you the optimum. Half of Nimbus’s functions landed between 512 MB and 1,536 MB. The resize function’s true optimum was 1,769 MB β a number no human would have guessed.
Still on x86 in 2026
Every Nimbus function ran on x86. Switching the runtime architecture to arm64 (Graviton2) is a one-line config change for most Python and Node functions, and Graviton Lambda is about 20% cheaper per GB-second β often faster too. On a $31K Lambda spend, that alone was roughly $6K/month for a checkbox. Deepak grumbled that PhotoBlob “might have a native dependency.” It had one. It had an ARM wheel. It shipped.
Chatty functions and the self-eating loop
The recursive thumbnail-fanout got a guard rail: resized output now writes to a different bucket than the trigger, plus an object-tag check so a reprocessed image can’t re-fire. Recursion loops are the serverless equivalent of leaving the tap running β AWS now has a recursion-detection kill switch, but relying on it is like relying on the smoke alarm to cook dinner.
The subtler waste was chattiness. A downstream notify function fired once per photo β 14 events, 14 cold-ish invocations, 14 sets of overhead. Priya put SQS in front of it and set a batch size of 100. Same work, roughly 1/100th the invocations. For true firehoses (the activity feed), the same idea with Kinesis batched records into fat, efficient calls. You pay per invocation and per GB-second; amortizing overhead across a batch cuts both.
Paying to wait
One function, await-encoding, called an external video service and then sat there polling for up to 45 seconds until the job finished. Forty-five seconds of billed Lambda duration doing nothing but sleeping on I/O. At 1,536 MB that’s about 0.066 GB-s per second of nap β pure waste, times millions of jobs.
The fix is a Step Functions state machine: kick off the encode, then wait using a .waitForTaskToken callback or a poll-and-wait loop between short Lambda invocations. Step Functions bills for state transitions, not for the wall-clock time you spend waiting. You stop renting compute to watch a progress bar.
When serverless is the wrong tool
Here is the heresy Marcus needed to hear. Lambda’s per-invocation model is a rounding error at spiky or low load and a liability at constant high load.
Nimbus’s feed-render function ran flat-out, 24⁄7, at steady throughput β it never scaled to zero because traffic never dropped to zero. Model the crossover:
- Lambda: ~40 billion GB-s/month at that sustained rate penciled out near $6,700/month, and that’s before request charges.
- Fargate/EC2 (Graviton): a right-sized, always-warm container fleet handling the same steady load came in around $2,400/month.
At sustained, predictable, high throughput, containers or EC2 win β often by 2β3x β because you stop paying the per-invocation premium for elasticity you aren’t using. The rule: model the crossover. Bursty and unpredictable β Lambda. Flat and constant β containers. Nimbus moved the feed renderer to Fargate and left the spiky upload path on Lambda.
Provisioned concurrency you’re not using
To kill cold starts on the checkout function, someone had set provisioned concurrency to 200. Actual peak concurrency was 35. Provisioned concurrency bills for reserved capacity whether or not it’s invoked β it’s a standing charge. Nimbus was paying to keep 165 execution environments warm for nobody. Priya dropped it to 50 with an application-auto-scaling schedule that ramps it up before the evening peak and back down overnight. Cold starts stayed rare; the standing bill dropped by two-thirds.
Fat packages, slow starts
The cold starts that remained were made worse by deployment package bloat. The upload function shipped a 240 MB zip that dragged in the entire AWS SDK, pandas (used by nobody), and three copies of an image library. Bigger packages mean slower cold-start initialization β you pay in latency and in init duration. Priya trimmed to the specific SDK v3 clients actually imported, moved shared code to a Lambda layer, and got the package under 20 MB. Cold starts dropped from ~2.1 s to ~400 ms.
The front door: REST vs HTTP vs URL
Every API at Nimbus went through API Gateway REST APIs at $3.50 per million requests. Most of those routes were dumb proxies that used none of REST’s fancy features (request validation, API keys, WAF wiring).
- HTTP API: about $1.00 per million β roughly 70% cheaper β and enough for straightforward proxy routes.
- ALB β Lambda: cheaper still at high volume, priced on LCUs rather than per-request.
- Lambda Function URLs: free front door for single-function internal endpoints.
Nimbus moved the high-volume public routes to HTTP API and internal webhooks to Function URLs. On a few hundred million requests a month, the delta was real four-figure money.
Standard vs Express Step Functions
Finally, the workflow bill. Step Functions Standard charges per state transition β great for long, durable, human-in-the-loop flows. Nimbus’s image pipeline was neither long nor durable: it ran in under a second, millions of times a day, and used Standard. At per-transition pricing that’s brutal.
Express workflows bill on duration and memory instead, and for high-volume, short-lived orchestration they’re often 10x+ cheaper. Priya converted the sub-second pipelines to Express and kept Standard only for the multi-day moderation-appeal flow that genuinely needs durability and full execution history.
Chad still tells the CFO serverless “scales to zero.” It does. The bill did too β from $31K to about $11K β once someone stopped it from eating itself.
Takeaways
- Right-size Lambda memory: memory also sets CPU, so too low means slow and longer billed duration while too high wastes money β tune empirically with AWS Lambda Power Tuning.
- Switch Lambda functions to arm64/Graviton2 for roughly 20% lower cost per GB-second, often with better performance.
- Kill over-invocation: batch events through SQS or Kinesis and add guard rails so a function can never recursively trigger itself.
- Don’t pay Lambda to wait on I/O β move long idle waits into Step Functions, which bills for transitions, not wall-clock sleep.
- For constant high throughput, model the Lambda-vs-container crossover; at sustained steady load, Fargate/EC2 (especially Graviton) often beats per-invocation Lambda by 2β3x.
- Match provisioned concurrency to real peak concurrency and schedule it β reserved warm capacity bills whether invoked or not.
- Slim deployment packages (specific SDK clients, layers for shared code) to cut cold-start latency and init duration.
- Pick the cheapest adequate front door: HTTP API and Lambda Function URLs undercut API Gateway REST; ALB/Lambda can win at high volume.
- Use Step Functions Express, not Standard, for high-volume short workflows β often 10x+ cheaper β and reserve Standard for long, durable orchestrations.
Part 8 β Cache Rules Everything Around Me
Jenny found it in Cost Explorer, which is where Jenny finds everything now. She had filtered by usage type, sorted by cost, and gone quiet.
“Priya,” she said. “We generate the same thumbnail forty thousand times a day. The same one. The cat picture on the trending page. Every single request re-reads the original from S3, resizes it, and throws it away.”
Priya looked. Jenny was right. The trending grid on the Nimbus home screen was rendering thumbnails on the fly, per request, from the original full-resolution masters in Big Bertha. Forty million users, one trending cat, and a Lambda dutifully re-resizing the identical 12-megapixel image tens of thousands of times because nobody had ever told the system to remember the answer.
Deepak’s take, as always, was that PhotoBlob “was designed for correctness, not caching.” Correctness is wonderful. Correctness that recomputes the same JPEG forty thousand times is just an expensive kind of amnesia.
The month’s data-transfer and compute bill for serving photos was the second-largest line item at Nimbus, right behind the NAT gateway from hell. Almost all of it was work the system had already done and immediately forgotten. This is the chapter about memory β the good kind.
No caching layer at all
Start with the database. Nimbus’s feed service queried DynamoDB and RDS for the same hot rows constantly: the trending post list, popular user profiles, the “photo of the day.” Every page load hit the database for data that changed maybe once an hour.
Repeated identical reads are exactly what an in-memory cache is for. Priya put ElastiCache for Redis in front of the hot read paths. The feed query β user profile plus recent-post metadata β went from a 40 ms round trip that consumed provisioned RDS capacity to a sub-millisecond Redis hit that consumed almost nothing.
The effect compounds two ways. Latency drops, and so does the database itself: RDS read replicas that had been added purely to absorb read load could be scaled down. For the DynamoDB paths, the same logic applied through DynamoDB Accelerator (DAX) β a purpose-built in-memory cache for DynamoDB that returns cached items in microseconds and lets you drop provisioned read capacity units. For simple string/blob key-value lookups with no clustering needs, Memcached would have done the job just as well; Redis won here because Nimbus wanted sorted sets for the leaderboard.
The DB layer’s read cost fell by more than half once the cache absorbed the repetitive reads.
No CDN in front of anything
Now the big one. Nimbus served photos, avatars, thumbnails, CSS, and JavaScript directly from S3 and from the application load balancer β straight out of the region, to users on six continents.
Every byte left an AWS region as egress, at full data-transfer-out rates. Every request hit the origin. A user in Jakarta pulling the home feed dragged image bytes all the way from us-east-1, and Nimbus paid regional egress for the privilege, over and over, for content that never changed.
Putting CloudFront in front of it changes the economics twice. First, egress: data transfer out to the internet through CloudFront is billed at CDN rates, which are lower than raw regional S3 egress, and β critically β origin-to-CloudFront transfer from S3 in the same partition is free. Second, offload: once an object is cached at the edge, the origin never sees the next request for it at all. S3 GET charges, ALB LCUs, and resize compute all evaporate for cache hits.
For a photo-sharing app this is not a nice-to-have. Static image bytes are the entire product. Nimbus fronted S3 and the app with CloudFront, and the origin request volume dropped by roughly 90% within a day as the edge warmed up.
A cache that doesn’t hit is just latency with extra steps
Except at first, it barely helped. Jenny pulled the CloudFront CacheHitRate metric and it was sitting at 41%. More than half of requests were still going to origin. A CDN with a bad hit ratio is decoration.
Two culprits, both classic:
Bad TTLs
The photos were being served with Cache-Control: max-age=0, inherited from a default nobody had touched. Immutable content β a thumbnail keyed to a specific photo version β should be cached for a long time. Priya set Cache-Control: public, max-age=31536000, immutable on versioned image URLs (the version is in the object key, so a new image gets a new URL and cache-busting is automatic). TTL went from zero to a year.
Cache keys that fragment
The other half of the problem was the cache key. CloudFront was including every query string and a pile of headers in the cache key. The trending cat thumbnail was requested as .../cat.jpg?utm_source=twitter, ?utm_source=email, ?ref=push… each a distinct cache entry, each a separate origin fetch, all returning the identical bytes. Marketing’s tracking parameters were shredding the cache.
Priya switched to a cache policy that keyed only on the parts that actually change the response and forwarded the tracking params without caching on them. Overnight the hit ratio climbed from 41% to 96%. Same infrastructure β the cache was just finally hitting.
Caching at the right tier
The deepest lesson was about where the cache lives. When Priya first started, the instinct was to cache inside the application β memoize the resized thumbnail in the Lambda, or in a Redis blob store the app checks before resizing.
That helps, but it still pays for the request to arrive: the user’s packet crosses the ocean, hits CloudFront (miss), hits the ALB, wakes a Lambda, which then checks a cache and returns bytes. You’ve saved the resize, but paid for the round trip, the load balancer, and the invocation.
An edge cache removes the request entirely. When the thumbnail lives in CloudFront’s edge cache, the request for the trending cat terminates in Jakarta. The origin β S3, the ALB, the Lambda, the whole us-east-1 machine β never hears about it. No egress from the region, no GET, no compute, no invocation. Nothing.
The rule of thumb Priya wrote on the whiteboard: cache as close to the user as the content’s freshness allows. Deep application caches save recomputation. Edge caches save the entire journey. For Nimbus’s thumbnails β immutable, versioned, wildly popular β the correct tier was the edge, not a clever memoization buried three layers down in PhotoBlob.
The trending cat is now generated exactly once, cached at 300-plus edge locations, and served to the world without ever waking the origin again. Chad described this to the CFO as “leveraging our global edge fabric.” It was, technically, true. Jenny just calls it remembering the answer.
Takeaways
- Add an in-memory caching layer (ElastiCache Redis/Memcached, or DAX for DynamoDB) in front of repeated identical DB and API reads β it cuts both database cost and latency and can let you scale down read replicas/RCUs.
- Put a CDN (CloudFront) in front of static and dynamic content to lower egress costs and offload origin GETs, compute, and load-balancer usage.
- Watch the CloudFront
CacheHitRatemetric and fix a low ratio by setting long TTLs on immutable content and trimming cache keys so tracking params don’t fragment the cache. - Cache at the right tier: an edge/CDN cache removes the request from the origin entirely, whereas caching deep in the app still pays for the round trip and invocation β for popular immutable content like thumbnails, cache at the edge.
- For a photo app specifically, serve versioned thumbnails from CloudFront edge caches instead of regenerating them per request against the origin.
Part 9 β Everything Is a Kafka, Nothing Is a Kafka
Priya found the cluster the way you find a raccoon in the attic: by following the noise of money disappearing.
It was line item 47 on the Cost Explorer breakdown β “Amazon MSK, us-east-1, *$1,190/month**” β and it belonged to a service called event-fanout-v3. She pulled the CloudWatch metrics. Over the trailing thirty days, the three-broker Managed Streaming for Kafka cluster had processed a grand total of roughly 15,000 messages. About 500 a day. One message every three minutes, give or take, humming along on infrastructure designed to move a million per second.
She walked over to Marcus’s desk. Marcus, Principal Architect, defender of Nimbus Platform 2.0, did not look up. “It’s the notification event bus,” he said. “We might need it at scale.”
“Marcus, it’s moving less traffic than my group chat.”
He gestured at the whiteboard, where a diagram of hexagons connected by arrows suggested a distributed system of real ambition. “Kafka gives us replayability, ordering guarantees, and durable retention. When we hit real scale, we’ll be glad it’s there.” Chad Ellison, passing in his Patagonia vest, nodded approvingly and said Nimbus was “building like a company that spends the GDP of a small island.” Which, Priya noted, was technically the problem.
The tool is not the trophy
Here is the uncomfortable truth about managed services: they are wonderful, and they are priced for the scale they were built to serve, whether or not you operate at that scale.
MSK bills you for broker instances that run 24⁄7 regardless of throughput. Marcus’s three kafka.m5.large brokers cost roughly $0.21/hour each, plus storage, plus data transfer β call it $1,190/month all-in β to deliver 500 messages a day. An SQS queue would have handled the same workload for essentially free: the first million requests per month are on the free tier, and after that it’s $0.40 per million requests. At 15,000 messages a month, Nimbus’s SQS bill would round to zero. Replace the whole thing with an SQS Standard queue feeding a Lambda consumer, and you keep durability, retries, and a dead-letter queue, and you delete $1,190/month.
This pattern repeats across the whole account like a recurring dream:
EMR for a spreadsheet’s worth of data
There was a persistent EMR cluster β three m5.xlarge nodes plus the EMR surcharge, about $600/month β running a nightly job that aggregated signup counts. The input was 2 GB of CSV. Two gigabytes. That is not a big-data problem; that is a laptop. The job was rewritten as a single Athena query (Athena is $5 per TB scanned, so this query cost about $0.01 per run) reading Parquet from S3. EMR cluster terminated. $600/month gone, and the job got faster because it stopped spending four minutes bootstrapping Hadoop.
OpenSearch as a very expensive grep
log-search-prod was a three-node OpenSearch domain, r6g.large.search instances plus EBS, around $900/month, that engineers used to grep application logs maybe twice a week during incidents. CloudWatch Logs Insights queries the same logs at $0.005 per GB scanned with zero standing cost. For a team that searches logs occasionally, the managed search cluster was a Ferrari kept idling in the driveway in case someone needed to check the mail.
Death by a thousand tiny services
The second act of over-engineering wasn’t a single monster β it was a swarm. Nimbus Platform 2.0 had been decomposed, per Marcus’s architecture, into 34 microservices. user-avatar-service. avatar-thumbnail-service. avatar-thumbnail-cache-invalidation-service. Each one deserving, in the diagram, of its own box.
The problem is that a microservice in the cloud is never just the code. Each of those 34 services had:
- Its own Application Load Balancer: $16.20/month minimum in hourly charges ($0.0225/hour) before a single LCU of traffic. That’s $550/month in ALBs just existing.
- Its own idle compute baseline β even at one small task each on Fargate, 34 services running
0.25 vCPU / 0.5 GBaround the clock is roughly $320/month of doing nothing in particular. - Its own path through the NAT gateway from hell for pulling images and calling AWS APIs, each dragging data-processing charges at $0.045/GB.
Fourteen of those 34 services averaged under one request per second. They were consolidated into three services grouped by actual bounded context, sharing a load balancer via path-based routing. Eleven ALBs deleted. The idle Fargate baseline dropped by two-thirds. The lesson: microservices are an organizational tool for letting teams deploy independently. If one team owns fourteen of them, that’s not microservices, that’s a monolith wearing eleven trench coats and paying eleven load-balancer bills.
Kubernetes, for four containers
Then there was EKS. Nimbus ran an EKS cluster to host β Priya counted twice β six containers. The EKS control plane alone is $0.10/hour, $73/month, before you run a single pod, plus the node group, plus the operational tax of an engineer who understands kubectl incidents at 2 a.m.
For six containers with no need for custom schedulers, service meshes, or multi-tenant namespacing, EKS was pure ceremony. Moving those workloads to ECS on Fargate eliminated the $73/month control-plane fee (ECS has no control-plane charge) and, more importantly, removed the person-cost of babysitting Kubernetes. This is not an anti-Kubernetes screed β at 300 containers with complex scheduling, EKS earns its keep and its control-plane fee vanishes into the noise. But “we might run it on Kubernetes someday” is not a reason to pay for Kubernetes today.
The observability tax
Nimbus was also being quietly billed for watching itself too closely.
- Detailed monitoring was enabled on every EC2 instance β 1-minute metrics at $2.10 per instance per month versus free 5-minute basic monitoring. Across 180 instances, that’s $378/month for granularity nobody was alerting on.
- X-Ray was on everything. Every request traced, at $5.00 per million traces recorded. A high-traffic ingest path was generating 400 million traces/month β $2,000 β when 5% sampling would have told the same story for $100.
- GuardDuty was running in throwaway sandbox accounts where it billed against VPC Flow Logs and CloudTrail event volume for accounts that got torn down weekly. Security theater, priced per GB analyzed.
None of this made the app better. It made the app observed, expensively, in places nobody was looking.
Managed vs. custom, evaluated honestly
The fair conclusion isn’t “managed bad, DIY good.” It cuts both ways.
Nimbus’s home-grown Redis-on-EC2 cluster β self-patched, self-monitored, with a failover script Deepak swore worked β cost about $400/month in instances but easily $2,000/month in the engineering time spent keeping it alive. Moving to ElastiCache Serverless cost more in raw AWS dollars and saved money on total cost of ownership. Managed won.
Meanwhile MSK, EMR, and OpenSearch lost on total cost, because the managed premium bought capabilities the workload never used. The rule is boring and unglamorous: match the tool to the actual scale and the actual operational burden, measured in real numbers, not in the scale you’re manifesting on a whiteboard.
Marcus kept the MSK cluster for exactly one more week, until the SQS-plus-Lambda replacement passed its load test at 200x the current traffic for $0/month. Then he quietly deleted it. On the whiteboard, one hexagon got erased. Chad never noticed.
Takeaways
- Match managed services to actual scale: MSK/EMR/OpenSearch/Kubernetes are priced for their target workloads β measure your real throughput before adopting them, and swap in SQS, Lambda, Athena, or a small RDS when the numbers don’t justify the cluster (500 messages/day belongs in a free-tier queue, not a $1,190/month Kafka cluster).
- Consolidate premature microservices: every tiny service carries its own ALB ($16+/month), idle compute baseline, and NAT path β group services by bounded context and share load balancers via path routing instead of paying per hexagon.
- Right-size the orchestrator: EKS charges $73/month for the control plane before any pods β for a handful of containers, ECS/Fargate or plain EC2 is cheaper in both dollars and operational toil; save Kubernetes for genuine scheduling complexity.
- Trim the observability tax: turn off EC2 detailed monitoring where 5-minute metrics suffice, sample X-Ray (5% often tells the same story at 1/20th the cost), and scope GuardDuty/tracing to accounts that actually matter.
- Kill unused premium features: don’t leave detailed monitoring, X-Ray-on-everything, or GuardDuty running in throwaway/sandbox accounts that get destroyed weekly.
- Evaluate managed vs. custom both ways: include total cost of ownership β sometimes managed (ElastiCache over self-run Redis) wins on saved engineering time; sometimes custom wins because the managed premium buys capabilities you never use. Decide with real numbers, not aspiration.
Part 10 β The Boring Stuff That Saves Millions
Jenny found it on a Tuesday, which is when Jenny found most things, because Tuesday was when she read Cost Explorer for fun.
“There’s an account called sandbox-test-delete-me,” she said, “and it’s spending $2,100 a month.”
Priya looked up. sandbox-test-delete-me had been created eighteen months ago by an intern who had since graduated, gotten a job, and possibly gotten married. Nobody had deleted it. Inside it, running patiently in the dark, was a p3.2xlarge GPU instances someone had spun up to try a machine-learning demo and never turned off. At roughly $3.06/hour, it had quietly burned about $44K across eighteen months to compute absolutely nothing.
This is the thing nobody tells you about cloud cost: the expensive part is rarely the clever architecture. It’s the boring part. It’s the stuff no one owns, no one tags, and no one is watching. This installment is about that stuff β governance, the quiet ancillary services, and the mess that multi-account setups create. None of it is glamorous. All of it saves millions.
The graveyard of** orphaned resources
Priya ran a sweep across every account. The findings read like an archaeological dig:
- 206 unattached EBS volumes. Volumes survive the instances they were attached to. gp3 storage is $0.08/GB-month, so 206 volumes averaging 100 GB was about $1,600/month of storage attached to nothing.
- 38 unassociated Elastic IPs. An EIP attached to a running instance is free; an EIP sitting unattached costs $0.005/hour β about $3.60/month each, $137/month total β a small charge whose only purpose is to punish forgetfulness.
- 9 empty load balancers with no healthy targets, billing their $16+/month minimum to route traffic to the void.
- ~2,800 stale EBS snapshots, some dating to 2024, at $0.05/GB-month of incremental storage.
- Idle VPC endpoints and dangling ENIs left behind by deleted Lambda functions and torn-down services, each interface endpoint quietly charging ~$7.20/month.
None of this was anyone’s fault, exactly. That was the point: it was nobody’s job.
You can’t manage what you can’t see
The root cause was that Nimbus had almost no cost governance. Specifically:
No tagging. Cost Explorer could show spend by service, but not by team, environment, or product feature, because resources weren’t tagged. Activate cost-allocation tags (team, env, service, owner) and suddenly you can answer “who spent this?” β which is the question that ends 90% of cost arguments.
No budgets, no anomaly detection. There was no AWS Budget alerting anyone when spend crossed a threshold, and no Cost Anomaly Detection β a free, ML-based service that would have emailed someone the day those GPU instances spun up in the sandbox, instead of eighteen months later. Both are effectively free. Not using them is like having a smoke detector and removing the battery.
Compute Optimizer and Trusted Advisor, ignored. Compute Optimizer had been sitting there for a year recommending downsizes on over-provisioned instances. Trusted Advisor’s cost-optimization checks had flagged the idle load balancers and low-utilization instances. Nobody read either. Free advice, unopened.
No CUR. Nimbus used Cost Explorer casually but had never enabled the Cost and Usage Report β the hour-by-hour, resource-by-resource billing firehose that lets you actually audit spend. Without it, deep analysis is guesswork.
No FinOps ownership. The real bug. Until Priya took it up, cost was everyone’s responsibility, which meant no one’s! And the support plan and AWS Marketplace charges β a $15,000/year Business Support tier and a forgotten $2,000/month third-party monitoring SaaS bought through Marketplace β sat unexamined because they weren’t anyone’s line to defend.
The silent killer: CloudWatch Logs
Then Priya found the single biggest surprise on the whole bill. Not compute. Logs.
CloudWatch Logs charges about $0.50/GB to ingest and $0.03/GB-month to store, forever, if you never set retention. Nimbus’s log groups had no retention policy, so logs from 2023 were still accruing storage charges. Worse, the applications logged at DEBUG in production, emitting full request/response bodies. Total ingestion was roughly 9 TB/month β about $4,500/month just to ingest logs, plus a growing storage tail.
The fix was unglamorous and enormous:
- Set retention on every log group (30 days for app logs, 90 for audit) β killing the infinite storage tail.
- Drop DEBUG logs in production and sample high-volume access logs.
- Route bulk logs to S3, where storage is $0.023/GB and you query with Athena on demand.
Ingestion dropped from 9 TB to under 2 TB/month. That one change saved more than the entire microservices consolidation from Part 9.
The other ancillary line items
Once you start looking at the boring services, they add up:
- KMS request costs. Every
Decrypt/GenerateDataKeycall is $0.03 per 10,000 requests. A service decrypting a config secret on every request was making 90 million KMS calls/month β $270 β when data key caching (encrypt once, reuse the plaintext key in memory for a bounded window) cut it by 95%. - Secrets Manager vs. SSM Parameter Store. Nimbus stored 300 non-rotating config values in Secrets Manager at $0.40/secret/month β $120/month β plus API charges. Static, non-rotating config belongs in SSM Parameter Store, where standard parameters are free. Reserve Secrets Manager for things that actually rotate (database creds), and the $120 becomes $0.
- GuardDuty, Config, CloudTrail scope. Three CloudTrail trails were logging the same events into three buckets β the first copy of management events is free, additional copies are not. CloudTrail data events (S3 object-level, Lambda invocations) were on for high-traffic buckets at $0.10 per 100,000 events, generating charges by the millions. AWS Config was recording every resource change in noisy dev accounts. Scope each to what you actually need to audit.
- WAF. $5/month per web ACL, $1/month per rule, plus $0.60 per million requests. Nimbus had 40 rules across web ACLs, many duplicating managed rule groups. Consolidating cut the standing rule cost by half.
- VPC Flow Logs. Enabled on every subnet at full verbosity, shipped to CloudWatch Logs (see: silent killer). Sampling and sending them to S3 instead cut this line by 80%.
The multi-account mess
Finally, the org-level waste β the kind you only see from the management account:
- Reserved Instances and Savings Plans weren’t shared org-wide. RIs and Savings Plans purchased in one account weren’t set to shared scope, so unused commitment in Account A didn’t cover matching usage in Account B. Turning on RI/Savings Plans sharing in the org billing settings applied idle reservations to hungry accounts β instant coverage, zero new spend.
- Consolidated billing volume tiers weren’t being fully exploited. Under consolidated billing, aggregated usage (like S3 storage across accounts) rolls up to hit lower per-unit tiers together. Fragmented, un-consolidated accounts were each buying at the higher first-tier rate.
- Untracked sandbox accounts like
sandbox-test-delete-mehad no SCPs restricting expensive instance types, no per-account budgets, and no auto-cleanup. Service Control Policies to block GPU/large instances in sandboxes, a $200 budget with an alert, and a scheduled Lambda that nukes untagged resources older than 7 days would have caught the $88K mistake on day one. - Duplicate services per account. Every account ran its own NAT gateways, its own logging pipeline, its own egress path. Centralizing networking and egress through a shared services account and aggregating logs to one destination removed dozens of redundant NAT gateways from hell.
- No centralized cost visibility. There was no management-account CUR aggregating all accounts, so no one could see total spend in one place. Enabling org-wide CUR into a single S3 bucket, queried with Athena, finally gave Priya one dashboard instead of forty tabs.
Priya deleted sandbox-test-delete-me herself. It felt, she said later, like closing a very expensive raccoon back out into the night.
Takeaways
- Sweep for orphaned resources regularly: unattached EBS volumes, unassociated EIPs ($3.60/mo each), empty load balancers, idle VPC endpoints, stale snapshots, and dangling ENIs cost real money for nothing β automate the sweep.
- Tag everything with cost-allocation tags (
team,env,service,owner) so Cost Explorer can attribute spend and end ownership arguments. - Turn on Budgets and Cost Anomaly Detection β both effectively free β so a runaway sandbox pings you on day one, not month eighteen.
- Actually read Compute Optimizer and Trusted Advisor, and enable the Cost and Usage Report for real auditing instead of guessing from Cost Explorer.
- Assign clear FinOps ownership β cost that is “everyone’s job” is no one’s β and review your support-plan tier and Marketplace subscriptions for forgotten charges.
- Tame CloudWatch Logs, the biggest silent cost: set retention on every log group, drop DEBUG in prod, sample high-volume logs, and offload bulk logs to S3/Athena.
- Use data key caching to slash KMS request costs, and put non-rotating config in free SSM Parameter Store instead of paid Secrets Manager.
- Scope GuardDuty, Config, and CloudTrail deliberately β avoid duplicate trails and unnecessary data events β and consolidate WAF rules and sample/redirect VPC Flow Logs.
- Share RIs/Savings Plans org-wide and exploit consolidated-billing volume tiers so commitment and volume discounts apply across every account.
- Govern sandbox/dev accounts with SCPs, per-account budgets, and auto-cleanup, and centralize logging/networking/egress plus a management-account CUR for one source of truth.
Part 11 β The Distributed Monolith and Other Horror Stories
Priya found the architecture diagram taped to the wall of the war room. Someone β the handwriting suggested Marcus β had drawn PhotoBlob in the center, a single friendly rectangle labeled “PhotoBlob (core).” Around it orbited eleven more rectangles, each connected to PhotoBlob by a bidirectional arrow, and each connected to the same database, drawn as a cylinder someone had lovingly shaded.
“So it’s a monolith,” said Priya.
“It’s a services architecture,” said Marcus.
“Every one of these services calls PhotoBlob synchronously, and they all share one Postgres instance.”
“We’re on a journey,” said Marcus.
Deepak, who had not looked up from his laptop, said: “PhotoBlob is fine. PhotoBlob has served us well.” He said this the way people talk about a dog that has bitten several children but is still, technically, a good boy.
Priya opened Cost Explorer. Compute was $180K a month. She had a feeling that most of it was going to waste in ways that had nothing to do with pricing and everything to do with the diagram on the wall.
The instances are mostly asleep
PhotoBlob’s image-processing worker ran on a fleet of c6i.8xlarge instances β 32 vCPUs each. Priya pulled up CloudWatch and found CPU utilization hovering around 4%. Not 40. Four.
The worker was a single-threaded Python process. One process, one core, on a 32-core machine. They were paying for 32 cores and using one.
One process is not a plan
Single-core executables on multi-core instances are one of the purest forms of cloud waste, because the cloud bills you for the whole instance whether you use it or not. On a laptop, an idle core is free. On a c6i.8xlarge, thirty-one idle cores cost real money, twenty-four hours a day.
The fix is not exotic. Run one process per core with a supervisor, or a worker pool, or β better β package the worker as many small containers and let the scheduler pack them onto instances until the cores are full. Bin-packing thirty-two single-threaded containers onto a 32-core box turns a 4% machine into a 90% machine, and 90% is the number you actually want to pay for.
Blocking, busy-waiting, and the polling tax
The worker also had no parallelism in a second, sneakier way: it processed images with blocking synchronous I/O. Fetch an image from S3 (wait), resize it (work), write it back (wait). During both waits, the core did nothing, but it was still checked out and unavailable.
Worse, PhotoBlob’s queue consumer didn’t use long-polling. It ran a busy-wait loop β poll SQS, get nothing, sleep(0.1), poll again β spinning the CPU ten times a second forever. On EC2 you at least pay a flat hourly rate for the spinning. Priya had also found the same pattern in a Lambda that polled a DynamoDB table in a loop until a flag flipped. In serverless, a busy-wait loop is a metered crime: you pay per millisecond of execution, so a function that spends four seconds spinning is four seconds of billed compute doing nothing. Blocking waits should yield; loops should long-poll or use event triggers, not spin.
Fat runtimes, slow starts
Half the “microservices” ran on the JVM inside 1.4 GB container images. The oversized runtime β a heavy JVM footprint for a service that did little more than validate a JSON body β meant every container reserved more memory than its work required, and every cold start dragged, because a fat container has more to pull and a JVM has more to warm up. On Lambda-backed services this showed up as cold-start latency; on ECS it showed up as slow, expensive scale-out during traffic spikes. Right-sizing the runtime to the work β a slim base image, a lighter runtime where the JVM wasn’t earning its keep β shrank both the memory reservation and the time-to-ready.
The diagram on the wall
Now to the wall.
One big server that does everything
Historically, PhotoBlob was one huge server running everything: uploads, resizing, feed generation, notifications, billing. This is a single point of failure β when it fell over, all of Nimbus fell over β and, just as expensive, you cannot scale its components independently. Feed generation needed more CPU during peak hours; notifications needed more memory during campaigns. But because they shared one process, the only lever was to scale the entire machine up, paying for peak feed capacity even at 3 a.m. when nobody was scrolling.
The distributed monolith: all the cost, none of the benefit
Then the team “broke it up,” which brings us to the diagram. What they built is the distributed monolith: services split across servers but still tightly coupled, sharing one database, calling each other in synchronous chains. You pay the full price of distributed systems β network hops, serialization, partial failures, cross-AZ data charges β and receive almost none of the benefits, because nothing can be deployed, scaled, or reasoned about independently. A change to the shared schema still requires coordinating eleven services.
Chatty chains multiply latency, failure, and cross-AZ bills
Because the calls are synchronous chains β the upload service calls the metadata service, which calls PhotoBlob, which calls the tagging service β latency adds up and failure multiplies. If each of five hops is 99.9% available, the chain is only about 99.5% available; each hop is another chance to time out. And on AWS every one of those hops that crosses an Availability Zone incurs cross-AZ data transfer charges at roughly one cent per gigabyte each way. Chatty services gossiping across AZs turn a rounding error into a line item.
But don’t overcorrect
Priya was careful here, because the opposite mistake was also on the table. Marcus’s “Platform 2.0” proposed splitting the tagging service into four services. That is premature, over-granular microservices β the cost of the network and the operational overhead without a real boundary to justify it. And the shared database across services is the coupling that makes the whole thing a monolith regardless of how many rectangles you draw. The rule Priya wrote on the whiteboard: split on real boundaries, give each service its own data, and prefer async messaging over synchronous chains.
Pets, snowflakes, and the local disk
Autoscaling was the next fight, and it was losing to biology.
Pets, not cattle
PhotoBlob’s servers were stateful. Each one held in-memory session state and cached user data that existed nowhere else. You could not just terminate one β it was a pet, with a name and feelings. Stateful application servers block autoscaling (you can’t scale in without losing state) and block Spot (you can’t tolerate a two-minute interruption notice). Nimbus was paying full On-Demand for a fleet it couldn’t safely shrink.
Two culprits made the servers stateful. First, session affinity / sticky sessions: the load balancer pinned each user to one server, so that server couldn’t be replaced without logging people out. Second, local filesystem dependence: uploads were written to the instance’s local disk before processing. On ephemeral compute, local disk vanishes when the instance does, so this both created state and lost data on every scale-in. Sessions belong in a shared store (Redis, DynamoDB); files belong in S3. Do that and the servers become cattle β interchangeable, disposable, Spot-eligible.
Snowflakes and the vertical-only ceiling
Every server was also a snowflake β hand-configured over years, no two quite alike, impossible to reproduce. The fix is Infrastructure as Code plus immutable images: bake an AMI or container, deploy it identically, replace rather than patch. And because the servers were pets, the team’s only scaling move had been vertical β buy a bigger box. Vertical scaling has a hard ceiling (there is a biggest instance) and it’s a SPOF by construction. Horizontal scaling β more small cattle behind a load balancer β is what actually lets you follow demand and turn machines off when it’s gone.
The database is doing chores it shouldn’t
Finally, the shared Postgres, which was on fire.
Pagination, scans, and the N+1
The feed API used application-layer pagination: it fetched every photo for a user into application memory, then sliced out page 3. For a power user with 40,000 photos, that’s 40,000 rows pulled across the network to return 20. The fix is to push the limit into the query β LIMIT/keyset (seek) pagination and projections that select only the needed columns β so the database returns the page, not the library. The same rule holds in DynamoDB: use a paginated Query with LastEvaluatedKey, not a Scan of the whole table.
Which was the next problem: full table scans and missing indexes. The “search by caption” feature had no index, so every search scanned the whole photos table. On a shared database, one unindexed query starves every other service on the box.
And the feed had a classic N+1 query pattern: one query to list 50 photos, then 50 more queries to fetch each photo’s author. Fifty-one round trips where a join or a batched IN query would do one. Every extra round trip is latency and, across AZs, transfer cost.
One database, no cache
There was no caching layer anywhere β every feed render hit Postgres cold, even though feeds barely change between refreshes. A read-through cache (ElastiCache) in front of the hot paths would have taken enormous load off the database. And underneath it all, the deepest anti-pattern: one database for every workload. Relational feed data, full-text search, session state, and a time-series of view counts were all crammed into one Postgres instance, so the instance was sized for the sum of every access pattern and tuned for none. Purpose-built stores β Postgres for relational, OpenSearch for search, DynamoDB for key-value sessions, a cache for hot reads β each right-sized, each independently scalable.
Priya took the diagram off the wall.
“We’re keeping PhotoBlob,” she told Deepak. “We’re just going to stop asking it to be the entire company.”
Deepak nodded slowly, like a man being told his dog would be allowed to stay, but must, from now on, stop biting the children.
Takeaways
- Single-core process on a multi-core instance: run one process per core, a worker pool, or pack many small containers so you pay for cores you actually use.
- No parallelism / blocking synchronous processing: use async or concurrent I/O so cores aren’t checked out waiting.
- Busy-wait / polling loops: replace with long-polling or event triggers β in serverless, spinning is billed millisecond-for-millisecond.
- Oversized runtime: right-size the runtime and slim the image to cut memory reservation and cold-start time.
- One huge server running everything: it’s a SPOF and forces whole-machine scaling; separate components so you can scale them independently.
- Distributed monolith: don’t pay distributed-systems cost (hops, cross-AZ transfer, partial failure) while staying tightly coupled through a shared DB and synchronous chains.
- Chatty synchronous call chains: collapse hops; each one multiplies latency, compounds failure probability, and racks up cross-AZ charges.
- Premature / over-granular microservices: split on real boundaries, not vanity β over-splitting adds network and ops cost with no payoff.
- Shared database across services: give each service its own data store; a shared DB is the coupling that makes it a monolith.
- Stateful application servers: move state out (Redis/DynamoDB/S3) so servers become disposable cattle β autoscaling- and Spot-eligible.
- Session affinity / sticky sessions: externalize sessions to a shared store so any node can serve any user.
- Local filesystem dependence: write to S3; local disk dies with ephemeral compute.
- Snowflake servers: use Infrastructure as Code plus immutable images; replace, don’t hand-patch.
- Vertical-only scaling: scale horizontally behind a load balancer to escape the single-box ceiling and follow demand.
- Application-layer pagination: push
LIMIT/keyset pagination and projections into the query; in DynamoDB, page withLastEvaluatedKey. - Full table scans / missing indexes: index your query paths; one unindexed scan starves a shared database.
- N+1 queries: batch with a join or
INclause instead of one query per row. - No caching layer: put a read-through cache (ElastiCache) in front of hot, rarely-changing reads.
- One database for every workload: use purpose-built stores, each right-sized and independently scalable.
Part 12 β It Compiles, Ship It
At 2:14 a.m., the entire Nimbus feed went dark because the thumbnail service was slow.
Not down. Slow. The thumbnail service took eight seconds to respond instead of the usual two hundred milliseconds, and because the upload path called it synchronously with no timeout, every upload request held its worker thread open for eight seconds, and the worker pool filled, and then the health check requests couldn’t get a thread either, and then the load balancer decided the whole fleet was unhealthy, and then 40 million people saw a spinner.
Priya got the page. She traced it in eleven minutes and then sat in the dark for a bit, because the thing that took down Nimbus was that one non-critical dependency β thumbnails β got a little slow, and nothing in the system was built to survive a dependency being a little slow.
Chad’s Slack message arrived at 2:40 a.m.: “why are we down, I told the board we spend the GDP of a small island on infra.” Priya typed, deleted, and retyped a reply four times before settling on: “Working on it.”
Things that fall over when one thing gets slow
No timeouts, no retries, no circuit breakers
The root cause was that the upload service called thumbnails with no timeout. A call with no timeout waits forever; forever is longer than your thread pool. Every outbound call needs a timeout shorter than the caller’s own budget.
Then there was the opposite failure. The notification service did retry β naively, immediately, three times, with no backoff. When the downstream service got slow, every caller retried, tripling the load on the thing that was already struggling. That’s a retry storm: retries kick a system while it’s down and keep it down. The cure is timeouts plus retries with exponential backoff and jitter, wrapped in a circuit breaker that trips open after repeated failures and stops calling the sick dependency entirely until it recovers.
Blocking the user on work the user doesn’t care about
The deeper sin was synchronous coupling to non-critical work. The upload request blocked on thumbnail generation. It also blocked on sending a confirmation email and on writing an analytics event. None of these are things the user is waiting for β they tapped “post,” they want to see “posted.” Emails, analytics, and thumbnails should be offloaded to a queue (SQS, EventBridge) and processed asynchronously. The user request returns in milliseconds; the slow, non-critical work happens later, on its own schedule, and its slowness can no longer take down the critical path.
One of everything
Priya drew the failure domains and found single points of failure stacked like dominoes: the primary database ran in one AZ with no replica; all egress went through one NAT gateway (the NAT gateway from hell returns); several services ran as a single instance with no second copy. Any one of them dying meant an outage. Multi-AZ databases, redundant NAT gateways per AZ, and at least two instances behind every load balancer are not gold-plating; they are the difference between a blip and an incident.
Dead nodes and dropped deploys
Two more gaps made every incident worse. There were no health checks deep enough to matter, so the load balancer kept routing traffic to nodes that were technically running but internally broken. And there was no graceful shutdown β on every deploy, instances were killed mid-request, so every routine deployment dropped a fistful of live requests. Real health checks pull dead nodes out of rotation; connection draining and graceful shutdown let a terminating node finish its in-flight work before it goes.
Idempotency, so retries are safe
Underlying all of it: the payment webhook was not idempotent. When a retry did fire, it could charge a card twice. Retries are only safe if the operation is idempotent β an idempotency key so that the same request processed twice has the same effect as processing it once. Without it, you must choose between resilience and correctness, and that is not a choice anyone should have to make at 2 a.m. **
Fighting the cloud instead of using it
The next morning, calmer, Priya went looking for why the bill was $340K, and found a theme: Nimbus kept rebuilding AWS inside AWS.
Reinventing managed services
The queue that should have decoupled uploads from thumbnails didn’t exist because someone had self-hosted RabbitMQ on a pair of EC2 instances β which they patched, monitored, and occasionally lost sleep over. Next to it: a self-managed Redis cluster and a self-hosted Postgres on raw EC2, backups scripted by hand. Every one of these reinvents a managed service. SQS, ElastiCache, and RDS exist, are cheaper once you count the human time, and don’t page anyone at 2 a.m. for a failed backup cron. Running your own message broker on EC2 is paying for the instance and paying an engineer to be the on-call for it.
Hardcoded secrets and hardcoded assumptions
Grepping the repo, Priya found database passwords committed in a config file β hardcoded configuration and secrets. These belong in SSM Parameter Store (for config) and Secrets Manager (for credentials), fetched at runtime, rotated without a redeploy, and kept out of git history forever.
She also found code that assumed instances live forever: cron jobs pinned to a specific hostname, an in-memory cache expected to survive for weeks, a “warmup” that ran once at boot and never again. This is ignoring ephemerality β on the cloud, instances come and go, and anything that assumes a long-lived host breaks the moment autoscaling or a Spot reclaim does its job.
A fleet sized for a peak that happens twice a day
The compute fleet was fixed: 80 instances, running 24⁄7, sized for the Friday-evening peak. Traffic at 4 a.m. was a tenth of that. This is hardcoded capacity with no elasticity β paying for peak around the clock. Autoscaling on a real metric (queue depth, CPU, request rate) would run 80 at peak and 12 overnight, and the difference is enormous when it’s compounded across 720 hours a month.
Not knowing where the bytes go
Finally, region and AZ blindness. Services were scattered across AZs with no thought to which called which, so requests hopped AZs constantly, each hop billing cross-AZ transfer and adding a millisecond of latency. One analytics job in us-east-1 pulled its source data from a bucket in eu-west-1 on every run β cross-region transfer, the most expensive kind, hidden in a nightly job nobody watched. Data has gravity and a price tag; keeping chatty components in the same AZ and data in the same region is free money.
Flying the plane with the instruments taped over
No pipeline, no IaC, no idea
Deployments were manual β someone SSH’d in and pulled the latest build, service by service. There was no CI/CD, so deploys were slow, inconsistent, and terrifying, and there was no Infrastructure as Code β the whole environment had been clicked together in the console over three years and had drifted so far that nobody could recreate it. Releases were big-bang and coupled: all eleven services shipped together because they were entangled, so any one bug rolled back all of them.
The fix is the boring, load-bearing stuff: a CI/CD pipeline that builds, tests, and deploys automatically; IaC (Terraform, CloudFormation, CDK) as the single source of truth so the console matches the code; and decoupled, independently deployable releases so a thumbnail fix doesn’t require redeploying billing.
Flying blind
The reason the 2 a.m. incident took eleven minutes to diagnose and could easily have taken two hours was no observability. No distributed tracing, no useful metrics, no dashboards. You cannot right-size what you cannot measure, and you cannot diagnose what you cannot see. Metrics, traces, and structured logs are how you find the idle instance and the slow dependency.
But observability has its own trap, which Priya found in the bill. CloudWatch Logs ingestion was a top-five line item. Every service logged every request at DEBUG, in unstructured free text, and it all shipped to CloudWatch at ingestion cost per gigabyte. Over-logging turns your telemetry into a cost center. Log at sensible levels, emit structured (JSON) logs you can actually query, sample the high-volume paths, and set retention so logs don’t live forever at full price.
The debt with interest
Two last categories, cheap to fix and expensive to ignore.
Security debt
IAM was a disaster of convenience: services shared over-broad roles (“just give it AdministratorAccess, we’ll tighten it later” β later never came) and authenticated with long-lived access keys hardcoded into containers. Least privilege β scope each role to exactly what it needs β and IAM roles over static keys (the container assumes a role and gets short-lived, auto-rotated credentials) close the biggest blast radius on the account.
The network was flat and public. Databases sat in public subnets with security groups open to 0.0.0.0/0, because it was easier during a hackathon in 2024. Data stores belong in private subnets with tight security groups, and traffic to AWS services should go through VPC endpoints β which also, pleasingly, keeps that traffic off the NAT gateway from hell and off the transfer bill.
Image hygiene
The container images were huge β the 1.4 GB JVM images from Part 11 β which slowed every pull, every scale-out, every cold start. And ECR had no lifecycle policy, so every image build for three years was still sitting in the registry, accruing storage charges for artifacts no deployment would ever use again. Slim base images and an ECR lifecycle policy that expires old, untagged images are ten minutes of work that pay rent every month.
Priya closed her laptop. The list was long, but none of it was mysterious. It was just debt β taken on one reasonable-at-the-time shortcut at a time, and now, at $340K a month, coming due all at once.
“It compiles,” Deepak had said, three years ago. “Ship it.”
It had, in fact, shipped.
Takeaways
- No timeouts / retries / circuit breakers: set timeouts on every call, retry with backoff and jitter, and trip a circuit breaker to stop hammering a sick dependency.
- Naive retries: immediate no-backoff retries cause retry storms; use exponential backoff with jitter.
- Synchronous coupling to non-critical work: offload emails, analytics, and thumbnails to a queue so slow side-work can’t block the user request.
- Single points of failure: eliminate the lone instance / AZ / NAT / DB with multi-AZ, redundant NATs, and at least two instances per service.
- No health checks / graceful shutdown: use deep health checks to pull dead nodes, and connection draining so deploys don’t drop live requests.
- Retry storms & missing idempotency: make operations idempotent with an idempotency key so retries are safe and can’t double-charge.
- Reinventing managed services: stop self-hosting queues/caches/DBs on raw EC2; use SQS, ElastiCache, and RDS and reclaim the on-call cost.
- Hardcoded config & secrets: externalize to SSM Parameter Store and Secrets Manager; rotate without redeploys and keep credentials out of git.
- Ignoring ephemerality: never assume long-lived hosts; design for instances that come and go.
- Hardcoded capacity / no elasticity: autoscale on real metrics instead of running a peak-sized fleet 24⁄7.
- Region/AZ blindness: keep chatty components in one AZ and data in one region to cut hidden cross-AZ/region transfer cost and latency.
- Manual deployments / no CI-CD: automate build, test, and deploy so releases are fast, consistent, and boring.
- No Infrastructure as Code: define everything in Terraform/CloudFormation/CDK so the console can’t drift from source.
- Big-bang coupled releases: decouple services so they deploy and roll back independently.
- No observability: add metrics, traces, and dashboards β you can’t right-size or diagnose what you can’t see.
- Over-logging / unstructured logging: log at sensible levels in structured JSON, sample hot paths, and set retention before CloudWatch ingestion becomes a top-five line item.
- Over-broad IAM / long-lived credentials: apply least privilege and use IAM roles with short-lived credentials instead of static access keys.
- Everything public / flat network: put data stores in private subnets behind tight security groups and reach AWS services via VPC endpoints.
- Poor container/image hygiene: slim your images and add ECR lifecycle policies so stale artifacts stop accruing storage charges.
Part 13 β The Real Reason - Humans not Servers
The lights come up on the main stage at CloudScale Summit, and there is Chad Ellison, CTO of Nimbus, in the vest. The Patagonia vest. He clicks to a slide that just says $340K/MONTH in a font large enough to read from orbit.
“People ask me how we know we’ve made it,” Chad says, pausing for the room. “And I tell them β we spend the GDP of a small island on AWS. Every single month.” Warm laughter. A ripple of applause. Someone in row three actually whoops.
In the fourth row, Priya Nair β Staff Engineer, freshly and involuntarily crowned “FinOps lead” β sits very still, the way you sit when you have just watched someone describe a fever as a personality trait. She has spent six weeks proving that at least a third of that number is pure waste: the NAT gateway from hell, PhotoBlob’s idle midnight fleet, Big Bertha the over-provisioned analytics cluster running at 4% utilization. And here is her CTO, on a stage, celebrating the symptom as if it were the diagnosis.
She writes one line in her notebook: The bill isn’t the problem. The applause is.
This is the last post in the series, and it’s the one no tool can fix. Every previous installment was about instances and gateways and lifecycle policies. This one is about people, incentives, and the strange gravity that keeps waste alive long after everyone can see it. Because here’s the uncomfortable truth: your architecture didn’t get expensive by accident. It got expensive because, for a lot of people, expensive was working exactly as intended.
The Ego Tax: When the Design Is Someone’s Identity
Marcus Webb designed Nimbus Platform 2.0. He will tell you this within ninety seconds of meeting him. He will also tell you, when Priya suggests collapsing three of its services into one, that “we might need it at scale” β a phrase that has never once been followed by evidence.
Sunk cost wearing a lanyard
The problem isn’t that Marcus is wrong. The problem is that Marcus can’t afford to be right about this. His comp band, his principal title, and roughly two years of his professional identity are welded to that design. Asking him to admit it’s over-built is asking him to file for a demotion. “We invested two years in this” is not an engineering argument; it’s grief. And sunk cost is the most expensive emotion in the building β those two years are gone whether you keep the design or not.
How to address it: separate the person from the artifact. Praise the ambition, retire the parts that didn’t pay off, and make “we simplified Marcus’s platform” a win Marcus gets credit for β not a verdict against him. Nobody rewrites their own monument. They’ll happily co-author the renovation.
Not-invented-here and the rΓ©sumΓ© that ate the roadmap
There are two related diseases here. The first is not-invented-here: “we built it, we keep it,” regardless of whether a managed service now does the same thing for a tenth of the cost and none of the on-call. The second is its louder cousin, rΓ©sumΓ©-driven development β the reason Nimbus runs Kubernetes to schedule what is functionally a nightly cron job, Kafka to move messages a queue would carry, and microservices to serve forty million users who would never notice a monolith.
None of these were chosen because the problem demanded them. They were chosen because they look magnificent on a CV, and because in 2023 you couldn’t get promoted talking about a well-tuned Postgres.
How to address it: require an ADR β an Architecture Decision Record β for anything load-bearing, and make one section mandatory: “Why not the simpler option?” If the honest answer is “because it’s boring,” that’s your signal. You’ve just caught a career decision masquerading as a technical one.
The Scoreboard Is Wrong
Which brings us back to Chad on that stage. Chad isn’t a villain. Chad is responding, rationally, to the scoreboard he’s been handed β and the scoreboard says a big AWS bill is a trophy.
Vanity metrics and the gross-spend brag
Gross spend is a vanity metric. It tells you nothing about health. A $340K bill serving 40M users at healthy margins is a triumph; the same bill serving them at 4% cluster utilization is a slow-motion fire. But “gross spend” is the number that makes it into the keynote, so gross spend is the number that gets optimized upward.
How to address it: change what’s on the scoreboard. Stop celebrating total spend; start celebrating unit economics β cost per active user, per upload, per GB served, per transaction. When Priya finally got Chad a slide that read “cost per active user: $0.0085, down 22% this quarter,” his brag changed overnight. Same ego, better target. He now brags about efficiency at scale, which, mercifully, is a thing you can be proud of without lighting money on fire.
Empire-building, use-it-or-lose-it, and the orphaned cost
Two more incentive traps live here. Empire-building β where status is measured in headcount and budget, so no leader ever has a reason to shrink either. And use-it-or-lose-it budgeting, where a team burns its remaining cloud allocation in Q4 specifically so it isn’t cut next year. Both reward consumption for its own sake.
And underneath all of it: at Nimbus, cost was nobody’s job. Everyone assumed someone else was watching the meter. Nobody was. That’s why Priya exists now.
How to address it: give cost an owner β a named FinOps function, even if it’s one person part-time. Reward returning budget, not spending it. And measure leaders on outcomes-per-dollar, not dollars commanded.
The Culture That Protects Waste
Deepak and the fortress of “it works”
Deepak maintains PhotoBlob, the original monolith, and his position is theological: “it works, leave it alone.” He’s not lazy β he’s scared and busy, which look identical from the outside. Refactoring PhotoBlob has no allocated time, no reward, and considerable risk. So of course it never happens.
How to address it: make improvement part of the actual job, not a heroic side quest. Allocate explicit time for it. Reward the refactor that halves PhotoBlob’s idle footprint the same way you’d reward a shipped feature. What gets celebrated gets done.
Fear-driven “don’t touch it”
Deepak’s fear is rational, though. The reason nobody touches Big Bertha or the NAT gateway from hell is that there’s no safety net β no tests, no canary, no clean rollback. In that world, “don’t touch it” is the correct strategy.
How to address it: build the net first. Tests, canary deploys, one-click rollback. Fear-driven caution isn’t a character flaw; it’s a missing seatbelt. Install the seatbelt and watch how much braver everyone gets.
Hero culture, blame culture, and alert fatigue
Nimbus has a hero culture: the engineer who stays up until 3 a.m. firefighting a cost spike gets the plaque. The engineer who quietly prevented the spike gets nothing, because you can’t throw a party for an incident that didn’t happen. This trains people to value the fire, not the fire-proofing.
It’s compounded by blame culture β when the post-incident question is “whose fault?”, people hide problems instead of surfacing them. And by alert fatigue: Jenny, the FinOps-curious junior, has been flagging anomalies in Cost Explorer for months. She was right about Big Bertha in March. Nobody listened, because the dashboards cry wolf so often that a real wolf reads as noise, and because Jenny is junior and the org has quietly normalized the waste as just “what the bill is.”
How to address it: reward prevention as loudly as you reward heroics. Run blameless postmortems so problems come out into the light. Tune alerts so the ones that fire actually mean something β and when your Jenny is right three times in a row, promote her, don’t ignore her.
The Governance Vacuum
Nobody can see their own bill
At Nimbus, no team could see what they cost. The bill arrived as one undifferentiated $340K boulder, owned by everyone and therefore no one. You cannot ask people to manage a number they never see.
How to address it: tag everything, then implement showback (here’s what your team cost) or chargeback (and here’s the invoice). Visibility alone changes behavior β teams optimize what they can see attributed to their own name.
Decisions made once, never revisited
The always-on schedule for PhotoBlob’s fleet was set in 2023, for a traffic pattern that no longer exists. Nobody chose to keep it running at 3 a.m.; they just never chose not to. Decisions calcify into facts.
How to address it: put a review cadence on the big ones. Quarterly, ask of every major cost: “if we were deciding today, would we decide this?”
Analysis paralysis and the approval swamp
The flip side of never-revisiting is never-deciding. Priya’s first proposal to right-size Big Bertha sat in an approvals queue for five weeks β three sign-offs for a change that saved $9K/month and could be undone in ten minutes.
How to address it: delegate with guardrails. For reversible decisions, default to yes and let engineers act. Save the ceremony for the one-way doors. An approval process that costs more than the thing it’s approving is itself waste.
Optimization as a project, not a practice
The deepest governance failure: Nimbus treated cost work as a project β a heroic quarter, a spreadsheet, a victory lap, then back to normal until the CFO panics again. Waste doesn’t work in sprints. It accrues continuously, so it has to be fought continuously.
How to address it: make it a standing practice. Continuous FinOps β a small recurring rhythm, not an annual crisis.
Skills, Silos, and the Translation Gap
Half of Nimbus’s “cloud architecture” is really a data-center mindset transplanted to AWS: lift-and-shift habits, always-on servers, capacity provisioned for a peak that arrives twice a year. That’s a skills gap, and it’s fixable with training, not blame β nobody unlearns fifteen years of on-prem instinct by osmosis.
Then there’s key-person risk: only Deepak understands PhotoBlob, only Marcus understands Platform 2.0’s wiring. Knowledge silos make every system too scary to change and every vacation a liability. The fix is unglamorous β docs, pairing, and deliberate rotation so no single brain is load-bearing.
And there’s the translation gap: engineers talk in vCPUs and IOPS; the CFO talks in margin. They are describing the same reality and cannot hear each other. Unit economics is the shared language β “this change improves gross margin by 1.5 points” lands in a boardroom in a way “we downsized the r5 fleet” never will.
Finally, vendor over-reliance: Nimbus pays a consultancy handsomely, and the consultancy is paid by the hour β which quietly rewards complexity, not thrift. Align those contracts to outcomes (a share of realized savings) and watch the recommendations get simpler.
Structure and Politics
Conway’s Law is having a field day here. Nimbus’s system mirrors its org chart, and the org chart is a mess, so there are three services doing image resizing because three teams each built their own rather than share. Turf wars over shared resources mean nobody will adopt a common component they didn’t own.
Above it all sits short-termism β next-quarter thinking that never funds the boring efficiency work β and reorg churn, which reassigns ownership so often that no one holds a system long enough to care about its bill.
How to address it: fund a standing efficiency-and-tech-debt allocation so the long game has a budget. Keep ownership stable long enough for pride to form. And occasionally, reshape a team to fix a redundant service β you can run Conway’s Law forwards instead of letting it run you.
Takeaways
- Separate the person from the design β retire flawed architecture without demoting its author; make simplification a win Marcus gets credit for.
- Kill sunk-cost logic β “we invested two years” is grief, not strategy; the past spend is gone either way.
- Cure not-invented-here β prefer managed services over homegrown when they’re cheaper and lower-toil.
- Require an ADR with “why not the simpler option?” β to catch rΓ©sumΓ©-driven Kubernetes/Kafka/microservices before they ship.
- Change the scoreboard to unit economics β measure cost per user / transaction / GB, not gross spend; stop bragging about the total.
- Name a cost owner β a FinOps function, so cost stops being nobody’s job.
- Reward returned budget, not burned budget β kill use-it-or-lose-it and empire-building by headcount.
- Make improvement part of the job β allocate time for refactors and reward them like features (Deepak’s PhotoBlob included).
- Build the safety net first β tests, canary, rollback β so “don’t touch it” fear stops protecting waste.
- Reward prevention over firefighting β retire the hero culture; celebrate the incident that never happened.
- Run blameless postmortems β so problems surface instead of hiding.
- Fix alert fatigue β tune signals so real anomalies stand out, and listen when your Jenny is right.
- Give teams cost visibility β tag everything, then showback/chargeback so each team sees its own number.
- Put a review cadence on standing decisions β “would we decide this today?”
- Delegate with guardrails β default-yes for reversible changes; save approvals for one-way doors.
- Treat FinOps as a continuous practice, not a project β waste accrues continuously, so fight it continuously.
- Close the skills gap β train away the always-on on-prem mindset.
- Reduce key-person risk β docs, pairing, rotation.
- Talk in unit economics β bridge the engineering/business translation gap.
- Align vendor contracts to outcomes β so consultants are paid for savings, not hours.
- Run Conway’s Law on purpose β reshape teams to eliminate redundant services and turf wars.
- Fund a standing efficiency and tech-debt allocation β and keep ownership stable through reorg churn.
The closing insight: almost every problem in this post reduces to just two root causes β misaligned incentives and invisible cost. Every villain in the series is really one of those two wearing a costume. So the two highest-leverage fixes aren’t technical at all. First, make cost visible and owned: showback, unit-economics dashboards, a named FinOps owner, anomaly alerts that people trust. Second, change what gets rewarded: efficiency, prevention, and simplicity β measure cost-per-unit instead of gross spend, and let teams keep a share of the savings they verifiably create. Get those two right and the individual behaviors start correcting themselves β not because anyone became a better person, but because, at long last, doing the right thing finally pays. Chad will still wear the vest. He’ll just brag about a different number.
Epilogue β Ninety Days Later
The second “AWS Spend” meeting was also on a Tuesday, because the universe enjoys symmetry.
This time Priya presented. The bill had gone from $340K/Month to $196K/Month. That is a 42% cut β and, crucially, Nimbus was serving more traffic than it had in the quarter before. The line that made the CFO smile wasn’t the gross number, though. It was the one Priya had learned to care about - cost per monthly active user had dropped from $0.0085 to $0.0049. They were serving each person for a little over half of what it used to cost.
Nobody had turned off anything customers noticed. There were no outages. The checkout path still had its multi-AZ redundancy, the critical database still had its standby, and the on-call rotation actually slept better than before, because half of the fragile, hand-tuned, snowflake infrastructure had been replaced with things that healed themselves.
Here’s what actually happened, and in roughly what order β because the order matters more than people expect.
The receipts
The quick, embarrassing wins came first. Big Bertha β the m5.24xlarge running a 3%-CPU cron β became a scheduled t4g.medium and saved more than most people’s salaries. Priya deleted 40 TB of orphaned EBS snapshots, released a dozen unattached Elastic IPs that had been billing hourly for a year, and flipped every gp2 volume to gp3 for the same performance at ~20% less. None of it required a meeting. All of it should have happened two years ago. This is the recurring lesson of cost work: the first 20% is almost free and slightly humiliating.
Then the structural stuff. The NAT gateway from hell got Gateway VPC Endpoints for S3 and DynamoDB β which are free β and its data-processing charge fell off a cliff. Right-sizing compute off Compute Optimizer’s recommendations, plus moving the fleet to Graviton, took a big bite. S3 lifecycle policies started aging cold photos down to Glacier instead of paying Standard rates to store a selfie from 2019 that no one will ever open again.
Then the commitments β but only after right-sizing. Priya covered the now-smaller steady baseline with Compute Savings Plans, deliberately not the peak, and let Spot and on-demand absorb the spikes. Buying commitments before right-sizing would have locked in the waste; she’d have been proud of a discount on things she was about to delete.
Then the slow, valuable refactors. PhotoBlob is still mostly a monolith β that’s fine, a well-run monolith is cheaper than a badly-run mesh of microservices β but its worst full-table-scan queries got indexes, its session state moved to Redis so the app tier could finally autoscale, and Marcus’s 3-broker MSK cluster serving 500 messages a day was quietly replaced with an SQS queue that costs approximately nothing. These didn’t land in week one. They landed over the quarter, on a standing “efficiency and debt” allocation that got its own slice of every sprint.
The two things that made it stick
If you take nothing else from the entire series, take these, because they’re the reason the savings didn’t evaporate the moment Priya looked away.
One: make cost visible and owned. Every team at Nimbus now sees its own spend, thanks to a tagging policy Priya enforced with the enthusiasm of a customs officer. There’s a dashboard showing cost-per-MAU, not just gross dollars. There are budgets and Cost Anomaly Detection alerts with a named human attached to each. Cost stopped being “finance’s problem” or “the platform team’s problem” and became everyone’s, visibly. You cannot fix what nobody can see and nobody owns.
Two: change what gets rewarded. This was the hard one, and it was Chad’s to solve, not Priya’s. The scoreboard changed. Nimbus stopped celebrating the size of the AWS bill and started celebrating unit economics. Teams that shipped efficiency wins got the same visible credit as teams that shipped features. And β the masterstroke β teams got to keep and redeploy a share of the savings they verified, which turned “cut the bill” from a chore imposed from above into a game people actually wanted to win. Chad, to his credit, updated his conference talk. The new version is about margin. It got fewer laughs and more job offers.
The prioritization framework, on one napkin
For anyone starting their own ninety days, here’s the order Priya wishes she’d known on day one:
- Kill waste. Orphaned resources, idle instances, unattached everything. Turn on Compute Optimizer, Trusted Advisor, budgets, and anomaly alerts. Free, fast, no risk.
- Grab the quick wins. gp2βgp3, NATβGateway Endpoints, non-prod scheduling, Graviton migration.
- Right-size compute and databases from real utilization data.
- Commit the stable baseline β never the peak β to Savings Plans, and only after right-sizing.
- Go structural. Networking topology, redundancy you don’t need, caching, and over-engineering. Durable savings, a bit more effort.
- Refactor the debt. Statelessness first (it unlocks autoscaling and Spot), then concurrency, then data-access patterns, then decoupling β scored by blast radius Γ recurring cost Γ effort, burned down over quarters.
- Fix the humans. Visibility and incentives. Without this, you’ll do steps 1β6 again next year.
Score each candidate by how much it costs you every month times how easy it is to fix, and start at the top of that list. Boring, repeatable, and it works.
Last word
Nimbus’s bill will creep back up β that’s what bills do, because businesses grow and entropy is undefeated. The difference is that now someone is watching, everyone can see, and the person who finds the next Big Bertha will get a high-five instead of a shrug.
Priya still doesn’t love being the FinOps lead. But she’s made her peace with it, mostly because of a sentence she now repeats in every architecture review, usually while looking directly at Marcus: “We can absolutely build it that way. Let’s just be honest about what it costs β and whether anyone will ever need it at scale.”
He hasn’t got a comeback for that one yet.
β End of series. Go read your bill. Bertha’s counting on you.