Previously, at Vandelay Industries#
To recap, Art Vandelay is vibe coding the app that runs Vandelay Industries’ potato chip import/export empire, or as we know it in the Godfather, the family business.
So far in this series, we’ve explained every piece of how Art gets to a working app:
- In Part 1 we split it into a frontend and a backend.
- Part 2 gave it a place to store data: a Postgres database on Railway.
- Part 3 locked the doors so that Newman can’t read the shipping schedule by adding in auth.
- Part 4 put it in the cloud and gave it a real public URL.
Art has stopped saying “well, it works on my machine”, which is growth (literally). But there’s danger on the horizon.
Unbeknownst to our proprietor, Kramerica Industries mentions Art’s app in a popular trade newsletter, and it goes viral: 4000 potato chip brokers frantically click the link to get in on the action. Good news for Art…or is it?
At first, brokers are getting through just fine, but then the app starts taking a while to load. 2 seconds, 5 seconds, 10 seconds, and then, finally:
The page stops working completely. Brokers lose patience and stop trying…and just like that, Art has fumbled the growth spike of a lifetime.
This exact story (well, mostly) happens to fast growing apps every week: an unexpected growth spike overloads your app, takes it down, and you lose your chance. More mechanically, this is what happens when the slice of the cloud you rented is too small for the traffic coming to your app. Nothing was hacked, the code wasn’t broken, we’ve just reached our cloud’s limits.
How could this happen to us? How can we avoid this scenario and make sure we’re prepared for demand spikes? And what's the deal with the cloud anyway?
Scaling limits in the cloud#
We’ve covered what the cloud is before at Technically, most recently in Part 4. When you use the cloud, you are renting a piece of a giant server in a data center. In most cases, you share this giant server with many, possibly hundreds, of other apps like yours.
There are tons of different possible ways to rent, just like apartments. The provider we used, Railway (who graciously sponsored this series), is the furnished-apartment version of that deal. Its batteries included, integrated with Github, creates a public domain for external access, and comes with observability included.
But slices, rooms, whatever metaphor you choose, come in different sizes, and getting that size right is important. Get the slice too small, and your servers can’t handle the traffic, get it too large, and you pay for capacity you don’t use. This problem is so pervasive that the infrastructure community has turned it, impossibly, into a verb: “right-sizing”.
In New York City it is fairly easy to right size your apartment (2nd bedroom if you can afford it). In the cloud, less easy.
When Art put his app on Railway’s hobby tier, he signed up for a small slice of the cloud, a certain fixed amount of CPU (compute power), memory, and bandwidth (moving data around) that he for sure clicked through quickly without reading. Railway bills by the minute based on what’s actually used, but the plan still has a ceiling; a maximum amount of computational resources it will allocate to the app. When all those brokers tried to log in simultaneously, the allocated servers’ capacity was reached, there was no compute left to go around, and performance tanked. There were too many people trying to live in the apartment.
The fact that a demand spike caused us to hit capacity constraints doesn’t mean that Art picked the wrong plan; when he started, he wasn’t anticipating this many users. What it does mean is that he needs to make some changes, and quick.
So what do we do when our application traffic overflows our servers?
We scale it.
Scaling Vandelay Industries#
Scaling is a broad term, but by scaling we mean upping the resources available to an application to handle more, or more complex, work. There are many ways to scale, some better than others.
Scaling can occur in two “directions”:
- Vertical Scaling
- Horizontal Scaling
You vertically scale infrastructure by upping the resources available to your workloads – in other words, beefing up the size and power of your existing server. You can think of this as simply moving into a larger apartment.
Horizontal scaling, on the other hand, is where you scale infrastructure by adding more instances of your workloads. You can think of this as keeping your existing apartment, but also renting another one across the hall for you to use as an office and/or Netflix depression den. The size of the apartment(s) in question stays the same, but the amount of them goes up.
Both of these methods have their time and place, neither is better than the other in an objective sense. You can only size up your server so much before it’s too big and expensive – they don’t make 7 bedroom apartments here. And having multiple different apartments introduces coordination problems between them.
For Art’s use case – a run of the mill web application with a database – vertical scaling to a larger server would be the right move at this point. And if Art were to vertically scale, his problems would go away, but a new one would come up: cost. See, more powerful servers cost more money.
In the old days, before the cloud, this was your only option. You needed to buy and run infrastructure that could accommodate your peak demand – it needed to be big enough to support your most demanding possible time, even if it’s wildly overpowered for your normal day to day activity.
But the modern cloud introduced the idea of elastic scaling - computational real estate that can automatically grow to meet demand, and shrink back down when the party’s over. Without elastic scaling, it would be like buying a mansion to throw a banger of a house party, and then paying for it to sit empty for the rest of the mortgage. Art aint broke, but Art aint got mansion money.
Art’s lucky because his platform of choice, Railway, automatically vertically scales his app’s memory and CPU and then only bills him for what he actually uses. Were Art’s app to get more complex and call for horizontal scaling, Railway allows him to create multiple replicas in multiple regions, allowing more concurrent users in more places to sign up and place orders.
The problem for Art is that he reached the ceiling of the hobby tier. This might sound like a small detail, but it’s a common mistake vibe coders make with their apps. You need to pay attention to plans and billing!
Scaling with Railway#
The hobby tier on Railway comes with up to 48 vCPU and 1GB RAM per service, which is our vertical ceiling, and limits him to 50 services including replicas, which is our horizontal ceiling (er… wall?).
For the purposes of this discussion, let’s say we need to address both ceilings 👍
- Registering new import/exporters is computationally intensive because it kicks off a recommendation machine learning algorithm to predict chip choice -- requiring us to vertically scale our machines.
- Additionally, the sheer number of concurrent users bombarding our application requires us to scale out, or horizontally, via more and more replicas of our service across the globe.
By upgrading to the pro tier, we automatically handle our vertical limit: Railway will now, with no additional configuration required, scale up our CPU and RAM in response to real demands on the system, up to 1000 vCPU and 1TB RAM, more than enough to handle the new traffic.
Critically, it will also scale things back down after the hype is over. Our allocation and utilization will look something like this:
Because Railway bills by the actual utilization of resources, not what’s allocated, you only pay for the actual workload (the blue line).
If ever Art wants to set limits for how much each service can consume, we can do so easily in his service’s settings:
The horizontal scaling process is less automatic because it requires greater judgement on part of the developer. We horizontally scale via replicas, exact copies of our application, over which the work of incoming requests can be shared. Art can create replicas easily by changing the number of instances next to his service.
But wait, what are those flags I see?
Art can deploy replicas in different regions around the globe, and he might want to do this for two reasons: response time and resiliency.
Response time gets quicker if a copy of the application is served closer to where users live and work. For Art’s global import/export operation, this is critically important. If a user in Italy, desperate to import Art’s fine chips to his home city of Milan, had to wait for his request to go all the way to Virginia and back, it might take upwards of a second, which feels sluggish and amateurish. By deploying a replica to the Netherlands, the app stays snappy and responsive.
By resiliency, we mean a software's ability to take damage and keep working. A non-resilient app is one where a single issue can take it down completely. If Art’s app is only deployed in one region and that region has an outage, which is rare, but does happen, then his users are locked out. If, on the other hand, he deploys replicas to multiple regions, Railway will automatically route requests to the next best region if one region goes down.
Where that leaves Art (and you)#
Practically, as a vibe coder, it’s unlikely that you’ll need to get into the weeds of colocating servers and configuring replicas. By the time your app gets popular enough for these kinds of champagne problems, you will hopefully have hired someone who is an actual expert in how the cloud works…they can do this for you. But knowing the basics of what’s available to you – and how to understand billing plans for your vibe coding platform of choice – will help you avoid a demand spike that leaves opportunity on the table.
So, where are we? The app now lives in a data center Art will never visit, on machines he rents by the second, in regions close to his prized users. When a trade newsletter, potato shortage overseas, or viral potato meme drives traffic to his site, he can make the slice bigger or make more slices, keeping his app alive and his business flowing.
But how does he know when to scale? In this case, Art found out the app was slow because Kramer called him. He reactively scaled his app up, but only after the app crashed and eager importer/exporters decided to take their business elsewhere. He’s flying the plane safely by listening for screams.
For Vandelay industries, and for any vibe coder serious about the success of their app, this is unacceptable.
To go from reactive scaling to proactive scaling, firefighting to fire prevention, we need to crack open a shiny new can of worms: observability. The data, dashboards, and alerts that show you inside the black box, before it's too late.
That’s Part 6.