Technically
AI Reference
Your dictionary for AI terms like LLM and RLHF
Company Breakdowns
What technical products actually do and why the companies that make them are valuable
Learning Tracks
In-depth, networked guides to learning specific concepts
Posts Archive
All Technically posts on software concepts since the dawn of time
Terms Universe
The dictionary of software terms you've always wanted

Explore learning tracks

AI, it's not that ComplicatedAnalyzing Software CompaniesBuilding Software ProductsWorking with Data Teams
Loading...
I'm feeling luckyPricing
Log In

The Vibe Coder's Guide to Cloud + Scaling

How to survive your app going viral.

Last updated Oct 6, 2026devops
Will Raphaelson
Will Raphaelson
Read within learning track:Software Engineering for Vibe Coders

Lots of you are probably building apps with AI, and we want to do our part to make sure those apps can thrive when they make contact with real users.

This 6-part series, Software Engineering for Vibe Coders, covers Art Vandelay’s sisyphean journey to ship an app to manage Vandelay Industries’ potato chip import & export business.

Thanks to our friends at Railway, who help you peacefully deploy those vibe coded apps, for sponsoring this series.

You can spin up anything your vibe coded app might need at railway.com/new.

ICYMI, check out:

  • Part 1 (on frontends + backends)
  • Part 2 (on databases + storage)
  • Part 3 (on auth + security)
  • Part 4 (on deployment)

The TL;DR#

Part 4 got your vibe coded app deployed into the cloud, AKA a computer that isn’t your laptop. This allows users other than you to actually access it, unless you want to open up your computer to public internet traffic, which, trust me on this, you do not. This part is about what's actually happening in your little slice of this cloud thing, and what happens when your app takes off and you outgrow that slice:

  • The mechanics of supporting many many users using your app at once, and where things tend to break
  • Scaling limits, and what happens when you hit them
  • How to scale your app, and the difference between vertical and horizontal scaling
  • Scaling with Railway, which makes it a breeze to keep your app alive and snappy, even under heavy load.

Terms Mentioned

Frontend

Server

Cloud

Infrastructure

Backend

Machine Learning

Database

Previously, at Vandelay Industries#

To recap, Art Vandelay is vibe coding the app that runs Vandelay Industries’ potato chip import/export empire, or as we know it in the Godfather, the family business.

Loading image...

So far in this series, we’ve explained every piece of how Art gets to a working app:

  1. In Part 1 we split it into a frontend and a backend.
  2. Part 2 gave it a place to store data: a Postgres database on Railway.
  3. Part 3 locked the doors so that Newman can’t read the shipping schedule by adding in auth.
  4. Part 4 put it in the cloud and gave it a real public URL.

Art has stopped saying “well, it works on my machine”, which is growth (literally). But there’s danger on the horizon.

Unbeknownst to our proprietor, Kramerica Industries mentions Art’s app in a popular trade newsletter, and it goes viral: 4000 potato chip brokers frantically click the link to get in on the action. Good news for Art…or is it?

At first, brokers are getting through just fine, but then the app starts taking a while to load. 2 seconds, 5 seconds, 10 seconds, and then, finally:

Loading image...

The page stops working completely. Brokers lose patience and stop trying…and just like that, Art has fumbled the growth spike of a lifetime.

This exact story (well, mostly) happens to fast growing apps every week: an unexpected growth spike overloads your app, takes it down, and you lose your chance. More mechanically, this is what happens when the slice of the cloud you rented is too small for the traffic coming to your app. Nothing was hacked, the code wasn’t broken, we’ve just reached our cloud’s limits.

How could this happen to us? How can we avoid this scenario and make sure we’re prepared for demand spikes? And what's the deal with the cloud anyway?

Scaling limits in the cloud#

We’ve covered what the cloud is before at Technically, most recently in Part 4. When you use the cloud, you are renting a piece of a giant server in a data center. In most cases, you share this giant server with many, possibly hundreds, of other apps like yours.

There are tons of different possible ways to rent, just like apartments. The provider we used, Railway (who graciously sponsored this series), is the furnished-apartment version of that deal. Its batteries included, integrated with Github, creates a public domain for external access, and comes with observability included.

But slices, rooms, whatever metaphor you choose, come in different sizes, and getting that size right is important. Get the slice too small, and your servers can’t handle the traffic, get it too large, and you pay for capacity you don’t use. This problem is so pervasive that the infrastructure community has turned it, impossibly, into a verb: “right-sizing”.

In New York City it is fairly easy to right size your apartment (2nd bedroom if you can afford it). In the cloud, less easy.

When Art put his app on Railway’s hobby tier, he signed up for a small slice of the cloud, a certain fixed amount of CPU (compute power), memory, and bandwidth (moving data around) that he for sure clicked through quickly without reading. Railway bills by the minute based on what’s actually used, but the plan still has a ceiling; a maximum amount of computational resources it will allocate to the app. When all those brokers tried to log in simultaneously, the allocated servers’ capacity was reached, there was no compute left to go around, and performance tanked. There were too many people trying to live in the apartment.

Loading image...
Loading image...

The fact that a demand spike caused us to hit capacity constraints doesn’t mean that Art picked the wrong plan; when he started, he wasn’t anticipating this many users. What it does mean is that he needs to make some changes, and quick.

So what do we do when our application traffic overflows our servers?

We scale it.

Scaling Vandelay Industries#

Scaling is a broad term, but by scaling we mean upping the resources available to an application to handle more, or more complex, work. There are many ways to scale, some better than others.

Scaling can occur in two “directions”:

  • Vertical Scaling
  • Horizontal Scaling

You vertically scale infrastructure by upping the resources available to your workloads – in other words, beefing up the size and power of your existing server. You can think of this as simply moving into a larger apartment.

Horizontal scaling, on the other hand, is where you scale infrastructure by adding more instances of your workloads. You can think of this as keeping your existing apartment, but also renting another one across the hall for you to use as an office and/or Netflix depression den. The size of the apartment(s) in question stays the same, but the amount of them goes up.

Both of these methods have their time and place, neither is better than the other in an objective sense. You can only size up your server so much before it’s too big and expensive – they don’t make 7 bedroom apartments here. And having multiple different apartments introduces coordination problems between them.

Loading image...

For Art’s use case – a run of the mill web application with a database – vertical scaling to a larger server would be the right move at this point. And if Art were to vertically scale, his problems would go away, but a new one would come up: cost. See, more powerful servers cost more money.

In the old days, before the cloud, this was your only option. You needed to buy and run infrastructure that could accommodate your peak demand – it needed to be big enough to support your most demanding possible time, even if it’s wildly overpowered for your normal day to day activity.

But the modern cloud introduced the idea of elastic scaling - computational real estate that can automatically grow to meet demand, and shrink back down when the party’s over. Without elastic scaling, it would be like buying a mansion to throw a banger of a house party, and then paying for it to sit empty for the rest of the mortgage. Art aint broke, but Art aint got mansion money.

Loading image...

Art’s lucky because his platform of choice, Railway, automatically vertically scales his app’s memory and CPU and then only bills him for what he actually uses. Were Art’s app to get more complex and call for horizontal scaling, Railway allows him to create multiple replicas in multiple regions, allowing more concurrent users in more places to sign up and place orders.

The problem for Art is that he reached the ceiling of the hobby tier. This might sound like a small detail, but it’s a common mistake vibe coders make with their apps. You need to pay attention to plans and billing!

Scaling with Railway#

The hobby tier on Railway comes with up to 48 vCPU and 1GB RAM per service, which is our vertical ceiling, and limits him to 50 services including replicas, which is our horizontal ceiling (er… wall?).

For the purposes of this discussion, let’s say we need to address both ceilings 👍

  • Registering new import/exporters is computationally intensive because it kicks off a recommendation machine learning algorithm to predict chip choice -- requiring us to vertically scale our machines.
  • Additionally, the sheer number of concurrent users bombarding our application requires us to scale out, or horizontally, via more and more replicas of our service across the globe.

By upgrading to the pro tier, we automatically handle our vertical limit: Railway will now, with no additional configuration required, scale up our CPU and RAM in response to real demands on the system, up to 1000 vCPU and 1TB RAM, more than enough to handle the new traffic.

Critically, it will also scale things back down after the hype is over. Our allocation and utilization will look something like this:

Loading image...

Because Railway bills by the actual utilization of resources, not what’s allocated, you only pay for the actual workload (the blue line).

If ever Art wants to set limits for how much each service can consume, we can do so easily in his service’s settings:

Loading image...

The horizontal scaling process is less automatic because it requires greater judgement on part of the developer. We horizontally scale via replicas, exact copies of our application, over which the work of incoming requests can be shared. Art can create replicas easily by changing the number of instances next to his service.

Loading image...

But wait, what are those flags I see?

Art can deploy replicas in different regions around the globe, and he might want to do this for two reasons: response time and resiliency.

Response time gets quicker if a copy of the application is served closer to where users live and work. For Art’s global import/export operation, this is critically important. If a user in Italy, desperate to import Art’s fine chips to his home city of Milan, had to wait for his request to go all the way to Virginia and back, it might take upwards of a second, which feels sluggish and amateurish. By deploying a replica to the Netherlands, the app stays snappy and responsive.

By resiliency, we mean a software's ability to take damage and keep working. A non-resilient app is one where a single issue can take it down completely. If Art’s app is only deployed in one region and that region has an outage, which is rare, but does happen, then his users are locked out. If, on the other hand, he deploys replicas to multiple regions, Railway will automatically route requests to the next best region if one region goes down.

Where that leaves Art (and you)#

Practically, as a vibe coder, it’s unlikely that you’ll need to get into the weeds of colocating servers and configuring replicas. By the time your app gets popular enough for these kinds of champagne problems, you will hopefully have hired someone who is an actual expert in how the cloud works…they can do this for you. But knowing the basics of what’s available to you – and how to understand billing plans for your vibe coding platform of choice – will help you avoid a demand spike that leaves opportunity on the table.

So, where are we? The app now lives in a data center Art will never visit, on machines he rents by the second, in regions close to his prized users. When a trade newsletter, potato shortage overseas, or viral potato meme drives traffic to his site, he can make the slice bigger or make more slices, keeping his app alive and his business flowing.

But how does he know when to scale? In this case, Art found out the app was slow because Kramer called him. He reactively scaled his app up, but only after the app crashed and eager importer/exporters decided to take their business elsewhere. He’s flying the plane safely by listening for screams.

For Vandelay industries, and for any vibe coder serious about the success of their app, this is unacceptable.

To go from reactive scaling to proactive scaling, firefighting to fire prevention, we need to crack open a shiny new can of worms: observability. The data, dashboards, and alerts that show you inside the black box, before it's too late.

That’s Part 6.

Up Next
The Vibe Coder's Guide to DeploymentFree

How your app gets off your laptop and onto the real internet, and what to do when a deploy breaks.

Software Eng for Vibe Coders: On Frontends + BackendsFree

A new series to help non-engineers build products that can handle going viral.

Software Eng for Vibe Coders: Auth + SecurityFree

How to make sure anyone on the internet can't just delete your app's data, and other security considerations.

Content
  • All Posts
  • Learning Tracks
  • AI Reference
  • Companies
  • Terms Universe
Company
  • Pricing
  • Sponsorships
  • Contribute
  • Contact
Connect
SubscribeSubstackYouTubeXLinkedInInstagram
Legal
  • Privacy Policy
  • Terms of Service

© 2026 Technically.