본문으로 건너뛰기
← Back to Blog
테크

How Blog Images Took Our Vercel Site Down

공유

On the morning of September 17, our entire site stopped loading. The home page, every blog post, and the studio booking pages all returned a 404, and the error name on the screen was DEPLOYMENT_NOT_FOUND.

It wasn't a hack, and nothing had crashed. The cause was something far more ordinary: where we were keeping the blog's images. If you run a site the same way, here is what happened, what tripped us up during recovery, and what we changed afterward.

A laptop on a desk with a tall, leaning stack of photos and boxes behind it
A laptop on a desk with a tall, leaning stack of photos and boxes behind it

What happened

The site is a Next.js app deployed on Vercel's free Hobby plan. Posts live in a database, but a large share of the images lived inside the site's code repository, in the public uploads folder. Most of them had come over from our old WordPress blog, and together they added up to roughly 1.5GB.

Here is the part we hadn't thought through: every time Vercel deploys, it copies that whole folder into the new deployment. One deploy meant about 1.5GB of new stored files.

Then automation sped everything up.

  • September 1 — We reworked our blog automation and rewrote the publishing steps. One line in the new steps said, in effect, "copy the image into the site folder and commit it." Before that, images had gone to a separate image storage bucket.
  • September 1 to 17 — That one line produced 26 image commits and 41 deploys. In all of August there had been 4 image commits.
  • Deployment Storage — climbed to 91GB, nine times the 10GB Hobby allowance. Looking back, it was already at 38GB on August 20. The number was there; nobody was looking at it.
  • September 16 — Vercel changed its Hobby policy: teams over 10GB of Deployment Storage can be blocked from deploying until they free up space. A structural problem that had been quietly growing for months became visible that day.
  • September 16 to 17 — While old deployments were being cleared out to get under the limit, the live production deployment was deleted along with them. By the morning of the 17th, the whole site was down.
  • Two traps during recovery

    Trap one: deleting does not lower the number right away. After 80 deployments were deleted, the storage graph barely moved: 93.02GB on September 15, 91.58GB on the 16th, 85.51GB on the 17th. Deleted deployments sit in "Recently Deleted" for about 30 days, still restorable, and they kept counting against storage the whole time. There was no button to delete them permanently. Our assumption that "once it's deleted, the block lifts" was simply wrong.

    Trap two: restoring the deployment does not bring the site back. The deleted production deployment was restored, and the site still returned 404. The deployment was back, but its domain assignments were not: the www address, the bare domain, and the default project addresses, four in all. Once those four domains were attached to the production deployment again, the site came back immediately, at 10:36 p.m. that night. Because this was a re-link and not a new deploy, the storage block didn't stand in the way.

    So now, after any cleanup or restore, we open the real public address and check it. A dashboard that says "restored" is not the same as a page your visitors can load.

    The three changes we made

    Fixing one line of instructions didn't feel like enough. One line of instructions was exactly how this started. So we put three layers in place.

    1. Process: images go to object storage only. Every image for the blog and social posts now goes to Cloudflare R2, and posts reference that address. We also standardized on a single upload tool. If a file with the same name already exists, it stops instead of overwriting. After uploading, it checks that the public address actually loads and that the file size matches.

    2. Guardrails: the machine blocks the mistake. The site repository now rejects any commit that adds files to the uploads folder. And our automated Friday check reports three numbers every week: the size of that folder, the number of image commits that week, and the total number of commits that week. Anything unusual triggers an alert. The point is to avoid another "38GB on August 20 that nobody saw."

    3. Structure: the images left the site. On September 29, we removed the uploads folder from the repository entirely.

  • The 703 images that posts actually use were moved to object storage under the exact same paths.
  • The 2,693 images nothing used were deleted. Most were the resized copies WordPress generates automatically, plus old media-library files.
  • We didn't edit a single image address in the old posts. Instead, the site configuration redirects any address starting with /uploads/ to the object storage bucket. That was much safer than editing more than 80 older posts one by one.
  • We compared the images on 419 pages before and after the move. Newly broken images: zero.
  • The site's public folder is now about 27MB. Each deploy carries tens of megabytes instead of 1.5GB.

    Five checks if your setup looks like ours

    Do your images and videos live alongside your site's code? If your host packages and stores your full site files on every deploy, then as long as images sit inside them, they get duplicated once per deploy. It looks harmless when the folder is small. It stops being harmless as posts pile up.

    How much has automation increased your deploy count? Publishing by hand might mean a few deploys a week. Add automation and it can become several a day. The same structure behaves very differently at a different frequency.

    Are you watching the by-products, not just the output? "Did today's post go up?" is easy to check every day. What piled up alongside that post, such as storage, deploy count, and cost, stays invisible unless you go looking. Three numbers once a week is enough.

    Would you know if your free tier's rules changed? In our case the site didn't change; the rules did, and that is the day things broke. It is worth having at least one way to see a provider's changelog or policy announcements.

    Before cleaning up, did you mark what must not be deleted? Freeing up storage tends to happen in a hurry. Before starting, write down what must never be touched, such as the live production deployment. When you finish, load the real address and confirm it works.

    The short version

    Automation repeats its instructions exactly. In our case, one line in a publishing procedure turned into 41 deploys in 17 days. Since then, whenever we review an automation, we ask two questions instead of one: "Did the output appear?" and "What else piled up along the way?"

    If you're setting up similar workflows, we've also written about checking the live result, not just the save and which automations we kept and which we turned off.

    Services by Botonglee

    Vercel Hobby Storage Limit: How Images Took Our Site Down | 보통리