본문으로 건너뛰기
← Back to Blog
테크

Why an Auto-Saved File Ends Up at 0 Bytes: Our AI Agents' Instruction File Sat Empty for Four Days

공유

I split my work between two AI coding agents, Claude Code and Codex. Before either one starts a task, it reads the same instruction file, AGENTS.md. That file tells them how to address me, what they must never do, and every fact I have had to correct for them in the past.

In August, that file was 0 bytes. A document of a little over 21,000 bytes had been emptied completely, and it stayed that way for four days. Ten days before that, the file holding the knowledge the agents had built up went from 349 entries to 1.

Neither was a hack or a hardware failure. Both happened during an ordinary save, the kind a script performs dozens of times a day without anyone looking. If your automation rewrites files on its own, the same path may be open on your machine, so here is exactly how each one happened and what we changed.

An open binder with every page gone blank, beside a box of index cards stored one by one — an illustration made for this post
An open binder with every page gone blank, beside a box of index cards stored one by one — an illustration made for this post

Incident one: the disk was full, and the instruction file came out empty

At the bottom of the instruction file there is a corrections section. Whenever I tell an agent "that's wrong, this is right," a small program rewrites that section and saves the whole file again. The idea was that the instructions would grow by themselves without me editing them by hand.

The save used the most common pattern there is. Open the file, which empties it first, then write the new content in. Most of the time this is invisible, because emptying and writing happen in the same instant.

On August 7 the C: drive was full. Emptying a file needs no space. Writing the new content does. If the write fails halfway, what remains is the emptied file. When we lined up the records afterward, the evidence pointed to exactly this path.

What made it worse came next.

  • An empty file raises no error. Whatever reads it simply concludes there are no instructions and carries on.
  • That noon, the empty file went into git as part of a snapshot commit bundling several other files. A commit does not check whether a file has suddenly lost everything.
  • For the next four days, both agents started their work from a shared instruction file with nothing in it.
  • On August 11 the file was restored from the last good version, saved on August 6.

    Incident two: several jobs, one file, and 349 entries became 1

    On the evening of July 28, three jobs were running in the same project folder at once: two Claude sessions and one Codex session. All the knowledge the agents had accumulated lived in a single file. The failure ran through five links.

  • While pulling the latest version from git during that parallel work, conflict markers ended up inside the knowledge file.
  • The code that loads the file could not parse it. It logged one warning and carried on with an empty graph.
  • That session added one new entry.
  • The save guard said "save if there is more than zero entries." One entry passed, and it overwrote the 349-entry file with a 1-entry file.
  • Failures in the save step were also set to pass silently, so the result flowed straight into a commit.
  • Break any one of those five links and nothing would have been lost.

    The way back was in git history. We pulled every past version of that file still stored in the repository, sorted them by entry count and save time, and found the newest version that was also the largest. Then we merged in only the entries created after the incident, matched by ID rather than overwritten. Nothing was lost, and the file came back with 351 entries.

    What we changed: guards first, then smaller units of storage

    The first fix was a pair of guards.

  • If the file failed to load at startup, saving is forbidden for that session.
  • If a save would shrink the entry count by more than 10 percent compared with what is on disk, the save is refused, and the refused content is written to a separate file so nothing disappears. Deliberately clearing the file requires an explicit override.
  • Guards stop the damage, but they leave the real cause in place: many writers sharing one file. That evening I asked whether we could change the structure so parallel jobs would be safe by design. We changed how knowledge is stored.

  • One entry, one file. Instead of 349 entries in one document, each entry is its own small card on disk.
  • A save only updates a card if it exists or adds it if it does not. Cards a session does not have in memory are never touched.
  • Deleting happens only through an explicit delete command.
  • With that, a session that starts with an empty view cannot erase anyone else's work, structurally rather than by luck. It turned out to be simpler to remove the thing jobs were fighting over than to build a better lock.

    For the instruction file, we changed the order of saving.

  • New content is written to a temporary file first, never directly over the original.
  • The temporary file is read back and compared, and only then swapped in for the original. If anything fails along the way, the original stays intact.
  • If the content to be written is empty, nothing is written at all.
  • If the instruction file is already empty, the update is skipped, because writing over an empty file would lock the loss in.
  • Every session now checks at startup whether the instruction file is 0 bytes. That check no longer depends on anyone remembering to look.
  • Five checks for any file your automation rewrites

    This applies the same way to an AI agent's instruction file, a config file a scheduled job updates, or a data file a script edits every night.

  • How does it save? Does it write directly over the original, or write a temporary file and swap it in? In Python, the baseline is to write to a temp file in the same folder and then call os.replace. Writing straight over the original leaves an empty file the moment something fails.
  • Refuse empty content. If what you are about to write is empty, stop. A healthy update almost never needs to produce an empty file.
  • Compare against the last state, not against zero. A rule like "save if count is greater than zero" is beaten by a count of one. Use "stop if this drops by more than X percent from what was there."
  • A failed read should stop the job loudly. Code that fails to read a file and continues with an empty state is, in practice, code that deletes. Do not let it off with a warning.
  • If several writers share a file, split it. When multiple agents or scheduled jobs write the same file, see whether it can be broken into separate units before you try to coordinate their timing.
  • One more thing. Both times, version history is what got us back. A commit carried the empty instruction file into the repository, but the earlier commits were exactly what we restored from. If a file is rewritten automatically, keeping it somewhere that records every change is the cheapest insurance you can buy.

    Closing

    Both incidents started with a single save that looked like nothing. The more you automate, the less often a person opens these files and looks inside. So instead of trusting that a save probably worked, we built in a save order that leaves the original standing when something fails, and a check that speaks up when a file comes back empty.

    How we confirm an automation's output all the way to the public page is in An Automation Is Not Finished When It Saves. The day blog photos stored inside the site's code took the whole site down is in How Blog Images Took Our Vercel Site Down.

    Services by Botonglee

    Why Auto-Saved Files End Up at 0 Bytes and How to Prevent It | 보통리