Guide

Why AI Keeps Breaking Your App, and What Stops It

You asked for a dark mode toggle. Your assistant built it in a few minutes, and it works. You flip it a couple of times and close the laptop.

Next morning a friend tries to sign up and never gets the confirmation email. You did not ask anyone to touch signup. Nobody touched signup. Something did.

Why the new change breaks the old one

Your assistant edits what it can see. To add the toggle it opened the settings page, the theme file, and a small helper that reads the user's preferences. The helper looked untidy, so it tidied it. It now hands back its answer in a slightly different shape.

Signup reads that same helper to decide which email to send. That dependency is real, and it is written down nowhere the assistant would think to look. From inside the session, the change was correct. It was just not the whole picture.

This is not the model being careless. It is a capable pair of hands in a house it has never seen, with no way to tell which walls hold the roof up. Your prompt said what to add. It never said where the change should stop. And the usual next message, “fix it”, has no edge either.

What everyone tries first

Checkpoints, and they are right. Commit before every change and roll back when it goes wrong. But a checkpoint is an undo button. It says nothing about why the break happened, and you usually find the break a day later, three changes on, when rolling back also throws away the work you liked.

Then the warnings. “Do not touch login.” “Leave payments alone.” They help for exactly one prompt. They are gone when the session ends, and they only cover what you remembered to worry about. Nobody would have written “leave the preferences helper alone”, because nobody knew signup depended on it.

Both share the same gap. Before the change, nothing says what it is allowed to touch. After it, nothing checks whether it stayed there.

A plan says where the change stops

The first half of the fix is a short brief, written before the assistant starts. Not a spec. A few lines: what this change is for, what is out of scope, how you will check it is done, and which files it will probably touch. The done check should name one or two things that already work and must keep working.

For the toggle, that might read: out of scope, anything to do with accounts, email or the database. Done when the setting survives a refresh and a brand new signup still gets its confirmation email. Likely files: the theme file and the settings page.

Now the tidy-up is not a quiet improvement. It is a file nobody listed, in an area the brief ruled out, and it shows up as exactly that.

Then review the edges, not the feature

You were always going to try the toggle. You asked for it.

A review that only checks the new feature is checking the one part of the app you already knew to look at.

The second half of the fix is a review that starts at the edges. Compare the files that changed with the ones the brief expected, and look hard at anything outside that list. Ask the assistant to end every build by saying what it noticed on the way that looked broken or odd. The warning you need is often a sentence it would not have written unless asked.

Then make the call yourself: accept it, send it back with a note, or throw it away. Handing that verdict to the assistant that made the change is how a bug gets marked as fixed. When something does break, give the fix its own small brief instead of a “fix it” in the session that broke it. What happens to those notes after review is a separate subject, covered in Loop Engineering.

Where PAPI fits

All of this works with a text file and some discipline. PAPI is one way to make it part of every task rather than a habit to keep. Each task gets a build handoff with the goal, what is out of scope, how to check it is done, the risks and the likely files, and your assistant builds from that. Every build files a report, including issues found along the way. Review is your verdict, task by task.

Decisions keep their reasoning too, so a deliberate choice that looks like a mistake from outside is less likely to get helpfully corrected (more on that in Architecture Decision Records Do Not Survive AI Speed). PAPI does not write code. Your assistant still does the building.

The honest version

None of this stops every break. A brief cannot list a dependency nobody knew existed, and some breaks come from outside your code entirely: a service changes, a model gets retired, and something stops working while you changed nothing. If you have tests, keep them. They catch things a person reading a change will miss.

And if you are still prototyping, throwing away most of what you build, a checkpoint before each prompt is plenty. Come back to this when the app holds something you would hate to lose. If you want to see a brief and a review on a real project first, walk through the demo project, no signup.

Your AI starts every session from zero. Your project stays on course.

Free on up to three projects. No card.