During recent work on the What Up Inc website, an AI coding agent reported that a set of changes had been completed and persisted. The work was bounded: make the requested website changes, leave unrelated areas alone, and report the result.

The report sounded like a reasonable place to stop. It described work that was finished, not work that still needed doing. But when I checked the actual repository and the rendered website, some of those changes weren’t there.

That was an uncomfortable moment. I wasn’t simply deciding whether I liked what the agent had built. I was trying to establish whether the work it described had happened at all.

It said it was done

There were now two versions of the project: the one described in the conversation and the one that existed in the files. Those versions did not agree.

Git, which records and shows changes to the website’s source, helped establish what had actually changed. Reading the source files and inspecting the rendered pages provided the rest of the picture. Some of the reported changes were not present. The completion report was not a reliable description of the current state.

This turned a small website task into a state-reconciliation problem. Before continuing, I needed to separate work that really existed from work that had only been reported. Otherwise, the next instruction could build on a change that wasn’t there, or repeat work that was already correct.

The distinction matters. If a task is unfinished and clearly reported as unfinished, I can decide what to do next. If it is unfinished but reported as complete, I may make the next decision using the wrong starting point.

The mistake wasn’t the whole problem

People make mistakes. Software tools fail. An AI agent can misunderstand a request or fail to carry out part of it. None of that, on its own, makes the tool useless.

What concerned me was the confidence of the completion report. It sounded authoritative even though the repository did not support it. A clear explanation of supposedly completed work can be persuasive, especially when it uses the language of checks, saved changes and successful delivery.

I don’t need to assign an intention to the agent to recognize the problem. I also don’t need to speculate about why each missing change was reported as finished. The practical issue was that I could not use its description as proof of the result.

An AI agent’s report is not evidence that the work happened.

That is the lesson I took from this experience. The report can help direct a review. It cannot replace looking at the work.

Trust the evidence, not the status report

For this website, evidence came from several places. Git status showed which files had changed. Git diff showed what was different inside them. The actual source files showed whether the requested content existed, and the rendered website showed what a visitor would see.

Those checks answer different questions. A changed file does not prove that the right change was made. Correct source text does not prove that the page presents it properly. A successful build shows that the site can be generated, but it does not decide whether the result meets the requirement.

Tests are useful for the same reason, and have the same boundary: they tell me about the behaviour they actually check. I still need to inspect the requested result. After a release, a production smoke test—a short check of the live website—helps confirm that the expected result reached the place visitors use.

This does not mean every wording edit needs an elaborate test programme. It means the evidence should match the change. For a small content change, reading the exact difference and checking the page may be enough. A more consequential change needs more substantial checks.

The useful question is not just, “Did the agent say it checked?” It is, “What did the check actually establish?”

A stronger model helped, but wasn’t the solution

I then used a more capable model to inspect the real repository state, reconcile the work and complete the changes. That recovery succeeded.

It would be easy to make the lesson about choosing the better model. Model capability mattered in the recovery, but that would be an incomplete conclusion. Simply buying or selecting a smarter, more expensive model is not a complete reliability strategy.

What made the recovery useful was that it started from the actual repository rather than accepting the earlier report as fact. The work could be compared with the requirement, missing changes could be identified, and the result could be checked again.

I still needed evidence from the stronger model’s work. Otherwise, I would only have replaced one completion report with another. A stronger model can help, but reliability cannot depend solely on choosing one. The process needs to make discrepancies visible regardless of which model is doing the work.

A delivery path with places to stop

The workflow emerging from this experience is approximately:

Baseline → Bounded change → Build/test → Inspect actual result → Git scope check → Human acceptance → Commit → Push → Production smoke test

The baseline gives me a known starting point, including any existing work that must be preserved. The bounded change makes the request reviewable. Build and test results provide technical evidence, while inspecting the actual result lets me compare what exists with what I asked for.

Before accepting the work, I also check its scope. Did only the intended files change? Was anything unrelated touched? That matters when there is already approved work in the repository that should not be altered.

Acceptance comes before recording and releasing the approved change. The final live check comes after release because a correct local result is not the same as a confirmed production result.

I don’t see this as bureaucracy around the agent. These are places where I can stop, resolve a disagreement and avoid carrying an unsupported assumption into the next step.

I still own the decision

In my earlier article, “My BA is ChatGPT. My developer is an AI agent. I still lead the project.”, I described how AI helps me move from an idea towards implementation while I retain responsibility for the outcome. This experience is a practical reason that responsibility matters.

I decide what should change, what evidence is required and whether the result meets the requirement. I also decide whether it is ready for production. An agent can help collect evidence and explain it, but its confidence does not make those decisions for me.

I continue to find AI agents extremely useful. This incident did not change that. It changed what I am prepared to accept as proof that their work is finished.

The question isn’t whether I trust AI. The question is what evidence I require before I trust its work.

What this means for a client

Using AI in delivery does not mean blindly accepting what an AI system reports. AI can help investigate, build, test and iterate. What Up remains responsible for reviewing the evidence and deciding whether the result actually solves the client’s problem.

That does not promise a mistake-free process. It means a completion report is something to verify, not a substitute for review. If the evidence and the report disagree, the discrepancy needs to be resolved before treating the work as complete.

Have a technology problem or an idea you want to work through? We can start with the problem and what a useful result would look like.

Start a Conversation