How to Give Feedback on AI-Generated Work Your Team Hands In

Written by Nick Baudoin

The most useful feedback on AI-generated work usually isn’t about accuracy, and it isn’t about whether the person should have used AI at all. It’s about what the tool couldn’t see. A lot of AI-assisted work that goes wrong reads fine on its own and only goes wrong next to the things it has to sit beside. So review the piece against its neighbors, find where it contradicts them, and coach the person on the context they gave the tool rather than on the output itself. “Check your work more carefully” doesn’t help. They did check it, and on its own, it was fine.

Why does AI-generated work pass review and still cause problems?

I’ve had this happen several times. Someone hands me a page design or a piece of code built with AI help, and if you look at it alone, there’s nothing wrong with it.

The trouble is that it doesn’t live alone. It sits among five or so related template pages that each position the same thing in their own way, and the new piece ignores all of them. Read together, the set says things that are incorrect or confusing. You can’t see that from the piece. You can only see it with the others open.

That pattern isn’t specific to websites. Think about what an HR team produces with AI help: a job description, a policy summary, a benefits explainer, an onboarding checklist, an internal announcement. Each one has neighbors. The job description has to agree with the careers page and the offer letter. The policy summary has to agree with the handbook section it summarizes. The onboarding checklist has to agree with what the hiring manager told the candidate. A draft can be well written, correctly formatted and plausible on every line, and still contradict any of those.

What’s wrong with the usual advice?

Most guidance on reviewing AI output tells managers to treat it as suspect for errors. Fact-check it, look for made-up details, and ask people to disclose when they’ve used a tool.

None of that is wrong, and fact-checking is worth doing. But the argument underneath it, that factual error is the main risk, sends reviewers looking in the wrong place. Sentence-level plausibility is exactly where these tools are strongest. A reviewer who reads the draft closely, finds nothing wrong, and approves it has done everything the usual advice asked, and missed the actual problem.

It also shapes the conversation badly. When the question is “did you use AI?”, feedback turns into a conversation about effort or honesty, and people get defensive. When the note is “be more careful”, the person has nothing to change. Their review had the same blind spot as the tool, because both of them were looking at one piece in isolation.

The better question is narrower and much easier to act on: what did the tool have in front of it when it produced this?

What should a manager look at instead?

Here’s the check I’d run on anything AI-assisted that has to fit into existing material. Call it the neighbor check. It has three steps, and once it’s habit it shouldn’t take long.

1. Name the neighbors before you read the draft. Ask what this piece has to agree with. Earlier versions of the same document, related pages, the policy it summarizes, whatever the recipient was told last week. Write the list down. If you genuinely can’t name anything, the work is standalone and a normal read is enough. Plenty of work is standalone, and this check isn’t for that.

2. Read it in company. Put the draft beside two or three of the neighbors and look for where it says something different. A different promise, a different term for the same thing, a different order of priorities, a different tone for the same audience. Any single difference might be defensible. The question is whether someone who reads both comes away confused about what’s true.

3. Trace the gap to the brief. When you find a contradiction, don’t start with the output. Ask two questions. Did the person know the neighbors existed? And did they give the tool visibility into them? The answers tell you what kind of feedback this is.

How do you tell whether the gap is theirs or yours?

Tracing it back usually lands in one of three places, and only one of them is a coaching conversation about the person’s work.

They didn’t know the neighbors existed. This is the organization’s gap, not theirs. The context lived in someone’s head or in a folder they were never pointed to. Newer people are the most exposed to this, because knowing what a document has to agree with is exactly the kind of knowledge that takes time on the team to build. The fix is to point them at it, and ideally to write the list down somewhere the next person will find it. Marking them down for it teaches the wrong lesson.

They knew, but the tool never saw them. This is the coaching case. The tool worked from what it was given, and it was given the task without the surroundings. So the feedback is the same two things I’ve ended up explaining more than once: communicate what you actually want, including what the piece has to stay consistent with, and give the model visibility into the related material. That’s a skill, it’s teachable, and it’s far easier to learn from a side-by-side than from a rule.

They gave it the context and it still drifted. Then the problem isn’t the person and it isn’t a one-off. It’s a process that needs a standing check. If the same kind of drift keeps showing up after good briefs, the fix belongs in that check, not in someone’s performance notes.

Sorting it this way does something useful for fairness. Two people can hand in the same flawed draft for completely different reasons, and only one of those reasons is about them.

What does the feedback conversation actually sound like?

Show, don’t describe. The problem only exists when two things are on screen together, so a written note that says “this conflicts with the benefits page” makes the person go find the conflict themselves, and they may find a different one.

A lot of our async explaining already happens in short recorded walkthroughs attached to the task, and this is the kind of feedback that format suits. Open the draft, open the neighbor, point at the line that disagrees, and say why it matters to the person who’ll read both.

Then give the why and the how, in that order. The why is almost always the same: the tool did a good job with what it had, and what it had was incomplete. The how is specific to the piece: here are the three documents it needed to see, here’s the sentence you could have added to the request so it knew to stay consistent with them.

Notice what that conversation doesn’t include. It doesn’t ask whether AI should have been used. It doesn’t imply the person was lazy. It treats the gap as a missing input, which is usually what it is, and it leaves the person with a concrete thing to do differently next time.

What changes for performance reviews?

If AI-assisted drafts are judged only on how good they look alone, you’ll find most of them look good. Standalone quality stops telling you much about who is doing strong work, because the tool has raised the floor for nearly everyone.

What still separates people is judgment about context. Who thinks to ask what a piece has to agree with. Who gathers the right material before starting. Who notices when their draft and an existing page tell different stories. Those are the things worth evaluating now, and they’re invisible unless you make them part of the expectation.

The simplest way to do that is to ask for the neighbor list up front. Before someone hands in anything that joins existing material, they note what it has to agree with and confirm they checked it. It should take a minute. It turns an invisible skill into a visible one, it gives newer people a prompt they wouldn’t otherwise have, and it gives you something concrete to coach against.

It also protects the reviewer. If the list was right and the draft still contradicts a neighbor, that’s a miss you can point to. If the list was wrong, the conversation is about what they didn’t know, which is often a gap you can close for the whole team at once.

Common questions

Should we require people to disclose when they’ve used AI? That’s a policy decision and a reasonable one to make, but it isn’t a substitute for feedback. Knowing a tool was involved doesn’t tell you whether the work fits with everything around it. The question that does is what the tool could see.

Isn’t this just ordinary editing? Partly. Consistency has always been part of good review. What’s changed is volume. Producing a polished, standalone draft is now fast and cheap, so more of them arrive, and each one arrives looking finished. Polish used to be a rough signal that someone had spent time with the surrounding material. It isn’t anymore.

What if the neighbors contradict each other too? Then the AI draft did you a favor. It exposed a disagreement that already existed in your own material. Fix the source documents first, or every future draft, human or not, will inherit the same confusion.

Author Bio:
Nick Baudoin is Founder and President of
Alkali, which builds websites and marketing programs for established B2B companies and publishes research on 55,000+ US B2B websites.

Share:

Leave a Reply

Your email address will not be published. Required fields are marked *