Team & capacity

How to spot AI-generated answers in take-home tests

You can't reliably detect AI-written work, and detector tools flag innocent people. The fix is task design: ask for decisions and process, then follow up live and have candidates explain their own work.

·5 min read

Start with the uncomfortable part: you cannot reliably tell whether a written take-home answer was produced with AI assistance. Not from the prose style, not from the em dashes, not from a detector tool.

Detection tools are not dependable enough to base a hiring decision on. They produce false positives, and the cost of a false positive is that you reject a good candidate for something they didn't do — and possibly tell them why. Nobody should be accused of dishonesty on the basis of a probability score from a tool that can't show its reasoning.

So stop trying to detect. Change what you're asking for instead.

The real problem isn't cheating

It's that your task no longer measures anything.

If the take-home asks for a polished artefact — a strategy summary, a landing page, a campaign concept, a working function, a content brief — then the artefact is now cheap to produce for everyone. A candidate who generates a first draft and refines it might be excellent. A candidate who generates something and submits it unread is not. Both submissions look similar, and the one that looks best may well be the second one.

That's the failure. Not that people use the tools — most of your team already does, and if the role involves this work you probably want someone who uses them well. The failure is that output-only tasks have stopped separating candidates.

Worth being clear with yourself about which you're testing:

Most agency roles need both, tested separately.

Design the task around decisions, not output

The reliable move is to make the thing you're grading the reasoning, because reasoning about a specific, messy, private situation is where generic output falls apart.

Give a deliberately flawed brief. A real one, anonymised. Contradictory goals, a budget that doesn't match the ask, a missing piece of information, a deadline that doesn't work. Ask what they'd do about it. Strong candidates notice the contradiction and say so; weak submissions answer the brief as written, confidently.

Ask for the trade-offs they rejected. "Give us two directions you considered and didn't pursue, and why." Rejected options require a point of view. Generic answers tend to produce plausible options and thin reasons.

Ask what they'd need from the client. "List the three questions you'd ask before starting." This is the single most revealing question in agency hiring and it barely requires any work to answer. The questions someone asks reveal exactly how much of this job they've done.

Constrain hard. "Two hours, maximum one page, one direction only." Volume of polished output stops being a proxy for effort, which removes the advantage of just generating more.

Use your own material. A situation from a real project, with names and numbers changed, is harder to answer generically than "how would you launch a product on social media."

Say what's allowed, out loud

Put it in the brief, plainly:

Use whatever tools you normally use, including AI. We care about the thinking and the decisions, not who typed the words. You'll walk us through this on a call, so be ready to explain any part of it.

Two things happen. Honest candidates stop worrying about a trap. And everyone now knows the live conversation is coming, which is what actually does the work.

Ambiguous rules only penalise the people who follow rules.

The follow-up call is where it resolves

Twenty to thirty minutes on the submission, with the candidate walking you through it. Not an interrogation — a normal work conversation, the kind you'd have with a colleague about their draft.

Useful moves:

You're not testing recall. Plenty of people who did all the work themselves will be vague about a detail from four days ago. What you're looking for is whether they can think inside the work — extend it, defend it, revise it, spot its weaknesses. That comes from having done it, however it was drafted.

What not to do

Don't use a detector as evidence. If a tool flags something, it has told you nothing you can act on.

Don't accuse. If a submission feels generic, that's a reason to ask better questions on the call, not a reason to write to someone about integrity. You will be wrong sometimes, and the harm is asymmetric.

Don't ban the tools. Unenforceable, and it filters for compliance rather than capability.

Don't grade on prose style. Non-native speakers, people who write formally, and anyone who uses a grammar checker will be punished for nothing.

Don't lengthen the task. Six hours instead of two doesn't improve the signal; it just costs honest candidates more of their weekend.

A version you can run next week

Replace your current take-home with this, whatever the discipline:

  1. A one-page brief from a real past project, anonymised, containing one genuine contradiction you remember arguing about at the time.
  2. Two hours, capped. Ask for: one recommended direction, two rejected options with reasons, three questions they'd ask the client, and one risk they'd flag.
  3. A note in the brief saying tools are fine and a call is coming.
  4. A 25-minute call where they walk you through it, ending with one changed constraint worked through together.

You'll find the conversation separates candidates far more sharply than the document does — which was always true, and is now the only part that reliably works.

Keep reading

All articles