Team & capacity
How to spot AI-generated answers in take-home tests
You can't reliably detect AI-written work, and detector tools flag innocent people. The fix is task design: ask for decisions and process, then follow up live and have candidates explain their own work.
·5 min read
Start with the uncomfortable part: you cannot reliably tell whether a written take-home answer was produced with AI assistance. Not from the prose style, not from the em dashes, not from a detector tool.
Detection tools are not dependable enough to base a hiring decision on. They produce false positives, and the cost of a false positive is that you reject a good candidate for something they didn't do — and possibly tell them why. Nobody should be accused of dishonesty on the basis of a probability score from a tool that can't show its reasoning.
So stop trying to detect. Change what you're asking for instead.
The real problem isn't cheating
It's that your task no longer measures anything.
If the take-home asks for a polished artefact — a strategy summary, a landing page, a campaign concept, a working function, a content brief — then the artefact is now cheap to produce for everyone. A candidate who generates a first draft and refines it might be excellent. A candidate who generates something and submits it unread is not. Both submissions look similar, and the one that looks best may well be the second one.
That's the failure. Not that people use the tools — most of your team already does, and if the role involves this work you probably want someone who uses them well. The failure is that output-only tasks have stopped separating candidates.
Worth being clear with yourself about which you're testing:
- Can they produce good work with the tools available? Then let them use anything, and judge the result at the standard you'd apply to a client deliverable.
- Can they think, decide and defend? Then the artefact was never the point, and you need to ask for something else.
Most agency roles need both, tested separately.
Design the task around decisions, not output
The reliable move is to make the thing you're grading the reasoning, because reasoning about a specific, messy, private situation is where generic output falls apart.
Give a deliberately flawed brief. A real one, anonymised. Contradictory goals, a budget that doesn't match the ask, a missing piece of information, a deadline that doesn't work. Ask what they'd do about it. Strong candidates notice the contradiction and say so; weak submissions answer the brief as written, confidently.
Ask for the trade-offs they rejected. "Give us two directions you considered and didn't pursue, and why." Rejected options require a point of view. Generic answers tend to produce plausible options and thin reasons.
Ask what they'd need from the client. "List the three questions you'd ask before starting." This is the single most revealing question in agency hiring and it barely requires any work to answer. The questions someone asks reveal exactly how much of this job they've done.
Constrain hard. "Two hours, maximum one page, one direction only." Volume of polished output stops being a proxy for effort, which removes the advantage of just generating more.
Use your own material. A situation from a real project, with names and numbers changed, is harder to answer generically than "how would you launch a product on social media."
Say what's allowed, out loud
Put it in the brief, plainly:
Use whatever tools you normally use, including AI. We care about the thinking and the decisions, not who typed the words. You'll walk us through this on a call, so be ready to explain any part of it.
Two things happen. Honest candidates stop worrying about a trap. And everyone now knows the live conversation is coming, which is what actually does the work.
Ambiguous rules only penalise the people who follow rules.
The follow-up call is where it resolves
Twenty to thirty minutes on the submission, with the candidate walking you through it. Not an interrogation — a normal work conversation, the kind you'd have with a colleague about their draft.
Useful moves:
- "Talk me through how this evolved." What was the first version, what changed, what got cut. Work that someone genuinely made has a history.
- "Why this and not the obvious alternative?" Pick something they did and name a plausible other route. You're looking for a considered reason, not the right answer.
- "This part — what would break it?" Someone who made a thing knows its weak points. It's the first thing they'll tell you if you ask.
- Change one constraint live. "The budget just halved" or "the client hates the third option, and it's Thursday." Then work through it together for five minutes. This is closest to the actual job, and it's the part no preparation covers.
You're not testing recall. Plenty of people who did all the work themselves will be vague about a detail from four days ago. What you're looking for is whether they can think inside the work — extend it, defend it, revise it, spot its weaknesses. That comes from having done it, however it was drafted.
What not to do
Don't use a detector as evidence. If a tool flags something, it has told you nothing you can act on.
Don't accuse. If a submission feels generic, that's a reason to ask better questions on the call, not a reason to write to someone about integrity. You will be wrong sometimes, and the harm is asymmetric.
Don't ban the tools. Unenforceable, and it filters for compliance rather than capability.
Don't grade on prose style. Non-native speakers, people who write formally, and anyone who uses a grammar checker will be punished for nothing.
Don't lengthen the task. Six hours instead of two doesn't improve the signal; it just costs honest candidates more of their weekend.
A version you can run next week
Replace your current take-home with this, whatever the discipline:
- A one-page brief from a real past project, anonymised, containing one genuine contradiction you remember arguing about at the time.
- Two hours, capped. Ask for: one recommended direction, two rejected options with reasons, three questions they'd ask the client, and one risk they'd flag.
- A note in the brief saying tools are fine and a call is coming.
- A 25-minute call where they walk you through it, ending with one changed constraint worked through together.
You'll find the conversation separates candidates far more sharply than the document does — which was always true, and is now the only part that reliably works.
Keep reading
- Screening candidates without wasting everyone's weekA five-stage funnel from 120 applications to one hire costs about 24 hours of senior time — $2,280 at a $95 loaded rate. Here's how to spend those hours where they actually change the decision.
- How to write a job post that gets good applicantsA job post's job is to filter, not to attract. Here's a reusable skeleton, the four things that make strong candidates skip your ad, and why leaving out the salary costs you the people you most want.
- When to hire your next person, and how to know you can afford themA hire whose fully loaded cost is $118,000 a year needs 66 billable hours a month just to break even. Here's the capacity test, the cash test, and the ramp curve most agencies forget to model.