Case study · 5 min read
Testing the thinking, now that AI can produce the artifact
- when the evidence is clear early
- 3 → 2 rounds
- one shared assessment system
- 4 disciplines
- the same criteria for every interviewer
- One standard
The problem: the artifact stopped being evidence
Our design hiring used a familiar format. Candidates received a brief, had two days to produce a solution, and then talked interviewers through it. It worked as long as the submission reflected the candidate's own thinking.
AI changed that. Almost overnight, submissions became uniformly polished. Interviewers were looking at robust, well-structured solutions with no reliable way to tell whether the candidate had actually reasoned their way there.
Two other gaps made it worse. Interviewers had no shared guide: each one decided alone what to ask and what "good" looked like at each level. And HR, which handled the final step before an offer, couldn't see what criteria had been applied at all.
The take-home had stopped being evidence. The question was what should replace it.
What I built
I kept what the candidate sees almost exactly as it was, and rebuilt everything they don't see.
For candidates: a deliberately simple brief. A problem statement, what to deliver, and how to submit. Nothing more.
For interviewers: a complete assessment kit, for each of four disciplines: UX, UI, research, and visual design. Each kit includes:
- A bank of assessment tasks across domains like fintech, edtech, e-commerce, and marketing sites, graded for junior, senior, and lead candidates.
- An evaluation guide with weighted dimensions and clear expectations for each level: potential and fundamentals in juniors, systems thinking in seniors, strategy and whole-ecosystem thinking in leads. It lists the red and green flags to watch for.
- Solution briefs for evaluators, describing strong approaches and common oversights. They're deliberately not answer keys.
- Question sets, both standard and task-specific. Some are designed so that only someone who genuinely reasoned through the problem can answer them well.
- Markers for spotting AI-generated work, which interviewers apply to the submission before the conversation starts.
For HR: evidence. Every interviewer scores against the same criteria in a shared sheet, which goes to HR. A hiring decision now comes with a record of how it was made.
The principle behind the design: the candidate knows the problem; only the interviewer knows the questions. We don't ban AI, and a candidate can use it to build their submission. What we test is whether the thinking behind it is theirs.
How adoption happened
I followed the same pattern I use for every system: train the people who will run it, then let them take it further.
For each discipline, we picked a set of interviewers from every location: senior designers for UX and UI, senior researchers for research, and senior graphic designers for visual design. I ran the first workshop myself, covering how to use the kit, and the design leads attended alongside them. The leads then ran their own smaller workshops within their units.
We also created assessor training for newly promoted senior designers, so the bench of people qualified to interview keeps growing as the team does.
The process itself became clear at every stage. Discipline interviewers assess the skill. Leads make the final hiring call. HR then assesses fit, which is deliberately separate from skill.
What changed
- One standard across every studio. The kit has been the organization's hiring process for about a year.
- Interviewers who know what they're looking for. Every interviewer now works from the same questions, criteria, and flags, instead of improvising.
- Fewer rounds, when the evidence allows it. In some hiring processes, interviewers could cut the process from three rounds to two. Sometimes they could tell from the portfolio alone that the work was largely AI-generated.
- Transparency for HR. For the first time, HR can see exactly what each candidate was assessed on and how they scored.
Time to hire didn't change much, and speed was never the goal. The goal was confidence in who we hired.
I'll also be honest about the limits of the evidence. Over the past year we hired far less than usual, by choice: as part of our AI transformation, we decided not to backfill attrition. So the system has run on fewer hires than I'd like. It's in place and ready for when hiring picks up.
What I'd do again, anywhere
- When the artifact stops being evidence, move the test to the conversation. Polished output proves little now. Live reasoning still proves a lot.
- Keep the brief simple and the questions private. Depth belongs with the evaluator, not in the brief.
- Test for thinking, not for AI use. Banning tools is unenforceable. Finding out whether the candidate owns the reasoning is not.
- Give every stakeholder the same criteria. Interviewers, leads, and HR should all know what "good" means before the first interview.
- Grow the interviewer bench on purpose. Train newly promoted seniors to assess, so hiring quality doesn't depend on a handful of people.