Skip to content
← Tools

Interview kits

Three kits, one per kind of role. Six questions, the two follow-ups that go underneath each one, and a scorecard you fill in during the interview rather than after it.

Print the one you need and take it into the room. Nothing to configure, no account, and no part of this needs the internet once it is on paper.

Where these came from

These questions are researched and structured. They are not a list I have personally used for years. They are built from the interview-validity literature and from published operator and design practice, then filtered against one rule: a question only survives if it still works after the candidate has read it here.

What is mine is the four-country experience the market notes are drawn from, and the decision about which questions made the cut.

How to use this

Takes about forty five minutes per candidate

What you need

  • The kit for the role you are hiring, printed
  • One interviewer, and a second person only if they will score independently
  • For the organisation-level kit, a real live situation from your business for question six
  • For the design kit, a screen and a real asset of yours to be criticised

What you get

  • Six questions in the order to ask them, with two planned follow-ups under each
  • What a strong and a weak answer actually contain, per question
  • A 0 to 3 scorecard with anchors, to fill in as you go
  • The four questions to ask about any market before you hire into it

Getting the most out of it

  1. 01

    Six questions is the limit for forty five minutes. If you are assessing nine things, you are running an unstructured interview wearing a form.

  2. 02

    Fill the scorecard during, not after. Scoring from memory is where the structure leaks back out.

  3. 03

    One interviewer, not three. A panel roughly doubles how much your interviewers agree with each other, and does nothing for whether the hire works out. If you want a second opinion, have someone score the recording independently.

  4. 04

    If your candidate is interviewing in their second language, do not add probes. Deep probing and think-aloud narration penalise the language, not the capability. These questions already lean toward past behaviour for that reason.

  5. 05

    Ask the same six, in the same order, every time. The comparison is the instrument. Changing the questions per candidate destroys it.

The scale

Four points, not five. A midpoint invites you to park every candidate in it. The strong and weak notes under each question tell you what a 2 and a 0 look like for that question.

0 · Absent
Nothing relevant offered, or entirely generic. Collapses at the first follow-up: no names, no numbers, no account of what was said.
1 · Weak
One real example, thin. Describes the situation but not their decision. Costs unnamed. Survives the first follow-up, fails the second.
2 · Solid
A specific event, a decision that was theirs, a named consequence, a stated reason. Answers the second follow-up with new material rather than retelling the first. This is a hire-able answer.
3 · Outstanding
All of that, plus one of: they name the cost or their own error before being asked; they can produce the artefact; the second follow-up reveals an operating principle the first only implied.

Kit one

Hospitality

Barista, floor and service, kitchen, shift lead, front desk, small-property operations.

Question 01 · Reading a room

Walk me through the last shift you worked, from the moment you arrived. Not a typical shift, the last one. What was the first thing you noticed that was off, and what did you do about it?

Whether they scan an environment or wait to be told. Also, indirectly, whether they have actually worked a shift recently.

Strong

A mundane, physical, verifiable detail. The under-counter fridge left ajar, two servers rostered on the same section, the grinder drifting coarse after a humid morning. Then what they did unasked, and what it cost or saved.

Weak

Habitual present tense instead of one past event: “I always check everything is clean.” Generic virtues, no physical detail. Or a heroic rescue with no mundane texture around it. Real shifts are mostly mundane.

  1. Follow-up one · Verify

    Who else noticed it, and what did they do?

  2. Follow-up two · Repeat

    That was the last shift. Now tell me the same thing about the shift before that.

Question 02 · Recovery

Tell me about a guest complaint where the guest was factually wrong and you knew it at the time. What did you do?

Judgment where truth and service conflict, and whether “the guest is always right” is a slogan or a considered position with limits.

Strong

Holds the tension rather than dissolving it. Separates the emotional claim from the factual one and answers the emotional one first. Reproduces roughly what they said. Names the cost of conceding.

Weak

“The customer is always right, so I apologised and comped it,” with no cost, no words, no aftermath. Or an answer where they won the argument and are still proud of it.

  1. Follow-up one · Verify

    What exactly did you say, as close to word for word as you can get?

  2. Follow-up two · Consequence

    What did your manager say when they found out, and did you agree with them?

Question 03 · Judgment

Tell me about a time you broke a rule at work. What was the rule, what did you do, and did you tell anyone?

Their actual decision architecture: whether they distinguish rules that protect the guest, rules that protect the business, and rules that are merely habit. And whether they disclose.

Strong

A real, small, specific rule. A comp limit, a closing time, a seating policy, a discount they were not authorised to give. A clear reason. A disclosure to someone, ideally the same day. An account of the consequence.

Weak

“I never break rules”, which is either untrue or operationally inflexible. A safety, allergen or cash-handling rule broken casually. A rule broken and concealed.

  1. Follow-up one · Verify

    Who did you tell, and how long after?

  2. Follow-up two · Invert

    Now tell me about a rule you kept even though you thought it was wrong, and what that cost.

Question 04 · Ownership

Tell me about the worst mistake you have made at work. Start with what it cost, in money, in time, or in a person.

Ownership. The specific diagnostic is whether they lead with the cost or with the mitigation.

Strong

Leads with the cost, in units. Uses “I”. Names who else was affected. Describes telling someone before being found out, or admits they did not and what that was like. Says what changed, and whether it survived.

Weak

A mistake that is secretly a strength. A cost never quantified. The passive voice: “mistakes were made”, “the order didn’t get placed”. A mistake attributed to a system with no personal share in it.

  1. Follow-up one · Verify

    Who found out, and how?

  2. Follow-up two · Deepen

    What is the mistake you made that nobody ever found out about?

Question 05 · Learning

What is the most recent thing you learned to do properly at work, and who taught you?

Learning as an active habit rather than accumulated tenure. Naming a teacher is also a proxy for coachability, and for having genuinely been on a working floor.

Strong

Recent, in weeks or months. Specific. Names a person. Describes the correction and how it felt. Honest about the practice interval: “I was bad at it for about three weeks.” Volunteers what they still cannot do.

Weak

Something from years ago. A certificate rather than a skill. Nobody taught them. Or instantaneous learning, “I picked it up straight away”, which is almost always false and always uninformative.

  1. Follow-up one · Verify

    What did they say to you that made it click?

  2. Follow-up two · Demonstrate

    Teach me the same thing now, in two minutes.

Question 06 · Behaviour unsupervised

Describe the last time you were the most senior person on site. What decision did you make that you would normally have escalated?

Behaviour in your absence. The single most valuable thing to know when you are personally present maybe a third of the time.

Strong

A concrete decision with money or safety inside it. The reasoning. Who they informed, and when. Whether the owner agreed afterwards. A working sense of where the escalation threshold sits, articulated rather than assumed.

Weak

Nothing ever happens when they are in charge. Every decision escalated, at two in the morning, including the trivial ones. Or a decision made and never reported.

  1. Follow-up one · Artefact

    What did you write down about it, and where does that live?

  2. Follow-up two · Invert

    Give me the decision you got wrong in that position.

Scorecard · Hospitality

Candidate Date
Question0–3Evidence, in their words
01 Reading a room
02 Recovery
03 Judgment
04 Ownership
05 Learning
06 Behaviour unsupervised

Score each question before you ask the next one. If someone else is scoring too, do not compare until you have both written yours down.

Kit two

Organisation-level

Operations lead, country manager, head of function, chief-of-staff EA, the first senior hire into a team under ten.

Before the room. Question six needs a real, currently live situation from your business, chosen before the interview. A hypothetical you invent on the spot is answerable from a script.

Question 01 · Borrowed judgment

Tell me about the last time you signed off on work you were not qualified to evaluate. What did you actually do to get comfortable?

Whether they have a repeatable method for borrowed judgment, or default to blind trust or to an attempt to become the expert, which at this level is a claim of infinite time.

Strong

Names the domain and admits the gap flatly. A specific mechanism: asked for the assumption behind one number, got a second read from someone outside the reporting line, asked what would have to be true for this to be wrong. Names what they deliberately did not check, and why that was acceptable.

Weak

“I trust my people” with nothing behind it. Or a description of the decision that never touches the verification. Or “I learned enough to check it myself” offered as a general policy.

  1. Follow-up one · Boundary

    What did you decide not to check, and why was that acceptable?

  2. Follow-up two · Failure

    When did that method last fail you? What got through?

Question 02 · Deciding without the facts

Tell me about a decision you made where you knew you were missing something important and chose not to wait. What were you missing?

Whether they can name the specific unknown. People who genuinely decide under uncertainty state the gap precisely. People who improvise rationalise afterwards.

Strong

Names the missing fact exactly. States the cost of waiting in units: days, money, a person about to resign, a window closing. Describes what they did to cap the downside given the gap.

Weak

A decision that was actually well-informed, reframed as brave. Or pure instinct, with no articulation of what was unknown at the time.

  1. Follow-up one · Counterfactual

    What would you have done differently if you had had that information?

  2. Follow-up two · Learning

    What did you find out afterwards? Were you right, and what changed in how you decide?

Question 03 · Authority they did not start with

Tell me about someone who reported to you who had been doing that job longer than you had been in the industry. What was the first thing you changed, and how did it land?

Whether they lead by earning or by asserting, and whether the relationship was real enough to be described in detail.

Strong

Names the change. Names the resistance concretely, including what the person actually said. Describes what they did about it, which in good answers is almost always a mix: conceded something real, held one line.

Weak

“I made sure to listen and learn from their experience”, with no change ever made. Deference as strategy. Or the change was imposed by authority and they never noticed how it landed.

  1. Follow-up one · Their words

    What did they say to you? Their words.

  2. Follow-up two · Cost

    What did you concede that you did not want to concede?

Question 04 · Conflict absorbed

Tell me about a conflict you absorbed, where you took the hit rather than passing it up or down. What did it cost you?

Whether they shield the organisation, and whether they understand that absorption has a price which has to be tracked.

Strong

Names the cost precisely: a relationship, a person who left thinking badly of them, a number they missed, months of their own time. Explains why absorbing was right in that instance rather than as a general posture.

Weak

A martyrdom narrative with no cost accounting. Or absorption that is really conflict avoidance with better branding.

  1. Follow-up one · Boundary

    When have you absorbed something you should have escalated?

  2. Follow-up two · System

    How does the founder find out about the things you absorb?

Question 05 · The first two weeks

In your last role, what did you actually do in your first two weeks? Not what you planned. What was in your calendar.

The gap between the theory everyone can recite and what they really did. That theory is now free to generate, which is exactly why it is worthless.

Strong

Concrete and boring, and therefore verifiable. Names of people, a specific document they demanded, a system they got access to, a customer they visited, a site they walked. Often includes something they now think wasted a week.

Weak

The textbook: listen, learn, meet stakeholders, find quick wins. No specifics attached to any of it.

  1. Follow-up one · Waste

    What did you spend time on in those two weeks that turned out not to matter?

  2. Follow-up two · Hindsight

    What did you miss then that you only saw at month four?

Question 06 · You, unreachable

I am on a plane, then out of signal, for two weeks. On day three, [insert a real, currently live situation from your business]. What do you do?

Judgment under your actual constraints. Also, immediately, whether they ask for information before answering.

Strong

Asks questions first. What does the contract say, who else knows, has this happened before, what did we do then. Separates what genuinely must be decided inside two weeks from what can wait. Chooses the option that preserves reversibility.

Weak

Decides instantly and confidently on no information. Or refuses to decide and waits for you, which is the exact failure you are hiring against. Or answers with process only, and never says what they would do.

  1. Follow-up one · Twist

    You do that. On day six it gets worse in this specific way [construct it live, so their first answer is now wrong]. Now what?

  2. Follow-up two · Disagreement

    I come back and tell you I disagree with what you did. What do you say to me?

Scorecard · Organisation-level

Candidate Date
Question0–3Evidence, in their words
01 Borrowed judgment
02 Deciding without the facts
03 Authority they did not start with
04 Conflict absorbed
05 The first two weeks
06 You, unreachable

Score each question before you ask the next one. If someone else is scoring too, do not compare until you have both written yours down.

Kit three

Design and brand

Brand and visual, content and social, and anyone whose work will be made with AI tools, which is now all of them.

Before the room. Questions one, two and three need a screen and a real asset of yours in the room. Book the time for it.

Question 01 · Whose judgment is this

Open the file for the piece you are proudest of. Show me the version history, or just the folder. Walk me through the three directions you killed.

Whether the finished artefact is evidence of this candidate’s judgment. Also, bluntly, whether the work is theirs.

Strong

They have the files. The killed directions are genuinely different ideas, not colour variants of one idea. They can say why each died and who decided. The survivor is visibly downstream of a real constraint.

Weak

No files, no versions, “I don’t keep drafts.” Three directions that are one idea in three palettes. A narrative about the final piece with no losers in it. Or no clarity on which decisions were theirs.

  1. Follow-up one · Regret

    Which killed direction do you still think was better, and why did it lose?

  2. Follow-up two · Origin

    Show me the earliest file. What was the brief at that point, and how did it change?

Question 02 · Live critique

Here is a piece of our actual work. Tell me what is wrong with it.

Critique on an unprepared stimulus, and whether they can be candid to the face of the person who commissioned the thing. Show a real asset: a menu, a sign, a room card, a recent post.

Strong

States what the piece is trying to do before criticising it. Separates brief failure from craft failure. Physical and specific: “the counters fill in at the size it actually appears on the door.” Offers a cheap fix.

Weak

Praises it. That is fatal: either dishonest or blind, and both disqualify in a role you are hiring for judgment. Or criticises everything with no priority. Or asserts taste with no reasoning: “it feels dated.”

  1. Follow-up one · Priority

    If you could change one thing and only one, which?

  2. Follow-up two · Invert

    Now argue the other side. Defend the version I have.

Question 03 · Directing the tools

Show me a piece where you used AI image or copy tools. Show me the prompts, the rejects, and the final. How many generations did it take, and what did you fix by hand?

The real working relationship with the tools. Whether they curate and finish, or accept the first output and ship it.

Strong

The reject pile exists and it is large. The prompts show iteration on a specific problem, not the same request reworded. They can name what the model kept getting wrong: hands, type, a brand element, consistency between frames.

Weak

No rejects kept. One prompt, one output, shipped. Cannot name a single failure mode, which means they have not looked closely at their own output.

  1. Follow-up one · Taste

    Which generation would most people have shipped? Why didn’t you?

  2. Follow-up two · Live

    Do the same brief now, in front of me, in fifteen minutes. Talk while you work.

Question 04 · Verification

Tell me about a time an AI tool gave you something that looked right and was wrong. How did you catch it?

Verification instinct. Plausibility is now free and accuracy is not, which makes this the most important competence on the list.

Strong

A specific instance. A fabricated statistic in generated copy, a citation that does not exist, a mark close enough to an existing one to be a real risk, type that reads at a glance and is gibberish at full size. And the check they now run every time.

Weak

It has never happened to them, which means they are not checking. Caught by luck, with no process consequence. Or an underlying assumption that the output is trustworthy by default.

  1. Follow-up one · Process

    What check do you run now, every single time?

  2. Follow-up two · Miss

    What did you ship with one of these in it, and only found out later?

Question 05 · Shipping under constraint

Tell me about a time you shipped work you knew was not good enough. What was the constraint, and what did you cut?

Production judgment. Whether they can degrade gracefully rather than either missing the date or shipping a mess.

Strong

Names the constraint precisely: a date, a budget, a missing asset, a late decision. Names what was deliberately sacrificed and what was protected. Articulates the rule used: “the thing most people will see gets the time.”

Weak

Never ships below standard, which means they will not survive a business with an opening date. Or shipped badly with no choice inside it. Or blames a stakeholder without ever making a decision.

  1. Follow-up one · Priority

    What did you protect at all costs?

  2. Follow-up two · Evidence

    Could anyone else tell? How do you know?

Question 06 · Interrogating a brief

I am going to give you a one-line brief and nothing else: we need something for the new property opening. Do not design anything. Give me the five questions you would ask me, in order.

Brief interrogation, which is the highest-leverage skill when the client is a founder who briefs badly. That describes almost every founder, including you.

Strong

Questions that reduce risk in descending order: who is it for, what does success look like, where does it physically appear, what already exists, what is the deadline and budget, who signs off, what must it not look like. Constraints before preferences.

Weak

Asks about visual preferences first: colours, style, references. Five questions that are really one question. Or starts designing despite being told not to.

  1. Follow-up one · Weighting

    Which of those, if I answered it badly, would waste the most of your time?

  2. Follow-up two · Scar

    Tell me about a project where you did not ask one of those, and paid for it.

Scorecard · Design and brand

Candidate Date
Question0–3Evidence, in their words
01 Whose judgment is this
02 Live critique
03 Directing the tools
04 Verification
05 Shipping under constraint
06 Interrogating a brief

Score each question before you ask the next one. If someone else is scoring too, do not compare until you have both written yours down.

Any market, four questions

Before you hire into a country, answer these four. The countries underneath are worked examples from the four I hire in. They are examples, not a live table, and the method is the part that lasts.

  1. What does it cost to run?

    The employer load on top of gross, not the salary. It is rarely the number you were quoted.

  2. What does it cost to change your mind?

    Not to exit. To move someone up a tier, to grow a team, to shrink one, to say out loud that it is not working.

  3. How fast do people leave?

    Monthly churn compounds into a retraining bill that never appears on the salary line.

  4. What can you not ask?

    The legal floor differs sharply, and the strictest of these four is criminally enforced.

Worked examples · correct as of August 2026

MarketCost to runCost to change your mindChurn
PhilippinesAbout 22% on top of minimum wage. The heaviest of the four at the bottom of the scale.No cap. Backwages accrue through two to five years of appeals, and the bond to appeal is the whole award.Not established.
IndiaFalls sharply as salary rises.Four Labour Codes commenced 21 November 2025. Verify against the current position.5–8% per month on the frontline. Over a year you rehire half to two thirds of them.
ThailandAbout 1.4% on a developer salary. Practically just the salary.180 days of severance at three years, and underperformance is not a ground you are allowed to give.Not established.
UAEFalls as salary rises, as elsewhere.Arbitrary dismissal capped at three months of wages.Not established.

The employer load falls as salary rises in all four, because contributions are capped in absolute money. The “add twenty percent” rule is only true at the bottom, which means the load is heaviest exactly where your margin is thinnest.

Emiratisation, the UAE national-hiring quota, starts at 20 staff. Below that the rules stay simple.

Nobody publishes credible time-to-hire data for any of these four. That market is served entirely by companies selling you speed. If you find a number, check who benefits from you believing it.

On AI in hiring

Every application arrives polished now, and the instinct is to buy a detector. Detectors are software you paste an application into. They return a score claiming how likely it is that a machine wrote it, and a growing number of employers screen people out on that score.

They wrongly flag around 61% of people writing English as a second language. Not people who cheated. People who wrote it themselves. If you are hiring in Manila, Gurgaon, Bangkok or Dubai, that is not a safeguard. It is a filter pointed at your own shortlist, and it never tells you it fired.

Two more findings change what to do instead. Graders systematically mark genuinely excellent work down as “too perfect” and assume a machine made it, so the false positive is the expensive error here. And in a controlled test, candidates instructed to cheat passed 73% of stock questions and 25% of novel ones, which is below the honest baseline. Nobody caught them. Writing your own question beat every detection method published.

So: write your own questions, score against anchors written beforehand, and interview live. That is what these kits are. If you want a work sample, let them use AI and then ask for a short extension without it, and score the gap. You cannot fake retention.