पाठशाला Pathshala · दल Dal, The team · Lesson 08 · Start
Interviewing for judgement, not pedigree
College names and the friendly hour predict little about how someone will do the job. How to build a scorecard, a structured interview, a paid work sample and reference checks that predict performance better.
Pathshala, The Founder Library · 11 October 2026 · 8 min read

Most early hiring decisions rest on two things that predict little: where a candidate studied or worked, and how an hour with them felt. A century of research on what predicts job performance has a plain message for a startup. Ask everyone the same questions, score the answers against a rubric written in advance, and watch the candidate do a piece of the real work.
This lesson sets out the evidence, then the instruments a small company can run without an HR team: the scorecard, the structured interview and the paid work sample, with reference checks that add information rather than ceremony.
What pedigree predicts, and what it does not
Pedigree is a proxy. An IIT degree or two years at a well-known product company tells you that someone cleared a hard filter once, usually years ago, under conditions unlike yours. It is useful at the top of the funnel when four hundred applications arrive and there is no time to read them. It is weak at the decision, because the question at the decision is whether this person will do this job well, in this company, starting next month.
The research has not measured college names directly, but it has measured the nearest thing a CV records: years of job experience. In the largest recent re-estimate of selection research, years of experience correlated with later performance at 0.07, close to nothing. The unstructured interview, the friendly hour most founders run, came in at 0.19. Between them they are what most startups rely on, and both sit near the bottom of the list.
There is a second cost. Hiring on pedigree in India shrinks the pool to the same few thousand people that every funded startup is chasing, at the prices they command. Judgement is distributed far more widely than pedigree. A process that can see it finds people the competition has filtered out.
The evidence on what predicts performance
For twenty-five years the standard reference was Frank Schmidt and John Hunter’s 1998 meta-analysis, which put work-sample tests at the top with a validity of 0.54 and cognitive ability and structured interviews at 0.51. In 2022 Paul Sackett and colleagues re-examined the same literature and showed that the earlier work had over-corrected for range restriction, inflating many estimates. Their revised figures changed the order: structured interviews first at 0.42, job knowledge tests at 0.40, empirically keyed biodata at 0.38, work samples at 0.33 and cognitive ability at 0.31. Unstructured interviews fell to 0.19 and years of experience to 0.07.
Two readings matter. The first is the ranking: instruments that look like the job and are scored consistently beat instruments that look like a conversation. The second is humility. A validity of 0.42 means the best single method explains under a fifth of the variation in later performance. No loop is a crystal ball, which is why the loop combines methods and why the first ninety days are the last stage of assessment rather than the first stage of the job.
Toggle between the two sets of estimates. The structured interview and the job knowledge test hold their places; the work sample falls but stays well above the unstructured interview. Switch to the share explained and the gap widens: in the 2022 figures the structured interview explains nearly five times as much of the variation in performance as the unstructured one, and years of experience explains less than one per cent.
Write the scorecard before the job post
Judgement cannot be assessed in general. It can be assessed against a specific job. Before anyone is interviewed, write one page: the mission of the role in a sentence; three to five outcomes the person must deliver in the first twelve months, each with a number or a date; and four to six competencies the outcomes require, each defined by what good looks like. For a first sales hire an outcome might be “₹60 lakh of new annual contract value closed by month twelve, at least half from accounts the founder did not source”, and a competency “runs a discovery call that ends with a next step the buyer has agreed to”.
The scorecard decides everything downstream. The questions test the competencies. The work sample produces a piece of an outcome. The rubric scores against the definitions of good. A candidate who is impressive but scores weakly against the scorecard is impressive at a different job. Write it with whoever will manage the hire, and do not change it after the first interview unless you are willing to restart every candidate on the new version.
The structured interview, in an hour
Google’s re:Work guide gives the working definition: the same questions for every candidate for a role, answers graded on the same scale, and decisions made against qualifications set in advance. It names two kinds of question. Behavioural questions ask what a candidate did (“tell me about a time when…”) and reveal patterns. Hypothetical questions put a situation from the job and show how the candidate handles something new. Google reports that prepared questions and rubrics save interviewers an average of forty minutes per interview, and that rejected candidates who had a structured interview were 35 per cent happier than those who did not.
For a startup the hour has four parts. Ten minutes on the candidate’s own account of the most relevant thing they built or sold, with follow-up on what they chose and why. Thirty minutes on four questions, one per competency, each asked identically: two behavioural and two hypothetical. Ten minutes on a problem from the actual business: a real customer email, a metric that moved last month, a trade-off you faced. Ten minutes for their questions, which are themselves evidence of judgement.
The rubric for each question describes a poor, an adequate, a strong and an exceptional answer in a sentence each, written before the first candidate. Each interviewer scores alone, in writing, before any discussion. The scores are the evidence; the debrief exists to resolve disagreements, not to discover what everyone felt. A founder who cannot resist saying “I really liked her” in the debrief should score first and speak last.
The work sample: paid, real and short
A work sample is a piece of the job, done by the candidate and judged against the scorecard. Sam Altman’s advice on hiring is to have people audition for roles rather than interview for them: a day or two of paid contract work, on a real but non-critical project for an engineer, or a press release and a list of reporters to pitch for a communications role. The 2022 estimates rank the work sample lower than the 1998 ones did, but it does what no interview can. It shows the candidate in the conditions of the job, with ambiguity and a deadline, and it shows the candidate the job, which means fewer declined offers and fewer departures in month three.

Rules that keep it fair. Pay for it at a day rate, always. Keep it to two days or less, and to a few hours for a candidate in a job. Use real work but not critical work, so the company is not extracting free labour and the candidate is not carrying risk. Give every candidate the same brief and the same materials. Weigh the questions they asked before starting as heavily as the output: the candidate who asks what the result is for has shown judgement before writing a line.
Examples by role. Engineer: a scoped feature on a copy of the codebase, with a short note on what they left out and why. Sales: a discovery call with a founder playing a real buyer from the pipeline, then the follow-up email. Operations: last month’s actual fulfilment exceptions, anonymised, and a proposal for the two to fix first. Marketing: a campaign brief for one real segment, with the metric it would move and the budget it would need.
Pedigree tells you someone passed a hard test years ago. A work sample tells you what they will do next month. Hire on the second.
References that tell you something
A reference check is a structured interview with a different subject. Ask for references who managed the candidate rather than peers they chose, and ask each the same three questions. What was this person’s best piece of work, and what made it good? Where did you have to support or correct them? Would you hire them again for this role, and what would you set up differently? Listen to the first two seconds of the answer to the third. Enthusiasm is immediate; hesitation is information.
In India the candidate’s current employer is usually off limits until an offer is accepted, so the useful references are the previous manager and a customer or partner who saw the work. Ask the candidate’s permission before calling anyone, tell them whom you will call, and record what you heard against the scorecard competencies the same day. References are the last step, never a tiebreaker for an interview that produced no evidence.
The loop, run every time
For every role: the scorecard, agreed by the hiring manager before the job is posted. A thirty-minute screen by phone against two competencies. The structured hour, with two interviewers scoring independently. The paid work sample, briefed identically. Two manager references. A decision within forty-eight hours of the last step, made from the written scores.
Then the quarterly check that makes the loop improve. For every hire made in the past year, put their interview scores beside their performance at six months. Where scores and performance agree, keep the questions. Where they disagree, rewrite the question or its rubric. Five hires are enough to expose a question that predicts nothing, and removing it improves the next ten. The [first engineer lesson](/library/hiring-first-engineer-below-market) applies the loop to the hardest single hire; the [first ten hires](/library/first-ten-hires-who-and-in-what-order) say which role it should run on next.
The validity figures are averages across many jobs and studies. They describe what tends to predict performance, not what will in any single hire.
Sources
- HumRRO, Is cognitive ability the best predictor of job performance? (on Sackett, Zhang, Berry and Lievens, Journal of Applied Psychology, 2022), September 2022
- Master International, New study providing updated validity estimates (Schmidt and Hunter 1998 against Sackett et al.), March 2022
- Google re:Work, A guide to structured interviewing
- Sam Altman, How to Hire