This is a viewer only at the moment see the article on how this works.
To update the preview hit Ctrl-Alt-R (or ⌘-Alt-R on Mac) or Enter to refresh. The Save icon lets you save the markdown file to disk
This is a preview from the server running through my markdig pipeline
Friday, 02 October 2026
In Part 1, I suggested a fairly simple approach to interviewing developers: prepare properly, make the process clear, and let candidates discuss code they understand. Then explore the decisions behind it.
I'm now researching how parts of that process could be automated using skill ledgers: a record of what a role needs, what the candidate claims, and what the interview has actually established.
Before building that, there's a more basic question. What does the research support? Can we reduce candidate stress, examine enough breadth, and still find out whether someone can do the job?
The evidence supports structured, job-related assessment. It also gives us reasons to question interview formats that add pressure unrelated to the work. But it does NOT establish that a pleasant conversation, a coding test, or an AI interviewer automatically predicts developer performance.
That distinction matters. I'm interested in collecting useful evidence with the least unnecessary burden. Automating a poor interview just makes it easier to inflict on more people.
NOTE: Draft This is another research piece, mainly reviewing evidence about developer interviews and proposing a skill-ledger approach. The proposed process has not been validated.
There are two overlapping bodies of research here. Personnel psychology has decades of work on selection, structured interviews and later job performance. Research specifically on developer interviews is smaller, with studies of stress, communication, preparation and candidate experience.
Those studies answer different questions. A candidate feeling more comfortable is one outcome. Producing a better explanation is another. Predicting how they will perform six months into a job is another again.
This is a selective research review for interview design, not a systematic review or a new validation study. I use primary research where available, distinguish software studies from broader employment research, and label the ledger design below as a proposal.
Sackett and colleagues' 2022 reanalysis challenged overcorrections in earlier estimates of how well selection methods predict job performance. Structured interviews still ranked highly, but the revised estimates were more modest. 1
| Selection method | Revised mean operational validity |
|---|---|
| Structured interview | 0.42 |
| Job knowledge test | 0.40 |
| Work sample test | 0.33 |
| General cognitive ability test | 0.31 |
| Unstructured interview | 0.19 |
These are estimated correlations, not percentages of correct hiring decisions. They combine occupational settings and depend on statistical corrections. The authors also report substantial variation between settings. You cannot assign your interview a validity of 0.42 because you called it structured.
For developer hiring, I take this as a reason to define the job and the assessment carefully. It doesn't settle whether your particular debugging exercise, portfolio conversation or system-design question works.
Campion, Palmer and Campion describe interview structure as a collection of practices affecting both question content and evaluation. Job analysis, consistent administration, anchored ratings and controlled follow-ups are part of that picture. 2
My proposed translation into a developer interview is straightforward:
Reading a script and then deciding whether you “liked their energy” leaves the evaluation largely unstructured.
The most directly relevant experiment is Does Stress Impact Technical Interview Performance?, by Behroozi, Shirolkar, Barik and Parnin. It compared private and observed whiteboard problem-solving in a randomised study with 48 computer science students. The observed condition also required thinking aloud. Median correctness was less than half that in the private condition, with higher reported stress and cognitive-load measures. 3
That's a serious finding, with important boundaries: one student sample, one algorithmic task, and a bundled change in observation and verbalisation. It does not isolate “being watched” from every other difference, establish the same effect for experienced developers, or measure later workplace performance.
It does demonstrate that assessment conditions can substantially change the performance you observe. Treating the score as an uncontaminated measure of programming ability is therefore a fairly large assumption.
Broader research points in the same direction. Powell, Stanley and Brown's meta-analysis reported an overall correlation of approximately −0.19 between self-reported interview anxiety and interview performance. That's an association, not proof that anxiety caused every lower score. 4
Schneider, Powell and Bonaccio then examined interview anxiety and later performance among applicants hired as residence assistants. Anxiety had near-zero correlations with job-performance ratings. Some moderation effects depended on the performance dimension and rater. This was not a software study, but it challenges the assumption that interview anxiety reliably signals poor workplace performance. 5
My implication for interviewing is simple: if you want to assess performance under pressure, define the pressure the job actually involves. An incident-response role may justify a realistic incident exercise. That does not establish the relevance of surprise algorithm questions and silent observation.
Klehe and colleagues studied transparency in structured interviews in two samples, with 123 and 269 participants. Revealing the dimensions being assessed improved interview performance and supported construct validity. The second study found no significant difference in criterion-related validity between transparent and nontransparent conditions. These were application-training settings, rather than developer hiring. 6
That supports disclosing the capabilities being assessed. It doesn't prove that publishing every exact question preserves validity, or that advance preparation never creates problems.
For a developer role, I'd send something like:
We'll discuss how you diagnose problems, make design decisions, test changes and work with other people. Please choose an example you know well. We'll ask about your contribution, constraints, alternatives and how you checked the outcome. You can use notes. We can provide a sample if you cannot share your own work.
That's a proposed briefing. It lets the candidate prepare relevant evidence instead of trying to guess the interviewer's favourite technology.
A follow-up software study, Asynchronous Technical Interviews, compared 24 recorded submissions with 24 earlier observed whiteboard sessions. The authors found improvements in communication informativeness and stress indicators, with technical performance preserved or improved on some measures. 7
This wasn't a fresh randomised comparison of identical conditions: cohorts, tasks, tools and time allowances differed. It provides promising evidence about removing live supervision, rather than a universal endorsement of asynchronous interviewing.
It also studied recorded technical explanations. It does NOT validate an AI avatar asking generic questions or grading personality from a video.
My design inference is to separate working time from explanation time where the job permits it. Let someone inspect the problem quietly, make notes, then discuss their approach. Offer an asynchronous technical route where appropriate, with a comparable alternative for candidates who find recording itself difficult.
“Tell me about a time...” looks simple until someone has to search years of experience, select an acceptable story and narrate it under evaluation.
Brosy, Bangerter and Ribeiro found that only half the responses to past-behaviour questions in their field sample were stories. In subsequent simulated interviews, probing elicited more stories and more narrative detail. Information about the expected format had mixed effects; retrieving a suitable example was a major difficulty. The studies concern eliciting responses, not predicting developer job performance. 8
My practical response would be to offer concrete prompts:
Pick a bug where your first explanation turned out to be wrong. What did you observe, what did you try, and what changed your mind?
If they struggle to retrieve one, let them choose another relevant case, use notes or return to it later. Keep track of whether you supplied clarification or a substantive technical hint.
An incomplete answer is a reason to clarify what evidence is missing. It needn't become an immediate verdict about competence.
Maras and colleagues evaluated adaptations to employment interview questions for autistic and non-autistic adults. More explicit questions and written prompts were associated with better answers in both groups, particularly autistic participants. The 50 participants undertook mock interviews around six months apart; this was an initial evaluation with a repeated-interview design, not proof of better hiring outcomes. 9
For the proposed system, I'd support written questions, one question at a time, explicit wording and pauses. Accessibility preferences should control presentation and pacing. They should not become hidden variables in a competence score.
The research does not give us a magic number of technologies, questions or interview rounds.
My proposal is to define breadth through the responsibilities of the vacancy. A backend developer might need to reason about implementation, data, testing, operation and collaboration. A compiler engineer needs a different profile. A senior engineer and an entry-level candidate need different evidence opportunities.
That is very different from accumulating everything anyone on the hiring panel happens to know.
For example, discussing a change to an order-processing service can provide several kinds of evidence:
| Capability | Evidence to explore |
|---|---|
| Understanding requirements | Clarifies duplicate orders and acceptable failure behaviour |
| Data and consistency | Explains transaction boundaries and concurrent updates |
| Verification | Proposes tests that distinguish success from plausible failure |
| Operation | Identifies useful logs, metrics and rollback conditions |
| Collaboration | Explains how assumptions and risks were communicated |
This is an illustrative role blueprint, not a validated instrument. It gives breadth a purpose.
One case can also hide gaps. If a candidate never touches concurrency, you need a planned opportunity to assess it. If the case is unusually familiar, use a bounded variation to explore transfer. The point is to sample the capabilities deliberately.
I would distinguish required on arrival, learnable during onboarding, and useful but optional. A missing library name should only be decisive when familiarity with that library is genuinely necessary for the role.
In Part 1, I advocated discussing familiar code. I still think it's a useful starting point. But personal experience with an approach is not a validation study, and candidates bring very different artefacts.
For comparison, I'd keep a shared role blueprint and a common small scenario alongside the candidate's own example. Offer a supplied artefact to anyone without shareable work. Confidentiality and lack of hobby time should not block access to the assessment.
A proposed sequence of probes is:
For an idempotent message handler, a useful variation might be: “Two workers receive the same message concurrently. Walk me through what happens.”
That probe has a stated purpose: exploring concurrency reasoning. Asking progressively obscure questions until everyone fails produces a less interpretable boundary.
Levashina and colleagues' review covers both past-behaviour and situational questions, rating scales, bias and follow-up questions. It identifies unresolved issues around probing. We should not turn the general success of structured interviews into a claim that unrestricted adaptive questioning has already been validated. 10
My proposed compromise is bounded adaptation: common competencies and anchors, a planned set of probe purposes, limits on hints and depth, and a record of which path each candidate encountered. Different wording and different examples still need evidence of comparable difficulty and opportunity.
Candidate experience does matter. Hausknecht, Day and Thomas reviewed 86 independent samples involving 48,750 participants. Positive perceptions of selection were associated with organisational attractiveness, intentions to accept offers and willingness to recommend the employer. Interviews and work samples were generally viewed favourably. These are associations, not proof that making an interview nicer causes better hires. 11
Preparation creates another burden. Bell and colleagues' survey of 131 candidates found that candidates rarely practised in authentic technical interview settings and reported stress and limited preparedness. It describes preparation experiences rather than establishing which method predicts job performance. 12
My implication is to count the candidate's whole investment: preparation, tasks, interviews and repeated explanation. Another round should resolve a specified uncertainty. “We always do five” is a process description, not evidence that the fifth contributes useful information.
We also need to keep the technical bar clear. A friendly interviewer can still favour someone who feels familiar. A polished explanation can still be wrong. Reducing unnecessary stress should make it easier to observe capability while preserving the same relevant standards.
This is the design I'm exploring. The studies above support several ingredients. They do not validate the complete automated system.
I'd begin with two linked records: the role's assessment requirements and the candidate's evidence. The candidate's claimed skills help locate relevant examples; they don't define the hiring standard.
For each capability, the assessment record should preserve:
I would keep evidence status separate from proficiency:
| Evidence status | Meaning |
|---|---|
| Not assessed | No adequate opportunity to demonstrate this capability |
| Partially assessed | Relevant material exists, but important evidence is missing |
| Assessed | Enough material exists to apply the rubric |
| Conflicting | Evidence points in different directions and needs review |
“Assessed” can produce a strong or weak rating. “Not assessed” should not silently become zero.
For example, “we used Kafka” is a self-reported technology claim. Explaining how a consumer handled duplicates provides reasoning evidence. Showing relevant code adds artefact evidence. Discussing a production incident adds further self-report. Those entries have different evidential strengths, and a ledger should retain the distinction.
A model's confidence that it extracted a statement correctly is also different from evidence that the candidate can perform the work. A fluent answer and a high classification score don't close that gap.
The next-question policy could use the ledger to identify a specific gap:
| Current evidence | Purpose of the next step |
|---|---|
| Broad claim without an example | Ask for one concrete case |
| Mechanism described but outcome unclear | Ask how it was checked |
| Personal contribution ambiguous | Clarify ownership and team involvement |
| Required competency not covered | Move to a planned scenario |
| Evidence conflicts | Ask a neutral clarification |
| Relevant evidence is sufficient | Move on |
These are proposed policies. They give automation something more concrete than “make the interview interesting”.
An LLM could help map answers to the ledger and suggest probes. Deterministic rules can enforce the required coverage, time budget and permitted probe types. A decision model could classify whether an answer contains a concrete example, a stated constraint or verification evidence. These extraction tasks would need their own labelled evaluations.
The system should retain source evidence so a reviewer can challenge the mapping. It should also let the candidate correct a transcription or an incorrect account of their contribution.
For an automated interviewer, the same controls need to govern the actual conversation. It should explain the format, allow pauses and clarification, and stop a topic once the agreed evidence requirement is met. A human-assisted version provides a useful first test because the interviewer can reject bad probes while we measure where the system fails.
The stopping rule needs care. Ending early for an apparently impressive candidate can deprive everyone else of equal opportunity. I'd require common minimum coverage and bounded probes, with a clear maximum duration. If important evidence is still missing, record that explicitly and arrange a targeted follow-up where justified.
Kuncel and colleagues' meta-analysis found that mechanical combination of assessment information generally predicted outcomes better than holistic combination, including decisions made by knowledgeable experts. This concerns combining evidence; it does not validate the measurements being combined or imply that an AI judge is accurate. 13
For this design, I'd have assessors score dimensions independently using agreed anchors, then combine the scores using a declared rule. For example, testing judgement might range from “cannot explain how correctness would be checked” to “proposes discriminating tests and identifies important failure cases”. The exact anchors need role-specific development and calibration.
Any critical minimums must be set in advance. Otherwise a weighted average can conceal a serious gap, or a panel can invent a new must-have after meeting a candidate it doesn't like.
Replacing the panel's intuition with an LLM's intuition would leave the measurement problem unresolved. Automatically populating a ledger is a separate claim from automatically making a sound hiring decision.
I'd evaluate the system in stages.
First, check whether the ledger accurately records the evidence. Have independent assessors annotate answer spans and artefacts. Measure incorrect mappings, missed evidence and disagreements. Include different communication styles, transcription quality and accessibility needs.
Second, compare the interview process with the existing one: candidate-reported stress, time burden, competency coverage, quality of elicited evidence and independent scorer agreement. A quieter interview with less evidence would be an incomplete success.
Third, examine later job outcomes using role-relevant measures and reviewers who don't know the interview score. Assess quality of changes, debugging, collaboration and operational judgement where relevant. Lines of code or ticket counts alone are inadequate for this proposed evaluation.
That last stage is difficult: we usually see outcomes only for hired candidates, teams differ, and hiring decisions restrict the range of scores. The validation plan must account for those limits. Agreement with historical hiring decisions would establish imitation of the old process, not predictive validity.
I'd also test whether each extra probe adds evidence, whether adaptive paths have comparable difficulty, whether assistance is distributed consistently, and whether AI-assisted answers change what we are actually assessing. Permitted tools should be declared before the interview and matched to the capability under assessment.
The research gives a defensible starting point: job-derived criteria, transparent expectations, accessible questions, deliberate scoring and opportunities to reason without unnecessary live performance pressure.
For the ledger system, I'd begin with a structured conversation around familiar work plus a common role-related scenario. Map answers to evidence, suggest bounded follow-ups and show what remains unresolved. Then test that process before giving automation responsibility for the final decision.
We can keep the conversation human while making the evidence inspectable.
And if a candidate has already shown you what you needed to know, you can stop asking them to prove it AGAIN.
Sources checked on 2 October 2026. Publication years refer to the research, not the date a repository copy was uploaded. The practical protocol and skill ledger above are design proposals, not findings from a trial of this system.
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.