Stress-Free Interviewing of Software Developers (Part 2): What Actually Works? (English)

Stress-Free Interviewing of Software Developers (Part 2): What Actually Works?

Friday, 02 October 2026

//

18 minute read

In Part 1, I suggested a fairly simple approach to interviewing developers: prepare properly, make the process clear, and let candidates discuss code they understand. Then explore the decisions behind it.

I'm now researching how parts of that process could be automated using skill ledgers: a record of what a role needs, what the candidate claims, and what the interview has actually established.

Before building that, there's a more basic question. What does the research support? Can we reduce candidate stress, examine enough breadth, and still find out whether someone can do the job?

The evidence supports structured, job-related assessment. It also gives us reasons to question interview formats that add pressure unrelated to the work. But it does NOT establish that a pleasant conversation, a coding test, or an AI interviewer automatically predicts developer performance.

That distinction matters. I'm interested in collecting useful evidence with the least unnecessary burden. Automating a poor interview just makes it easier to inflict on more people.

NOTE: Draft This is another research piece, mainly reviewing evidence about developer interviews and proposing a skill-ledger approach. The proposed process has not been validated.

What the research can actually tell us

There are two overlapping bodies of research here. Personnel psychology has decades of work on selection, structured interviews and later job performance. Research specifically on developer interviews is smaller, with studies of stress, communication, preparation and candidate experience.

Those studies answer different questions. A candidate feeling more comfortable is one outcome. Producing a better explanation is another. Predicting how they will perform six months into a job is another again.

This is a selective research review for interview design, not a systematic review or a new validation study. I use primary research where available, distinguish software studies from broader employment research, and label the ledger design below as a proposal.

Structure has evidence behind it

Sackett and colleagues' 2022 reanalysis challenged overcorrections in earlier estimates of how well selection methods predict job performance. Structured interviews still ranked highly, but the revised estimates were more modest. 1

Selection method Revised mean operational validity
Structured interview 0.42
Job knowledge test 0.40
Work sample test 0.33
General cognitive ability test 0.31
Unstructured interview 0.19

These are estimated correlations, not percentages of correct hiring decisions. They combine occupational settings and depend on statistical corrections. The authors also report substantial variation between settings. You cannot assign your interview a validity of 0.42 because you called it structured.

For developer hiring, I take this as a reason to define the job and the assessment carefully. It doesn't settle whether your particular debugging exercise, portfolio conversation or system-design question works.

Structure includes how you score

Campion, Palmer and Campion describe interview structure as a collection of practices affecting both question content and evaluation. Job analysis, consistent administration, anchored ratings and controlled follow-ups are part of that picture. 2

My proposed translation into a developer interview is straightforward:

  • Decide which capabilities the role actually requires.
  • Give candidates comparable opportunities to demonstrate them.
  • Define what adequate evidence looks like before the interview.
  • Record evidence against each capability.
  • Score using those standards before discussing an overall decision.

Reading a script and then deciding whether you “liked their energy” leaves the evaluation largely unstructured.

Interview stress can change the result

The most directly relevant experiment is Does Stress Impact Technical Interview Performance?, by Behroozi, Shirolkar, Barik and Parnin. It compared private and observed whiteboard problem-solving in a randomised study with 48 computer science students. The observed condition also required thinking aloud. Median correctness was less than half that in the private condition, with higher reported stress and cognitive-load measures. 3

That's a serious finding, with important boundaries: one student sample, one algorithmic task, and a bundled change in observation and verbalisation. It does not isolate “being watched” from every other difference, establish the same effect for experienced developers, or measure later workplace performance.

It does demonstrate that assessment conditions can substantially change the performance you observe. Treating the score as an uncontaminated measure of programming ability is therefore a fairly large assumption.

Broader research points in the same direction. Powell, Stanley and Brown's meta-analysis reported an overall correlation of approximately −0.19 between self-reported interview anxiety and interview performance. That's an association, not proof that anxiety caused every lower score. 4

Schneider, Powell and Bonaccio then examined interview anxiety and later performance among applicants hired as residence assistants. Anxiety had near-zero correlations with job-performance ratings. Some moderation effects depended on the performance dimension and rater. This was not a software study, but it challenges the assumption that interview anxiety reliably signals poor workplace performance. 5

My implication for interviewing is simple: if you want to assess performance under pressure, define the pressure the job actually involves. An incident-response role may justify a realistic incident exercise. That does not establish the relevance of surprise algorithm questions and silent observation.

What helps candidates show their ability

Make the assessment criteria visible

Klehe and colleagues studied transparency in structured interviews in two samples, with 123 and 269 participants. Revealing the dimensions being assessed improved interview performance and supported construct validity. The second study found no significant difference in criterion-related validity between transparent and nontransparent conditions. These were application-training settings, rather than developer hiring. 6

That supports disclosing the capabilities being assessed. It doesn't prove that publishing every exact question preserves validity, or that advance preparation never creates problems.

For a developer role, I'd send something like:

We'll discuss how you diagnose problems, make design decisions, test changes and work with other people. Please choose an example you know well. We'll ask about your contribution, constraints, alternatives and how you checked the outcome. You can use notes. We can provide a sample if you cannot share your own work.

That's a proposed briefing. It lets the candidate prepare relevant evidence instead of trying to guess the interviewer's favourite technology.

Give people time to think before explaining

A follow-up software study, Asynchronous Technical Interviews, compared 24 recorded submissions with 24 earlier observed whiteboard sessions. The authors found improvements in communication informativeness and stress indicators, with technical performance preserved or improved on some measures. 7

This wasn't a fresh randomised comparison of identical conditions: cohorts, tasks, tools and time allowances differed. It provides promising evidence about removing live supervision, rather than a universal endorsement of asynchronous interviewing.

It also studied recorded technical explanations. It does NOT validate an AI avatar asking generic questions or grading personality from a video.

My design inference is to separate working time from explanation time where the job permits it. Let someone inspect the problem quietly, make notes, then discuss their approach. Offer an asynchronous technical route where appropriate, with a comparable alternative for candidates who find recording itself difficult.

Help people retrieve a specific example

“Tell me about a time...” looks simple until someone has to search years of experience, select an acceptable story and narrate it under evaluation.

Brosy, Bangerter and Ribeiro found that only half the responses to past-behaviour questions in their field sample were stories. In subsequent simulated interviews, probing elicited more stories and more narrative detail. Information about the expected format had mixed effects; retrieving a suitable example was a major difficulty. The studies concern eliciting responses, not predicting developer job performance. 8

My practical response would be to offer concrete prompts:

Pick a bug where your first explanation turned out to be wrong. What did you observe, what did you try, and what changed your mind?

If they struggle to retrieve one, let them choose another relevant case, use notes or return to it later. Keep track of whether you supplied clarification or a substantive technical hint.

An incomplete answer is a reason to clarify what evidence is missing. It needn't become an immediate verdict about competence.

Make questions accessible

Maras and colleagues evaluated adaptations to employment interview questions for autistic and non-autistic adults. More explicit questions and written prompts were associated with better answers in both groups, particularly autistic participants. The 50 participants undertook mock interviews around six months apart; this was an initial evaluation with a repeated-interview design, not proof of better hiring outcomes. 9

For the proposed system, I'd support written questions, one question at a time, explicit wording and pauses. Accessibility preferences should control presentation and pacing. They should not become hidden variables in a competence score.

Breadth should come from the job

The research does not give us a magic number of technologies, questions or interview rounds.

My proposal is to define breadth through the responsibilities of the vacancy. A backend developer might need to reason about implementation, data, testing, operation and collaboration. A compiler engineer needs a different profile. A senior engineer and an entry-level candidate need different evidence opportunities.

That is very different from accumulating everything anyone on the hiring panel happens to know.

For example, discussing a change to an order-processing service can provide several kinds of evidence:

Capability Evidence to explore
Understanding requirements Clarifies duplicate orders and acceptable failure behaviour
Data and consistency Explains transaction boundaries and concurrent updates
Verification Proposes tests that distinguish success from plausible failure
Operation Identifies useful logs, metrics and rollback conditions
Collaboration Explains how assumptions and risks were communicated

This is an illustrative role blueprint, not a validated instrument. It gives breadth a purpose.

One case can also hide gaps. If a candidate never touches concurrency, you need a planned opportunity to assess it. If the case is unusually familiar, use a bounded variation to explore transfer. The point is to sample the capabilities deliberately.

I would distinguish required on arrival, learnable during onboarding, and useful but optional. A missing library name should only be decisive when familiarity with that library is genuinely necessary for the role.

Depth should establish something

In Part 1, I advocated discussing familiar code. I still think it's a useful starting point. But personal experience with an approach is not a validation study, and candidates bring very different artefacts.

For comparison, I'd keep a shared role blueprint and a common small scenario alongside the candidate's own example. Offer a supplied artefact to anyone without shareable work. Confidentiality and lack of hobby time should not block access to the assessment.

A proposed sequence of probes is:

  1. Contribution: What did you personally change or decide?
  2. Mechanism: How does it work, including the relevant boundary conditions?
  3. Choice: Which alternative did you consider, and why reject it?
  4. Verification: What evidence showed the change worked?
  5. Transfer: What would you reconsider if one important constraint changed?

For an idempotent message handler, a useful variation might be: “Two workers receive the same message concurrently. Walk me through what happens.”

That probe has a stated purpose: exploring concurrency reasoning. Asking progressively obscure questions until everyone fails produces a less interpretable boundary.

Levashina and colleagues' review covers both past-behaviour and situational questions, rating scales, bias and follow-up questions. It identifies unresolved issues around probing. We should not turn the general success of structured interviews into a claim that unrestricted adaptive questioning has already been validated. 10

My proposed compromise is bounded adaptation: common competencies and anchors, a planned set of probe purposes, limits on hints and depth, and a record of which path each candidate encountered. Different wording and different examples still need evidence of comparable difficulty and opportunity.

A pleasant experience is not the whole measurement

Candidate experience does matter. Hausknecht, Day and Thomas reviewed 86 independent samples involving 48,750 participants. Positive perceptions of selection were associated with organisational attractiveness, intentions to accept offers and willingness to recommend the employer. Interviews and work samples were generally viewed favourably. These are associations, not proof that making an interview nicer causes better hires. 11

Preparation creates another burden. Bell and colleagues' survey of 131 candidates found that candidates rarely practised in authentic technical interview settings and reported stress and limited preparedness. It describes preparation experiences rather than establishing which method predicts job performance. 12

My implication is to count the candidate's whole investment: preparation, tasks, interviews and repeated explanation. Another round should resolve a specified uncertainty. “We always do five” is a process description, not evidence that the fifth contributes useful information.

We also need to keep the technical bar clear. A friendly interviewer can still favour someone who feels familiar. A polished explanation can still be wrong. Reducing unnecessary stress should make it easier to observe capability while preserving the same relevant standards.

Turning this into a skill ledger

This is the design I'm exploring. The studies above support several ingredients. They do not validate the complete automated system.

I'd begin with two linked records: the role's assessment requirements and the candidate's evidence. The candidate's claimed skills help locate relevant examples; they don't define the hiring standard.

For each capability, the assessment record should preserve:

  • The role requirement and expected level.
  • The question or task used to assess it.
  • The candidate's exact answer span or artefact reference.
  • The evidence type, such as self-report, observed reasoning or inspected code.
  • What that evidence supports, including its limits.
  • Contradictory evidence, assistance supplied and unresolved questions.
  • A rubric rating, recorded separately from confidence in that rating.

I would keep evidence status separate from proficiency:

Evidence status Meaning
Not assessed No adequate opportunity to demonstrate this capability
Partially assessed Relevant material exists, but important evidence is missing
Assessed Enough material exists to apply the rubric
Conflicting Evidence points in different directions and needs review

“Assessed” can produce a strong or weak rating. “Not assessed” should not silently become zero.

For example, “we used Kafka” is a self-reported technology claim. Explaining how a consumer handled duplicates provides reasoning evidence. Showing relevant code adds artefact evidence. Discussing a production incident adds further self-report. Those entries have different evidential strengths, and a ledger should retain the distinction.

A model's confidence that it extracted a statement correctly is also different from evidence that the candidate can perform the work. A fluent answer and a high classification score don't close that gap.

Choosing the next question

The next-question policy could use the ledger to identify a specific gap:

Current evidence Purpose of the next step
Broad claim without an example Ask for one concrete case
Mechanism described but outcome unclear Ask how it was checked
Personal contribution ambiguous Clarify ownership and team involvement
Required competency not covered Move to a planned scenario
Evidence conflicts Ask a neutral clarification
Relevant evidence is sufficient Move on

These are proposed policies. They give automation something more concrete than “make the interview interesting”.

An LLM could help map answers to the ledger and suggest probes. Deterministic rules can enforce the required coverage, time budget and permitted probe types. A decision model could classify whether an answer contains a concrete example, a stated constraint or verification evidence. These extraction tasks would need their own labelled evaluations.

The system should retain source evidence so a reviewer can challenge the mapping. It should also let the candidate correct a transcription or an incorrect account of their contribution.

For an automated interviewer, the same controls need to govern the actual conversation. It should explain the format, allow pauses and clarification, and stop a topic once the agreed evidence requirement is met. A human-assisted version provides a useful first test because the interviewer can reject bad probes while we measure where the system fails.

The stopping rule needs care. Ending early for an apparently impressive candidate can deprive everyone else of equal opportunity. I'd require common minimum coverage and bounded probes, with a clear maximum duration. If important evidence is still missing, record that explicitly and arrange a targeted follow-up where justified.

Scoring needs structure too

Kuncel and colleagues' meta-analysis found that mechanical combination of assessment information generally predicted outcomes better than holistic combination, including decisions made by knowledgeable experts. This concerns combining evidence; it does not validate the measurements being combined or imply that an AI judge is accurate. 13

For this design, I'd have assessors score dimensions independently using agreed anchors, then combine the scores using a declared rule. For example, testing judgement might range from “cannot explain how correctness would be checked” to “proposes discriminating tests and identifies important failure cases”. The exact anchors need role-specific development and calibration.

Any critical minimums must be set in advance. Otherwise a weighted average can conceal a serious gap, or a panel can invent a new must-have after meeting a candidate it doesn't like.

Replacing the panel's intuition with an LLM's intuition would leave the measurement problem unresolved. Automatically populating a ledger is a separate claim from automatically making a sound hiring decision.

How we would know whether it works

I'd evaluate the system in stages.

First, check whether the ledger accurately records the evidence. Have independent assessors annotate answer spans and artefacts. Measure incorrect mappings, missed evidence and disagreements. Include different communication styles, transcription quality and accessibility needs.

Second, compare the interview process with the existing one: candidate-reported stress, time burden, competency coverage, quality of elicited evidence and independent scorer agreement. A quieter interview with less evidence would be an incomplete success.

Third, examine later job outcomes using role-relevant measures and reviewers who don't know the interview score. Assess quality of changes, debugging, collaboration and operational judgement where relevant. Lines of code or ticket counts alone are inadequate for this proposed evaluation.

That last stage is difficult: we usually see outcomes only for hired candidates, teams differ, and hiring decisions restrict the range of scores. The validation plan must account for those limits. Agreement with historical hiring decisions would establish imitation of the old process, not predictive validity.

I'd also test whether each extra probe adds evidence, whether adaptive paths have comparable difficulty, whether assistance is distributed consistently, and whether AI-assisted answers change what we are actually assessing. Permitted tools should be declared before the interview and matched to the capability under assessment.

What I would build first

The research gives a defensible starting point: job-derived criteria, transparent expectations, accessible questions, deliberate scoring and opportunities to reason without unnecessary live performance pressure.

For the ledger system, I'd begin with a structured conversation around familiar work plus a common role-related scenario. Map answers to evidence, suggest bounded follow-ups and show what remains unresolved. Then test that process before giving automation responsibility for the final decision.

We can keep the conversation human while making the evidence inspectable.

And if a candidate has already shown you what you needed to know, you can stop asking them to prove it AGAIN.

Research sources

Sources checked on 2 October 2026. Publication years refer to the research, not the date a repository copy was uploaded. The practical protocol and skill ledger above are design proposals, not findings from a trial of this system.

  1. Sackett, P. R., Zhang, C., Berry, C. M., and Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107, 2040–2068. DOI · Institutional record.
  2. Campion, M. A., Palmer, D. K., and Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50, 655–702. DOI.
  3. Behroozi, M., Shirolkar, S., Barik, T., and Parnin, C. (2020). Does stress impact technical interview performance? ESEC/FSE. DOI · Author manuscript.
  4. Powell, D. M., Stanley, D. J., and Brown, K. N. (2018). Meta-analysis of the relation between interview anxiety and interview performance. Canadian Journal of Behavioural Science, 50, 195–207. DOI. The reported coefficient is also discussed in the open-access study below.
  5. Schneider, L., Powell, D. M., and Bonaccio, S. (2019). Does interview anxiety predict job performance and does it influence the predictive validity of interviews? International Journal of Selection and Assessment, 27, 328–336. Open-access article.
  6. Klehe, U.-C., König, C. J., Richter, G. M., Kleinmann, M., and Melchers, K. G. (2008). Transparency in structured interviews: Consequences for construct and criterion-related validity. Human Performance, 21, 107–137. DOI · Institutional record.
  7. Behroozi, M., Brown, C., and Parnin, C. (2022). Asynchronous technical interviews: Reducing the effect of supervised think-aloud on communication ability. ESEC/FSE. DOI · Author manuscript.
  8. Brosy, J., Bangerter, A., and Ribeiro, S. (2020; published online 2019). Encouraging the production of narrative responses to past-behaviour interview questions: Effects of probing and information. European Journal of Work and Organizational Psychology. DOI · Institutional manuscript.
  9. Maras, K., Norris, J. E., Nicholson, J., Heasman, B., Remington, A., and Crane, L. (2021; published online 2020). Ameliorating the disadvantage for autistic job seekers: An initial evaluation of adapted employment interview questions. Autism, 25, 1060–1075. DOI.
  10. Levashina, J., Hartwell, C. J., Morgeson, F. P., and Campion, M. A. (2014). The structured employment interview: Narrative and quantitative review of the research literature. Personnel Psychology, 67, 241–293. DOI.
  11. Hausknecht, J. P., Day, D. V., and Thomas, S. C. (2004). Applicant reactions to selection procedures: An updated model and meta-analysis. Personnel Psychology, 57, 639–683. DOI · Cornell manuscript.
  12. Bell, B., Thomas, T., Lee, S. W., and Brown, C. (2025). How do software engineering candidates prepare for technical interviews? Research manuscript. Findings cited here concern the reported survey; they do not establish job-performance prediction.
  13. Kuncel, N. R., Klieger, D. M., Connelly, B. S., and Ones, D. S. (2013). Mechanical versus clinical data combination in selection and admissions decisions: A meta-analysis. Journal of Applied Psychology, 98, 1060–1072. DOI · PubMed record.
Finding related posts...
logo

© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.