Cited Works: Unstructured Interviews

Decades of peer-reviewed research — including work by a Nobel laureate — converge on a single finding: traditional unstructured interviews are foundationally flawed as a method of evaluating candidates, frequently performing worse than random selection from a qualified pool.

What follows is not an opinion. It is a conclusion drawn from decades of peer-reviewed, rigorously conducted research published in the most prestigious academic journals in psychology, behavioral economics, and organizational science. The studies cited below span more than eighty-five years of data and include work by researchers who have gone on to receive the Nobel Prize.

The evidence converges on a finding that is as close to conclusive as the social sciences can produce: traditional, unstructured interview techniques are foundationally flawed as a method of evaluating candidates. In many documented cases, they perform not merely poorly, but worse than random selection from a qualified pool — actively degrading the quality of hiring decisions compared to simply reviewing existing credentials alone.

Despite the overwhelming weight of this evidence, the vast majority of modern organizations continue to rely on these methods as a primary selection mechanism. The research below explains why that is a mistake.


The “Random Answers” Study

Jason Dana, Robyn Dawes & Nathaniel Peterson — Belief in the Unstructured Interview: The Persistence of an Illusion, 2013

Researchers at Yale and Carnegie Mellon designed an elegant experiment to test whether interviewers could distinguish between genuine candidate responses and pure noise. In the critical condition, interviewees were explicitly instructed to answer questions at random — selecting options through a hidden system that had no relationship to the truth of their answers.

The results were striking. Interviewers who spoke with candidates generating random answers reported the same level of confidence in their assessments as those who conducted genuine interviews. The human compulsion to construct a coherent narrative was so powerful that it operated even when the raw material was meaningless.

Most critically, the study demonstrated a “worse than random” effect: subjects who conducted an interview made significantly worse predictions about a candidate’s future GPA than subjects who simply reviewed background data — prior grades and course load — without meeting the candidate at all. The interview did not merely fail to add value; it actively diluted valid information, producing outcomes inferior to what a simple statistical model or random selection from the qualified pool would have achieved.

The University of Texas Natural Experiment

Medical School Performance of Initially Rejected StudentsJournal of the American Medical Association (JAMA), 1987

In 1979, the Texas state legislature directed the University of Texas Medical School at Houston to increase its incoming class from 150 to 200 students — after the admissions committee had already conducted interviews and sent rejection letters. The school was forced to admit fifty additional students drawn from the pool of applicants it had previously interviewed and explicitly rejected, creating a rare and ethically unplanned natural experiment.

Researchers followed both cohorts — the original “top 150” and the previously rejected “bottom 50” — through their entire medical school careers. The findings were unambiguous: there was no statistically significant difference between the two groups in academic grades, clinical rotation performance, or graduation rates.

The interview-based evaluation process, which the admissions committee had relied upon with confidence, had failed entirely to distinguish future high performers from those it had deemed unfit for admission. Fifty capable future physicians were nearly excluded by a selection method that, in retrospect, carried no meaningful predictive power.

Harvard Business Review: Taking the Bias Out of Interviews

Iris BohnetHow to Take the Bias Out of Interviews, Harvard Business Review, 2016

Professor Iris Bohnet, a behavioral economist at Harvard Kennedy School, synthesized decades of research — including the landmark work of Frank Schmidt and John Hunter — into an accessible and widely cited Harvard Business Review article. Her analysis presents the accumulated evidence in terms that are difficult to dismiss.

Bohnet reports that unstructured interviews exhibit a predictive validity of approximately 0.38 on a 0-to-1 scale — a figure that is remarkably low for a method that consumes enormous organizational time and resources. She explains the underlying mechanism: interviewers are prone to confirmation bias, typically forming a judgment within the first ten seconds and then spending the remainder of the conversation unconsciously seeking evidence to confirm that initial impression. This process systematically causes interviewers to discount more valid and reliable data points such as work samples and cognitive assessments.

Bohnet’s conclusion is unequivocal: “Unstructured interviews are essentially a waste of time.” She goes further, arguing that organizations unwilling to adopt structured evaluation formats would achieve comparable outcomes by replacing the interview process with a lottery.

The Illusion of Validity

Daniel KahnemanThinking, Fast and Slow, 2011 — Nobel Memorial Prize in Economic Sciences, 2002

Nobel laureate Daniel Kahneman drew on his own formative experience to articulate one of the most enduring concepts in the study of human judgment. As a young psychologist serving in the Israeli Defense Forces, Kahneman was tasked with predicting which military recruits would become successful officers, based on unstructured observation of an obstacle course exercise followed by a free-form interview.

Upon examining the data, Kahneman discovered that his team’s predictions were statistically no better than a coin flip. Yet despite repeated exposure to this evidence, the interviewers — himself included — could not shake the powerful subjective feeling that they were perceiving something real and meaningful about the candidates. He termed this phenomenon the “illusion of validity”: an intense, visceral sense of insight that persists even in the complete absence of predictive accuracy.

This work has become a cornerstone of the broader argument against intuitive judgment in high-stakes selection. The illusion of validity explains why organizations continue to defend interview-based hiring with conviction — the process feels informative even when the evidence consistently demonstrates that it is not. Confidence, Kahneman showed, is not a proxy for correctness; it is a psychological trap that provides the sensation of accuracy without the substance.

The Meta-Analysis: 85 Years of Selection Methods

Frank L. Schmidt & John E. HunterThe Validity and Utility of Selection Methods in Personnel Psychology: Practical and Theoretical Implications of 85 Years of Research Findings, 1998 (updated 2016)

This landmark meta-analysis is widely regarded as the gold standard in the field of personnel selection research. Schmidt and Hunter systematically evaluated the predictive power of nineteen different selection methods using eighty-five years of accumulated data — an evidence base of extraordinary breadth and depth.

Their findings established a clear hierarchy. General Mental Ability (GMA) tests and work sample tests emerged as the strongest predictors of job performance — the methods most closely aligned with evaluating what a candidate can actually do. Unstructured interviews, by contrast, ranked remarkably low, frequently providing less incremental validity than simply reviewing a candidate’s existing credentials on paper.

The researchers concluded that most interviewers systematically over-interpret social cues — warmth, articulateness, eye contact, perceived “culture fit” — that have little to no correlation with actual job success. The result is a selection process that rewards performance in the interview rather than performance in the role, and one that introduces substantial noise into decisions that have lasting organizational consequences.


The Consensus

Taken together, these works — spanning decades, institutions, and disciplines — converge on a single, well-substantiated conclusion. The “worse than random” phenomenon documented across this body of research arises from two deeply rooted psychological mechanisms:

Sensemaking. The human brain is extraordinarily skilled at constructing coherent narratives. In an interview setting, this means an evaluator will unconsciously assemble a compelling story of a candidate’s potential even from random, contradictory, or irrelevant information. The result is a confident assessment built on a foundation that does not exist.

The Dilution Effect. Irrelevant information — a shared hobby, a charming anecdote, a perceived “culture fit” — does not simply sit alongside diagnostic data; it actively weakens it. An interviewer who learns that a candidate shares their enthusiasm for a sports team will, on average, weight that candidate’s ten years of relevant technical experience less heavily than a reviewer who never encountered that irrelevant detail. The interview introduces noise that degrades the signal already present in the candidate’s record.

These are not theoretical risks. They are empirically documented, repeatedly replicated, and acknowledged by researchers at the highest levels of their fields — including a Nobel laureate. The conclusion is not that interviews can sometimes go wrong; it is that their fundamental structure, absent rigorous controls, is systematically incompatible with accurate evaluation.