Cited Works: Behavioral Interviews

Over 90% of candidates engage in faking behavior during behavioral interviews. Self-presentation tactics inflate ratings without predicting job performance. Interview ratings correlate twice as strongly with impression management as with job-related competence. The STAR method, now universally taught, has made the format so transparent that it grades preparation skill, not work ability. Four decades of research converge on a single finding: behavioral interviews measure the wrong thing, and they measure it reliably.

The behavioral interview was the industry’s corrective. After decades of research demonstrated that unstructured interviews are foundationally flawed — performing worse than random selection in many documented cases — organizations adopted the structured behavioral format as a remedy. “Tell me about a time when…” replaced open-ended conversation. The STAR method (Situation, Task, Action, Result) replaced intuitive assessment. The assumption was straightforward: if the problem with unstructured interviews was the interviewer’s unchecked subjectivity, then standardizing the questions and grounding them in past behavior would fix the signal.

It did not.

What the research demonstrates — consistently, across multiple laboratories, methodologies, and decades — is that behavioral interviews did not eliminate the distortion. They shifted it. Where unstructured interviews allowed the interviewer to construct a false narrative about a candidate, behavioral interviews invite the candidate to construct one about themselves. The failure mode changed. The failure did not.

The studies cited below examine what behavioral interviews actually measure. The answer is not job competence. It is impression management, coaching fluency, verbal performance, and the ability to deliver a rehearsed story under pressure. For roles where those traits are irrelevant to the work — software engineers, systems architects, data engineers, analysts — the behavioral interview evaluates a dimension that has no bearing on the job.


“Over Ninety Percent”

Julia Levashina & Michael A. CampionMeasuring Faking in the Employment Interview: Development and Validation of an Interview Faking Behavior Scale, Journal of Applied Psychology, 2007

Levashina and Campion set out to measure something that personnel psychologists had long suspected but rarely quantified: the extent to which candidates fabricate, embellish, or strategically distort their responses in behavioral interviews. Across six studies involving 1,346 participants, they developed and validated the Interview Faking Behavior (IFB) scale — a taxonomy of deception in the interview setting.

The taxonomy is revealing. Faking is not a single behavior; it is a spectrum with four distinct factors: Slight Image Creation (embellishing accomplishments, tailoring responses to match perceived job requirements), Extensive Image Creation (constructing fictional experiences, inventing achievements, borrowing others’ accomplishments), Image Protection (masking weaknesses, distancing from failures, omitting damaging information), and Ingratiation (conforming to the interviewer’s apparent values, enhancing the interviewer’s self-image). Eleven subfactors in total. The researchers did not discover a gap in the system; they mapped an ecosystem.

The central finding was stark: over 90% of undergraduate job candidates engaged in some form of faking during employment interviews. Fewer — between 28% and 75% depending on the subfactor — engaged in behaviors that crossed the line into outright fabrication. But the milder forms, the embellishments and omissions, were nearly universal. Most critically, faking behavior was positively correlated with receiving second interviews and job offers, but showed a low-to-zero correlation with subsequent job performance. The interview was reliably selecting for the skill it was inadvertently testing: the ability to fake.

When over ninety percent of candidates are performing, the performance is the test. The behavioral interview does not detect faking because faking is not an aberration within the format — it is the format’s dominant mode of engagement.

“What You See May Not Be What You Get”

Murray R. Barrick, Jonathan A. Shaffer & Sandra W. DeGrassi — What You See May Not Be What You Get: Relationships Among Self-Presentation Tactics and Ratings of Interview and Job Performance, Journal of Applied Psychology, 2009

The title of this paper is its thesis. Barrick, Shaffer, and DeGrassi examined two specific self-presentation tactics — self-promotion (verbal claims of competence, skill, and past achievement) and ingratiation (efforts to make the interviewer personally like the candidate) — and measured their relationship to two separate outcomes: the rating the candidate received in the interview, and the performance they subsequently delivered on the job.

The results confirmed what the Levashina and Campion taxonomy predicted. Candidates who engaged in high levels of self-promotion received significantly higher interview ratings. Their objective technical qualifications did not account for the difference — the inflation came from the presentation itself. Ingratiation produced a similar, if somewhat weaker, effect. The tactics worked. They raised scores.

But when the researchers followed through to actual job performance, the relationship collapsed. Self-presentation tactics that inflated interview ratings did not predict performance in the role. The person the interviewer evaluated was, in a measurable sense, not the person who showed up to do the work. The interview captured a performance — a literal one, in the theatrical sense — and the performance had no reliable connection to the competence it purported to demonstrate.

Barrick and colleagues noted that even structured behavioral formats were susceptible. The structure constrained the interviewer’s subjectivity, but it did nothing to constrain the candidate’s self-presentation. The guardrails were on the wrong side of the table.

“Twice the Wrong Signal”

Allen I. HuffcuttAn Empirical Review of the Employment Interview Construct Literature, International Journal of Selection and Assessment, 2011

If the preceding studies demonstrated that behavioral interviews are susceptible to distortion, Huffcutt’s 2011 review answered a deeper question: what are interviews actually measuring? Not what they intend to measure — what they demonstrably, empirically measure, as determined by the statistical relationship between interview ratings and independent assessments of candidate attributes.

Huffcutt’s theoretical model identified three sources of construct-related variance in interview ratings: job-related content (knowledge, competence, technical skill), interviewee performance (impression management, social skill, verbal fluency, composure), and personal/demographic characteristics (attractiveness, demographic similarity to the interviewer). These are distinct dimensions. A valid hiring instrument would weight the first heavily and the others not at all.

The data showed the opposite. The mean correlation between interview ratings and interviewee performance constructs was approximately twice as large as the correlation between interview ratings and job-related content. The interview was not failing to measure the right thing — it was measuring something else with remarkable consistency. It was a high-fidelity instrument pointed at the wrong target.

This finding reframes the entire debate. The problem with behavioral interviews is not that they are noisy or unreliable. They are, in fact, reasonably reliable. They produce consistent ratings across interviewers and sessions. The problem is that the construct they reliably measure — the candidate’s ability to perform well in the interview — is not the construct that predicts job performance. Reliability without validity is not a defense; it is an indictment.

“The Construct Taxonomy”

Allen I. Huffcutt, James M. Conway, Philip L. Roth & Nancy J. Stone — Identification and Meta-Analytic Assessment of Psychological Constructs Measured in Employment Interviews, Journal of Applied Psychology, 2001

A decade before his solo review, Huffcutt and colleagues conducted the foundational meta-analysis that mapped the construct landscape of employment interviews. Analyzing 338 construct ratings across 47 studies, they built a taxonomy of seven construct categories and asked a straightforward empirical question: when interviewers rate candidates, which psychological constructs are they actually assessing?

The answer was unambiguous. Basic personality traits and applied social skills were the most frequently assessed constructs across the interview literature. Mental capability appeared as a secondary dimension. Job knowledge and job-specific skills — the constructs most directly relevant to whether a candidate can perform the work — were assessed far less frequently.

The study also revealed a structural divergence between high- and low-structure interviews. Low-structure interviews (the conversational style) gravitated toward personality assessment. High-structure interviews (the behavioral style) shifted the emphasis toward applied social skills and mental capability — an improvement, but one that still placed the center of gravity on how the candidate presents rather than what they know or can do.

For technical roles, the implications are damning. When the empirical map of what an interview measures shows personality and social performance at the top, and job knowledge at the bottom, the instrument is grading an audition for the wrong role. A software engineer’s ability to narrate a conflict resolution story in the STAR format has no demonstrated relationship to their ability to design a distributed system, debug a concurrency issue, or reason about failure modes in production. The interview is testing the candidate’s skill at a task they will never perform on the job.


The Consensus

These four studies — spanning two decades, multiple research teams, and thousands of participants — converge on a conclusion that the behavioral interview’s proponents have been reluctant to confront. The format was designed to solve the problems of unstructured interviewing: interviewer subjectivity, narrative construction from noise, the illusion of validity. It addressed none of them. It relocated them.

The Coaching Effect. When a test’s methodology is publicly documented, universally taught, and endlessly rehearsed, the test ceases to measure competence and begins to measure preparation. The STAR method is the most widely disseminated interview framework in existence. It is taught in every career coaching book, every MBA program, every university career center, and on thousands of YouTube channels with millions of cumulative views. Over 90% of candidates engage in faking behavior within this framework. Coaching correlates with interview scores; interview scores do not correlate with job performance. The behavioral interview now measures whether a candidate has been coached — not whether they can do the work.

Construct-Irrelevant Variance. The traits that behavioral interviews reward — verbal fluency, extraversion, narrative construction, social perception, composure under performative pressure — are not the traits required to write correct software, design reliable systems, debug production failures, or reason about complex domains. For technical roles, the interview is measuring a dimension that is orthogonal to the work. Interview ratings correlate twice as strongly with impression management as with job-related competence. The instrument is reliable. It is reliably measuring the wrong thing.

The behavioral interview was the industry’s attempt to fix what the research on unstructured interviews had broken. It preserved the form — a human conversation, an evaluator’s subjective assessment, a narrative exchanged across a table — and believed that constraining the questions would constrain the distortion. The research is clear that it did not. The distortion simply moved from one side of the table to the other, and the result is a selection process that, for technical roles, evaluates performance art where it should be evaluating engineering.