What forensic science can and cannot prove

A witness in a white coat says the print at the scene belongs to the defendant. Not that it's similar. Not that it's consistent with. Belongs to, to the exclusion of every other person on earth.
For most of the twentieth century that sentence was allowed in court, and jurors had no reason to doubt it. It sounded like chemistry. It was presented like chemistry.
It wasn't chemistry. And the people who eventually said so out loud were not defence lawyers or campaigners but the American scientific establishment, in two documents that changed what expert witnesses are allowed to claim.
The two reports that broke the spell
In 2009 the National Research Council, the working arm of the United States National Academy of Sciences, published Strengthening Forensic Science in the United States: A Path Forward. Congress had asked for it. What came back was considerably more brutal than anyone expected.
Its central finding, stated plainly and repeated ever since, was that with the exception of nuclear DNA analysis, no forensic method had been rigorously shown to be capable of consistently and with a high degree of certainty linking a piece of evidence to a specific individual or source. Not fingerprints. Not toolmarks. Not bite marks, hair, handwriting, shoeprints or tyre impressions.
The committee's objection wasn't that these methods were useless. It was that nobody had done the work to find out how often they were wrong. There were no error rates, because the studies that would produce error rates had never been run, and in several disciplines the practitioners had actively resisted the idea that error rates applied to them at all.
In 2016 the President's Council of Advisors on Science and Technology went back over the same ground and published Forensic Science in Criminal Courts, which asked a sharper question: for each comparison discipline, do the empirical studies exist that would establish it as valid, and what do they show?
The answers were uneven. Some methods came through with qualifications attached. Others didn't come through at all.
Fingerprints are better than most and worse than you think
Start with the good news, because there is some. PCAST concluded that latent fingerprint analysis is foundationally valid — examiners comparing a mark from a scene against a known print really can do the task, at a rate far better than chance, and there are now properly designed studies proving it.
The trouble is what the studies found when they measured how often examiners get it wrong.
A large black-box study run with the FBI and published in 2011 produced a false positive rate of roughly one in three hundred. Another study, run in a Florida laboratory, produced a considerably worse figure — closer to one in eighteen. PCAST reported both, because both were legitimate studies and the honest answer is that the range is wide and depends heavily on the quality of the mark.
One in three hundred is not zero. It's also not the number a jury has in mind when someone says the prints match.
The confusion sits in the difference between two very different tasks that share a name. Taking a full set of ten prints from a cooperative person, under controlled conditions, and matching them against a database record is close to trivially reliable. Comparing a latent mark — a partial, smeared, distorted fragment lifted from a doorframe, possibly overlapping another print, possibly deposited days earlier — is a hard perceptual judgment made by a human being.
The method used to make that judgment is called ACE-V: Analysis, Comparison, Evaluation, Verification. It's a sensible workflow. What it isn't is a measurement. There's no threshold number of matching features that constitutes identification in most systems, no statistical model producing a probability, and the verification step has historically been performed by a colleague who already knew what the first examiner concluded.
The uniqueness claim needs care too. It's likely enough that no two people have identical friction ridge detail across a whole finger, but that was never the question at issue. The question is whether a smudged fragment of a fingertip contains enough information to distinguish one person from everybody else, and that's a completely different problem which the uniqueness assertion was quietly standing in for.
Bite marks did not survive at all
Bitemark comparison rests on two assumptions, and PCAST found essentially no support for either.
The first is that human dentition is unique in a way that expresses itself in a bite. The second is that skin can record that expression faithfully. Skin is a hopeless impression medium — it's elastic, it moves, it swells, it bruises unevenly, and the mark changes over the hours and days after it's made.
Studies asking examiners to work on known samples found poor agreement between them, and in some tests examiners couldn't reliably agree on whether a given injury was a human bite at all, let alone whose. PCAST concluded the discipline was not scientifically valid and, unusually, went further: it said the prospects of establishing validity through future research looked poor.
The Texas Forensic Science Commission recommended a moratorium on bitemark testimony in 2016. A number of convictions in which bitemark evidence featured have since been vacated after DNA testing identified someone else.
This is the part worth sitting with. It isn't that a technique was oversold at the margins. An entire category of expert testimony was admitted in serious criminal trials for decades on the strength of professional confidence, and when somebody finally checked, there was nothing underneath it.
DNA is genuinely strong, and it has a hard edge
Nuclear DNA analysis is the exception in every one of these reviews, and deservedly.
Standard profiling looks at short tandem repeats — stretches of DNA where a short sequence repeats a variable number of times, at locations chosen because they vary a lot between people and don't code for anything. The American CODIS system used thirteen core locations for years and expanded to twenty in 2017. For a good quality sample from a single person, the probability that an unrelated individual would produce the same profile by chance can be smaller than one in a trillion. The underlying biology is understood, the chemistry is quantitative, and the statistics were built into the method from the start rather than bolted on afterwards.
Then you leave the laboratory conditions and it gets difficult very quickly.
Real samples are often mixtures. Two people, three, sometimes more, in unknown proportions, and you don't necessarily know how many contributors there are — that's an inference, and it's frequently wrong. Real samples are often tiny, and at low quantities the chemistry becomes unstable: some of a person's alleles fail to amplify and vanish from the result, artefacts appear that look like alleles but aren't, and the profile you're interpreting is a partial and noisy shadow of the truth.
Probabilistic genotyping software was developed to handle exactly this, and it does help. It produces a likelihood ratio rather than a yes or no — how much more probable the observed data is if a particular person contributed than if they didn't. That's a much more honest output. But different programs, given the same complex mixture, can return meaningfully different answers, and the assumptions built into each are not always transparent to the court.
Interlaboratory studies run in the United States asked many laboratories to interpret the same mixture. The spread of conclusions was uncomfortable. In one widely discussed exercise, a substantial number of participating laboratories reported that a person was included as a possible contributor when the reference answer said otherwise.
DNA is the strongest thing in the toolbox. That doesn't make every DNA result strong.
DNA can tell you who, not how or when
This is the most consistently misunderstood point in the whole field, and it survives even among people who understand everything else.
A profile establishes that biological material from a person is present. It says nothing whatsoever about how that material arrived, when it arrived, or what the person was doing.
DNA transfers. You shed it constantly. It moves from your hand to a door handle to another person's hand to an object you have never touched — secondary and even tertiary transfer have been demonstrated experimentally. It can be carried on clothing, on tools, by insects, by ambulance crews and by investigators. Some people shed far more than others, consistently, for reasons that aren't fully understood. Material can persist on a surface for a long time, so the presence of a profile carries no reliable timestamp.
Every improvement in sensitivity makes this worse rather than better. Techniques that can build a profile from a few dozen cells will pick up traces that arrived through entirely innocent routes, and the more sensitive the method, the higher the proportion of what it detects that has nothing to do with the crime.
Locard's exchange principle, the old maxim that every contact leaves a trace, is now the problem rather than the promise. Every contact leaves a trace, and so does every contact with every contact.
The examiner is part of the instrument
Feature comparison is a human judgment, which means it's subject to everything that affects human judgment.
Itiel Dror ran a set of studies from 2006 onwards that ought to be better known. He took fingerprint comparisons that experienced examiners had previously decided, presented the same pairs back to the same examiners some time later with different contextual information attached, and found that a number of them reached different conclusions on identical evidence. Telling an examiner that the suspect had confessed, or that the case was a high-profile one, shifted the analysis.
None of this involves dishonesty. It's ordinary cognition operating in an ambiguous task, and it happens outside awareness, which is exactly why confidence is no protection against it.
The structural fixes are known and only partly implemented. Blind verification, so the second examiner doesn't know what the first concluded. Context management, so the analyst receives the evidence and nothing else — not the confession, not the suspect's record, not the detective's theory. Sequential unmasking, so the mark from the scene is documented in full before the known print is ever shown. Laboratories independent of the investigating police force, which the 2009 report recommended and which many jurisdictions still haven't done.
There's also a language problem the United States Department of Justice has been working through since 2018, restricting the phrases examiners may use in testimony. Claims of individualisation to the exclusion of all others, assertions of zero error rate, and expressions of one hundred per cent certainty are the ones that have had to go.
Whole disciplines have quietly collapsed
Fire investigation used to run on a list of indicators supposedly proving that an accelerant had been used — crazed glass, distinctive burn patterns on floors, a particular alligatoring of charred wood. Controlled burn experiments established that a room reaching flashover, where everything ignites more or less at once, produces all of those features with no accelerant present at all. The National Fire Protection Association's guide, first published in 1992 and revised repeatedly since, reset the discipline on physics rather than lore. A number of arson convictions rested on the older indicators.
Comparative bullet lead analysis measured trace elements in a bullet to link it to a batch. The FBI used it for decades, a National Research Council review in 2004 found the inferences drawn from it were not supportable, and the Bureau stopped using it in 2005.
Microscopic hair comparison produced the starkest number of the lot. A joint review by the FBI and the Department of Justice, working with the Innocence Project and the National Association of Criminal Defense Lawyers, examined trial transcripts in which examiners had given hair testimony. Of the first several hundred cases reviewed and reported in 2015, the overwhelming majority contained testimony that overstated what the evidence could support. Hair comparison isn't worthless — it can exclude, and it can indicate what to test with DNA — but the claims made about it in court had drifted far beyond it.
Firearms and toolmark comparison is still in argument. PCAST found only one appropriately designed study at the time and judged the foundation thin. Testimony asserting a match to a particular weapon to the exclusion of all others is much less common than it was.
What forensic science does well
A list of failures gives a distorted picture, so here's the other side.
Analytical chemistry is genuinely excellent. Identifying a drug, quantifying alcohol in blood, detecting a poison — these use mass spectrometry and chromatography, with calibration standards, blanks, controls and quantified uncertainty. It's real measurement.
Exclusion is powerful and undersold. A method that can't reliably tell you who did something can often tell you with confidence who didn't, and that asymmetry has freed a great many people.
Digital forensics deals in records with timestamps and cryptographic properties, which is a fundamentally different evidentiary situation from a smudge on a windowsill, though the interpretation of who was at the keyboard remains as human as ever.
Fire dynamics, ballistics trajectory reconstruction, blood pattern physics at the level of fluid mechanics, and forensic anthropology's estimates of age and stature from bone all rest on physical principles that can be tested.
And some things everyone assumes are precise really aren't. Estimating time of death is a good example — body cooling, rigor and lividity are affected by temperature, clothing, body mass, illness and a dozen other variables, and honest estimates come as wide windows rather than the confident hour named on television. Forensic entomology, which uses insect development to bracket a minimum time since death, is often better than the pathology, and it still produces ranges.
The CSI effect, which may not be a thing
The standard worry is that crime dramas have taught jurors to expect conclusive forensic evidence in every case, and that they now acquit when prosecutors can't produce it.
It's a plausible story. The evidence for it is weak.
Researchers have surveyed jurors about their expectations and then compared those expectations to actual verdicts. The expectations are real enough — people who watch a lot of forensic drama do expect to be shown scientific evidence. The link to how they actually vote has been much harder to find, and studies looking for it have largely come up empty or produced effects too small to matter.
There's a competing worry that points the other way and has better support. Jurors tend to over-value forensic testimony rather than under-value it. Evidence presented by a confident expert in a technical vocabulary carries enormous weight, and the difficulty is getting jurors to discount it appropriately when it deserves discounting.
Which is the more serious problem, and it's the reason all of the above matters. The questions worth asking of any forensic claim are simple enough. How often is this method wrong, and how do you know? Was the analyst told anything about the case beyond the evidence itself? Was the verification blind? And if the answer is that the material is present, does anyone actually know how it got there?


