“They Are Not Tests of Driving” — What the Field Sobriety Tests Actually Measure

Forensic toxicology laboratory with chromatogram, evidence vials, legal books, and courthouse

Okorie Okorocha, JD, MS, MSForensic Toxicology · Expert Witness Services

Sworn testimony · April 17, 1998

“Are these tests of driving? They are not.

Dr. Marcelline Burns — who built them — Depo. 39:23–25

What the field sobriety tests actually measure, according to the people who made them

A review of the science behind Horizontal Gaze Nystagmus, the Walk-and-Turn, and the One-Leg Stand — and of the sworn record that has been sitting in the file since 1998.

Three exercises decide a great many driving-under-the-influence cases: follow a light with your eyes, walk a line heel to toe, stand on one leg and count. They are administered at night, on the shoulder of a road, by a person with no medical training, to a stranger the officer has known for about five minutes. They carry the initials of a federal agency and the word standardized, and juries and hearing officers hear both of those as a guarantee.

They are not a guarantee. The battery has never been validated as a measure of driving impairment. The researcher who developed it said so under oath. Sober people fail it at rates approaching one in three. And every scoring threshold in it was calibrated to a legal limit that no longer exists.

What follows is the scientific record, drawn from the primary research and from the 1998 examination under oath of Dr. Marcelline Burns — founder and former director of the Southern California Research Institute, project director on the NHTSA contracts that produced the battery, and the government’s witness of choice in this field for decades.

Section 01The battery was built to increase arrests, not to measure driving

In 1975 NHTSA issued a request for proposals because the average blood alcohol concentration of impaired-driving arrestees nationwide was 0.17 percent while the prevailing statutory limit was 0.10 percent. The agency’s concern, in Burns’s account, was that a great many drivers who should have been arrested were not being detected. The contract awarded to her institute called for a battery officers could use at roadside to help them make arrest decisions.

That is a detection-yield objective. A screening instrument built to produce more positives produces more false positives in proportion — and, as Section 05 shows, that is exactly what happened.

There was also no existing science to build on:

A.  At that time, I had several years’ background in studying the effects of alcohol and other drugs. I didn’t have any background in roadside tests, nor do I think anybody in this country did at that time. It’s not a research topic that has gotten a lot of attention worldwide.

Examination under oath of Marcelline Burns, Ph.D.8:9–14

Burns testified that the word “standardized” had not entered law enforcement practice in 1975, and that before her work there was no research on the validity or reliability of roadside tests at all.

The three surviving exercises were not chosen because they are the best available measures of human neurological function. They were chosen because, in a 1977 regression analysis run on 238 laboratory subjects, they correlated most efficiently as a set with a breath instrument reading. The Romberg test — which Burns called “a very good test, an excellent test” — was cut not for weakness but for redundancy with the One-Leg Stand. No candidate exercise was ever evaluated against driving.

Section 02The tests are not tests of driving

The inference everyone draws from a poor performance is that the driver was too impaired to drive safely. The developer of the battery rejected that inference in plain language.

A.  What you’re asking is, are these tests of driving? They are not. If they were tests of driving, they would be field driving tests. I can elaborate on the reasons and everything behind that if you want, but they are not tests of driving. They are tests of sobriety.

Burns, examination under oath39:23–40:3

She was equally direct about the limits of a roadside encounter:

A.  The officer is not charged with making a decision about driving skills at roadside. There’s no way you can judge somebody in five minutes at roadside that you never saw before to make a decision about their driving skills.

Burns, examination under oath40:16–21

The 1998 San Diego validation report she co-authored says the same thing in print. It concedes that driving is a complex activity, that it is “unlikely that complex human performance, such as that required to safely drive an automobile, can be measured at roadside,” and that the tests assist arrest decisions “even though SFSTs do not directly measure driving impairment.” The report also admits that the exercises merely seem relevant on the face of it — which is an admission of facial plausibility, the opposite of validation.

Facial plausibility is the hazard. An observer watching someone lose balance on a line believes he is watching impaired driving. What he is actually watching is an unfamiliar athletic task, scored by eye, under conditions no one has ever validated against driving performance.

Section 03Every threshold was set for a limit that no longer exists

All of the scoring in the battery traces to one calibration exercise: officers in the 1977 and 1981 laboratory studies estimating whether a subject was above or below 0.10 percent. Burns described the task expressly — the officer had to record a decision whether the subject was “above or below .10, which was the statute in California at that time” — and described the function of the cut-points just as plainly: “What you’re trying to do is predict accurately whether this person is going to have a breath test that shows above or below .10.”

The clue counts everyone recites — four of six on HGN, two of eight on Walk-and-Turn, two of four on One-Leg Stand — were derived to sort subjects around that threshold. A decision cut-point calibrated to one criterion does not transfer to another without revalidation, and Burns testified in 1998 that she knew of no research revalidating the battery.

The limit moved by relabeling

The 1998 San Diego study did not re-derive anything for a 0.08 criterion. Its executive summary states that scoring “was changed slightly” so that four HGN clues would now indicate 0.08 percent rather than 0.10, and two clues would indicate 0.04 percent. The identical physical observation was assigned a new meaning by administrative decision.

And the criterion floats with the examiner

Asked about the false-positive rate in the first experiment, Burns explained that the officers “made a lot of false alarms” — calling subjects above 0.10 when they were not — and that on review of the data, “their criterion was really .08. In other words, they were saying arrest at the point they saw significant impairment.”

Read that as an admission rather than a defense. If the threshold each examiner applies is invisible and personal, the output is not an objective measurement. It is a judgment wearing a number.

Section 04What “NHTSA-validated” actually consists of

The reliability of the battery is usually asserted by that one phrase. Here is what stands behind it.

Two laboratory studies, run on weekends, in one building. The first used 238 subjects, roughly fifteen to twenty a day, because the study “completely took over our facility.” The second used 297. Subject blood alcohol concentrations ran from zero to 0.15 percent. That is the entire experimental foundation for an instrument now administered millions of times a year.

Examiners trained for one afternoon.

A.  We recruited ten police officers from law enforcement agencies in and around Los Angeles, and brought them in for one session which was about four hours long, and we trained them on how we wanted them to administer these six tests. … But it was a short training, and given that police officers had not had any experience with standardized testing methods, I feel fairly confident saying they hadn’t developed any particular confidence themselves in what they were doing.

Burns, examination under oath23:25–24:11

With one exception, none of them had previously heard of horizontal gaze nystagmus. As Burns put it, “it takes a period of learning to believe what you really see.”

A laboratory with no consequences, which she concedes distorted the data.

A.  Certainly, in the controlled environment where there was no consequence to an officer’s error, that had to affect the data. If you look at the data, you can see it did. … So you can’t recreate all the same variables in the laboratory that you have at roadside, which is one of the reasons I wanted to do a field study.

Burns, examination under oath48:10–49:3

A field study its own author disowned.

A.  Also, we did a small field study. Not a good field study, not big enough. There were a lot of things that we didn’t like about it, and reported that we didn’t like it because there weren’t funds to do it.

Burns, examination under oath44:13–17

And no sober control group — ever. The 1998 San Diego study enrolled only motorists officers had already stopped on suspicion of impaired driving. Seven officers completed 298 forms on patrol, and the celebrated 91 percent accuracy figure describes how often their BAC estimates matched the later breath test. Nobody sober was tested. Without a control group, an accuracy figure measures the agreement of two determinations on a pre-selected population of suspects and nothing more.

The Stuster and Burns data set contains its own refutation. Of the 83 subjects who actually measured below 0.08 percent, 24 — twenty-nine percent — were classified by the officers as at or above it.

Finally, in 1998 Burns testified that she knew of no ongoing research on field sobriety tests anywhere, and that when officers proudly told her they had modified or hardened the exercises, her answer was that the modified version may well be better, “but we don’t know that.”

Section 05Sober people fail these tests, and we now know how often

The question the validation studies never asked has since been asked three separate times, by independent researchers, with consistent results.

Failure rates among unimpaired subjects

Cole & Nowaczyk (1994)46%

Of officers’ judgments watching 21 subjects at 0.00 BAC perform field sobriety tests, 46% concluded the subject had had too much to drink. Watching the same sober people do ordinary tasks — reciting an address, walking normally — the figure was 15%.

Yoshizuka et al. (2014)26%

Of 185 drug-naïve subjects given the full battery at baseline in three double-blinded trials, 49 failed. Failures clustered in the balance tests: in the largest cohort, 1 subject failed HGN, 18 failed Walk-and-Turn, 5 failed One-Leg Stand.

Marcotte et al. (2023), JAMA Psychiatry49.2%

In a randomized, placebo-controlled trial, highly trained officers classified 49.2% of the placebo group as impaired. Specificity to actual exposure was roughly that of a coin flip.

NHTSA’s own data (Stuster & Burns, 1998)29%

Of 83 drivers who measured below 0.08 percent, 24 were classified by trained officers as at or above it — a false-positive rate nearly identical to the sober failure rate found in independent research.

Cole and Nowaczyk also reported the reliability coefficients that rarely come up in court. Test-retest reliability ran from .61 to .72 for individual exercises. Inter-rater reliability — whether two officers scoring the same performance agree — ran from .34 to .60. The accepted standard for a reliable standardized clinical instrument is about .85. The battery falls short of the ordinary psychometric threshold, and it falls shortest exactly where it matters most.

Their explanation of the mechanism is the whole argument in one sentence: because the exercises demand unfamiliar, unpracticed motor sequences, and because as few as two miscues produce a failing classification, the tests may do nothing but shift the officer’s decision criterion downward rather than actually separate the impaired from the unimpaired.

The confirmation problem

The Marcotte trial found something worse than a high placebo failure rate. Of the 128 participants officers classified as impaired, officers suspected 127 — 99.2 percent — of having received the active compound. The officers knew a large share of the sample would get placebo. Poor performance was still attributed to drugs in essentially every case. The authors name the mechanism: confirmation bias, which they note persists despite advanced training, including in law enforcement and the forensic sciences.

This is structural, not personal. The examiner who administers the battery has already formed the hypothesis the battery is supposed to test. He decided to detain the driver, selected the exercises, scored them by eye, recorded the clues, and decided what the performance meant. Nothing in the procedure checks that hypothesis.

Section 06The balance tests are designed to defeat balance

Human balance runs on three inputs working together: vision, proprioception, and the vestibular apparatus of the inner ear. Degrade one and sway increases. Degrade two and instability is the expected result. The Walk-and-Turn and One-Leg Stand degrade all three at once.

On the Walk-and-Turn, the subject is told to watch his feet — tilting the head down and disrupting vestibular function — while walking heel to toe on a line, which eliminates the normal base of support. On the One-Leg Stand, he looks down at the raised foot, again disrupting vestibular input, while standing on one leg, which removes half the available proprioceptive input. On the Romberg and finger-to-nose exercises used in drug evaluations, the head goes back and the eyes close, leaving one of three systems in operation.

Hansson, Beckman and Håkansson measured this directly. Thirty healthy adults with normal vision, normal vestibular function and normal cervical range of motion stood on a force plate while the researchers manipulated exactly these variables. Sway increased significantly with eyes closed and with an unstable surface on every measure, and head position significantly affected anteroposterior sway and sway area.

Those were screened, sober people standing still in a clinic. The roadside battery does the same thing to them on uneven pavement, in the dark, near traffic, under detention.

Add to this what clinical practice requires and roadside practice omits. A clinician evaluating gait and balance stratifies by age and screens for the long list of conditions that independently degrade performance — diabetes, peripheral neuropathy, vestibular disorders, cerebellar dysfunction, orthopedic injury, medication effects, obesity, visual impairment. The roadside examiner stratifies by nothing and screens for nothing. Burns’s own testimony anticipates the problem: candidate tests were cut from the original battery precisely because they were “not suitable for certain ages or certain conditions.” The survivors are administered to everyone anyway.

Why the One-Leg Stand runs thirty seconds

Burns explained the duration herself: “it turns out that people at .10 very often can hold it to 20 or 25 seconds. It’s only when the attention begins to waiver that the balance gets messed up. So it’s critical to hold it for 30 seconds.”

The test was extended until the target population began to fail. That is calibration to a desired yield, not to a physiological threshold — and it drags sober people with ordinary attentional variation across the same line.

Section 07HGN is a medical examination done without medicine

Horizontal gaze nystagmus is the most dangerous component, because it is the only one that sounds clinical. An officer describes involuntary eye movement, invokes the nervous system, and the listener hears diagnosis.

The differential diagnosis is enormous. Duane’s Clinical Ophthalmology records that precise eye-movement recording has established forty-nine distinct types of nystagmus and their causes. Normal people commonly show physiologic end-point nystagmus. Low doses of tranquilizers that do not affect driving ability can produce nystagmus. It can arise from neurologic disease. It can be congenital. Pathology cannot be sorted out at roadside; it requires a neuro-ophthalmologist or an oculographer. The text itself observes that it seems unreasonable for such judgments to belong to cursorily trained law officers, however well-meaning, and that careful history-taking and drug screening are often essential.

Clinicians use instruments. Officers use a penlight. Clinical evaluation employs infrared technology, magnetic search coils, video-oculography and nystagmometry, and produces a recording that can be reviewed for accuracy. The roadside examination produces no recording of any kind. There is nothing to review, nothing to reproduce, and no way for anyone to verify what the officer says he saw. The evidence exists only as his recollection of a fleeting observation made while forming the belief that the driver was intoxicated.

It is routinely done wrong. Booker reported in 2001 that more than ninety-five percent of officers improperly conducted the HGN examination they then used as a basis for arrest. Combine that with Burns’s testimony that a non-standard administration cannot be related to the research data at all, and the conclusion follows on its own.

The angle of onset is a guess. Burns described the third clue as determining “the angle of gaze when there’s the first onset of jerking — in other words, has the individual deviated his eyes 40 degrees, 45 degrees or 30 degrees? Because it’s the relationship between that and the BAC.” The officer is estimating a fifteen-degree difference in eye deviation, by unaided sight, in low light, at the instant jerking first becomes perceptible. No instrument. No measurement recorded. No second observer.

And Burns conceded the obvious limit. Describing a subject with “really strange eyes for some reason that I don’t know,” she said: “if that’s the only test you have, you really don’t have any basis for a decision.”

Section 08Three tests, one measurement

The battery is presented as three independent confirmations. It isn’t.

Walk-and-Turn and One-Leg Stand are both balance tasks measuring the same underlying construct — which is precisely why Romberg was dropped: “if you’ve already measured balance, you don’t gain much by measuring it again.” Failing both is one construct observed twice, not two data points.

Meanwhile HGN dominates. Burns testified that “Horizontal Gaze Nystagmus is almost as good alone a predictor as all three tests.” The San Diego study puts numbers on it: HGN alone correlates at r = 0.65, and all three combined reach r = 0.69. Adding two entire exercises improved the correlation by four one-hundredths — and a correlation of 0.69 leaves well over half the variance in outcome unexplained by test performance.

So the battery functions as a single test with two correlated balance tasks attached. And Burns’s own standard is that a single marker is not enough: “There’s always a risk if you rely on a single marker.” She drew the analogy of a physician who would want the blood work and the EKG all pointing the same direction before making a diagnosis.

Section 09Accuracy percentages don’t describe a person

Numbers get attached to the individual exercises in testimony — typically 77 percent for HGN, 68 percent for Walk-and-Turn, 65 percent for One-Leg Stand. They fail on three separate grounds.

First, they describe groups, not individuals. A percentage from group data tells you how often a determination was correct across a sample. It tells you nothing about the probability that this determination was correct. California authority is directly on point. In People v. Wilson (2019) 33 Cal.App.5th 559, the Court of Appeal held it was error to admit statistical-likelihood evidence, reasoning that testimony conveying a numerical probability effectively tells the trier of fact there is an overwhelming likelihood as to the disputed matter, and thereby invades the province of the trier of fact, whose job it is to draw the ultimate inferences from the evidence. That rests on People v. Collins (1968) 68 Cal.2d 319, where the Supreme Court held that mathematical-probability testimony infected the case with fatal error and warned that mathematics “must not cast a spell” over the fact-finder.

Second, they don’t measure impairment. They measure agreement between an officer’s BAC estimate and a later chemical test. Not impairment, and certainly not impaired driving.

Third, they omit the denominator that matters. An accuracy figure reported without its false-positive rate is scientifically incomplete — and the false-positive rate here runs from twenty-three to forty-nine percent.

Section 10What the record has to show before any of this means anything

Because the exercises have no intrinsic meaning — they aren’t driving tests and they aren’t clinical neurological examinations — their entire interpretive claim rests on comparison to the research data. Burns was unequivocal: “If the tests are going to have meaning as objective measures, they have to be administered in a standardized way.”

And when an officer improvises, the score cannot be salvaged. Asked whether a partially standardized administration could be discounted proportionally or instead simply had no backing, she answered “Neither of the above,” and added: “I would not try to adjust it by any percentage.” There is no correction factor. When an officer scores something that isn’t a scored error, she agreed predictive power is diminished — and when it was put to her that it could actually be getting worse because nobody has studied it, she answered: “Could be.”

Two further points rarely surface at hearing. The scoring is not hers — “NHTSA developed the scoring; I didn’t” — and she reviewed only the first manual before release, not the later editions officers are actually trained from.

So before any weight is assigned to a result, the record should show, and exclude, the following:

  • The examiner’s certification date, the manual edition he trained under, and his most recent refresher.
  • The verbatim instructions actually delivered — not an assertion that instructions were given per training.
  • Whether the subject held the starting stance throughout the instructions, and confirmed understanding.
  • Any audio or video of the encounter, in-car and body-worn, plus the retention policy for each.
  • The original scoring sheet, and whether clues and timing were recorded contemporaneously or reconstructed.
  • Roadway conditions: surface, grade and cross-slope, wet or dry, ambient and artificial lighting, weather, passing traffic, emergency lights, other personnel.
  • Footwear, and whether the option to remove it was offered.
  • Age, height, weight, and any orthopedic, vestibular, neurologic, ophthalmologic, endocrine or medication-related condition affecting balance, gait, coordination or eye movement.
  • Primary language, English proficiency, hearing, and when any interpreter arrived relative to each exercise.
  • Elapsed time between the end of driving and each exercise.
  • For HGN: the stimulus, its distance and elevation, pass speed and count, hold time at maximum deviation, the light source, and how any onset angle was estimated.

Where those are missing from the file, the alternative explanations have not been excluded — and the performance is equally consistent with an unimpaired person attempting an unfamiliar task under bad conditions.

Section 11The bottom line

The Standardized Field Sobriety Tests do not measure impaired driving, and were never designed to. They were built between 1975 and 1981 to help officers estimate whether a driver was over a statutory threshold that no longer exists, using three exercises picked by regression analysis on 238 laboratory subjects, administered by officers trained for about four hours, in a facility the researcher conceded could not reproduce roadside conditions.

Every load-bearing claim has been contradicted by the people who built the thing. Burns testified the tests are not tests of driving. She testified nobody can assess driving skills in five minutes at roadside. She testified a non-standard administration cannot be related to the research data and cannot be adjusted for. She testified the scoring was NHTSA’s, not hers, and that she didn’t review the later manuals. She testified the consequence-free laboratory distorted the data, that the field study was not good and not big enough, and that no further research was underway. The validation report bearing her name states in print that the tests do not directly measure driving impairment.

Independent researchers supplied the control group NHTSA never did. When sober, drug-free people take these tests, between twenty-three and forty-nine percent of them fail. Inside NHTSA’s own flagship data set, twenty-nine percent of drivers below the limit were called over it. And when trained officers see poor performance, they attribute it to intoxication almost every time, whether or not anything is there.

These exercises do not separate the impaired from the sober. They separate the practiced from the unpracticed, the young from the old, the healthy from the unwell, the fluent from the foreign, and the calm from the frightened. Presented with a federal agency’s imprimatur and the word standardized, they convert an officer’s existing suspicion into something that looks like proof.

ReferencesEverything above, sourced

  1. Examination Under Oath of Marcelline Burns, Ph.D., April 17, 1998 (Lori Raye, CSR No. 7052), pp. 1–62.
  2. Stuster, J., & Burns, M. Validation of the Standardized Field Sobriety Test Battery at BACs Below 0.10 Percent. Final Report to NHTSA (1998).
  3. Cole, S., & Nowaczyk, R. H. Field Sobriety Tests: Are They Designed for Failure? 79 Perceptual & Motor Skills 99 (1994).
  4. Yoshizuka, K., Perry, P. J., Upton, G., Lopes, I., & Ip, E. J. Standardized Field Sobriety Test: False Positive Test Rate Among Sober Subjects. 3 J. Forensic Toxicology & Pharmacology 2, art. 120 (2014).
  5. Marcotte, T. D., et al. Evaluation of Field Sobriety Tests for Identifying Drivers Under the Influence of Cannabis: A Randomized Clinical Trial. 80 JAMA Psychiatry 914 (2023).
  6. Rose, S., Bartell, D., & Hensel, D. SFSTs: Neurologic Failure Guaranteed. University Medical & Forensic Consultants, Inc. (2015).
  7. Hansson, E. E., Beckman, A., & Håkansson, A. Effect of Vision, Proprioception, and the Position of the Vestibular Organ on Postural Sway. 130 Acta Oto-Laryngologica 1358 (2010).
  8. Duane’s Clinical Ophthalmology (Tasman & Jaeger eds., 2007).
  9. Booker, J. L. End-Position Nystagmus as an Indicator of Ethanol Intoxication. 41 Science & Justice 113 (2001).
  10. Okorocha, O. Field Sobriety Tests: Criminal Injustice. 5 LSD Journal 298 (2013).
  11. Burns, M., & Moskowitz, H. Psychophysical Tests for DWI Arrest. DOT-HS-802-424, NHTSA (1977).
  12. Tharp, V., Burns, M., & Moskowitz, H. Development and Field Test of Psychophysical Tests for DWI Arrest. DOT-HS-805-864, NHTSA (1981).
  13. Salzman, B. Gait and Balance Disorders in Older Adults. 82 American Family Physician 61 (2010).
  14. People v. Wilson (2019) 33 Cal.App.5th 559.
  15. People v. Collins (1968) 68 Cal.2d 319.

About the author

Okorie Okorocha, JD, MS, MS is a forensic toxicologist and expert witness. He holds master’s degrees in forensic science and toxicology and a juris doctor, and has testified on alcohol and drug pharmacology, blood and breath testing, and field sobriety testing.

Note on citations

Page and line citations to the Burns examination under oath were taken from the transcript on file. Confirm each pin cite against the certified transcript before use in any filing or at hearing.

Disclaimer

This article is general scientific and educational commentary. It is not legal advice, it does not create an attorney-client or expert-client relationship, and it is not a substitute for case-specific analysis by qualified counsel and a retained expert.

Spread the love


The National Black Lawyers

top 40 lawyers

civil trial law

Lawyers of Distinction

Loading...

Recent Blog Articles

“They Are Not Tests of Driving” — What the Field Sobriety Tests Actually Measure

The woman who built the Standardized Field Sobriety Tests testified under oath that they do not measure driving. A forensic toxicology review of the science behind HGN, the Walk-and-Turn, and the One-Leg Stand.

Spread the love

Read More

California Dual Employment: Working for a Competitor

California generally voids employee noncompete agreements, but that does not necessarily prevent an at-will employer from terminating a worker who simultaneously works for a competitor. The result depends on duties of loyalty, trade-secret conduct, contract terms and other statutory protections.

Spread the love

Read More

The Okorocha Mid-Year Street Drug Report of 2026

No credible dataset measures the quantity of illicit drugs actually sold in the United States at the state or national level. Seizures reflect enforcement activity, forensic identifications reflect laboratory submissions, and the only defensible national market-size estimate covers 2006–2016. The…

Spread the love

Read More

Speak with an expert today!

Contact the offices of Okorie Okorocha for professional and reliable advice which you can trust.

Call (424) 363-3347 Contact Us