Orthopedics · Part Four · Seeing and Ruling Out
Lesson 28 / 44
The Clinicians Advantage: What an Examination Reads That a Scan Cannot
A scan finding is a clue to interpret, never a verdict to fear.
A physical examination applies a load to the body and records how the system answers, which is what separates it from a scan. In 80 medical outpatients, the history alone led to the final diagnosis in 61 cases, or 76 percent. Ten imaging centers reading the same patient within three weeks agreed at a Fleiss kappa of 0.20. The Unified Model of Tone places accuracy in the assessment rather than in the picture.
History alone against the final diagnosis, 80 outpatients
61 of 80, or 76 percent
Agreement across 10 centers reading one lumbar spine
Fleiss kappa 0.20, miss rate 43.6 percent
Same films read with a suggestive clinical history
true positives rose from 38 to 84 percent
Centralization on repeated loading against discography
specificity 94 percent, likelihood ratio 6.9
What an orthopedic examination consists of
The interview comes first, covering mechanism of injury, what worsens and eases the complaint, and how it behaves across a real day. Then the hands: observation, active and passive motion, provocation tests that load a suspected structure, testing of power, sensation and reflexes, and functional tasks such as standing from a chair or walking a corridor.
Why loading reads what a picture cannot
An image records geometry in one position, with the person lying still. Muscle guarding, reflex threshold and pain modulation are properties that only appear when the body is asked to do something. Loading a tissue and watching what happens next samples the regulator rather than the shape, and the return to baseline is part of the reading.
01Reading a scan in context
The most useful reading of a spine scan comes from the clinician who examined the person
A lumbar MRI answers one question above all others: does this anatomy actually explain this presentation. A radiologist performs a genuinely valuable service, yet works almost entirely from the study itself, describing every structure faithfully without ever meeting the individual, examining the joints, or watching how the person moves. The treating clinician holds a different and larger set of information: the mechanism of injury, the symptom pattern, the neurological findings, the functional limitations, and the way the complaint behaves across a real day.
The effect of that context on a reading has been measured. Doubilet and Herman slipped test films carrying eight subtle but unambiguous abnormalities into the daily workload of readers who did not know a study was running. Each abnormality was read four times with a suggestive clinical history and four times without (Doubilet 1981).
True-positive readings rose from 16 to 72 percent among residents. Once staff radiologists had reviewed the interpretations, the rate rose from 38 to 84 percent. False positives rose somewhat as well, which is the honest shape of the effect. The same picture yields a different report depending on what the reader has been told.
Where diagnostic responsibility sits
Diagnostic responsibility ultimately belongs to the clinician who treats the person, not to the description on a page. That is why the report is read as one input rather than the last word, and why the person, not the picture alone, sits at the center of the interpretation. Clinical suspicion can direct attention toward a structure the report did not emphasize, and a careful reader will examine the exact region that the history and examination point to.
The advantage is measurable under controlled conditions too. Two radiologists read mammograms from 480 women blind, then read them again knowing the type and site of symptoms. Sensitivity rose from 71.3 to 75.8 percent for one reader, and overall accuracy improved by about 2 percent for both (Houssami 2004). Imaging is meant to support clinical reasoning, never to replace it.
02Findings
What the research shows
From two diagnostic-yield studies, two reader-context experiments, a ten-center imaging comparison, a palpation reliability review and two centralization studies.
03What the history contributes
The history carries most of the diagnosis and the examination decides what it means
Most diagnoses are made from the interview. Peterson and colleagues studied 80 medical outpatients with new or undiagnosed conditions. Internists listed their differential and rated their confidence three times: after the history, after the physical examination, and after the laboratory investigation. The history led to the final diagnosis in 61 patients, or 76 percent (Peterson 1992).
The physical examination led to the diagnosis in 10 patients and the laboratory investigation in 9. An earlier study of 80 new outpatients found the same order. A diagnosis matching the one finally accepted followed the referral letter and history in 66 cases, while the examination was useful in seven and the laboratory in a further seven (Hampton 1975).
Both were general medical clinics rather than spine clinics. In musculoskeletal care the interview carries mechanism of injury, the direction that provokes, the position that eases, night pain, and what the complaint does across a working day. Those items decide which examination is worth performing.
What the examination adds
The later steps buy exclusion and confidence rather than new diagnoses. In the same 80 patients, confidence in the correct diagnosis rose from 7.1 on a ten-point scale after the history to 8.2 after the examination and 9.3 after the laboratory work (Peterson 1992). The authors noted that examination and laboratory findings were instrumental in excluding diagnostic possibilities.
Exclusion is the point in a spine clinic. Ruling out fracture, infection, inflammatory disease and progressive neurological loss is what makes a conservative plan safe to start, and The Red Flags Clinicians Screen For sets out the items that carry real diagnostic weight. Part of the beauty of this work is ruling out the worse things.
04Overreading and underreading
Two opposite errors define the skill of interpretation
Skilled interpretation lives between two opposite mistakes, and recognizing both is what keeps a reading honest. The first error is overinterpreting, treating disc degeneration, mild bulges, small protrusions, or ordinary signal changes as proof of the pain when many of these appearances are age related and common in comfortable people. Findings in People Without Pain carries the age-stratified table.
The second error is underinterpreting, dismissing a subtle appearance that carries real clinical weight, so that meaningful nerve compression or progressive change is overlooked. Balance is the entire skill.
What ten reports of one spine showed
Both errors have been counted in the same experiment. One patient with low back pain and right L5 radicular symptoms was scanned at 10 different imaging centers over three weeks. Across the ten reports, 49 distinct findings were named at specific motion segments. Not one finding was reported by all 10 centers, and 32.7 percent appeared in a single report only (Herzog 2017).
Agreement across all reported findings reached a Fleiss kappa of 0.20. The average report carried 12.5 interpretive errors, made up of 10.9 false negatives out of 25 reference findings and 1.6 false positives. That works out to a sensitivity of 56.4 percent and a miss rate of 43.6 percent.
The reference standard in that study was two subspecialist spine radiologists reading by consensus, which places the finding where it belongs. Expert reading is a demanding skill, and a single report is one reading rather than a measurement of the spine. An abnormality on a film does not automatically become the pain generator, and a label such as mild does not automatically mean harmless.
The finding earns its meaning only when it fits the history, the examination, and the way symptoms actually present.
05The regions a scan covers
A lumbar scan answers one practical question: is neural tissue compressed along its path
Two regions carry that answer. The central canal houses the cauda equina and the descending nerve roots, and narrowing there can produce bilateral symptoms, neurogenic claudication, weakness, and in severe situations progressive deficit. The neural foramina are the side doors where exiting roots leave the spine, and encroachment there tends to produce radiculopathy, dermatomal pain, sensory change, and focal weakness.
The organization is predictable. At the L4 to L5 level a foraminal herniation typically involves the L4 root, while a paracentral herniation typically involves the descending L5 root. That consistency lets the clinician match an appearance to the exact complaint in front of them, and The Nerve Root carries the anatomy behind it.
Stenosis rarely arrives from one cause alone. It usually results from a combination of disc bulging, facet hypertrophy, and thickening of the ligamentum flavum, so the reader weighs several contributors together rather than fixing on a single label.
Rebuilding the person from the slices
A scan is a stack of flat slices, and the real work is reassembling them in the mind into one three dimensional person. Sagittal, axial, and coronal views each show a partial truth. The clinician tracks how the vertebral bodies, discs, facet joints, ligamentum flavum, central canal, foramina, and the exiting and descending roots line up from plane to plane.
Learning how these structures align across slices turns a confusing set of images into fast, reliable pattern recognition. A methodical, repeatable search pattern is how misses are avoided, and in the ten-center study the average report missed 10.9 of 25 reference findings (Herzog 2017).
This is where knowing the individual changes everything. When an appearance sits in a protective, threat driven state of the nervous system, guarding and symptom amplification can dominate the picture. Signal change interpreted in isolation misleads; interpreted against symptom location, neurological findings, and function, it becomes genuinely useful.
The same scan that answers the compression question also quietly rules out the conditions that genuinely warrant escalation, including tumor, infection, fracture, inflammatory disease, and severe stenosis. Ruling those out is itself reassuring information. Excluding the Dangerous on Imaging sets out the sensitivities, and What MRI Is Really For covers when a scan is warranted.
A clue, not a verdict
In practice the finding is treated as a clue to be interpreted, never a sentence to be handed down, and that stance is both more accurate and more reassuring. The systematic read moves through sequence, alignment, vertebral bodies, discs, central canal, foramina, and posterior elements. It ends on the step that matters most: correlating every appearance with the actual complaint.
Does the anatomy explain what this person feels, or is the abnormality incidental and better left alone. Pain in one region does not automatically identify the source, and an ache felt near the sacroiliac area can originate from lumbar structures. The image is captured elsewhere and read here with the full context that only the treating relationship provides.
That is how an incidental finding stays incidental instead of becoming a source of fear. Understood this way, a scan does not shrink a person into a defect. It gives the clinician one more instrument for building a confident, conservative plan aimed at restoring movement and calm, not at chasing shadows.
06The examination under load
A hands-on examination applies a load and records the answer, which a still picture cannot do
The examination samples something the image does not contain. A picture holds shape and position. An examination applies a mechanical demand and records what the body does with it, then how quickly it settles. Guarding, reflex threshold and symptom migration only exist while the system is being asked a question.
Which of those questions can be asked reliably has been studied directly. A systematic review screened 797 primary research articles and included 49. Among the studies using kappa statistics, acceptable reliability was reached by 64 percent of pain provocation studies and 58 percent of motion studies. Landmark location studies reached it in 33 percent and soft tissue studies in none (Seffinger 2004).
What the hands can and cannot reproduce
Regional range of motion proved more reliable than segmental range of motion. A single examiner repeating a test also agreed more closely than two examiners comparing their results. Discipline, experience level, and training immediately before the study did not improve reliability.
A meta-analysis of 48 studies published between 1965 and 2005 reached the same split. There was strong evidence that agreement between examiners on osseous pain and soft tissue pain is clinically acceptable, at a kappa of 0.4 or above (Stochkendahl 2006). Other manual procedures were either not reproducible or the evidence was conflicting.
The dividing line is consistent across both reviews. Tests that ask the tissue to answer a load are reproducible. Tests that ask for a static position or a texture are not. That is a property of what is being measured rather than a verdict on hands.
07Challenge and return
How symptoms move under repeated loading predicts recovery better than any static finding
The most informative examination variable is a change, not a starting value. Werneke and Hart followed 223 consecutive adults with acute low back pain for one year. They assessed 23 psychosocial, clinical and demographic variables, and recorded how pain location shifted with repeated trunk movements at every visit (Werneke 2001).
Failure of pain to centralize, together with leg pain at intake, were the strongest predictors of chronic pain and disability at 12 months. The authors state the contrast plainly: measures from physical examination rarely predict chronic behavior, while dynamic assessment of change in pain location during treatment did.
Centralization is the progressive retreat of referred pain toward the spinal midline under repeated end-range loading. Across 29 studies and 4,745 patients with back and neck pain, it appeared in 44.4 percent overall, in 74 percent of acute cases and 42 percent of subacute or chronic ones. Prognostic validity was supported in 21 of 23 studies (May 2012).
Two moderate-quality studies in that review did not support it, and reliability across five studies ranged from a kappa of 0.15 to 0.9. The response is powerful when it appears and weak as a rule-out. Tested against provocation discography in 107 patients with chronic low back pain, centralization reached a sensitivity of 40 percent, a specificity of 94 percent and a positive likelihood ratio of 6.9 (Laslett 2005).
Specificity fell to 80 percent with severe disability and 89 percent with psychological distress, and reached 100 percent in patients with minimal disability or distress. Thirty-eight patients could not tolerate a full examination and were excluded from the main analysis, which is part of the finding.
When the test is also the treatment
Long and colleagues gave 312 acute, subacute and chronic patients a standardized mechanical assessment. A directional preference, defined as an immediate lasting improvement in pain from repeated lumbar flexion, extension or sideglide testing, was elicited in 230 of them, or 74 percent (Long 2004).
Those 230 were randomized to exercises matching their direction, exercises opposite to it, or nondirectional exercises. The matched group improved significantly more on every outcome, including a threefold decrease in medication use. A third of both other groups withdrew within two weeks from no improvement or worsening, and no matched patient withdrew.
The early response also forecasts the late one. Among 100 patients treated with manual therapy, a change of 2 points or more on an 11-point pain scale by the second visit was associated with functional recovery at discharge. That threshold described the outcome correctly in 67 percent of cases (Cook 2012).
08Claims removed from this page
Two claims from the earlier version were removed
A gold pull-quote attributed to Dr. Jason Dulberg came off the page, because its wording could not be matched to any recorded source.
The claim that the treating clinician produces the most accurate reading of a scan was restated as what has actually been measured. A suggestive clinical history raised true-positive readings on identical films from 38 to 84 percent, and it raised false positives too (Doubilet 1981). The word most was doing work no study supports. What survives is the more useful reading, with the measured size of the effect attached.
Nothing else in the earlier text asserted a figure, so nothing else required removal.
09The model on assessment
What the Unified Model of Tone claims about clinical assessment
Everything above is established science, including the palpation reviews that came back unfavorable and the centralization studies that disagreed. What follows is this model’s reading, stated as ours rather than drawn from the papers cited.
Tone is the integrated organization through which the body’s interacting processes relate to one another at a given moment. It gathers mechanical tension, neural excitability, autonomic regulation, circulation, sensory gain, prediction and behavioral readiness. Our model holds that assessment, and not delivery, is the true seat of accuracy, because an input can only correspond to a pattern someone has already read.
From that follows a second claim. The act of reading tone changes tone. A hand loading a guarded segment alters what that segment is doing while the reading is being taken. That is why the same contact can be both an assessment and a treatment in the same motion.
The directional preference trial demonstrates it in numbers. The preference was defined by an immediate lasting improvement in pain produced by the testing itself, and it appeared in 230 of 312 patients (Long 2004). The assessment that identified the direction had already begun the treatment.
Why the challenge outperforms the number
The most informative variable is frequently not the starting number but the ability to change appropriately and return. Our model predicts the split the reliability reviews found. A test that asks the body to hold a position or present a texture requests a value the system does not keep still. None of the soft tissue palpation studies reached acceptable agreement (Seffinger 2004).
A test that applies a load and reads the answer reached acceptable agreement in 64 percent of studies, and the meta-analysis found strong evidence for pain provocation across 48 studies (Stochkendahl 2006). Reproducibility tracks whether the examiner asked the system to respond.
The same reading explains why clinical context bought only about 2 percent of accuracy in mammography (Houssami 2004). That question is whether a mass is present, and the image answers it directly. The spine question is whether a regulating system explains a symptom, and no still image holds that information.
The prediction
From this follows a claim the diagnostic literature does not make. Our model predicts that two people with matched imaging findings and matched pain scores separate on how they answer a load. Four measures recorded together in the same people will share one underlying factor rather than varying independently.
The four are pressure pain threshold at the tested segment, active range of motion in the loaded direction, resting heart rate variability, and time to return to baseline after a standardized repeated-loading test. Our model further predicts that this set forecasts function at 12 months more accurately than any finding on the image. An input that restores regulation moves high and low starters toward the same middle.
This is a claim about how accuracy is organized rather than a claim about what treatment does. If pressure pain threshold, active range of motion in the loaded direction, heart rate variability and time to return to baseline are shown to move together, the unification claim is confirmed.
10The tone reading
How clinical examination expresses tone
Every topic in this library expresses all of tone. In examination three aspects carry the signature, because the informative reading is what the body does under a demand rather than how it sits at rest.
Load
The examination is a graded load. Repeated end-range loading elicited a lasting direction of preference in 230 of 312 patients, and that direction guided what followed.
Time course
The return is the reading. A two-point change in pain by the second visit described the outcome at discharge correctly in 67 percent of 100 patients.
Input quality
Hands read what a picture cannot. Pain provocation tests reached acceptable agreement in 64 percent of studies, and soft tissue palpation reached it in none.
The remaining foundations run through the examination as well. Constraint: guarding narrows the range an examiner can move, and the narrowed range is itself a finding. Coupling: breathing, trunk motion and pelvic position change together under a load test, so one segment never answers alone. Gain: a segment that reports pain to light pressure is running a raised amplifier, which is why provocation thresholds carry information. Set point: the resting level of protection decides how much load the examination needs before anything answers. Prediction: what a person expects a movement to do shapes what the movement produces, so the order of tests changes the result. Oscillation: a complaint that follows a daily rhythm points the examination toward the regulator rather than the tissue. These are readings of one organization rather than separate systems, which is the core claim of the Unified Model of Tone.
11Across the library
How this page relates to the rest of the library
The examination is what turns every imaging page in this section into a decision.
What happens when a picture arrives before the question, and how the cascade toward surgery begins.
The age-stratified prevalence of bulges, protrusions and degeneration in people with no symptoms at all.
When a scan is warranted, and the narrow set of questions it answers better than any examination.
Which history and examination items carry diagnostic weight, and the false-positive burden they bring.
Specificity as correspondence between the input and the pattern, which is what an assessment exists to find.
The thresholds at which a reassessment stops being a reassessment and becomes a referral.
The reflex, motor and sensory testing that turns a symptom description into a localized finding.
12Frequently asked
Questions patients ask about the exam and the scan
Why does the examination matter more than the scan?
Because the examination measures what the person can do, and the scan records what the spine looks like at one instant. In 80 medical outpatients, the history alone led to the final diagnosis in 61 cases, and the physical examination in 10 more. In back pain specifically, how symptoms move under repeated loading predicted chronic pain and disability at one year, while static examination findings rarely did. The scan then answers a narrow question: is neural tissue compressed anywhere along its path.
Can a chiropractor read my images?
Yes. Interpreting imaging is a core clinical skill. The study itself is performed elsewhere, at an imaging center, and the clinician decides when a scan is warranted, refers for it, then reads the report and the images against the history and examination. That correlation is the work, and it matters because reports vary. When one patient was scanned at 10 different centers within three weeks, agreement across the 10 reports reached a Fleiss kappa of 0.20, and the average report missed 10.9 of 25 reference findings.
Does a scan finding explain my pain?
Only if it matches the clinical picture. A finding is a clue to interpret, never a verdict on its own, which is why correlation with symptoms is essential. Disc degeneration, mild bulges and ordinary signal change are common in people who feel completely well. The question a clinician asks is narrow: does this anatomy explain this presentation, or is the abnormality incidental and better left alone. Pain in one region does not automatically identify the source of that pain.
What is a challenge and return test?
It is any examination step that applies a load, records how the body answers, then records how quickly it returns. Repeated end-range movement testing is the best-studied version in back pain. Symptoms that retreat toward the spinal midline under repeated loading, a response called centralization, appeared in 44.4 percent of 4,745 patients across 29 studies and in 74 percent of acute cases. That response carried prognostic weight in 21 of 23 studies. The starting number matters less than the ability to change and return.
Why do two radiologists describe the same scan differently?
Because a report is a reading rather than a measurement. One patient was scanned at 10 imaging centers within three weeks. Across the 10 reports, 49 distinct findings were named at specific motion segments, no finding appeared in all 10 reports, and 32.7 percent appeared only once. Agreement reached a Fleiss kappa of 0.20. Expert reading is a demanding skill, which the same study showed by using two subspecialist spine radiologists reading by consensus as its reference standard for that spine.
If the history gives the diagnosis, what does the examination add?
Exclusion and confidence. In one prospective study of 80 outpatients, the history led to the final diagnosis in 76 percent of cases and the physical examination in 12 percent. The examination still raised the confidence of the internists in the correct diagnosis from 7.1 to 8.2 on a ten-point scale. It also ruled out possibilities the history had left open. In musculoskeletal care that exclusion is the point, because ruling out fracture, infection and progressive neurological loss is what makes a conservative plan safe to start.
What does the Unified Model of Tone say about clinical examination?
That assessment, and not delivery, is the true seat of accuracy, because an input can only correspond to a pattern someone has already read. The act of reading tone changes tone, so the same contact can be both an assessment and a treatment in the same motion. Repeated movement testing shows this directly. The assessment that identified a direction of preference in 230 of 312 patients did so by producing an immediate lasting improvement in pain. The model predicts that challenge measures share one underlying factor.
13The sources
References
12 primary sources, each linked to its record. Figures quoted on this page were checked against the published abstract.
Related evidence