Orthopedics · Part Four · Seeing and Ruling Out

28PART IV

Lesson 28 / 44

The Clinicians Advantage: What an Examination Reads That a Scan Cannot

A scan finding is a clue to interpret, never a verdict to fear.

A physical examination applies a load to the body and records how the system answers, which is what separates it from a scan. In 80 medical outpatients, the history alone led to the final diagnosis in 61 cases, or 76 percent. Ten imaging centers reading the same patient within three weeks agreed at a Fleiss kappa of 0.20. The Unified Model of Tone places accuracy in the assessment rather than in the picture.

History alone against the final diagnosis, 80 outpatients

61 of 80, or 76 percent

Agreement across 10 centers reading one lumbar spine

Fleiss kappa 0.20, miss rate 43.6 percent

Same films read with a suggestive clinical history

true positives rose from 38 to 84 percent

Centralization on repeated loading against discography

specificity 94 percent, likelihood ratio 6.9

What an orthopedic examination consists of

The interview comes first, covering mechanism of injury, what worsens and eases the complaint, and how it behaves across a real day. Then the hands: observation, active and passive motion, provocation tests that load a suspected structure, testing of power, sensation and reflexes, and functional tasks such as standing from a chair or walking a corridor.

Why loading reads what a picture cannot

An image records geometry in one position, with the person lying still. Muscle guarding, reflex threshold and pain modulation are properties that only appear when the body is asked to do something. Loading a tissue and watching what happens next samples the regulator rather than the shape, and the return to baseline is part of the reading.

01Reading a scan in context

The most useful reading of a spine scan comes from the clinician who examined the person

A lumbar MRI answers one question above all others: does this anatomy actually explain this presentation. A radiologist performs a genuinely valuable service, yet works almost entirely from the study itself, describing every structure faithfully without ever meeting the individual, examining the joints, or watching how the person moves. The treating clinician holds a different and larger set of information: the mechanism of injury, the symptom pattern, the neurological findings, the functional limitations, and the way the complaint behaves across a real day.

The effect of that context on a reading has been measured. Doubilet and Herman slipped test films carrying eight subtle but unambiguous abnormalities into the daily workload of readers who did not know a study was running. Each abnormality was read four times with a suggestive clinical history and four times without (Doubilet 1981).

True-positive readings rose from 16 to 72 percent among residents. Once staff radiologists had reviewed the interpretations, the rate rose from 38 to 84 percent. False positives rose somewhat as well, which is the honest shape of the effect. The same picture yields a different report depending on what the reader has been told.

Where diagnostic responsibility sits

Diagnostic responsibility ultimately belongs to the clinician who treats the person, not to the description on a page. That is why the report is read as one input rather than the last word, and why the person, not the picture alone, sits at the center of the interpretation. Clinical suspicion can direct attention toward a structure the report did not emphasize, and a careful reader will examine the exact region that the history and examination point to.

The advantage is measurable under controlled conditions too. Two radiologists read mammograms from 480 women blind, then read them again knowing the type and site of symptoms. Sensitivity rose from 71.3 to 75.8 percent for one reader, and overall accuracy improved by about 2 percent for both (Houssami 2004). Imaging is meant to support clinical reasoning, never to replace it.

02Findings

What the research shows

From two diagnostic-yield studies, two reader-context experiments, a ten-center imaging comparison, a palpation reliability review and two centralization studies.

The history carried 61 of 80 diagnoses
Internists recorded their differential at three points in the workup. The history led to the final diagnosis in 61 of 80 patients, the examination in 10 and the laboratory in 9 (Peterson 1992). Talking is the highest-yield step.
The examination changed few diagnoses and excluded many possibilities
In 80 new medical outpatients, the referral letter and history produced the accepted diagnosis in 66 cases. The physical examination was useful in seven and the laboratory in a further seven (Hampton 1975). Yield and value are different quantities.
Confidence climbed even when the diagnosis did not
On a scale of 1 to 10, confidence in the correct diagnosis rose from 7.1 after the history to 8.2 after the examination (Peterson 1992). Exclusion is what the later steps buy.
A clinical history more than doubled what readers saw
Eight subtle abnormalities were slipped into the daily workload of unaware readers. With a suggestive history, true-positive readings rose from 38 to 84 percent for combined resident and staff readings (Doubilet 1981). Context decides what a reader sees.
Ten centers, one spine, poor agreement
One patient was scanned at 10 imaging centers within three weeks. Across the reports, 49 distinct findings were named at specific motion segments, none appeared in all 10, and agreement reached a Fleiss kappa of 0.20 (Herzog 2017). A report is a reading.
Provocation tests are reproducible, texture is not
Among spinal palpation studies using kappa statistics, acceptable reliability was reached by 64 percent of pain provocation studies and none of the soft tissue studies (Seffinger 2004). What can be reproduced is a response to a load.
Symptom migration under load predicted the year ahead
In 223 adults followed for one year after acute low back pain, failure of pain to centralize during treatment was among the strongest predictors of chronic pain and disability (Werneke 2001). Static examination findings rarely predicted it.
Matching the exercise to the tested direction changed everything
Of 312 patients, 230 showed a directional preference. Those given exercise matching it improved more on every outcome, and a third of the other two groups withdrew within two weeks (Long 2004). No matched patient withdrew.

03What the history contributes

The history carries most of the diagnosis and the examination decides what it means

Most diagnoses are made from the interview. Peterson and colleagues studied 80 medical outpatients with new or undiagnosed conditions. Internists listed their differential and rated their confidence three times: after the history, after the physical examination, and after the laboratory investigation. The history led to the final diagnosis in 61 patients, or 76 percent (Peterson 1992).

The physical examination led to the diagnosis in 10 patients and the laboratory investigation in 9. An earlier study of 80 new outpatients found the same order. A diagnosis matching the one finally accepted followed the referral letter and history in 66 cases, while the examination was useful in seven and the laboratory in a further seven (Hampton 1975).

Both were general medical clinics rather than spine clinics. In musculoskeletal care the interview carries mechanism of injury, the direction that provokes, the position that eases, night pain, and what the complaint does across a working day. Those items decide which examination is worth performing.

What the examination adds

The later steps buy exclusion and confidence rather than new diagnoses. In the same 80 patients, confidence in the correct diagnosis rose from 7.1 on a ten-point scale after the history to 8.2 after the examination and 9.3 after the laboratory work (Peterson 1992). The authors noted that examination and laboratory findings were instrumental in excluding diagnostic possibilities.

Exclusion is the point in a spine clinic. Ruling out fracture, infection, inflammatory disease and progressive neurological loss is what makes a conservative plan safe to start, and The Red Flags Clinicians Screen For sets out the items that carry real diagnostic weight. Part of the beauty of this work is ruling out the worse things.

04Overreading and underreading

Two opposite errors define the skill of interpretation

Skilled interpretation lives between two opposite mistakes, and recognizing both is what keeps a reading honest. The first error is overinterpreting, treating disc degeneration, mild bulges, small protrusions, or ordinary signal changes as proof of the pain when many of these appearances are age related and common in comfortable people. Findings in People Without Pain carries the age-stratified table.

The second error is underinterpreting, dismissing a subtle appearance that carries real clinical weight, so that meaningful nerve compression or progressive change is overlooked. Balance is the entire skill.

What ten reports of one spine showed

Both errors have been counted in the same experiment. One patient with low back pain and right L5 radicular symptoms was scanned at 10 different imaging centers over three weeks. Across the ten reports, 49 distinct findings were named at specific motion segments. Not one finding was reported by all 10 centers, and 32.7 percent appeared in a single report only (Herzog 2017).

Agreement across all reported findings reached a Fleiss kappa of 0.20. The average report carried 12.5 interpretive errors, made up of 10.9 false negatives out of 25 reference findings and 1.6 false positives. That works out to a sensitivity of 56.4 percent and a miss rate of 43.6 percent.

The reference standard in that study was two subspecialist spine radiologists reading by consensus, which places the finding where it belongs. Expert reading is a demanding skill, and a single report is one reading rather than a measurement of the spine. An abnormality on a film does not automatically become the pain generator, and a label such as mild does not automatically mean harmless.

The finding earns its meaning only when it fits the history, the examination, and the way symptoms actually present.

05The regions a scan covers

A lumbar scan answers one practical question: is neural tissue compressed along its path

Two regions carry that answer. The central canal houses the cauda equina and the descending nerve roots, and narrowing there can produce bilateral symptoms, neurogenic claudication, weakness, and in severe situations progressive deficit. The neural foramina are the side doors where exiting roots leave the spine, and encroachment there tends to produce radiculopathy, dermatomal pain, sensory change, and focal weakness.

The organization is predictable. At the L4 to L5 level a foraminal herniation typically involves the L4 root, while a paracentral herniation typically involves the descending L5 root. That consistency lets the clinician match an appearance to the exact complaint in front of them, and The Nerve Root carries the anatomy behind it.

Stenosis rarely arrives from one cause alone. It usually results from a combination of disc bulging, facet hypertrophy, and thickening of the ligamentum flavum, so the reader weighs several contributors together rather than fixing on a single label.

Rebuilding the person from the slices

A scan is a stack of flat slices, and the real work is reassembling them in the mind into one three dimensional person. Sagittal, axial, and coronal views each show a partial truth. The clinician tracks how the vertebral bodies, discs, facet joints, ligamentum flavum, central canal, foramina, and the exiting and descending roots line up from plane to plane.

Learning how these structures align across slices turns a confusing set of images into fast, reliable pattern recognition. A methodical, repeatable search pattern is how misses are avoided, and in the ten-center study the average report missed 10.9 of 25 reference findings (Herzog 2017).

This is where knowing the individual changes everything. When an appearance sits in a protective, threat driven state of the nervous system, guarding and symptom amplification can dominate the picture. Signal change interpreted in isolation misleads; interpreted against symptom location, neurological findings, and function, it becomes genuinely useful.

The same scan that answers the compression question also quietly rules out the conditions that genuinely warrant escalation, including tumor, infection, fracture, inflammatory disease, and severe stenosis. Ruling those out is itself reassuring information. Excluding the Dangerous on Imaging sets out the sensitivities, and What MRI Is Really For covers when a scan is warranted.

A clue, not a verdict

In practice the finding is treated as a clue to be interpreted, never a sentence to be handed down, and that stance is both more accurate and more reassuring. The systematic read moves through sequence, alignment, vertebral bodies, discs, central canal, foramina, and posterior elements. It ends on the step that matters most: correlating every appearance with the actual complaint.

Does the anatomy explain what this person feels, or is the abnormality incidental and better left alone. Pain in one region does not automatically identify the source, and an ache felt near the sacroiliac area can originate from lumbar structures. The image is captured elsewhere and read here with the full context that only the treating relationship provides.

That is how an incidental finding stays incidental instead of becoming a source of fear. Understood this way, a scan does not shrink a person into a defect. It gives the clinician one more instrument for building a confident, conservative plan aimed at restoring movement and calm, not at chasing shadows.

06The examination under load

A hands-on examination applies a load and records the answer, which a still picture cannot do

The examination samples something the image does not contain. A picture holds shape and position. An examination applies a mechanical demand and records what the body does with it, then how quickly it settles. Guarding, reflex threshold and symptom migration only exist while the system is being asked a question.

Which of those questions can be asked reliably has been studied directly. A systematic review screened 797 primary research articles and included 49. Among the studies using kappa statistics, acceptable reliability was reached by 64 percent of pain provocation studies and 58 percent of motion studies. Landmark location studies reached it in 33 percent and soft tissue studies in none (Seffinger 2004).

What the hands can and cannot reproduce

Regional range of motion proved more reliable than segmental range of motion. A single examiner repeating a test also agreed more closely than two examiners comparing their results. Discipline, experience level, and training immediately before the study did not improve reliability.

A meta-analysis of 48 studies published between 1965 and 2005 reached the same split. There was strong evidence that agreement between examiners on osseous pain and soft tissue pain is clinically acceptable, at a kappa of 0.4 or above (Stochkendahl 2006). Other manual procedures were either not reproducible or the evidence was conflicting.

The dividing line is consistent across both reviews. Tests that ask the tissue to answer a load are reproducible. Tests that ask for a static position or a texture are not. That is a property of what is being measured rather than a verdict on hands.

07Challenge and return

How symptoms move under repeated loading predicts recovery better than any static finding

The most informative examination variable is a change, not a starting value. Werneke and Hart followed 223 consecutive adults with acute low back pain for one year. They assessed 23 psychosocial, clinical and demographic variables, and recorded how pain location shifted with repeated trunk movements at every visit (Werneke 2001).

Failure of pain to centralize, together with leg pain at intake, were the strongest predictors of chronic pain and disability at 12 months. The authors state the contrast plainly: measures from physical examination rarely predict chronic behavior, while dynamic assessment of change in pain location during treatment did.

Centralization is the progressive retreat of referred pain toward the spinal midline under repeated end-range loading. Across 29 studies and 4,745 patients with back and neck pain, it appeared in 44.4 percent overall, in 74 percent of acute cases and 42 percent of subacute or chronic ones. Prognostic validity was supported in 21 of 23 studies (May 2012).

Two moderate-quality studies in that review did not support it, and reliability across five studies ranged from a kappa of 0.15 to 0.9. The response is powerful when it appears and weak as a rule-out. Tested against provocation discography in 107 patients with chronic low back pain, centralization reached a sensitivity of 40 percent, a specificity of 94 percent and a positive likelihood ratio of 6.9 (Laslett 2005).

Specificity fell to 80 percent with severe disability and 89 percent with psychological distress, and reached 100 percent in patients with minimal disability or distress. Thirty-eight patients could not tolerate a full examination and were excluded from the main analysis, which is part of the finding.

When the test is also the treatment

Long and colleagues gave 312 acute, subacute and chronic patients a standardized mechanical assessment. A directional preference, defined as an immediate lasting improvement in pain from repeated lumbar flexion, extension or sideglide testing, was elicited in 230 of them, or 74 percent (Long 2004).

Those 230 were randomized to exercises matching their direction, exercises opposite to it, or nondirectional exercises. The matched group improved significantly more on every outcome, including a threefold decrease in medication use. A third of both other groups withdrew within two weeks from no improvement or worsening, and no matched patient withdrew.

The early response also forecasts the late one. Among 100 patients treated with manual therapy, a change of 2 points or more on an 11-point pain scale by the second visit was associated with functional recovery at discharge. That threshold described the outcome correctly in 67 percent of cases (Cook 2012).

08Claims removed from this page

Two claims from the earlier version were removed

A gold pull-quote attributed to Dr. Jason Dulberg came off the page, because its wording could not be matched to any recorded source.

The claim that the treating clinician produces the most accurate reading of a scan was restated as what has actually been measured. A suggestive clinical history raised true-positive readings on identical films from 38 to 84 percent, and it raised false positives too (Doubilet 1981). The word most was doing work no study supports. What survives is the more useful reading, with the measured size of the effect attached.

Nothing else in the earlier text asserted a figure, so nothing else required removal.

09The model on assessment

What the Unified Model of Tone claims about clinical assessment

Everything above is established science, including the palpation reviews that came back unfavorable and the centralization studies that disagreed. What follows is this model’s reading, stated as ours rather than drawn from the papers cited.

Tone is the integrated organization through which the body’s interacting processes relate to one another at a given moment. It gathers mechanical tension, neural excitability, autonomic regulation, circulation, sensory gain, prediction and behavioral readiness. Our model holds that assessment, and not delivery, is the true seat of accuracy, because an input can only correspond to a pattern someone has already read.

From that follows a second claim. The act of reading tone changes tone. A hand loading a guarded segment alters what that segment is doing while the reading is being taken. That is why the same contact can be both an assessment and a treatment in the same motion.

The directional preference trial demonstrates it in numbers. The preference was defined by an immediate lasting improvement in pain produced by the testing itself, and it appeared in 230 of 312 patients (Long 2004). The assessment that identified the direction had already begun the treatment.

Why the challenge outperforms the number

The most informative variable is frequently not the starting number but the ability to change appropriately and return. Our model predicts the split the reliability reviews found. A test that asks the body to hold a position or present a texture requests a value the system does not keep still. None of the soft tissue palpation studies reached acceptable agreement (Seffinger 2004).

A test that applies a load and reads the answer reached acceptable agreement in 64 percent of studies, and the meta-analysis found strong evidence for pain provocation across 48 studies (Stochkendahl 2006). Reproducibility tracks whether the examiner asked the system to respond.

The same reading explains why clinical context bought only about 2 percent of accuracy in mammography (Houssami 2004). That question is whether a mass is present, and the image answers it directly. The spine question is whether a regulating system explains a symptom, and no still image holds that information.

The prediction

From this follows a claim the diagnostic literature does not make. Our model predicts that two people with matched imaging findings and matched pain scores separate on how they answer a load. Four measures recorded together in the same people will share one underlying factor rather than varying independently.

The four are pressure pain threshold at the tested segment, active range of motion in the loaded direction, resting heart rate variability, and time to return to baseline after a standardized repeated-loading test. Our model further predicts that this set forecasts function at 12 months more accurately than any finding on the image. An input that restores regulation moves high and low starters toward the same middle.

This is a claim about how accuracy is organized rather than a claim about what treatment does. If pressure pain threshold, active range of motion in the loaded direction, heart rate variability and time to return to baseline are shown to move together, the unification claim is confirmed.

10The tone reading

How clinical examination expresses tone

Every topic in this library expresses all of tone. In examination three aspects carry the signature, because the informative reading is what the body does under a demand rather than how it sits at rest.

Load

The examination is a graded load. Repeated end-range loading elicited a lasting direction of preference in 230 of 312 patients, and that direction guided what followed.

Time course

The return is the reading. A two-point change in pain by the second visit described the outcome at discharge correctly in 67 percent of 100 patients.

Input quality

Hands read what a picture cannot. Pain provocation tests reached acceptable agreement in 64 percent of studies, and soft tissue palpation reached it in none.

The remaining foundations run through the examination as well. Constraint: guarding narrows the range an examiner can move, and the narrowed range is itself a finding. Coupling: breathing, trunk motion and pelvic position change together under a load test, so one segment never answers alone. Gain: a segment that reports pain to light pressure is running a raised amplifier, which is why provocation thresholds carry information. Set point: the resting level of protection decides how much load the examination needs before anything answers. Prediction: what a person expects a movement to do shapes what the movement produces, so the order of tests changes the result. Oscillation: a complaint that follows a daily rhythm points the examination toward the regulator rather than the tissue. These are readings of one organization rather than separate systems, which is the core claim of the Unified Model of Tone.

11Across the library

How this page relates to the rest of the library

The examination is what turns every imaging page in this section into a decision.

Why MRI Misleads

What happens when a picture arrives before the question, and how the cascade toward surgery begins.

Findings in People Without Pain

The age-stratified prevalence of bulges, protrusions and degeneration in people with no symptoms at all.

What MRI Is Really For

When a scan is warranted, and the narrow set of questions it answers better than any examination.

The Red Flags Clinicians Screen For

Which history and examination items carry diagnostic weight, and the false-positive burden they bring.

What an Adjustment Is Really Doing

Specificity as correspondence between the input and the pattern, which is what an assessment exists to find.

When Conservative Care Stops

The thresholds at which a reassessment stops being a reassessment and becomes a referral.

The Neurological Exam

The reflex, motor and sensory testing that turns a symptom description into a localized finding.

12Frequently asked

Questions patients ask about the exam and the scan

Why does the examination matter more than the scan?

Because the examination measures what the person can do, and the scan records what the spine looks like at one instant. In 80 medical outpatients, the history alone led to the final diagnosis in 61 cases, and the physical examination in 10 more. In back pain specifically, how symptoms move under repeated loading predicted chronic pain and disability at one year, while static examination findings rarely did. The scan then answers a narrow question: is neural tissue compressed anywhere along its path.

Can a chiropractor read my images?

Yes. Interpreting imaging is a core clinical skill. The study itself is performed elsewhere, at an imaging center, and the clinician decides when a scan is warranted, refers for it, then reads the report and the images against the history and examination. That correlation is the work, and it matters because reports vary. When one patient was scanned at 10 different centers within three weeks, agreement across the 10 reports reached a Fleiss kappa of 0.20, and the average report missed 10.9 of 25 reference findings.

Does a scan finding explain my pain?

Only if it matches the clinical picture. A finding is a clue to interpret, never a verdict on its own, which is why correlation with symptoms is essential. Disc degeneration, mild bulges and ordinary signal change are common in people who feel completely well. The question a clinician asks is narrow: does this anatomy explain this presentation, or is the abnormality incidental and better left alone. Pain in one region does not automatically identify the source of that pain.

What is a challenge and return test?

It is any examination step that applies a load, records how the body answers, then records how quickly it returns. Repeated end-range movement testing is the best-studied version in back pain. Symptoms that retreat toward the spinal midline under repeated loading, a response called centralization, appeared in 44.4 percent of 4,745 patients across 29 studies and in 74 percent of acute cases. That response carried prognostic weight in 21 of 23 studies. The starting number matters less than the ability to change and return.

Why do two radiologists describe the same scan differently?

Because a report is a reading rather than a measurement. One patient was scanned at 10 imaging centers within three weeks. Across the 10 reports, 49 distinct findings were named at specific motion segments, no finding appeared in all 10 reports, and 32.7 percent appeared only once. Agreement reached a Fleiss kappa of 0.20. Expert reading is a demanding skill, which the same study showed by using two subspecialist spine radiologists reading by consensus as its reference standard for that spine.

If the history gives the diagnosis, what does the examination add?

Exclusion and confidence. In one prospective study of 80 outpatients, the history led to the final diagnosis in 76 percent of cases and the physical examination in 12 percent. The examination still raised the confidence of the internists in the correct diagnosis from 7.1 to 8.2 on a ten-point scale. It also ruled out possibilities the history had left open. In musculoskeletal care that exclusion is the point, because ruling out fracture, infection and progressive neurological loss is what makes a conservative plan safe to start.

What does the Unified Model of Tone say about clinical examination?

That assessment, and not delivery, is the true seat of accuracy, because an input can only correspond to a pattern someone has already read. The act of reading tone changes tone, so the same contact can be both an assessment and a treatment in the same motion. Repeated movement testing shows this directly. The assessment that identified a direction of preference in 230 of 312 patients did so by producing an immediate lasting improvement in pain. The model predicts that challenge measures share one underlying factor.

13The sources

References

1
Hampton JR, Harrison MJ, Mitchell JR, Prichard JS, Seymour C. Relative contributions of history-taking, physical examination, and laboratory investigation to diagnosis and management of medical outpatients. Br Med J. 1975. PMID 1148666
2
Peterson MC, Holbrook JH, Von Hales D, Smith NL, Staker LV. Contributions of the history, physical examination, and laboratory investigation in making medical diagnoses. West J Med. 1992. PMID 1536065
3
Doubilet P, Herman PG. Interpretation of radiographs: effect of clinical history. AJR Am J Roentgenol. 1981. PMID 6975000
4
Houssami N, Irwig L, Simpson JM, McKessar M, Blome S, et al. The influence of clinical information on the accuracy of diagnostic mammography. Breast Cancer Res Treat. 2004. PMID 15111760
5
Herzog R, Elgort DR, Flanders AE, Moley PJ. Variability in diagnostic error rates of 10 MRI centers performing lumbar spine MRI examinations on the same patient within a 3-week period. Spine J. 2017. PMID 27867079
6
Seffinger MA, Najm WI, Mishra SI, Adams A, Dickerson VM, et al. Reliability of spinal palpation for diagnosis of back and neck pain: a systematic review of the literature. Spine (Phila Pa 1976). 2004. PMID 15454722
7
Stochkendahl MJ, Christensen HW, Hartvigsen J, Vach W, Haas M, et al. Manual examination of the spine: a systematic critical literature review of reproducibility. J Manipulative Physiol Ther. 2006. PMID 16904495
8
Werneke M, Hart DL. Centralization phenomenon as a prognostic factor for chronic low back pain and disability. Spine (Phila Pa 1976). 2001. PMID 11295896
9
May S, Aina A. Centralization and directional preference: a systematic review. Man Ther. 2012. PMID 22695365
10
Laslett M, Oberg B, Aprill CN, McDonald B. Centralization as a predictor of provocation discography results in chronic low back pain, and the influence of disability and distress on diagnostic power. Spine J. 2005. PMID 15996606
11
Long A, Donelson R, Fung T. Does it matter which exercise? A randomized control trial of exercise for low back pain. Spine (Phila Pa 1976). 2004. PMID 15564907
12
Cook CE, Showalter C, Kabbaz V, O’Halloran B. Can a within/between-session change in pain during reassessment predict outcome using a manual therapy intervention in patients with mechanical low back pain?. Man Ther. 2012. PMID 22445052

12 primary sources, each linked to its record. Figures quoted on this page were checked against the published abstract.

Related evidence

← All 44 lessons