Orthopedics · Part Four · Seeing and Ruling Out
Lesson 33 / 44
The Red Flags Clinicians Screen For: What They Catch, and What They Miss
Most back pain is benign, and the real question a patient carries into the room is when back pain is something serious rather than routine.
Red flags are the warning features clinicians screen for when back pain might have a rare serious cause: fracture, malignancy, infection, cauda equina compression, myelopathy, or vascular disease. Specific spinal pathology accounts for under 1 percent of primary care presentations. A 2013 systematic review of 53 red flags found that only a small subset changes the probability of disease meaningfully. The Unified Model of Tone reads a red flag as one input to a judgment rather than a trigger.
Red flags evaluated in the 2013 diagnostic accuracy review
53, across 14 studies
Post-test probability of spinal fracture when several red flags appear together
90 percent
Patients with spinal malignancy who carried no red flag at all
64 percent
Serious spinal pathology in emergency department series
2.5 to 5.1 percent
Red flags and diagnostic triage
A red flag is a feature of the history or the physical examination that raises the probability of a specific serious cause of back pain. Diagnostic triage is the sorting step it belongs to. Each presentation is placed in one of three groups: specific spinal pathology, radicular syndrome, or nonspecific back pain. The sorting is repeated at every visit.
How a screening question moves a probability
A red flag works by revising an estimate that already exists. The clinician holds a starting probability built from age, history and setting. A finding with a high likelihood ratio pushes that estimate sharply upward. A finding with a ratio near one leaves it almost where it was. The lower the starting probability, the more work a single finding has to do.
01The real question
Serious causes of back pain are rare, and screening is how a clinician confirms it
Specific spinal pathology accounts for less than 1 percent of low back pain presenting to primary care. Diagnostic triage sorts each presentation into three groups. Radicular syndrome accounts for roughly 5 to 10 percent, and nonspecific low back pain for the remaining 90 to 95 percent (Bardin 2017).
Most back pain is benign, mechanical and self-limiting. A clinician answers the question a person actually brought into the room by screening a short list of rare dangerous categories, then treating with confidence once they are excluded. Finding nothing concerning is the ordinary, expected result. That negative screen is what allows conservative care to proceed with confidence.
The value of a trained portal-of-entry clinician is not that danger is common. It is that a rare dangerous cause is caught early and referred without hesitation.
Red flags are best understood as the exception a careful chiropractor is built to detect, not a checklist for a person to run on themselves. The screening happens inside the consultation, in the questions asked and the examination performed. Nothing on this page is a test for a reader to apply to their own body.
The base rate depends on the room
Setting changes the arithmetic before a question is asked. A review of 22 emergency department studies covering 41,320 patients found serious spinal pathology in 2.5 to 5.1 percent of prospective series (Galliker 2020). Vertebral fracture ran from 0.0 to 7.2 percent and spinal cancer from 0.0 to 2.1 percent.
Infectious disorders ran from 0.0 to 1.9 percent, cord or cauda equina compression from 0.1 to 1.9 percent, and vascular pathology from 0.0 to 0.9 percent. Every one of those figures is higher than the primary care rate. The same question carries different weight in different rooms.
Age moves it too. Among 669 adults over 55 consulting a general practitioner for back pain, 6 percent had a serious underlying cause diagnosed within a year, and 33 of them had a vertebral fracture (Enthoven 2016). The prior probability shifts before the first question.
02Findings
What the research shows
From a systematic review of diagnostic accuracy, a 9,940-patient clinical series, two guideline reviews and an emergency department meta-analysis.
03Six categories, four in the guidelines
Six serious categories deserve deliberate attention on every presentation
The categories are fracture, malignancy or tumor, infection, cauda equina syndrome, myelopathy, and vascular causes such as an abdominal aortic aneurysm. Each is uncommon. Each has a recognizable signature. Each is exactly the kind of thing a diagnostically literate clinician is trained to separate from the mechanical back pain that fills most of the day.
Published guidelines formalize four of the six. A review of 16 national guidelines from 15 countries, plus one written for Europe as a whole, found 46 discrete red flags grouped under malignancy, fracture, cauda equina syndrome and infection (Verhagen 2016). Most guidelines carried two flags for fracture and two for malignancy.
The other two categories carry their own lessons here. Fracture accuracy sits on The Missed Fracture, and imaging for malignancy and infection on Excluding the Dangerous on Imaging. Cauda Equina and the Spinal Emergencies holds the emergency, and The Slow Cord Compression holds myelopathy.
Vascular causes run on Vascular Mimics and Claudication. Two neighbors apply the same logic outside the low back. When a Headache Is Dangerous screens head pain, and Inflammatory Back Pain separates an inflammatory pattern from a mechanical one.
Where the guideline lists came from
The provenance of those 46 flags is thin. Eight of the 16 guidelines based their choice on consensus or on earlier guidelines, and five supplied no reference at all (Verhagen 2016). Accuracy data was rarely presented.
A separate review traced the origin of individual malignancy flags to case reports or to no identifiable source (Verhagen 2017). The lists circulated faster than the evidence that would have justified them.
04How the red flags perform
Most individual red flags barely change the probability of fracture or malignancy
The definitive test of that question was published in the BMJ in 2013. Fourteen diagnostic studies met the criteria, eight from primary care, two from secondary care and four from tertiary care, together evaluating 53 red flags (Downie 2013). Only five of the 14 studies examined combinations of flags.
Pooling was not possible, because the index tests differed too much from study to study. The reviewers therefore reported diagnostic accuracy and post-test probability flag by flag. Their summary was blunt. Many red flags in current guidelines provide virtually no change in the probability of fracture or malignancy, or have untested diagnostic accuracy.
They also named the ones that work. For fracture, the presence of a contusion or abrasion reached a post-test probability of 62 percent, prolonged corticosteroid use 33 percent, severe trauma 11 percent, and older age 9 percent. For malignancy, a history of malignancy reached 33 percent, and no other single flag came close.
Combinations changed the picture entirely. Where multiple red flags were present together, the probability of spinal fracture reached 90 percent. The discipline here is combination, not alarm.
Why a weak flag cannot rescue a low starting probability
The arithmetic behind that result is ordinary diagnostic reasoning. A likelihood ratio near unity has little effect on decision making, while a high or low ratio can greatly shift a clinician’s estimate of the probability of disease (Grimes 2005).
Start from a prior under 1 percent, which is where primary care back pain begins (Bardin 2017). A flag with a ratio near one leaves that estimate essentially untouched. This is why a long list of weak questions produces a long list of unchanged probabilities.
The malignancy literature says the same thing from a different angle. One review identified 13 red flags endorsed across 16 guidelines and examined 33 publications. Only seven supplied diagnostic accuracy data at all (Verhagen 2017). Two flags survived: a history of malignancy, and strong clinical suspicion.
That second one deserves a pause. Strong clinical suspicion is the clinician’s own integrated judgment, and it performed as well as any item on the list. The reviewers put the incidence of malignancy in primary care back pain at 0 to 0.7 percent.
05False positives and false comfort
A negative answer to one or two screening questions does not rule serious disease out
The largest clinical test of the questions themselves reviewed 9,940 patients presenting with low back pain to a multidisciplinary academic spine center (Premkumar 2018). Each completed a red flag questionnaire at the first physician visit. Diagnoses were taken from the medical record and corroborated against imaging reports.
Some flags earned their place. Recent trauma and an age over 50 were associated with vertebral fracture, as the accuracy review had also found. Others were associated with nothing. Night pain, one of the most widely circulated flags of all, was unrelated to any particular diagnosis.
The false positive burden is the number worth carrying. In patients with no recent history of infection and no fever, chills or sweating, night pain was a false positive finding for infection more than 96 percent of the time. Acting on that flag alone would investigate a great many people who are well.
The reciprocal error is larger still. Of the patients who did have spinal malignancy, 64 percent carried no associated red flags. The absence of red flag responses did not meaningfully decrease the likelihood of a red flag diagnosis, and the authors advised caution in using these questions as screening tools.
Adding more flags to the model did not help
A prospective cohort of older adults tested whether stacking flags improves prediction. Trauma had the highest positive predictive value for vertebral fracture at 0.25, with a positive likelihood ratio of 6.2 (Enthoven 2016). Four other associated flags were added to a prediction model.
The model did not improve on trauma alone. Age of 75 or over, osteoporosis, a pain intensity score of 7 or more and thoracic pain were each associated with fracture, and combining them raised neither the predictive value nor the likelihood ratio. The authors noted that low prevalence could have produced findings by chance.
None of this argues for a lighter screen. It argues for a screen read as a whole. Pain location alone should never exclude a fracture, and easy reassurance is never a substitute for a deliberate look.
06The signature of each category
Each dangerous category of back pain announces itself through pattern rather than intensity
A vertebral fracture can hide inside what looks like ordinary mechanical back pain, which is why bone is always considered a possible pain generator. In weakened bone, a vertebra can fail under a trivial load. A clinician weighs advancing age, prolonged corticosteroid use, prior fracture and sudden onset after a minor event rather than resting on any single finding.
Malignancy announces itself through pattern rather than through pain intensity alone. The features weighed together include a prior history of cancer, unexplained weight loss, pain that is unrelenting and present at rest, and night pain that does not settle with position. Only the first of those reached a post-test probability above 30 percent (Downie 2013).
No one red flag confirms anything on its own. A cluster prompts a careful clinician to investigate rather than to keep treating. The constellation, read against age and history, is what shifts a thoughtful clinician from routine care toward imaging or referral.
Infection, and what actually raises the suspicion
Spinal infection is rare, and its warning features are fever, night pain, and pain that behaves differently from a mechanical strain, sometimes on a background of immune compromise or recent illness. The emergency department review identified the features that genuinely raise the likelihood of an epidural abscess (Galliker 2020).
They are intravenous drug use, an indwelling vascular catheter, and another site of infection elsewhere in the body. Each names an actual route by which organisms reach the spine. A clinician who sees fever paired with escalating spinal pain treats that combination as a reason to look further rather than to adjust and reassess in a week.
Cauda equina, myelopathy, and the vessel
Cauda equina syndrome sits at the top of the urgency list because the window to act is short. The features screened for are a change in bladder or bowel control, numbness across the saddle region, and progressive weakness in the legs. These point to compression of the descending nerve roots within the central canal.
Myelopathy is compression of the spinal cord itself, and its signature is progressive neurological deficit rather than local ache. A clinician watches for worsening coordination, changes in gait and balance, altered reflexes, and weakness that advances rather than fluctuates. Progressive is the operative word.
A deficit that deepens over days or weeks is treated very differently from one that improves as tissues settle. Vascular causes are the sixth category, and the one held in mind for the older patient is an abdominal aortic aneurysm. It can present as deep, unrelenting pain that does not track with movement or posture.
Pain that ignores the usual mechanical rules, especially in an at-risk person, is a prompt to broaden the assessment rather than to narrow it. A person is not asked to self-diagnose bladder change or saddle numbness. A trained clinician asks the right questions on every visit, so the rare case is recognized the moment it appears.
Across all six categories, the same protective logic applies. When pain persists, escalates, or refuses to behave mechanically, a diagnostically literate clinician reads it as the nervous system holding the body in a protective, threat-driven state. That state deserves explanation rather than dismissal. Screening is how it is interpreted correctly and, when necessary, acted upon fast.
07Screening across visits
Screening runs across the course of care, and that is what makes conservative care safe
An international framework published in 2020 settled how the flags should be used. The International Federation of Orthopaedic Manipulative Physical Therapists stated plainly that high-quality evidence for the diagnostic accuracy of most red flags is absent. In its place the framework supplies a clinical reasoning pathway (Finucane 2020).
A review of the 2020 to 2025 literature reached the same conclusion from the other direction. Serious spinal conditions are rare, and clinicians should assess overall concern from a combination of alerting features rather than from isolated red flags (Pinto 2026).
The practical form of that is safety netting. Across 47 studies it is defined as a consultation technique that communicates uncertainty, gives information on warning symptoms, and plans follow-up so a condition is reassessed in time (Jones 2019). Applied to back pain, the screen has a schedule.
The response to care is itself information
A presentation that behaves mechanically responds to mechanical inputs. One that does not is a reason to reconsider the category it was placed in. The Clinician’s Advantage sets out the challenge-and-return test that turns each visit into another reading.
Screening for red flags is not a search for danger in every patient. It is the quiet expertise that lets a chiropractor treat the vast majority with total confidence, precisely because the rare exception will not slip past.
Rule out, then treat
Ruling out the dangerous is not a limitation on chiropractic. It is the credential that lets a chiropractor function as a genuine portal-of-entry clinician who commands the right diagnostic tools and knows exactly when to use them.
Imaging is one of those tools, and it is ordered rather than performed here. The practice refers out for radiographs, computed tomography and magnetic resonance imaging, then reads the report against the history and the examination and coordinates with radiology and medicine. What MRI Is Really For sets out which question each scan answers.
The reassuring truth remains first and last. Most back pain is benign, most people recover, and a confident clinician can say so plainly because the screen has already done its work. Rule out the dangerous, then treat the rest with conviction. Escalation thresholds sit on When Conservative Care Stops.
08Claims removed from this page
Three red flag claims from the earlier version were changed or removed
A gold pull-quote attributed to Dr. Jason Dulberg came off the page, because its wording could not be matched to any recorded source. The figure stating that fewer than one third of osteoporotic vertebral compression fractures are correctly diagnosed came off as well. It carried no citation here, and The Missed Fracture now carries the cohort behind it.
The osteoporosis threshold of a T-score below minus 2.5 moved to the same page, where the fracture risk data sits. Night pain kept its place in the list and gained the number that belongs to it, a false positive rate above 96 percent for infection (Premkumar 2018).
The short memorable set of red flags survives unchanged, and so does the order of the argument. What is new is that each warning feature now travels with the accuracy figure attached to it.
09The model on screening
What the Unified Model of Tone claims about red flag screening
Everything above is established science, including the finding that most red flags perform poorly in isolation. What follows is this model’s reading, stated as ours rather than drawn from the papers cited.
Our model puts accuracy in assessment rather than in delivery. Assessment is the true seat of accuracy, and that single move reframes what a red flag is. A red flag is an input to a judgment, never a trigger that fires an action on its own.
The diagnostic accuracy data behaves exactly as that reading predicts. Tested as individual triggers, 53 flags mostly failed (Downie 2013). Tested as contributors to one combined judgment, the same features reached a post-test probability of 90 percent for fracture. The flags did not change. The judgment did.
Why a single finding rarely names the process
Every body carries a different structure of constraint, and a disturbance surfaces wherever that particular body has the least reserve to absorb it. Our model holds that the site of expression is not reliably the site of origin. One finding therefore rarely identifies the process that produced it.
Night pain is the clean example. In a body whose sleep regulation is already strained, night pain reports the regulator. In a body harboring an infection, it reports the infection. The feature itself carries no way to tell those apart, which is what a false positive rate above 96 percent looks like in a clinic (Premkumar 2018).
Delay is its own harm
Our model states the counterweight just as firmly. When a distortion has descended past what any surface input can reach, the larger magnitude is the correct one, and delay becomes its own kind of harm. Escalation is correct dosing.
Referral, imaging and surgery are inputs of larger magnitude, and they are right when the situation calls for them. A screen that finds something is a success of exactly the same kind as a screen that finds nothing. That is why the 2.5 to 5.1 percent found in emergency departments carries as much weight in the reasoning as the large majority who are well (Galliker 2020).
The prediction
From that follows a claim the screening literature does not make. Our model predicts that the informativeness of any single alerting feature depends on the regulatory state it appears in, and that this state is measurable at the first visit.
Four measures recorded together in the same people carry it. They are pressure pain threshold at a site away from the painful region, active lumbar range of motion, resting heart rate variability, and time to return to baseline after a standardized load test.
The concrete prediction is this. Take people presenting with back pain and one isolated red flag. Those whose four measures have drifted together will show a higher rate of confirmed serious pathology than those whose measures sit near their own baseline. Adding the four to a flag list should lift positive predictive value above the 0.25 that trauma alone achieved (Enthoven 2016).
This is a claim about how a disturbance organizes itself rather than a claim about what treatment does. If pressure pain threshold, active lumbar range of motion, resting heart rate variability and time to return to baseline after a load test are shown to move together, the unification claim is confirmed.
10The tone reading
How red flag screening expresses tone
Every topic in this library expresses all of tone. In red flag screening three aspects carry the signature, because 64 percent of the people found to have spinal malignancy reached the clinic with no red flag at all.
Constraint
Serious spinal pathology surfaces where reserve has already run out. Among adults over 55 consulting a general practitioner, 6 percent carried a specific cause.
Time course
Progression separates dangerous from mechanical. A deficit that deepens over weeks reads differently from one that settles, which is why screening repeats at every visit.
Input quality
A single screening question carries thin information. Night pain was a false positive for infection more than 96 percent of the time in one 9,940-patient series.
The remaining foundations run through screening as well. Prediction: the clinician holds a prior probability, and every question either revises it or leaves it where it was. Gain: a protective nervous system amplifies ordinary sensation, which is why pain severity ranks so poorly as a screening feature. Set point: fever and unexplained weight loss are regulated values that have moved, which is what makes them worth asking about. Load: severe trauma reached a post-test probability of 11 percent for fracture because it names an input the tissue may not have absorbed. Coupling: bladder, bowel and leg strength fail together in cauda equina compression because one structure serves all three. Oscillation: pain that keeps a strict nightly rhythm is reporting a regulator, and pain that ignores rhythm entirely is reporting something else. These are readings of one organization rather than separate systems, which is the core claim of the Unified Model of Tone.
11Across the library
How this page relates to the rest of the library
The reasoning framework sits here, and each dangerous category is carried in full by its own lesson.
Vertebral fracture carries the strongest accuracy data in the section, including the decision rules and the underdiagnosis figures.
Which modality answers which question once malignancy or infection is suspected, and how sensitive each one is.
Incidence, the full feature set, and why time to decompression is the variable that decides the outcome.
Myelopathy graded on the mJOA scale, and the tensile stress reading of why the cord loses function.
Abdominal aortic aneurysm and peripheral arterial disease, and how neurogenic claudication is separated from vascular.
The same screening logic applied to head pain, with the decision rules for subarachnoid hemorrhage.
The criteria that separate an inflammatory pattern from a mechanical one, read from time course rather than from a single feature.
12Frequently asked
Questions patients ask about red flags and serious back pain
What are the red flags for back pain?
The warning features a clinician screens for are unrelenting night pain, unexplained weight loss, fever, progressive neurological deficit, change in bladder or bowel control, and a history of cancer. They point toward six categories: fracture, malignancy, infection, cauda equina syndrome, myelopathy, and vascular disease. A review of 16 national guidelines counted 46 discrete red flags in circulation, grouped under four of those categories. The screening happens inside a consultation, in the questions a clinician asks and the examination performed. None of it is a self-test.
How accurate are red flags at detecting serious disease?
Individually, most of them are poor. A systematic review of 14 diagnostic studies evaluated 53 red flags and concluded that many provide virtually no change in the probability of fracture or malignancy, or have never been tested. The informative ones are few. A contusion or abrasion reached a post-test probability of 62 percent for fracture, and a history of malignancy reached 33 percent for cancer. Where several flags appeared together, the probability of fracture reached 90 percent. Combination is what carries the accuracy.
If someone has a red flag, does that mean something is seriously wrong?
Usually not. Red flags are common in people who turn out to be well, which is why single findings mislead. Among patients with no recent infection and no fever, chills or sweating, night pain was a false positive finding for infection more than 96 percent of the time. A single feature in isolation is common and usually benign. What shifts a clinician toward imaging or referral is a cluster of features read against age and history, together with how the pain has behaved so far.
Can a screen miss a serious cause?
It can, which is why screening repeats. In a series of 9,940 patients with low back pain, 64 percent of those who had spinal malignancy reported no red flags at their first visit. The absence of red flag answers did not meaningfully reduce the likelihood of a serious diagnosis. A clinician therefore treats the first screen as a starting reading rather than a verdict. The questions are asked again at every visit, and how the presentation responds to care becomes part of the evidence.
How often is back pain caused by something dangerous?
Rarely. Specific spinal pathology accounts for under 1 percent of low back pain presenting to primary care, with roughly 90 to 95 percent classed as nonspecific. The figure rises with setting and age. Across 22 emergency department studies covering 41,320 patients, serious pathology ran at 2.5 to 5.1 percent in prospective series. Among 669 adults over 55 consulting a general practitioner, 6 percent had a serious cause diagnosed within a year, most of them vertebral fractures. Age and setting both move the starting probability.
Why keep screening if the individual questions perform badly?
Because the combination performs well and the cost of missing a case is high. The 2013 accuracy review found that probability of spinal fracture reached 90 percent when multiple red flags were present, against 9 percent for older age and 11 percent for severe trauma taken alone. An international framework published in 2020 turned this into a clinical reasoning pathway rather than a checklist. Strong clinical suspicion, the clinician’s own integrated judgment, was one of only two red flags for malignancy with evidence of acceptable accuracy.
What does the Unified Model of Tone say about red flag screening?
That accuracy lives in assessment rather than in delivery, so a red flag is an input to a judgment rather than a trigger. A disturbance surfaces wherever a particular body has the least reserve, so one finding rarely names the process behind it. The model also states the counterweight: when a problem has descended past what any surface input can reach, the larger magnitude is correct, and delay becomes its own kind of harm. Escalation is correct dosing rather than a failure of conservative care.
13The sources
References
11 primary sources, each linked to its record. Figures quoted on this page were checked against the published abstract.
Related evidence