Our Approach · The History · Act V
1991 to 1993 · Evidence
The Evidence Era
When the manipulation of the spine met the methods of measurement
The Evidence Era names the window from 1991 to 1993 when the RAND Corporation rated 1,570 indications for spinal manipulation. In the same window the Manga Report costed chiropractic management of low back pain for the Ontario Ministry of Health. Both reported in its favor, both found modest average effects, and chiropractic entered the mainstream by measurement rather than by argument. Why careful trials keep reading a large effect as a small one is the trial prediction stated by the Unified Model of Tone.
forthcoming
Key dates
RAND R-4025 series 1991 to 1992; Manga Report August 1993
Place
RAND Corporation, Santa Monica; Pran Manga and Associates, Ottawa
Known for
1,570 indications rated on a nine-point scale; Manga Report, 104 pages
Legacy
AHCPR Guideline 14, December 1994; 1998 audit found 46% appropriate
The claim
Two institutions outside the profession measured manipulation by their own rules
Between 1991 and 1993 the RAND Corporation in Santa Monica and a health economics team in Ottawa examined spinal manipulation for low back pain and reported in its favor. Neither was a chiropractic body. RAND published the R-4025 series under Paul G. Shekelle, funded through the Consortium for Chiropractic Research and the Foundation for Chiropractic Education and Research under grants 89-038 and 89-039. Ottawa produced the document everyone now calls the Manga Report, written by Pran Manga, Douglas E. Angus, Costa Papadopoulos and William R. Swan, funded by the Ministry of Health of the Government of Ontario and published in August 1993 at 104 pages.
Ask what that combination actually establishes. It establishes that outsiders holding formal methods, asked to look at a manual therapy, did not dismiss it. It does not establish that the profession had its own account of why manipulation works put to the test. Standing and vindication are not the same thing. That distinction runs through this entire page. The Evidence Era gave chiropractic standing in the language the health system respects. It also fixed the terms on which chiropractic would be judged for the next thirty years, and those terms were pain in the lower back, measured in weeks.
The method
RAND did not run a trial, it ran a rating
The RAND appropriateness method asks expert panels to judge described clinical scenarios, not to generate new patient data. A procedure counts as appropriate when the expected health benefit exceeds the expected negative consequences by a wide enough margin that the procedure is worth doing, exclusive of cost. Benefits in that calculation include pain relief, reduced anxiety and improved functional capacity. Harms include mortality, morbidity, pain and disruption of work. Panellists rate each scenario on a nine-point scale, once privately and once after group discussion.
Ask what such a method can and cannot deliver. It can deliver a disciplined map of expert opinion, indication by indication, in a field where trials are thin. It cannot deliver proof. RAND has always been explicit that appropriateness ratings encode judgment, informed by published evidence where published evidence exists. The chiropractic profession accepted that deal and paid for it. Two chiropractic research bodies handed an outside institution the power to publish an unflattering answer, and RAND eventually published one. A profession expecting to fail an audit does not commission the audit. That decision, taken in 1989, is the quiet act of nerve underneath everything that followed.
The literature
Fifty-eight articles, twenty-five controlled trials, one modest number
The first volume, R-4025/1, Project Overview and Literature Review, ran to twenty-nine pages of text and appeared in 1991 with the identifier 0-8330-1150-2 (Shekelle 1991). Its findings reached a far wider readership the following year in the Annals of Internal Medicine, where Shekelle, Adams, Chassin, Hurwitz and Brook reviewed fifty-eight articles including twenty-five controlled trials (Shekelle 1992). For acute uncomplicated low back pain, the difference in probability of recovery at three weeks favoring manipulation was 0.17. Where sciatic nerve irritation was present, the difference fell to 0.098.
Hold those two numbers side by side. Seventeen percentage points at three weeks for simple mechanical back pain. Under ten points once a nerve root is involved. The signal was real, modest and narrow. The same review recorded that serious complications including paraplegia and death had appeared in case reports, while judging the rate probably low. Both facts were published in one document and both should travel together. That is honest reporting and this page will not soften it. Manipulation is a mechanical input delivered into a loaded structure. Mechanical inputs carry mechanical risk. The evidence said the risk was small and the benefit was small, in a condition that mostly resolves on its own.
short-term benefit in some patients, particularly those with uncomplicated, acute low-back pain
Shekelle, Adams, Chassin, Hurwitz and Brook · Annals of Internal Medicine, 1992, volume 117, pages 590 to 598The first panel
Nine experts from four disciplines rated the same list twice
R-4025/2, Indications and Ratings by a Multidisciplinary Expert Panel, appeared in 1991 at eleven front pages plus ninety-nine, identifier 0-8330-1145-6 (Shekelle 1991). The panel seated nine back pain experts drawn from orthopedics, chiropractic, family medicine and neurology. They worked through structured clinical presentations assembled from history, examination findings, imaging, duration of symptoms and response to prior treatment. RAND describes 1,570 indications for the use of spinal manipulation for low back pain rated on the nine-point scale.
Ask why the list had to be that large. Because appropriateness is never a verdict on a technique. It is a verdict on a technique applied to a described patient in a described state. Manipulation for four weeks of mechanical low back pain in an otherwise healthy adult is a different question from manipulation for the same pain accompanied by a progressive neurological deficit. Every clinician knows this instinctively. The presentation is the question. RAND simply wrote the instinct down in a form a health system could audit. This is the most underrated gift of the Evidence Era. It moved the argument from does this work to for whom, in what state, and for how long.
The second panel
An all-chiropractic panel rated identical indications and disagreed
RAND then repeated the whole exercise with a panel drawn entirely from chiropractic. That volume appeared in 1992 at 117 pages, identifier 0-8330-1264-9 (Shekelle 1992). The design was deliberate and unusually clever. Run one list of indications past a mixed panel and a single-profession panel, then compare the maps. RAND reported that the chiropractic panel was able to formulate detailed lists of indications and rate their appropriateness, which was itself a finding about professional maturity rather than about spines.
The comparison is where the story turns. Two competent panels, one instrument, one list, two different pictures of appropriate care. Ask what that means. It means the instrument was not measuring the spine at all. It was measuring the assumptions each discipline carries to the spine. Nobody cheated. Both panels read the same thin literature and filled the gaps with clinical judgment, and the judgment diverged because the underlying models of the problem diverged. Evidence-based medicine runs cleanly where trials are dense. Where trials are sparse, appropriateness ratings become a high-resolution portrait of the priors of the raters.
The pilot
The same care was 38 percent appropriate or 74 percent appropriate depending on the judge
In 1995 the RAND team tested both rating sets against real practice. Shekelle, Hurwitz, Coulter, Adams, Genovese and Brook published a pilot study in the Journal of Manipulative and Physiological Therapeutics, volume 18, issue 5, pages 265 to 270 (Shekelle 1995). Eight of thirteen eligible chiropractors took part, a participation rate of 62 percent. Judged against the multidisciplinary criteria, 38 percent of care was appropriate and 26 percent inappropriate. Judged against the all-chiropractic criteria, 74 percent was appropriate and 7 percent inappropriate. The two panels agreed on 48 percent of cases.
Read that as a measurement problem rather than a political one. One system, two instruments calibrated to different reference frames, readings that differ by roughly a factor of two. The honest conclusion is not that one panel was right and the other wrong. The honest conclusion is that appropriateness is frame dependent in a way that length and mass are not. It also tells you how much of the disagreement in this field is empirical and how much is doctrinal. The 1995 paper said so in plain language, reporting the rate of appropriate care as lying somewhere between 38 and 74 percent. Few findings in the history of this field are more useful and fewer still are quoted less often.
The field test
Turned on real charts in 1998, the method found 29 percent inappropriate
The full field study appeared in the Annals of Internal Medicine on 1 July 1998, volume 129, pages 9 to 17, authored by Shekelle, Coulter, Hurwitz and colleagues (Shekelle 1998). Of 1,310 patients who sought chiropractic care for low back pain, 1,088, or 83 percent, received spinal manipulation. Among the 859 cases documented well enough to rate, 46 percent were judged appropriate, 25 percent uncertain and 29 percent inappropriate. A companion paper by Hurwitz and colleagues in the American Journal of Public Health in May 1998 recorded 101.2 chiropractic visits per 100 person-years across the United States sites and 140.9 in Ontario (Hurwitz 1998).
The authors made the fair comparison themselves. They noted that the proportion of care congruent with appropriateness criteria resembled proportions already described for accepted medical procedures. They also noted that more than a quarter of patients were treated for indications judged inappropriate. Both halves of that sentence matter and both belong in any honest retelling. A profession that quotes only the first half is not being audited, it is being advertised. Cite the whole finding or cite none of it.
Ottawa
The Manga Report was 104 pages of health economics commissioned by a provincial government
In August 1993 Pran Manga and Associates Inc. of Ottawa published A Study to Examine the Effectiveness and Cost-Effectiveness of Chiropractic Management of Low-Back Pain, funded by the Ontario Ministry of Health (Manga 1993). It runs 104 pages with a bibliography from page 87 to page 104, identifier 0-9697615-0-3. Pran Manga held a chair in health economics at the University of Ottawa. Douglas E. Angus directed the Masters Program in Health Administration at the same institution. Ontario had reason to ask the question. Musculoskeletal disorders rank first as a reason for consulting a health professional in the province and first as a cause of long-term disability. The report surveyed the treatment literature, compared the cost of medical and chiropractic management, and closed with ten numbered recommendations.
Those recommendations were structural rather than clinical. Full insurance of chiropractic services under the Ontario Health Insurance Plan. Chiropractors employed in hospitals. Hospital privileges and access to patient records. A seat in Workers Compensation Board policy. Public funding for university-based chiropractic education. Ask what a health economist has to believe before writing that list. He has to believe the cheaper provider is also the safer and the more effective one for the condition under review. Manga wrote it down without decoration.
There is no clinical or case-control study that demonstrates or even implies that chiropractic spinal manipulation is unsafe in the treatment of low-back pain.
Manga, Angus, Papadopoulos and Swan · The Effectiveness and Cost-Effectiveness of Chiropractic Management of Low-Back Pain, 1993, executive summaryThe limits
Manga synthesized evidence and modeled policy, it did not measure patients
The 1993 report was a review and a policy analysis, not a clinical trial. Every effectiveness claim inside it rests on studies conducted elsewhere. Chief among them was the randomized comparison by Meade, Dyer, Browne, Townsend and Frank published in the British Medical Journal on 2 June 1990, which followed 741 patients aged 18 to 65 for up to two years and found roughly a seven point advantage on the Oswestry disability scale for chiropractic over hospital outpatient management, concentrated in those with chronic or severe pain (Meade 1990).
Two cautions belong on the record. First, the savings figure most often attached to the Manga Report does not come from the Manga Report. The estimates of 770 million, 548 million and 380 million dollars in direct savings, together with 3.775 billion, 1.849 billion and 1.255 billion in indirect savings, come from a separate paper by Manga and Angus dated February 1998 on enhanced chiropractic coverage under OHIP (Manga and Angus 1998), and they are projections conditional on a proposed policy change costing an additional 200 million dollars by the year 2000. Second, the 1993 report speaks of savings in the hundreds of millions annually without publishing an audited figure. Use it as an argument. Do not use it as an accounting.
The guideline
In December 1994 a United States federal panel wrote manipulation into national guidance
The Agency for Health Care Policy and Research published Clinical Practice Guideline Number 14, Acute Low Back Problems in Adults, in December 1994 under panel chair Stanley J. Bigos, carrying AHCPR Publication Number 95-0643 (Bigos 1994). The guideline placed manipulation among the recommended options for the first month of acute symptoms in the absence of radiculopathy. It also drew boundaries. Beyond one month, efficacy was called unproven. If four weeks of manipulation produced no symptomatic and functional improvement, clinicians were told to stop and reassess the patient.
That paragraph is the shape of the entire Evidence Era in miniature. Approval inside a narrow window. Silence outside it. Guidelines are not theories. They are permissions with an expiry date, and this permission expired at four weeks. The mainstream did not adopt a theory of health. It adopted a procedure for a symptom, with a stopwatch attached. A profession that had spent ninety years arguing about the nervous system was admitted on the strength of four weeks of back pain. Standing was granted. The argument was postponed.
Manipulation, defined as manual loading of the spine using short or long leverage methods, is safe and effective for patients in the first month of acute low back symptoms without radiculopathy.
Bigos and the AHCPR panel · Acute Low Back Problems in Adults, Quick Reference Guide Number 14, AHCPR Publication 95-0643, December 1994The Evidence Era and the model
The model predicts the small averages the Evidence Era reported
The Evidence Era measured honestly and averaged the wrong variable. That is the reading of the Unified Model of Tone, and the model states it as a prediction rather than a complaint. In the model, the outcome of a manual input depends on the correspondence between the input and the constraint structure of each individual body. A trial that delivers the same predetermined input to everyone, a standardized manipulation of a fixed segment, therefore averages a well-matched intervention and a mismatched one across a sample that was never stratified by tone. The responders for whom the input fit and the non-responders for whom it did not are collapsed into a modest mean that describes neither. A genuinely large effect reads as weak.
Hold the numbers the era produced against that account. A probability difference of 0.17 at three weeks, thinning to 0.098 once a nerve root is involved, is the exact shape the model expects a pooled mean to take. So is a literature of 25 controlled trials that keeps finding the same small edge. Every trial RAND reviewed chose its input the same way for every back, and none recorded the state of regulation in the person before delivering it. The model reads the modest mean as an artifact of the design rather than as the size of the effect.
The model turns that reading into a trial anyone can run. Stratify a sample by a tone measure recorded before any input is given, and specify a leverage point for each person from that measure. Then randomize between an input delivered there and the identical input delivered to a site the measure did not select. The model predicts a substantially larger effect in the matched arm and a modest one in the mismatched arm. It predicts that pooling the two reproduces the small averages the literature keeps reporting. The separation between the two arms is what confirms it, and confirms with it the claim that correspondence rather than force is the active ingredient.
Fixing the site in advance is what keeps this a prediction. A leverage point identified after the result is known explains everything and forecasts nothing, and the model claims no such privilege. Say plainly where the record ends and the reading begins. Shekelle, Manga and Bigos measured what they were asked to measure, and nothing in their reports supports or refutes a tone model. What they proved is that this field can be audited by outsiders and survive the audit. Wilk v. AMA secured the right to practice, and the Evidence Era secured the right to be counted. The field lacks evidence organized around the correct variable, and that is a different scarcity from lacking evidence. The audit belongs to the era. The trial that reads the variable belongs to the model.
What the record shows
The measuring of manipulation in seven dated findings
- 2 June 1990. Meade, Dyer, Browne, Townsend and Frank in the British Medical Journal followed 741 patients aged 18 to 65. Chiropractic beat hospital outpatient management by roughly seven Oswestry points, concentrated in chronic or severe pain (Meade 1990).
- 1991 to 1992. The RAND R-4025 series under Paul G. Shekelle rated 1,570 indications for spinal manipulation on a nine-point scale, first by a multidisciplinary panel of nine, then by an all-chiropractic panel (Shekelle 1991).
- 1992. The review in the Annals of Internal Medicine covered 58 articles including 25 controlled trials. The probability of recovery at three weeks favored manipulation by 0.17 in acute uncomplicated low back pain and by 0.098 with sciatic nerve irritation (Shekelle 1992).
- August 1993. The Manga Report, 104 pages funded by the Ontario Ministry of Health, closed with ten structural recommendations, including full OHIP insurance of chiropractic services and hospital privileges (Manga 1993).
- December 1994. AHCPR Guideline 14 under Stanley J. Bigos recommended manipulation in the first month of acute low back symptoms without radiculopathy, and instructed clinicians to reassess after four weeks without improvement (Bigos 1994).
- 1995. The RAND pilot judged the same practice records 38 percent appropriate by multidisciplinary criteria and 74 percent appropriate by all-chiropractic criteria. The two panels agreed on 48 percent of cases (Shekelle 1995).
- 1998. The July field study rated 859 documented cases: 46 percent appropriate, 25 percent uncertain, 29 percent inappropriate (Shekelle 1998). In February, Manga and Angus projected direct savings of as much as 770 million dollars under enhanced OHIP coverage, conditional on a proposed policy change costing 200 million dollars a year by 2000 (Manga and Angus 1998).
Questions people ask
Did the RAND reports prove that chiropractic works?
No, and RAND never said so. The R-4025 series ran expert appropriateness panels, which rate described clinical scenarios rather than generate patient outcomes. Evidence of effect came from the separate 1992 review in the Annals of Internal Medicine, which found a difference in probability of recovery at three weeks of 0.17 for acute uncomplicated low back pain and 0.098 where sciatic nerve irritation was present. Real, modest, and confined to a narrow clinical window.
Where does the 770 million dollar savings figure actually come from?
Not from the 1993 Manga Report, despite decades of citation saying otherwise. It comes from a February 1998 paper by Pran Manga and Doug Angus on enhanced chiropractic coverage under OHIP (Manga and Angus 1998), which projected direct savings of as much as 770 million dollars, very likely 548 million and at least 380 million, conditional on a proposed policy change costing an extra 200 million dollars a year by 2000. The 1993 report speaks only of savings in the hundreds of millions annually.
Why did the two RAND panels disagree so sharply?
Because appropriateness ratings measure the raters as much as the procedure. Applied to the same practice records in the 1995 pilot study, the multidisciplinary criteria called 38 percent of care appropriate while the all-chiropractic criteria called 74 percent appropriate, and the two panels agreed on 48 percent of cases. Where controlled trials are sparse, expert panels fill the gaps with clinical judgment, and judgment follows the model of the problem each discipline already holds.
What did the 1994 AHCPR guideline actually recommend?
Manipulation during the first month of acute low back symptoms without radiculopathy, described as safe and effective. Beyond one month the guideline called efficacy unproven, and it instructed clinicians to stop and reassess if four weeks of manipulation produced no symptomatic and functional improvement. It is a narrow endorsement stated precisely, and quoting the first clause without the second misrepresents a public document that is easy to check.
How does the Unified Model of Tone read the Evidence Era?
As a prediction about trial design. A trial that delivers one predetermined input to every participant pools the people it matched with the people it missed, so the mean understates what a matched input can do. The model specifies the remedy in advance: record a tone measure before any input, derive a leverage point for each person from it, then randomize matched against mismatched delivery. It predicts a substantially larger effect in the matched arm, and that separation between the two arms is the finding that confirms it.