A peer-reviewed study in a recognised journal has survived a process that most research does not complete. From the formation of a testable question through ethics approval, data collection, statistical analysis, and the scrutiny of independent reviewers, the path to publication is designed to filter out weak evidence before it reaches the reader. Understanding that path — and the places where it can falter — makes it possible to distinguish a preliminary finding from a conclusion built on solid ground.
Before Anyone Collects a Single Data Point

The Research Question Has to Earn Its Place
A research study does not begin with a laboratory or a grant cheque. It begins with a question — one specific enough to test. Chiropractic researchers, like researchers in any health discipline, start by surveying what is already known. They search databases of published studies, looking for gaps: questions that have not been asked, populations that have not been studied, comparisons that have not been made. A researcher might notice, for example, that spinal manipulation has been examined in dozens of trials for chronic low back pain but that its effects on a related condition — say, referred leg pain — remain largely untested.
The question that emerges from this gap must be precise. “Does chiropractic work?” is not a research question. “Does six weeks of spinal manipulation reduce disability scores in adults with chronic non-specific neck pain compared to supervised exercise?” is. That specificity is what makes the question testable. It defines who will be studied, what intervention they will receive, what will be measured, and over what timeframe. Without that precision, there is no study to design — only a topic to have opinions about.
Ethics Approval Is Not a Rubber Stamp
Once a research question and study protocol exist, they must pass through an ethics committee before a single participant is recruited. In New Zealand, health research involving human participants is reviewed by a Health and Disability Ethics Committee, part of a national system administered by the Ministry of Health. The committee is not there to judge whether the research is interesting. It is there to protect the people who will take part.
The review examines several things at once. Are participants fully informed about what the study involves and what risks it carries? Is informed consent genuinely voluntary, with no pressure to participate? Are the potential benefits proportionate to the risks? How will personal health data be stored and protected? If the study involves populations that require additional safeguards — children, people with cognitive impairments, pregnant women — are those safeguards in place?
This process is not fast. An initial submission may come back with requests for clarification or protocol changes. Researchers might need to revise their consent forms, adjust their inclusion criteria, or strengthen their data security arrangements. Several rounds of revision are common. The months spent on ethics approval often surprise people outside research, but the process exists for a reason that most would agree is worth the delay.
Choosing a Study Design That Fits the Question
The form a study takes depends on what it is trying to find out. A randomised controlled trial, where participants are allocated by chance to either a treatment or a control group, is considered the strongest design for testing whether an intervention works. But it is not always the right choice, and in chiropractic research it carries particular complications. Blinding — keeping participants unaware of which group they are in — is straightforward in drug research, where a placebo pill can look identical to the real one. It is considerably harder when the intervention involves a practitioner placing their hands on someone and performing a physical adjustment.
Cohort studies, which follow groups of people over time and compare outcomes between those who received a treatment and those who did not, are suited to questions about long-term effects and real-world practice patterns. Case series — detailed accounts of outcomes in a defined group of patients — can be valuable for documenting an emerging treatment approach or an unusual presentation, though they cannot establish cause and effect.
Before data collection begins, researchers are increasingly expected to register their trial on a public registry such as the Australian New Zealand Clinical Trials Registry. Registration records the study design, outcome measures, and analysis plan in advance. It is a form of pre-commitment: the researchers declare what they will measure and how they will analyse it before they know what the results will be. This makes it much harder to quietly adjust the goalposts after the data come in.
The Long Middle: Recruitment, Treatment, and Data
Finding Participants Is Harder Than It Sounds
Study protocols specify exactly who is eligible to participate, and the criteria can be narrow. A trial investigating spinal manipulation for chronic neck pain might require participants to be adults aged 18 to 65, with non-specific neck pain persisting for at least twelve weeks, who have not had cervical spine surgery and have no diagnosed neurological condition. Each criterion has a methodological reason: age limits control for developmental and degenerative variables, the twelve-week threshold distinguishes chronic from acute pain, and exclusions remove conditions where the intervention might be unsafe or the results difficult to interpret.
Finding enough people who meet these criteria, who live near enough to attend treatment sessions and follow-up appointments, and who are willing to accept that they might be allocated to a control group rather than the treatment group — this is one of the most common bottlenecks in clinical research. Recruitment timelines routinely stretch well beyond projections. A trial planned for eighteen months of recruitment may take three years. The problem is not a shortage of people with neck pain. It is a shortage of people with exactly the right kind of neck pain who are available, willing, and eligible.
Chiropractic trials face an additional wrinkle: participants often have strong treatment preferences about whether they receive manual therapy, and those preferences can affect both recruitment and results.
Controlling for What You Cannot See
In a drug trial, neither the participant nor the clinician administering the treatment needs to know whether the capsule contains the active ingredient or an inert filler. That kind of double blinding is the gold standard for minimising bias, and it is genuinely difficult to achieve in manual therapy research. A chiropractor delivering spinal manipulation knows they are delivering it. A participant receiving a high-velocity thrust to the thoracic spine can usually tell they have received something different from a light touch to the same area.
Researchers have developed sham manipulation protocols — procedures that mimic some aspects of an adjustment without delivering the therapeutic component — but these remain an imperfect solution. Participants in sham groups sometimes guess correctly that they are in the control arm, and that knowledge can influence their reported outcomes. This is not a flaw unique to chiropractic research; it affects physiotherapy, acupuncture, and surgical trials as well. It is a methodological reality that researchers must acknowledge and account for in their analysis.
Randomisation — the process of allocating participants to groups by chance rather than choice — addresses a different kind of bias. It ensures that the groups being compared are as similar as possible at the start of the trial, so that any differences in outcome can more confidently be attributed to the intervention rather than to pre-existing differences between participants. Outcomes are measured using validated instruments: standardised questionnaires, pain scales, and functional assessments that have been tested for reliability and consistency across populations.
When the Numbers Come Back
After the last participant completes their final follow-up visit, the database is locked and the analysis begins. The statistical analysis follows a plan that was specified before data collection started — ideally the same plan filed with the trial registry. This pre-specification matters because data can be interrogated in many ways, and without a pre-committed plan, it is tempting to keep testing until something appears significant.
Statistical significance, expressed as a p-value, is widely used but widely misunderstood. A p-value below 0.05 does not mean there is a 95 percent chance the treatment works. It means that if the treatment had no real effect at all, data this extreme would arise less than 5 percent of the time by chance. It is a statement about the data under a specific assumption, not a probability that the treatment is effective. Confidence intervals provide richer information: they estimate a range within which the true effect is likely to fall, giving a sense of both the direction and the precision of the finding.
A critical distinction that is often lost in reporting is the difference between statistical significance and clinical significance. A treatment might produce a statistically significant reduction in pain scores — meaning the difference is unlikely to be due to chance — while the actual size of that reduction is too small for a patient to notice. Researchers are expected to report all outcomes they pre-specified, including the ones that showed no effect. Selective reporting — highlighting the results that look good and burying the rest — is a recognised form of bias that trial registration is designed to prevent.
The Gauntlet of Peer Review

What Happens After You Hit Submit
A completed study is written up as a manuscript, typically following a structured reporting guideline. Randomised controlled trials use the CONSORT checklist, which specifies what must be reported: how participants were recruited, how randomisation was conducted, how many dropped out and why, what the primary and secondary outcomes were, and how the statistics were handled. Observational studies follow a similar framework called STROBE. These guidelines exist because incomplete reporting makes it impossible for readers to assess the quality of the evidence.
The manuscript is submitted to a journal, and the first person to read it is usually the editor. Many papers are rejected at this stage — a desk rejection — without being sent for review. The editor may judge the paper to be outside the journal scope, methodologically inadequate, or insufficiently novel. If the paper passes this initial screening, the editor selects two or three reviewers: researchers with expertise relevant to the study topic who have agreed to evaluate the work. In single-blind review, which is the most common model, the reviewers know the authors identity but the authors do not know who reviewed their paper.
Reviewers Are Not Looking to Be Impressed
Peer reviewers are asked to evaluate the manuscript on several fronts. Is the research question clearly stated and justified? Is the study design appropriate for that question? Were the methods executed rigorously? Is the statistical analysis sound? Do the conclusions follow from the data, or do they overreach? Is the reporting complete, or are important details missing?
The reviewers write detailed reports, often several pages long, identifying strengths and weaknesses. They may question the choice of control group, challenge the interpretation of a non-significant secondary outcome, or point out that the sample size was too small to detect the effect the authors claim to have found. The process is not adversarial in intent, but it is rigorous by design. Reviewers are looking for what might be wrong, not for reasons to be impressed.
The most common outcome is neither outright acceptance nor outright rejection. Most manuscripts receive a decision of major revisions or minor revisions. Authors are given the reviewer comments and asked to respond — revising the manuscript where they agree with the criticism and providing a reasoned rebuttal where they do not. The revised manuscript goes back to the reviewers, who assess whether their concerns have been addressed. This cycle can repeat more than once. From initial submission to final acceptance, the peer review process for a single paper commonly takes six months to a year.
Peer Review Has Limits Worth Knowing About
Peer review is the best quality filter academic publishing has, but it is not infallible and no one who understands the system would claim otherwise. Reviewers are unpaid volunteers, often reviewing manuscripts alongside their own research, teaching, and clinical responsibilities. They have limited time and cannot replicate the statistical analyses or verify the raw data. Honest errors, and occasionally deliberate ones, do get through.
Publication bias presents a subtler problem. Journals are more likely to accept studies that report positive or novel findings, and researchers are more likely to write up and submit studies that produced statistically significant results. The consequence is that the published literature can overrepresent the effectiveness of treatments, because studies finding no effect are less likely to see print. Trial registration helps — a registered trial that never publishes its results is a visible gap — but it does not fully solve the problem.
Then there are predatory journals: publications that charge authors a fee and provide little or no genuine peer review. They exist across all fields of health research, including chiropractic. Their papers may look superficially similar to those in legitimate journals, complete with official-sounding names and reference lists. For readers trying to evaluate research quality, knowing the difference between an indexed, peer-reviewed journal and a predatory outlet is a genuinely useful skill. Databases such as PubMed index only journals that meet specific quality criteria, which is one practical way to filter.
Not All Published Evidence Carries the Same Weight
From Case Reports to Systematic Reviews
Published research exists on a spectrum. At one end is the case report: a detailed account of what happened to one patient under specific circumstances. A case report of cervical pain resolving after a course of chiropractic adjustment is clinically interesting — it may suggest a hypothesis worth testing — but it cannot establish that the adjustment caused the improvement. The patient might have recovered without any treatment at all.
A case series extends this to a group of patients, providing more data points but still lacking a comparison group. Cohort studies track larger groups over time and can compare those who received an intervention with those who did not, but participants are not randomly assigned, leaving room for confounding variables — differences between groups that might explain the outcomes instead of the treatment.
The randomised controlled trial addresses confounding through random allocation. An RCT comparing spinal manipulation to a structured exercise programme for low back pain, with participants assigned to each group by chance, provides stronger evidence because the randomisation process distributes both known and unknown confounders roughly equally between groups. At the top of the hierarchy sit systematic reviews and meta-analyses, which locate and synthesise the results of all available studies on a specific question, weighing each study according to its methodological quality and combining the data to estimate an overall effect.
Where a Single Study Ends and Consensus Begins
A single study, no matter how well conducted, is a single data point. Results might be specific to the population studied, the setting, the practitioners, or the precise intervention protocol. A trial conducted in one country with one group of clinicians tells us something, but it does not tell us everything.
Replication — repeating a study with a different group of participants, ideally by independent researchers — is how findings gain credibility. When multiple trials, conducted by different teams in different settings, converge on a similar finding, the evidence becomes substantially more persuasive. When they diverge, that divergence is itself informative: it suggests the effect may depend on factors that varied between studies, or that the original finding was less robust than it appeared.
Systematic reviews formalise this process. A Cochrane review, for instance, uses a pre-specified protocol to search for all relevant studies on a question, assess their quality, extract their data, and where possible combine the results statistically in a meta-analysis. The output is not just a summary of what individual studies found; it is a quantitative estimate of the overall direction and magnitude of the evidence. A treatment supported by one positive randomised trial is a promising lead. The same treatment supported by a Cochrane review synthesising data from a dozen trials across multiple countries is a substantially firmer foundation for clinical decision-making.
Reading Published Research with Appropriate Scepticism
Understanding the publication process does not require a statistics degree, but it does reward a few habits of careful reading. When encountering a published study, checking who funded the research is a reasonable starting point. Industry-funded studies are not automatically invalid, but the funding source is context worth having.
Sample size matters. A trial with 20 participants has wide confidence intervals and limited ability to detect real effects; the same trial with 200 participants produces more precise estimates and more reliable conclusions. Readers do not need to calculate statistical power themselves, but noticing whether a study enrolled dozens or hundreds of people provides a rough gauge of how much weight to place on the results.
The conclusion section of a paper deserves particular attention. Researchers sometimes draw conclusions that are more confident than their data support — concluding that a treatment is effective when their results showed only a non-significant trend in that direction. Comparing the results section with the conclusion is a simple but effective check.
Finally, considering where the study was published matters. A paper in a journal indexed in PubMed or the Cochrane Library has met minimum quality thresholds that a paper in an obscure, unindexed journal may not have. None of these checks requires specialist expertise. They are habits of attentive reading — the kind of informed scepticism that strengthens, rather than undermines, trust in evidence.
The machinery behind a published study is largely invisible to the people it is ultimately meant to serve. Knowing that it exists, and knowing roughly how it works, changes the way research findings land. Not with less trust, but with better-calibrated trust — the kind that asks what was measured, by whom, and how confidently the results can be generalised beyond the study that produced them.
4 Comments
The bit about predatory journals is something I wish more people understood. My sister shared a study on Facebook last month that turned out to be from one of those pay-to-publish outlets. Looked totally legit at first glance.
Really clear explanation of the difference between statistical and clinical significance. I work in health policy and even some of my colleagues mix those up.
How long does the ethics approval process usually take in NZ? You mention several rounds of revision but is that like 3 months or more like a year?
I did my masters thesis in exercise science and the recruitment section hit home hard. We planned for 6 months of recruitment and it took nearly 14. The eligibility criteria were tight and people just didnt want to risk getting the control group. Nobody tells you how much of research is basically begging people to sign up.