← Back to blog

Clinicians: Therapy Outcome Measures, 18 New 2026 Scales, ROOT Checks

September 10, 2026
Clinicians: Therapy Outcome Measures, 18 New 2026 Scales, ROOT Checks

Therapy Outcome Measures (TOMs) are a clinician-rated framework that captures change across four domains: impairment, activity, participation and wellbeing. Their value lies in what they refuse to ignore: participation and wellbeing sit alongside physical or cognitive change, not below it. The 2025 Therapy Outcome Measure Handbook and the RCSLT Online Outcome Tool (ROOT) are now the two reference points any service using TOMs needs to check.


TL;DR:

  • Different domains, including wellbeing and participation, are scored on a detailed 11-point scale to capture nuanced patient progress.
  • Services should validate support for adapted scales and non-standard domains in their reporting tools before full implementation.
  • Inter-rater reliability practices are essential every two to three months, and licensing for descriptor wording must be carefully managed.
  • Tailoring scale choice to the patient's specific therapy goal and verifying system support effectively optimize outcome measurement adoption.
  • Digital tools like emotional check-ins can supplement TOMs by providing high-frequency wellbeing data between formal assessments.

Mysafetherapy
Support Beyond Formal Assessments
MySafeTherapy connects UK adults with confidential online therapy, flexible formats, and tools supporting ongoing mental health management.
Visit MySafeTherapy

Table of Contents

What do therapy outcome measures actually assess?

TOMs assess change across four distinct domains, and each one asks a different clinical question.

  • Impairment looks at the underlying condition itself, such as reduced muscle strength or a language deficit.
  • Activity asks whether the patient can perform a task, such as walking to the shop or holding a conversation.
  • Participation captures whether the patient can take part in life roles, such as returning to work or attending social events.
  • Wellbeing records the patient's own sense of distress, confidence or adjustment to their condition.

Each domain uses an 11-point ordinal scale running from 0 to 5 in 0.5 increments, so a score of 3.5 sits between "moderate" and "mild" difficulty rather than forcing a clinician into a coarser judgement. Lower scores indicate greater severity; a 0 usually reflects the most profound level of difficulty a scale describes, while 5 reflects no detectable problem in that domain. Completing all four domains typically takes under five minutes, and it's the clinician, not the patient, who assigns the rating based on assessment and clinical judgement.

When and how should TOMs be administered?

Getting the timing right matters as much as getting the score right. Most services follow a simple structure:

  1. Baseline. Score all four domains at the first contact, before any intervention begins, to establish the starting point against which everything else is judged.
  2. Interim. Re-score at a clinically meaningful midpoint, particularly for longer episodes of care, so drift or plateau shows up before discharge.
  3. Discharge. Score again at the end of the episode to quantify change and feed service-level audit data.
  4. Mid-treatment starts. Where a patient joins a service partway through an established episode (a transfer between teams, for example), flag this using an episode marker rather than treating the score as a fresh baseline, since that distorts change calculations at aggregate level.
  5. Variable presentations. When performance fluctuates within a session or across the week, rate the prevalent performance rather than the best or worst moment observed, and use NA where wellbeing genuinely cannot be assessed rather than forcing a score.

Retrospective scoring is acceptable where clinical notes are detailed enough to support it, but it should be the exception, not the routine. Wherever possible, discuss the score with the patient or their carer before finalising it. That conversation often surfaces a discrepancy between what the clinician observes in clinic and what the patient experiences day to day, which is exactly the kind of gap TOMs are designed to catch.

Adapted scales and the 2025 updates: what to check in your service

Adapted scales exist because a generic four-domain rating doesn't capture the nuance of, say, dysphagia recovery in the same way it captures aphasia recovery. There are now well over 60 adapted scales in circulation, covering conditions from voice disorders to autism, and the library keeps growing.

The 2025 Therapy Outcome Measure Handbook added 18 new scales and updated 7 existing ones, a meaningful expansion for services that had been working around gaps in the previous list. Several of these newer scales carry additional "non-standard" domains beyond the usual four, built to capture something specific to that condition.

That extra granularity comes with a practical catch worth checking before you roll a new scale out to your team:

  • Not every non-standard domain imports cleanly into ROOT reporting yet; some functionality is still in development.
  • A scale that looks perfect on paper may generate incomplete reports until ROOT catches up.
  • Services should check which adapted scales their ROOT or electronic patient record (EPR) actually supports before committing a caseload to one.
  • Where a scale's reporting behaviour is unclear, contact ROOT support directly at root@rcslt.org rather than guessing.

If a patient's presentation doesn't match any adapted scale, use the core TOM. Waiting for a bespoke scale, or forcing a diagnosis-specific one onto a patient it wasn't built for, does more harm than defaulting to the four core domains.

Reliability, licensing and reporting implications (practical checks)

Inter-rater reliability is the quiet failure point in most TOMs implementations. Two clinicians can watch the same session and land on different scores for the same domain, and that drift compounds across a service if nobody checks for it.

Statistic to note: guidance recommends inter-rater reliability practice sessions every two to three months to keep scoring consistent across a team. That's a modest commitment for the confidence it buys when comparing scores across clinicians or feeding aggregate data into service audits.

Illustration of inter-rater scoring calibration

Licensing is the other detail that catches services out. Reproducing TOMs descriptors, the actual wording of each 0.5 increment, inside an EPR requires a licence from J&R Press. Copying descriptor text into a local template without checking this is a common and avoidable compliance gap.

On the IT side, ROOT's import behaviour for non-standard domains varies by scale, and some SNOMED CT code mappings are still being requested and built out.

Pro Tip: Before you commit a whole caseload to a newly adapted scale, run one patient through it end to end, score, save, generate a report, so you catch any ROOT import quirks before they affect a full audit cycle.

How to choose and implement the right TOMs scale in your service

Rolling out TOMs well is less about the scales themselves and more about the sequence you introduce them in.

  1. Match the scale to the clinical focus. Start with the therapy goal, not the diagnosis label, and pick the adapted scale that reflects it. If nothing fits, use the core TOM rather than delaying treatment while you search for a perfect match.
  2. Verify ROOT and EPR support first. Check whether the scale you want is fully supported and whether it carries extra domains that could affect how reports render.
  3. Build reliability checks into the rota. Set inter-rater practice as a fixed calendar item, not an optional extra, and document your team's NA rules and episode-flag conventions so new starters follow the same logic.
  4. Plan aggregation from day one. Decide how scores will roll up into service-level audit and benchmarking before you've collected six months of data you can't easily compare.

Pro Tip: Assign one person on the team as the "ROOT champion" for non-standard domains. It saves the whole service from independently discovering the same reporting quirk three separate times.

Resources, training and official materials to consult

Start with the primary sources rather than secondary summaries, since scale details change between handbook editions.

  • The RCSLT ROOT scales page lists which adapted scales are live and which carry non-standard domains.
  • The detailed TOM scale list PDF flags scales still in development and licensing requirements.
  • ROOT's welcome page covers reporting features used for audit and benchmarking.
  • Contact root@rcslt.org for technical support on non-standard domains or reporting queries specific to your service.
  • Set inter-rater training on a fixed two to three month cycle rather than an annual one-off session.

Patient and clinician feedback on the use of TOMs and their impact on therapy

Clinicians who've used TOMs across a full caseload tend to describe the same shift: the wellbeing and participation domains change how a discharge conversation goes. A patient with good physical recovery but a wellbeing score stuck at 2 is a very different discharge decision to one where both scores have moved together. Without that domain sitting alongside impairment data, that gap is easy to miss.

The friction clinicians report is almost always about time and consistency, not the concept itself. Rating four domains per patient per contact adds up across a busy caseload, and teams without a fixed inter-rater schedule notice scores drifting between colleagues within a few months. Services that treat the reliability checks as a rota item rather than an afterthought tend to report far less of this friction.

Patients rarely see their own TOMs score directly, since it's a clinician-rated tool, but the domains shape the conversation they do have. Asking someone to reflect on how confident they feel about a work meeting, rather than simply how their speech sounds, tends to surface concerns a purely impairment-focused conversation misses. That's the practical argument for keeping participation goals explicit rather than folding them into a general sense of "doing better".

The consistent feedback theme across services adopting the 2025 handbook updates is cautious enthusiasm: the new adapted scales fit specific presentations better, but the ROOT reporting catch-up has meant some early frustration for teams who rolled them out before checking import behaviour.

Examples of case studies or scenarios illustrating effective TOMs application

Consider a patient recovering from a stroke with residual dysarthria. At baseline, impairment sits at 2 (moderate speech difficulty), activity at 2.5, participation at 1.5, and wellbeing at 2. The impairment score improves steadily over eight weeks of therapy to 3.5. But participation stays flat at 1.5 because the patient has stopped attending their weekly community group, embarrassed by how they now sound. Without the participation domain, that case reads as a clear success. With it, the clinician has a specific, actionable gap to address, perhaps a graded return to the group with the patient's consent, rather than closing the episode on impairment data alone.

A second scenario: a service transitions to an adapted voice-disorder scale from the 2025 handbook update mid quarter. The scale includes a non-standard domain tracking vocal fatigue across the working day, something the core TOM doesn't capture at all. The clinical team flags this domain to ROOT support before rollout and discovers it isn't yet fully reportable at aggregate level. Rather than abandoning the scale, they continue recording it manually for six months alongside standard ROOT output, giving them both the richer patient-level data and a paper trail ready for when full import support arrives.

Both examples point to the same lesson: the extra domains and adapted scales earn their place only when a service checks, rather than assumes, how the data will actually surface later.

Comparison of Therapy Outcome Measures with other assessment tools or methods

Patient-reported outcome measures (PROMs) ask the patient to self-rate their symptoms or function directly, often through a standardised questionnaire. TOMs differ in one structural way: a trained clinician assigns the score, drawing on observation and assessment rather than self-report alone. That distinction cuts both ways. TOMs benefit from clinical judgement filtering out reporting bias, but they lose the direct patient voice that a well-designed PROM captures.

Impairment-only tools, the kind that measure grip strength, range of motion or a specific cognitive test score in isolation, give precise, reproducible numbers but say nothing about whether that improvement changes how someone lives their week. TOMs were built specifically to close that gap by carrying participation and wellbeing alongside impairment in a single, comparable framework, which is why services choose them for cross-disciplinary audit rather than relying on impairment metrics alone.

Generic quality-of-life scales cover breadth that TOMs don't attempt: they're often disease-agnostic and comparable across entirely different patient populations. TOMs sacrifice that breadth for clinical specificity and speed, a five-minute clinician rating rather than a lengthy patient questionnaire. The realistic answer for most services isn't choosing one over the other. It's using TOMs for routine clinician-rated tracking and pairing them with a PROM or quality-of-life measure where patient-reported nuance genuinely changes the clinical picture.

Comparison of TOMs and assessment tools

Clinical perspective: integrating TOMs into everyday practice

What TOMs get right is forcing wellbeing into the same conversation as impairment, rather than treating it as an afterthought. That only holds up with consistent training behind it. Digital supports that let patients reflect between formal ratings can add real texture to that picture, provided nobody mistakes them for a replacement.

— MySafeTherapy

An alternative option: how MySafeTherapy complements outcome measurement

Formal TOMs ratings happen at fixed points, baseline, interim, discharge, which leaves real gaps in what's actually happening to a patient's wellbeing week to week. Mood tracking and emotional check-ins can give patients a way to log how they're feeling between those formal ratings, without replacing the clinician-rated assessment itself.

Mysafetherapy

Paired with one-to-one therapist sessions delivered by video, chat, phone or avatar, this kind of high-frequency monitoring can surface a dip in wellbeing long before the next scheduled TOMs interim score would catch it. It's a complement to formal outcome measurement, not a substitute for it: the structured clinical rating still does the heavy lifting for audit and benchmarking, while the day-to-day check-ins fill in what happens between appointments.

Clinicians and service leads curious about how this kind of digital support sits alongside structured outcome tracking can explore the emotional check-in tool directly, or look at how therapy sessions are structured for patients who need more frequent contact between formal review points.

Sources