Practice note
Choosing Outcome Measures a Commissioner Will Actually Accept
A guide to picking loneliness and connection outcome measures that survive commissioner scrutiny, drawn from how social prescribing evaluations have approached the same problem.
Institute for Social Connection

A commissioner reading your evaluation plan is not asking whether your programme feels valuable. They are asking whether the outcome you report will hold up next to the other seventeen bids on their desk, and whether it maps onto something they can defend to their own finance committee. Get the outcome measure wrong and everything else in the report — however well the programme actually ran — becomes noise.
Social prescribing programmes have been through this problem publicly, because the referral pathway forces the question early: a GP or link worker has to justify the spend to a clinical commissioning group, and “people said they felt better” does not survive that conversation. The pattern that has emerged from published evaluations is instructive for anyone choosing outcome measures for a connection-focused programme, whether or not social prescribing is the delivery model.
The three things commissioners actually want to see
Strip away the polite framing and a commissioner wants three things from an outcome measure, in this order.
- A number that moved, not a description of activity. Attendance figures and satisfaction scores describe delivery, not effect.
- A number that means something outside your programme. If nobody but you uses the instrument, nobody but you can judge whether the change is large.
- A number that plausibly connects to the cost line they are worried about — usually GP attendance, A&E use, staff sickness absence, or a budget line with a name attached to it.
Almost every outcome measure fails on one of these. Custom-built “connectedness” surveys fail on the second. Loneliness scores fail on the third unless you can show the chain to a cost outcome, which the evidence mostly cannot yet do convincingly. Service-use data fails on the first if your sample is too small to show a credible movement.
What the social prescribing evaluations settled on, and why
A 2021 systematic review of social prescribing and loneliness found that all nine included studies reported positive individual-level impacts, and three of those went further and reported reductions in GP, emergency, social worker, or inpatient service use. That third group is doing something the other six are not: connecting a self-report change to a service-use change, which is the chain a commissioner can actually take upstairs.
A separate systematic review the same year, looking at wellbeing outcomes more broadly, reported self-esteem and self-confidence as recurring benefits — but flagged limited trial evidence and heterogeneity across programmes as a persistent weakness. That is a polite way of saying the outcome measures varied so much between programmes that pooling them into anything comparable was difficult. If your evaluation uses a measure nobody else uses, you inherit that same problem: your result becomes an island.
A 2022 qualitative meta-synthesis of how people on social prescribing pathways describe the benefit themselves found something worth building a measurement strategy around: participants describe the gain as extending past social contact itself, into restored meaningful participation and a sense of purpose. Structured, purposeful activity reads as more effective in these accounts than unstructured social contact. That is a finding about what to measure, not just what happened — it argues for outcome measures that capture participation and role, not only frequency of social contact.
Put together, the pattern across these reviews points to a hierarchy, not a single answer.
| Outcome measure | What it captures | Commissioner acceptance | Where it fails |
|---|---|---|---|
| A validated loneliness scale (e.g., the UCLA scale used in AARP’s national survey) | Subjective loneliness, comparable across studies | Moderate — recognised in the literature, rarely tied to budget | No direct cost linkage; commissioners often ask “so what happens to our spend” |
| Service-use change (GP contacts, A&E visits, inpatient episodes) | Downstream health system cost | High, when the sample is large enough | Needs a big enough cohort and a defensible before/after window; small pilots usually can’t show it |
| Custom wellbeing or connectedness survey | Whatever the programme designer wanted it to capture | Low | Not comparable to anything; commissioners have seen this fail before |
| Participation/role indicators (volunteering, group attendance sustained over time, self-reported purpose) | Restored social function, not just contact | Growing acceptance, evidence still developing | Weak causal chain to any single hard outcome; best used as a secondary measure |
| Social isolation prevalence (e.g., National Academies’ quarter-of-older-adults figure) | Population-level framing for need, not for programme effect | High for a needs case, wrong tool for an outcomes case | Describes the problem you’re funded to address, not whether your programme fixed it |
The mistake many programmes make is picking from the top of this table because it is the most published, when the commissioner in front of them is actually asking a question that sits in the second or fourth row.
The failure mode: the borrowed instrument problem
The most common error is what might be called the borrowed instrument problem: a programme adopts a validated scale — the UCLA Loneliness Scale, say, which underpins AARP’s national estimate that one in three U.S. adults aged 45 and older report loneliness — because it is credible, without checking whether their programme can plausibly move it within the funded period, or whether their sample size can detect the movement if it happens. The scale is legitimate. The mismatch between scale and programme is not.
A related version shows up with mortality and morbidity framing. Meta-analyses linking social isolation and loneliness to a roughly 26–29% increase in mortality risk are genuinely strong evidence, and the National Academies’ 2020 consensus report is a credible basis for arguing that a quarter of older adults are socially isolated and that health systems should be routinely assessing for it. But none of that evidence was generated by measuring the effect of a specific twelve-week programme. Citing population-level mortality risk in a funding bid as though it will show up in your evaluation is a category error commissioners have started to catch. The Surgeon General’s 2023 advisory makes the population case at national scale — it is exactly the right document for a needs section and the wrong document to imply as an outcome claim.
What this means in practice: decide first whether you are making a needs case or an outcomes case, and use different sources for each. Prevalence and mortality-risk findings — AARP’s national figures, the National Academies’ isolation estimates, the mortality meta-analyses — belong in the section that justifies why the programme should exist. Service-use change and participation indicators belong in the section that reports what it did. Mixing the two in a single measurement plan is the single fastest way to lose a commissioner’s confidence.
Building the measurement plan backward
Start from the commissioner’s cost line, not from the available instruments. If the funder is a clinical commissioning body, service-use data — however hard to collect — outranks a loneliness scale, because it is the number three studies in the 2021 review managed to connect to positive outcomes and it’s the number that maps to the budget being defended. If the funder is a local authority weighing social value against a needs assessment, prevalence framing plus a participation indicator carries more weight. If the funder is philanthropic and interested in individual transformation, the qualitative account of restored purpose and role — the finding from the 2022 meta-synthesis — may be exactly what they want, provided you are honest that it is descriptive, not a magnitude-of-effect claim.
What this does not solve
None of this fixes the underlying evidence gap: robust intervention trials measuring loneliness or isolation outcomes over time, with adequate control groups, remain scarce, and most reviews cited here flag heterogeneity and weak trial design as limits, not footnotes. A well-chosen outcome measure makes your evaluation legible to a commissioner. It does not make the underlying causal evidence stronger than it is, and no measurement strategy will substitute for a properly powered study that this field, on the whole, still lacks.
Sources
- Loneliness and Social Isolation as Risk Factors for Mortality: A Meta-Analytic Review
- Loneliness and Social Connections: A National Survey of Adults 45 and Older
- Social Isolation and Loneliness in Older Adults: Opportunities for the Health Care System
- Can Social Prescribing Foster Individual and Community Well-Being? A Systematic Review of the Evidence
- Understanding Loneliness: A Systematic Review of the Impact of Social Prescribing Initiatives on Loneliness
- Do People Perceive Benefits in the Use of Social Prescribing to Address Loneliness and/or Social Isolation? A Qualitative Meta-Synthesis
- Our Epidemic of Loneliness and Isolation: The U.S. Surgeon General Advisory on the Healing Effects of Social Connection and Community