What the numbers reveal, what they conceal and what responsible leaders do next
Statistics can change a conversation. They can take an experience that has been dismissed as isolated and show that it belongs to a pattern. They can reveal who is being suspended, who is attaining, who is promoted, who leaves and where public resources are producing results. Used well, numbers create accountability. They allow us to move beyond “we think” and “we feel” towards a clearer account of what is happening.
But a statistic is also an act of compression. It turns complicated lives into categories, counts and averages. That compression is necessary if we want to see scale, compare outcomes and detect inequality. It is dangerous if we forget what was compressed. A suspension figure does not show the conversation that preceded removal from class. An attainment score does not record whether a pupil felt recognised by the curriculum. A programme total does not tell us whether participation changed a life or simply generated activity.
To go beyond the statistics is not to reject evidence. It is to use evidence more intelligently. The number is a doorway, not a destination. It tells us where to look, whose experience requires attention and which assumptions need testing. Then leaders must return to the people, places and systems from which the data came.
This principle has particular meaning for me. I founded an organisation called Bouncing Statistics because I understood that young people are regularly reduced to the numbers attached to them. The name carries a refusal: a young person may appear within statistics on deprivation, exclusion, attainment or crime, but no statistic is a complete account of who they are or what they can become. The task is not to bounce evidence away. It is to prevent a measured outcome from hardening into an identity.
When data becomes destiny
Institutions need categories. Schools track attainment, attendance, behaviour, special educational needs and disadvantage because they must understand pupils and allocate support. Public bodies cannot identify inequality without grouping and comparison. The problem begins when a descriptive category becomes a prediction, and the prediction begins to shape treatment.
A pupil with poor attendance can quickly become “an attendance problem”. A recorded incident can become evidence of character. A postcode can be read as risk. Once the label is established, new information is interpreted through it. Vigilance is read as aggression; withdrawal as lack of interest; improvisation as defiance. The institution no longer measures the pupil’s interaction with a system. It imagines it has measured the pupil.
I have lived the difference between being seen and being classified. At my multicultural primary school in Newtown, Birmingham, teachers combined expectation with care. I was given opportunities to speak and lead, and I achieved strongly. At grammar school, I gained formal access to a prestigious environment but increasingly felt that my cultural background and the realities surrounding me were not understood. My behaviour became more visible while my potential became less visible. The person had not suddenly changed in isolation; the relationship between person and institution had changed.
This is why outcomes should never be interpreted without context. Context does not erase responsibility or lower standards. It helps us understand the mechanisms producing the outcome. If a pupil is late, the school still needs punctuality. It also needs to know whether the cause is indifference, caring responsibilities, housing instability, transport or fear on the journey. Different mechanisms require different responses. The same number can conceal several problems.
The pattern matters
Personal experience can be dismissed as anecdote. Statistics make patterned inequality harder to ignore. England’s education data show why disaggregation is essential. Department for Education figures for 2022/23 recorded 787,000 suspensions in state-funded schools, equivalent to 933 suspensions for every 10,000 pupils. The rate differed considerably between ethnic groups. Black Caribbean pupils had 1,358 suspensions per 10,000 pupils, while pupils from mixed White and Black Caribbean backgrounds had 1,736. White Gypsy or Roma pupils and Traveller of Irish heritage pupils experienced still higher rates.
Those figures do not establish the cause of each disparity. They do establish a question that institutions are responsible for examining. Are pupils exposed to different circumstances? Are similar behaviours interpreted differently? Do curriculum, relationships, special educational needs, poverty and school practices interact? Which schools achieve fairer outcomes for comparable pupils, and what are they doing differently?
Attainment data add another layer. For 2022/23, the average Attainment 8 score across England was 46.3. The reported averages for Black Caribbean boys and girls were 36.0 and 43.8 respectively. Aggregating all Black pupils would obscure important differences between Black African, Black Caribbean and other Black groups; aggregating by ethnicity alone would obscure gender, disadvantage and special educational needs. An average can illuminate and hide at the same time.
Leaders should resist two opposite errors. The first is denial: finding a limitation in the data and using it to avoid the pattern. All evidence has limits, but limitations are a reason for care and further investigation, not automatic dismissal. The second is determinism: treating a group average as a forecast for an individual. Population patterns should guide institutional questions, never cap personal expectation.
What is missing from the spreadsheet
Administrative data record what systems are designed to record. This means silence in the dataset is not evidence that something does not matter. Belonging, trust, humiliation, cultural recognition and the quality of a relationship may be central to an outcome while remaining absent from the standard dashboard.
Consider the entrepreneurial activity I showed at school. I sold sweets and chocolates, managing stock and customers while responding to financial pressure at home. The school record could capture the rule breach and sanction. It had no obvious field for initiative, mental arithmetic, market awareness or the reason money mattered to me. The data was not false. It was radically incomplete.
The same problem appears in programme evaluation. A mentoring project may report sessions delivered, pupils reached and satisfaction scores. These measures are useful for checking activity and reach. They do not tell us whether the young person developed a trusted relationship, returned to learning, avoided exclusion or gained a stronger sense of agency. Nor do they reveal whether the school changed the environment to which that young person returned.
What is easiest to count can therefore become what receives attention. This is the measurement trap. Organisations optimise visible indicators while the underlying purpose recedes. Schools may focus on moving a headline number; providers may focus on volumes required by a contract; funders may prefer immediate outputs because long-term outcomes are harder to attribute. Everyone can meet a measure while the young person’s experience remains unchanged.
Responsible measurement begins with purpose. What human or social outcome are we trying to produce? What would meaningful change look like from the participant’s perspective? Which indicators offer credible evidence of that change, and what might they miss? The dashboard should serve the mission, not replace it.
Stories are evidence, but not the whole evidence
If statistics compress, stories expand. A lived account can reveal sequence, meaning and mechanism. It can show how repeated small interactions accumulate, why a service feels inaccessible or why an apparently reasonable policy creates harm. Stories allow leaders to encounter consequences that a percentage cannot convey.
But a story carries different limits. One vivid account cannot establish prevalence. The most available speaker may not be representative. Organisations may select the stories that confirm their existing narrative and ignore quieter or contradictory experiences. A success story can become marketing material while the conditions that made success exceptional remain untouched.
The answer is not to choose between data and stories. It is to make them interrogate each other. If exclusion data reveal a disparity, speak with pupils, families and staff to explore the processes underneath it. If interviews reveal a repeated experience, examine administrative records to estimate its distribution. If sources disagree, treat that disagreement as information. Perhaps the measure is poor, participation is selective or an average is hiding local variation.
Lived-experience work must also be ethical. Asking people to disclose difficulty is not automatically empowering. Participants should understand why their experience is being gathered, how it will be used and what influence it can have. Organisations should offer support, protect privacy, compensate substantial expertise where appropriate and return with an account of what changed. Extracting testimony without transferring influence is another form of taking.
The difference between activity and impact
Organisations often describe what they did when they are asked what changed. They trained 200 staff, mentored 100 pupils, held 20 listening sessions or created a new policy. These are outputs. They establish that activity occurred, not that the intended outcome followed.
Impact requires a plausible causal account. What problem was the intervention designed to address? Through which mechanism should the activity produce change? What assumptions need to hold? What other factors could explain the result? HM Treasury’s Magenta Book advises organisations to make this theory of change explicit and plan evaluation from the design stage. The Education Endowment Foundation similarly emphasises implementation as a process rather than a single event.
This discipline matters in alternative provision and inclusion work. An Inclusion Hub may reduce immediate pressure, improve engagement and create a route back into mainstream education. But a favourable outcome cannot be assumed from attendance alone. Leaders should examine the quality and consistency of delivery, pupil experience, academic progress, reintegration and longer-term outcomes. They should also ask whether the intervention changes mainstream practice or becomes a room into which the institution exports difficulty.
Attribution must be proportionate. Not every community programme can run an experimental trial, and demanding an unrealistic standard of proof can privilege large organisations with analytical resources. Smaller providers can still establish baselines, record reach and delivery, collect outcomes over time, use credible comparisons where possible and be explicit about uncertainty. Honest contribution is more valuable than an inflated claim of transformation.
Reading numbers structurally
Statistics are often presented as facts about people when they may be facts about systems. A high exclusion rate can be described as the behaviour of a pupil group. It can also be read as the outcome of interactions among behaviour, need, teacher judgement, school policy, resource constraints and social conditions. The first reading locates the problem in the child. The second creates multiple points for action.
This structural reading does not deny agency. Pupils make choices, and harmful behaviour has consequences. It asks whether institutions apply judgement consistently, whether early support is available and whether people have credible routes back after failure. It also asks who receives the benefit of ambiguity. One pupil’s mistake may be treated as an exception; another’s as confirmation of a category.
Data should therefore be examined across the journey. Who receives early warnings, support, sanctions, managed moves, alternative provision and permanent exclusion? How long do processes take? Which staff or settings produce markedly different outcomes? What happens after intervention? A final statistic is the residue of many earlier decisions. If leaders measure only the endpoint, they miss where change is possible.
Place matters too. National averages can conceal differences across local authorities, neighbourhoods and schools. A local team should compare itself with relevant contexts rather than using a national trend to dismiss a local problem. Equally, a highly visible local incident should not be presented as proof of a national pattern without supporting evidence. Scale must match the claim.
Measure with people, not only about them
The people represented in data should have a role in deciding what success means. Co-design can improve the relevance of measures, identify unintended consequences and reveal barriers that professionals overlook. This does not mean every preference becomes a target or technical judgement is abandoned. It means those who live with the consequences are treated as sources of knowledge.
For young people, this might involve asking what changed in their relationship with learning, which adult they trust, whether they feel able to recover from mistakes and what helped or blocked progress. Their accounts can be connected to attendance, attainment and exclusion data rather than placed in a separate “voice” section. Families and frontline staff contribute different perspectives. None should be treated as complete on its own.
Participation should also extend to interpretation. Organisations often collect feedback and analyse it behind closed doors. Returning emerging findings to participants can test whether conclusions recognise their reality. Where leaders choose a different course, they should explain the trade-off. Voice becomes meaningful when it enters the reasoning, not merely the appendix.
A better evidence culture
Going beyond the statistics requires organisational habits. First, separate observation from interpretation. “Suspensions increased” is an observation; “behaviour worsened” is one possible explanation. Second, disaggregate data while protecting privacy. Third, combine quantitative patterns with qualitative inquiry. Fourth, record what the measure excludes. Fifth, connect outputs to outcomes through a theory of change. Sixth, create review points at which evidence can stop or reshape activity.
Leaders must make it safe to report inconvenient findings. If funding, reputation or personal identity depends on success, teams will unconsciously polish results. A learning culture rewards the discovery of a weak assumption before it becomes an expensive failure. It treats a programme that needs adaptation as information, not betrayal.
Language matters. Avoid describing a community as “hard to reach” before examining whether the organisation is difficult to access. Avoid using “data-driven” as though numbers drive decisions without human values. Evidence informs judgement; it does not eliminate it. Leaders still decide which outcomes matter, how risk is distributed and what level of inequality is unacceptable.
Beyond prediction, towards possibility
The deepest purpose of measurement is not to predict who will fail. It is to identify where systems must improve. Statistics should enlarge institutional responsibility, not narrow individual possibility.
For a young person, being recognised within a troubling pattern can unlock support and accountability. Being reduced to that pattern can close the future. The distinction depends on how adults read the evidence. Do they ask, “What is wrong with this child?” or “What has happened, what strengths are present and what must change around them?”
Bouncing Statistics was built in that space between pattern and person. Our work with schools, families, public agencies and young people begins from the evidence that outcomes are unequal. It continues beyond the number because intervention requires relationship, cultural understanding, professional skill and reflection. The statistic establishes urgency; the person establishes purpose.
We should count what matters, improve what we count and remain honest about what no count can contain. We should use averages to expose systems and refuse to use them as ceilings. We should measure impact without allowing the demand for measurement to strip work of care.
Beyond the statistics lies neither sentiment nor guesswork. It lies a fuller form of evidence: numbers with context, stories with scrutiny, action with evaluation and leadership with accountability. That is how data becomes more than a description of inequality. It becomes part of changing it.
References
Department for Education (2025), Suspensions and Permanent Exclusions in England: 2023 to 2024.
Department for Education (2024), GCSE Results (Attainment 8): 2022 to 2023, Ethnicity Facts and Figures.
Education Endowment Foundation (2024), A School’s Guide to Implementation.
HM Treasury and Evaluation Task Force (2026), The Magenta Book: Central Government Guidance on Evaluation.
Strand, S. (2021), Ethnic, Socio-economic and Sex Inequalities in Educational Achievement at Age 16, Commission on Race and Ethnic Disparities supporting research.
