The Comfort of a Number

What we gain, and sometimes lose, when complexity becomes measurable

James Hochreutiner

8/11/202614 min read

There is something reassuring about a number. It gives something complicated a boundary, makes comparison possible and, particularly in a large organisation, allows information to travel. A sales leader cannot participate in every customer conversation, an executive cannot personally understand thousands of customer relationships, and the leadership team of a global business cannot directly observe what is happening in every market.

So we measure. Sales has pipeline and activity, operations has SLAs, customer teams have satisfaction scores and NPS, HR has engagement surveys, and transformation programmes have milestones and adoption measures. Trying to manage a complex organisation without any of them would leave us relying on little more than instinct, anecdote and whichever voice happens to carry the most weight in the room.

What interests me is what happens afterwards, when a measure created to help us understand something gradually starts to become the thing we think we are understanding. The distinction can be surprisingly difficult to notice because nothing necessarily looks wrong. The dashboard still works, the calculations are correct and the underlying data may be perfectly accurate. What can change is the meaning we attach to it.

What we see

Much earlier in my career, I worked for a founder who would occasionally appear at my office door late in the evening. I was usually one of the first people in and often one of the last out, partly because I genuinely enjoyed the quiet hours at either end of the day. The early morning ritual of arriving before everyone else, switching things on and having uninterrupted time to think suited me.

By the evening, most people had gone home. Looking across the empty office, she would occasionally make a slightly barbed observation along the lines of, "You would think they had all reached their targets. Where is everyone?"

I understood what she meant and probably didn't think particularly deeply about it at the time. Years later, I find the comment much more interesting because what she could see was perfectly real: the office was empty and I was still working. What neither of us could derive from that observation was how productive everyone had been, how much value they had created that day, or whether my longer hours were producing better outcomes than somebody else's shorter ones. In my case, they reflected at least partly a personal working rhythm that happened to suit me.

Visible behaviours are seductive management information because the things we actually care about are often considerably harder to observe. Presence can begin to stand in for commitment, meetings for sales activity, pipeline for future revenue, and milestones for transformation progress. The proxy is useful precisely because the underlying thing is difficult to see, but usefulness and equivalence are not the same thing.

There is another side to visibility, however. Sometimes the problem is not that we attach too much meaning to what we can see, but that some of the most important things happening inside an organisation are remarkably difficult to see at all.

I was once leading a geographically dispersed team when it became apparent that something wasn't quite right in one of our offices. A relatively new colleague was struggling to become part of an established group. From a distance, it was difficult to understand exactly what was happening. Meetings took place, work continued and, viewed through the normal mechanisms available to a remote leader, nothing adequately explained the tension.

Eventually I decided that I needed to be there in person, and what appeared remotely to be a fairly ordinary interpersonal difficulty looked rather different up close. There were established loyalties and behaviours at play, and the newcomer was being excluded in ways that felt uncomfortably close to the social dynamics of a school playground.

I got the people involved into a room and we had an unusually frank conversation. One of the leaders later joked that they had always considered themselves fairly direct, but apparently I had raised the bar. Things improved for a while, although over time some of the old behaviour returned.

That experience has stayed with me because it challenged two assumptions at once. Being physically present allowed me to see something that the formal mechanisms of remote management had largely hidden, yet my presence also changed the environment I was observing. I could intervene in the local dynamic, but I couldn't become part of it simply by visiting. It left me wondering how much of what leaders believe they are managing is simply what becomes visible when they are looking, and how much survives when they are not.

There is an uncomfortable limitation to leadership hidden in that question. We can influence what happens while we are in the room, but the more meaningful test may be what happens after we leave it. In a geographically dispersed organisation, that reality quickly becomes unavoidable because no leader can observe enough of the organisation directly to construct a complete picture from personal experience.

What the numbers show

The larger and more geographically dispersed an organisation becomes, the more dependent we therefore become on representations of reality. Meetings, reports, dashboards, surveys and KPIs allow something happening elsewhere to become visible to someone who wasn't there, and sales pipeline is one of those representations.

I once inherited a salesperson with an impressive pipeline. There was no shortage of recorded opportunity or activity, but there was one increasingly obvious problem: very little business was actually being won.

My response at the time was what I suspect most sales leaders would have done. I started looking more closely at the opportunities themselves. What was the customer actually trying to achieve? Who were we talking to? Was there a genuine buying process? Could we deliver what was being discussed? Was there a realistic route from the conversation taking place today to somebody eventually signing a contract?

A considerable amount of pipeline did not survive those questions, which created an uncomfortable situation for me as well. I was responsible for the team's number, and by improving the quality of the pipeline I had simultaneously removed a substantial amount of our reported future business. Nothing had happened in the market overnight. We had not lost those opportunities to competitors, nor had a group of customers suddenly changed their minds. The number had fallen because I had changed the standard by which I was prepared to call something an opportunity.

That distinction matters. On one day the dashboard showed a large pipeline and on the next it showed a substantially smaller one, yet the commercial reality outside the company was essentially unchanged. What had improved was our understanding of that reality, even though the conventional performance indicator had moved sharply in the wrong direction. Better information had made the business look worse.

At the time, I saw that largely as good sales discipline, although I see it somewhat differently now. The salesperson had not invented the behaviour that created the pipeline. He had previously been tasked with building it, and the threshold had deliberately been broad. Conversations that could conceivably develop into business were expected to be captured because pipeline was doing more than helping the sales organisation understand what might close. It also contributed to a broader view of the company's future growth potential.

There was another complication. Some of what he was expected to sell was not sufficiently mature, proven or referenceable in his market, so there could be considerable distance between an interesting customer conversation and something that could credibly be turned into a sale. A large pipeline combined with low conversion could therefore support several interpretations. It might indicate weak sales execution, but it could equally be telling us something about qualification standards, proposition maturity, product-market fit or the purpose for which the pipeline had originally been constructed.

He eventually lost his job, and with the benefit of hindsight I am no longer convinced that was entirely fair. That is uncomfortable for me to acknowledge because I participated in the process that led there. I don't think the pipeline should have remained untouched, and I would make the same decision to clean it up today, but I would question the conclusion we subsequently drew about the person who had created it.

He had been operating according to one definition of what belonged in the pipeline, while I arrived and applied another. What looked initially like an individual performance problem was at least partly a problem of what the organisation had asked a number to represent. The data had not necessarily deceived us. We had changed the question we expected it to answer.

There are well-established ideas behind this. Charles Goodhart originally made his observation in the rather different world of monetary policy, but the principle associated with his work has travelled much further: once an indicator becomes an instrument of control, the relationship between the measure and what it originally represented can begin to break down.[1] Donald Campbell approached a related problem from the social sciences, arguing that the more heavily quantitative indicators are used for decision-making, the greater the pressure placed on both the indicator and the behaviour surrounding it.[2]

It is tempting to interpret both ideas simply as warnings about people gaming metrics, and sometimes they undoubtedly do. My experience with that pipeline makes me less comfortable with such a simple explanation. Nobody needed to game anything. The salesperson could follow the instruction he had been given, management could look at the pipeline it had asked him to build, and I could subsequently apply a more rigorous qualification standard. Each decision made sense in isolation, and the problem only became apparent when we expected the resulting number to mean the same thing throughout.

When the number comes from the customer

Pipeline is created inside the organisation, which makes it relatively easy to imagine how incentives might influence it. Customer feedback appears to offer something different because the organisation is no longer measuring itself. We ask the customer, presumably because their answer gives us an external view of what is actually happening.

I have participated in more than one large-scale NPS exercise, and the apparent simplicity of Net Promoter Score can hide an extraordinary amount of work. Before a single response arrives, there can be lengthy discussions about who should receive the survey, which questions should be asked, how many questions customers might tolerate, how frequently it should be sent and how to improve response rates among people who increasingly seem to be surveyed about everything.

Once the results arrive, we are presented with what appears to be a refreshingly simple answer. The customer has spoken and the organisation has a number, but interpretation starts almost immediately.

When the result is good, there can be considerable back-slapping. When it is bad, we become remarkably interested in methodology and context: response rates were low, the timing wasn't ideal, a difficult implementation affected the result, or one unhappy customer distorted the sample. Some of those explanations may be entirely legitimate. What interests me is why we are not always equally curious about those limitations when the score tells us something we wanted to hear.

Research has questioned some of the stronger claims made for NPS. A 2022 study in the Journal of Business Researchcompared NPS with alternative ways of calculating customer mindset metrics and found that a simpler measure focused on the most enthusiastic customers was a better predictor of future sales growth in its analysis. The familiar separation into Promoters, Passives and Detractors did not provide additional predictive insight.[3] That does not make NPS useless, but it does make the confidence sometimes attached to a single number worth examining.

My own doubts came from something much less academic. I remember individual poor survey responses triggering slightly panicked follow-up conversations, only for the customer effectively to say that the issue had happened months earlier, they had been frustrated when they completed the survey and, generally speaking, they were quite happy with the service.

Nothing about the original response needed to be inaccurate. The customer really could have felt that way when answering the question, while the organisation could also be wrong to treat the resulting score as a durable description of the relationship. The measurement captured a real moment; our interpretation quietly extended that moment across a much longer period.

Even the way people use the scale introduces context. In international customer relationships, I have encountered people who are perfectly comfortable using the very top of a rating scale to describe an excellent experience, while others regard a seven or eight as an excellent score precisely because there should always be room for improvement. I would be wary of turning those experiences into national stereotypes, but the broader phenomenon is well established in cross-cultural survey research. Researchers have identified systematic differences in response styles, including tendencies towards more extreme or more moderate responses.[4]

The number we receive can therefore contain more than the underlying sentiment we intended to measure. It can contain timing, circumstance, interpretation and the respondent's own approach to the scale. None of those things makes the response invalid, but each affects what we can reasonably infer from it.

Once the number enters a dashboard, much of that context disappears. A ten is a ten, an eight is an eight and a five-million opportunity is a five-million opportunity. Aggregation requires that simplification, otherwise comparison becomes almost impossible, but the arithmetic can remain flawless while the meaning becomes considerably less precise.

Where we disagree

Perhaps that is why I found a rather different approach to measuring customer relationships so useful. As part of the strategic account-planning process, we introduced a two-way relationship scorecard in which customer and provider independently assessed the same relationship across dimensions including programme value, operational performance, responsiveness, problem resolution, communication, KPIs, continuous improvement, account management, thought leadership and business impact.

It still produced numbers and it certainly wasn't immune from subjectivity. What made it useful was that we weren't particularly interested in pretending otherwise, because some of the most revealing conversations began precisely where the scores did not agree.

If we thought our thought leadership deserved a four while the customer gave us a two, averaging the scores would have produced a perfectly respectable three and removed almost everything worth knowing. The difference created a reason to explore what each side thought thought leadership actually meant, what the customer expected from us and whether the value we believed we were creating was the value they experienced.

Perhaps we believed we were bringing ideas to the customer while they experienced those ideas as little more than incremental improvements. Perhaps we believed we understood their strategic priorities while they still saw us primarily as an operational supplier. Equally, perhaps they thought they were giving us the access and support required to create greater value while we experienced the relationship quite differently.

The score gave both sides something concrete to react to without pretending that either side possessed the objective truth. Instead of eliminating subjectivity, it made the differences in our subjective interpretations visible, and those differences were often more useful than agreement because they exposed assumptions that might otherwise have remained comfortably hidden.

That seems particularly appropriate for B2B relationships because research suggests that relationship quality is itself multidimensional. Trust, commitment, satisfaction and service quality have all been identified as distinct components associated with customer loyalty, with research also distinguishing between a customer's relationship with individual supplier employees and the relationship with the supplier organisation.[5]

A customer can therefore be satisfied with service delivery while having little strategic trust in the supplier, just as another can be frustrated by a particular operational problem while retaining considerable confidence in the overall relationship. Neither position is contradictory. They are different aspects of something considerably more complicated than a single satisfaction score can comfortably contain.

The two-way nature of the exercise added something else because the customer was not simply being asked to judge us while we waited for the verdict. Both sides had to consider the same relationship and then sit together with the differences. Measurement became useful not because it supplied an answer, but because it provided enough structure to make a more difficult conversation possible.

When the conversation changes

One comment from those relationship discussions has stayed with me. A customer described the evolution of our relationship roughly as moving from them pushing us to us challenging them, which struck me because it had very little to do with conventional customer satisfaction and a great deal to do with how the relationship itself had changed.

The conversation had moved beyond whether we were doing what we had been contracted to do. Operational credibility had created the opportunity to talk about improvement, improvement had created room for ideas, and eventually the relationship had developed enough trust for disagreement itself to become valuable.

There is an obvious connection with the Challenger approach to complex B2B sales, which emphasises bringing insight, tailoring it to the customer's circumstances and creating constructive tension that encourages customers to reconsider existing assumptions.[6] What interests me more than the methodology, however, is what has to happen before that challenge becomes welcome.

A supplier cannot simply decide one morning to become more challenging. Constructive tension without sufficient understanding can easily become irritation, particularly when it comes from someone who has not yet demonstrated that they understand the customer's business or earned credibility through delivery. Before a customer values being challenged, there has usually been a long accumulation of interactions through which competence, judgement and trust have been established.

That progression is remarkably difficult to measure because the observable activity does not necessarily tell us what has changed underneath it. We can count the calls and meetings that happened along the way, but the count tells us very little about when a customer begins to trust someone's judgement or becomes comfortable exposing a problem that has not yet been neatly defined.

A customer involving a supplier before a requirement has been formally created may represent far greater relationship progress than another scheduled account meeting. An executive asking for an opinion rather than a proposal tells us something about how they value the relationship. A stakeholder introducing us to another part of the organisation because they believe the conversation itself will be useful can matter more than another opportunity appearing in the CRM.

The amount of observable activity may actually remain unchanged or even decline as the relationship develops, while the quality and consequence of the conversations increase substantially. Fewer conversations with greater candour, earlier access and more consequential stakeholders can represent far more commercial progress than a calendar full of meetings that never move beyond the transactional.

This is where I find myself returning to the pipeline I cleaned up years earlier. I had been right to question whether those opportunities represented credible future business, but the more interesting question was always what was happening in the customer relationships underneath them. Were the conversations becoming more consequential? Were we gaining access to the right people? Were customers beginning to involve us earlier? Were we learning enough about their business to challenge their assumptions rather than simply responding to requirements?

Those questions do not make pipeline or activity measures redundant. They simply remind me that in complex solution sales, the things we can most easily count are not necessarily the things that tell us whether a relationship is moving somewhere. Measurement gives us evidence of activity and progress, but understanding the quality of that progress still requires judgement.

What the number is for

None of this leaves me wanting fewer numbers. Complex organisations could not function without measurement, and instinct is hardly a superior alternative. Human judgement brings its own collection of biases, assumptions and convenient interpretations, and my own experiences are evidence of that as much as anyone else's.

What I have become less certain about is what we expect measurement to do for us.

Pipeline can help us understand future opportunity, NPS can tell us something about customer sentiment, activity can provide visibility into sales effort, SLAs can tell us whether agreed service levels are being achieved, and a relationship scorecard can help us understand how two parties perceive the same relationship. Each gives us information that would be difficult to manage without, but none contains the full reality it is being used to represent.

Perhaps this is why the most useful part of that two-way scorecard was never the number itself. The value appeared when the numbers did not agree and we had to understand why, because disagreement forced both sides to put context back around the measurement.

The same may be true elsewhere. A pipeline that refuses to convert is telling us something, but not necessarily that the salesperson needs to make more calls. An unexpectedly poor customer score deserves attention, but not necessarily panic. A transformation dashboard showing every milestone in green may still tell us surprisingly little about whether people are actually behaving differently. In each case, the measure can direct our attention without being capable of completing the diagnosis.

Sometimes the number confirms what we already understand, sometimes it challenges it, and sometimes it exposes that two people have been looking at the same situation and seeing something entirely different. If we treat the number as the conclusion, that difference becomes a problem to be reconciled. If we treat it as a starting point, the difference can become the most interesting piece of information we have.

The number gives us something to examine, while the difference, anomaly, unexpected movement or disagreement gives us somewhere to look. What we find there still requires context, curiosity and conversation, precisely because those are the things that become harder to preserve as organisational complexity is compressed into something measurable.

Perhaps the danger begins not when we measure too much, but when the measurement becomes sufficiently reassuring that we stop asking what it actually means.

References

[1] Goodhart, C. A. E. (1975). Problems of Monetary Management: The U.K. Experience. Papers in Monetary Economics, Reserve Bank of Australia. Goodhart's original observation concerned the instability of statistical relationships once they are used for control purposes.

[2] Campbell, D. T. (1976). Assessing the Impact of Planned Social Change. Campbell's work describes the tendency of quantitative indicators to become subject to distortion pressures when they are heavily relied upon for social decision-making.

[3] Baehre, S., O'Dwyer, M., O'Malley, L. & Story, V. M. (2022). Customer mindset metrics: A systematic evaluation of the net promoter score (NPS) vs. alternative calculation methods. Journal of Business Research, 149, 353–362.

[4] Van Vaerenbergh, Y. & Thomas, T. D. (2013). Response Styles in Survey Research: A Literature Review of Antecedents, Consequences, and Remedies. International Journal of Public Opinion Research, 25(2), 195–217.

[5] Rauyruen, P. & Miller, K. E. (2007). Relationship quality as a predictor of B2B customer loyalty. Journal of Business Research, 60(1), 21–31.

[6] Dixon, M. & Adamson, B. (2011). The Challenger Sale: Taking Control of the Customer Conversation.Portfolio/Penguin.

Reach out for tailored support across services procurement, external workforce strategies, workflow optimisation and SaaS evaluation.

© 2026. All rights reserved.

JH Workforce Labs
JH Workforce Labs