JUDGMENT DESIGN

The Vendor Was Chosen Before the Evaluation Began.

The evaluation ran for six weeks. Six vendors, a weighted scorecard with fourteen criteria, three rounds of demos, reference calls, a security review, and a cross-functional committee that gave up the better part of four afternoons. The deck that came out the other end was immaculate. It had a heat map. It had a recommendation. It had, on slide nine, a winner separated from the runner-up by 0.3 points on a five-point scale, which everyone in the room understood to be the sound a foregone conclusion makes when it is asked to show its work.

Because the decision had been made in week zero. One executive had walked out of a conference the previous quarter having already chosen the logo he wanted on the transformation slide. Everything that followed was not a search for the right answer. It was the construction of a defensible paper trail for an answer that already existed. The scorecard was not a decision tool. It was an alibi with columns.

The weights get tuned to fit the winner.

Watch how it actually works, because the mechanics are almost admirable. The criteria are drafted before the demos, which feels rigorous. Then the demos happen, and the favorite is strong on integration but weak on price, so the weight on integration quietly climbs and the weight on price quietly falls, and nobody records the adjustment as anything other than refinement. By the time the scores are totaled, the model has been reverse-engineered from its own conclusion. The number at the bottom of the column is real. The process that produced it was theater with good lighting.

This is not fraud. Nobody lied. Every score was defensible in isolation, every meeting happened, every reference was called. That is precisely what makes it durable. An organization that wanted to cheat would be easy to catch. An organization that has learned to launder a made decision through a legitimate-looking process leaves no fingerprints, because there was no single moment where anyone chose to deceive. The deception is distributed across fourteen criteria and four afternoons, and no one holds enough of it to feel responsible for the whole.

Ask why the machine exists at all and you get closer to the real function. It is not there to find the best vendor. It is there to make sure that if the vendor turns out to be wrong, no human being is standing in the open when the blame arrives. The committee decided. The scorecard decided. The process decided. Those are all ways of saying nobody decided, which is the entire point. Accountability has been diluted to the concentration of a homeopathic remedy, and the executive who actually made the call in week zero is now, on paper, just one voice among many who deferred to the model.

Notice the runner-up. The 0.3-point margin is not an accident of scoring; it is a requirement of the genre. A blowout would look rigged, so the process needs a credible loser: a vendor good enough to make the contest appear real and conveniently flawed in exactly the dimension the winner is strong. The runner-up is not a competitor. It is set dressing. It exists so the winner can be described as chosen rather than installed, and it is thanked warmly for its time and shown the door it was always going to be shown.

A scorecard cannot be wrong. That is not a strength. It is the reason it was built.

The cost is not the vendor. It is the reflex.

Sometimes the pre-made decision is the right one. The executive who chose the logo may have better instincts than the committee that ratified it, and often does, which is how he became the executive. So the immediate cost, the wrong vendor, is frequently zero. That is why the practice survives. It looks harmless because the outcome is usually fine.

The real cost accrues somewhere the procurement file will never show. Every time an organization runs this ritual, it teaches its people a lesson more durable than any training: that the way to be safe is to make the decision invisible. That judgment is a liability to be distributed rather than an act to be owned. That the smart move is never to be the name attached to a call, but to be one signature among eleven on a process nobody can pin down. You are not just selecting a vendor. You are running a masterclass, quarterly, in how to never be accountable for anything, and your most ambitious people are the best students in the room.

The problem was never the pre-made decision.

Here is the turn, and it inverts the obvious complaint. The failure is not that a person decided before the process ran. Decisive instinct informed by experience is not a bug in leadership. It is most of the job. The failure is that the organization built an elaborate apparatus to erase the fact that a person decided, and in erasing the person, it erased the one thing that makes judgment improve over time, which is a name attached to a call and a consequence attached to the name. A pre-made decision owned openly by a named leader is a healthy thing. A rigorous-looking decision owned by no one is a slow poison, because it cannot be learned from. You cannot get better at a call you have arranged for no one to have made.

What it looks like when judgment is allowed to exist.

Rebuilding this does not mean abolishing the scorecard. It means demoting it from oracle to instrument. An organization that takes judgment seriously does something that feels dangerous the first time: it lets the executive say, out loud and on the record, that he has a strong prior toward a particular vendor and here is why. Then it points the entire evaluation not at confirming that prior but at trying to break it. The demos become a search for the reasons the favorite is wrong. The references are called to surface the failure modes, not the applause. The committee's job stops being ratification and becomes adversarial pressure-testing of a named person's stated bet.

The difference is total. In the first version, the process exists to hide a decision so no one is accountable. In the second, the process exists to stress a decision so the accountable person makes it better. Same six weeks. Same fourteen criteria. Opposite relationship to the truth. One produces a defensible file. The other produces a decision someone will actually own when the vendor goes sideways in month nine, which means it also produces an organization that gets sharper every time it chooses, instead of one that just gets better at not being caught.

This is the work SSUNDAR does when a leadership team realizes its decision-making has quietly optimized for defensibility instead of quality. Not a better scorecard. A redesign of how a decision is owned, so that the instinct of the person with the most context is put on the table and tested rather than smuggled through a spreadsheet. The organizations that make this shift stop asking who can we blame if this fails and start asking who is making this call and what would change their mind. The ones that do not keep generating immaculate decks and keep wondering why no one ever seems responsible when the immaculate decision falls apart.

The scorecard said 0.3 points. What it meant was that somebody had already decided, and the whole building had agreed to pretend otherwise.

TEST YOUR OWN JUDGMENT

Theory is interesting. Data is better.

Five cascading crises. AI-generated. Your decisions compound. Get your personalized Leadership Architecture Report in under 4 minutes.

Run the Simulation.

WEEKLY INTELLIGENCE

One insight on leadership systems. Every Monday.

No fluff. No spam. Unsubscribe anytime.