The project has been green for six months. Every milestone on track, every box checked, nothing in the status report worth a follow-up question.
Then one review it isn't. The date slips, or the number the whole thing was supposed to produce comes in nowhere near plan, and a project nobody was worried about last month is the emergency at the top of the agenda. It never went amber first. It went from fine to broken between one update and the next.
That's the watermelon effect. Green on the outside, red the whole way through, and you only find out when something forces you to cut into it. You might have heard it called the iceberg effect, or the stoplight problem, but the name matters less than how ordinary it is.
We ran a webinar on it recently, co-hosted with the Balanced Scorecard Institute. Alex Lee, our Chief Customer Officer, hosted the session with David Wilsey, CEO of BSI, who has spent twenty years helping organizations work out what to measure and whether they're actually getting anywhere. Neither of them treated this as an exotic failure. David opened by saying that almost every client he works with wrestles with some version of it, and that finding a watermelon in your plan doesn't mean you're bad at this.
Separate the Fear-Driven Watermelons From the Accidental Ones
There are two kinds of watermelons, and they need different fixes.
"I always categorize these as the intentional watermelon effect versus the unintentional. The intentional one is the fear-based one, where people are afraid of sharing bad news. The unintentional one is just that measurement is hard. And there are a lot of unintentional watermelons out there." - David Wilsey
The fear-based version gets all the attention, because it's the one with a villain. The second category is bigger and much less dramatic. People measure what's easy to measure, teams end up on parallel tracks, and nobody sees the gap until it surfaces on its own, which it does late and usually in front of an audience.
His example was a large nonprofit working on outcomes for women and children worldwide. Reams of data on the situation they were trying to change, country by country, and a full set of internal measures on what their own teams did every day.
"The fundamental problem, most of the time, is when our measurements focus on what we do rather than what we're trying to achieve. In this case they had two different teams measuring two different scenarios, and they weren't connecting what we're doing day in and day out with the outcome they were ultimately trying to achieve." - David Wilsey
Quarter after quarter the outcome metrics came back red while the internal reviews stayed full of green, and when a new leader arrived and looked at both, he couldn't work out how they could both be true. That's what brought BSI in, on the assumption that the organization was measuring the wrong things. What they actually had was an alignment problem: the day-to-day work was never mapped to the change it was supposed to produce, so nobody could tell whether it was working.
Nobody in that organization was lying about their performance. They were solving different problems on different floors of the same building.
Notice Who Gets Grilled When a Report Comes In Red
The unintentional version of the fear is worth sitting with, because most leaders assume they've already handled it. Watch what actually happens in the review.
"Susie gets up, gives her presentation, she's in the green, and we say nice job and move on. Then Fred gets up, his report is in the red, and we grill him. Suddenly we're treating Fred like somebody who needs to be managed. Nobody gets fired, so we think we've avoided the punishment issue. We're still creating the fear." - David Wilsey
Alex has watched that play out one target at a time:
"You might be in a culture where you don't get yelled at, but you get put down, or singled out for being in the red. So you end up moving your target a little bit. You keep guiding yourself toward a green." - Alex Lee
Fred learns the lesson either way. His report stays green right up to the point where it can't be held green any longer, which is usually the week the work was supposed to have delivered something.
Push that dynamic far enough and the data stops meaning anything at all. A stranger once told David, on learning what he did for a living, that a balanced scorecard had cost him his job. His employer, a large telecommunications company, ran layoffs every quarter against whether people hit the numbers on their scorecards. Everyone understood the rule, so everyone sandbagged their targets and softened their reporting. The measurement system did exactly what it had been built to do, and by the end nothing in it was worth reading.
Alex ran into the same pattern earlier in his career in retail banking. A California bank had built its sales culture around one internal slogan, "eight is great," meaning eight products per customer household. The target was easy to communicate and close to impossible to hit honestly, so branch staff started opening accounts their customers had never asked for. The sales numbers stayed green the entire time, right up until the scandal broke and the fired employees started explaining why. Setting the goal was never the hard part. What nobody asked was whether hitting that number would produce the outcome the bank actually wanted.
Tell the Difference Between a Milestone and a Measure
We polled the audience during the session, and the most common answer to "what's your biggest barrier to seeing the true status" was that metrics measure activity rather than outcomes. It's the most common cause and the easiest one to miss, because activity metrics are genuinely satisfying. The work is real. Somebody did it.
Sometimes leadership encourages the confusion outright. Alex watched a CEO do it in a single sentence:
"I once worked with an organization where the CEO literally told them: if you get the work done, we've accomplished the strategy. That was his mantra. So everything was always going to be green or yellow, right up until it wasn't." - Alex Lee
In David's classes, activities don't get to count as measures at all.
"Every employee creating an OKR should be able to differentiate between an activity with a milestone and an outcome you're actually measuring. Something you're counting over time that you can trend, versus a milestone you'll be done with on a certain date, yes or no. We drill that home all day, all week in our classes." - David Wilsey
Some of his OKR clients mix the two, and that's workable as long as the balance holds: mostly outcomes, a few activities, and every objective with a milestone attached also carrying a measure of whether the milestone changed anything.
Check What Your Index Score Is Averaging Away
The other source of watermelons has nothing to do with culture. It's arithmetic.
Take the county government David worked with. It reported quality of life to residents and commissioners as a single index, one number rolled up from everything the county measured about what it was like to live there. Six or seven of the underlying measures covered parks and recreation, and one covered the local economy. Weighted that way, the entire economic health of the county was worth about a seventh of the score.
Then a couple of the county's largest employers packed up and left. Jobs went with them and the local economy fell apart. The index barely moved, because the parks were still lovely and the parks were most of the number.
Nothing about the index was inaccurate. It was doing exactly what an index does, which is compress a lot of measures into one and report a direction. What it can't do is tell you that one component has collapsed while the others hold steady, so the score stays green and the collapse stays invisible until somebody goes looking for it.
David's fix isn't to abolish composite scores. It's to demote them: use the index for high-level directional shifts, and put the underlying measures in front of the people who are responsible for doing something about them.
Alex offered a test for when a blended number has stopped deserving your trust:
"Even if the index looks good, you want it tied to an outcome you actually care about. If market sentiment is positive and revenue is going up with it, great. But if all the indicators are positive and you're still not seeing the outcome, that gives you a reason to dig deeper, or it tells you you're blending too many things into one number." - Alex Lee
Any of this only works if somebody in the room is data literate. Averaging averages, normalizing badly, missing the sampling bias sitting in your own numbers, these are ordinary mistakes and they don't announce themselves. The version everyone has met personally is the review score:
"I used to rely on Yelp scores to figure out if things were good. But you don't get the whole distribution. You get people who rate really high and people who rate really low, and nobody rates in the middle. The average gives you an indicator, but the metric you probably want is how many people really loved it, and what they loved. Even when you're blending data, the indicator might not be useful while the core data underneath it still is." - Alex Lee
Data also has to match the altitude of the person reading it. Senior leaders should be looking at high-level outcomes defined in a way that can actually be measured. Someone three levels down solving a scrap problem on a production line needs to know which machines have dull tooling. Handing either one the other's dashboard produces the same green blur.
Spot the Warning Signs Before It Blows Up
More than half the audience told us they'd had a project reported as on track that turned out to be in trouble, and for 44% of them it blew up before anyone caught it. So the practical question is what the early tells look like.
Everything Is Green
An all-green board sounds like the goal, which is exactly why it's the hardest sign to take seriously. Look at what a plan actually contains, though. Some of it has never been attempted before, and much of the rest depends on another team, a supplier, or a market that nobody in the room controls, which makes all of it landing on plan in the same month an unlikely outcome. Unlikely outcomes deserve a second look before anyone accepts them.
The other problem is what an all-green report leaves you with. A status review earns its place by showing you where to spend attention over the next month, and a board with no variance in it gives you nothing to act on. A little yellow is what makes the green believable.
The Plan Is All Milestones and No Measures
You can catch this one upstream, before a single status update comes in. If everything you've written down can be finished rather than moved, you've built a to-do list and called it a strategy.
The Leading Measures Are Green and the Lagging Ones Are Red
When everything at the bottom of the strategy map is healthy and the outcomes at the top are red, something in the logic connecting them is broken. Sometimes that's only a timing problem, since operational measures move faster than the outcomes they feed and the two reconnect after a quarter or so. When the gap persists, the honest conclusion is that the causal logic was wrong: you said these activities would produce those results, they didn't, and it's the strategy that needs to change rather than the reporting.
Every Initiative Is a Band-Aid
The term comes from measurement specialist Stacey Barr. A band-aid initiative treats the symptom that showed up in the meeting rather than whatever is causing it, and organizations reach for them because they're quick to approve and easy to be seen doing: run a training program, ask for two more headcount, move the team under a different leader.
What makes this a warning sign rather than just a weak plan is how neatly they report. Each one has a clear finish line, so it sits green on the board the whole way through and closes out on time, while the problem it was meant to fix stays exactly where it was. When the entire initiative list looks like that, you're holding a plan full of activity with nothing in it capable of moving a number.
The Number Was Never Real
David was reviewing a stalled program with a client when somebody explained why nobody had questioned it sooner: the performance report from the month before had shown it doing well. So David asked the person who prepared that report to pull up the data behind it, expecting to find a figure that had been misread somewhere along the way. There was no figure. The status hadn't come from a system, a spreadsheet, or a measure of any kind: he'd formed an impression that things were going well and typed that onto the slide.
Slides make that easy in a way a live dashboard doesn't, and it usually isn't dishonesty. Somebody has to fill a status box the night before the meeting, they know roughly how the work is going, and what they know goes in. The trouble is that once it's in the deck, nothing marks it out from the figures that came off a system, so it gets quoted in the next meeting and the one after that, and decisions start getting made against it.
Fix the Culture Before the Dashboard
So where does he tell people to start? Not with the measures.
"Performance is about a conversation. It's never about that one metric. It's about getting people in a room, sitting around a table, talking about things." - David Wilsey
He invokes Goodhart's law here, the one that says a metric stops being a good measurement the moment it becomes a target. The point isn't that targets are bad. It's that the metric was only ever a proxy for the objective, useful for starting the discussion about what's actually happening.
Which means the shift most organizations need is in what a red status means when it lands on the table. Treated as a verdict, it gets managed away long before anyone acts on it.
Not everyone gets to set that tone, though. If you're the one reporting up to a manager who expects green, David's advice is to change the shape of the report rather than its color:
"Make sure your performance report is presented in a way that says: I've identified a problem, and here's what we're doing about it. Not just that we're in the red and we're failing, and that's the end of the conversation." - David Wilsey
He also gave the managers demanding green more credit than they usually get. What they're reaching for is discipline, and that instinct is sound, because nobody should be able to say "we're almost there" for six months running and have it go unchallenged. The discipline just has to attach to the decision rather than to the color.
The cross-functional version is the hardest to catch. Work that runs across departments gets reported department by department, so a red sitting inside one team can disappear into a healthy-looking roll-up before anyone outside that team sees it, and the functions depending on that work carry on as though everything is fine. Inconsistent definitions make it worse. His example: two departments reporting the same scrap-per-employee metric, one green and one red, until somebody noticed they were counting employees differently and one of them had included contractors. Correcting the definition made the gap disappear. Treating it as a process problem is what surfaced the definition error in the first place. Handled as a performance problem, the meeting would have been spent watching both departments defend their denominators.
Decide in Advance What Failing Fast Means
"Fail fast" gets used to justify abandoning things after three weeks, which isn't what David means by it.
"We want to share bad news, because most improvement initiatives are not effective. The sooner we get to that bad news, the sooner we can move the resources to something new." - David Wilsey
The version that works has a target set in advance, a genuine attempt at hitting it, and an honest read on whether the number moved. If it didn't, the effort goes somewhere it can. What you're protecting against is both bailing too early and the far more common failure of letting something run all year because stopping it would require admitting it didn't work.
Alex has seen organizations that handle this well, and they tend to treat a red the same way:
"The successful organizations I've seen don't treat everything that's off track as bad and everyone being in trouble. A red means there's some action we have to take, so let's figure out what that action is. Sometimes the action is fixing it. Sometimes it's deciding we're not going to work on that anymore and changing the priority. The reds are opportunities to fix something that isn't working." - Alex Lee
Question the Greens as Hard as You Question the Reds
David's test for a performance report is short. It should describe truthfully what's happening now, prompt a conversation about why the results look the way they do, and leave the room with a view on what to do next. A report that manages none of those is decoration.
Alex ended the session on the line the whole hour reduces to:
"Green doesn't mean good. It just means somebody decided to call it good." - Alex Lee
That points at the smallest useful intervention available to anyone reading this. Every color on a status board is a claim, and every claim has a source: a system, a survey, a spreadsheet, or an impression somebody formed on the way to the meeting. Asking which one takes a few seconds, and it works on both kinds of watermelon, because the fear-driven version and the badly-measured version each come apart the moment somebody has to say where the number came from.
The question does get asked, of course. It gets asked in the post-mortem, once the project has gone red and everyone in the room wants to know how nobody saw it coming. Same question, ten months late.
Somewhere in your plan right now is something reporting green that isn't. Asking about it this month costs a few minutes of an otherwise comfortable meeting.
___________________________________________
This article is based on a webinar co-hosted by Cascade and the Balanced Scorecard Institute, featuring David Wilsey (CEO, Balanced Scorecard Institute) and Alex Lee (Chief Customer Officer, Cascade). You can watch the full recording here.




.avif)
%20(1).avif)
%20(1).avif)
.png)



