Eval harness
Inside our eval harness: How we know the advice isn’t generic
Advice that would have been the same for anyone is worthless, and reading it will not tell you. Generic advice is fluent, specific-sounding and confident. So we stopped reading and started measuring. This is the test rig we built, and what it found.
10 min read
The short version
- The same feature scored 75 for one business and 7 for another. Identical words, different company research. Generic advice cannot do that.
- Asked about an idea that is bad for everyone, it said no to all of them. Every answer landed within five points. It disagrees for a reason, not at random.
- The signal is now 3.3× the noise, up from 1.6×. A gap between two businesses used to be a maybe. Now it is a fact.
- Update, 13 September 2026: the “this can’t be built” alarm no longer flickers. It fires for hard facts, like a law, and stays quiet on matters of opinion.
Spongeware answers one question: is this feature worth your team’s next six weeks? It researches your company once, then argues both sides of every feature idea you submit and returns a recommendation with the reasoning attached.
All of which is worth nothing if the answer would have been the same for anybody. That is the failure mode we care about most, and it is close to invisible from the inside: a report written for nobody in particular still reads as thoughtful, still cites your industry, still sounds like it was written about you. You cannot catch it by rereading the output. You can only catch it by asking the same question on behalf of businesses that should get different answers, and checking that they do.
That is what an eval harness is for. It runs the product against a fixed set of cases and measures what comes back, so quality becomes a number that moves rather than a feeling. Ours runs on every change. In September 2026 it ran 158 evaluations across 53 combinations of feature and business, three repeats each. Here is what it found.
Test one: does the business change the answer?
We wrote one feature description: a compliance reporting tool, the kind enterprise buyers demand before they sign. Then we submitted that identical text on behalf of five businesses with almost nothing in common. Every evaluation ends in a score out of 100, and only the company research changed between these runs, so anything that moves is the research doing its job.
Compliance reporting, scored for five businesses
68-point spreadFor a company selling to enterprise security teams, compliance reporting unblocks revenue. For a website builder whose customers are individuals with a landing page, it is six weeks spent on something nobody asked for. Same feature, opposite answers, and nobody had to tell the system which was which. Generic advice cannot produce that chart.
Test two: is it disagreeing for a reason?
Here is where a harness earns its keep, because the first chart on its own is not proof. Something completely broken, returning random numbers, would also produce five different answers, and it would look exactly as convincing. So the second test is the mirror image of the first: an idea that is bad for everyone. We asked each business about giving the product away with no limits at all. A trustworthy adviser should say no to every one of them, and say it about as firmly each time.
Unlimited free tier, scored for four businesses
5-point spreadIt disagrees when the business genuinely changes the answer, and agrees when it doesn’t. Either test alone is easy to pass by accident. Passing both is the whole product.
The number the harness actually optimises
Two measurements decide whether any of this is trustworthy, both in points on the same 0–100 scale. Neither is a score. Each one is a gap.
| Previous run | This run | |
|---|---|---|
| Ask the identical question twice. How far apart are the answers?The noise | 12 | 6 |
| Ask about two different businesses. How far apart?The signal | 20 | 21 |
| How many times bigger the signal is than the noiseSignal divided by noise. Below 2× a gap between two businesses could just be wobble; above 3× it is a real difference. | 1.6× | 3.3× |
Neither number means anything alone. Something that gave every business a wildly different answer would score brilliantly on the second row and still be worthless, because it would just be rolling dice. Something that gave everyone the same answer would score brilliantly on the first and be worthless for the opposite reason. It only works when the signal is much bigger than the noise.
One cycle ago they were close enough that you could not honestly separate a real difference between two businesses from the system wobbling. Now the signal is more than three times the noise: the company effect held steady while the wobble halved.
The reason for that improvement is the best argument for running a harness at all. It showed that two of the factors going into the score were quietly measuring almost the same thing, so one judgement was being counted twice and carrying roughly half the final number. Nobody reading a report would have spotted it. No customer would ever have filed it as a bug. It was only ever going to surface as a number that would not move.
Test three: is it steady where it counts?
An average can hide a lot, so the harness checks the individual runs too, and this is where the picture gets genuinely encouraging. In 35 of the 53 combinations, all three runs landed on the same recommendation on their own, with no averaging required. 70% of individual runs came within five points of their own average.
The clearest cases are the steadiest, which is the behaviour you want. Across the 22 features it judged decisively not worth building, not one changed its answer between runs, with a typical run-to-run gap of 3 points on a 100-point scale. When it tells you to stop, it means it, and it will say the same thing tomorrow.
The features that move between runs are the genuinely close calls, the ones sitting near the line between build and defer. That is the honest place for uncertainty to live. We would rather the system be least certain exactly where a good product manager would be, than manufacture a confident-looking number for a decision that is actually finely balanced. Narrowing that middle band is what the next cycle of work is for.
Update · 13 September 2026
What changed this cycle
Three fixes, each found by the harness rather than by reading reports. The runs behind them are narrower than the one above, so their numbers stand on their own rather than being compared with it.
An alarm that only fires for facts
Before recommending anything, Spongeware checks whether an idea simply cannot work: a law forbids it, the company lacks something it depends on, a contract rules it out. When that check fires, the idea is sent back for rethinking however good it otherwise looks. So it has to be steady.
It wasn’t. We asked about the same hard ideas 10 times each, the way ten different people might describe them, and on 3 of 10 the alarm went off for some descriptions and not others. In the worst case it fired on 6 of 10. Every one of those flickering alarms was an opinion, like “this doesn’t really fit the company”, not a fact. An opinion that half the readings hold is a judgement call, and a judgement call should nudge a score, not flip a switch.
Now the alarm only counts when the reason is a hard fact, the fact is backed by the company’s own record, and the check reaches the same conclusion independently. Opinions about fit still count, through the score, where they belong. Across 30 repeated cases in the latest run, the alarm flickered on 0.
A quiet alarm is only good news if it still rings. So we planted ideas with real blockers: a healthcare software vendor selling patients’ diagnoses to advertisers, and a cloud infrastructure company copying the secrets its customers store into an analytics database. The alarm caught them on 6 of 6 runs. In an earlier, smaller check the second one got through 2 times in 5, because the supporting evidence was quoted with words left out. Tightening that is next.
Must-have features stopped losing points for not being unique
Part of every score asks whether a feature would set the company apart from its competitors. That is the right question for most ideas and the wrong one for features buyers simply require. For a healthcare software vendor, single sign-on and SOC 2 compliance reporting are the price of a sales conversation, and every competitor already has them, so they rated between 10 and 40 out of 100 on uniqueness. That one answer held both below the score of 65 needed for a “build it” recommendation.
We rebalanced the score so uniqueness counts for less, and demand and fit with the company’s direction count for more, then re-scored the recorded runs before changing anything live. Single sign-on moved from 59 to 68, and compliance reporting from 62 to 68. Across 53 feature and business combinations, ideas we had judged worth building got a “build it” recommendation 9 times instead of 7, bad ideas stayed at 1, and the gap between how different businesses scored the same feature widened from 22 to 25 points. The company research is doing more of the work, not less.
We fixed our own answer key
A harness can only grade against the answers someone wrote for it, and ours had mistakes. Checking every “good idea” and “bad idea” label against what each business already offers turned up 7 good-idea labels asking the product to recommend building something the company already had or was already building, such as one-click custom domains for a website builder that already offers them. Every time Spongeware correctly answered “you already have this”, the harness had been scoring it as a miss.
We relabelled those 7, rewrote the reasoning behind 8 more that cited details the research no longer contained, and added 14 new cases no test had seen. On the corrected key, the latest run ranked the better idea above the worse one for the same business in 91% of 46 comparisons. That is the unglamorous lesson of this cycle: a test is only as trustworthy as the answers it grades against, so the answers get audited too.