Inside the Model · 6 September 2026

An 80% probability didn't mean we were more confident than 64%.

tactica. discovered that an 80% protected-market probability wasn't necessarily stronger football conviction than a 64% outright probability. Our first attempt to fix it didn't survive independent review.

tactica. was comparing different betting probabilities as though the bigger number represented stronger football conviction. Fixing that exposed a deeper problem and our first implementation didn't survive independent review.

An 80% probability sounds stronger than a 64% probability.

Usually, if two percentages are measuring the same thing, that's a reasonable conclusion.

Ours weren't.

During an audit of tactica.'s Match Intelligence, we found situations where the model could prefer a betting expression showing around 80% over one showing around 64%, even though both came from exactly the same underlying football opinion.

The bigger number looked more convincing.

It wasn't necessarily more conviction.

It was answering a different question.

That distinction ended up changing how tactica. chooses between related betting markets.

It also produced an implementation that we thought was ready, sent for independent review and had rejected.

That rejection turned out to matter as much as the original finding.

One football opinion. Different ways to bet it.

Before tactica. decides how a football opinion might be expressed as a bet, it has a simpler view of the match:

Home win. Draw. Away win.

Those three probabilities form the underlying football opinion.

Suppose tactica. thinks a team has:

64% chance of winning
19% chance of drawing
17% chance of losing

That 64% is an unconditional win probability.

Now consider Draw No Bet.

With DNB, a draw refunds the stake. So the probability displayed for the winning part of that market is effectively asking:

If this match doesn't finish as a draw, how often does this team win?

Naturally, that percentage can be much higher.

The underlying football opinion hasn't suddenly become more bullish.

We've changed the event we're describing.

Double Chance changes it again: team-or-draw covers two outcomes instead of one.

So tactica. could have three related expressions of the same football view:

Outright win
Draw No Bet
Team or Draw

Each can legitimately have a different probability.

But those percentages are not one universal scale of confidence.

That was our problem.

Arsenal showed it clearly

One of the audit examples involved Arsenal against Chelsea.

tactica.'s unconditional view gave Arsenal a 62.7% probability of winning.

The Arsenal Draw No Bet expression showed 80.5%.

Under the old selection logic, the DNB expression won.

It had the larger transformed probability and the stronger resulting tier:

Arsenal outright — 62.7%, Lean

Arsenal DNB — 80.5%, Elite

At first glance, selecting the 80.5% Elite expression looks obvious.

But Arsenal hadn't somehow gone from a 62.7% chance of winning the match to an 80.5% chance.

The 80.5% figure was conditional on the match not finishing as a draw.

Same football opinion.

Different event.

And, crucially, different price.

At the prices available in that audited example, the outright expression offered an expected net return of roughly +0.041u per 1u, compared with approximately +0.005u for DNB.

The protected market had the bigger probability.

The outright had the stronger payoff at that price.

We had been allowing the former to determine the latter.

Bigger wasn't necessarily better

This gave us a fairly precise question.

Once tactica. has decided what it thinks about the football match, and several eligible betting expressions represent the same underlying opinion, how should it choose between them?

Not by pretending their event probabilities are directly comparable.

Not by saying DNB is always preferable because it protects against the draw.

Not by blindly choosing the highest tier.

And not by rewriting the underlying football prediction.

We eventually settled on a cleaner separation.

First:

What does tactica. think will happen in the football match?

That remains the unconditional Home/Draw/Away opinion.

Then:

Which already-qualified expression represents that opinion most effectively at the available price?

For that second question, tactica. now compares the payoff-correct expected net return of the eligible expressions.

That means price can influence how an opinion is executed.

It cannot change what tactica. thinks about the football.

That boundary is important.

We didn't add a value filter

There is a subtle point here.

We did not turn expected value into a new gate deciding what tactica. is allowed to recommend.

Existing probability, confidence, tier, market, price, freshness and availability rules still determine whether an expression qualifies.

Only after those existing rules have been satisfied does the new comparison happen.

So an eligible expression can still have a negative calculated EV.

The comparison asks:

Among the expressions that have already earned eligibility for this same football opinion, which is the best execution at these prices?

It doesn't ask:

Can EV override the rest of Match Intelligence?

It can't.

Then independent review rejected our first fix

This is where the story became more useful.

We implemented the new policy.

Tested it.

Submitted it for independent review.

And the reviewer requested changes.

Three material problems had survived our first implementation.

None proved that the original football probabilities were wrong.

They showed that we had made dangerous assumptions about identity, provenance and compatibility while trying to use those probabilities differently.

So the implementation didn't ship.

Dundee wasn't necessarily Dundee United

The first defect sounds almost comical until you see the consequence.

Consider:

Dundee

and:

Dundee United

A loose piece of team-matching logic could find the word “Dundee” inside “Dundee United” and decide it had identified the correct side.

In one review case, it hadn't.

The model's underlying probabilities were:

Dundee 20%
Draw 20%
Dundee United 60%

At the available prices:

Dundee United outright had an expected return of +0.08u.

Dundee United-or-Draw had an expected return of +0.12u.

The Double Chance expression should have won.

Instead, the first implementation associated it with Dundee, calculated the wrong side of the fixture and produced -0.44u.

It therefore selected the outright expression incorrectly.

The new policy idea hadn't failed.

Our identity assumption had.

So we replaced loose matching with complete, exact team and fixture identity.

If identity isn't proven, tactica. now fails closed rather than guessing.

“Latest” didn't mean “caused this prediction”

The second review failure was more important philosophically.

To compare betting expressions properly, tactica. needs the original Home/Draw/Away probabilities that produced the relevant opinion.

Our first implementation could fall back to the latest simulation for a fixture when that source relationship wasn't fully available.

That sounds practical.

It's also exactly the kind of shortcut we had just spent W1 teaching tactica. not to make.

The latest simulation for a match is not necessarily the simulation that caused an earlier published prediction.

Current state is not provenance.

Two sets of probabilities looking identical is not provenance either.

So the corrected contract became stricter:

If tactica. cannot prove the relationship between the projection, simulation and event, the execution context remains unavailable.

It doesn't reverse-engineer Home/Draw/Away from DNB.

It doesn't subtract one protected probability from another.

It doesn't borrow the latest state because it looks plausible.

Unknown remains unknown.

Twelve historical Result contexts still lack their original source relationship today.

W2 deliberately left them unavailable.

That's a feature, not unfinished cleanup.

Then we broke an existing caller

The third review issue was less philosophically interesting but equally release-critical.

One existing part of tactica. could legitimately produce a Result recommendation using information it already possessed during the same calculation.

Our first implementation required new execution context but didn't ensure this existing caller supplied it.

Before the change:

1 Result recommendation.

With the rejected implementation:

0.

The candidate had still passed all the existing qualification rules.

We had simply broken the route by which its legitimate context reached the selector.

Again, weakening the new provenance requirement would have been the wrong fix.

The correct solution was to make the caller supply the authoritative information it already had, through the same ordinary selection process.

After correction:

1 Result recommendation again.

No special bypass.

No relaxed evidence requirement.

So we submitted it again

We fixed all three defects.

Exact team identity.

Proven source lineage.

Existing-caller compatibility without weakening the new rules.

Then the corrected implementation went through another independent review.

This time it was approved.

The reviewed version passed the relevant test suites, the full backend suite, frontend checks and exact probability comparisons.

More importantly, the numerical football model stayed unchanged.

Across the captured probability checks:

xG didn't change.

Home/Draw/Away didn't change.

DNB and Double Chance transformations didn't change.

BTTS didn't change.

Corners and Bookings didn't change.

We weren't changing tactica.'s football opinion.

We were changing how already-qualified expressions of that opinion compete for execution.

Before release, four decisions changed

We ran a bounded comparison across 40 upcoming fixtures while holding the captured eligibility and prices fixed.

There were 70 eligible candidate inputs.

Before W2, 65 expressions were selected.

After W2, 65 were selected.

61 stayed exactly the same.

Four changed from DNB to outright.

There were:

no additions
no removals
no opposite-side changes

Every change stayed within the same underlying football opinion.

That was exactly the sort of narrow delta we wanted.

But it still wasn't production evidence.

So, once again, we waited.

Football eventually gave us the example we needed

After deployment, an ordinary scheduled tactica. production run generated fresh Match Intelligence.

No manual test run.

No hand-picked production write.

Normal operation.

Two persisted decisions showed the new policy clearly.

Barnet v Cheltenham

tactica.'s unconditional football opinion gave Barnet a 64.2% chance of winning.

The DNB expression showed 78.8% and carried an Elite tier.

The outright was only Lean.

Under the old way of looking at those numbers, DNB appeared stronger.

But at the available prices:

Barnet outright — 64.2%, Lean, +0.037u expected net return

Barnet DNB — 78.8%, Elite, -0.012u expected net return

tactica. selected the Lean outright.

Not because tactica. suddenly became less interested in its tier system.

Not because Elite had become worse than Lean.

And not because the 78.8% probability was wrong.

They described different things.

The Elite label described the DNB event under its existing rules.

The 64.2% described Barnet winning the match.

And the downstream expression selector decided that, at those prices, the outright was the better way to execute the same Barnet football opinion.

Port Vale v Crewe

The second production example did the same thing.

Port Vale outright — 58.5%, Edge

Port Vale DNB — 79.5%, Elite

Again, DNB had the larger probability.

Again, DNB had the higher tier.

But at the available prices:

Outright: +0.288u

DNB: +0.200u

tactica. selected the outright.

Those examples don't prove outright bets will outperform DNB in future.

They prove something narrower:

The deployed system was now making the comparison we intended it to make.

We didn't rewrite history either

Changing today's decision policy raised another question.

What should happen to Match Intelligence tactica. had already published under the old rules?

Nothing.

For Barnet and Port Vale, previous Case Files retained the DNB expressions tactica. had actually published at that time.

The newly generated current Bets surface showed the newly selected outright expressions.

Historical tracking, settlements and Performance remained attached to their original recommendations.

That distinction is part of a larger tactica. rule:

Case File = what tactica. published then.

Bets = what tactica. considers executable now.

Performance = what happened to the historical decisions.

Correcting current policy does not grant us permission to rewrite the record.

This wasn't a profitability optimisation

It would be easy to tell this story incorrectly.

DNB had returned -4.99% in our early observational Performance sample.

That did not cause W2.

The sample was small, contained a large number of voids and had substantial uncertainty.

Similarly, we didn't choose expected net return because historical analysis had proved this policy would make more money.

It hadn't.

W2 came from a conceptual problem:

We were allowing probabilities describing different events to compete as though they measured the same conviction.

Once we stopped doing that, we needed a coherent objective for choosing among the already-qualified expressions of one football opinion.

Payoff-correct expected return became that downstream execution objective.

Whether it produces better long-term outcomes still has to be observed.

The bigger number can be answering the easier question

That's probably the most useful lesson from W2.

Numbers gain authority from context.

78.8% looks stronger than 64.2%.

But only if you ignore what each percentage means.

One asks how often Barnet wins the football match.

The other asks how often Barnet wins after removing draws from the denominator.

Both can be mathematically correct.

Neither should impersonate the other.

That is why tactica. now separates:

football belief

from:

betting expression

from:

event tier

from:

execution choice

The distinction makes the system slightly harder to explain.

It also makes it more honest.

And the first time we tried to implement it, independent review stopped us from shipping three assumptions we hadn't earned.

That's part of the story too.

Because an independent review process isn't valuable when it agrees with you.

It's valuable when it doesn't.