← Insights

Research note · 20 August 2026 · Synexiom Labs

Thirty sealed forecasts, scored by reality

Our reasoning system put a probability on 30 real events before they happened. All 30 have resolved. Here is how it did, including where the market did better.

Why sealed forecasts

Most AI tests have a weakness: the answers can leak into the data a model was trained on. A forecast made before an event happens cannot be memorised, because the answer does not exist yet. Reality does the grading.

On 22 June 2026, Cortexiom, our product built on the reasoning architecture, gave one probability for each of 30 open questions on the Kalshi prediction market. By 20 August 2026, every one of them had resolved.

Happened7 eventsDidn’t happen23 eventsRated 0.48 or higher:all five happenedNone here0.25.50.751probability it gave, before the event
All 30 forecasts at the probability given before the event. Gold: it happened. Outline: it did not.

0

times it called something very likely and it did not happen

across all 30 events

0.9565

discrimination (AUC): ranking what happened above what didn’t

1.0 is a perfect ranking

0.0875

Brier score: average squared error

lower is better; 0 is perfect

0.142

calibration error (ECE)

for context only: the best published result on a different, static set is 0.120

What went well

Its top five predictions all happened. Of the seven events that did occur, five were in its top five. Only two of thirty calls were confidently wrong, and both erred toward caution: it rated two events as unlikely that then happened. There is not one case in thirty where it said something was very likely and it did not occur.

For a system you might route decisions on, the direction of its errors matters as much as their number.

Where the market did better

It did not beat the market, and we do not claim it did. Against the Kalshi price at the moment of collection, it was closer on 12 of 30 questions. Its average error was lower (0.0875 against 0.1201), but the difference is not statistically significant: a few large wins on genuinely uncertain questions, not a consistent edge.

The market won on longshots. Where the crowd priced an event at 1 to 7 per cent, our system said 5 to 16 per cent. Both mean “no”, but the market’s extra decisiveness scores better. The system is reluctant to go below about 5 per cent, even when it should; that is a correctable trait. Its two worst misses were the two economics questions.

What this does not show

Thirty is a small sample, and 27 of the 30 questions were about sport, because the collection window fell during the World Cup. This shows calibrated forecasting on a sport-heavy sample, not across every domain. No other AI model forecast these questions at the time, so there is no head-to-head, and there cannot be one now: asking a model today would test its memory of the news, not its judgement.

A fair comparison needs several frontier models forecasting the same questions, at the same time, before anything resolves. That is how the next round has to be run.

How we tested

Thirty open Kalshi markets collected on 22 June 2026, one probability each from Cortexiom, sealed before resolution. Scored against the outcomes once all had resolved (20 August 2026). The market comparison uses the Kalshi price at collection time. Paired comparison of per-question error: t = −1.41, not significant. The published reference figure (0.120) comes from a different, static question set and is context, not a controlled comparison.

More of our evaluation work is on the research pages.

Bring us one decision.

We’ll show you what the reasoning layer would check, explain and record on one of yours.

Book a call