Research note · 20 August 2026 · Synexiom Labs
Thirty sealed forecasts, scored by reality
Our reasoning system put a probability on 30 real events before they happened. All 30 have resolved. Here is how it did, including where the market did better.
Why sealed forecasts
Most AI tests have a weakness: the answers can leak into the data a model was trained on. A forecast made before an event happens cannot be memorised, because the answer does not exist yet. Reality does the grading.
On 22 June 2026, Cortexiom, our product built on the reasoning architecture, gave one probability for each of 30 open questions on the Kalshi prediction market. By 20 August 2026, every one of them had resolved.
0
times it called something very likely and it did not happen
across all 30 events
0.9565
discrimination (AUC): ranking what happened above what didn’t
1.0 is a perfect ranking
0.0875
Brier score: average squared error
lower is better; 0 is perfect
0.142
calibration error (ECE)
for context only: the best published result on a different, static set is 0.120
What went well
Its top five predictions all happened. Of the seven events that did occur, five were in its top five. Only two of thirty calls were confidently wrong, and both erred toward caution: it rated two events as unlikely that then happened. There is not one case in thirty where it said something was very likely and it did not occur.
For a system you might route decisions on, the direction of its errors matters as much as their number.
Where the market did better
It did not beat the market, and we do not claim it did. Against the Kalshi price at the moment of collection, it was closer on 12 of 30 questions. Its average error was lower (0.0875 against 0.1201), but the difference is not statistically significant: a few large wins on genuinely uncertain questions, not a consistent edge.
The market won on longshots. Where the crowd priced an event at 1 to 7 per cent, our system said 5 to 16 per cent. Both mean “no”, but the market’s extra decisiveness scores better. The system is reluctant to go below about 5 per cent, even when it should; that is a correctable trait. Its two worst misses were the two economics questions.
What this does not show
Thirty is a small sample, and 27 of the 30 questions were about sport, because the collection window fell during the World Cup. This shows calibrated forecasting on a sport-heavy sample, not across every domain. No other AI model forecast these questions at the time, so there is no head-to-head, and there cannot be one now: asking a model today would test its memory of the news, not its judgement.
A fair comparison needs several frontier models forecasting the same questions, at the same time, before anything resolves. That is how the next round has to be run.
How we tested
Thirty open Kalshi markets collected on 22 June 2026, one probability each from Cortexiom, sealed before resolution. Scored against the outcomes once all had resolved (20 August 2026). The market comparison uses the Kalshi price at collection time. Paired comparison of per-question error: t = −1.41, not significant. The published reference figure (0.120) comes from a different, static question set and is context, not a controlled comparison.
More of our evaluation work is on the research pages.
Bring us one decision.
We’ll show you what the reasoning layer would check, explain and record on one of yours.