← Insights

Research note · 30 September 2026 · Synexiom Labs

Where checking beats generating

Rules checked in code, evidence verified word for word, and where our own method did not help: notes from testing a review step on a case it had never seen.

When an AI drafts something an organisation will act on, such as a Board paper, a plan or an approval, what actually makes the result safe to act on? A stronger model? More steps? We tested our own answer, including on a case we had not tuned on. Some of it held up. Some of it did not.

1. Hard rules belong in code

On a fictional harbour, a capable model planning vessel movements broke at least one hard rule in every first plan, six out of six. Checking each plan against the rules in plain code caught every violation, and plans built that way held. A model can be asked to follow rules; it cannot be relied on to. The full field note has the details.

2. Evidence should be checked word for word

When a review cites a source, plain code checks that the quoted words actually appear in it. That matters because a fluent review can misquote a policy as confidently as it quotes one correctly.

The check caught a mistake of ours first. An early version of our matcher handled special characters, like “m³”, on one side of the comparison only, and in one run it rejected 27 of 74 correct quotes as fabricated. We fixed it, kept the affected runs on record, and did not score them. A check you cannot audit is just another opinion.

3. More steps did not find more problems

We built a multi-step review that works through a draft claim by claim and verifies the evidence before it reports. On our development cases it looked better than a single review on one of the models. Then we wrote a held-out case, a fictional dredging contract award, after the method was tuned and before any review of it existed.

On that case, the multi-step review did not catch more of the problems in the answer key than one well-written review request, on any of the three models we tried: Anthropic’s Claude, Cohere’s Command A+ and Moonshot’s Kimi K3. The improvement we had seen in development did not carry over. It also showed a design limit: on a briefing that was not itself a decision, the multi-step review pulled in the conditions of a different decision, a scope error the single review did not make.

Test on a case you have not tuned on. Our development results looked better than the held-out one.

4. What the extra steps did change was the form

The multi-step review did not find more, but what it produced was different: every claim triaged, every quote verified against its source, an owner for each problem, a worklist, a routing decision made by rule, and the order of steps to get to a valid decision. That is what makes a review fast to act on, and it is what we build on.

Time per review

One well-written review≈ 70 s
Multi-step review≈ 4 min

Tokens per review

One well-written review≈ 6K tokens
Multi-step review≈ 45K tokens
On Claude, held-out case, September 2026.

It costs more. On Claude, a multi-step review took about four minutes and 45,000 tokens; the single review took about 70 seconds and 6,000 tokens. Whether that is worth it depends on what the decision costs to get wrong.

What we take from it

The model

  • Reads the situation
  • Drafts the plan or the review
  • Explains it in your words

Plain code

  • Checks every hard rule
  • Verifies every quote against its source
  • Does the arithmetic
A decision you can act on, with its record

Use the model for judgement and language. Use code for rules and evidence. Spend extra steps on making the output something people can act on, not on hoping they find more. That division is the core of the reasoning layer.

How we tested

Fictional organisations and documents. Answer keys were written and committed before any draft or review existed; the held-out case was written after the method was frozen. Every review was scored by a judge model from a different lab than any model under test (OpenAI’s GPT-5.6 Sol), and the judge first recorded which problems the draft already handled, so a review earned nothing for repeating them. September 2026. All runs are kept, including failed ones.

Bring us one decision.

We’ll show you what the reasoning layer would check, explain and record on one of yours.

Book a call