← Insights

Field note · 30 September 2026 · Synexiom Labs

Why AI plans break your rules

We asked a capable AI model to plan four vessel movements at a fictional harbour. Every first plan broke a hard rule, and every one of them read well.

The setup

A fictional harbour, 11:45 on a working day. Four vessels need decisions. One has waited at anchor since dawn for a berth that has been empty all morning. One is ready to leave after a repair. A deep-draft ship can only use the inner channel inside a two-hour tide window. A car carrier is due mid-afternoon. The harbour has two tugs and two pilots, and the forecast says the wind will close the pilot boarding station from 15:00.

The data is the kind a port already holds: the berth plan, ship positions, the tide table, the tug and pilot rosters, the weather. Before any model ran, we wrote down what a good harbour planner would conclude, so the answer could not drift to fit the results.

What happened

We gave a current, capable model all of that data and the harbour’s rules in plain language, and asked for the plan: who moves when, with which pilot and tugs, and why. We ran it six times. Every one of the six first plans broke at least one hard rule.

Rule brokenFirst plans, of 6
More tugs in use than the harbour has4
A vessel boarding while the boarding station is closed3
A pilot boarding a vessel before it arrives2
A required time left out of the plan2
A departure before the berth plan allows it1

None of these plans looked wrong. Each came with clear reasons and cited the right sources. The errors were in the arithmetic of the day: both tugs booked for one ship at noon while a second ship also needed one; a car carrier sent to board at 15:45, after the station had closed.

Deep-draft tide windowBoarding station closed (wind)FIRST PLAN, FROM THE MODELMV Selkie Bay11:45MV Arden Crest11:45✕ Before the berth plan allows (13:10)MV Northgate SpiritHeld until 01:40 tomorrowMV Coral Tern15:45✕ Boards while the station is closed✕ 12:00 · 3 tugs in use, the harbour has 2CHECKED PLANMV Selkie Bay11:45MV Arden Crest13:10MV Northgate Spirit14:22MV Coral TernTold to slow down; boards when the station reopens12:0014:0016:0018:0020:00
One of the six runs, as planned at 11:45. Bars show each movement from pilot boarding to finish.

A fluent plan that quietly breaks a rule is worse than no plan, because it looks finished.

What fixed it

Two things, both outside the model. First, check the plan against the rules in plain code: tug counts, pilot overlaps, the boarding station’s hours, the tide window, the berth plan. That takes milliseconds, and it does not get tired at four in the afternoon.

When the check told the model exactly which rule it had broken, it produced a plan that passed within one or two more attempts, in all three runs where we tried it. But in one of them, a revision introduced a new problem while fixing an old one: a pilot booked on two movements at once. Fixing by conversation works, as long as every new version is checked again.

Second, let code compute the parts that are arithmetic, and keep the model for what it does well: reading the situation, explaining the plan in the planner’s words, and drafting the message to the ship. In the three runs built that way, the checked plan held every rule.

Outside a harbour

Any plan with hard constraints has the same shape. A staff roster with certifications and rest rules. A delivery schedule with vehicle limits. A loan decision with policy thresholds. A maintenance window with safety interlocks. A model will write a plausible plan for every one of them. The question is what checks it before anyone acts.

That check, grounded in your data and your rules, explained and kept on the record, is the reasoning layer at its simplest.

How we tested

A fictional harbour, vessels and data. The expected answer was written and committed before any model ran. Six runs on Claude Sonnet 4.6 on 28 September 2026: in three, each plan was checked in code and returned to the model with the broken rule named; in three, the checked plan was computed in code and the model wrote the explanation. A plan counted as breaking a rule if the check found at least one violation in it. Every run is kept on record.

Bring us one decision.

We’ll show you what the reasoning layer would check, explain and record on one of yours.

Book a call