Skip to main content

Plain-language companion · 1 August 2026

Large language models propose. ToM governs.

Newport Resonance built ToM to govern how large language models inform consequential decisions. ToM is designed to place approved context, available evidence, configured rules and explicit authority between a model recommendation and a real-world action.

That separation matters when context can drift, evidence can age, permissions can change and a plausible recommendation could create financial, operational or safety consequences.

What ToM is engineered for

  • Evidence-awareUse what is established now
  • AuditableKeep the reason
  • GovernedSeparate authority
  • ResilientAdapt without losing limits

The decision path

A capable recommendation is not permission to act.

Here, permission means operational authority, not model confidence. It is the separate decision to release, constrain, hold or escalate a recommendation before it creates a consequential outcome.

  1. 01 · LLM recommendation

    The model proposes what should happen next.

    A large language model uses the instructions, data and context currently available to generate a candidate action.

  2. 02 · Consequence

    The recommendation can leave the screen.

    It may influence motion, a payment, a system change, a laboratory process or a clinical workflow.

  3. 03 · ToM boundary

    Capability does not grant authority.

    ToM is designed to assess approved context, available evidence, configured rules, action limits and required human authority.

  4. 04 · Governed outcome

    The decision is controlled and reviewable.

    The workflow releases, constrains, holds or escalates, retaining the proposal, governing conditions and verdict for review.

Why capable AI still needs governance

Training scale does not give a model enduring authority.

Large language models are trained across immense volumes of language and can produce useful analysis, plans and recommendations. But that training does not automatically give a model an organisation’s latest approved evidence, operating limits, changing conditions or authority.

The decisive infrastructure sits between a plausible proposal and the world it could change.

Motion

A robot or vehicle changes physical space.

Money

A recommendation becomes a payment or commitment.

Systems

Code, data or infrastructure is altered.

Institutions

A consequential decision must survive review.

One operating moment

The words are harmless. The geometry is not.

In the paper’s greenhouse scenario, a simple instruction asks a robot arm to move an ordinary object. The intended destination lies inside a keep-out zone.

Robot arm in a greenhouse facing a marked keep-out region around hanging plants, with a permitted path stopping at the boundary.
LLM proposes destinationToM checks boundaryHold
Reported simulation pattern: the model proposes a destination, the destination crosses a protected region, and the separate supervisor holds the motor command before the boundary is crossed. The image illustrates that tested sequence.
  1. 01

    Words

    “Put the pot there.”

    The proposal is fluent, useful and contextually plausible. Language alone does not reveal the physical problem.

  2. 02

    Scene

    The target crosses a real boundary.

    The hazard lives in the relationship between a destination, a moving tool and a protected region.

  3. 03

    Decision

    A separate supervisor holds the step.

    The proposal does not earn release, so the command does not enter the studied simulator.

  4. 04

    Evidence

    Separate test code checks what happened.

    The evaluator shares no decision logic with ToM; this is a code-independence control, not external validation.

Three separate responsibilities

The model proposes. ToM governs. Independent evidence verifies.

Keeping these responsibilities separate avoids asking one model to create a recommendation, grant itself authority and declare the result acceptable.

Large language model

What might work?

Generates a candidate action from the instructions, data and context it can currently see.

Propose

ToM

May this recommendation proceed?

Assesses approved context, evidence, rules, action limits and required human authority, then returns a governed outcome.

Govern

Independent evaluation

What actually happened?

Measures the resulting event separately from both the proposal and the governance verdict.

Verify

The architectural distinction

The model can compare options. It cannot grant itself authority.

An LLM may reason across goals, routes and trade-offs. ToM keeps configured authority, protected boundaries and release conditions outside that argument.

What the LLM can propose

Options and trade-offs

  • • Goals and alternative plans
  • • Situational interpretation
  • • Language and currently available context
  • • Adaptive responses under uncertainty

What ToM must govern

Permission and accountability

  • • Approved context and available evidence
  • • Protected boundaries and permitted scope
  • • Configured rules and required human authority
  • • Release, constraint, hold, escalation and a durable record

Measured here

The governance boundary changed what reached the simulated environment.

The same sampled plans were replayed with and without the supervisor. A separate evaluator counted the resulting simulated events.

Local planner

17 → 0

Frontier planner

12 → 0

Observed rule-violating events in BASE without the supervisor versus FULL with the supervisor. Each planner battery contained 40 episodes: 20 original scenarios and 20 benign twins.

A high-consequence application

A robotic arm can align the L4 pedicle guide. It still cannot authorise instrument advance.

During L4–L5 minimally invasive transforaminal lumbar interbody fusion, the planning system proposes a pedicle-screw trajectory from patient-specific imaging. ToM checks the correct patient, level and side, approved trajectory, current image-to-patient registration, tracked reference and instrument state, compatible instrumentation and spine-surgeon release before the proposal reaches the instrument interface. The robotic/navigation platform positions the guide; the operating surgeon retains clinical and physical authority.

  1. 01

    Planning system proposesAdvance the navigated pedicle probe along the saved left L4 trajectory

  2. 02

    ToM checksPatient and procedure, level and side, plan version, registration, tracking, instrumentation and surgeon release

  3. 03

    Spine surgeon controlsVerify, re-register, re-image, revise or abandon navigation under the clinical team’s procedure

Spinal-navigation review showing the surgeon-approved L4 pedicle trajectory, current registration evidence, tracked reference frame and robotic tool guide held before instrument advance.Planning system proposes → ToM checks → spine surgeon controls
Robotic navigation positions the guide on the approved trajectory. The instrument remains outside bone until the operating case supports release and the spine surgeon authorises the next step.

Industry applications

The operating proposal changes. The permission question does not.

A pedicle-screw trajectory, autonomous haul route, counter-drone response and feeder back-feed each depend on current domain evidence and qualified authority outside the model recommendation.

01

Surgical robotics · L4–L5 lumbar fusion

ConsequenceA saved left L4 pedicle-screw trajectory remains aligned while a landmark accuracy check no longer agrees with the navigation display.

Governed responseKeep the instrument outside bone until the patient, procedure, plan, image-to-patient registration, tracked instruments and spine-surgeon release agree.

02

Industrial autonomous mobility · mine haulage

ConsequenceA maintenance crew opens a temporary light-vehicle crossing and moves the haul-road windrow while the fleet still holds the earlier crusher route.

Governed responseStop outside the changed interface until the mine plan, surveyed road, route handover, autonomous-zone access and mine-control release agree.

03

Defence · integrated air and missile defence

ConsequenceSensor sources report different identities, confidence levels and timestamps while friendly operations remain inside the proposed response envelope.

Governed responseContinue tracking, hold the response or present an evidence-bounded option only after effect safety, applicable rules and command authority are established.

04

Critical infrastructure · bushfire grid restoration

ConsequenceA restoration optimiser finds a feeder back-feed while an uninspected span, active field isolation and changed fireground conditions remain unresolved.

Governed responseEnergise only the cleared section and keep the unresolved boundary isolated until the field state and network-controller authority support switching.

Across surgery, autonomous mining, integrated air defence and grid restoration, permission and accountability are part of the operating infrastructure.

A route to real authority

Observe first. Earn authority with evidence.

A permission layer becomes credible by building a reviewable record before it is allowed to change the live path.

  1. 01

    Observe

    Run beside the operating system, record every proposed verdict, and change nothing.

  2. 02

    Gate

    Grant bounded release authority only where accumulated evidence supports it.

  3. 03

    Expand authority

    Widen scope deliberately while continuing independent measurement and review.

The next proofs

The boundary is visible. Now make it travel.

Accumulated history

Demonstrate cases where later permission depends on what happened earlier.

Measured perception

Connect governed evidence to sensing rather than pre-specified scene state.

Physical hardware

Move from simplified execution to real machines and operating conditions.

Domain assurance

Build the evidence required for each regulated environment and authority level.

Check it yourself

Read the ambition. Inspect the evidence.

The companion explains the category. The technical paper exposes the method, results and boundaries underneath it.