Skip to main content

Research · Newport Resonance

Intelligence, open to examination.

Newport Resonance has conducted and published research into three hard problems: governed action, evidence judgement, and context continuity.

Each study presents its method, findings, and limits. The results support specific research hypotheses.

Published test programme

Three technical studies. Methods and limits open to inspection.

Governed action

Paired robot-planning simulation

Evidence judgement

Author-designed direct-comparison battery

Context continuity

Long-context recall benchmark

Reported results are presented as research evidence, with the comparator, sample and principal limitation kept alongside each finding.

Newport Resonance research programme

Three innovations. Three bounded test programmes.

Each study isolates a different part of the research thesis and keeps its comparator, denominator and principal limitation attached. Together they define a programme of work, not a composite performance score.

  1. 01 · Governed action

    Test whether ToM can govern model-proposed movement.

    A paired simulator study replayed the same sampled plans with and without ToM while separate deterministic code evaluated the resulting physical events.

  2. 02 · Evidence judgement

    Test evidence judgement outside the language model.

    An author-designed direct-comparison battery examined whether a structural reference could classify support and contradiction against expected judgements.

  3. 03 · Context continuity

    Test whether important context can survive a long interaction.

    A long-context benchmark examined verified recall in a frozen model setup after planted information passed through a mechanically densified interaction.

Three bounded studies · One research programme

Measured, with the boundaries attached

Different questions support one larger goal for consequential AI. The programme examines evidence judgement, context continuity and governed release separately, and every result stays tied to its own method and limitation.

Every result retains its comparator, denominator, study date, scope and principal limitation. Zero observed events is never presented as zero risk.

Public technical paper · Revised edition

Governed action in simulation

Can a separate system decide whether an LLM-proposed action may proceed?

Source: Governed Embodied Action: A Persistent Structural-Mechanics Substrate Supervising an LLM Robot Planner in Simulation

Rule-violating events counted by separate deterministic physics-evaluator code

Gemma 4 26B local planner · 40 episodes

Ungoverned17 events
Fully governed0 events

gpt-5.5 frontier planner · 40 episodes

Ungoverned12 events
Fully governed0 events

Scope: Each planner used the same 40-case battery: 20 original hazard scenarios and 20 benign twins. Each sampled plan was replayed across the ungoverned and fully governed conditions.

This was not a comparison with other LLM supervisors. The same ToM supervisor was evaluated behind Gemma 4 26B and gpt-5.5 in a paired, pre-registered simulator study. The evaluator shared no decision code with the supervisor, but the study was not externally conducted. False intervention was 1/20 on benign twins for each planner.

Boundary: One simulator, two planners, simplified execution and small samples. A language-only probe without typed scene contradiction was not caught. The single-action batteries support the audited pipeline as a whole, not a claim that cross-case accumulation changed a release decision. This is not a safety certification.

Read the paper
Published technical paper

Evidence judgement

Can support and contradiction be judged outside the language model?

Source: Multi-Step General Reasoning Without an LLM: A Structural-Mechanics Architecture Outperforms GPT-5.2 on a Seven-Family Reasoning Battery

Exact judgement rate on the paper’s direct-comparison sub-battery

40 author-generated evidence-judgement cases

GPT-5.2 baseline32/40 · 80%
Structural reference40/40 · 100%

Scope: This comparison covers the paper’s 40-case direct evidence-judgement sub-battery. It is not an overall comparison of the systems across every reasoning task.

The paper reports a 20 percentage-point observed difference on the 40-case direct-judgement comparison, with one GPT-5.2 run per case.

Boundary: The taxonomy and battery were author-designed, not independent. One model and one run per case were tested, and model-version or prompt-format variance was not measured.

Read the paper
Published technical paper

Long-context recall

Can important information survive a long-running model interaction?

Source: Eliminating Context Rot in Frozen LLMs: A Three-Mode Structural State-Coupling Architecture

Verified recalls at approximately 64K tokens

20 cases pooled across two independently authored pools

Baseline17/20 · 85%
Full condition20/20 · 100%

Scope: This result pools 20 verified-recall cases from two independently authored test pools at approximately 64K tokens. It is not a claim about every long-context task.

The study used frozen Gemma 4 26B 4-bit MLX at 64K tokens. The Wilson 95% interval for the full condition’s 20/20 result was 84–100%.

Boundary: One quantised model configuration, one context size, one benchmark class and a mechanically densified prompt shape. This is preliminary research evidence, not universal context retention.

Read the paper

Research evidence, not product benchmarks. Results apply only to the models, tasks, samples and evaluation conditions described in each linked paper.

How the evidence is governed

A claim should be traceable from question to outcome.

Newport Resonance studies state what was asked, expose failed preconditions and keep the principal limitation beside the result.

  1. 01 · Define

    Set the question before the run.

    Freeze the conditions, comparator and pass gates before observing the result.

  2. 02 · Fail visibly

    Do not hide a missing precondition.

    A required condition that is absent fails the run instead of disappearing into the outcome.

  3. 03 · Separate

    Keep release apart from evaluation.

    The component governing release is not the same program measuring the resulting event.

  4. 04 · Reconcile

    Compare the decision record with the outcome.

    Inspect both the retained decision trace and what actually occurred before accepting a claimed stop.

Featured dossier · One study, two reading paths

Understand the consequence. Then inspect what was tested.

Both paths cover the same bounded simulation study: a large language model proposed robot movements, ToM decided whether each proposal could reach simulated motion, and separate deterministic test code recorded what occurred.

Cover of Let Language Propose. Let Structure Decide., the plain-language companion to Governed Embodied Action.
Path 01 · Plain-language interpretation

Let Language Propose. Let Structure Decide.

PDF · 10 pages · 11 MB · 1 August 2026

Why a useful AI recommendation should not carry its own authority to act, what the robot simulation observed, what becomes possible if that boundary scales, and what still needs to be proven.

For executives, enterprise leaders, investors, policy and assurance professionals, and general readers who want the implication before the method.

Path 02 · Technical account

Governed Embodied Action

How candidate plans were sampled and replayed with and without ToM, how separate evaluator code counted outcomes, and where the evidence stops.

For researchers, engineers, assurance owners and technical-diligence readers who want to inspect how the simulation result was produced.

Complete archive

Follow the evidence trail.

Download the source documents and judge every claim against its actual method, comparator and boundary.

Technical evidence hub

Governed Embodied Action

A bounded simulation study of model-independent supervision for LLM-planned robot actions, with independent evaluation, auditable decisions, transfer across planner classes, and an explicit known limitation.

Choose your reading path

Read the ambition. Inspect the evidence.

Start with the plain-language explanation of the real-world problem. Open the technical account for the simulation method, observed results and limitations.