Test autonomous agents before their experience becomes your policy. Begin with one measurable capability.
A score alone does not explain how an agent responds to conflicting evidence, changing conditions or unreliable peers. Work with us to define a bounded evaluation: the capability, the conditions, the evidence and the limits.
Public simulation access and external-agent attachment are in development.
The problem
No single agent has complete truth.
A real organisation is distributed by nature: each department holds local information, goals are not fully aligned, communication is costly and lags, and decisions cannot wait. What matters is not one optimal answer but good-enough judgment that keeps forming as information changes, and that can be explained afterwards. The outcome matters. So does the evidence behind it.
The same structure holds wherever there is no centre and no complete picture: disaster response, drone formations, edge-device networks, cross-team emergency response. That is the condition XMesh is built to test, not a corner case.
The solution
From bounded work to controlled simulation.
The developer runtime and private simulation engine provide different parts of the foundation. External-agent simulation and memory qualification are separate development steps.
- 01
Bounded work
The 0.10 runtime preview records what was asked, what the agent reported, and any checks, review decisions or approvals recorded for that mission.
- 02
Controlled simulation
The private grid engine uses local observations, selected event records and replay of recorded inputs. An outside-agent host is not available yet.
- 03
Memory qualification
Our research asks whether a candidate memory change improves performance on held-out missions. Retaining a record is not proof that it should change future behaviour.
For an organisation of agents
Organizational Cognition.
Organizational Cognition is the direction: independent agents contributing to real work while retaining local control. Start with one workflow, explicit permissions and a measurable capability. Extend the evaluation as the evidence supports it.
Design principles
- 01
Experience
Define the conditions in which an agent acts.
- 02
Influence
Exchange selected cognitive projections without a shared central mind.
- 03
Admission
Let each receiver evaluate what enters its cognition.
- 04
Audit
Distinguish recorded claims, available checks and approvals.
- 05
Adaptation
Evaluate a candidate change before adopting it.
- 06
Transfer
Test an apparent improvement beyond the cases used to develop it.
Evaluation evidence
Agree the evidence before the evaluation.
Define the baseline, success criteria and records needed to judge the capability. Distinguish what was observed, what was reported, what was checked and what remains unverified. Agree the deliverable and its limits before work begins.
How it starts
Start with one agent and one capability.
Tell us the capability you want to evaluate. We’ll discuss the conditions and evidence needed to assess it, and whether the current system is ready for that evaluation.