Infrastructure to build and harden multi-agent systems
Multi-agent AI systems are the next major unlock. Coding agents are already interacting with each other, and interacting agents will soon touch every area of the economy: banks will use agents for customer service, healthcare will use them for triage, and personalized secretaries will interact with each other on the internet to perform tasks for their owners.
We are not prepared for this.
Research on safety in multi-agent systems shows that these agents are Agents of Chaos: red-teaming multi-agent AI systemsShapira et al. Northeastern UniversityInstagram AI chatbot tricked by hackers to give access to others'
accountsReported by Liv McMahon for BBC News, and reliable systems for making interacting agents robust have yet to be developed. This bottlenecks deployment and creates massive societal risk.
We fix this bottleneck.
We custom-build realistic simulations of economically valuable environments that agents will run in. We then use these environments to test and red-team groups of agents, iteratively improving both the agents and the environments. This process predictably removes sources of harm, which we measure as we go. The result is groups of interacting agents that can be safely deployed in high-risk settings.
02/ backgroundWe were the first to publish work on red-teaming multiple agents, Agents of Chaos (Shapira et al., 2026), covered by Science and Wired, documenting how the harms due to interacting agents can manifest. We began this work immediately after OpenClaw was introduced. We were then hired by a frontier lab to design and lead internal red-teaming campaigns.
03/ approachCustom agentic infrastructure. We adapt your business's native environment and workflow into a realistic testbed for agents to work in, and for us to stress-test.
Red-teaming. We research your agent-integrated workflow, then attack it to uncover agentic vulnerabilities — leveraging our track record of finding the vulnerabilities that cause real harm.
Integrated harm reports. We generate a structured plan for you to integrate AI agents safely, and we act as third-party auditors for your integrations. We'll produce a daily-updated harm taxonomy, severity ratings, sanitized artifacts, and recommended mitigations so that your business is free to act immediately.