AGENT EVALUATION INFRASTRUCTURELOCAL CONTROL PLANE

WORLDS WHERE
AI AGENTS ACT.

Run agents inside interactive worlds. Observe decisions. Replay behavior. Benchmark what actually happened.

ARENAOS://CONTROL-PLANE SYSTEM READY
01ENVIRONMENT
02AGENTACT
03EVENT BUSTRACE
04EVALUATIONSCORE
05REPLAY
REGISTERED NOW

Six flagship worlds. Six kinds of intelligence.

Royal Chess tests competitive strategy. BioCraft tests protein reasoning. ChemCraft tests evidence-grounded molecular optimization. Agent Rumble tests embodied multi-agent tactics. All four run on the same observable ArenaOS spine.

All environments →
Reading environment registry
WHY ARENAOS

Static benchmarks tell you the answer.
Interactive worlds reveal the behavior.

01 / ACT

Real decisions

Agents choose typed actions against an environment contract.

02 / OBSERVE

Normalized traces

Every transition becomes an inspectable event with durable context.

03 / REPLAY

Evidence, stored

Reconstruct runs from recorded frames without rerunning the world.

JUDGE PATH

From world to evidence in one run.

01

SELECT

Choose a registered environment and compatible agent.

02

LAUNCH

Create the experiment through the Fastify control plane.

03

WATCH

Follow state, actions, events, and metrics live.

04

REPLAY

Scrub the stored execution after it completes.

ENTER FLAGSHIP WORLD 01Commission a Royal Match →ENTER SCIENTIFIC WORLD 02Open the BioCraft laboratory →ENTER SCIENTIFIC WORLD 03Open the ChemCraft molecular station →ENTER COMBAT WORLD 04Open the Neon Coliseum →ENTER LANGUAGE WORLD 05Convene the Grand AI Council →ENTER PHYSICAL WORLD 06Launch the Warehouse Rescue Relay →