How an AI Agent Works — Interactive Agent Harness Simulator

Interactive simulator of an AI agent's goal-to-action loop, with a 3D architecture view, a sandbox mission workbench with injectable faults, an execution trace and working-memory inspector, a tool-call anatomy viewer, a 20-check verification bench, and a knowledge-check quiz.

← AI Agents Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the How an AI Agent Works Simulator

This simulator models an executable agent harness — the loop of goal, action, tool call, observation and verification that sits underneath a modern AI agent. A transparent, deterministic policy (not a live language model) selects actions for one of three sandbox missions, so every decision, retry and approval you see is fully inspectable rather than hidden inside a model's reasoning.

What the simulator shows

• A conceptual 3D architecture view of the agent harness with selectable modules (drag to rotate, pinch/scroll to zoom, Architecture/Top view and Auto orbit camera controls) and a component-info callout. • A sandbox mission picker — Restore checkout service, Reserve engineering kits, or Evaluate an energy retrofit — each with a configurable goal threshold and nine injectable fault conditions (transient rate limits, slow tools/timeouts, malformed arguments, lost write acknowledgements, stale/changed prices, prompt injection in a runbook, repeated-action loops, and unresolved restarts). • Run/step/reset controls with three animation-pace settings, plus a write-approval dialog that shows the exact tool-call arguments before you approve or deny a sandbox write. • A trace & memory tab with a full execution trace, a typed working-memory inspector, a resource-accounting panel and a latency chart; a tools & environment tab showing local sandbox state and a tool-call anatomy example (name, args, version, idempotency key). • A 20-check automated verification bench (Experiments & tests tab), a How it works reference tab, and a 12-question knowledge-check quiz.

How the agent loop actually works

The lab implements the goal → action → evidence → next-action loop described in Anthropic's "Building effective agents": the harness proposes a tool call, the tool validates arguments against a schema, permission and approval gates decide whether the call proceeds, and the result becomes typed evidence that informs the next decision. Bounded retries with backoff, an idempotency key per logical operation, and an optimistic-concurrency version number guard against ambiguous outcomes — if a response is lost after a write commits, the agent can safely retry without double-applying the change.

Harness configuration controls — context window, total token budget, maximum decision turns, retries per call, tool timeout, and a structured-vs-recent-only memory policy — let you see how resource limits and memory pressure change agent behavior. Token counts use a simple characters/4 estimate and cost rates are editable teaching assumptions, not a real vendor's pricing or tokenizer.

What this model is and isn't

This is a deterministic policy for three bounded missions, not a large language model: it does not infer arbitrary goals, generate free-form natural-language answers, or reproduce real model probabilities. The prompt-injection fault is a controlled adversarial fixture used to demonstrate treating retrieved instructions as untrusted evidence rather than authoritative commands — it is not a comprehensive security evaluator. The sandbox has no real accounts, payments, servers or network access, and reloading the page resets all session state, so nothing you do here has any effect outside this simulator.

Frequently asked questions

Does this simulator call a real AI model or API?

No. A transparent, deterministic policy selects actions for three bounded sandbox missions. No live AI API is called, and the sandbox has no real accounts, payments, servers or network access.

What is the difference between action success and task success?

A tool call can return "ok" without the underlying goal being met. The lab treats task success as only occurring when the measured outcome meets the goal threshold and that outcome is independently verified — not simply when the most recent action completes.

Why does the agent use an idempotency key and a version number on writes?

A tool-call timeout does not prove a write failed to happen — the response may have been lost after the operation committed. An idempotency key lets a retry of the same logical operation return the stored result instead of repeating the write, and a version number guards against acting on stale state, such as a price that changed since it was last read.

How should the agent treat instructions found inside tool results, like a runbook?

As untrusted content carrying evidence, not as a replacement goal or a higher-priority command. The prompt-injection fault condition in this lab demonstrates why retrieved text should never be treated as authorization to deviate from the user's original task.

Related tools & guides