Dashboard
Use the dashboard to create a test environment, connect your agent, define a successful outcome, and inspect what changed. Test files and run history stay in your project.
Start with the installation guide. For an existing installation, see launching and reopening the dashboard.
Check connector scope for available resources, imports and outcomes before building a test.
Create your first test
1. Create a world
On Overview, choose Create a world. Select the systems your agent uses, choose starting data, and give the world a recognizable name.
- Templates provide ready-made starting data you can preview and customize.
- Describe my workflow helps you prepare your own case.
- Import my data starts from your cases or a supported export.
Review the included records, relationships and starting time, then choose Create and explore world. A world is an isolated working copy of that starting data. Creating one does not require an agent or a test suite.
To learn the workflow first, choose Try a guided sample on Overview. The billing, support and combined examples use scripted agents and need no provider key. Compare a failed run with its correction, then return to testing your agent.
2. Inspect the starting state
The world opens on State. Browse its records and relationships to check that your test has the information the agent needs.
Use Failure conditions to add a supported failure, such as a rate limit. Use Clock to advance logical time. Available controls depend on the connector.
Evidence lets you create a mark, a baseline for comparing later changes. A mark does not restore state. Manage → Reset world restores the world's starting state; creating another world leaves this one intact.
3. Connect your agent
Open Connect, then select an agent or choose Create an agent profile. Save its command, working directory and timeout. Follow the connector's instructions to point its SDK at the local world.
Profiles can reference a private environment file and map your agent's endpoint variables. Keep credentials out of version control. Source-account import credentials are separate from the local twin credentials.
Connection checks diagnose endpoint configuration. If your agent waits for a task, use a test to check its normal behavior. Copy developer setup instructions provides a handoff without credentials or answer keys.
4. Define the outcome
Choose Create test in the world. Select the original starting data and clock, or preview an export of its current state. Saving creates a suite: the tasks, starting data and expected outcomes that future runs will use.
In Test suites, edit each task's expected changes and protections. For example, a support task might require a particular status and assignee while protecting unrelated tickets. A billing task might require one specific refund and no changes to another customer.
Combined tests need explicit links between the relevant records. Matching names or email addresses do not establish those links. Written instructions alone are not grading rules; finish the structured expectations or supported YAML rules. Tasks without expectations remain labeled as ungraded.
5. Run fresh tests
Choose Run this suite, select the agent, and review the starting data, outcomes and run settings.
| Mode | When to use it |
|---|---|
| Independent tasks | Test each task in its own fresh world. |
| Shared workflow | Test an ordered set of tasks in the same fresh world. |
Set Repetitions to run the test more than once. Save test keeps the configuration for later; Save and run fresh tests saves it and starts execution. You can save a test before an agent is connected.
Read results and try again
A completed process does not necessarily mean the test passed. Start with Outcome to see the evaluation, then open Evidence for task grades, recorded changes and available reports. A failed protection can fail the overall world even when an individual task was completed correctly.
Open Reproduce to choose the next run:
| Action | What it uses |
|---|---|
| Rerun saved inputs | The captured test data and execution settings. |
| Review current configuration | Current project files and agent settings, reviewed before running. |
| Rerun selected failures from saved inputs | Selected tasks with their saved inputs and protections. |
Saved inputs do not freeze agent code or installed dependencies. A run of selected failures does not establish that the full suite passes. Compare runs identifies compatible tests; when inputs differ, inspect the evidence side by side.
If a submission loses its response, use the displayed recovery action before starting another. Reloading does not rerun work automatically. Keep the launcher terminal open while working; closing the browser leaves running work active.
Custom workflows from your own cases
Under Create a world, choose Describe my workflow or Import my data. Name the preparation so you can resume it later.
- Add the request and starting situation. Choose an example, project dataset or supported source export for the records the task needs.
- For CSV, JSON or JSONL cases, review the field mapping and select the cases to include. Imported request text does not supply missing business records.
- Review each case's relevant objects, proposed outcomes and protections. Confirm relationships explicitly, especially across connectors.
- Choose Review prepared test files, inspect them, then Save prepared workflow. Continue to Create and explore world and the test flow above.
Preparation is private and resumable. Changed starting data requires reviewing affected outcomes and relationships again.
Account imports and support exports
A scoped Stripe import can collect selected customers, invoices, charges or subscriptions with their supported dependencies. Review the sanitized records, omissions and starting date. The saved data is frozen; reruns do not contact the account again.
Native Zendesk accepts supported offline snapshots. Its Full JSON export path uses extracted ticket envelopes, users and optional organizations. Review the conversation cutoff, starting status, assignment, groups and agent identity. For resolved cases, exclude later resolution messages from the starting context. Live Zendesk ingestion is not supported. See the native workflow reference.
Combined datasets must have compatible starting times and non-overlapping connector data. Imports copy the supported state; they do not reconstruct a complete account history.
Optional model drafting and private data
You can prepare tests manually without a model. Optional OpenAI or Anthropic drafting sends only the content you review and explicitly approve to the selected provider. Supply a private environment-file reference or a session-only key. Review the proposed revision, its assumptions and its outcomes before accepting it.
Accepted private test files live under worlds/private/. Keep environment files, identity maps and source-account credentials out of version control. Default evidence and CI exports exclude private text-bearing inputs; review download previews and prepare de-identified inputs before making a portable CI configuration.
Private reports may omit text or process output. Missing evidence is shown as missing, rather than inferred. Optional advisory analysis remains separate from the test's authoritative verdict.
Return to saved work
Overview shows unfinished setups, saved tests and recent outcomes. Use Review and run to revisit a saved test after changing your agent.
In Worlds → Saved setups, reuse named starting data to create another world. Save as starting setup on a live world previews a reusable export of its current state; this does not preserve its full history or alter the live world.
Runs → Saved tests and batches runs several saved configurations in order with one agent profile. Each result remains inspectable, and a failed test does not automatically cancel later tests. Preview the CI export to run reviewed configurations outside the dashboard.
| Destination | Use it for |
|---|---|
| Overview | Resume work and find recent outcomes. |
| Worlds | Templates, saved starting setups and live world inspection. |
| Test suites | Tasks, outcomes and protections. |
| Runs | History, batches, comparisons and recorded evidence. |
| Failure conditions | Configure supported scenarios on a world. |
| Settings | Agent profiles, imports, project setup and runtime connection. |