Zendesk test files and CLI workflows

Use this reference to connect an agent, import a snapshot, author native task files and run them in CI. Native Zendesk and combined support-and-billing workflows are available in the public beta. Live Zendesk API qualification remains pending; the API reference defines the supported subset.

On this page

To use the dashboard, start with the dashboard guide. Installation is covered in the quickstart.

Start in an agent project

Use a pinned beta runtime for the commands and generated CI configuration. The example below creates a native support project; replace node agent.mjs with your agent command. A deterministic test agent needs no Zendesk account or model key.

You can also set WORLDS_RUNTIME to a versioned tarball URL. A local tarball path works on your machine, but must also exist on the CI runner if retained in its configuration.

Keep the JSON requested by --json-file alongside the HTML requested by --report. In 0.9.0-beta.2, native standalone HTML shows the overall result and task findings, but details of world-level protected or unexplained changes remain in the safe JSON. A passing task does not override a failed overall result. The same limitation applies to beta.1. Newer source builds include an HTML rendering correction that is not part of either beta archive.

Init creates:

Output Purpose
Project seed Starting records and clock.
worlds/tasks.yaml Tasks and expected outcomes.
Policy and connection guides Agent setup and test policy.
Node, Vitest or pytest harness Run the suite with the project's test runner.
.github/workflows/worlds-agent-tests.yml Run the same tests in CI.

Init requires an agent command and refuses existing scaffold destinations before writing. It preserves unrelated files. The harness and workflow pin the project seed directory so another library cannot replace a same-named seed. For an external server, start that server with --seeds-dir seeds; a client-side fallback does not override the server's library.

Choose --connector composite for the bundled support-and-billing example. The CLI defaults to Stripe, so include the connector option for native tests. Native init does not read ambient Stripe import keys; it refuses Stripe-only import and policy options.

Connect an agent and check traffic

worlds/connect.md contains the pinned node-zendesk installation and configuration example. The agent receives:

binding meaning
WORLDS_BASE_URL this world's origin
WORLDS_API_KEY this world's API key
WORLDS_ZENDESK_URL origin followed by /api/v2
WORLDS_ZENDESK_TOKEN the same world key
WORLDS_WORLD_ID the isolated world's identity

Configure endpointUri from WORLDS_ZENDESK_URL, token from WORLDS_ZENDESK_TOKEN, and enable oauth: true. See the SDK example.

Use numeric ticket IDs. A private note is a ticket update with comment: {body: "Internal note", public: false}. Set public: true for a public reply. Repeating a successful update can append another comment.

For composite agents, Stripe uses the same world origin and world key. The existing Stripe preload redirects a stock supported Stripe SDK unless --no-preload or WORLDS_PRELOAD=0 disables it; an explicit host remains respected. Zendesk needs its explicit endpoint configuration. There is no Zendesk preload or hosted MCP routing promise.

Without an agent command, doctor performs an HTTP check and reports qualification: http-only. With a command, it checks requests observed while that process runs; combined diagnostics require both connectors.

The check does not pass when traffic is absent, a route is unsupported, execution fails or times out, state is unavailable, or cleanup fails. Its report identifies the reason and reached/missing/unsupported surfaces without child output, credentials or raw provider errors.

An agent that waits for task input may make no requests during this diagnostic; use a fresh test to check its normal behavior. Connectivity does not establish provider fidelity.

Import an offline snapshot

The CLI accepts either a native seed with all Stripe/legacy data collections empty, or the worlds-zendesk-native-export wrapper in the seed format. It does not fetch Zendesk or turn arbitrary vendor CSV into a native snapshot. For the reviewed Full JSON export flow, use the dashboard importer.

Relationships and supported ticket/comment/audit history must be complete and valid. Import preserves the supported structure while transforming identifying data:

Data Import behavior
IDs and relationships Remapped deterministically; comments and audit events share one identity namespace.
Timestamps, privacy, memberships and declared lifecycle Retained.
Free text and identities Pseudonymized.
Attachments and unsupported fields Omitted and counted.

A combined seed with populated Stripe records is refused by this importer. Import a separate sanitized Stripe seed, then compose the two.

The default outputs are seeds/my-support.json and the mandatory worlds/my-support.manifest.json. Review the manifest's omissions before using the seed.

Optional --identity-map worlds/my-support.identities.json writes a private authoring map with restrictive permissions and a Git ignore entry. The map must be inside the selected project, outside Git metadata and untracked. Existing destinations and descendant symlinks are refused. The map is never given to an agent, native compiler, CI workflow or public report.

Init copies a supplied seed unchanged, so use an already sanitized input, normally the importer's output. Structural validation does not anonymize it. If the input is already the regular seeds/<name>.json destination, init keeps it; otherwise the destination must be absent.

The starter task uses the lowest-ID active ticket and requires a valid assignment and a private principal note. Invalid relationships or a seed with no active ticket are refused.

Custom composite input needs explicit bindings:

Both snapshots must already be sanitized and share an epoch. The Stripe fixture must not contain legacy desk tickets. tasks.seed names the new composed seed, and each monetary task supplies its explicit stripe_customer binding. Identity is never joined by a guessed matching email. The composed destination must be absent and distinct from both inputs.

Version 2 task authority

Native tasks use version: 2 and connector: zendesk, including combined tasks. They reference tickets already present in the seed. Start from the task file generated by native or composite init; its compiled rubric is a separate answer key. Version-1 task files do not accept native worlds.

Each task binds a unique id, numeric ticket_id and requester_id.

Field What it grades or permits
expected.ticket Final status and assignment; allowed tag additions/removals.
expected.comments Exact new public/private comment counts and author counts.
allowed_ticket_fields Only the listed ticket changes.
allow_rule_closure Explicit permission for the configured simulated closure.
forbidden_changes Protected resource IDs or a resource wildcard.
forbidden_public_text Local text checks on new public comments; the text is excluded from reports.

Optional Stripe expectations use the existing money matcher and explicit customer authority. Shared-customer tasks share accounting: one refund cannot satisfy two promises, and an extra refund remains overreach.

Filtering with --task retains protections for unselected tasks. Version 2 refuses tasks build --rewrite, --keep-seed-tickets and legacy tasks import --into.

The compiler preserves every seed record and keeps the answer key out of world state. The runner sends this shape through both stdin and WORLDS_TASK_JSON:

{
  "version": 2,
  "connector": "zendesk",
  "mode": "isolated",
  "tickets": [{ "id": 101, "task_id": "assign-ticket" }]
}

Here, id is the connector ticket ID and task_id is the suite task ID. Your agent fetches task context from the local API. Expected outcomes and identity maps are not included.

--mode isolated gives each task its own world; --mode shift gives the selected tasks one shared world. Each repetition starts fresh.

Grading validates before/after state, persisted comments and audits, clocks and the diff. Missing or inconsistent state is inconclusive. Completed isolated-task grades remain available if another world cannot be graded.

Exit code Result
0 Passed.
1 Graded failure.
2 Invalid configuration.
3 Setup failure.
4 Inconclusive or interrupted execution.

Reports and CI

Native grading discards agent stdout/stderr and uses persisted outcomes. JSON and self-contained HTML contain the closed result projection, not comment bodies, expected text, source identities or raw exceptions. The native runner does not use an LLM judge. Explicit private evidence remains local and subject to its destination policy; generated workflows upload only worlds-results.json and worlds-report.html.

When using the setup action, keep the immutable ref generated by init and capture-evidence: 'false'. Teardown still runs, while raw world/request captures and server log tails are suppressed. Branch names, tags, short hashes and unreviewed refs are refused by native setup.

Local tests and synthetic fixtures establish the declared behavior. Live headers and tenant-dependent behavior still require qualification; a passing local suite does not establish complete Zendesk compatibility.