Skip to content

Factories > Managed self-hosting

Running agents with an external orchestrator

Open in ChatGPT ↗
Ask ChatGPT about this page
Open in Claude ↗
Ask Claude about this page
Copied!

Connect an external scheduler to self-hosted agents with one-shot Direct or the Command backend.

Use an existing CI system, Kubernetes controller, or internal scheduler to allocate compute while the Automation Platform routes and tracks each run. Choose one-shot Direct when your scheduler starts a worker for each job. Choose the Command backend when a connected worker should hand tasks to an existing runtime.

PatternExternal orchestrator responsibilityWorker behaviorUse when
One-shot DirectAllocates a host, starts a worker, and creates a run for its unique worker IDRuns one accepted task on the allocated host, then exitsYour scheduler creates a VM, pod, or CI runner for each job
Command backendProvides a command or API that accepts a task and owns its compute lifecycleStays connected and invokes the dispatch command for each taskYour runtime already exposes a job API, queue, or scheduler

Both patterns use the managed architecture: Warp accepts triggers and routes runs, while your orchestrator owns compute allocation. If your system should also create and start each agent process without Warp routing, use unmanaged execution.

Complete the shared managed prerequisites. The one-shot example below uses one Default Service Account agent key for both the worker and run-cloud. When those actions run in different processes, use a self-hosted worker key for the worker and an agent key for the run creator.

Install oz-agent-worker from a published release. One-shot Direct also requires the Oz CLI on the allocated host. The Command backend worker requires the tools used by your dispatch and cancellation commands.

Start a Direct worker with --one-shot when your scheduler allocates one host per job. One-shot mode accepts one task, forces max_concurrent_tasks to 1, and exits after that task succeeds, fails, or is cancelled. If no task arrives, the worker stays connected until your orchestrator stops it.

Use a unique worker ID for every job so another worker cannot claim the run.

The agent process normally remains available for follow-up prompts for 45 minutes after a conversation completes. Timeout settings use this precedence:

  1. The task-level config.idle_timeout_minutes value, including defaults applied by the task creator. Runs created directly through the Agent API use 10 minutes when this field is omitted.
  2. The worker’s idle_on_complete setting or --idle-on-complete flag when the task reaches the worker without a task-level value. Unlike a direct Agent API request, oz agent run-cloud uses the CLI task-creation path, which does not apply the API’s 10-minute default and leaves this value unset.
  3. The 45-minute default when neither value is set.

For a job created directly through the Agent API that must exit immediately, set config.idle_timeout_minutes to 0. If the task reaches the worker without this field, start the worker with --idle-on-complete 0s. A task-level value overrides --idle-on-complete 0s.

Add this script to the job your orchestrator starts. CI_JOB_ID must be unique for each job.

run-agent.sh
#!/usr/bin/env bash
set -euo pipefail
: "${WARP_API_KEY:?Set WARP_API_KEY in the job environment}"
: "${CI_JOB_ID:?Set CI_JOB_ID to a unique job identifier}"
worker_id="external-${CI_JOB_ID}"
oz-agent-worker \
--worker-id "$worker_id" \
--backend direct \
--one-shot \
--idle-on-complete 0s &
worker_pid=$!
trap 'kill "$worker_pid" 2>/dev/null || true' EXIT
oz agent run-cloud \
--host "$worker_id" \
--prompt "Run the test suite, fix failures, and open a pull request."
wait "$worker_pid"
trap - EXIT

Expected outcome: The run appears in the dashboard, the worker accepts it, and the worker exits when the conversation completes. The run can enter the queue before the worker connects. If no matching worker claims it, Warp eventually fails the run. Use Direct backend setup and teardown commands to prepare and clean up the allocated host.

Use the Command backend when your runtime already owns job creation and cleanup. The worker sends each assigned task to dispatch_command as JSON on standard input. Exit code 0 means the runtime durably accepted the task; a nonzero exit or dispatch timeout fails it.

The public command-backend example includes Python dispatch and cancellation scripts for an HTTP runtime. Copy and adapt those scripts, then configure the worker:

worker.yaml
worker_id: "external-runtime"
backend:
command:
dispatch_command: "python3 /opt/warp/dispatch.py"
cancel_command: "python3 /opt/warp/cancel.py"
dispatch_timeout: "60s"
environment:
- name: OZ_DISPATCH_URL
value: "https://runtime.internal.example.com/agent-runs"
- name: OZ_CANCEL_URL
value: "https://runtime.internal.example.com/agent-runs/cancel"
- name: OZ_DISPATCH_AUTH_HEADER

Supply the runtime credential and Warp API key through your secret manager, then start the worker:

Terminal window
export OZ_DISPATCH_AUTH_HEADER="Bearer YOUR_RUNTIME_TOKEN"
export WARP_API_KEY="YOUR_SELF_HOSTED_WORKER_API_KEY"
oz-agent-worker --config-file worker.yaml

Your dispatch implementation must handle the task payload as follows:

  • Start command - Run the supplied base_args with env.
  • Runtime configuration - Use docker_image and sidecars to prepare the runtime.
  • Task payload - Keep it out of logs because it contains task credentials. See the example’s task payload reference for each field.
  • Shutdown report - After the CLI exits, run oz harness-support --run-id RUN_ID report-shutdown so Warp receives the terminal state. Set RUN_ID to the run_id field from the dispatch payload.

The worker does not wait for the remote agent to finish. Stopping or restarting the worker does not cancel tasks that the external runtime already accepted. max_concurrent_tasks limits concurrent dispatch calls, not agents running in the external runtime. Configure the external runtime’s capacity separately.