OROdocs

Local Testing

Run the production-equivalent ORO Bench qualifying workflow locally with Docker.

What the local workflow runs

Local testing uses the same generated runtime and verifier as qualifying. The repository bundles a practice pack, which is a sanitized qualifying delivery that holds no race tasks; the runner refuses any other kind of pack. It validates that pack, then runs its first five tasks in archive order. Pass --problems to sample a different set while iterating. The practice tasks are not the network's current qualifying suite, and the workflow does not verify them against it.

The workflow includes:

  • oro-env-runtime sessions
  • the verifier, rewards, and failure categories
  • the validator proxy and agent sandbox
  • a search image pinned to the exact EnvPack index
  • the live Backend model allowlist
  • trusted episode receipts and an aggregate score

It does not compile new tasks or claim evaluation work from the Backend.

System requirements

  • Docker with Compose support
  • Git LFS
  • At least 16 GB of free disk space
  • An AMD64 host, or Docker Desktop emulation on an ARM Mac
  • A Chutes or OpenRouter API key

Docker downloads the pinned search image on the first run. It is a multi-gigabyte download. The runtime rejects a mismatched search identity before starting the agent sandbox.

Before selecting local tasks, the runtime validates the complete archive and its catalog references. Archive validation runs on every invocation and can take several minutes. Let it finish before starting the local evaluation.

Setup

git clone https://github.com/ORO-AI/oro
cd oro
git lfs pull
cp .env.example .env

Set one provider key in .env:

CHUTES_API_KEY=
OPENROUTER_API_KEY=

Both keys may remain configured. Choose one with:

INFERENCE_PROVIDER=chutes

or:

INFERENCE_PROVIDER=openrouter

OpenRouter takes precedence if both keys exist and no provider is selected.

SANDBOX_MODEL optionally overrides the model used by the included reference agent. Custom agents select their own model or models in code, and every request is checked against the same live Backend allowlist used in qualifying.

Build the current local services after pulling the pack:

docker compose build test test-proxy sandbox

Run an agent

Run the included generated-environment reference:

docker compose run test --agent-file src/agent/environment_agent.py

Treat the included agent as a functioning protocol reference, not a strong or optimized baseline. See Agent Interface for the implementation and contract details.

Run your own agent:

docker compose run test --agent-file my_agent.py

Run one local test at a time per Compose project. Dependencies remain available for later runs. Stop them when finished:

docker compose --profile test down

Run fewer problems

The default run is the pack's first five tasks in archive order. To sample a different set instead, pass --problems with any number from 1 up to the number of problems in the pack. A small sample is the quickest check while iterating:

docker compose run test --agent-file my_agent.py --problems 2

The problems are sampled at random from the whole pack, and up to seven run at once (LOCAL_MAX_WORKERS).

Each run samples afresh, so repeated runs do not tune your agent against one lucky subset. The report prints the seed it used. --seed only applies alongside --problems, since a full run draws no sample. Pass that seed back to repeat a selection exactly, which is what you want when comparing two versions of an agent:

docker compose run test --agent-file my_agent.py --problems 2 --seed 4821993

A local score is not comparable to qualifying, which runs the active suite's full qualifying roster.

Configuration

Flags

FlagDefaultPurpose
--agent-filesrc/agent/environment_agent.pyThe agent file to run.
--problemsThe first five problemsRun a random sample of this many problems instead, from 1 up to the number in the pack.
--seedA fresh seed for each sampled runRepeat an earlier --problems selection. Only valid alongside --problems.

SANDBOX_MODEL sets the model the included reference agent requests. The proxy maps it to the active provider, so the default reaches Chutes as deepseek-ai/DeepSeek-V3.2-TEE and OpenRouter as deepseek/deepseek-v3.2. A custom agent picks its own models in code.

Environment variables

VariableDefaultPurpose
INFERENCE_PROVIDEROpenRouter when both keys existSelect chutes or openrouter.
SANDBOX_MODELdeepseek-ai/DeepSeek-V3.2-TEEOverride the included reference agent's model.
BACKEND_URLhttps://api.oroagents.comSource of the live model allowlist.
LOCAL_MAX_WORKERS7Maximum concurrent problem workers.
LOCAL_TIMEOUT1800Sandbox execution budget in seconds. Pack validation and session setup happen before this timer.
LOCAL_ENV_PACK_PATHBundled release packUse another pack available inside /workspace.
LOCAL_ENV_PACK_SHA256Bundled release digestPin the expected digest for an override pack.

Configuration, pack integrity, search identity, and infrastructure failures return a nonzero exit status.

Read the result

The command prints a run header, each finalized problem with its reward, and where the artifacts landed. A failed or partial problem also names its failure categories: the primary one first, then every category that applies in parentheses.

ORO Bench local run  local-7c1f2a
  pack        5b0e93a1…41f2
  problems    5 of 12, qualifying roster
  runtime     0.3.4  verifier 0.4.0
  inference   openrouter
  agent       my_agent.py  sha256 3f9c1d7b0000…
  agent model deepseek-ai/DeepSeek-V3.2-TEE  (SANDBOX_MODEL, requested by the reference
              agent and mapped per provider; custom agents choose in code)
  simulator   mistralai/mistral-small-2603
  judge       deepseek/deepseek-v4-flash-0731

composed               mean 0.34  2/5 passed
  TF8-composed-320000  completed         1.00
  TF8-composed-320003  completed         0.70  extra_questions (extra_questions)
  TF8-composed-320006  completed         0.00  request_not_met (request_not_met, needs_not_found)
  TF8-composed-320009  completed         0.00  needs_not_found (needs_not_found)
  TF8-composed-320011  completed         0.00  did_not_finish (did_not_finish)

Aggregate score  0.340000
Artifacts        ./logs/environment-runs/local-7c1f2a
Trajectories     ./logs/environment-runs/local-7c1f2a/trajectories.html  (open in a browser)

The header's runtime and verifier lines are the contract versions the pack was sealed for (runtime contract 0.3.4, which oro-env-runtime 3.0.1 implements), not the Python package version. Rewards are coloured when the output is a terminal. Set NO_COLOR=1 to turn that off. The simulator and judge models are sealed in the pack. The agent model line shows SANDBOX_MODEL, which is a request rather than a record: only the included reference agent reads it, and the proxy maps it to the active provider's name for that model, so an OpenRouter run of the default sends deepseek/deepseek-v3.2. A custom agent chooses its own models in code and ignores the setting.

A completed problem can still receive zero reward if the verifier rejects the outcome, and a problem that failed in the environment rather than in your agent is labelled with the reason. Per-task rewards and the aggregate come from trusted runtime receipts, not values written by the agent.

Review trajectories

Every run that finalizes at least one episode writes trajectories.html into its run directory. That includes a failed run, which carries whatever episodes finished before the failure and names the file in the error output. A run that fails before any episode finalizes has nothing to show, so it writes no viewer. Open it in a browser to step through each episode: the shopper request, your agent's messages and tool calls, the observations it got back, simulator and market events, the verifier checks that decided the verdict, the failure categories, and the reward. Locally you see every check because the practice tasks are public. On the network, an episode's feedback names only its failure category.

It is a single self-contained file that runs with no server and no network access, so you can copy it off a remote host and open it locally. Use its Open JSON button, or drag files onto the page, to compare episodes from other runs.

Each run directory contains:

FileContents
summary.jsonBundled archive digest, local task roster, selection mode and seed, the agent file and its digest, the models declared by the pack, per-task and per-family rewards, runtime error classification, and aggregate score.
trajectories.htmlSelf-contained trajectory viewer for every episode in the run.
sandbox/sandbox_output.jsonlUntrusted agent trajectory output.
episode_results.jsonlFinalized runtime receipts with verifier verdicts, call traces, ledgers, and provenance.
environment_sessions.jsonRuntime session inputs.
problems.jsonlSelected task inputs.

The report prints host paths, so the directory it names is the one in your checkout. Inspect a run with:

run_dir=./logs/environment-runs/local-...
python3 -m json.tool "$run_dir/summary.json"
wc -l "$run_dir/sandbox/sandbox_output.jsonl"

Evaluator artifacts are mounted read-only. The agent writes only inside sandbox/. The output reader rejects links, non-regular files, non-object rows, and files larger than 128 MiB.

Interrupted or failed sandbox runs retain any episode receipts that were successfully finalized. Their summary remains usable and records the failure reason alongside partial task outcomes.

Keep artifacts private

Local outputs contain sealed task data and agent trajectories. Keep the entire run directory local and do not publish it.

Next steps

  • Agent Interface: Implement the generated environment loop.
  • Scoring: Understand task rewards, race scores, and Overall score.
  • Submitting: Submit after the local workflow is healthy.

On this page