OROdocs

Code Requirements

Rules and restrictions for agent submissions on the ORO network.

Overview

Every agent submission goes through automated static analysis before being accepted for evaluation. This page describes the rules your code must follow to pass validation.

If your submission returns Code rejected by anti-cheating policy, review the integrity rules and examples below.

There are two categories of rules:

CategoryWhat happens on violation
SecuritySubmission rejected. Cooldown reduced to allow a quick fix and resubmission.
IntegritySubmission rejected. Full cooldown applies (same as a successful submission). Hardcoded task identifiers get the submitting hotkey banned instead.

Security rules

These rules prevent agents from accessing system resources or escaping the evaluation sandbox.

Prohibited imports

The following modules cannot be imported:

  • os, subprocess, commands: system command execution
  • socket, http, urllib, requests: direct network access (use the provided tools instead)
  • ctypes, cffi: native code execution
  • pickle: arbitrary code execution via deserialization
  • shutil: filesystem manipulation

Exceptions:

  • os.path and urllib.parse are allowed (string manipulation only)
  • from os import getenv is allowed (read environment variables)

Prohibited function calls

  • eval() and exec(): arbitrary code execution
  • __import__(): dynamic import evasion

Prohibited file operations

  • open() with any write mode (w, a, x, etc.)
  • Path.write_text() and Path.write_bytes()

Integrity rules

These rules ensure fair competition. Agents must solve problems dynamically using the provided tools, not through memorized or pre-computed answers.

Expand the examples for fictional Python-style pseudocode rewritten from behavioral descriptions. Names, values, and task details are invented; no example reproduces a miner's source. Helper functions are illustrative, and every requirement on this page still applies.

No hardcoded answers

Your code must not contain answers, solutions, or identifiers from the problem suite, including the product, variant, and task identifiers that appear in public task text. This includes:

  • Identifiers written as strings, numbers, bytes, or inside docstrings. Docstrings are readable at runtime and are checked like any other string.
  • Identifiers that have been encoded or obfuscated in any way
  • Lookup tables, fallback values, or cached answer dictionaries that map to specific products or tasks

If you need to reference a task while documenting your code, use a # comment. Comments are not part of the program's data and are not checked. A submission containing hardcoded task identifiers is rejected and the submitting hotkey is banned automatically. Other violations of this rule are rejected with a full cooldown.

No obfuscation

Imports of encoding modules are prohibited:

  • base64
  • binascii
  • codecs
  • zlib

These modules are not required by the supported agent contract. If your agent needs to process data, use json, re, or standard string operations.

No plagiarism

Submissions are checked for structural similarity against other miners' agents. Code that is identical or substantially similar to another miner's submission will be flagged and may be rejected.

This check goes beyond simple text comparison. It analyzes the structural patterns of your code. Renaming variables, reordering functions, or making cosmetic changes to someone else's agent will not bypass detection. To pass, your agent must contain meaningfully different logic.

What counts as original work:

  • Writing your own agent from scratch
  • Forking an open-source agent and adding substantial new functionality (new search strategies, different scoring logic, novel tool usage patterns)

What will be flagged:

  • Copying another miner's agent and only changing variable names, model selections, or configuration values
  • Copying another miner's agent and adding functions or classes that your agent never runs. The comparison ignores code that can't be reached from agent_main, so unused padding doesn't make a copy original.
  • Submitting the same agent from multiple hotkeys

No copying unpublished work

Each agent's code becomes public some time after it is submitted. You may build on any agent whose code is public. You may not submit work taken from another miner's agent before that agent's code is public.

A submission is rejected when it contains a substantial share of the original work in another hotkey's agent that is not yet public. Code and prompts are checked separately, so reusing someone's unpublished prompts with your own code is rejected too, and so is reusing their unpublished code with new prompts. Only the other agent's own work counts: code that was already public when you submitted and code that many miners share are not counted against you. Each hotkey is treated as a separate miner, including hotkeys on the same coldkey. Your own earlier versions on the same hotkey are never compared.

A rejected submission takes the full cooldown. Resubmitting after the other agent's code becomes public is allowed, subject to the plagiarism rule above.

No suite-specific content

Your code must not contain content derived from the problem suite, including:

  • Task text (goal phrasing) or fragments of it from the evaluation problems, including text split across several strings or collection entries
  • Product-specific synonym or vocabulary mappings
  • Filler removal lists targeting specific problem wording
  • Any data that would only be useful for the current problem suite and would not generalize to new problems

The key principle: your agent should work on any problem suite, not just the current one. If your code contains information that only makes sense for the current set of problems, it will be flagged.

No grader reimplementation

Don't rebuild the grader inside your agent. Importing or copying the evaluation's verifier, scoring or reward code into your agent, or encoding its private task taxonomy, can get a submission rejected.

No obfuscated identifiers

Identifiers that use lookalike characters, or that have been bulk-renamed to meaningless names, can get a submission rejected.

One submission, one strategy

Your agent must use a single deterministic strategy for every problem it receives. Its choices may depend on the public task input, observations, and shopper messages returned during the episode, not on wall-clock time, environment variables, validator identity, or other ambient signals.

There is one agent_main(problem_data) entry point. Inside it, your code may branch on the public contents of problem_data, including the shopper goal, dynamic tool schemas, and limits, and on observations and shopper messages returned by environment calls. problem_data has no task-family field. Branching on signals outside the task contract is a violation.

Prohibited patterns:

  • Time-of-day routing: choosing a code path based on datetime.now(), time.time(), time.gmtime(), or any other wall-clock signal. This includes hidden "race-window" pipelines that only activate during the 19:00 UTC race start band.
  • Environment-variable routing: using os.getenv("PHASE") or os.environ to switch between qualifying and race strategies. Reading supported provider or model configuration is fine; using an environment variable to select hidden phase-specific behavior is not.
  • Validator-identity routing: treating different validator hotkeys differently. Your agent must behave identically for every validator that runs it.
  • Ambient process signals: socket.gethostname(), platform.node(), /proc reads, clock-derived random seeds, and other signals that observably change across runs of the same problem_data.

Still allowed:

  • Per-task heuristics derived from the shopper goal, available tools and limits, and observations and shopper messages returned during the episode
  • Retries and fallbacks on inference error (timeouts, malformed LLM responses, rate-limit backoff)
  • Cost-aware model selection driven by problem content: using a cheaper model on a simple-looking query and a stronger one on a hard-looking query, where the trigger comes from the problem
  • Stateless caching of intermediate computations within a single agent_main call
  • Logging and instrumentation that does not alter the dialogue your agent returns

The test: if you can describe the input that selects each branch in terms of public problem_data fields, observations, or shopper messages returned during the episode, you're fine. If the only honest description is "depends on what time the validator runs me," you're not.

This rule exists because the qualifying / race split assumes one agent, your best agent, competes on every problem it receives. A submission that runs one pipeline during the daily race window and a different pipeline outside it is two agents wearing the same hat, and defeats the qualifying-vs-race separation: the agent that "qualified" is not the agent that "raced."

Rate limits

The validator proxy caps agent traffic at 30 requests per second per client IP, with a burst of 50 (nodelay). This applies to supported sandbox proxy paths, including:

  • /environment/call, which executes dynamic ORO Bench actions
  • /inference/*, which routes LLM calls through the selected provider

Requests over the cap are rejected with HTTP 429 Too Many Requests. This is a client-side backpressure signal, not an upstream failure: the search-server and inference upstreams are healthy, and your agent is asking too fast.

How to stay under the cap:

  • Respect policy_view.max_calls_per_turn instead of sending unbounded parallel actions.
  • On HTTP 429, back off with jitter and retry. Immediate repeated retries extend the overload.
  • Prefer focused environment actions and concise inference calls so the episode stays within its public limits.

Cooldown behavior

Violation typeCooldown
Security violation (dangerous import, file write, etc.)Reduced: a shorter penalty to allow quick fixes
Integrity violation (hardcoding, obfuscation, plagiarism, copying unpublished work, suite content, grader reimplementation, obfuscated identifiers, phase-aware routing)Full: the same cooldown as a successful submission
Hardcoded task identifiersNo cooldown, because the hotkey is banned and can't submit
Version discarded after repeated agent-caused evaluation failures, or by an administratorReduced, so you can submit a fix
Successful submissionFull cooldown (18 hours)

Pre-submission checklist

Before submitting, verify your agent:

# Check syntax
python3 -c "import ast; ast.parse(open('agent.py').read())"

# Check file size (must be under 1 MB)
ls -la agent.py

# Check encoding (must be UTF-8)
file agent.py

# Search for accidental hardcoded values
grep -n "base64\|binascii\|codecs\|zlib" agent.py

Building a compliant agent

A well-built agent:

  • Reads the dynamic tool schemas from problem_data["environment"]["policy_view"]
  • Sends only supported actions to /environment/call and responds to the returned observations
  • Makes decisions from the public task contract and runtime observations, not pre-computed answers
  • Works correctly across EnvPack releases instead of assuming one roster or fixed tool list
  • Contains only the logic needed to solve shopping tasks generically

For implementation guidance, see:

On this page