Task package spec
A task package is one problem a model can attempt, plus everything needed to grade it without a human. Delta Grounds reads the common open package layout (task, environment, verifier, reference solution) unchanged, so existing packages import as they are, and adds an optional assay block for enterprise metadata.
Package layout
invoice-match-0412/ ├─ task.md frontmatter + "## prompt" ├─ environment/ │ ├─ Dockerfile the sandbox image │ └─ seed/inputs.json copied into the workspace ├─ verifier/ │ ├─ test.sh runs pytest, writes reward.txt and reward.json │ ├─ test_outputs.py the checks │ ├─ verifier.md strategy and rubric │ └─ data/expected.json private ground truth, uploaded after the attempt └─ oracle/ └─ solve.sh reference solution
The workspace is $BENCHFLOW_WORKSPACE (default /root); Delta Grounds also sets ASSAY_WORKSPACE to the same path. The verifier directory is never visible during the attempt.
task.md frontmatter
--- version: "1.0" metadata: author_name: Acme Finance Ops author_email: finops@acme.example category: data-processing # one of 18 standard categories license: Apache-2.0 # SPDX identifier origin: original # original | adapted | generated difficulty: medium # easy | medium | hard tags: [accounts-payable, three-way-match] agent: timeout_sec: 900 verifier: timeout_sec: 180 environment: build_timeout_sec: 600 cpus: 1 memory_mb: 2048 storage_mb: 10240 allow_internet: false ---
The assay block
assay:
domain: finance-ops # finance-ops, support-crm, legal-compliance, hr-people,
# procurement-supply, data-bi, it-ops, sales-revops
family: invoice-three-way-match # tasks in a family share one skill
harness: program # answer | program | agent
workspace_files: [inputs.json] # inlined into the prompt for single-turn harnesses
outputs: [answer.json] # files the verifier reads
split_hint: train # advisory; project suites decideFamilies matter for proof. The transfer suite is made of families no training task belongs to. If every task were its own family, held-out improvement could only ever mean memorization.
Writing the prompt
Put the instructions under ## prompt. Name every input path and the exact output path and schema. The verifier is mechanical, so an ambiguous prompt produces ambiguous grades. Write the business rule the way the policy document states it, including ordering and tie-breaks.
Never put the answer, or anything that determines it without doing the work, in the prompt or in the environment.
Environment
- Start from a small pinned base image and preinstall
pytest==8.4.1andpytest-json-ctrf==0.3.5. - Pin every verifier dependency, so grading works with the network off.
- Copy seed data only. Never copy the verifier, expected outputs or the oracle into the image.
- Generate seed data deterministically from a seed, so reruns start from the same state.
Verifier
test.sh runs test_outputs.py with pytest and writes /logs/verifier/reward.txt (1.0 if every test passes, else 0.0), reward.json and a CTRF report. Delta Grounds also records partial credit as the share of tests passed.
A good verifier:
- recomputes the expected answer from its private copy of the seed data instead of hard-coding it;
- checks values, not just that an output file exists;
- is deterministic: no clocks, randomness, network or model calls;
- scores the oracle 1.0 and an untouched workspace 0.0;
- rejects plausible-but-wrong outputs. Delta Grounds measures this with mutation testing: it perturbs the oracle output (a changed amount, a flipped flag, a dropped row) and counts how often the verifier notices.
Oracle
oracle/solve.sh solves the task in the same sandbox the model gets. It must score 1.0 on all 8 reruns. Compute the answer from the seed data and keep it simple: it documents the intended method and seeds the mutation tests.
Harnesses
| Harness | The model receives | The model returns | Use for |
|---|---|---|---|
| answer | Prompt plus inlined workspace files | One or more <file path=…> blocks | Short structured answers |
| program | Prompt plus inlined workspace files | One Python program, run once in the sandbox | Training (default): single-turn and procedural |
| agent | Prompt and tools: list, read, write, run | Tool calls over many turns, then submit | Evaluating API models and long workflows |
Quality gates
| Severity | Triggers | Effect |
|---|---|---|
| Block | Task name matches a sealed task; a prompt shares half its 13-grams with a sealed task | The whole collection fails validation |
| Reject | Image copies verifier data or answers; verifier reads answer-like files; any shared 13-gram with a sealed task | The task is excluded |
| Control | No working oracle | Eligible only if a no-op scores 0 on 8 reruns and the base model solves it at least once |
| Review | Existence-only assertions; remote ADD; oracle downloads; caches or .git in the image | Advisory, shown to the reviewer |
Dynamic gates run in the sandbox: the image builds and the oracle scores 1.0 on 8 reruns, an untouched workspace scores 0.0 on 8 reruns, and the base model solves the task 1 to 3 times out of 4.
Collections
my-collection/ ├─ submission.yaml team_name, contact_email, track: environments └─ envs/ ├─ invoice-match-0412/ └─ … 1 to 200 packages
Upload a collection as a tar.gz, point at a public Hugging Face dataset or GitHub repository, or create one from the Task Builder. Delta Grounds pins the revision, runs the gates, and then it can be trained against a challenge in the Arena.