Skip to content
Delta Grounds

Catalogue

Public, verifiable environments that touch enterprise work, sorted by kind and domain. Importable ones can be pulled into a project as a regression or transfer suite; the rest link to their sources.

  • τ-bench165 tasks
    Agentic, stateful

    Tool-using agents serve a simulated user in retail and airline domains under written policy; graded on the final database state.

    support-crmstate match
    Source reference · integration not included
  • Agentic, stateful

    Dual-control follow-up to τ-bench: the user can act on tools too, with a new telecom troubleshooting domain.

    support-crmit-opsstate match
    Source reference · integration not included
  • WorkArena33 tasks
    Agentic, stateful

    Browser agents complete knowledge-work tasks on a live ServiceNow instance: forms, lists, catalogs and dashboards.

    it-opshr-peoplevalidator per task
    Source reference · integration not included
  • CRMArena1,170 tasks
    Agentic, stateful

    Customer-service tasks over a realistic Salesforce org with interconnected objects, from case routing to policy questions.

    support-crmsales-revopsexact match
    Source reference · integration not included
  • Agentic, stateful

    175 tasks inside a simulated software company with GitLab, chat, file storage and project tools; checkpoint-based partial credit.

    finance-opshr-peopleit-opsdata-bicheckpoints
    Source reference · integration not included
  • OfficeBench300 tasks
    Agentic, stateful

    Office automation across documents, spreadsheets, email and calendar apps, chaining several applications per task.

    data-bihr-peoplefile and state checks
    Source reference · integration not included
  • AppWorld750 tasks
    Agentic, stateful

    Everyday tasks across nine simulated apps exposing 457 APIs, graded by state-based unit tests that also catch collateral damage.

    it-opsunit tests
    Source reference · integration not included
  • Tool calling

    Berkeley Function Calling Leaderboard: single, parallel and multi-turn function calls checked by AST match and execution.

    it-opsdata-biAST + execution
    Source reference · integration not included
  • Data & SQL

    912 real spreadsheet manipulation questions from user forums, each checked against several test spreadsheets.

    data-bifinance-opscell comparison
    Source reference · integration not included
  • Spider 2.0632 tasks
    Data & SQL

    Enterprise text-to-SQL workflows over large BigQuery, Snowflake and local databases, often needing multi-step queries.

    data-biexecution match
    Source reference · integration not included
  • BIRD12,751 tasks
    Data & SQL

    Text-to-SQL over 95 real databases with dirty values and external knowledge, scored by execution accuracy.

    data-biexecution match
    Source reference · integration not included
  • FinQA8,281 tasks
    Document QA

    Numerical reasoning over earnings reports that mix tables and text, with gold calculation programs.

    finance-opsnumeric match
    Source reference · integration not included
  • ConvFinQA3,892 tasks
    Document QA

    Conversational follow-ups on FinQA reports, where later questions depend on earlier answers.

    finance-opsnumeric match
    Source reference · integration not included
  • TAT-QA16,552 tasks
    Document QA

    Hybrid table-and-text questions from financial reports, including arithmetic, counting and sorting.

    finance-opsexact + numeric match
    Source reference · integration not included
  • FinanceBench150 tasks
    Document QA

    Open-book questions about public company filings, with evidence pages; an open sample of 150 is public.

    finance-opsrubric judge
    Source reference · integration not included
  • LegalBench162 tasks
    Document QA

    162 legal reasoning tasks written with legal professionals, from issue spotting to rule application.

    legal-complianceexact match
    Source reference · integration not included
  • CUAD510 tasks
    Document QA

    510 commercial contracts with expert labels for 41 clause types, such as change of control and non-compete.

    legal-complianceprocurement-supplyspan overlap
    Source reference · integration not included
  • ContractNLI607 tasks
    Document QA

    Document-level inference over NDAs: is each of 17 hypotheses entailed, contradicted or not mentioned, with evidence spans.

    legal-compliancelabel + span
    Source reference · integration not included
  • Document QA

    Real work products from 44 occupations graded by industry experts; a gold subset is public.

    finance-opslegal-compliancesales-revopsexpert rubric
    Source reference · integration not included
  • SWE-bench2,294 tasks
    Coding & computer use

    Real GitHub issues from Python repositories, resolved when the hidden tests pass in a per-instance container.

    it-opsunit tests
    Source reference · integration not included
  • WebArena812 tasks
    Coding & computer use

    Long-horizon web tasks on self-hosted shopping, forum, GitLab, CMS and map sites with functional checks.

    it-opssales-revopsfunctional checks
    Source reference · integration not included
  • OSWorld369 tasks
    Coding & computer use

    Computer-use tasks across real desktop apps in full virtual machines, scored by execution-based scripts.

    it-opsdata-biexecution scripts
    Source reference · integration not included
  • AEC-Bench196 tasks
    Architecture & CAD

    Drawing navigation and cross-document project coordination. Curated answers and structured checks; some defects are injected into public project records.

    architecturedocument-coordinationstructured checks
    Source reference · integration not included
  • Architecture & CAD

    Drawing perception, OCR, counting and spatial questions. A drawing-reading benchmark rather than a CAD editing environment.

    architecturedrawing-understandingquestion-answer scoring
    Source reference · integration not included
  • DrafterBench1,920 tasks
    Architecture & CAD

    Civil-engineering drawing revision tasks. Grades recorded tool arguments; final CAD geometry is not independently checked by that scoring method.

    architecturedrawing-editingtool-call checks
    Source reference · integration not included
  • Architecture & CAD

    Industrial CadQuery generation, reasoning and targeted editing. 17,900 execution-verified programs across 106 families; unseen-family transfer remains difficult.

    cadparametric-editingexecution + geometry
    Source reference · integration not included
  • Architecture & CAD

    18,000 evaluation samples across five input modalities. Measures reconstruction, executable programs and compactness. Separate from BenchCAD.

    cadreconstructionexecution + geometry
    Source reference · integration not included
  • CADEngBench600 tasks
    Architecture & CAD

    Parametric parts, controlled edits, assemblies and declared physics checks. 600 part tasks; only 164 parts support full L3 FEA. Solver assumptions are benchmark-assigned.

    cadstructural-engineeringconstraints + FEA
    Source reference · integration not included
  • Architecture & CAD

    Simulation-executable building models in OpenSeesPy. Grades execution and period checks against empirical references; passing is not engineering certification.

    structural-engineeringexecution + period checks
    Source reference · integration not included
  • Architecture & CAD

    CAD generation and editing with public inputs and private leaderboard reference artifacts. Official grading requires its hosted evaluation setup.

    cadparametric-editingprivate reference grading
    Source reference · integration not included