Catalogue
Public, verifiable environments that touch enterprise work, sorted by kind and domain. Importable ones can be pulled into a project as a regression or transfer suite; the rest link to their sources.
- τ-bench165 tasksAgentic, stateful
Tool-using agents serve a simulated user in retail and airline domains under written policy; graded on the final database state.
support-crmstate matchSource reference · integration not included - Agentic, stateful
Dual-control follow-up to τ-bench: the user can act on tools too, with a new telecom troubleshooting domain.
support-crmit-opsstate matchSource reference · integration not included - WorkArena33 tasksAgentic, stateful
Browser agents complete knowledge-work tasks on a live ServiceNow instance: forms, lists, catalogs and dashboards.
it-opshr-peoplevalidator per taskSource reference · integration not included - CRMArena1,170 tasksAgentic, stateful
Customer-service tasks over a realistic Salesforce org with interconnected objects, from case routing to policy questions.
support-crmsales-revopsexact matchSource reference · integration not included - TheAgentCompany175 tasksAgentic, stateful
175 tasks inside a simulated software company with GitLab, chat, file storage and project tools; checkpoint-based partial credit.
finance-opshr-peopleit-opsdata-bicheckpointsSource reference · integration not included - OfficeBench300 tasksAgentic, stateful
Office automation across documents, spreadsheets, email and calendar apps, chaining several applications per task.
data-bihr-peoplefile and state checksSource reference · integration not included - AppWorld750 tasksAgentic, stateful
Everyday tasks across nine simulated apps exposing 457 APIs, graded by state-based unit tests that also catch collateral damage.
it-opsunit testsSource reference · integration not included - Tool calling
Berkeley Function Calling Leaderboard: single, parallel and multi-turn function calls checked by AST match and execution.
it-opsdata-biAST + executionSource reference · integration not included - SpreadsheetBench912 tasksData & SQL
912 real spreadsheet manipulation questions from user forums, each checked against several test spreadsheets.
data-bifinance-opscell comparisonSource reference · integration not included - Spider 2.0632 tasksData & SQL
Enterprise text-to-SQL workflows over large BigQuery, Snowflake and local databases, often needing multi-step queries.
data-biexecution matchSource reference · integration not included - BIRD12,751 tasksData & SQL
Text-to-SQL over 95 real databases with dirty values and external knowledge, scored by execution accuracy.
data-biexecution matchSource reference · integration not included - FinQA8,281 tasksDocument QA
Numerical reasoning over earnings reports that mix tables and text, with gold calculation programs.
finance-opsnumeric matchSource reference · integration not included - ConvFinQA3,892 tasksDocument QA
Conversational follow-ups on FinQA reports, where later questions depend on earlier answers.
finance-opsnumeric matchSource reference · integration not included - TAT-QA16,552 tasksDocument QA
Hybrid table-and-text questions from financial reports, including arithmetic, counting and sorting.
finance-opsexact + numeric matchSource reference · integration not included - FinanceBench150 tasksDocument QA
Open-book questions about public company filings, with evidence pages; an open sample of 150 is public.
finance-opsrubric judgeSource reference · integration not included - LegalBench162 tasksDocument QA
162 legal reasoning tasks written with legal professionals, from issue spotting to rule application.
legal-complianceexact matchSource reference · integration not included - CUAD510 tasksDocument QA
510 commercial contracts with expert labels for 41 clause types, such as change of control and non-compete.
legal-complianceprocurement-supplyspan overlapSource reference · integration not included - ContractNLI607 tasksDocument QA
Document-level inference over NDAs: is each of 17 hypotheses entailed, contradicted or not mentioned, with evidence spans.
legal-compliancelabel + spanSource reference · integration not included - Document QA
Real work products from 44 occupations graded by industry experts; a gold subset is public.
finance-opslegal-compliancesales-revopsexpert rubricSource reference · integration not included - SWE-bench2,294 tasksCoding & computer use
Real GitHub issues from Python repositories, resolved when the hidden tests pass in a per-instance container.
it-opsunit testsSource reference · integration not included - WebArena812 tasksCoding & computer use
Long-horizon web tasks on self-hosted shopping, forum, GitLab, CMS and map sites with functional checks.
it-opssales-revopsfunctional checksSource reference · integration not included - OSWorld369 tasksCoding & computer use
Computer-use tasks across real desktop apps in full virtual machines, scored by execution-based scripts.
it-opsdata-biexecution scriptsSource reference · integration not included - AEC-Bench196 tasksArchitecture & CAD
Drawing navigation and cross-document project coordination. Curated answers and structured checks; some defects are injected into public project records.
architecturedocument-coordinationstructured checksSource reference · integration not included - Architecture & CAD
Drawing perception, OCR, counting and spatial questions. A drawing-reading benchmark rather than a CAD editing environment.
architecturedrawing-understandingquestion-answer scoringSource reference · integration not included - DrafterBench1,920 tasksArchitecture & CAD
Civil-engineering drawing revision tasks. Grades recorded tool arguments; final CAD geometry is not independently checked by that scoring method.
architecturedrawing-editingtool-call checksSource reference · integration not included - Architecture & CAD
Industrial CadQuery generation, reasoning and targeted editing. 17,900 execution-verified programs across 106 families; unseen-family transfer remains difficult.
cadparametric-editingexecution + geometrySource reference · integration not included - CADBench · MIT DeCoDE18,000 tasksArchitecture & CAD
18,000 evaluation samples across five input modalities. Measures reconstruction, executable programs and compactness. Separate from BenchCAD.
cadreconstructionexecution + geometrySource reference · integration not included - CADEngBench600 tasksArchitecture & CAD
Parametric parts, controlled edits, assemblies and declared physics checks. 600 part tasks; only 164 parts support full L3 FEA. Solver assumptions are benchmark-assigned.
cadstructural-engineeringconstraints + FEASource reference · integration not included - BMEval / AutoBM128 tasksArchitecture & CAD
Simulation-executable building models in OpenSeesPy. Grades execution and period checks against empirical references; passing is not engineering certification.
structural-engineeringexecution + period checksSource reference · integration not included - Architecture & CAD
CAD generation and editing with public inputs and private leaderboard reference artifacts. Official grading requires its hosted evaluation setup.
cadparametric-editingprivate reference gradingSource reference · integration not included