GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Artificial Analysis Launches AA-Briefcase: 91-Task Private Benchmark for Agentic Knowledge Work

Artificial Analysis has launched AA-Briefcase, a private agentic benchmark designed to measure frontier model performance on long-horizon professional knowledge work. It is the lab’s second purpose-built agentic evaluation after AA-AgentPerf, and it covers a category that SWE-bench and tau-bench do not: multi-week business workflows with deliverable outputs.

What It Tests

AA-Briefcase comprises four scenarios:

  • Data Science — multi-step analysis workflows producing structured outputs
  • Product Management — decision documentation, roadmap artifacts, stakeholder memos
  • Banking Operations — financial workflow execution across structured datasets
  • Heavy Industry / Corporate Strategy — complex strategy deliverables across extended timelines

Each scenario is structured as a sequence of weeks, with each week holding up to five tasks. In total, the benchmark covers 91 tasks across thousands of input files. Grading uses a rubric of checks applied against each deliverable. Models currently complete each task independently — prior submissions do not carry over between weeks, a constraint Artificial Analysis flags as an open limitation.

Why It Matters

Coding benchmarks like SWE-bench test models on well-defined, verifiable tasks with binary pass/fail signals. Tau-bench tests service-domain agentic loops. AA-Briefcase occupies a different tier: open-ended professional work where the output is a business deliverable — a spreadsheet, a presentation deck, a strategy memo — graded not by whether tests pass but by whether the document meets a rubric.

This is closer to the actual work enterprises are deploying AI agents to do. Most real-world agentic failures are not code execution failures; they are failures of judgment, coherence, and sustained professional quality across many steps.

AA-Briefcase is private, meaning Artificial Analysis runs it rather than accepting self-reported submissions. That eliminates the benchmark saturation problem that has made public leaderboards increasingly unreliable. Results will be published alongside AA Intelligence Index updates as models are evaluated.

Context

The launch follows Artificial Analysis’s overhaul of its Intelligence Index methodology in the AA v4.1 release, which introduced tau3-bench banking as a core component. AA-Briefcase goes further: instead of a single-turn service scenario, it chains tasks across a simulated project lifecycle.

No model scores have been published yet. Artificial Analysis has indicated results will appear alongside future leaderboard updates as evaluations complete.