Artificial Analysis Launches AA-Briefcase: 91-Task Private Benchmark for Agentic Knowledge Work
Artificial Analysis has launched AA-Briefcase, a private agentic benchmark designed to measure frontier model performance on long-horizon professional knowledge work. It is the lab’s second purpose-built agentic evaluation after AA-AgentPerf, and it covers a category that SWE-bench and tau-bench do not: multi-week business workflows with deliverable outputs.
What It Tests
AA-Briefcase comprises four scenarios:
- Data Science — multi-step analysis workflows producing structured outputs
- Product Management — decision documentation, roadmap artifacts, stakeholder memos
- Banking Operations — financial workflow execution across structured datasets
- Heavy Industry / Corporate Strategy — complex strategy deliverables across extended timelines
Each scenario is structured as a sequence of weeks, with each week holding up to five tasks. In total, the benchmark covers 91 tasks across thousands of input files. Grading uses a rubric of checks applied against each deliverable. Models currently complete each task independently — prior submissions do not carry over between weeks, a constraint Artificial Analysis flags as an open limitation.
Why It Matters
Coding benchmarks like SWE-bench test models on well-defined, verifiable tasks with binary pass/fail signals. Tau-bench tests service-domain agentic loops. AA-Briefcase occupies a different tier: open-ended professional work where the output is a business deliverable — a spreadsheet, a presentation deck, a strategy memo — graded not by whether tests pass but by whether the document meets a rubric.
This is closer to the actual work enterprises are deploying AI agents to do. Most real-world agentic failures are not code execution failures; they are failures of judgment, coherence, and sustained professional quality across many steps.
AA-Briefcase is private, meaning Artificial Analysis runs it rather than accepting self-reported submissions. That eliminates the benchmark saturation problem that has made public leaderboards increasingly unreliable. Results will be published alongside AA Intelligence Index updates as models are evaluated.
Context
The launch follows Artificial Analysis’s overhaul of its Intelligence Index methodology in the AA v4.1 release, which introduced tau3-bench banking as a core component. AA-Briefcase goes further: instead of a single-turn service scenario, it chains tasks across a simulated project lifecycle.
No model scores have been published yet. Artificial Analysis has indicated results will appear alongside future leaderboard updates as evaluations complete.