AriseKit

Benchmark Tasks

The full list of 117 tasks involved in building a clinical LLM benchmark, tagged by whether AriseKit can automate the work, serve as a tool for a human decision, provide a scaffold of templates, or sits outside our scope.

Flag distribution
A
Automate
AriseKit does it end-to-end
30 25.6%
T
Tool
AriseKit helps a human decide faster/better
69 59.0%
S
Scaffold
AriseKit ships templates/checklists
8 6.8%
N
Not us
Outside AriseKit scope
10 8.5%
Per-phase density
Phase A T S N Total Mix
1. Conception 91 10
2. Scope 12 12
3. Data 3726 18
4. Items 1061 17
5. Rubric 515 20
6. Panel 2632 13
7. Eval Run 64 10
8. Release 221 5
9. Governance 51 6
10. Maintenance 231 6

All tasks

A · Automate T · Tool S · Scaffold N · Not us