Briefcase AI

The Open-Weight LLM Benchmark Suite

Which downloadable-weights model to run, and what it costs. Four instruments over one 28-model roster, scored 0–5 from public benchmarks and fact-checked — read the maps, or just describe the job and get a stack.

Start here
Stack Recommender

Describe your use case in plain language. Get a recommended open-weight stack — a primary model, a budget workhorse, an escalation tier — and a cost estimate per task, not per token, with editable per-model pricing.

Open the recommender →
Map 01
Domain Proficiency

28 models across 18 professional domains — legal, medical, finance, software, data, safety, license — with per-cell evidence and cited benchmarks.

Open map →
Map 02
Agentic Tasks

The same roster as agent cores across 20 common service-desk tasks, tiered L1 routine, L2 escalated, L3 expert. Tool use and reliability, not chat knowledge.

Open map →
Map 03
Knowledge vs Agentic

One roster, two maps, side by side. The best knowers are not the best doers — a scatter and twin per-model profiles show the divergence.

Open view →