We gave Claude Code the same seven Slack tasks through Slack's hosted MCP server and through this CLI, 35 runs each. With the changes described below, the CLI answered every task correctly at a third less cost per task.
33% lower
mean cost per task: $0.033 with the CLI, $0.050 with MCP
60% less
tool output read by the model: 4,144 vs 10,466 characters per task
35 / 35
tasks answered correctly by each arm in the final run
Final run, 7 tasks × 5 repeats per arm, Claude Sonnet 5 in Claude Code. Costs are the
total_cost_usd Claude Code reports for each session.
This is our own benchmark of our own tool, on one workspace. The harness is in the repository so you can run it against yours; the caveats say where the numbers are weak.
Slack's hosted MCP server at mcp.slack.com, connected with the OAuth client
from Slack's own Claude Code plugin. Its read and search tools are available; its write
tools are denied.
The slack CLI and the skill that slack skills add installs.
The agent runs it from a shell. Commands that write to Slack are denied.
claude -p) on Claude
Sonnet 5, in an empty directory with no user plugins or other MCP servers. The harness
stops if an arm sees the other arm's tools.
| Task (median of 5) | CLI cost | MCP cost | CLI context | MCP context | CLI turns | MCP turns |
|---|---|---|---|---|---|---|
| Search for an update, extract facts | $0.049 | $0.054 | 101,872 | 81,134 | 7 | 5 |
| Answer from a support thread | $0.024 | $0.127 | 62,572 | 72,362 | 5 | 4 |
| Look up a person by job title | $0.020 | $0.026 | 61,402 | 60,577 | 5 | 4 |
| Read a message from its link | $0.014 | $0.019 | 45,660 | 44,360 | 4 | 3 |
| Two-hop thread question | $0.032 | $0.066 | 78,937 | 118,788 | 6 | 6 |
| Read a bot deployment notice | $0.045 | $0.036 | 71,228 | 67,338 | 5 | 4 |
| Summarize a window of a channel | $0.024 | $0.026 | 47,050 | 68,594 | 4 | 4 |
The CLI costs less on six of seven tasks. MCP costs less reading the bot's deployment notice: the notice keeps some fields only in its attachment, and the CLI skill tells the agent to read such a message unfiltered.
The CLI as released in 0.3.0 cost $0.057 per task, more than MCP's $0.049 in the same run. The changes in 0.4.0, described below, brought it to $0.033.
| Mean per task | CLI 0.3.0 | CLI 0.4.0 | Slack MCP |
|---|---|---|---|
| Cost | $0.057 | $0.033 | $0.050 |
| Context tokens | 109,876 | 66,827 | 77,037 |
| Output tokens | 987 | 617 | 754 |
| Tool output (characters) | 10,553 | 4,144 | 10,466 |
| Model turns | 8.0 | 5.1 | 4.5 |
| Tool calls that errored (all 35 runs) | 15 | 0 | 4 |
| Passed | 35 / 35 | 35 / 35 | 35 / 35 |
CLI 0.3.0 is from the first run; the other columns are from the final run. Context tokens count everything the model read across its turns, cached or not. Fixed overhead is the same for both arms: the "OK" baseline used 13,546 context tokens with the CLI and 13,726 with MCP.
Most of an agent's Slack cost is the tool output it reads back. In 0.4.0:
--filter-output
field lists for search, history, and threads; 53 of 68 read commands in the final run used
one.
--json prints one line when piped to a
program.
--oldest 2026-09-24T21:00, UTC). In 0.3.0, 13 of 15 failed tool calls were date math with the wrong
platform's flags.
MCP's search tools produced 84% of its tool output in the final run, as formatted text the agent cannot trim. In exchange, MCP messages name their authors; the CLI returns user ids.
The harness is
evals/ab
in the repository. Authorize the MCP arm once, sign in with the CLI, and write a task file
with prompts and the facts each answer must contain:
# Once: authenticate the MCP arm with /mcp inside this session $ claude --strict-mcp-config --mcp-config evals/ab/slack-mcp.json $ slack login $ node evals/ab/run.ts --tasks evals/ab/tasks/mine.local.json --repeats 5 $ node evals/ab/report.ts evals/ab/results/<run>/runs.jsonl
Files named *.local.json and all results are ignored by git, because tasks and
transcripts contain your workspace's messages. The report prints medians per task and arm
and counts denied commands and calls to the other arm's tools, so a skewed run is visible.