Claude Code Benchmark

Latest internal run: secret-scanner CLI. Use the Distill.codes CLI to repeat a paired local comparison.

Fable 5 xhigh

Same task, same prompt, same model, separate fresh directories. Both runs produced implementations that passed the task smoke verifier.

RouteVerifiedTimeSource LOCOutput tokensEst. Claude cost
Directyes945.0s47972,088$6.16
Distill.codesyes343.8s24924,876$2.97

Main reductions

63.6%

faster

65.5%

fewer output tokens

48.0%

less source LOC

51.8%

lower estimated provider cost

Supporting clean result

Fable 5 and Opus 5 are recommended for evaluation. Sonnet 5 showed a smaller and unstable effect in our runs.

74.9%

faster

76.5%

fewer output tokens

62.2%

less source LOC

Methodology

Distill.codes does not guarantee smaller diffs, lower token use, or faster runs on every task. Short or simple tasks may show little difference; longer and more complex tasks tend to reveal workflow waste more clearly.

Run it locally

Use your proxy URL to run direct and Distill.codes Claude Code tasks in separate local folders. The CLI saves the report locally.

npx distill-codes bench "https://proxy.distill.codes/<proxy-key>/essential/anthropic"

Try it on your workload

Use the trial to compare review time, output tokens, files touched, LOC, and whether the final patch still satisfies your task.

Choose plan and start 3-day trial