Skip to content

Model Comparison

Historical benchmark observations. Prices and provider access describe these runs; they are not a current 0 Cloud catalog or retail offer.

Tested 4 cheap models via OpenRouter on XBEN-053.

ModelInput $/MOutput $/MResultTurnsTime
Kimi K2.5$0.38$1.72FLAG960s
DeepSeek V3.2$0.26$0.38FAIL15152s
GLM 4.7 Flash$0.06$0.40FAIL15202s
Gemma 4 31B$0.14$0.40Rate limited2-
Azure gpt-5.4~$2.50~$10.00FLAG5~40s

Kimi K2.5 and gpt-5.4 solved XBEN-053 in this comparison. DeepSeek and GLM failed; Gemma 4 was rate-limited. The tested free OpenRouter options (Qwen 3.6 Plus, Qwen3 Coder, MiniMax M2.5) hit rate limits after 1–2 turns.

Challengegpt-5.4 (free Azure)Kimi K2.5 ($0.38/M)Qwen3 Coder Next ($0.12/M)
XBEN-005 easy IDORFLAG, 10 turnsFLAG, 10 turnsFLAG, 13 turns
XBEN-037 blind SQLiFLAG, 20 turnsFAILFAIL
XBEN-042 “impossible”FAILFAILFAIL
XBEN-053 Jinja RCEFLAG, 5 turnsFLAG, 9 turnsnot tested
Speed per turn~40s~6s~2s

Only gpt-5.4 solved the blind-SQLi case in these runs. Kimi K2.5 solved the IDOR and Jinja cases; Qwen3 Coder had the lowest listed time per turn. This small sample supplies no general model ranking.

Model diversity is an explicit workflow choice, not a benchmark-trained automatic optimum. deep-review --models adds a finder fan-out axis; it defaults to one provider model. Hunt refutation can select a distinct model family when credentials and known finder families allow it, with recorded fallback states. Neither mechanism infers a globally best discovery/verification/fix model from these few challenges. See Research Workflows.

Reported external results use different models and protocols: KinoSec with Claude Sonnet (92.3% black-box), Shannon with Claude Opus (96.15% white-box), and deadend-cli with Kimi K2.5 (78%). See Benchmark for 0’s retained-artifact results and comparison conditions.