Scrabble LLM Benchmark

Immediate-score benchmark for model move quality.

Benchmark

Rank models by how close they get to the perfect immediate Scrabble move.

Each run sums model points across fixed benchmark boards and divides by the exact solver total. The benchmark ignores exchange strategy and leave value on purpose.

Total runs36
Completed runs32
Best score94.3%

Token Range

044 K88 K132 K176 K220 KGrok 4.6Kimi K3Grok 4.5DeepSeek V4 Pro 0813Codex CLI (gpt-5.6-sol) [xhigh]Codex CLI (gpt-5.6-terra) [high]Codex CLI (gpt-5.6-sol) [low]GPT-5.5DeepSeek V4 ProGemini 3.1 Pro PreviewGrok 4.20GPT-5.4Codex CLI (gpt-5.6-luna) [low]DeepSeek V3.2GLM 5 TurboQwen3.6 PlusGLM 5.1o3Step 3.7 FlashQwen3.7 MaxDeepSeek V4 Flash 0731DeepSeek V4 FlashGLM 5.2Composer 2.5Owl AlphaNemotron 3 Ultra (free)Qwen3.7 PlusMiniMax M3Step 3.5 FlashKimi K2.6Gemini 2.5 ProKimi K2.6

Average Tokens vs Score

026 K52 K78 K103 K129 K0%25%50%75%100%Average tokensScore
X-
ST
CUOP
ST

Release Timeline

Apr 16, 2025Aug 15, 2025Dec 14, 2025Apr 13, 2026Aug 12, 20260%25%50%75%100%Release dateScore
X-
ST
CUOP
ST