Scrabble LLM Benchmark

Immediate-score benchmark for model move quality.

Benchmark

Rank models by how close they get to the perfect immediate Scrabble move.

Each run sums model points across fixed benchmark boards and divides by the exact solver total. The benchmark ignores exchange strategy and leave value on purpose.

Total runs32
Completed runs30
Best score94.3%

Token Range

044 K88 K132 K176 K220 KGrok 4.6Kimi K3Grok 4.5Codex CLI (gpt-5.6-sol) [xhigh]Codex CLI (gpt-5.6-terra) [high]Codex CLI (gpt-5.6-sol) [low]GPT-5.5DeepSeek V4 ProGemini 3.1 Pro PreviewGrok 4.20GPT-5.4Codex CLI (gpt-5.6-luna) [low]DeepSeek V3.2GLM 5 TurboQwen3.6 PlusGLM 5.1o3Step 3.7 FlashQwen3.7 MaxDeepSeek V4 FlashGLM 5.2Composer 2.5Owl AlphaNemotron 3 Ultra (free)Qwen3.7 PlusMiniMax M3Step 3.5 FlashKimi K2.6Gemini 2.5 ProKimi K2.6

Average Tokens vs Score

026 K52 K78 K103 K129 K0%25%50%75%100%Average tokensScore
X-
ST
CUOP
ST

Release Timeline

Apr 16, 2025Aug 8, 2025Nov 30, 2025Mar 24, 2026Jul 16, 20260%25%50%75%100%Release dateScore
X-
ST
CUOP
ST