Redactle
Compare model performance across Redactle evaluation configurations without revealing the test articles.
About Redactle
3 Gemini 3.8 Flash · medium 12 12 / 0 / 0 100% 0.0 $0.006 7s
20 DeepSeek V4 Pro 0813 · low 12 2 / 1 / 9 25% 45.1 $0.015 62s
21 DeepSeek V4 Flash 0731 · low 12 1 / 0 / 11 8% 55.1 $0.001 266s
22 Gemini 2.5 Flash Lite · none 12 0 / 0 / 12 0% — $0.006 87s
Each label is a complete model configuration. Lower and farther left is better; the shaded quadrant is below the median on both measures.
A model must make Redactle word guesses until every meaningful title word is revealed. Scores measure accepted guesses above the article title's word-count par, so a model that reveals every title word with exactly one guess per title word scores 0. In the hint-enabled evaluations, every hint adds 50 points.
The 500-word evaluation stops after 30 accepted guesses. If a model is still solving at that point, the attempt receives a score of 60 guesses. That is a cost-saving shortcut, not an estimate of how many guesses the model would have needed to finish. Other configurations use the same rule at twice their own guess limit.
Results from incomplete configurations remain visible but are not ranked. Article titles, excerpts, and per-article outcomes are deliberately omitted so this page does not spoil Redactle puzzles.
Discussion
Sign in to join the discussion.
Loading comments…
Tagged
Alternatives to Redactle
Tools in the same space, ranked by how they are performing in the directory.
Otter
Experimental JS runtime with ability to run 1k+ threads - gi-dellav/otter
/Developer Tools- Subscription/Developer Tools
- Open source/Productivity
Loofah
Free, open-source meeting notes for macOS with private on-device transcription, Markdown storage, and optional bring-your-own AI—no accounts, subscriptions, or…
Open source/Developer Tools- /Developer Tools
Saccade
Closed-loop browser control runtime for AI agents.
/Developer Tools
