The Puzzle Dataset
Published on September 23, 2026
While there are millions of chess puzzles available online, offering an effective and structured training experience requires careful curation and a deep understanding of puzzle difficulty. In this post, I want to shed some light on the dataset that powers ChessKinetics, how the puzzles are picked, and who the platform is built for.
The Target Group and Rating Limits
When you first initialize your account, you'll notice that the selectable rating limits range from 900 to 2700. This specific range defines the target audience of the app.
When selecting your initial rating, you might be tempted to use your Lichess puzzle rating. However, Lichess puzzle ratings often sit significantly higher than actual playing strength. It is highly recommended to start a bit lower, using your FIDE rating (or a conservative estimate of it) rather than an inflated online puzzle projection. This ensures a smoother progression curve as the algorithm calibrates to your real calculation ability.
Why Aren't the Rating Limits Broader?
You might wonder why the platform doesn't cater to ratings below 900 or above 2700. The answer lies in data quality and the mechanics of the wave algorithm.
To provide a consistent experience, the system relies on distinct "buckets" spanning 50 rating points each. We need every single one of these rating buckets to contain a large enough pool of high-quality puzzles without introducing bias. Pushing the limits too far in either direction makes it difficult to maintain this strict quality control.
Furthermore, the wave progression requires "space" from the minimum and maximum absolute puzzle ratings. Since the difficulty curve pushes you slightly above your baseline target, we need a buffer of challenging puzzles beyond the 2700 selectable limit, as well as easier warm-up puzzles below the 900 limit.
How the Puzzles Are Picked and Rated
The foundational dataset comes from the open Lichess puzzle database.
The rating of each puzzle is based on how it performs against players on Lichess. However, a puzzle's rating alone doesn't guarantee it's a good training tool. Some puzzles are confusing, have multiple seemingly winning lines, or rely on computer-like moves that are uninstructive.
To solve this, the dataset was filtered. Puzzles with a high rating deviation, meaning players of the same strength have wildly inconsistent results solving them, were completely removed. The remaining pool was further narrowed down by setting minimum thresholds for popularity and the number of times they had been played to ensure a baseline of quality. Finally, from this qualified pool 2,000 puzzles were selected entirely at random for each 50-point rating bucket. This random selection is crucial because it prevents any bias that would be introduced if the system simply cherry-picked the absolute most popular or highest-rated puzzles.
Dataset Size and Repetition
Through this strict curation process, the final dataset was refined down to 96,000 puzzles.
A common concern with puzzle platforms is memorization: Will I see the same puzzle twice in the same session or day?
The answer is No. Due to the bucket structure, you are guaranteed not to reach the exact same puzzle again until you have played at least 2,000 other challenges. In practice, depending on how your rating fluctuates, it will likely take several thousands of challenges before a puzzle repeats, making memorization virtually impossible.
Will the Dataset Change?
Currently, there is no immediate plan to change or expand the available puzzles. The curated puzzles provide a massive, balanced, and rigorously tested training ground. This stability ensures that your progress through the waves is an accurate reflection of your improving skill, rather than fluctuations in the underlying dataset.