RSS Amplifier

Learn Through Stories · Jul 13, 2026

I Tested 885 Duolingo Chess Puzzles. 24% Do Not Make Sense.

0
Sign in to vote or save

Tyler Schwartz · Learn Through Stories

Oscar told me to ignore his rook on c1.

There is no rook on c1.

Oscar, buddy. I’ve stared at this board for five minutes. I’ve counted my rooks. I even considered that maybe, from Oscar’s side of the board, some other square was secretly c1. There is no universe in which I have a rook on c1.

The more I did Duolingo Chess, the more issues I noticed.

Share

I’m a chess nerd, professionally. My first job out of college was teaching chess in a preschool. I turned my method into a company, Story Time Chess, which won Toy of the Year in 2021 and 2025. I once beat US Chess Champion Hikaru Nakamura, which I mention at every possible opportunity. I will mention it three more times in this blog. Consider yourself warned.

So when Duolingo, one of the biggest education companies on Earth, launched a chess course in June 2025, I was thrilled. Teaching chess to total beginners is my life’s work, and now the biggest name in learning apps was going to do it at a scale I could only dream about. I dove in immediately.

And a lot of it is great. The gamification is excellent, the pacing between lessons feels right, and some of the puzzles are so good I screenshotted them just to admire them later.

But, I kept hitting puzzles that felt... off. Lessons that didn’t teach what their titles promised. Solutions a beginner couldn’t reasonably find. And Oscar confidently describing pieces that did not exist.

I wanted to investigate further. Did I mention, I love doing chess puzzles?

I went through Section 7 of Duolingo Chess, all 885 puzzles, one by one, and evaluated each on two questions:

  1. Does the puzzle teach what it claims to teach? If the lesson says “Skewers,” is a skewer actually the point of the solution?

  2. Could a beginner realistically solve it? Duolingo’s audience isn’t tournament players. It’s people learning chess on the bus.

I sorted every puzzle into Effective, Not Effective, or Borderline. When a puzzle didn’t work, I coded why, using a rubric of 14 issue types I built as patterns emerged. Every puzzle got a difficulty rating and an expected success rate. Yes, there’s a spreadsheet. Yes, it has multiple tabs. I told you I was a nerd.

  • 672 (76%) are effective. Three out of four puzzles do their job.

  • 178 (20%) are not effective.

  • 35 (4%) are borderline.

That’s 213 puzzles, 24% of the section, with problems.

Now the painful part. Let’s do some math. Duolingo reports 7 million daily active chess users. Assume just 10% of them are in Section 7 on any given day, it’s the final section, where every learner eventually lands and stays. That’s 700,000 people. Say each one does a modest 10 puzzles a day. That’s 7 million Section 7 puzzles served daily, and if 24% have problems, that’s roughly 1.7 million broken puzzles handed to unsuspecting beginners. Per day. Over 600 million a year, from one section of one course. Cut my assumptions in half if you like. It’s still 400,000 a day.

What is wrong with these puzzles? I coded 181 issues across those puzzles, spanning 14 issue types, and just two issue types account for nearly two-thirds of everything wrong.

That’s not 181 random problems. That’s two problems, repeating. And that’s great news, because two problems can be fixed.

Imagine a Where’s Waldo book where, on some of the pages, Waldo just isn’t there. Nobody tells you this. You search the beach scene for twenty minutes. You check every stripe, every hat. Nothing.

Here’s the sad consequence: you don’t conclude the book is defective. You conclude you’re bad at finding Waldo. A kid will sit there feeling dumber and dumber, hunting for a man who was never on the page.

That’s Theme Not Present, and it happens 60 times in Section 7. Duolingo tells you to find the fork, and there’s no fork on the page.

My favorite example is from a lesson called “Disadvantage Promotions.” (I’ve taught chess for over a decade and have never heard this term, but let’s roll with it.) Oscar tells you to try to promote a pawn. Small problem: White doesn’t have a single pawn past the third rank. Oscar, do you know how pawns move?

The actual solution is a queen sacrifice that sets up a knight fork. It’s a great puzzle! It’s just filed in the wrong drawer. Promotion never happens, never threatens to happen, and was never the point.

For the chess players: 1.Qxe7 Qxe7 2.Ng6+ Kh7 3.Nxe7 winning a piece. A clean attraction-into-fork combo. Not a pawn promotion in sight.

Good beginner puzzles run on forcing moves. I check you, you must respond. I take your piece, you must take back.

A Branching Variation puzzle has quiet moves. Moves that don’t force the opponent’s response. Because of that, they’re really difficult to calculate. This happens 57 times.

Here’s a glaring example, from 7.1.3.

I got this puzzle wrong when I did it. I came back weeks later for this screenshot and got it wrong again and I have beaten Hikaru Nakamura (that’s one).

The problem: White’s first move is a quiet pawn capture. Nothing is forced, so Black has at least four reasonable replies, and each one is its own puzzle. You’re expected to have solved all of them before touching a piece. Even White’s second move only works because of a hidden bishop trick most masters would need a minute to spot.

For the chess players: 1.hxg4 Nf5, but ...Qd6, ...f6, and ...Ng6 all demand their own calculation. Then 2.Qh5 is the only move, since 2.Qh3 runs into 2...Nh6, unveiling the c8 bishop onto the queen. This puzzle starts a move too early.

From my notes, verbatim: “no way, 2nd move not forced, no one getting this, bad puzzle.” Beginner success rate, my honest estimate: under 5%. This is an expert puzzle wearing a beginner costume.

The remaining third of the issues spread across 12 categories, some of which are honestly pretty funny:

  • Weird AI Puzzle (8): positions that feel machine-generated and human-unreviewed. One puzzle asks you to win a queen. The solution is: you take the queen. It’s just sitting there. Black responds with two pointless moves and the puzzle ends. I studied this position for several minutes convinced I was missing something. I was not. I once outplayed Hikaru Nakamura and this puzzle nearly got me (that’s two).

  • Confusing Character Prompt (8): Oscar’s hints contradicting the board, including my beloved phantom rook on c1.

  • Opaque Payoff (19): you play the right moves, win the position, and have no idea why you won. Beginners need the payoff to be visible: material, or mate.

  • Plus a long tail: puzzles that were too easy, too long, out of order, dependent on untaught endgame theory, one with a genuine red herring, and exactly one puzzle that was just plain bugged.

Every good puzzle tells a story. There’s a threat, a response, a resolution. The moves aren’t just legal, they’re motivated. That story is the signal. Everything else that happens along the way is noise.

My theory: Duolingo has a shallow tagging system. It knows the result of a puzzle, but not its narrative.

A shallow tagger sees “queen captured” and files the puzzle under Win the Queen. But sometimes Black throws the queen incidentally, a last-resort move to delay a mate threat. The mate threat is the signal. The queen is noise. Duolingo tagging can’t tell the difference, so the learner is told to hunt the queen and misses the mate driving every move. That’s TNP, 60 times over.

Same blindness explains BV. The tagger confirms a winning line exists, but “a win exists” is a result, not a narrative. It says nothing about whether a beginner can follow the path. The engine says solvable. A teacher says solvable by whom?

If you only know results, you can’t know where a puzzle belongs, not its theme, not its difficulty. One shallow tagger, two symptoms, 117 of my 181 coded issues.

And here’s the part Duolingo should lose sleep over: it poisons their own data. When 24% of puzzles have validity problems, a low completion rate stops meaning anything. Hard puzzle, or broken puzzle? The analytics can’t tell, so every difficulty calibration and A/B test built on this section is partly optimizing static. Fixing the puzzles isn’t just good pedagogy, it’s data hygiene.

The concentration of errors is the gift here. You don’t rebuild 885 puzzles. You build one thing: an agent that doesn’t just play chess but understands it. An engine tells you the best move. This agent tells you the story, the signal driving the solution, not the noise around it. A queen falling as a throwaway? Not a queen-winning puzzle. Four unforced branches? Not a beginner puzzle.

And here’s the beautiful part: the training data already exists. Lichess maintains an open database of millions of puzzles, tagged by theme and refined by the votes of actual humans who solved them. Benchmark the agent there until its labels match human consensus, then point it at Duolingo’s catalog. Every puzzle where the agent’s theme disagrees with the label gets flagged for human review. You’re not reviewing 885 puzzles anymore, you’re reviewing the flags.

That approach catches TNP and BV, two-thirds of everything I found. Add one human curation pass for the phantom rooks, and Section 7 goes from 76% effective to the mid-90s without designing a single new puzzle.

I want to be clear about where I stand: Duolingo Chess is the best thing to happen to beginner chess since The Queen’s Gambit (aside from lichess.org). Millions of people who would never touch a chess book are doing tactics puzzles on their phones because a green owl motivated them into it. That is a miracle and I am fully in favor of it. Thank you Oscar!

The gamification is world-class, like Hikaru Nakamura’s chess rating, whom I beat (Boom goes the dynamite, that’s three). The content just needs to match it. And unlike most 24%-failure-rate headlines, this one comes with an unusually cheap fix, because the failures aren’t random. They’re systematic, and systematic means solvable.

I have 885 rows of data, a 14-category issue rubric, and a suggested fix for every single category. If anyone in Pittsburgh wants them, my inbox is open.

Thanks for reading Learn Through Stories! This post is public so feel free to share it.

Share

No posts

Read the original on learnthroughstories.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.