If you’ve ever sat through a training session where the instructor throws a dozen practice problems at you and says “learn by doing,” Sweller would like a word. In the late 1980s, educational psychologist John Sweller was studying how people solve problems — algebra, geometry, the kind of thing that makes most adults flinch. He noticed something strange. When learners used a strategy called means-ends analysis (working backwards from the goal, checking “where am I now vs. where I need to be”), they were so busy managing the problem-solving process that they never actually learned the underlying rules. They could solve the problem in front of them, but give them a slightly different variation and they were lost. The more they practiced, the worse they got at transferring knowledge.
🧠 The Problem-Solving Trap
Sweller’s insight came from watching how novices approach a new domain. Here’s what happens: you see a problem, you work backwards from the desired outcome, you track sub-goals, you monitor progress. All of this eats up working memory — that tiny scratchpad in your brain that can hold maybe four things at once. The problem is, working memory is also where learning happens. You need it to build schemas, to recognize patterns, to connect the new thing to something you already know. When all your working memory is consumed by the mechanics of solving, there’s nothing left for learning.
Sweller called this the worked-example effect. Learners who studied fully solved problems — worked examples — consistently outperformed learners who practiced solving the same problems themselves. The act of solving created what he called extraneous cognitive load: mental effort spent on task-irrelevant processes. Study the example, and every mental resource goes into pattern recognition. Solve it yourself, and half your brain is busy navigating the maze.
🤔 Why “Learn by Doing” Isn’t the Whole Story
This is deeply counterintuitive because “practice makes perfect” is practically a cultural axiom. We assume that struggle builds character and that more reps equal more learning. Sweller showed that for novices, struggle often builds nothing but cognitive exhaustion. The expertise that makes problem-solving valuable — the ability to see the deep structure, to recognize which approach fits — that only comes after you’ve built a schema. And schemas are built by studying examples, not by thrashing through exercises.
There’s a twist, though. The worked-example effect flips as you gain expertise. Once you have a solid mental model, studying worked examples becomes redundant — even boring — and solving your own problems becomes the efficient path. This is called the expertise reversal effect. What’s good for a beginner is bad for an expert, and vice versa. Instructional design that ignores where the learner is on this curve is design that’s working against the brain.
🔗 Why Your UX Has a Cognitive Load Problem
Every interface you design is either reducing or adding to your user’s cognitive load. Think about any digital companion’s chat interface: when the AI returns a dense wall of text with multiple layers of interpretation, the user has to split their working memory between parsing the structure and understanding the meaning. That’s extraneous load — the presentation format is consuming resources that should go toward the experience itself. The fix isn’t “make it shorter.” The fix is chunked presentation: first the headline, then the reasoning, then the nuance. Let the user build a schema for each layer before moving to the next.
The same principle applies to pricing pages, onboarding flows, and feature documentation. Every time you force the user to hold information in one place while looking for related information in another (the split-attention effect), you’re burning their cognitive budget on navigation instead of comprehension. The best interfaces feel effortless because they don’t ask the user to juggle.
🎲 The Magic Number Before the Magic Number
Everyone knows Miller’s Law: we can hold 7 ± 2 items in working memory. But here’s the less-known follow-up. Sweller’s work, combined with later research by Cowan (2001), suggests the real limit is closer to 4 chunks for novel information — especially when you’re also doing something with that information, not just passively holding it. That 7 ± 2 number came from experiments where participants simply recalled digits. In real learning — where you need to process, connect, and apply — the bottleneck is much tighter. Next time you’re designing a step in a workflow that asks the user to consider more than four things at once, pause. Your user’s working memory is already full.