The Science · Memory and learning
Dual Coding: Why Words Plus an Image Beats Either Alone
Allan Paivio's theory explains why a quip floating over a meaningful photograph sticks in a way that a Post-it never does. The science of two channels at once.
May 29, 2026 · 8 min read
In 1971, a Canadian psychologist named Allan Paivio published a book called Imagery and Verbal Processes. It was a quiet event. The book was technical, the claims were modest in tone, and the broader field of cognitive psychology was still working out whether mental images were even a respectable thing to study. Behaviorism had spent forty years arguing that internal pictures were a fiction, or at least not the business of science. Paivio disagreed, politely and at length, and ended up proposing one of the few ideas in memory research that has survived essentially intact for more than fifty years. He called it dual coding.
The claim is simple enough to fit on an index card. The mind has two parallel systems for handling information. One is verbal and processes language. The other is nonverbal and processes images, sounds, and other sensory traces. The two systems run in parallel and can refer to each other, but they are distinct. Information encoded in only one channel is fragile. Information encoded in both channels is roughly twice as likely to be found again later. That redundancy is not waste. It is what memory looks like when it is working.
Image slot
Two channels, one memory: words and image arriving in parallel.
Image prompt: Editorial illustration on warm cream paper. Two horizontal channels flow from the left edge toward a single point on the right where they converge into a small glowing icon labeled 'memory'. The top channel is labeled 'verbal' and shows a faint stream of typeset words. The bottom channel is labeled 'nonverbal' and shows a faint stream of small images: a tree, a bird, a face. Subtle paper grain. Stone-50 background palette, muted ink lines.
What Paivio actually proposed
Paivio gave his two systems precise names. The verbal system was made of units he called logogens, modeled loosely on John Morton's earlier work on word recognition. A logogen fires when you hear or read a word, and it carries connections to other logogens through grammar, semantics, and association. The nonverbal system was made of imagens, which fire when you perceive an image, hear a non-speech sound, or call up a sensory memory. The two systems are linked by what Paivio called referential connections, the cross-links that let the word dog evoke the picture of a dog and let the picture evoke the word.
The architecture was tidy enough to be testable, and Paivio spent the next thirty years testing it. The standard finding, replicated in dozens of studies and reaffirmed in his 1986 book Mental Representations, was that words rated as easy to picture (apple, hammer, sunset) were remembered far better than words rated as abstract (justice, freedom, tendency). The advantage was not subtle. Concrete words were often recalled at twice the rate of abstract ones, and the gap held across recognition, recall, paired associates, and free recall tasks. The dual-channel account predicted exactly this. A concrete word got encoded twice, once verbally and once as a picture. An abstract word got encoded once.
The picture superiority effect
The dual-coding story would be elegant on its own, but it gained a lot of force from a parallel literature about pictures themselves. The single most striking result is Lionel Standing's 1973 paper, charmingly titled Learning 10,000 Pictures. Standing showed subjects ten thousand pictures, one after another, for a few seconds each. Later he asked them to discriminate the ones they had seen from new ones they had not. Recognition accuracy was around 83 percent. Ten thousand pictures. Ten thousand. The comparable rate for words was substantially lower.
Standing's result was not a one-off. Nelson, Reed, and Walling in 1976 carefully dissected what was happening and showed that the advantage was not just about visual distinctiveness. Pictures were remembered better than the same items presented as words, and the gap widened the longer the delay between study and test. The effect was named the picture superiority effect and has been replicated so consistently that it is now a default assumption in textbook treatments of memory. The mechanism most often proposed is the dual-coding one. A picture invites both an image trace and an automatic verbal label. A word, on its own, mostly just gets the verbal label.
Memory for concrete items is superior to memory for abstract items because concrete items activate both the verbal and the nonverbal symbolic systems, while abstract items activate only the verbal.
Mayer's multimedia principles
Paivio set the theory in motion. Richard Mayer, working at UC Santa Barbara from the 1990s onward, turned it into a design discipline. Mayer's research program, summed up in The Cambridge Handbook of Multimedia Learning (2005, with updated editions since), ran hundreds of experiments on a single question. When you are trying to teach someone something, when do words and pictures work together, and when do they get in each other's way? The answers came out as a set of principles, each backed by replications.
Three principles bear directly on the case for putting a quote over a photograph. The multimedia principle states the headline finding: people learn better from words and pictures together than from words alone. The contiguity principle adds an important constraint: the words and the picture have to be presented close together in space and time, not on separate pages or with a delay between them. The redundancy principle warns about the failure mode: if you also pile on a third channel that duplicates one of the first two (say, a narration that reads the on-screen text aloud), you start to overload working memory and learning gets worse, not better. Two channels, well composed, is the sweet spot. Three is often a mess.
Why working memory matters here
The architecture under all of this is John Sweller's cognitive load theory, which dates to a series of papers starting in 1988. Sweller's starting point was a fact that Paivio took for granted: working memory is small. Most estimates put it at something like four chunks held active at once, in the manner of George Miller's famous 1956 paper. Anything you want to learn has to pass through that bottleneck before it can land in long-term memory. The bottleneck is the choke point of most teaching.
Sweller's contribution was to notice that the working memory bottleneck is not a single channel. It is at least two. There is a verbal/auditory channel and a visual/spatial channel, and each has its own small capacity. If you try to load four things into the verbal channel you will probably lose one. If you load two into the verbal channel and two into the visual channel, you may keep them all. Dual coding is, in Sweller's framing, an architectural feature of the brain that smart instructional design can take advantage of. It is also why a phrase laid over a meaningful photograph is more compact, in working memory terms, than the same phrase shown alone on a card. The phrase loads one channel. The photograph loads the other. Together they cost less than you would expect.
Image slot
The same six words, on a blank card and on a meaningful photograph.
Image prompt: A two-panel comparison on warm cream paper. Left panel: a small index card with the typeset phrase 'today is a good day to begin' centered on plain white. Caption beneath: 'one channel'. Right panel: the same phrase rendered over a quiet photograph of a fog-lit field at dawn, the typography sitting comfortably in the upper-left third. Caption: 'two channels'. Soft paper grain, muted ink. Both panels at the same size for clean comparison.
The blank-card problem
If you keep all of this in mind, a familiar object starts to look strange: the motivational poster, or the Post-it stuck to the bathroom mirror, or the quote-of-the-day calendar with a blank background. They are all single-channel objects. They give the verbal system something to do and leave the nonverbal system idle. The dual-coding account predicts, and the empirical literature largely confirms, that these objects are not very memorable. People walk past their own affirmations every morning and could not recite them at the end of the week.
The fix is not to write better affirmations. The fix is to give the second channel something to hold. A quote rendered over a photograph that genuinely belongs to the quote, not a stock image bolted on, becomes a dual-coded object. The verbal system processes the words. The nonverbal system processes the scene. The two traces land together, with the contiguity principle satisfied, and the chance of finding the memory again later goes up substantially. The poster does not just look nicer. It works differently.
What makes a pairing actually dual-coded
Not every image-plus-text combination triggers the effect. Mayer's contiguity principle is the first constraint: the picture and the words have to share a frame and a moment. The second constraint is that the picture has to be referentially connected to the words, in Paivio's sense. A photograph of a mountain trail paired with a line about persistence works. The same line laid over a stock photograph of a model laughing at a salad does not, because the verbal and nonverbal traces do not refer to each other and so the cross-links never form. The image becomes decoration rather than encoding, and decoration is, from the memory system's point of view, almost noise.
The third constraint is calmness. If the image is so busy that the words become hard to read, the redundancy principle starts to bite and the whole composition costs more in working memory than it pays back. Dual coding works best when the two channels are clear, legible, and clearly related, and when neither is asked to fight the other for attention. That is partly a typography problem, partly an image-selection problem, and partly a surface problem. A scrolling phone is bad at all three. A still frame on a wall is good at all three.
Further reading
The primary sources are worth tracking down. Allan Paivio, Imagery and Verbal Processes, 1971, for the original statement of the theory; Mental Representations: A Dual Coding Approach, 1986, for the mature version. Lionel Standing's 1973 paper Learning 10,000 Pictures in the Quarterly Journal of Experimental Psychology, for the picture superiority effect in its most vivid form, and Nelson, Reed, and Walling, 1976, for the careful follow-up. Richard Mayer's Cambridge Handbook of Multimedia Learning, 2005, with later editions, for the design principles. John Sweller's 1988 paper Cognitive Load During Problem Solving in Cognitive Science, for the architecture of working memory that makes the rest of it make sense. Read in that order, they tell one coherent story about why two channels remember what one channel forgets.
Related reads
How to Actually Remember the Things You Read
Dual coding is one mechanism. Retrieval, generation, interleaving, and elaboration are the others. The complete map.
Vision Boards That Don't Just Sit Still
Dual coding is also the reason vision boards feel so promising. Why most still fail, and what would actually work.
Why E-Paper Feels Different From a Screen
Dual coding needs a surface that can carry both image and text in a calm, ambient way. The physics that makes that possible.
Bring it home
A quiet device that uses what the research already knows.
QuipCast is a reflective e-paper frame that cues short, meaningful content on the cadence you set. Launches summer 2026. Waitlist signups get founder pricing and a three-year price lock.
Join the waitlist