Mayer's 12 Principles of Multimedia Learning (With Examples)
By Devlin Peck · Updated
Part of the Instructional Design Fundamentals guide
Mayer's 12 principles of multimedia learning are research-backed guidelines for combining words, visuals, and audio so people actually learn from them. They come from psychologist Richard Mayer, whose lab has run hundreds of controlled experiments on how people learn from words and pictures. If you build eLearning, slides, or training videos, then these principles are the closest thing our field has to a set of dos and don'ts backed by data.
I use these principles to guide a huge number of the decisions I make when creating eLearning, and they sit near the core of instructional design fundamentals. They remove guesswork. Instead of debating whether background music "adds energy," you can point to the experiment that says it hurts learning and move on.
This guide covers all 12 principles with eLearning examples, the violations I see most often, how strong the evidence is for each principle, the three newer principles Mayer added in 2021, and whether any of this still holds up in AI-powered learning experiences.
If you prefer video, my walkthrough below covers the principles in condensed form (it groups the two contiguity principles together and skips a couple of the less-studied ones):
What are Mayer's 12 principles of multimedia learning?
The 12 principles are evidence-based rules for designing multimedia lessons, developed by Richard Mayer from his cognitive theory of multimedia learning. The theory rests on three assumptions about how people learn. First, we process words and pictures through two separate channels (auditory and visual). Second, each channel has limited capacity. Third, learning happens when people actively select, organize, and integrate information, not when they passively receive it.
Every principle follows from those three assumptions. Here are the 12 principles from Mayer's Multimedia Learning at a glance. This table summarizes Mayer's framework; the sections below add the examples, violations, and exceptions.
| Principle | What it says | Do this instead |
|---|---|---|
| Coherence | Extraneous material hurts learning | Cut background music, decorative images, and fun-but-irrelevant facts |
| Signaling | Cues that highlight key material help learning | Use headings, arrows, highlights, and vocal emphasis on essentials |
| Redundancy | Narration plus identical on-screen text overloads learners | Pair narration with visuals, not with a transcript on screen |
| Spatial contiguity | Words belong near the graphics they describe | Put labels on the diagram & feedback next to the answer |
| Temporal contiguity | Spoken words and graphics belong together in time | Sync narration to the animation step it describes |
| Segmenting | Learner-paced chunks beat one continuous stream | Break lessons into short segments with a continue button |
| Pre-training | Knowing key terms first frees capacity for the lesson | Teach names and characteristics of key concepts up front |
| Modality | Narration beats on-screen text for explaining graphics | Describe complex visuals with audio, not paragraphs |
| Multimedia | Words plus relevant pictures beat words alone | Replace the bullet wall with a diagram |
| Personalization | Conversational style beats formal style | Write in first and second person ("you," "we") |
| Voice | Human voice beat machine voice in the original studies | Use a friendly human voice (more on AI voices below) |
| Image | The speaker's static image on screen doesn't necessarily help | Don't add a talking head just to fill space |
One point of confusion worth settling: you'll see sources claim 7, 12, 14, or 15 principles, and they're all describing the same body of work at different stages. The first edition of Multimedia Learning (2001) presented 7 principles, the second edition (2009) expanded to the 12 above, and the third edition (2021) added three more for 15. Lists claiming 14 are usually partial updates caught between editions.
These are specifically multimedia design principles, by the way. They tell you how to present words and pictures, not how to analyze a performance problem or structure a curriculum. For the wider set, see the broader principles of instructional design.
Mayer groups the principles under three goals, which is how I've organized the rest of this guide: reduce extraneous processing, manage essential processing, and foster generative processing.
Which principles reduce extraneous processing?
Five principles cut the material and layout choices that waste learners' limited processing capacity: coherence, signaling, redundancy, spatial contiguity, and temporal contiguity. Extraneous processing is mental effort spent on things that don't contribute to learning, and it's the easiest problem to fix because the fix is usually deletion.
Coherence principle
People learn better when extraneous words, pictures, and sounds are excluded. This is the "less is more" principle, and it's one of the better-studied principles in the literature per the 2022 systematic review in Smart Learning Environments.
The classic violations: background music under narration, decorative stock photos that relate to the topic only in mood, and interesting-but-irrelevant tangents ("fun fact!"). Each one competes for the same limited capacity your essential content needs. If an element doesn't support the learning objective, cut it. Your course will feel less "produced" and teach more.
Signaling principle
People learn better when cues highlight the organization and location of essential material. Arrows, highlights, bolded key terms, clear headings, and vocal emphasis all tell the learner where to spend their attention.
Two cautions here. First, cueing everything cues nothing. If half the screen is highlighted, the highlight carries no information. Second, cues cut both ways, including ones you didn't design. Back in 2021, while building an interactive 360° scene in Storyline, I noticed the cursor turned into a pointer whenever it passed over a hotspot. That subtle change gave away exactly which parts of the scene were interactive, which I hadn't intended. Everything on screen cues something. Audit the cues your authoring tool adds for you, not just the ones you place deliberately, and decide which ones serve the experience.
Redundancy principle
People learn better from graphics plus narration than from graphics plus narration plus identical on-screen text. Reading and listening to the same words simultaneously forces the learner to reconcile two streams of the same content, which adds cognitive load instead of reinforcement.
This is the principle I see violated most often, by new and experienced designers alike. The pattern is always the same: the designer writes a script, pastes it onto the slide, and then narrates it word for word. Some people defend this as an accessibility measure, but vision-impaired learners have screen readers for exactly this reason, and voice-acted audio can be offered as an opt-in for learners who need it rather than forced on everyone by default.
I'll admit I violate this principle to a degree myself. When I want to plan and record a YouTube video quickly, I put bullet points on slides that double as cue cards for me and visual references for the audience. With more production time, those videos would use more visuals and less text. Which is the practical lesson: most redundancy violations are production shortcuts, not design decisions. You should know when you're making the trade.
Spatial contiguity principle
People learn better when printed words appear near the graphics they describe, rather than far from them. Every time a learner's eyes travel between a diagram and a separate legend, they're spending capacity on visual search instead of learning.
In eLearning this means labels directly on the diagram instead of a key below it, feedback that appears right next to the question the learner just answered, and tooltips anchored to the element they explain. If the learner has to scan to connect a word with its picture, move the word.
Temporal contiguity principle
People learn better when corresponding narration and animation happen at the same time rather than one after the other. "Let me explain the process, and then I'll show you" splits one integrated idea into two disconnected presentations.
The fix is synchronization: as the narration describes step three, show the animation for step three. My video above treats this and spatial contiguity as a single "contiguity principle," which is a fair shorthand, but they're distinct in Mayer's framework and worth checking separately, because a course can pass one and fail the other.
Which principles manage essential processing?
Three principles help learners handle material that is inherently complex: segmenting, pre-training, and modality. You can't delete essential complexity the way you delete background music, but you can portion it out so it never exceeds the learner's capacity.
Segmenting principle
People learn better when a lesson comes in learner-paced segments rather than one continuous unit. The key word is learner-paced: segmenting isn't just chopping content shorter, it's giving the learner a 'Continue' button so they process each chunk before the next one arrives.
Instead of a 20-minute screencast of an entire workflow, break it into short segments at natural boundaries and let the learner advance when ready. This is also why a well-designed click-through course can beat a video of the same content: the pacing belongs to the learner.
Pre-training principle
People learn better from a complex lesson when they already know the names and characteristics of its key concepts. If a learner finds five unfamiliar terms inside an already-difficult explanation, they're learning vocabulary and process at the same time (and likey doing both badly).
For example: a short "learn the terms" screen or intro module before the complex scenario, a glossary the learner works through first, or a labeled diagram of the system before the animation of how it works would all work well.
Modality principle
People learn better from graphics plus narration than from graphics plus on-screen text, especially when the material is complex and fast-paced. Explaining a visual with audio spreads the work across both channels; explaining it with a paragraph piles everything onto the visual channel.
Worth knowing: modality is the most studied principle in the entire literature according to the 2022 Smart Learning Environments review, so this one rests on a deep evidence base. It also has real boundary conditions. When the material is full of unfamiliar technical terms, when learners aren't native speakers of the narration language, or when learners control the pacing and can re-read, then on-screen text can work fine.
Which principles foster generative processing?
Four principles encourage learners to actively make sense of the material: multimedia, personalization, voice, and image. Reducing load isn't enough if the learner doesn't engage; these principles are about motivating the mental work that produces learning.
Multimedia principle
People learn better from words and pictures than from words alone. This is the founding principle and the reason the field is called multimedia learning. A relevant diagram, chart, or animation gives the learner a second representation to integrate with the words, and that integration is where understanding forms.
The operative word is relevant. A process deserves a flowchart, a comparison deserves a table, a piece of equipment deserves a labeled photo. A slide of bullet points with a stock photo of smiling coworkers totally unnecessary though (and fails coherence at the same time).
Personalization principle
People learn better when words are in a conversational style rather than a formal style. "Your lungs move oxygen into your blood" beats "The lungs facilitate oxygen transfer" because first- and second-person language makes learners treat the lesson as a conversation partner worth listening to.
This one shaped my very first client project. It was an anti-phishing course, and I was nervous, but I knew how to write learning objectives and I knew these multimedia principles. So instead of a dry "here's what you're going to learn" open, I got creative and had a phisher hacker character present the content himself. That choice was personalization and a bit of embodiment (more on that principle below) doing real work in a beginner project. Conversational tone is also the cheapest principle to apply. If your course copy sounds like a policy manual, start with writing conversationally for eLearning.
Voice principle
People learn better when narration comes from a friendly human voice rather than a machine voice. That was the finding in the original studies, anyway, and it made obvious sense at the time: early text-to-speech sounded robotic and grating.
Modern AI voices are a different story, and the original research predates them entirely. Per the 2022 systematic review, voice is also one of the two least-studied principles (along with pre-training), so the evidence base here was thin even before AI voices arrived. Treat this one as a live question. I'll give you my current read in the AI section below.
Image principle
People do not necessarily learn better when the speaker's static image is on the screen. This is the weakest, most conditional principle in the set: a talking-head box or presenter photo doesn't reliably help, and it can pull attention from the content it's supposed to support.
The practical takeaway isn't "never show a face." It's that a face is not automatically valuable, so it has to earn its screen space like everything else. A presenter demonstrating a physical task earns it; a floating headshot next to a chart usually doesn't.
How strong is the evidence for each principle?
The principles are research-backed, but unevenly: some rest on dozens of experiments and others on a handful. This matters because the principles get taught as commandments when they're better treated as strong defaults with known exceptions.
Here's what the research base actually looks like:
| Finding | What the research says | Source |
|---|---|---|
| Overall evidence base | The current 15 principles are grounded in 200+ experimental comparisons from Mayer's research program | Multimedia Learning, 3rd ed. (Cambridge University Press, 2021) |
| Most-studied principles | Modality leads the literature, followed by redundancy, multimedia, signaling, and coherence | 2022 systematic review of 136 journal articles, Smart Learning Environments |
| Least-studied principles | Pre-training and voice have the thinnest evidence base of the twelve | Same 2022 Smart Learning Environments review |
| New environments | Studies testing the principles in VR and AR remain limited; most evidence comes from traditional screen-based lessons | Same 2022 Smart Learning Environments review |
| Theoretical organization | All 15 principles map to three goals: reduce extraneous, manage essential, foster generative processing | Mayer (2024), Educational Psychology Review |
| Boundary conditions | Effects often shrink or disappear for learners with high prior knowledge; each principle carries documented exceptions | Multimedia Learning, 3rd ed. (Cambridge, 2021) |
Here are two takeaways: first, if you're going to argue with a stakeholder over one of these principles, then modality, redundancy, and coherence are the hills worth dying on; voice and pre-training deserve more humility. Second, the most consistent boundary condition across principles is prior knowledge: experts can handle (and sometimes benefit from) presentations that would overload novices. A principle "failing" with your expert audience isn't evidence the research is wrong. It's the research working as documented.
If you want to see how this framework sits alongside the field's other big ideas, I compare it with other instructional design theories and models in a separate article.
What new principles did Mayer add after the original 12?
The third edition of Multimedia Learning (2021) expands the set to 15, adding the embodiment, immersion, and generative activity principles. Most articles on this topic still stop at 12, but these three are important, especially if you work with video, VR, or interactive practice.
Embodiment principle
People learn better when on-screen instructors display human-like gesture, movement, and eye contact. A static cartoon narrator adds little; an agent that gestures toward the diagram it's explaining acts as a social partner and a signaling device at once.
Immersion principle
People do not necessarily learn better in immersive 3D virtual reality than from a 2D desktop version of the same lesson. Immersion is engaging, but engagement isn't learning, and the extra realism can add extraneous load. If you're pitching VR, then pitch it for what genuinely needs spatial presence (e.g. physical procedures or spatial layouts).
Generative activity principle
People learn better when guided through generative activities during learning: summarizing, self-explaining, drawing, teaching back, or enacting. Of the three additions, this one has the biggest implications for course design, because it shifts the question from "how do I present this?" to "what will the learner do with it?" It's also the natural bridge to the next section.
Do Mayer's principles apply to AI-powered learning?
Yes. The principles describe how human cognition processes words and pictures, not any particular delivery technology, so they apply to AI-generated and AI-driven experiences too. But some principles need re-examination, and the research is racing to catch up: the 2022 systematic review found studies in VR and AR environments still limited, so anything newer than that is extrapolation territory.
Start with voice. The original voice principle compared human narration to early machine synthesis, and modern AI voices have closed most of that quality gap. My expectation is that updated studies will show the gap between human and AI-delivered voice on learning outcomes is very small. AI-generated faces are the more interesting question. My hypothesis is that learning suffers with AI avatars because of their uncanny quality, which is unfortunate given how popular they've become for rapid content creation in the corporate space. If I'm right, the embodiment principle cuts against the current avatar trend, not in favor of it. Watch for the studies; until then, I'd bet on AI voice and stay cautious with AI faces.
The bigger AI story isn't production speed, though. Most of the AI conversation in L&D right now is about making the same stuff faster: faster storyboards, faster scripts, faster slide builds. I get the appeal, but I think the real opportunity is that AI makes entirely new categories of learning experience possible. Consider the generative activity principle. The best generative activity for a conversation skill is having the conversation, and before AI, the closest we could get in self-paced eLearning was the multiple-choice branching scenario or live facilitator time. Those were the best tools we had, but they were never realistic: nobody facing an angry customer on the job has three options floating in front of them.
AI conversation simulation closes this gap, and it's what I work on now. My company devlin.ai builds AI text and voice conversation simulations that embed in Storyline, Rise, or any LMS, with the AI's evaluation and scores reporting back to the course. The classic principles still work inside them: coherence (no extraneous fluff in the scenario setup), segmenting (one exchange at a time, paced by the learner), pre-training (brief the learner on the situation before the conversation starts), and personalization (a conversation is conversational by definition).
Here's what this looks like in practice. Try a short text-based simulation where you play a manager coaching an underperforming employee:
How do you apply Mayer's principles in your next project?
Run every screen through a quick three-pass review that mirrors the three cognitive-load goals: cut, then structure, then engage. You don't need to memorize 15 principles to do this; you need three questions and the discipline to ask them before publishing.
Cut: what can I remove?
Apply coherence, redundancy, and image. Delete background music, decorative visuals, tangents, on-screen text that duplicates narration, and any face that isn't earning its space. Deletion is the highest-leverage edit in eLearning.
Structure: is the essential material easy to process?
Apply signaling, both contiguity principles, segmenting, pre-training, and modality. Cue the essentials (and audit accidental cues), put words next to their graphics, sync narration with animation, chunk with learner-paced controls, front-load key terms, and let audio carry explanations of complex visuals.
Engage: does the learner have to think?
Apply multimedia, personalization, voice, embodiment, and generative activity. Pair words with relevant visuals, write conversationally, use a warm voice, and build in at least one moment where the learner produces something: an answer, a summary, a decision, a conversation.
There's a career payoff to this discipline, too. A lot of instructional designers today still violate these principles, so if you can speak fluently about them and use them to justify your design decisions, you will stand out in portfolio reviews, interviews, and stakeholder conversations alike.
If you can speak fluently about Mayer's principles and use them to justify your decisions, you are going to be gold.Devlin Peck
This is why these principles are part of the foundational theory we teach at Peck Academy, my licensed career school for people transitioning into instructional design. And whether or not you ever take a course, get the applied companion to all of this research: e-Learning and the Science of Instruction by Ruth Clark and Richard Mayer, now in its 5th edition. It translates the lab findings into working guidance for course designers, and it has a permanent spot on my list of the best eLearning books.
Frequently asked questions
Which of Mayer's principles is the most important?
Mayer doesn't rank them, but modality has the deepest research base (it's the most-studied principle per the 2022 Smart Learning Environments review), and in practice the biggest wins usually come from the cutting principles: coherence and redundancy. Most real-world courses fail by including too much, so start there.
Are there 12 or 15 principles of multimedia learning?
Both numbers are correct for different editions. The second edition of Mayer's Multimedia Learning (2009) presented 12 principles, and the third edition (2021) added embodiment, immersion, and generative activity for a total of 15. The first edition (2001) had 7.
Do Mayer's principles apply to PowerPoint presentations and video?
Yes. The principles apply to any medium that combines words and pictures, including slides, explainer videos, eLearning courses, and job aids. A slide deck benefits from coherence, signaling, and redundancy checks exactly the way a Storyline course does.
When is it OK to break one of Mayer's principles?
When a documented boundary condition or a real constraint applies: accessibility needs (captions), learners with high prior knowledge, non-native speakers who benefit from on-screen text, or production limits you've weighed deliberately. Break principles on purpose and be able to say why, never by default.
What theory are the principles based on?
Mayer's cognitive theory of multimedia learning, which holds that people process words and pictures in two limited-capacity channels and learn by actively selecting, organizing, and integrating information. Every principle is a practical consequence of those three assumptions.