
AA HL papers haven’t gotten harder because the syllabus expanded—they’ve gotten harder because the framing changed. Extended questions now require students to identify which mathematical domain governs a problem before selecting any technique. When a question is pre-labeled by topic, that label makes the domain decision for the student, meaning the recognition step never actually occurs in practice; remove the labels and that step becomes the exam. A candidate in the r/IBO May 2026 exam discussion thread described Paper 1 as “completely different compared to previous years,” moving away from “schematic papers”—difficulty came from interpreting prompts and choosing an approach, not from encountering new content. Other candidates in the same thread pointed specifically to Paper 1 Section B as unusually hard, linking that difficulty to the extended, multi-step structure of the questions rather than any new topic content. That recognition-and-setup step is where 2024 examiner commentary and community observation both identified the most consequential gaps, and it’s what every recommendation in this guide targets.
Where the 2024 Papers Exposed Gaps—and What to Do About Each
Paper 1 exposed two related failures: procedural errors under non-calculator conditions in Section B and incomplete written justification of reasoning steps. The remediation combines a preparation rule with a writing habit—run every Paper 1 practice session without a calculator throughout, and write one justification sentence per reasoning step regardless of whether the answer feels self-evident. Neither adjustment requires new material, only that these conventions are enforced from session one rather than introduced in the final weeks.
Correct numerical answers with no visible analytical setup—GDC output appearing on the page without a written model, variable definitions, or governing equation—were another pattern 2024 examiner commentary specifically flagged. That omission costs method marks regardless of whether the number is right. The fix is a single non-negotiable habit: before any GDC input enters the working, write one setup line naming the variables and the governing equation. Practiced consistently on every extended-question attempt, it becomes automatic under exam pressure.
Paper 3 carry-forward failures amplified single early errors across entire multi-part investigations. A carry-forward signpost—”Using result from part (b)…”—written at the start of each dependent sub-part prevents this and must be practiced as a writing convention in every timed Paper 3 attempt rather than added during review. Separately, 2024 examiner reports identified presentation requirements—including one extra significant figure beyond any given approximation and specific Part B answer-booklet conventions—that should be confirmed against published reports and, once verified, embedded as session-opening checklist items from the first week.

The Core Preparation Adjustment—Building Domain Recognition
Scattered as those 2024 failures appear—missing justification, absent GDC setup, weak carry-forward discipline—they share a common root: students reaching for a procedure before correctly identifying what kind of problem they’re solving. That recognition step, not procedural execution alone, defines expert performance in mathematics. A 2022 review, Developing Problem-Solving Expertise for Word Problems, synthesizes decades of research showing that experts categorize problems by underlying structure rather than surface features like topic label or story context and that conceptual and procedural knowledge reinforce each other—gains in one tend to support gains in the other. The ability to translate an unfamiliar worded scenario into a mathematical formulation early in the solution process is what the 2022 review identifies as one of the most reliable markers distinguishing successful from unsuccessful problem solvers.
Topic-labeled question banks suppress that recognition step at a structural level. When a question arrives pre-tagged, the domain call is already made, not earned under any meaningful conditions. Strip topic headers from an existing bank and treat domain identification as the first deliberate task on every question—no new material required, just a different constraint applied to existing practice.
Apply a brief pre-solve commitment on each unlabeled question: write what it’s asking for (target quantity and constraint), name one or two candidate domains before any algebra begins, then write one line translating the words into mathematics. After marking, log two outcomes separately—domain call (correct or wrong) and execution (clean or sloppy). If domain calls are frequently wrong across the last twenty questions, continue with unlabeled mixed-topic drilling even if procedural accuracy looks fine; if calls are usually right but execution is failing, shift into targeted skill blocks while keeping some mixed sets in rotation.
How to Phase Your Preparation and When to Use Recent Papers
Effective preparation follows three sequential phases: Phase 1 is diagnostic topic work with all Section 2 conventions embedded from session one; Phase 2 is unlabeled mixed-topic drilling to build the domain-recognition layer; Phase 3 is full simulation using recent authenticated papers. Research on deliberate practice—specifically the 2022 review Developing Problem-Solving Expertise for Word Problems—finds that targeted, weakness-focused work is more effective than undirected repetition, which is why IB Math AA HL practice exams from the most recent sessions should be reserved for Phase 3. Two readiness conditions apply: stable topic-level accuracy under timed, non-calculator conditions and consistently correct domain calls in unlabeled sets before technique selection begins. Each paper then functions as a dress rehearsal reviewed against the 2024 error categories—and the value is almost entirely in how you process the attempt, not just in sitting it:
- Sit the paper as a true simulation—timed, with the correct calculator rule and all Section 2 writing conventions active from the first line.
- On the same day (30–45 min), mark each lost mark or point of uncertainty with one primary tag: (1) domain or approach choice, (2) setup or model line missing or wrong, (3) algebra or arithmetic slip, (4) notation or communication, or (5) time and pacing.
- Within 24 hours (45–60 min), identify your three highest-cost tags and run a short mixed mini-block built only around those fixes—after each question, stop and state the correction rule explicitly before moving on.
- Within 72 hours (15–30 min), re-test one representative item per fix, cold. Decision rule: if the same tag reappears twice in the re-test, it stays on next week’s Fix List; if it does not, replace it with the next-highest-cost tag from the simulation.
- Weekly cadence: one full simulation per week in Phase 3; everything else that week is driven by the Fix List produced by that simulation, not by what feels familiar. If a markscheme or reliable worked solution is unavailable, tag uncertainty points instead of lost marks, resolve them via a trusted solution, then build the Fix List from those tags—the loop is less precise but still produces directed practice.
Paper 3—What Changed and How to Prepare
The five-step loop applies to Papers 1 and 2 directly. Paper 3 adds a dimension those papers don’t test, and that dimension needs its own preparation logic. Drilling archived investigation scenarios can push students toward surface familiarity with past contexts rather than the underlying structures those tasks share. As the 2022 review Developing Problem-Solving Expertise for Word Problems notes, expert problem solvers group problems by deep principles rather than story details, so familiarity with specific previous scenarios often won’t transfer when a new investigation frames the mathematics in a genuinely novel context. What the opening scenario of a Paper 3 investigation tests is not pattern-matching against a memorized scenario library but productive engagement with an unfamiliar structure—the ability to form a working conjecture when the mathematical context itself is new.
Building that capacity means supplementing archived investigations with additional exploration problems to ensure exposure to unfamiliar structures alongside recognizable ones. Practice the opening-scenario reading and conjecture-formation stages as timed, discrete skills: give yourself a fixed window to form a creditable conjecture before proceeding to the guided sub-questions. Apply the carry-forward signpost from Section 2 in every timed attempt. The target is flexible schemas—conceptual hooks paired with usable procedures, as the 2022 research on problem-solving expertise suggests—rather than a memorized library of investigation templates. This work belongs in Phase 2, running alongside unlabeled mixed-topic drilling rather than being saved for Phase 3; the two practices reinforce each other before full simulations begin.
A Twelve-Week Preparation Calendar
The three phases translate into a twelve-week arc. Weeks 1–4 (Phase 1): diagnostic topic work with all Section 2 conventions active from session one—written justification, GDC setup protocol, carry-forward signpost, and the significant-figure rule. Weeks 5–9 (Phase 2): unlabeled mixed-topic sessions as the primary practice mode, with Paper 3 exploration running alongside archived investigations. Weeks 10–12 (Phase 3): full authenticated-paper simulations reviewed against the 2024 error categories, not a generic accuracy percentage.
The calendar is a phase-transition sequence, not a volume prescription. Phase 2 is ready when topic-level accuracy holds without a calculator across the Phase 1 material; Phase 3 is ready when domain calls in unlabeled sets are consistently correct before technique selection begins. A student who hits those marks ahead of schedule should move, not wait for the week number.
Phase 3 weeks are structured by the simulation loop: tag lost marks by primary cause, build a Top 3 Fix List, run a targeted mini-block within 24 hours, and re-test within 72 hours to confirm each weakness is actually closing. The loop—not the paper count—keeps Phase 3 honest. Doing more papers is only useful if each one tells you something different about what to fix next.