A ACT GAME PREPBuild skill. Keep momentum.
SAT prepACT — currently selected

MA6 Statistics & Probability

Lesson

Data questions — a mean, a table, a probability, a count, a line of best fit — are a steady presence on the ACT Math section, and they are among the quickest points on it. They are also thrown away in the same few ways every year: a median read off an unsorted list, a probability divided by the wrong total, a count that was added when it should have been multiplied.

The four questions that decide almost every item here: Is the list sorted? Is the denominator the whole group or a subgroup? Does the question want both things (multiply) or either thing (add)? Does order matter (permutation) or not (combination)?

Sort first. Every single time.

The median is the middle value of the ordered list, so ordering the list is step one, not an optional tidy-up. ACT writes these lists deliberately out of order, because the middle number as printed is a number — it just is not the median. With an even count there is no single middle value, so the median is the average of the two middle ones, and it need not be a value in the data at all.

Mean and median answer different questions. The mean is the total shared out equally, so every value affects it: change one value and the total changes, so the mean changes. The median is a position, so it moves only when the ordering around the middle moves. That is what people mean when they say the median resists extreme values: in the list 4, 6, 7, 9, 40, pushing the 40 up to 400 leaves the median at 7, because the middle position of the sorted list has not changed — while the mean jumps from 13.2 to 85.2. It is not that the median ignores large values; it is that it only cares where they sit in the order.

A mean problem is a total problem

Whenever a question gives you an average and asks for anything else, turn the average back into a total: total = mean × how many. That one move handles the three shapes ACT actually asks.

  • What do I need on the last test? Five tests averaging 88 means 440 points in total. Subtract what you already have; what is left is the score you need. If that comes out above 100, the target is not reachable — which is sometimes the point of the question.
  • Adding or removing a value. Rebuild the total, adjust it, then divide by the new count. Forgetting to change the count is the single most common error in this family.
  • Weighted averages. You cannot average two averages when the groups are different sizes. 12 students averaging 78 and 18 averaging 88 gives (12×78 + 18×88) ÷ 30 = 84, not 83. The bigger group pulls harder.

Two questions decide every probability item

First: what is the denominator? A probability is favourable outcomes over the group the selection is actually made from. "One of the 150 students" means 150 on the bottom. "One of the seniors" means the senior total on the bottom — the same numerator over a different whole. That single switch is what a two-way-table question is testing, and it is why a conditional probability and a joint probability built from the same cell are different numbers.

Second: are the events independent? Independent means knowing one happened does not change the probability of the other. Two spins, two coin flips, two draws with replacement: independent. Two draws without replacement: not independent — after the first marble leaves the bag, both the favourable count and the total have dropped by one. And a two-way table almost never shows independent events; when the table prints the count of people in both categories, use that count instead of multiplying.

Add or multiply?

  • "And", both, in a row → multiply. P(A and B) = P(A) × P(B) when the events are independent; otherwise the second factor is the probability of B given A.
  • "Or", either → add. If the two outcomes cannot both happen (one marble is not both red and blue), just add them. If they can both happen, add and then subtract the overlap once: P(A or B) = P(A) + P(B) − P(A and B), because everyone in both groups has been counted twice.
  • "At least one" → use the complement. "At least one" covers many cases; "none" covers one. Compute P(none) as a product and subtract from 1.
  • Sanity check: requiring both makes an event rarer, so a multiplication answer must be smaller than either probability you started with. If your "both" answer is bigger, you added.
The trap ACT sets most often: the same cell of a table with three different denominators. From a table of 150 students, 24 of the 80 juniors are in band, and 55 students in all are in band. The chance a random student is a junior in band is 24/150. The chance a random junior is in band is 24/80. The chance a random band member is a junior is 24/55. Three correct answers to three different questions, and the only thing that tells them apart is the phrase naming the group being chosen from. Underline that phrase before you compute anything.

Counting: order is the whole question

If choices are made from separate categories — a bread, a filling, a drink — multiply the counts. That is the fundamental counting principle, and it is most of what the ACT asks.

When you are choosing a group from one pool, ask whether rearranging the chosen people gives a different outcome. Gold, silver and bronze are ranked, so ABC and CBA are different results: that is a permutation, 8 × 7 × 6 = 336 ways from eight runners. A three-person committee is not ranked, so ABC and CBA are the same committee: that is a combination, and you divide the 336 by the 6 orders of any three people to get 56. Same start, different question, and whenever you are choosing more than one thing, the permutation count is the larger of the two.

Expected value

Expected value is an average weighted by probability: multiply each outcome by its probability and add the products. Do not divide at the end — the probabilities already add to 1. If a game charges a fee, the fee is paid on every play, win or lose, so subtract it once from the expected winnings. The result is a long-run average per play and usually is not any prize actually available.

Samples, correlation and lines of best fit

  • A random sample supports an estimate about the population it was drawn from — and about no wider group. A random sample of one school's students says something about that school and nothing about the state. A sample also supports an estimate, not an exact count: another random sample would give a slightly different figure.
  • A sample that is not random cannot be fixed by being large. Volunteers, whoever showed up first, whoever stayed late: each over-collects a particular kind of person, and adding more of them keeps the same tilt.
  • A correlation says two measurements move together, and stops there. It does not establish that one causes the other. In observational data a third variable may drive both, and the same numbers are equally consistent with the causal arrow pointing the other way. The defensible choice is the one phrased as a tendency — "tend to", "is associated with", "on average". A choice that says one thing causes the other is the trap unless the study actually assigned people to groups.
  • Reading a line of best fit y = mx + b: m is the predicted change in y for each additional unit of x; b is the predicted y when x is 0. Swapping their roles is the standard wrong answer.

Predicting inside the range of the data you fit is interpolation, and it is what a line of best fit is for: observed values sit on both sides of your estimate. Predicting well outside that range is extrapolation, and it is the weakest thing a line of best fit can do — it assumes a straight-line pattern continues where nothing whatsoever was measured. A plant measured for ten weeks will not typically keep growing at that same rate for a year — growth levels off. When a question asks which prediction is least reliable, look for the one furthest outside the range of the data.

💾 Create a free account (or log in) to save your XP, streak, and progress across devices.