MYP 4 Mathematics · Statistics and Probability

Data Manipulation and Misinterpretation

Get started

What Is Data Manipulation?

Data Manipulation

The process of changing, adjusting, or selectively presenting data in a way that influences how people interpret the information — sometimes honestly (for clarity) and sometimes dishonestly (to mislead).

Data is everywhere: in news articles, advertisements, social media posts, scientific studies, and political campaigns. But not all data presentations tell the full story. Data manipulation occurs when someone — intentionally or unintentionally — presents statistics in a way that creates a misleading impression.

This doesn't always mean someone is lying. Sometimes data is misrepresented because the person analysing it made a poor choice of statistical measure. Other times, it's deliberately engineered to persuade you of something that isn't quite true.

Analogy

Imagine taking a photograph of a messy room, but only from the one clean corner. The photo is real — it's not photoshopped — but it gives a completely false impression. Data manipulation works the same way: the numbers can be real, but the way they're presented can be deeply misleading.

As an IB MYP student, developing critical thinking about data is one of the most important skills you can build. Let's explore how data can be manipulated and how you can spot it.

Choosing the Wrong Average: Mean, Median, or Mode

One of the most common ways data is misinterpreted is through the choice of average (measure of central tendency). The three main averages — mean, median, and mode — can give very different impressions of the same data set.

Mean

The sum of all values divided by the number of values. It is commonly used but is affected by extreme values (outliers).

Median

The middle value when data is arranged in order. It is not affected by extreme values, making it more robust for skewed data.

Mode

The most frequently occurring value in a data set. It is useful for categorical data but may not represent the overall distribution.

Example

Salary data at a small company:

Six employees earn: 32,000, 36,000, 250,000

  • Mean = 70,500$
  • Median = 35,500$
  • Mode = No mode (all values are unique)

The CEO might say: *"The average salary at our company is 35,500 is far more representative of what most employees actually earn.

This is a classic example of how choosing the wrong average can mislead.

Warning

Common Misconception: Students often think "the average" always means the mean. In fact, mean, median, and mode are ALL types of average. When someone says "the average is..." without specifying which one, be suspicious — they may have chosen the one that best supports their argument.

Free preview

10 more sections in this topic

Pick this up in your Library: it holds the whole topic, notes, cheatsheet and questions.