Collecting and Organising Data

Website: Young Education
Kurs: Statistics
Buch: Collecting and Organising Data
Gedruckt von: Invitado
Datum: Freitag, 25. September 2026, 01:53

1. Types of Data

Learning outcomes
  • I can distinguish between qualitative and quantitative data.
  • I can distinguish between discrete and continuous variables.
  • I can identify levels of measurement.
  • I can classify data appropriately.
  • I can select appropriate methods for recording data.

Introduction

Data is collected whenever we make observations, conduct experiments, carry out surveys, or record measurements. Scientists, mathematicians, businesses, governments, and researchers all rely on data to answer questions, solve problems, and make informed decisions. However, not all data is the same. Different types of data require different methods of collection, organisation, and analysis.

Understanding the different types of data helps us choose the most appropriate way to record information, display results, and draw conclusions. In this lesson, you will learn how to classify data as qualitative or quantitative, distinguish between discrete and continuous variables, recognise the levels of measurement, and select suitable methods for recording different kinds of data.


What Is Data?

Data is a collection of facts, observations, measurements, or information gathered for analysis.

Data can come from:

  • Scientific experiments.
  • Surveys.
  • Observations.
  • Measurements.
  • Questionnaires.
  • Sensors and instruments.

Before analysing data, it is important to identify what type of data has been collected.


https://images.openai.com/static-rsc-4/FXYfIBVdWAw5z1r--Rutj8Zb2L3F45bo17Y4ulk6R7utEBBGmKEjjw11WMC00mSCsQ-OmgoE9vq2c6sY-shZpWBPuC0hx4sxp3t3CsVPLreH5b3ps6ubeFnMsymFsXsMgTxSSCbCDLri5F_vv9XHZwH9EtfyLJb9MA2XsNJ8FupqzzddiK2c0tFArVaQ5225?purpose=fullsize
 
https://images.openai.com/static-rsc-4/9bblNAwtSvwc2kc8dmZIICy7D7TFaYzTnIGvQHX7b_gGQiOjbDK0JXdsIADQFebVG2fUVmQeetaHbpXHkQV8_DyKhNWxEemWOUzS-rYxLRlPJeaNGYD0_MA7E3mtB55blijyZKEx9mvwj_QlG5X2oYPjgOdE43dfHzlRcY3cRBbwAKWtIhP8XCyF5gWCGI_L?purpose=fullsize
 
https://images.openai.com/static-rsc-4/ViBfHEeyi7wZz98mdaTxQ1dZsFIqL-dgId8Hj0pIhNuQXxuJ2BghqwnJeW5g3VlsMWkTbtulW7g5ap1Mm7CzhtLIB3Ukx7fZ5rbqvGT2PyETEnK6i3Xm05pzkhTf257qFPCGu4Z-7ngZiPOnRWaIU835kbf2pz9rgNLxtc9Q_FwSnRs-h1zgqOcUxi_kGpal?purpose=fullsize
5

Figure 1. Data can be collected in many different forms depending on the investigation.


Qualitative Data

Qualitative data describes qualities or characteristics.

It is usually non-numerical.

Examples include:

  • Eye colour.
  • Favourite sport.
  • Type of pet.
  • Weather description (sunny, cloudy, rainy).
  • Blood type.

Qualitative data is often grouped into categories.


Quantitative Data

Quantitative data consists of numbers that represent measurements or counts.

Examples include:

  • Height.
  • Age.
  • Temperature.
  • Number of siblings.
  • Test scores.

Quantitative data can be analysed using mathematical and statistical methods.


Comparing Qualitative and Quantitative Data

Qualitative Data Quantitative Data
Describes qualities Measures quantities
Usually non-numerical.    Numerical
Categories Counts or measurements
Example: Eye colour Example: Height

Both types of data are useful depending on the question being investigated.


https://images.openai.com/static-rsc-4/kqtVfaCOn8DQdPAFuZ9BvboQUOPFkQ7q49FKmsht8kjQd1zIgWkh_F2sluKvC-sj1wWsTwbVXFBqT2ehxJSKhbcHXomlWrRlwF8DGkRkP6gzKeZW1ywSFmZ8Fhg5Etw1gzoIN7jKcVbA-CfhXSEJ4lj1kAW5dRi6r0V_g0vimdYZ6-HWxZnIwzyh4Qs2UtGL?purpose=fullsize
 
https://images.openai.com/static-rsc-4/UyLM9hrZgyeThdjnZ25HsKZ8m7OYVih8d5NQXtrx2rZYFfMXxtsMGASgEW6hBQtTVwjyZjYdYhLh3r_zRI-vwhacvdQgGqrXxWzFFY1Y3LI0Pd-mi0ZnQ8lEGIjwu0oQ_Ev3CDoyoYFst1aud9gZUSC2TK4u2yU9VEqb1dyfzFnKJNPPlwmGGv9srtsfZCvb?purpose=fullsize
 
https://images.openai.com/static-rsc-4/IcqF6Zehf5-_yEwl3tTuN0Y-9RUNPMt3j2ER6KGexpoAfe0g5tOQ-WRdoKJ7VWGX-kt4tRmhF7Ks8jTI9YvuJ34AOnVdLY_IffuX36KmyrEdax-tbZgj8q2BvJ1g46IeqM799NeYb7cTEMtSmGYLO823-OUVcq6c7gdDDu8VqoBw2UrgxZaWScsm0Pf7NLSf?purpose=fullsize
4

Figure 2. Qualitative data describes categories, while quantitative data measures quantities.


Discrete Variables

A discrete variable can take only specific, separate values.

It usually involves counting.

Examples:

  • Number of students in a class.
  • Number of books on a shelf.
  • Goals scored in a football match.
  • Number of pets owned.

Discrete values are usually whole numbers.

You cannot have 3.7 students or 5.2 bicycles.


Continuous Variables

A continuous variable can take any value within a range.

It usually involves measurement.

Examples:

  • Height.
  • Mass.
  • Time.
  • Temperature.
  • Distance.

Continuous variables may include decimal values.

For example:

  • 1.62 m
  • 25.4 °C
  • 3.78 kg

Comparing Discrete and Continuous Variables

Discrete Variable Continuous Variable
Counted Measured
Separate values Any value within a range
Usually whole numbers Often decimals
Example: Number of cars.  Example: Speed

Understanding this difference helps determine which graphs and statistical methods are appropriate.


https://images.openai.com/static-rsc-4/n9QN5I1J2Hj-RX8akYRJr-ec06vaHbBFl9kXRqPoqAvQdZcBXeWXoUmb8kCauWFvDfT9neTK2h-qJ9yGA-NCeKuuLHM8JQIabzxaeG9qcC77YJC0ig0kvwNlw908i98UPWR9BFV5hYWXmOhvbFFqcXeWgfyLlJO3rZbky0r4tMpnviBpTx1F-Ba4u1YJE5Z1?purpose=fullsize
 
https://images.openai.com/static-rsc-4/kYP09-K6_uA46STKUzH8GYQ62Y33Sp0hEmRazROUxYHnQGwprJLIxbRLyUbM-CxJxmGBVj2IUPnoK3x0X_B7aRjWCCsFNGkFtJ0TtbS8DVSPoKfg3RDZtGk8m36XTbNrixtCnRRCVJ5IV8K-DMIQrIJ-CfpGsZ0dtr_BV_fIfR-uLlbIrxENainY-Kd1RTq0?purpose=fullsize
 
https://images.openai.com/static-rsc-4/hf_3jRtU2RlAYJ1VGECFGV7dt_zgxyAMJY1-pjTFUSLy-fa5OF5AIihpV3Ljem4T3YNj6nImP0yTP8WLkUXsJyaypqEnKxJFN-NTru7s9p5fHlXUUatggca-PDDk1a8tmBuBTmegUlemRigAHa8kCD9FZCCBBrQDUCepiJ7w9asfJ9edEWz7_N8KXOsjK3DW?purpose=fullsize
5

Figure 3. Discrete variables are counted, while continuous variables are measured.


Levels of Measurement

Data can also be classified according to its level of measurement.

There are four commonly recognised levels.


Nominal

Data consists of categories with no natural order.

Examples:

  • Eye colour.
  • Favourite fruit.
  • Nationality.
  • Blood type.

Ordinal

Categories have a meaningful order, but the differences between them are not necessarily equal.

Examples:

  • Small, medium, large.
  • Class rankings.
  • Customer satisfaction ratings.
  • Letter grades.

Interval

Numerical data with equal intervals, but no true zero.

Examples:

  • Temperature in degrees Celsius.
  • Calendar years.

A temperature of 20 °C is not "twice as hot" as 10 °C because the zero point is arbitrary.


Ratio

Numerical data with equal intervals and a true zero.

Examples:

  • Height.
  • Mass.
  • Time.
  • Distance.
  • Age.

Because ratio data has a true zero, meaningful comparisons such as "twice as much" can be made.


Summary of the Levels of Measurement

Level Ordered?    Equal Intervals?    True Zero?    Example
Nominal.   No No No Eye colour
Ordinal Yes No No Competition ranking
Interval Yes Yes No Temperature (°C)
Ratio Yes Yes Yes Height

These levels help determine which statistical methods are appropriate.


https://images.openai.com/static-rsc-4/R4kP14bcaTYlxS1icUu8U8h_oZ4-BhJ5XwYkKy-WuTi09DEYce32KcDojXJ_FUpUef-jgrRAhRjZiUPSxS9ct7YM_l3nEjHNMw9HH2DL6_Nv5ZivjW5k4GPMh8DxYPAxYztFV0V1E3hi1BUIEea1h0UHduhSglUunNblUuKITGbDfSTBIgbDPVf8tWcOfb-_?purpose=fullsize
 
https://images.openai.com/static-rsc-4/KMmn8S0Rqi--w92UU-cmUGLWgmjUkro2Lz3IMrf0pmMTnPJ811vQmSJjC8PTROaCNsAo_4NEnRq2FmlGMB7o_KGEkKgu3uhd1Xt-OGmLWUWGZ1JL8EYcNoiAvj8mDIb21NKSf_MRR57dX5b2fJOm1gfa3O_cwpDxery06AY76HFLsv6_IGxXKCZ1pLb2zJ9v?purpose=fullsize
 
https://images.openai.com/static-rsc-4/C3hTCsjycbSrbcMtuj_EfywsNqRxD98qfM1ZJaOWbEknsSVrUTvRwd24KMWFYOKJDNHq_hrHSnibQBezQgjWAAdu6lopITO6FOrnBf8jghP-jkiK9yFkzoPOkInStbmOxtpjPXKZ422qNH94tBbj30OlBltxBIzSjSvlyQ_1M6dv7wHOaCmMks1BjmAFuzci?purpose=fullsize
5

Figure 4. The four levels of measurement classify data according to how it can be compared and analysed.


Recording Data

Different types of data are recorded in different ways.

Qualitative Data

Often recorded using:

  • Category tables.
  • Frequency tables.
  • Checklists.
  • Surveys.

Quantitative Data

Often recorded using:

  • Data tables.
  • Measurement sheets.
  • Spreadsheets.
  • Scientific notebooks.

Good data recording should always be:

  • Accurate.
  • Organised.
  • Clearly labelled.
  • Include units where appropriate.

Choosing an Appropriate Recording Method

Type of Data Suitable Recording Method
Favourite colour Frequency table
Height of students.   Measurement table
Number of pets Frequency table
Daily temperature Data table with units
Survey responses Spreadsheet or tally chart

Choosing the correct recording method makes later analysis much easier.


https://images.openai.com/static-rsc-4/Afv13lBhZkxDLWyy_9Jk0Vh1z6bltLkUpkALDxQtDdK6Cm58fspgcG4wshaNMVlQzDfTqas8LCBiHe7JFlAGSS-lVaBLSItUuKgQroKRTn1rLzIqhmE_6X73ShWBwDYDc1XewmD9To-27XBHzVmWSk97svE2qBho9lmkOzClmRb9fuIdgLDmwhrZcyVOhhx3?purpose=fullsize
 
https://images.openai.com/static-rsc-4/1-Tw4fCN9kbcQvlDeYgZaRRtSgOEdMbx3ujz5c3gcK6J_VTl00_oJi8NFwMDsI5VxjhLt-GwSzlwzD_BGOOMdISKkqom0UIlYEA3zuRk6uOABxF3qvzE--azrfOkA1ELkpmgrh-qsI9fi5B67V--XH2MXHqjxT-WinTBWSQZLaiqDqpw9tnaa58vDkaLmPPv?purpose=fullsize
 
https://images.openai.com/static-rsc-4/Yu0Is-3VKLZoov19ET2F1-nPbE2M2D1oKYbUjeIwa7BlcyXx-qHdWqFYhHmxbIJZ1iD8El33SAQkb2xfL7fDw2G6Ucp82F1pp9z433w5Mc7F53swiNg7Zv1A0hzz7vn7mhpvj5IWwOLzIzj5xs5r78r6In6NzIoEELpXSKdyFY4m7rz_hJ0FGkJ4JWiJi2WS?purpose=fullsize
4

Figure 5. Different types of data are best recorded using different methods.


Worked Example

Question

Classify each variable.

Variable Classification
Number of siblings.  ?
Height ?
Eye colour ?
Temperature ?

 

Solution

Variable Classification
Number of siblings.  Quantitative, discrete, ratio
Height Quantitative, continuous, ratio
Eye colour Qualitative, nominal
Temperature (°C) Quantitative, continuous, interval

Real-World Connection

A school conducting a student survey might collect several different types of data. Students' favourite subject is qualitative (nominal) data, while their height is quantitative (continuous) data. The number of siblings is quantitative (discrete), and a satisfaction rating such as "poor," "fair," "good," or "excellent" is qualitative (ordinal). Correctly identifying each type of data helps researchers choose appropriate graphs and statistical analyses.


Did You Know?

Many smartphones collect continuous data such as your location, speed, and altitude using sensors, while apps often ask for qualitative data such as your preferred language or favourite music genres. Modern data science combines many different types of data to improve navigation, weather forecasting, health monitoring, and personalised recommendations.


Key Terms

Continuous variable – A variable that can take any value within a range and is usually measured.

Data – Facts, observations, or measurements collected for analysis.

Discrete variable – A variable that takes separate, countable values.

Interval level – A level of measurement with equal intervals but no true zero.

Nominal level – A level of measurement consisting of unordered categories.

Ordinal level – A level of measurement with ordered categories.

Qualitative data – Descriptive, non-numerical data grouped into categories.

Quantitative data – Numerical data obtained by counting or measuring.

Ratio level – A level of measurement with equal intervals and a true zero.


Key Takeaways

  • Qualitative data describes categories or qualities, while quantitative data consists of numerical measurements or counts.
  • Discrete variables are counted and take separate values, while continuous variables are measured and can take any value within a range.
  • The four levels of measurement are nominal, ordinal, interval, and ratio.
  • Correctly classifying data helps determine the most appropriate methods for recording, displaying, and analysing it.
  • Different types of data are best recorded using methods such as frequency tables, measurement tables, tally charts, or spreadsheets.
  • Understanding data types is an essential foundation for statistics, scientific investigations, and data analysis.

2. Data Collection Methods

Learning outcomes
  • I can identify different methods of collecting data.
  • I can distinguish between observational and experimental studies.
  • I can explain the importance of unbiased data collection.
  • I can identify sources of bias.
  • I can evaluate data collection methods.

Introduction

Before data can be analysed, it must first be collected. The quality of any investigation depends on the quality of the data gathered. If data is collected carefully and fairly, the conclusions are more likely to be accurate and reliable. However, if data is collected poorly or unfairly, the results may be misleading.

Researchers use different methods to collect data depending on the question they are trying to answer. Some investigations involve simply observing what happens, while others involve changing variables in carefully controlled experiments. Understanding these methods helps us recognise reliable evidence and identify possible sources of error or bias.


What Is Data Collection?

Data collection is the process of gathering information for analysis.

Researchers collect data to:

  • Answer questions.
  • Test hypotheses.
  • Identify patterns.
  • Make predictions.
  • Support conclusions.

Choosing an appropriate method of data collection is one of the most important parts of any investigation.


https://images.openai.com/static-rsc-4/3rYspsr-t-5Wuz2aFM-oPK6kVRnO_L9GrscS9nCc6nurf08Bh_gh0Y4_ZPQ3RNMXkA4MNu3CjZGDp1hHxwPXiA1QDU281K7I9QSFP_qra3Nznb5nfUJAAnx0ppu9fBoTEASa-2UBeGdJZ6ND6L2PEo-ShmEkOpqLado0Mle62TA5OYkkfTr0b1KiDyDnj6vV?purpose=fullsize
 
https://images.openai.com/static-rsc-4/VPrnfJNI0eu79baMJlxq2p7uDWneNutz_ejtKNFH7HDtBEQqr97yddCFg1zblZMmht7MpygmSajJ3nBx0PW34tha3kc2CXQaqawIEcrAnAvsfghp9WKvD9XzusniP4j9gHTrOu16AKe_35hDJXReG_3l7N3kbupBJIWJigN2tfW68AIWbZKsa8743gRcjo5x?purpose=fullsize
 
https://images.openai.com/static-rsc-4/4n4GlJ7_4NEl-z_DdUPYuwtoNJl9uT_7wcVofWeBYeRBj0fFM_kZozja2eHXLGgdtttY3IS7BsaoKvjbiWcdysUStO7p4Yr0rCIcvd1XhKnTIZbXBtgSu-qoMOESiLVmP7xXPDsqReRc-loYWKgJ0TuAhu_NX3l9yhzgIE5mGN_0tmch5dIXYGEAud3xIZI4?purpose=fullsize
5

Figure 1. Data can be collected in many different ways depending on the investigation.


Common Data Collection Methods

Several methods are commonly used.

Observation

Researchers record what they see without changing the situation.

Examples:

  • Watching animal behaviour.
  • Counting birds in a park.
  • Recording weather conditions.

Experiments

Researchers deliberately change one variable to investigate its effect on another.

Examples:

  • Testing how light affects plant growth.
  • Comparing different fertilisers.
  • Measuring the effect of temperature on reaction rate.

Surveys and Questionnaires

People answer questions about opinions, behaviours, or experiences.

Examples:

  • Favourite school subject.
  • Exercise habits.
  • Customer satisfaction.

Measurements

Researchers use instruments to collect numerical data.

Examples:

  • Measuring height.
  • Recording temperature.
  • Timing a race.
  • Measuring rainfall.

Existing Data

Researchers analyse information that has already been collected.

Examples:

  • Government census data.
  • Weather records.
  • Hospital statistics.
  • Scientific databases.

https://images.openai.com/static-rsc-4/5k4Hy2E2iNtu7mTp-UtfZT6pQk42OTc2BgfLeL9wAZLZAYiA6i0s6eSE8QS6a4GSx12-cT8CL6rL3Cdx1WGLOY2otAcAamq260VaWO6bdtJaiRqHPhh1NbHwSMXi2ZNBoODOeM4bbVsk84dZPGdS2nQITe_fk19RtlaO9KTB1zlt7cevvonooghsC7q3MseM?purpose=fullsize
 
https://images.openai.com/static-rsc-4/ZU8NqGBJcaUsEjY2VntAiiwh5AUXPfDKjtehzuid8Dao3FWOOza5HgjA1BzySeU3t2Yh9n5N6XMMf7e4ZkymeCERDdeSF7l4FbrsOdB2ciKeZNBvsKBZcT1TJxqLILndjNe_VO11ZzH2D0pSWAxeAWcuVj5_PZDz9futvh2cUDZwqjc56Cev86fQkIdLt9oG?purpose=fullsize
 
https://images.openai.com/static-rsc-4/tvm6I8QmO8NZzLyBBi3nKdQKAq7lijaVPVx-lnqbTFM3465FifvsxCVKYqxwBQWvDVFV0JF2OUfjXLrtFKRCvlxFFMy_nCNU6xTglwjOm2SfyZihnTj94lXrMnmxY38HYxrgke6uOSZL6lQv9qsCkwcGKwsKhBVVaTMgIfx7EXmc7iL18gIcF524GMGpO_yD?purpose=fullsize
5

Figure 2. Different investigations require different methods of collecting data.


Observational Studies

An observational study records information without interfering with the subjects or environment.

Researchers simply observe and record.

Examples:

  • Studying animal migration.
  • Recording traffic flow.
  • Monitoring weather patterns.
  • Observing classroom behaviour.

Observational studies are useful when experiments are impractical or unethical.

However, they cannot usually demonstrate cause and effect, only relationships or patterns.


Experimental Studies

An experimental study investigates the effect of changing one variable while keeping other variables controlled.

Experiments involve:

  • An independent variable (the variable being changed).
  • A dependent variable (the variable being measured).
  • Controlled variables (variables kept constant).

Example:

Investigating how different amounts of sunlight affect plant growth.

Experiments are useful because they can provide evidence for cause-and-effect relationships.


https://images.openai.com/static-rsc-4/P8hB5Y4DQdttZmRAF6TiDxSgCO8_X0O5svLwhTAByqXiUtjHtqy0rpq7O3oZ_T_xzxkhVuQ9nM3HcJmPOtY5NJzbR7C05qnB1Y6uquZUU_B-cGHQ-pucfQl3HaJ3yyOTP46yO1Ocj8PHgZDVjHTrEyySxHC46oPUTGjFtXyABi9ALVkr_wkgdzHVvNMG2Y32?purpose=fullsize
 
https://images.openai.com/static-rsc-4/ny3wEm8HqQdThs7hxib6iWev_pxjqREh2tkKXTr3TKUkfz1oTeGEBnZZ2Eosy6p5mClgJRXCK6D7mQ2HvITghJrXGRYA8RdCH4QaExj8friWAWeBFLvJEInptCcaa3AhzkiZkV1bz7VesuSJfJ37jeNqrecgm7dT0ftt0tqwsA9JnBUbNF1oLPyqGihbT2Io?purpose=fullsize
 
https://images.openai.com/static-rsc-4/3BFAvuwj6PqGUwu_7fov2TX0tS9AQTUUfcxErcZXGqsq2bxVYr4t41HpWs7ag4eaEkSKt7Z__g6jOQsvTjvw-Vf22E50lJTM_cv-LQLnLBxNANm32WnQKyMpZyDvEOA8VWFLW78cReHgRlfU0Ag9NDWnVG26aMBfHwq2z_zoHWn47WXPagR5WucnuhKmrhdz?purpose=fullsize
5

Figure 3. Observational studies record natural events, while experiments involve controlled changes to variables.


Comparing Observational and Experimental Studies

Observational Study Experimental Study
No variables changed Variables deliberately changed
Records natural events.   Tests cause and effect
Usually less control Highly controlled
Often used in ecology Often used in laboratories

Both methods are valuable, depending on the research question.


Why Unbiased Data Collection Matters

Bias is anything that unfairly influences the results of an investigation.

Unbiased data collection helps ensure that:

  • Results are fair.
  • Conclusions are accurate.
  • Data represents the population being studied.
  • Other researchers can trust the findings.

Scientists work carefully to reduce bias whenever possible.


Sources of Bias

Bias can occur in many ways.

Sampling Bias

The sample does not fairly represent the population.

Example:

Surveying only students from one class when studying the entire school.


Measurement Bias

Measuring instruments are inaccurate or used incorrectly.

Example:

Using an incorrectly calibrated thermometer.


Observer Bias

The researcher's expectations influence observations.

Example:

Recording only behaviours that support a hypothesis.


Question Bias

Survey questions influence participants' answers.

Example:

"Don't you agree that school uniforms improve behaviour?"

This wording encourages a particular response.


Selection Bias

Participants are chosen unfairly.

Example:

Allowing volunteers to participate instead of selecting people randomly.


https://images.openai.com/static-rsc-4/U7cSNTWTJW5VuK5wLRlEgQ80Z3YOy0JO2xY9bBQkpv0Vc7ZvNBtUPXJm7ngR9rsW0SffrTq682DGWu6wmrBbcw6NDxbM1uWEOHYtRwsV92HZgos4EAhBu6VGFHmMAlVzafOjotGAcfANjwJOHmrS5qMHjft4rkIMbjSaZS1c6KfSbnhQRYQ5FHDuQeKW-zC1?purpose=fullsize
 
https://images.openai.com/static-rsc-4/15ph3uvQeKypp_AcAQ12Ybuyfr_hzW7Gfa1xNoQJK4cnDsvFVDXs7bkUsdd4ZSSypbiReci8Pze1cZwig37bjMmGYBQYYAsha-yAGUQjl4Ep0wGaePnRR6Mm6_QOgXHqvOn8YqJ-TBnaprLHta8xyWGnHEpFGq5gS_u7BO5nYpAC2jqQ50lPQXbswf_4Wn88?purpose=fullsize
 
https://images.openai.com/static-rsc-4/_afvWS4X4faS8apU5z8fSs14ExUUbuIFPa9ijg6T2YvJ1zyL8dUuINPI62jU8ANNlc4BPuBzkIai4BHumMY9_HLYZLW8mxrE1d5c_MUy_m0MF2wXn8CS4V_MliMVBMpHFNLv9hmVgljnMsjYfmJL0f2xzRH-h675LdKS9emdPWtmoH_6OWP9a4M6k5ZGxjpE?purpose=fullsize
4

Figure 4. Identifying sources of bias helps improve the quality of investigations.


Reducing Bias

Scientists reduce bias by:

  • Using random sampling.
  • Increasing sample size.
  • Using calibrated equipment.
  • Following standard procedures.
  • Asking neutral survey questions.
  • Repeating measurements.
  • Allowing independent verification.

These methods improve the reliability of the data.


Evaluating Data Collection Methods

When evaluating a method, consider:

  • Is it appropriate for the question?
  • Is the sample representative?
  • Are measurements accurate?
  • Could bias affect the results?
  • Can the investigation be repeated?

Good investigations collect data that is:

  • Accurate.
  • Reliable.
  • Fair.
  • Relevant.

Choosing the Best Method

Investigation Suitable Method
Measuring plant growth Experiment
Counting birds in a forest Observation
Student opinions about homework.  Survey
Daily rainfall Measurement
Population trends over 50 years Existing data

The best method depends on the type of question being investigated.


https://images.openai.com/static-rsc-4/Tp-ENpNBzey4abI8-i-rvJUqBhsHJhRQyVRzX2KizmgSzkaToRL4f-fYUZpbW1zoxMTUrzZPf0ojpNgQ4MXbTTJmmM_aiKDiG588wxTqWt0_Zj5Leswdr_q2eFa4AevY-cYE6iRzzHQtbGEcSNdjdWYaBZkww63A-jJj46lEJkg52Xr30pL6gyQkrFFYJhK_?purpose=fullsize
 
https://images.openai.com/static-rsc-4/4WWIqIQ-dXpwjydnD9gH48fZf4FPROnCnSWPbEEKAeFHKQKOrzFQJhib_9_nujAawMUcCSt7CzQsohb4piCbPYwGwlnjDvB0_EKM5JkSioxbmD8Hp0n5nAqryeTfZHkcfJ1SS6ZJeGYXhYEyejxW5iJKA6VuFfNIDuxc74fQRbfOARiIKYsbAoJSnwA-gcME?purpose=fullsize
 
https://images.openai.com/static-rsc-4/uhKrgRPMMEh-L_csdbR5DZJ4o2-FSPJc2LZPYSvU_49I5Cedid3d5K2B2Uk0ngLQhSY0pVirV2xdgjsgaFps5BGasC0XPU6lp-0g2I895x80vFHRCyWSEfU_3gJ0jJaiFRHy8eETQcYretQcpN_WVTlnbjOnlNy-bkDHo1IvitQb5pfzOZrijBuyFA4Cdry-?purpose=fullsize
5

Figure 5. Different research questions require different methods of collecting data.


Worked Example

Question

A scientist wants to determine whether a new fertiliser increases plant growth.

Should they use an observational study or an experimental study?

Solution

An experimental study is appropriate.

The scientist can:

  • Apply the new fertiliser to one group of plants.
  • Compare it with a control group.
  • Keep all other variables constant.

This allows the scientist to determine whether the fertiliser causes increased growth.


Real-World Connection

Health researchers often use observational studies to investigate relationships between lifestyle and disease because it would be unethical to deliberately expose people to harmful conditions. For example, researchers may compare the health of people with different exercise habits over many years. In contrast, scientists testing a new medicine usually conduct controlled experiments or clinical trials to determine whether the treatment causes improvements in health.


Did You Know?

Many weather forecasts rely on millions of observations collected every day from satellites, weather stations, ocean buoys, aircraft, and balloons around the world. By combining these observations, scientists build computer models that predict future weather with increasing accuracy.


Key Terms

Bias – A systematic influence that causes results to be unfair or unrepresentative.

Controlled variable – A factor kept constant during an experiment.

Data collection – The process of gathering information for analysis.

Dependent variable – The variable that is measured in an experiment.

Experimental study – A study in which researchers deliberately change one variable to investigate its effect on another.

Independent variable – The variable that is deliberately changed in an experiment.

Observation – Collecting data by watching and recording without interfering.

Observational study – A study that records information without changing the conditions being investigated.

Random sampling – Selecting participants so that every member of the population has an equal chance of being chosen.

Survey – A method of collecting information by asking questions.


Key Takeaways

  • Data can be collected through observations, experiments, surveys, measurements, and existing data sources.
  • Observational studies record natural events without changing variables, while experimental studies investigate cause-and-effect relationships by controlling variables.
  • Unbiased data collection is essential for producing fair, accurate, and trustworthy results.
  • Common sources of bias include sampling bias, measurement bias, observer bias, question bias, and selection bias.
  • Scientists reduce bias by using random sampling, accurate equipment, standard procedures, and repeated measurements.
  • Evaluating data collection methods helps ensure that investigations produce reliable and meaningful conclusions.
 
 
 

3. Sampling Techniques

Learning outcomes
  • I can describe different sampling methods.
  • I can distinguish between random and non-random samples.
  • I can identify sampling bias.
  • I can evaluate sample quality.
  • I can select appropriate sampling techniques.

Introduction

In many investigations, it is impossible or impractical to collect data from every member of a population. Imagine trying to measure the height of every person in a country or count every tree in a forest. Instead, researchers collect data from a sample, which is a smaller group selected from the larger population.

A good sample should represent the population as accurately as possible. If the sample is chosen fairly, the results are more likely to reflect the true characteristics of the population. Understanding different sampling techniques helps researchers collect reliable data and avoid sampling bias.


Population and Sample

A population is the entire group being studied.

Examples:

  • All students in a school.
  • Every tree in a forest.
  • All households in a city.

A sample is a smaller group selected from the population.

Researchers analyse the sample and use the results to make conclusions about the whole population.


https://images.openai.com/static-rsc-4/fwP2Msy5zTgRilLKhtkVZ-Inqt59BArT8wsdW6WNYnC7UAUtHIqWH-TPYt1uTNF5mxZwy_1sLiuKlVispvu6_WHkTnfTOS0BK9fYaC07Q8ushfQP5oMAPOxi9upfM0uR4119gdlan4IbbOdr__CuHJmv7rRLdl5fWQh5tjkBL7E06Z5lhCizlFySHHJvhnJD?purpose=fullsize
 
https://images.openai.com/static-rsc-4/2iRg6hW8PGl9b3eAkyGeFwCP_4xpevsKsYDd6YMKI-saepYJgmJOsEBFPerdvSxN1D0zZrF2OysUkN6mDP7-7qr2DuLXMctecsU_AlN79sDlw_sSOg9FGQTTSdqPhxUWmC8ITngmxk4e9zJo66m2o07ku14Uq66GSJjwM8E-8jN8N9TdAlTDTklxcbuCpsiD?purpose=fullsize
 
https://images.openai.com/static-rsc-4/zEiP5izwmXOIkLzCNJaZOC-CJP0W_iFgTHOP_Zm7GNcPgFMdjY-T4BFFZQAz45lj6us_arUay_j4Lb9CvJcQRbOKV5yPZuB2nmV3pWBCmv7TFVnccvrAqpsAMZUXegCXJYj6cKjyNghW1Jj-l9Fv_FwzegNAZHnHyNsbj_LbYqtIrKBCixWB-wAtdCRsEexE?purpose=fullsize
5

Figure 1. A sample is a smaller group selected to represent the larger population.


Why Do We Use Samples?

Sampling saves:

  • Time.
  • Money.
  • Effort.

It also makes large investigations practical.

For example:

Instead of surveying every student in a school, researchers may randomly select 100 students.

If the sample is representative, the conclusions are likely to be reliable.


Random Sampling

In a random sample, every member of the population has an equal chance of being selected.

Examples:

  • Drawing names from a hat.
  • Using a random number generator.
  • Selecting random student ID numbers.

Random sampling helps reduce bias.


Non-Random Sampling

A non-random sample is selected using methods that do not give everyone an equal chance of being chosen.

Examples include:

  • Asking only nearby people.
  • Choosing volunteers.
  • Selecting the easiest people to reach.

Non-random samples are often quicker to collect but may not represent the population accurately.


https://images.openai.com/static-rsc-4/tdkUpzz3jk-OLKzf_aFpOttWPyqO4o0TR3oYbQMDYLyw6C8du3odsSGOm1PiVTPKouGbhlvsungKYx3FmLksLhOpvQyIDG0JIKt0AenkewzCzHgxCA7XklDWr1G3m2cBwPMvrzQM7M5HGMOldc849X7EKks2E7Y_IwEUacHqA2qQnKcQjsyG_j2qnpSPhHPF?purpose=fullsize
 
https://images.openai.com/static-rsc-4/rXMUsRRkouu2tudPfzHJhkBYTKkjJiTIq2yJDb-9dbNlXbi0hV0LweAtcMBknD8Dvv8R53ev9C9w0CNd8jARtHyONZTVEg5E15yCloq_ljC-HlPcTepYf6oBCmuGD7UaHVChi14M3WVQZ87Yqa524sMVUnGIpEPBi3d15Zq8LRGhKZ8EKSaHng2S7RHcF3XW?purpose=fullsize
 
https://images.openai.com/static-rsc-4/kxUsynDFnac7hS0mcLOB5fta7Pqi0riJm-DECeqyQJuAfuwyPVNbC3ukH5fGWDQv4_LiRG6cZ-8v0ails7W-yQbYR3btV-XY10NbivUWT6tYlVBJ2Bq4r4xonkeHQeGeuKNEVkmEN1aH3qzHslDLYr6Pw8hymjI6mksoSWS_HC6UoN0EjEcNC2o50YW6hZC1?purpose=fullsize
5

Figure 2. Random sampling gives every member an equal chance of selection, reducing bias.


Common Sampling Methods

Simple Random Sampling

Every member has an equal chance of selection.

Example:

Randomly selecting student numbers from the school register.


Systematic Sampling

Researchers select every nth member after a random starting point.

Example:

Choosing every 10th customer entering a shop.


Stratified Sampling

The population is divided into groups (strata) with shared characteristics.

Samples are then selected from each group in proportion to their size.

Example:

Sampling students from each grade level in a school.

This often produces a more representative sample.


Convenience Sampling

Researchers select individuals who are easiest to reach.

Example:

Surveying students who happen to be in the school cafeteria.

This method is quick but often introduces bias.


Voluntary Response Sampling

People choose whether to participate.

Example:

An online poll where anyone can vote.

People with strong opinions are often more likely to respond, increasing the risk of bias.


https://images.openai.com/static-rsc-4/kxUsynDFnac7hS0mcLOB5fta7Pqi0riJm-DECeqyQJuAfuwyPVNbC3ukH5fGWDQv4_LiRG6cZ-8v0ails7W-yQbYR3btV-XY10NbivUWT6tYlVBJ2Bq4r4xonkeHQeGeuKNEVkmEN1aH3qzHslDLYr6Pw8hymjI6mksoSWS_HC6UoN0EjEcNC2o50YW6hZC1?purpose=fullsize
 
https://images.openai.com/static-rsc-4/0_vhSHW5Kl9fkdPaE6MLKUKF44XNfHv3F996dXWygJlOFtZf9LXzWsbwEBLa3K_4nUIMZCy-h--IbF-R90ktVnayk4nB5et9Hunhh_dIoZjBTLzKe4rW1AJDZob-VDRtrlgGSGvRAcOp7ZXnxKMUopc8zorr4i6FbSdTcBki3wCgp9aJXGyPBRLGdfPLo_nm?purpose=fullsize
 
https://images.openai.com/static-rsc-4/LGuITBn-gmnwG6RKkUm8sKsNr03eJZzYitrcXZv0JUq6vw1aV1U5Cxv8MMqzftIRPYsnYrRFp6lXxLZTPJCYB84VxbXFeFnDudq6IOJBPU-WyQhb5g-9phw0yHWuy_edU93-9H0Cj7Wcy4Pr41uQjrydTyQeu7FdknsY_ZGkGMuduoiCkTwEtvpu_OzJu8Z9?purpose=fullsize
4

Figure 3. Different sampling methods have different strengths and weaknesses.


Random vs Non-Random Samples

Random Sampling Non-Random Sampling
Equal chance of selection Unequal chance of selection
Less bias More potential bias
More representative May not represent the population well
Better for scientific investigations.   Often quicker and easier

Researchers usually prefer random sampling whenever possible.


Sampling Bias

Sampling bias occurs when the sample does not fairly represent the population.

This can lead to misleading conclusions.

Examples:

  • Surveying only morning students about school transport.
  • Studying exercise habits using only members of a sports club.
  • Asking only adults about video game preferences.

Sampling bias reduces the reliability of the investigation.


https://images.openai.com/static-rsc-4/J-ArxMoQOf2oll_lWl9UO4Uu_TiqdZsTpBbRuugRskNJRnlB3JLeJLUdqaCV812b4PgTcicGo5Or0euVRVK-UrI3M3myEoNyXSIAlSaiAhYg5mOBU4kdMiZ24St7HkmfVt7_com65RYZXCfQDTYyHjnmTqRfBjySk3UL-saIeaHT8ZNi69aKgJLI9vRnjeOT?purpose=fullsize
 
https://images.openai.com/static-rsc-4/hhyf0gpg-nJZceZ-9P17Vf2zqnRpo_BPabOSIaMf5ewoXCQv_glOjjwTSBqBFDZCAIqReu3iJ29bnMhlMt2xBuoLe3cngw8FTjDvibSQOH-Z8qRypYZK66X3LhMDWT8iHvW9okiOLLsr7zgCOkjaH1hdBLMVoQIbn8w_ajGi2CkxX4mcrRLQJfKDdlKVzYVq?purpose=fullsize
 
https://images.openai.com/static-rsc-4/wK8klDyAOOCCa3ZbMsZbJUGkVK0OJPQpinQXjLCnmTrqXd5vEXgfuZJU1gNqkmxBpAqD6o-IW1S4Bzc2pY9gssync3XUUd9CUa29NuD2svda0TS8gVZC78mKtzhJNK1YAAZpOE2N9v_Amu5uMl-6RVyEuymRYeQzoUuuo1AA91zmjbWMZSmKkwoQzjYBk1BV?purpose=fullsize
5

Figure 4. Sampling bias occurs when the selected sample does not accurately represent the population.


Evaluating Sample Quality

A high-quality sample should be:

  • Representative of the population.
  • Large enough to reduce random variation.
  • Selected fairly.
  • Free from unnecessary bias.
  • Appropriate for the research question.

Researchers should ask:

  • Does the sample represent the whole population?
  • Was the sample chosen fairly?
  • Is the sample size sufficient?
  • Could bias affect the conclusions?

Choosing the Best Sampling Technique

Different investigations require different sampling methods.

Investigation Suitable Sampling Method
Surveying student opinions across an entire school.   Stratified random sampling
Measuring tree heights in a forest Systematic sampling
Selecting participants for a medical study Simple random sampling
Quick classroom opinion poll Convenience sampling
Online public opinion poll Voluntary response sampling

The best method depends on the purpose of the investigation and the characteristics of the population.


Why Sampling Matters

Good sampling allows researchers to:

  • Save time and resources.
  • Draw reliable conclusions.
  • Reduce bias.
  • Improve the accuracy of investigations.
  • Make fair comparisons between groups.

Poor sampling can produce misleading results, even if the data is analysed correctly.


https://images.openai.com/static-rsc-4/0bAT3mbpT00qdKWPXvgFFCH0mZnB8u4LCrcdn74e5Xn9IWBdfvIUo_ekY2WngAXFRmEAYqsQNpx3dZTw-z-_hE1fsw9fmzi1yLM_Q-ef-mV728fWKt4p0l-4-xE3P-y40jw4CzxE9C7Swv1bfvE1RbQNNhPW3DoYu4HlMiuNQiN0NfhXdUWjpbuxKTCmPU5X?purpose=fullsize
 
https://images.openai.com/static-rsc-4/xh3bwwW_egCdu13pGkyeYstc645OEjj30PmX4vIbmw9Q2lhV3FYK99Pqc7WFQo3uYaLOGqMJQd5SJkwHWzN0m08mc3da_InSxnuBA3N49O-rQbGD7-_Z4fW-msmyvFUdy3QKLo98y4lXmAA5OlcwA4tlIyeDLAFue_lbX_OWdPHW1KvoclXVUxKtAjCQmBeH?purpose=fullsize
 
https://images.openai.com/static-rsc-4/uIguOPfg6TTlaZwsC2KiI1UNOvlr1d8RfxdYwKSAX-VPUJqhqikP0ANL9_QK76PAKJjx7kWeR4esmM81NQi8jD6VJCYlgcIjdEBocTGxS_n5aMSfJ2HtrgZK1Y5hJtQ9v7tAplywCbTd-IPpudopiEvn2FUAnJ_H3AhpAtoSATmYy4M96bv2ngFW4FFDHUOF?purpose=fullsize
6

Figure 5. Careful sampling improves the accuracy and reliability of statistical investigations.


Worked Example

Question

A researcher wants to survey the opinions of students in a school with Grades 7–12.

Which sampling method is most appropriate?

Solution

Stratified sampling is the best choice.

The researcher divides students into grade levels (the strata) and randomly selects students from each grade in proportion to the size of that grade.

This ensures that all year levels are fairly represented.


Real-World Connection

Before national elections, polling organisations survey a relatively small number of voters instead of asking every eligible citizen. To make their predictions as accurate as possible, they use carefully designed sampling methods that include people of different ages, regions, occupations, and backgrounds. If important groups are left out, the poll may not accurately represent the opinions of the entire population.


Did You Know?

Although some opinion polls survey only 1,000–2,000 people, they can often estimate the views of millions with remarkable accuracy—provided the sample is large enough, randomly selected, and representative of the population. In statistics, how the sample is chosen is usually more important than simply making it larger.


Key Terms

Convenience sampling – Selecting individuals who are easiest to reach.

Non-random sample – A sample in which not every member of the population has an equal chance of selection.

Population – The entire group being studied.

Random sample – A sample in which every member of the population has an equal chance of being selected.

Sample – A smaller group selected from a population for study.

Sampling bias – Bias that occurs when a sample does not fairly represent the population.

Simple random sampling – Selecting participants entirely by chance.

Stratified sampling – Dividing a population into groups and sampling proportionally from each group.

Systematic sampling – Selecting every nth member after a random starting point.

Voluntary response sampling – A sampling method in which individuals choose whether to participate.


Key Takeaways

  • A population is the entire group being studied, while a sample is a smaller group selected from that population.
  • Random sampling gives every member of the population an equal chance of being selected and generally produces more representative results than non-random sampling.
  • Common sampling methods include simple random, systematic, stratified, convenience, and voluntary response sampling.
  • Sampling bias occurs when the sample does not accurately represent the population, leading to unreliable conclusions.
  • A good sample is representative, fairly selected, large enough, and appropriate for the research question.
  • Choosing the correct sampling technique improves the quality, accuracy, and reliability of statistical investigations.
 
 
 

4. Organising Data

Learning outcomes
  • I can organise data into frequency tables.
  • I can construct grouped frequency tables.
  • I can create cumulative frequency tables.
  • I can interpret organised datasets.
  • I can prepare data for analysis.

Introduction

Raw data collected from experiments, surveys, or observations is often difficult to understand at first glance. A long list of numbers may contain useful information, but patterns are not always obvious. Before analysing data, it is important to organise it into a clear and systematic form.

One of the most useful ways to organise data is with frequency tables. These tables show how often different values or groups of values occur, making it much easier to identify patterns, compare results, and prepare data for graphs and statistical calculations.


Why Organise Data?

Organising data helps us:

  • Find patterns.
  • Compare results.
  • Identify unusual values.
  • Reduce errors.
  • Prepare data for graphs and statistical analysis.

Well-organised data is easier to understand and communicate.


https://images.openai.com/static-rsc-4/gJWY95jJIhHJGUc6nuOdM8AOM9TdtZEeCHDrRu1gbVUEAtr_Gauv39vHDtju0baFCNKyYIxUteqD4mOmVq1ob8JV-1UmAl2A9iSr-7CiPOL0nUnp7p_H-P9C2H7YdIT3ZkcGSPjHsEhVzxnmsXSqMY2Oes1qLNUELmPV-9iTR8jC0exIwRRtXdYHiB22TAGQ?purpose=fullsize
 
https://images.openai.com/static-rsc-4/d5n1z3o1OjsspSqiuoGYZn5i4VjfFjsNzcDSSAoxhbOdKw8gw9jCTixNjGu0EZ2CSKyUK3WweKk71uUg28MunF1D-PszUnqs-LBuV5nRDjrmQLN2A4kGXdg4PRbPU-acppB_mRJhCu3ONe-eBbgJ7vAKiBSguH8x_Iwc-V_Cz-zVZqaa6n2i94UePB01gpEe?purpose=fullsize
 
https://images.openai.com/static-rsc-4/MCrv7DvrOhLn9lPzqNthhINOwMS6YWXzVSe7dVadYiGsJatO78SqVTCnONz2EYkn_Gm0qDLLHv0UvpbrfVDxG8TKBEzI_-Hhy2JDZDjY9OvCaU6_l9_hKi3cWI7fBTOlBSbBvyVArHBOj68p-JLCyDa73ZUXSHuG0adymcz_leKwPXHFhF52KJdNCj4rs_99?purpose=fullsize
6

Figure 1. Organising raw data into tables makes patterns easier to identify.


Raw Data

Raw data is data collected in its original, unorganised form.

Example:

18, 16, 20, 18, 17, 19, 18, 16, 20, 17, 18, 19

Although the information is complete, it is difficult to see which values occur most often.

Organising the data makes interpretation much easier.


Frequency Tables

A frequency table lists each value and the number of times it occurs.

Example

 Score   Frequency 
16 2
17 2
18 4
19 2
20 2

The frequency is simply the number of times each value appears.

Frequency tables work well for small datasets or discrete data.


Constructing a Frequency Table

To create a frequency table:

  1. List all the different values in order.
  2. Count how many times each value occurs.
  3. Record the frequencies.
  4. Check that the total frequency equals the number of data values collected.

Organising values in ascending order makes the table easier to read.


https://images.openai.com/static-rsc-4/gJWY95jJIhHJGUc6nuOdM8AOM9TdtZEeCHDrRu1gbVUEAtr_Gauv39vHDtju0baFCNKyYIxUteqD4mOmVq1ob8JV-1UmAl2A9iSr-7CiPOL0nUnp7p_H-P9C2H7YdIT3ZkcGSPjHsEhVzxnmsXSqMY2Oes1qLNUELmPV-9iTR8jC0exIwRRtXdYHiB22TAGQ?purpose=fullsize
 
https://images.openai.com/static-rsc-4/4KVPymmfLvECEDDsnaZclKnK2WPGxhKQHoRLo5BdPEqDrR7CEaPrJ72jVeYm5jg9uwgDje_ov_Aah5WmBt22dPZL58BROdIvUCL5sdCeCAQFxcdj3XLDfXQzUCxZalPfgWuASbVbk69KsSx2TLt-mirltz8tgun4PayzcK37xitnfg7UFA1UVq_guIBslb38?purpose=fullsize
 
https://images.openai.com/static-rsc-4/Uo-JFaDItndUaO2tXKh2hqVMAVvGDalWHxAhVjcmYztJONNMAqBWYYJX_lHiEZY-spHlhx0-0971mnEoL18LZ6kg4q85LTfTvmYpgV7GE_R7xoC-3UsYKHAdk5VCs_zXgIE1IkN4qmNvE051KwrqtaMRgYu1DFuhrrXgIJoW48IIp-1ZM4B-N7CLJJoo9PBj?purpose=fullsize
5

Figure 2. A frequency table summarises how often each value occurs.


Grouped Frequency Tables

Large datasets are often organised into groups, called class intervals.

Instead of listing every individual value, similar values are grouped together.

Example

 Height (cm)   Frequency 
140–149 3
150–159 7
160–169 11
170–179 6
180–189 3

Grouped frequency tables make large datasets much easier to summarise.


Choosing Class Intervals

Good class intervals should:

  • Cover all data values.
  • Be equal in width whenever possible.
  • Not overlap.
  • Be easy to interpret.

For example:

✔ 0–9, 10–19, 20–29

✘ 0–10, 10–20, 20–30

The second example overlaps because the value 10 belongs to two groups.


Cumulative Frequency Tables

A cumulative frequency table shows the running total of frequencies.

Example

 Score   Frequency   Cumulative Frequency 
16 2 2
17 2 4
18 4 8
19 2 10
20 2 12

The cumulative frequency tells us how many observations are at or below each value.


https://images.openai.com/static-rsc-4/ZUCGynAUHB9rivmeS6KwsRiEwBSlD7cPpKF2NljWxhPqSVFTLMCKWakeVhLjl_9o-l6pFwj-7lBOVVeTXtiWBJrG4iLnmsJv2Nr9N6l9bcKyEG3gDN7vp5Iv9s9zt754VMYqd9tPzasEBDOyDf7uQ-4Z3LYwqZbXn7UvytFPzcFbMtVAIvwfTyPum-CS97TY?purpose=fullsize
 
https://images.openai.com/static-rsc-4/RCo8CD-AyeTbXGvtuudrSQI-4WA-c7t_-X7BA7ldgecB6KB-mGe1l_D5OVPx8hHFD04X1qCldtmmcEAZXTMztVUiWN3owM-HxBptZg2PVuOKT1yWqRZibuNnbUgdhm6lTJCey-0Zj3U8aeZQwqqHNaIrSm90CX_Vr4WOC2e9NrB1Y5UulDpqczx8GH7So1qa?purpose=fullsize
 
https://images.openai.com/static-rsc-4/Bq_AC75gqktonPlT9SKr02nMWUWdGv9ZDQ1TyMIn9fEfOauJJjEsuX09Qvda94xNv1tJSxsXJLIKoI9aZrmUNZb3jg2bQd-jjm0yjr1jVN_NZVPg-jvcf2WV8AAw1JyB7eirmfLdPh1zUpMF26od98Km-X9YE1IsSqjeF8yIqk03sdASncJBLKk1MM8ohDNE?purpose=fullsize
5

Figure 3. Cumulative frequency tables show the running total of observations.


Interpreting Organised Data

Organised data helps answer questions such as:

  • Which value occurs most often?
  • What is the smallest value?
  • What is the largest value?
  • How many observations are below a certain value?
  • Are there any unusual values?

Looking for patterns becomes much easier once the data has been organised.


Preparing Data for Analysis

Before calculating averages or drawing graphs, data should be checked.

Good preparation includes:

  • Removing obvious recording errors.
  • Checking for missing values.
  • Ensuring units are consistent.
  • Organising values into tables.
  • Labelling headings clearly.

Well-prepared data leads to more reliable analysis.


Choosing the Right Table

Situation Best Table
Small list of test scores Frequency table
Heights of 500 students Grouped frequency table
Finding the number of students scoring below 70%.    Cumulative frequency table
Counting different eye colours Frequency table

Choosing the appropriate table makes the data easier to interpret.


https://images.openai.com/static-rsc-4/ZUCGynAUHB9rivmeS6KwsRiEwBSlD7cPpKF2NljWxhPqSVFTLMCKWakeVhLjl_9o-l6pFwj-7lBOVVeTXtiWBJrG4iLnmsJv2Nr9N6l9bcKyEG3gDN7vp5Iv9s9zt754VMYqd9tPzasEBDOyDf7uQ-4Z3LYwqZbXn7UvytFPzcFbMtVAIvwfTyPum-CS97TY?purpose=fullsize
 
https://images.openai.com/static-rsc-4/-E8WVtEFOtfYLb9DTUbOwpD39JphFFIVVfCKSVRXqEzbEltoM19daFVpNHY895wm_XBFG7kfJjxueCjxL3cdcAqtAd75lB9TFr2YpEcq7qix9BVKtHqziUMR47hzCLlxvVky000uWleBbF9Sq0bQLOawnUzUHnK_ZD1oBv3GuzPQS4mbQ4NN-tA3XOuDa9lf?purpose=fullsize
 
https://images.openai.com/static-rsc-4/GyawzIwA1vJUlX5_f7hdnoBmBwtwHdasBe6_LGPEe5m3S-gciN8MZRoLkaL757cAm15gEtheW3Pkx8ybxai0hI72_4WMSw0dlHqMvRab1AHVWkDgekSxopq5uqebDUW2WpTRwYk9_A0jebR_MGavuHMu1ubVbf5pyXSp6dHpAk-nnYDEPWjpdi29syu3DqwU?purpose=fullsize
5

Figure 4. Different types of frequency tables are suited to different types of datasets.


Why Organising Data Is Important

Scientists and statisticians organise data because it:

  • Reveals patterns.
  • Simplifies calculations.
  • Makes graphs easier to construct.
  • Improves communication.
  • Reduces mistakes.

Organised data provides the foundation for all statistical analysis.


https://images.openai.com/static-rsc-4/YJHfMxHtie5FzXOUC3lwsbZtD9tsw38WJw8fKq71U9bYdI3RwQU4jvZpQ3jc0FgUgLfddx5at0oOAa3uWzjmFSlUnT3ntVxQ42T4rsIuBbLdnpnm8u3cssoU5IsPNuChRxkIwAFXBcxi4imHITzXE4PsHIPjMqhLEYDtD96g3MO8RsXqnuYX5G4zzpwqA-Ku?purpose=fullsize
 
https://images.openai.com/static-rsc-4/5Zgwt-9BgV8kdlmTCUrSzVMYvM5M9ckb4dkysxPe032LAUJ-mdDlY-eS0WHR5u4gyOUjyjpOh3EAM9cWKs652oOjEtNgdowH5u2VjUSOBles3gim3gbmNcWdkiTRx7q1YMySxzCrbxT5nL9NaC6IOaJ3ErvAppjssWpE1oi5lqytG8a9J6SFIs5B1TX-P1d0?purpose=fullsize
 
https://images.openai.com/static-rsc-4/iIYMj971koxbAXD-aNNxl5GSWvyB0PC-gA7Sp_J99tYTfqkM5J6aLG95S701iwnTInMlcsE5AEr9soY2oEVkmEjRueTTHnxPW0iEelVqT4kxEIhdcvqc9tazAtYair1uHiyrjhe9UlzQIDpvxnWYtSNUb2TU3AP-RHVi1OM-7JTvgcy_n8zbxCziyHDjTzd3?purpose=fullsize
4

Figure 5. Organised data is the starting point for graphs, averages, and statistical analysis.


Worked Example

Question

The following test scores were recorded:

15, 17, 18, 16, 18, 17, 19, 18, 20, 17

Construct a frequency table.

Solution

 Score   Frequency 
15 1
16 1
17 3
18 3
19 1
20 1

The total frequency is 10, which matches the number of scores collected.


Real-World Connection

Schools often organise examination results into frequency tables before analysing student performance. A grouped frequency table can quickly show how many students scored within different mark ranges, while a cumulative frequency table helps identify the percentage of students who achieved a particular grade or higher. This information supports teachers in evaluating class performance and planning future lessons.


Did You Know?

Before computers became common, statisticians organised thousands of data values by hand using tally marks. Even today, tally marks remain a quick and effective way to record frequencies during classroom experiments, surveys, and field investigations before transferring the data into frequency tables.


Key Terms

Class interval – A range of values used in a grouped frequency table.

Cumulative frequency – The running total of frequencies up to and including a given value or class.

Cumulative frequency table – A table showing frequencies and their running totals.

Frequency – The number of times a value or group of values occurs.

Frequency table – A table showing each value and its frequency.

Grouped frequency table – A frequency table in which values are organised into class intervals.

Raw data – Data in its original, unorganised form.


Key Takeaways

  • Raw data should be organised before it is analysed.
  • A frequency table shows how often each value occurs and is suitable for smaller datasets.
  • A grouped frequency table summarises larger datasets by placing values into class intervals.
  • A cumulative frequency table shows the running total of observations and is useful for determining how many values fall below a given point.
  • Organised data makes patterns easier to identify and prepares the data for graphs and statistical calculations.
  • Careful organisation improves the accuracy, clarity, and usefulness of statistical investigations.
 
 
 

5. Data Visualisation

Learning outcomes
  • I can construct appropriate statistical graphs.
  • I can choose suitable graphical representations.
  • I can compare different graph types.
  • I can interpret graphical displays.
  • I can communicate data effectively using graphs.

Introduction

Large tables of numbers can be difficult to understand, but graphs make patterns much easier to see. By displaying data visually, graphs help us identify trends, compare groups, recognise relationships, and communicate information clearly.

Different types of data require different types of graphs. A bar chart is useful for comparing categories, while a line graph shows changes over time. Choosing the correct graph is an important statistical skill because the wrong graph can make data confusing or even misleading.


What Is Data Visualisation?

Data visualisation is the presentation of data using graphs, charts, or diagrams.

It helps people:

  • Understand information quickly.
  • Identify patterns.
  • Compare values.
  • Detect trends.
  • Communicate results clearly.

Scientists, businesses, governments, and researchers all use data visualisation to present information effectively.


https://images.openai.com/static-rsc-4/5hzp42jzM3VLYcxyKFx979KWxwYa6J-obtuL-d73okP9uapvzfLyOoRZuXWLZycgSlxoOhC9pZQd_Zr7oC3BIRHKJzflYkzBRBzJ2XqEiJ2yz_xCC0OWmvCgrUZzO804YylvB-0I-S6AIll0nvt_dcYbAX0D74hjycxSbGow4AONcLMBfmE5jwK6zR5kupS-?purpose=fullsize
 
https://images.openai.com/static-rsc-4/BLwAxXmln6QNnkJOrRhB3CEoAI-U28-voc9Ra7GrqgamjAis6tWNEtYATdKchNm17arMwaVUHnLyJk2u7j1qU4wpOdMlKRn_LKRFLWKH8N1bvK5cPy1b7-BI3aqN0_WJ0lnLAarBGTqe0r5PLnOvQZ7p1rK3s96sgQQIKEYgdDJmJq57dsHcNRZWGuIgMLun?purpose=fullsize
 
https://images.openai.com/static-rsc-4/abS0QeB4XaGF8K7tzHf2rmHq7yGfdCx_7D21B-m1sMu0QbL21lNHolu_8mGSCNW7gg-h7OwrssJe8JB-ikxcjUbeTdiXA3bJZ0JuBf7OqxoV2ycZW5BHj4PqI5eY06uGU9fpTEJ-DSwtxxMEdCv27h19agGpvHGfy3_9Jrm_0S2oUwEBlYTj3A5IwV-dEnej?purpose=fullsize
5

Figure 1. Different graph types are used to display different kinds of data.


Choosing the Right Graph

Different graphs are designed for different purposes.

Graph Type.    Best Used For
Bar chart Comparing categories
Line graph Showing changes over time
Pie chart Showing parts of a whole
Histogram Continuous numerical data
Scatter plot Relationships between two variables

Selecting the correct graph makes the data easier to interpret.


Bar Charts

A bar chart compares quantities in different categories.

Features:

  • Separate bars.
  • Equal bar widths.
  • Categories on one axis.
  • Frequency or value on the other axis.

Examples:

  • Favourite sports.
  • Number of students in each grade.
  • Types of pets owned.

Bar charts are best for qualitative or discrete data.


https://images.openai.com/static-rsc-4/KlxFJaGGbKXJ-UvE49FHKDpHY4MwanVWwO_ZBb41x0mbLrRpQbs-adoJ8km71CwE4DZPfE4helxNbX5wDy73hlxwagnMwb3sRPoBxMbjl14hOICCl_EO-OCuV0RKJ22tm4eAeytZ6fbl8Q8q8a3sBs9m20rKkYjR1COZ_r9gs6uyvtPgBWkEvBSR5khaZY4t?purpose=fullsize
 
https://images.openai.com/static-rsc-4/Z7-AW6vyTsvY-Hikd-6Y2KdKvyGHKOCNg6jgyjI2OgFEE68rBJ_mN-bRAwXdaPIjbPCScczWYcZ7_hbgJniST264ydXtq7bp7-YKbkuqkZbZuLYUXJzQz7_QUAD4Ftbmf5380cSKTbOd2fLxrzI1NiA2LuiJnix3stxtaGoiIRUuta_q26k6XlqVDMkaTyeT?purpose=fullsize
 
https://images.openai.com/static-rsc-4/6EknmnIcBc_JZxuYkTyZW8-DvHycSZkRINOQqCQkHQz21RNmXPwYAH4P3xD-9aRJxThAnJ0JKiR8A4PYga05vtXv2ocbO3jU0aW2l0meNvsyEoarcD6STgk-48aWwQVjIYMuPO7nvtqeGCbFa0YCRsD4yFNHEr99dwr_WVPW41juygfAsyhaLrsboQ1t_2g2?purpose=fullsize
5

Figure 2. Bar charts compare values across different categories.


Line Graphs

A line graph shows how a quantity changes over time or another ordered variable.

Features:

  • Data points joined by lines.
  • Ordered horizontal axis.
  • Useful for identifying trends.

Examples:

  • Daily temperature.
  • Population growth.
  • Monthly rainfall.
  • Stock prices.

Line graphs are best for continuous data collected over time.


Pie Charts

A pie chart shows how a whole is divided into different parts.

Each slice represents a proportion of the total.

Examples:

  • Household spending.
  • Survey responses.
  • Market share.
  • Favourite school subjects.

All slices together represent 100% of the data.

Pie charts work best when there are only a few categories.


https://images.openai.com/static-rsc-4/y0RNvxkl8p9ScdLhiWsbSB08WscK3Mp8aGiiWJrVAdkaro-1pwGCAEsh_j4lqhpPJlRoKUk8AF_rkF8LLgHDxgDvENbbpqOU9hXkZ7FjG7lOHlJoxcGEISHP4fAceP3FC2v9lqmmYNCABJcO_Q3sxN6EJwScmSeRHfW5UCtSUe5b-Rng4cmiwQ8_rYo4pVrN?purpose=fullsize
 
https://images.openai.com/static-rsc-4/Yu_N6k3J4sw5yw9IFngqFXMJlKjEU7zKjxOPwxRGFYN4A5nkO7NmK964qRog7_VP9zD4pFXv-0G4QcPqnfvLBECYqFiFI0eeOf99ZMTJ2efXIfAL1-YwGTcZNkn7-OJgTWSneWCR8W45IavvDhRqYXRf2qZuSZwODNNV7TeYAM1mwrLmsnyiy6ICVyASErUc?purpose=fullsize
 
https://images.openai.com/static-rsc-4/Ldg3lsa5TYPLqT7GWNVr9mnsJbRWWK9Pk2dtEyzJ43vhwlogJHgRYV8-6_JoKLqJHc_oYjYl9yjmSzdYLCeZ8iCaK9kfIV8os6PKEjnvjdlt5ZSJmE2LtI61SdvuuCmpVD3-yvjrxYJTqCKvviYICMJ0EzhEESpbNFRPxdyfm84hj6xN3bcr3kkJIy929Rtp?purpose=fullsize
5

Figure 3. Pie charts show how categories contribute to a whole.


Histograms

A histogram displays the distribution of continuous data.

Unlike a bar chart:

  • Bars touch each other.
  • Data is grouped into class intervals.
  • The horizontal axis represents numerical ranges.

Examples:

  • Heights.
  • Test scores.
  • Masses.
  • Ages.

Histograms help reveal the shape of a distribution.


Scatter Plots

A scatter plot shows the relationship between two numerical variables.

Each point represents one observation.

Examples:

  • Height versus weight.
  • Hours studied versus exam score.
  • Temperature versus electricity use.

Scatter plots help identify:

  • Positive relationships.
  • Negative relationships.
  • No relationship.

https://images.openai.com/static-rsc-4/_T9q2fvEROSAzrXRvHYYNo0BaEMpTSMTBhj2dbsJS3g0-WDZqlTNQ14txJzSuRDcGeVww44SWDsuWSxqRyNf6OAHRfD0sED5EuDAeHgEaiYcWcv612SkGCEoelreG5H2TfyLS0CEVfbdQcXuUYn2lhta3Eis4nMA7NHuugJU395WZgk6M49SAxnD5ziwMs2y?purpose=fullsize
 
https://images.openai.com/static-rsc-4/mOxVBx7ML5f-Hns0KNcLqFEJux-IDBplkMsWOXxBj-w6YqDsgxUIphFjxi-D8SNg88tLKgAcC0NckpzJ_3gBaKMgoJh_cpg6u9ZNVxsMSUlbFsLfesklLwQHkZA4cEU186eZQhCr9vVFJd6E9nXwIVo0ebRaeWwhSfomXe6SiQmml4AkweJY7T06UmF81w-j?purpose=fullsize
 
https://images.openai.com/static-rsc-4/015XVmneEWRf3E4t0GyE_esEPVeYShE7g1ki_cdUHbPaEbhz4CdaKrvbdzXDufD9YR-YodweAtSH9OtLzUGEzNCfvvcamAaJFQVyHT8drbUTBsuvLJdph7jrfilfYT-p-lqI_nZujBLF5D2HhkDUFG5knTbMWqQ6IHeDIvdavbg97oklsUWpmLxyQ25iU9xH?purpose=fullsize
5

Figure 4. Scatter plots help identify relationships between two numerical variables.


Comparing Common Graph Types

Graph Best For Data Type
Bar chart Comparing categories.   Qualitative or discrete
Line graph Trends over time Continuous
Pie chart Parts of a whole Categorical
Histogram Distribution Continuous
Scatter plot.   Relationships Two numerical variables

Each graph highlights different features of the data.


Constructing Good Graphs

A well-designed graph should include:

  • A clear title.
  • Labelled axes.
  • Appropriate units.
  • Even scales.
  • Accurate plotting.
  • A legend (if needed).

Graphs should be neat and easy to read.


Interpreting Graphs

When reading a graph, ask questions such as:

  • Which category is largest?
  • Is there an increasing or decreasing trend?
  • Are there any unusual values?
  • Is there a relationship between variables?
  • What conclusions can be drawn?

Always consider the scale and labels before interpreting the data.


Avoiding Misleading Graphs

Graphs can be misleading if they use:

  • Unequal intervals.
  • Missing labels.
  • Truncated axes.
  • Inappropriate graph types.
  • Distorted scales.

Accurate graphs communicate information honestly.


https://images.openai.com/static-rsc-4/Wm1Fw45e5OUQb9ImEqvrxPihBG-ZLbwoOouhs0mCX5PYJW7subTr3EbR8SJNHbCRH0zgLm4yis1fMgVFGzXTAxSaeei7f-Ip_N4TRDBr4K3OsFFQaeUr6jL_k5RCmzr9XCxYHvo8B9_nt5IMn_PWnxZ0mLpbo-98c6wR_fptoUf70bSV7A7LZSJZZdUYwGKV?purpose=fullsize
 
https://images.openai.com/static-rsc-4/DJT2I4KXSqfvelX0dDZkn1aMUfk0_3THaob6w2WLduxYPapzVuwRqe8xE3a1z4h2Y-VaUgEAzlFi4QBrOjfRRavi_n5T6LXTg2YwuELcdASCWiEVWHipMkt9ix2QO35jhXUuO-ofPU4b7hzRGi1n2J7xYH1IBIkU6JM3NFmLZ89xTp-4BeCaBP6HOxFBtuR8?purpose=fullsize
 
https://images.openai.com/static-rsc-4/9ACQluhTdoYpzqsRjOrzxqLSv25rZPlsQnHX8bEfXT5X-VfPXPxgw-9tc9ax8KjIqXXCQRLYbkOZBbPrx-KU-GI5dyaILQUjKEfCZecxGlVoW6kfnDULOkfEBDVqP4rVc29Cw6b9R4lImSdWiAmfpkrZ6rrxPaWNTNVY5gxj029D_EUdFf8T5NftEuPxuFPD?purpose=fullsize
5

Figure 5. Clear scales, labels, and graph choices help prevent misleading interpretations.


Communicating Data Effectively

Graphs help communicate information because they:

  • Summarise large datasets.
  • Highlight important patterns.
  • Make comparisons easier.
  • Support conclusions.
  • Improve presentations and reports.

Choosing the correct graph allows others to understand the results quickly.


Worked Example

Question

Choose the most appropriate graph for each dataset.

Dataset Best Graph
Favourite fruit of students ?
Daily temperature for one month ?
Heights of 200 students ?
Hours studied versus exam marks.   ?

 

 

Solution

Dataset Best Graph
Favourite fruit Bar chart
Daily temperature Line graph
Heights of students Histogram
Hours studied versus exam marks Scatter plot

Real-World Connection

Weather services use line graphs to show changes in temperature over time, while governments often use bar charts to compare population sizes between regions. Scientists commonly use scatter plots to investigate relationships between variables, such as the connection between air pollution and respiratory illnesses. Choosing the correct graph helps people understand complex information quickly and make informed decisions.


Did You Know?

Many misleading graphs exaggerate differences by starting the vertical axis well above zero. A small increase can then appear much larger than it really is. Whenever you look at a graph, always check the scale before drawing conclusions.


Key Terms

Bar chart – A graph that compares values across different categories using separate bars.

Data visualisation – The presentation of data using graphs, charts, or diagrams.

Histogram – A graph showing the distribution of continuous data using touching bars.

Legend – A key that explains the symbols, colours, or patterns used in a graph.

Line graph – A graph showing changes in a variable over an ordered scale, usually time.

Pie chart – A circular chart showing how categories make up a whole.

Scatter plot – A graph that displays the relationship between two numerical variables.

Trend – The general direction in which data changes over time or across values.


Key Takeaways

  • Data visualisation presents information in graphical form, making patterns and relationships easier to understand.
  • Different datasets require different graphs: bar charts compare categories, line graphs show trends, pie charts display proportions, histograms show distributions, and scatter plots reveal relationships.
  • Good graphs include clear titles, labelled axes, appropriate scales, units, and legends where necessary.
  • Interpreting graphs involves identifying patterns, trends, comparisons, and unusual values.
  • Poor graph design or inappropriate scales can produce misleading conclusions.
  • Choosing the most appropriate graph helps communicate statistical information accurately and effectively.