Collecting and Organizing Data
5. Evaluating Data Collection
Learning outcomes
- I can identify sources of bias in data collection.
- I can explain how sample size affects reliability.
- I can distinguish between reliable and unreliable data.
- I can evaluate the validity of a data collection method.
- I can suggest improvements to data collection procedures.
Can We Trust the Data?
Imagine a school wants to answer the question:
"What is the most popular school lunch?"
A student asks five friends.
All five choose:
pizza.
The student concludes:
"Pizza is the most popular lunch in the entire school."
Is that conclusion justified?
Probably not.
The problem is not necessarily the calculation. The problem is:
how the data were collected.
Before accepting conclusions based on data, we should ask:
- Who was surveyed?
- How were participants selected?
- How many people participated?
- Were the questions fair?
- Were measurements made consistently?
- Does the sample represent the population?
- Could bias have influenced the results?
Evaluating data means thinking critically about:
where the data came from.
Population and Sample
Two important terms in data collection are:
population
and:
sample.
The population is the entire group we want information about.
The sample is the smaller group from which data are actually collected.
For example:
Population: all 800 students at a school
Sample: 80 students surveyed
Researchers often use samples because collecting information from every member of a population can be:
expensive, difficult, or time-consuming.
A Representative Sample
A good sample should be reasonably:
representative of the population.
This means that important characteristics of the population should not be systematically excluded or overrepresented.
Suppose a school contains students from Grades 6–12.
If researchers survey only:
Grade 12 students
the sample may not represent:
the entire school.
A larger sample is not automatically good if it is selected in a:
biased way.
What Is Bias?
Bias is a systematic influence that tends to push data or conclusions in a particular direction.
Bias can occur during:
- participant selection
- question design
- measurement
- observation
- recording
- reporting
Bias does not necessarily mean someone deliberately:
cheated.
It can occur unintentionally because of a poorly designed:
data collection method.
Sampling Bias
Sampling bias occurs when some members of a population are more likely to be included than others in a way that affects representativeness.
Suppose researchers want to estimate how much exercise students get.
They survey students leaving:
the school gym.
This group may exercise more than the average student.
The sample is therefore likely to be:
biased.
Example: The School Cafeteria
Suppose a school wants to know whether students like cafeteria food.
The survey is conducted only among students who:
eat in the cafeteria every day.
There is a problem.
Students who dislike cafeteria food may be more likely to:
bring food from home or eat elsewhere.
Their opinions could be underrepresented.
The data collection method therefore contains possible:
sampling bias.
Convenience Samples
A convenience sample consists of participants who are easy to reach.
Examples include:
- asking your friends
- surveying your own class
- questioning people standing nearby
- using the first people who respond
Convenience samples are:
easy and fast,
but they may not represent the:
target population.
Voluntary Response Bias
Imagine an online news website asks:
"Click here to tell us whether you support our new website design."
People choose whether to participate.
This is a:
voluntary response sample.
People with particularly strong opinions may be more likely to:
respond.
Therefore, the responses may not accurately represent all:
website users.
Non-Response Bias
Sometimes researchers select an appropriate sample, but many participants:
do not respond.
Suppose 500 people receive a survey but only:
75 respond.
If those 75 people differ systematically from the people who did not respond, the results may suffer from:
non-response bias.
The problem is not simply that responses are missing.
The problem is that the missing responses may not be:
random.
Question Wording Can Create Bias
The way a question is written can influence:
responses.
Consider:
Question A:
"Do you support extending the school lunch period?"
Compare it with:
Question B:
"Don't you agree that students deserve a longer and more relaxing lunch period?"
Question B encourages a particular:
answer.
This is called a:
leading question.
Neutral Questions
Good survey questions should avoid encouraging participants toward a particular response.
Instead of:
"How much do you enjoy our excellent new cafeteria menu?"
use:
"How satisfied are you with the new cafeteria menu?"
Possible responses might be:
- Very satisfied
- Satisfied
- Neither satisfied nor dissatisfied
- Dissatisfied
- Very dissatisfied
Neutral wording helps reduce:
response bias.
Loaded Questions
A loaded question contains assumptions or emotionally charged wording that may influence responses.
For example:
"Should the school stop wasting money on unnecessary decorations?"
The words:
"wasting" and "unnecessary"
already suggest a conclusion.
A more neutral question might be:
"Should the school increase, decrease, or maintain its current spending on decorations?"
Double-Barrelled Questions
A question should generally ask about:
one issue at a time.
Consider:
"Are you satisfied with the school's cafeteria food and prices?"
What if someone likes the food but thinks it is:
too expensive?
The respondent cannot answer accurately with a single:
yes or no.
Instead, ask two questions:
"How satisfied are you with the cafeteria food?"
and:
"How satisfied are you with cafeteria prices?"
Measurement Bias
Bias can also occur when measurements are collected:
incorrectly or inconsistently.
Suppose students measure plant height.
One student measures from:
the soil surface.
Another measures from:
the bottom of the pot.
The measurements are not being collected using the same:
method.
This reduces the quality of the data.
Instrument Problems
Data quality can also be affected by the measuring:
instrument.
Examples include:
- an incorrectly zeroed balance
- a damaged ruler
- an uncalibrated thermometer
- a stopwatch used inconsistently
- a sensor with insufficient precision
If an instrument consistently gives values that are too high or too low, it may introduce:
systematic error.
Observer Bias
Sometimes the person collecting data can influence the:
observations.
Suppose a researcher expects plants receiving fertilizer to grow better.
If plant health is judged only as:
"good" or "poor"
the researcher's expectations might unintentionally affect how plants are classified.
Using objective measurements such as:
height, mass, or leaf number
can help reduce this problem.
Recording Errors
Even if measurements are made correctly, errors can occur when data are:
recorded.
For example:
A measurement of:
12.6 cm
might accidentally be entered as:
126 cm.
This is not sampling bias.
It is a:
recording or transcription error.
Researchers should check unusual values and verify data before:
analysis.
What Is Sample Size?
The sample size is the number of observations or individuals included in a sample.
It is often written as:
n.
For example:
n = 10
means the sample contains:
10 observations.
Sample size can have a major effect on the:
reliability and precision of estimates.
Small Samples
Suppose we want to estimate the favourite sport of 1,000 students.
We ask:
3 students.
Their choices are:
Football
Football
Football
Can we confidently conclude that nearly everyone prefers football?
No.
Three students provide very limited information about a population of:
1,000.
Larger Samples
Now suppose we survey:
300 students
selected appropriately from different grades.
The sample includes students with different:
- ages
- classes
- interests
- backgrounds
This sample is more likely to provide a stable estimate of the population than a sample of:
three students.
However, size alone does not guarantee:
good data.
Bigger Does Not Automatically Mean Better
Suppose a school has 2,000 students.
Researchers survey:
1,000 members of the school football fan club
to determine the school's favourite sport.
That is a:
large sample.
But it is also:
biased.
Now suppose researchers randomly select:
300 students from the entire school.
The smaller sample may provide a much more useful estimate because its selection method is:
more representative.
The key idea is:
sample size and sampling method both matter.
Random Sampling
One way to reduce selection bias is:
random sampling.
In a simple random sample, each member of the population has an equal chance of:
being selected.
For example, researchers could assign each student a number and use a random-number generator to choose:
participants.
Random selection does not guarantee a perfect sample, but it helps reduce:
systematic selection bias.
Stratified Sampling
Sometimes a population contains important:
subgroups.
Suppose a school contains:
- 30% Grade 9
- 25% Grade 10
- 25% Grade 11
- 20% Grade 12
A researcher might deliberately select participants from each grade in similar proportions.
This is an example of:
stratified sampling.
This can help ensure that important groups are:
represented.
Reliability
In data collection, reliability relates to the consistency of measurements or results.
A reliable method tends to produce:
consistent results when repeated under similar conditions.
For example, suppose the same object is measured five times:
25.1 cm, 25.0 cm, 25.1 cm, 25.0 cm, 25.1 cm
These measurements show strong:
consistency.
Reliability Does Not Always Mean Accuracy
Consider a scale that always adds:
2 kg
to the true mass.
You weigh the same object repeatedly and obtain:
12.0 kg, 12.0 kg, 12.0 kg, 12.0 kg.
The results are highly:
consistent.
But if the true mass is:
10.0 kg,
they are not:
accurate.
A measurement system can therefore be:
reliable but inaccurate.
Accuracy
Accuracy refers to how close a measurement is to the accepted or true value, when such a value is meaningful and available.
Imagine the true length is:
20.0 cm.
Measurements of:
20.0, 20.1, 19.9 cm
are both reasonably consistent and:
accurate.
Measurements of:
24.9, 25.0, 25.0 cm
may be consistent but are:
inaccurate.
Reliability vs Validity
Reliability and validity are related, but they are not the:
same thing.
Reliability
Does the method produce consistent results?
Validity
Does the method actually measure or investigate what it is supposed to measure?
A method can be reliable without being:
valid.
Example: Measuring Fitness
Suppose researchers want to measure:
cardiovascular fitness.
They ask participants:
"How many sports do you enjoy watching?"
The answers might be recorded very reliably.
But the method is not a valid measure of:
cardiovascular fitness.
The data do not directly address the intended:
question.
Valid Data Collection
A valid data collection method should produce information relevant to the:
research question.
Suppose the question is:
"Does increasing light exposure affect plant growth?"
A reasonable method would involve:
- changing light exposure
- measuring plant growth
- keeping other important conditions similar
- using repeated measurements
- comparing results
Measuring the colour of the pots would not directly answer:
the investigation question.
Reliability in Experiments
Suppose a student measures the time required for a toy car to travel down a ramp.
One trial gives:
1.42 s.
Can we be confident in that measurement?
A better approach is to perform:
several trials.
For example:
1.42 s
1.38 s
1.40 s
1.41 s
1.39 s
Repeated trials allow us to assess:
consistency.
Repeated Measurements
Repeating measurements can help:
- identify unusual results
- estimate natural variation
- reduce the influence of random errors
- calculate a representative average
- improve confidence in the result
This is why repetition is common in:
scientific investigations.
Reliable and Unreliable Data
Consider two datasets measuring the same quantity.
Dataset A
15.1, 15.0, 15.2, 15.1, 15.0
Dataset B
12.4, 17.8, 14.1, 19.3, 11.5
Dataset A is much more:
consistent.
Dataset B shows much greater:
variation.
Before deciding why, we would need to consider the measurement method, natural variability, instruments, and experimental conditions.
Variation Is Not Automatically Bad Data
Not all variation means that data are:
unreliable.
Human height varies naturally.
Heart rate varies.
Weather varies.
Animal populations vary.
Biological measurements often contain substantial:
natural variation.
The important question is whether the data collection method is sufficiently consistent and appropriate to distinguish real variation from:
measurement problems.
Random Error
Random errors cause measurements to vary unpredictably.
Examples include:
- small reaction-time differences when using a stopwatch
- slight changes in environmental conditions
- uncertainty when reading a scale
- natural variation between samples
Repeating measurements can help reduce the influence of random error when estimating a:
typical value.
Systematic Error
A systematic error pushes measurements consistently in the same:
direction.
For example, a balance that reads:
+5 g
when empty could make every mass measurement too:
high.
Repeating the measurement does not automatically solve this problem.
The instrument needs to be:
checked, zeroed, or calibrated.
Random vs Systematic Error
| Random Error | Systematic Error |
|---|---|
| Causes unpredictable variation | Shifts results consistently |
| Can affect precision | Often affects accuracy |
| Repetition can help estimate/reduce its effect on an average | Repetition alone usually does not remove it |
| Example: stopwatch reaction time | Example: incorrectly calibrated balance |
Recognizing the difference helps us choose the correct:
improvement.
Evaluating a Data Collection Method
When evaluating a method, ask several questions.
Who?
Who provided the data?
How many?
Was the sample large enough for the purpose?
How?
How were participants or measurements selected?
What?
Does the data actually address the question?
How measured?
Were measurements collected consistently?
Repeated?
Were enough observations or trials collected?
Bias?
Could anything systematically influence the results?
These questions provide a useful framework for:
evaluation.
Worked Example 1: School Sleep Survey
A student wants to determine how much sleep teenagers get.
They ask:
six friends
how long they slept last night.
Problem 1: Sample size
Six people provide:
limited evidence.
Problem 2: Sampling method
Friends are a:
convenience sample.
Problem 3: Time period
One night may not represent:
usual sleep habits.
Improved method
Survey a larger, appropriately selected sample of teenagers and collect sleep information across:
several nights.
This would provide more useful evidence about:
typical sleep patterns.
Worked Example 2: Favourite School Subject
A school wants to determine the most popular subject.
Researchers survey:
200 students leaving a mathematics competition.
The sample is fairly:
large.
But it may still be:
biased.
Students attending a mathematics competition may have unusually positive attitudes toward:
mathematics.
Improvement:
Select students randomly or proportionally from:
the entire school population.
Worked Example 3: Plant Growth
A student investigates fertilizer and plant growth.
They use:
one fertilized plant
and:
one unfertilized plant.
After two weeks, the fertilized plant is taller.
Can the student confidently conclude that fertilizer caused the difference?
The evidence is:
weak.
Individual plants naturally vary.
A stronger experiment would use:
multiple plants in each condition.
Other important variables should also be controlled, such as:
- plant species
- starting size
- water
- light
- soil
- temperature
- growing time
Worked Example 4: Reaction Time
A student investigates whether caffeine affects reaction time.
They test themselves:
once before
and:
once after
drinking a caffeinated beverage.
Potential problems include:
- extremely small sample
- only one trial per condition
- practice effects
- expectations
- uncontrolled conditions
A stronger design could include more participants, repeated trials, standardized conditions, and appropriate comparison procedures.
Validity and Fair Tests
In experiments, validity often depends on controlling variables that could otherwise provide:
alternative explanations.
Suppose we investigate whether temperature affects reaction rate.
If we also change:
concentration
then we cannot easily determine whether the change in rate resulted from:
temperature or concentration.
A valid experiment should isolate the factor being:
investigated, as far as practical.
Control Variables
Control variables are factors deliberately kept as consistent as possible.
For a plant-growth experiment, they might include:
- plant species
- amount of water
- soil type
- pot size
- temperature
Controlling important variables helps make the comparison:
fairer and more interpretable.
Improving Sample Size
A common improvement is:
"Use a larger sample."
But this explanation should go further.
Instead of writing:
Use more people.
write:
Increase the sample size so that individual unusual responses have less influence and the sample provides a more stable estimate of the population.
Good scientific evaluation explains:
why the improvement helps.
Improving Sampling
Instead of:
"Ask different people."
write:
Use a random or appropriately stratified sample from the target population to reduce selection bias and improve representativeness.
Again, the improvement should address a specific:
weakness.
Improving Measurements
Instead of:
"Measure better."
write:
Use the same calibrated measuring instrument and measurement procedure for every trial.
Or:
Repeat each measurement several times and calculate an appropriate average.
Specific improvements are more useful than:
vague suggestions.
Match the Improvement to the Problem
| Problem | Possible Improvement |
|---|---|
| Sample too small | Increase sample size |
| Convenience sample | Use random or stratified sampling |
| Leading question | Rewrite using neutral wording |
| Inconsistent measurement | Standardize the procedure |
| Instrument offset | Calibrate or zero the instrument |
| One experimental trial | Repeat trials |
| Uncontrolled variable | Keep important variables constant |
| Recording mistakes | Check and verify data entries |
| High non-response | Improve follow-up or survey accessibility |
The best improvement directly addresses the:
identified limitation.
Correlation Does Not Automatically Mean Causation
Suppose data show that students who exercise more also report:
better sleep.
This demonstrates an:
association in the collected data.
It does not automatically prove that exercise caused the difference.
Other factors could be involved.
Evaluating data means being careful not to make conclusions that are stronger than:
the evidence supports.
Data Can Be Accurate but Unrepresentative
Imagine researchers accurately record the opinions of:
500 professional athletes.
The measurements may be perfectly:
accurate.
But if the question concerns the opinions of:
all adults,
the sample may not be representative.
Accurate recording cannot fix a poor:
sampling design.
Data Can Be Representative but Poorly Measured
The opposite can also occur.
Researchers might select an excellent random sample but use:
confusing survey questions.
Or they might use:
poorly calibrated equipment.
Good data collection requires attention to both:
sampling and measurement.
Primary and Secondary Data
Primary data are collected directly for a particular investigation.
Examples:
- conducting your own survey
- measuring plant height
- timing a moving object
Secondary data were collected previously by someone else.
Examples:
- government statistics
- published research
- historical records
- databases
Both can be useful, but both should be:
evaluated critically.
Evaluating Secondary Data
When using existing data, ask:
- Who collected it?
- Why was it collected?
- When was it collected?
- How was it collected?
- What population was studied?
- How large was the sample?
- Are definitions and units clear?
- Is the source credible?
- Is the information current enough for the question?
Never assume data are trustworthy simply because they appear:
online.
Data Collection in Science
Scientists spend considerable effort designing how data will be:
collected.
Good scientific data collection may involve:
- repeated trials
- standardized procedures
- calibrated instruments
- control groups
- randomization
- sufficient sample sizes
- careful recording
- uncertainty estimates
Strong conclusions depend on the quality of the:
evidence.
Data Collection in Everyday Life
The same principles apply outside laboratories.
Consider:
product reviews
A few extremely positive reviews may not represent all customers.
online polls
Participants may choose themselves.
fitness trackers
Sensors can have measurement limitations.
news surveys
Sampling methods matter.
advertisements
Companies may present selected statistics.
Being able to evaluate data collection is therefore an important part of:
data literacy.
Spotting Warning Signs
Be cautious when you see claims based on:
very small samples
unknown samples
self-selected participants
leading questions
unclear measurement methods
single trials
missing comparison groups
unexplained exclusions
unidentified data sources
These do not automatically make a conclusion false.
They mean we should be more cautious about:
how strongly the evidence supports it.
A Useful Evaluation Structure
When evaluating data collection, use:
Weakness → Effect → Improvement
For example:
Weakness: Only five students were surveyed.
Effect: The sample may not adequately represent the school population, and the estimate may be unstable.
Improvement: Survey a larger sample selected from across the school.
Another example:
Weakness: The question used emotionally loaded language.
Effect: Participants may have been encouraged toward a particular response.
Improvement: Rewrite the question using neutral wording.
This structure produces much stronger:
scientific evaluations.
The BIAS Check
A simple way to evaluate data collection is to remember:
B — Bias
Could the method systematically favour certain outcomes?
I — Individuals
Who was included, and do they represent the population?
A — Amount
Was enough data collected?
S — System
Was the data collected using a consistent and appropriate method?
This provides a quick first check of:
data quality.
Worked Example 5: Evaluate the Investigation
A student wants to determine the average amount of screen time among students at a school.
They ask:
10 members of the gaming club
how many hours they spend using screens each day.
The average is:
7.2 hours.
Evaluation
Sample size:
Ten students is a relatively small sample for estimating the whole school.
Sampling bias:
Gaming-club members may use screens differently from the general student population.
Validity:
Screen time is relevant to the research question, but the sampling method limits how well the result represents the whole school.
Improvement:
Survey a larger, randomly or appropriately stratified sample of students from different grades and activities.
Therefore, the calculated average may accurately describe those ten participants, but it should not automatically be generalized to:
the entire school.
Worked Example 6: Evaluate an Experiment
A student investigates whether water temperature affects how quickly sugar dissolves.
They use:
cold water in a 100 mL beaker
and:
hot water in a 250 mL beaker.
They stir the hot water continuously but do not stir the cold water.
They perform:
one trial.
There are several problems.
The student changed:
- temperature
- container
- stirring
Therefore, temperature is not the only factor that differs.
Only one trial also provides limited evidence about:
repeatability.
A better investigation would:
- use identical containers
- use equal water volumes
- use equal sugar masses
- standardize stirring
- change only temperature
- repeat each condition several times
This would improve both the:
validity and reliability of the investigation.
Reliable Data Support Stronger Conclusions
Good data do not guarantee that every interpretation will be:
correct.
However, weak data severely limit the conclusions we can reasonably:
draw.
The strength of a conclusion should match the strength of:
the evidence.
That is one of the central principles of scientific and statistical:
reasoning.
Check Your Understanding
1. What is the difference between a population and a sample?
2. What does it mean for a sample to be representative?
3. Define bias.
4. What is sampling bias?
5. Why might surveying your friends produce biased results?
6. What is voluntary response bias?
7. Explain how a leading question can affect data.
8. Rewrite this question more neutrally:
"Don't you agree that our excellent new school schedule is better?"
9. What is sample size?
10. Why can a larger sample improve an estimate?
11. Why does a large sample not automatically guarantee good data?
12. Explain the difference between reliability and validity.
13. Can measurements be reliable but inaccurate? Explain.
14. Explain the difference between random and systematic error.
15. Why are repeated measurements useful?
16. A survey of school exercise habits is conducted only among sports-team members. Identify the problem.
17. A thermometer consistently reads 3°C too high. What type of problem is this?
18. A student performs an experiment only once. Suggest an improvement and explain why it helps.
19. Why should important control variables be kept constant?
20. Use Weakness → Effect → Improvement to evaluate a survey that asks five students from one class about the opinions of an entire school.
Key Terms
- Population: Entire group about which information is wanted.
- Sample: Smaller group from which data are collected.
- Sample size: Number of individuals or observations in a sample.
- Representative sample: Sample that reasonably reflects relevant characteristics of the target population.
- Bias: Systematic influence that tends to push data or conclusions in a particular direction.
- Sampling bias: Bias caused by how participants or observations are selected.
- Convenience sample: Sample chosen because participants are easy to access.
- Voluntary response: Sampling method in which individuals choose whether to participate.
- Non-response bias: Bias that can occur when responders differ systematically from non-responders.
- Leading question: Question worded in a way that encourages a particular response.
- Loaded question: Question containing assumptions or emotionally influential language.
- Double-barrelled question: Question asking about more than one issue at once.
- Random sample: Sample selected using a random process.
- Stratified sample: Sample constructed to represent important subgroups.
- Reliability: Consistency of measurements or results under similar conditions.
- Validity: Extent to which a method appropriately measures or investigates what it is intended to.
- Accuracy: Closeness of a measurement to an accepted or true value where applicable.
- Random error: Unpredictable variation between measurements.
- Systematic error: Consistent shift in measurements caused by a method or instrument.
- Control variable: Factor deliberately kept consistent during an investigation.
- Primary data: Data collected directly for the current investigation.
- Secondary data: Data previously collected by another source.
- Calibration: Checking or adjusting an instrument against an appropriate reference.
Key Takeaways
- The quality of a conclusion depends strongly on how the data were collected.
- A population is the complete group of interest, while a sample is the smaller group actually studied.
- Samples should be selected so they reasonably represent the target population.
- Bias is a systematic influence that can distort data or conclusions.
- Sampling bias occurs when some members of the population are systematically more likely to be represented than others.
- Convenience samples are easy to collect but can be unrepresentative.
- Voluntary-response surveys can overrepresent people with strong opinions.
- Non-response can create bias when responders differ systematically from non-responders.
- Leading and loaded questions can influence participants' answers.
- Questions should generally be neutral, clear, and focused on one issue.
- Measurement procedures should be consistent.
- Instruments should be appropriate, checked, and calibrated when necessary.
- Sample size affects how stable and precise estimates can be.
- Larger samples generally provide more information, but size cannot repair a fundamentally biased sampling method.
- Random and stratified sampling can help improve representativeness.
- Reliability concerns the consistency of results.
- Validity concerns whether the method appropriately addresses the intended question.
- Reliable measurements are not necessarily accurate.
- Random errors produce unpredictable variation.
- Systematic errors consistently shift measurements in a particular direction.
- Repeated measurements help assess consistency and reduce the influence of random variation on estimates.
- Repetition alone does not remove systematic error.
- Experimental validity is improved by controlling important variables.
- Natural variation should not automatically be mistaken for unreliable data.
- Secondary data should be evaluated by examining their source, date, purpose, population, and collection method.
- Strong evaluations identify a specific weakness, explain its effect, and propose a realistic improvement.
- The strength of a conclusion should match the quality and quantity of the evidence available.