Nurses Revision

Sampling Methods, Disease Rates and Surveys

Sampling Methods, Disease Rates and Surveys

Sampling Methods, Disease Rates and Surveys
Learning Outcomes

By the end of this session, you should be able to:

  • Explain why sampling is used in epidemiology and biostatistics.
  • Describe simple random, systematic, stratified, and cluster sampling.
  • Identify strengths and limitations of non-probability sampling.
  • Calculate simple incidence, prevalence, ratios, and proportions.
  • Interpret disease measures for practical nursing decisions.

🧠 Core Question: "If we cannot study everyone, how do we select people fairly?" This is the central challenge of sampling. The answer determines whether your findings are trusted or dismissed.

Session 1: Introduction to Sampling
Why Do We Sample?

We sample because studying an entire population is usually impossible, impractical, or unnecessary. Here is why sampling is essential:

  • A population may be too large to study completely. You cannot interview all 40 million Ugandans about malaria knowledge. But you can interview 400 carefully selected people and learn a great deal.
  • Sampling saves time, money, and staff effort. A census (studying everyone) takes years and costs millions. A well designed survey takes weeks and costs thousands.
  • A good sample gives useful information about the wider group. If the sample truly represents the population, the findings apply to everyone not just those interviewed.
  • Poor sampling can produce misleading findings. If you only survey clinic attenders, you will overestimate service use. If you only survey urban areas, you will miss rural realities.

⚡ Golden Rule: Good sampling is not about studying many people only; it is about studying the right people. A sample of 80 well chosen mothers is more valuable than a sample of 800 poorly chosen ones.

From Population to Sample: The Flow
  • TARGET POPULATION: The full group we want to understand
  • SOURCE POPULATION: The accessible subset we can reach
  • SAMPLING FRAME: The list or method to identify eligible people
  • SELECTED SAMPLE: The smaller group actually studied
  • COLLECTED DATA: The information we analyse and interpret
Key Sampling Terms You Must Know
Term Definition & Example
Sampling Unit The individual person, household, school, or facility that is selected. Example: One mother with a child under one year.
Eligibility / Inclusion Criteria The rule that says who can be included in the study. Example: "Mothers with children aged 0-11 months living in the catchment area for at least 6 months."
Representativeness How well the sample reflects the characteristics of the whole population. A representative sample has the same age, sex, and socioeconomic distribution as the population.
Sampling Error The natural difference between a sample statistic and the true population value. Even a perfect random sample will not exactly match the population but the error is predictable and measurable. Larger samples have smaller sampling error.
Sampling Bias Systematic error caused by poor selection methods. Bias means the sample consistently overrepresents or underrepresents certain groups. Unlike sampling error, bias does not decrease with larger sample size.
📝 Exam Tip Sampling Error vs. Sampling Bias: This is a favourite exam distinction. Sampling error is random and natural it happens even with perfect methods. Sampling bias is systematic and caused by poor methods it happens because you selected the wrong way. Error can be reduced by increasing sample size. Bias can only be reduced by improving the sampling method.
Representativeness Matters: Weak vs. Strong Samples
❌ Weak Sample (Biased) ✅ Stronger Sample (Representative)
Only easy to reach households (those near the road) Includes different villages, including remote ones
Only clinic attenders (already using services) Uses household sampling to find non-attenders too
Excludes remote villages (no transport to reach them) Allocates resources to reach remote areas
Findings may be biased and not generalisable Findings are more credible and applicable to the whole population
The Sampling Frame

A sampling frame is the practical list or source from which sampling units are selected. It is the bridge between the theoretical population and the actual sample.

Examples of sampling frames:

  • Village register (list of all households).
  • School attendance list.
  • ANC (Antenatal Care) register at a health facility.
  • Facility list of all health centres in a district.
  • Household list from a recent census.

A weak frame is dangerous:

  • It may leave out eligible people (e.g., a village register that was last updated 3 years ago misses new households).
  • It may include ineligible people (e.g., the ANC register includes women who have since moved away or delivered).
  • It may be incomplete (e.g., no register exists for informal settlements).

⚠️ Critical Rule: A sample cannot be better than the frame used to select it. If your frame is missing half the population, your sample will miss them too no matter how fancy your randomisation method is.

Sampling Bias in Practice

🩺 Example: A survey about immunisation interviews only mothers who came to the clinic today.

  • Problem: It may miss mothers whose children are most likely to have missed vaccines the very group the survey wants to understand. Mothers who do not come to clinic may be the ones with transport barriers, misinformation, or cultural objections.
  • Likely effect: Coverage may appear higher than it truly is in the community. The survey concludes "90% coverage" when the real coverage is 60%.
  • Better approach: Sample households or use outreach lists across the entire catchment area. Include mothers who have never attended the clinic.
Mini Case: Immunisation Survey

🩺 The Situation: A health centre serves 1,200 mothers with children under one year. The team wants to interview 120 mothers about missed vaccines. They have village registers from 10 villages.

Task: How should they select mothers fairly?

Step by step thinking:

  • Population: 1,200 mothers with children under one year in the catchment area.
  • Sampling frame: Village registers from 10 villages. Check: Are the registers complete? Do they include all mothers? Are they up to date?
  • Sample size: 120 mothers (10% of the population a reasonable proportion for a survey).
  • Selection method:
    • Option A Simple Random Sampling: Combine all 10 village registers into one master list of 1,200 mothers. Number them 1 to 1,200. Use a random number table or computer to select 120 numbers. Interview those mothers.
    • Option B Systematic Sampling: Sampling interval = 1,200 ÷ 120 = 10. Choose a random start between 1 and 10 (e.g., 7). Then select every 10th mother: 7, 17, 27, 37... up to 1,197.
    • Option C Stratified Sampling: If some villages are much larger or poorer than others, divide the 120 sample proportionally by village size. Sample 12 from each village if equal, or proportionally if unequal. This ensures no village is overrepresented or underrepresented.
  • Possible bias: Village registers may miss mothers who recently moved in, or mothers who live in informal settlements not on any register. The team should plan for "non response" what if a selected mother is not home? Have a replacement rule (e.g., interview the next household) or revisit later.
Session 2: Probability Sampling

Core question: How do we give eligible people a known chance of selection?

Probability sampling means every eligible unit has a known, non zero chance of being selected. The selection uses a random or rule based method. This is the gold standard for surveys that need to estimate population levels.

📝 Exam Tip: In an exam, if you are asked to design a survey that estimates prevalence or compares groups, always choose a probability sampling method. Non-probability methods are only acceptable for exploratory or qualitative work.
Simple Random Sampling (SRS)

How it works:

  • Make a complete list of all eligible units in the population.
  • Number all eligible units (1 to N).
  • Use a random number table, lottery, or computer to select the required sample size.
  • Every unit has an equal chance of being selected.

Example: A health centre has a complete ANC register of 500 mothers. The team needs to select 100 for a satisfaction survey. They number the mothers 1-500, use a random number generator to pick 100 numbers, and interview those mothers.

Advantages:

  • Simple to understand and explain.
  • Every unit has equal chance no subgroup is favoured or ignored.
  • Statistical formulas work perfectly (standard errors, confidence intervals).

Limitations:

  • Requires a complete list of the population often unavailable in community settings.
  • Can be expensive and logistically difficult if selected units are scattered across a wide area.
  • May miss small subgroups by chance (e.g., only 2 elderly people selected in a sample of 100).
Systematic Sampling

How it works:

  • Calculate the sampling interval (k) = population size ÷ sample size.
  • Choose a random starting point between 1 and k.
  • Select every kth unit from the ordered list.
Systematic Sampling Formula: k = N ÷ n
Where N = population size, n = sample size, k = sampling interval

Example: 1,000 households ÷ 100 = k = 10. Random start = 3 (chosen between 1 and 10). Selected households: 3, 13, 23, 33, 43... 993.

Advantages:

  • Easier to implement than simple random sampling no need for a random number table.
  • Spreads the sample evenly across the list.
  • Often used in community surveys with household lists.

Limitations:

  • Hidden patterns in the list can introduce bias. Example: If a list is ordered by household head (male, female, male, female...) and k = 2, you might select only males or only females.
  • If the list has a periodic pattern that matches k, the sample is not random.
  • Less flexible than simple random sampling if you need to adjust mid study.

⚠️ Watch Out For: Always check the ordered list for hidden patterns before using systematic sampling. If the list is ordered by age, sex, or village in a repeating pattern, systematic sampling may be biased. In that case, use simple random or stratified sampling instead.

Stratified Sampling

How it works:

  • Divide the population into subgroups (strata) based on important characteristics. Strata should be mutually exclusive and collectively exhaustive (everyone fits in one and only one stratum).
  • Sample from each stratum separately. You can use simple random or systematic sampling within each stratum.
  • Combine the samples from all strata to form the total sample.

Common strata in health surveys:

  • Sex (male / female).
  • Age group (under 5, 5-14, 15-49, 50+).
  • Village or urban/rural.
  • School or facility type.
  • Socioeconomic status (wealth quintile).

Example: A district has 10 villages. 3 are near the main road (urban like), 7 are remote (rural). If you sample randomly, you might by chance select mostly road side villages. Instead, stratify by location: sample 30 from road side villages and 70 from remote villages, proportional to their population sizes.

Advantages:

  • Ensures representation of all important subgroups.
  • Allows separate analysis for each stratum (e.g., compare urban vs. rural vaccination rates).
  • More precise than simple random sampling when strata are internally similar but different from each other.

Limitations:

  • Requires knowledge of the population structure before sampling.
  • More complex to plan and analyse.
  • If strata are chosen poorly, it adds complexity without benefit.
Cluster Sampling

How it works:

  • First stage: Divide the population into natural groups called clusters (villages, schools, zones, parishes).
  • Second stage: Randomly select some clusters (not all).
  • Third stage: Within selected clusters, sample all individuals or a random subset.

Example: A district has 50 villages. You need to survey 500 households. Instead of listing all households in all 50 villages (impossible), you randomly select 10 villages, then survey 50 households in each selected village.

Advantages:

  • Practical and cheap for large, dispersed populations. No need for a complete list of all individuals.
  • Reduces travel costs interviewers stay in one area rather than travelling across the entire district.
  • Widely used in national surveys (DHS, MICS, SMART surveys).

Limitations:

  • People within a cluster tend to be similar (homogeneous). This increases sampling error compared to simple random sampling.
  • To compensate, you need a larger sample size than simple random sampling.
  • If clusters are selected poorly (e.g., only easy to reach villages), bias is introduced.
💡 Key Point: Cluster sampling is practical for community surveys, but people within a cluster may be similar. This is called the design effect (DEFF) a statistical penalty for using clusters. In exams, know that cluster sampling is cheaper but less precise than simple random sampling.
Choosing a Probability Method: Decision Guide
Situation Best Method Why
Complete list of all individuals exists Simple Random Sampling Every unit has equal chance; most statistically pure.
Ordered list exists; no hidden patterns Systematic Sampling Easy to implement; spreads sample evenly.
Subgroups differ in risk or access; equity matters Stratified Sampling Ensures all subgroups are represented; allows subgroup comparison.
Population is large and dispersed; no complete list of individuals Cluster Sampling Practical and cost-effective; only need lists of clusters (villages, schools).
Equity is important; want to compare urban vs. rural Stratified + Cluster First stratify by location, then cluster-sample within each stratum. Common in national surveys.
Class Activity: Pick the Method

Scenario A: 600 ANC clients are listed in a register; select 60.

Answer: Simple Random Sampling or Systematic Sampling. A complete list exists, so either works. Systematic might be easier: k = 600 ÷ 60 = 10. Random start between 1-10, then every 10th client.

Scenario B: A district has 12 villages; fieldwork can visit only 4 villages.

Answer: Cluster Sampling. Villages are the clusters. Randomly select 4 of 12 villages, then survey all or a sample of households within those 4. This is practical because visiting all 12 villages is too expensive.

Scenario C: The study must compare males and females fairly.

Answer: Stratified Sampling. Divide the population into male and female strata. Sample proportionally from each stratum. This guarantees enough males and females for statistical comparison.

Session 3: Non Probability Sampling

Core question: When random selection is not possible, what are the trade offs?

Non probability sampling means selection does not give every eligible person a known chance of being selected. It is faster, cheaper, and useful for hard to reach groups but it carries a higher risk of bias and weaker generalisation.

⚠️ Critical Rule: Use non probability sampling carefully and describe its limitations honestly in any report. Never claim that a convenience sample represents the whole population.

Convenience Sampling
  • Meaning: Select those who are easiest to reach.
  • When it is used: Pilot studies, practice exercises, rapid assessments, student research with limited time.
  • Example: A nursing student interviews patients sitting in the clinic waiting room because they are available right now.
  • Limitations:
    • May exclude the absent or remote the people who most need to be heard.
    • Can overrepresent service users people already in the clinic are not the same as people who never come.
    • Weak for population estimates. You cannot say "30% of the district has hypertension" based on a convenience sample of clinic patients.
Purposive Sampling
  • Meaning: Select people because they have specific knowledge or experience. The researcher deliberately chooses participants who can provide rich, relevant information.
  • When it is used: Qualitative research, key informant interviews, expert consultations, programme evaluations.
  • Example: Interviewing TB focal persons about case detection challenges, or interviewing traditional birth attendants about home delivery practices.
  • Quality depends on: Clear, transparent selection criteria. The researcher must explain why each person was chosen.
💡 Key Question for Purposive Sampling: "Who can provide the information needed?" Not "Who is easiest to find?" but "Who knows what we need to know?"
Quota Sampling
  • Meaning: Set required numbers for categories before data collection, then fill each quota with convenient respondents.
  • How it works: Decide you need 50 men and 50 women. Interview the first 50 men and 50 women you meet who fit the criteria.
  • Example: A rapid assessment in a market decides to interview 20 vendors, 20 shoppers, and 20 passers by.
  • Limitation: Quota controls numbers in groups, but it does not remove selection bias by itself. The interviewer still chooses which men and women to interview usually the easiest to approach. It is non random unless selection within quotas is randomised.
Snowball Sampling
  • Meaning: Initial participants help identify other eligible participants. Like a snowball rolling downhill it grows as it goes.
  • When it is used: Hidden or hard to reach populations: commercial sex workers, drug users, undocumented migrants, men who have sex with men, people with rare diseases.
  • Example: A researcher interviews one person living with HIV who then introduces 3 others in their support group, who each introduce more.
  • Limitations:
    • May overrepresent connected social networks. If the first participant only knows people from one church or one neighbourhood, the sample is biased.
    • Weak for estimating true population prevalence. You cannot calculate how common a behaviour is in the whole population from a snowball sample.
    • Ethical concerns: participants may feel pressured to recruit others.
Strengths and Limitations of Non Probability Sampling
Strengths Limitations
Fast and practical no need for complete lists. Unknown selection chance you cannot calculate the probability that any person was selected.
Useful for pilot studies and pretesting tools. Higher risk of bias certain groups are systematically overrepresented or excluded.
Good for qualitative depth rich, detailed information from key informants. Weak generalisation findings cannot be confidently applied to the wider population.
Can reach special groups that probability sampling cannot (hidden populations). Requires transparent reporting you must openly state the limitations in any report or publication.
Probability or Non Probability? Decision Guide
Use Probability When... Use Non Probability When...
Estimating prevalence or incidence in a population. Exploring experiences, beliefs, or perceptions (qualitative research).
Comparing population groups (e.g., urban vs. rural). Finding key informants with specialised knowledge.
Informing district or national planning. Pretesting questionnaires or data collection tools.
Generalisation to the wider population is important. Time, money, or lists are severely limited.
Group Task: Design a Sampling Plan

Study question: Why are some children missing immunisation?

Population: Mothers of children under one year in a catchment area.

Task: Choose one sampling method. State the sampling frame and one likely source of bias. Prepare a two minute explanation.

Example response:

  • Method: Stratified random sampling.
  • Strata: Urban and rural mothers (because access barriers differ).
  • Frame: Village registers for rural areas; facility ANC registers for urban areas.
  • Sample size: 100 mothers total 40 urban, 60 rural (proportional to population).
  • Likely bias: Village registers may miss mothers who recently moved in or who live in informal settlements not on any register. Urban ANC registers may miss mothers who never attended ANC.
  • Mitigation: Use community health workers to identify unregistered mothers. Plan for non response by selecting replacement households.
Session 4: Surveys and Disease Occurrence

Core question: After sampling, how do we count and interpret disease occurrence?

What Is a Survey?

A survey is a systematic method of collecting standard information from a defined group of people. It is not just a questionnaire it is a planned method for answering a health question.

What surveys can measure:

  • Health status: Prevalence of disease, nutritional status, disability.
  • Behaviour: Handwashing practices, net use, sexual behaviour, dietary habits.
  • Service use: ANC attendance, vaccination coverage, facility delivery rates.
  • Knowledge: Awareness of danger signs, understanding of disease transmission, health literacy.

What makes a good survey:

  • Clear questions every question has a purpose and is understood the same way by all respondents.
  • Appropriate sampling method matches the study question and population.
  • Quality control training interviewers, pretesting tools, supervising data collection, checking for completeness.
  • Practical decisions survey results should lead to action, not just sit in a report.
Basic Survey Steps
  1. DEFINE QUESTION
  2. DEFINE POPULATION
  3. CHOOSE SAMPLE
  4. COLLECT DATA
  5. ANALYSE MEASURES
  6. USE FINDINGS

⚠️ Critical Rule: Each step should be planned before fieldwork begins. The analysis plan should match the study question. Quality control starts before data collection not after you realise your questionnaire is confusing.

Counting Disease Occurrence: The Three Components

Every disease measure has three essential parts. Without all three, the number is meaningless:

Component What It Means
Numerator The number of cases or events counted. Example: 15 new malaria cases.
Denominator The population or group from which the cases came. Example: 120 hostel students.
Time Period The period during which new cases or events occurred. Example: During the month of July 2026.
📝 Exam Tip: A number becomes meaningful only when we know the denominator and time period. "15 malaria cases" tells you almost nothing. "15 new malaria cases among 120 students in July" tells you the risk is 12.5% actionable information.
Incidence

Definition: Incidence measures the number of new cases that develop in a population during a specific time period. It tells us about risk the probability that a healthy person will develop the disease.

Incidence Formula:
Incidence = New Cases During a Period ÷ Population at Risk
Usually expressed as a percentage or per 1,000 population

Key rules for incidence:

  • The numerator must include new cases only people who did not have the disease at the start of the period.
  • The denominator should include people at risk those who could have developed the disease. People who already have the disease should be excluded (unless studying recurrence).
  • Incidence must have a time period. Without time, it is not incidence it is just a count.

Incidence Example

Scenario: 15 new malaria cases occurred among 120 hostel students during the month of July.

Incidence = 15 ÷ 120 = 0.125 = 12.5%

Interpretation: About 13 out of every 100 students developed malaria during July. This is the risk of getting malaria in that hostel during that month.

Nursing action: A 12.5% monthly incidence is high. The nurse should investigate: Are nets being used? Is there stagnant water near the hostel? Are students seeking treatment promptly? Consider a mass net distribution or environmental clean up.

Prevalence

Definition: Prevalence measures the total number of existing cases (both new and old) in a population at a specific point in time (point prevalence) or over a period (period prevalence). It tells us about burden how widespread the condition is.

Prevalence Formula:
Prevalence = Existing Cases at a Point or Period ÷ Total Population
Includes people who already have the condition + new cases

Key rules for prevalence:

  • The numerator includes all existing cases both new and old. A person who has had diabetes for 10 years is still counted in prevalence.
  • The denominator is the total population not just those at risk. Everyone in the population could potentially be a case.
  • Prevalence is especially useful for chronic conditions (hypertension, diabetes, HIV) and for planning services (how many beds, drugs, or clinics are needed?).

Prevalence Example

Scenario: During a community screening, 18 adults out of 80 screened have high blood pressure.

Prevalence = 18 ÷ 80 = 0.225 = 22.5%

Interpretation: About 23 in every 100 screened adults had high blood pressure readings. This is the burden of hypertension in the screened population.

Nursing action: A 22.5% prevalence suggests hypertension is common in this community. The nurse should: confirm readings with repeat measurements, counsel on lifestyle, refer high readings, plan follow up clinics, and consider community education on diet and exercise.

Incidence vs. Prevalence: Side by Side
Feature Incidence Prevalence
Counts New cases only All existing cases (new + old)
Measures Risk how likely is a healthy person to get the disease? Burden how widespread is the disease right now?
Needs time? Yes must specify the time period Can be a point in time or a period
Denominator Population at risk (those who could get the disease) Total population (everyone in the group)
Best for Outbreaks, acute diseases, studying causes Chronic diseases, planning services, resource allocation
Example "10% of students got malaria in July" "22.5% of adults screened had high BP"
📝 Exam Tip: Incidence = new. Prevalence = existing. Incidence asks "How many got sick?" Prevalence asks "How many are sick?"
Incidence vs. Prevalence: The Relationship

Prevalence depends on both incidence and duration of disease:

Prevalence ≈ Incidence × Average Duration of Disease

This means:

  • If incidence is high and duration is long → prevalence is very high (e.g., HIV in high burden areas before ART scale up many new infections, and people lived with the disease for years).
  • If incidence is high but duration is short → prevalence may be lower than expected (e.g., acute diarrhoea many new cases, but they recover within 3-5 days, so at any single point, few people are sick).
  • If incidence drops but duration stays long → prevalence may remain high for years (e.g., diabetes fewer new cases due to prevention, but existing cases live for decades with the condition).
💡 Nursing Implication: A high prevalence of hypertension does not necessarily mean many new cases are appearing. It may mean people are living longer with the disease (good chronic care) or that detection has improved. Always ask: "Is prevalence high because of new cases, long duration, or better detection?"
Common Calculation Mistakes
Mistake Why It Is Wrong How to Fix It
Using total cases when the measure requires new cases only This gives prevalence, not incidence Check: Are these new cases or all cases?
Forgetting the time period for incidence Without time, it is not a rate it is just a count Always state: "per month," "per year," "during the outbreak"
Using the wrong denominator Comparing apples to oranges Ensure denominator matches the population at risk
Reporting a percentage without explaining what it means Numbers without context are useless Always interpret: "X out of every 100..."
Comparing groups without considering group size 10 cases in 50 vs. 10 cases in 500 are very different Always calculate rates, not just counts
📝 Exam Tip: Always ask yourself: "Numerator of what? Denominator among whom? During what time?" If you cannot answer all three, your measure is incomplete. Examiners love to give you a number and ask "What is missing?" the answer is usually the denominator or the time period.
Session 5: Ratios, Proportions and Practice

Core question: How do we calculate and interpret basic disease measures accurately?

Ratio

A ratio compares two quantities where the numerator is not necessarily part of the denominator. The two quantities are independent.

Ratio = One Quantity ÷ Another Quantity

Key feature: The numerator and denominator are separate groups. One is not a subset of the other.

Example: 30 male patients and 60 female patients attended the clinic.
Male to female ratio = 30 : 60 = 1 : 2
Interpretation: There is 1 male patient for every 2 female patients.

Other nursing examples:

  • Nurse to patient ratio: 5 nurses for 50 patients = 1 : 10
  • Doctor to nurse ratio: 2 doctors for 10 nurses = 1 : 5
  • Bed to population ratio: 100 beds for 50,000 people = 1 : 500
⚠️ Important: A ratio does NOT tell you what fraction of the whole has a condition. It only compares two groups. "1:2 male to female ratio" does not mean 33% are male it means for every male, there are 2 females.
Proportion

A proportion compares a part to the whole, where the numerator is included in the denominator. It is always expressed as a decimal or percentage.

Proportion = Part ÷ Whole

Key feature: The numerator is a subset of the denominator. The result ranges from 0 to 1 (or 0% to 100%).

Example: 20 diarrhoea cases among 200 children screened.
Proportion = 20 ÷ 200 = 0.10 = 10%
Interpretation: 10% of screened children had diarrhoea.

Other nursing examples:

  • Proportion of ANC attendees who are HIV positive: 15 HIV+ women ÷ 200 ANC attendees = 7.5%
  • Proportion of deliveries by caesarean section: 30 C-sections ÷ 300 deliveries = 10%
  • Proportion of children fully immunised: 85 fully immunised ÷ 100 children = 85%
📝 Exam Tip Ratio vs. Proportion: This is a classic exam trap. Ratio = compares two separate groups (male:female). Proportion = part of a whole (males ÷ total patients). If the numerator is included in the denominator, it is a proportion. If not, it is a ratio.
Rate

A rate describes how fast events occur in a population over time. It is the most informative measure in epidemiology because it combines count, population, and time.

Rate = Occurrence ÷ Population at Risk over Time

Key features:

  • Rates must state the time period.
  • Rates allow comparison between groups of different sizes.
  • Incidence is the most common type of rate.
  • Rates are often expressed "per 1,000" or "per 100,000" for rare diseases.

Example: 40 new malaria cases in a village of 500 children during August.
Rate = 40 ÷ 500 = 0.08 = 8% per month (or 80 per 1,000 per month).

Why "per 1,000" is useful: For rare diseases, percentages are tiny and hard to interpret. Saying "0.002% got Ebola" is confusing. Saying "2 cases per 100,000 population" is clear and standard for international comparison.

💡 Mnemonic Rate vs. Ratio vs. Proportion: "Rate has Time, Ratio has Two groups, Proportion has Part of whole." = R-T, R-T, P-P
Worked Example: Village Diarrhoea

Scenario: A village has 500 children under five. During August, 40 new diarrhoea cases are recorded. At the end of August, 25 children still have diarrhoea.

Question: Calculate August incidence and end of month prevalence.

Answer:

Measure Formula Calculation Result Interpretation
Incidence New cases ÷ Population at risk 40 ÷ 500 8% 8 out of every 100 children developed diarrhoea in August
Prevalence Existing cases ÷ Total population 25 ÷ 500 5% 5 out of every 100 children had diarrhoea at the end of August

Why the difference?

  • Incidence (8%) counts all new cases that occurred during August including those who already recovered by month end.
  • Prevalence (5%) counts only those still sick at the end of the month.
  • The gap (8% − 5% = 3%) represents children who got diarrhoea but recovered before month end.
⚡ Key Principle: Incidence tells you how fast the disease is spreading. Prevalence tells you how much disease is in the community right now. For acute diseases (diarrhoea, malaria attack), incidence is usually higher than point prevalence because people recover quickly. For chronic diseases (diabetes, hypertension), prevalence is much higher than incidence because cases accumulate over years.
Practical Exercise: Calculate and Interpret
Scenario Measure Calculation Interpretation
Village A: 30 new malaria cases among 300 people in July Incidence (risk) 30 ÷ 300 = 10% High risk — 1 in 10 people got malaria that month. Needs urgent vector control.
Village B: 30 new malaria cases among 1,500 people in July Incidence (risk) 30 ÷ 1,500 = 2% Lower risk — 1 in 50 people got malaria. Still monitor, but less urgent.
Health centre: 18 high BP readings among 80 adults screened Prevalence 18 ÷ 80 = 22.5% About 1 in 4 screened adults has high BP. Plan NCD follow up clinic.
Clinic register: 12 males and 36 females attended ANC education Ratio 12:36 = 1:3 For every male companion, 3 female companions attended. Male involvement is low.
📝 Exam Tip: Same number of cases can mean very different risk when denominators differ. Village A and Village B both had 30 cases but Village A's risk was 5 times higher. This is why denominators are essential. In an exam, never just compare counts. Always calculate rates.
Interpreting the Numbers: A 5 Step Framework

When you calculate a disease measure, follow these five steps to interpret it meaningfully for public health action:

Step What to Do Example
1. Name the measure Is it incidence, prevalence, ratio, or proportion? "This is an incidence measure..."
2. State the group Among whom was it calculated? "...among hostel students..."
3. State the time When or over what period? "...during the month of July..."
4. Translate to plain language "X out of every 100..." "...about 13 out of every 100 students..."
5. Suggest one action What should be done? "...suggests the need for improved net use and environmental clean up."

Full example interpretation:
"The incidence of malaria among hostel students was 12.5% during July. This means about 13 out of every 100 students developed malaria that month. This high risk suggests the need for improved insecticide treated net use, removal of stagnant water near the hostel, and prompt testing and treatment of febrile students."

📝 Exam Tip: In exams, marks are awarded for calculation AND interpretation. Many students calculate correctly but lose marks because they do not explain what the number means in plain language. Always finish with: "This means..." and "Therefore, we should..."
Group Assignment Brief

Task: Choose one health problem: malaria, diarrhoea, missed immunisation, or hypertension.

  • Define the population and sampling frame.
  • Choose a sampling method and justify it.
  • Create a small dataset and calculate one disease measure (incidence, prevalence, ratio, or proportion).
  • Present findings in three minutes using the 5 step interpretation framework.

Assessment focus: Clear sampling plan, correct calculation, and practical interpretation.

Example response structure:
"We studied [population] using [sampling method] because [justification]. Our sampling frame was [frame]. We found a [measure] of [X%], which means [interpretation]. Therefore, we recommend [action]. One limitation is [bias/limitation]."

Quick Self Check
Question Answer
Why is sampling used in epidemiology? Because populations are often too large, expensive, or time consuming to study completely. A good sample gives valid information about the wider group without the cost of a census.
What is the difference between stratified and quota sampling? Stratified sampling uses random selection within each stratum (probability method valid for generalisation). Quota sampling sets numbers for categories but uses convenience selection within quotas (non probability method faster but biased).
When would cluster sampling be practical? When the population is large and dispersed, no complete list of individuals exists, and travel costs must be minimised (e.g., national immunisation coverage surveys, DHS, MICS).
How is incidence different from prevalence? Incidence counts new cases over a time period (measures risk "how many got sick?"). Prevalence counts all existing cases at a point or period (measures burden "how many are sick?").
Why must every rate have a denominator and time period? Without a denominator, you cannot compare groups of different sizes. Without a time period, you cannot distinguish rapid outbreaks from slow trends. A rate without both is just a number not actionable evidence.
What is the formula for systematic sampling? k = N ÷ n (population size ÷ sample size = sampling interval). Choose random start between 1 and k, then select every kth unit.
What is sampling bias, and how is it different from sampling error? Sampling bias is systematic error caused by poor selection methods it does not decrease with larger sample size. Sampling error is random natural variation between sample and population it decreases with larger sample size.
Why is a sampling frame important? A sample cannot be better than the frame used to select it. If the frame is incomplete, outdated, or excludes certain groups, the sample will be biased no matter how random the selection method is.
When should you use non probability sampling? For exploratory research, qualitative depth, pilot studies, pretesting tools, or when studying hard to reach/hidden populations where probability sampling is impossible.
How do you interpret a ratio of 1:3 male to female? For every 1 male, there are 3 females. This does NOT mean 25% are male (that would be a proportion). It only compares the two groups.
References
  • Gordis, L. (2014). Epidemiology. Elsevier Saunders.
  • Bonita, R., Beaglehole, R., & Kjellström, T. (2006). Basic Epidemiology. World Health Organization.
  • Webb, P., Bain, C., & Page, A. (2017). Essential Epidemiology: An Introduction for Students and Health Professionals. Cambridge University Press.

Quick Quiz

Sampling Methods Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Leave a Comment

Your email address will not be published. Required fields are marked *

Want notes in PDF? Join our classes!!

Send us a message on WhatsApp
0726113908

Scroll to Top
Enable Notifications OK No thanks