Nurses Revision

nursesrevision@gmail.com

Healthcare Team and Their Roles & Responsibilities

Healthcare Team and Their Roles & Responsibilities
Medical Staff: Physician
Assessment:
  • Performing complete health assessments including: Taking a full medical history including presenting complaint, past illnesses, social history, family history, and performing a complete physical examination.
  • Screening patients at risk for hereditary conditions and potentially preventable disorders.
  • Assessment, diagnosis, primary medical treatment and advice for management of acute medical conditions and injuries.
  • Assessment of the exacerbations and complications of chronic medical problems.
Treatment/Management:

Provision of continuous care to patients over their lifetime based on the delivery of the following services:

  • Acute medical treatment for a range of medical problems from minor ambulatory care visits to severe life threatening illness presenting to emergency rooms, in hospitals, in the home and in long term care facilities.
  • Provide primary reproductive care including maternal and newborn care.
  • Provide screening for and treatment of sexually transmitted diseases (STDs).
  • Provide primary mental health care.
  • Provide palliative care.
  • Provide hospital care where required.
  • Provide early intervention and counseling to reduce risk or development of harm from disease.
  • Provide appropriate immunizations.
  • Provide care and monitoring of chronic illnesses, including patients with complex co-morbidities.
  • Provide early access for assessment of episodic illness or injury with provision of diagnosis, primary medical treatment and advice on self-care and prevention.
  • Maintain and keep safe the medical record of each patient.
  • Perform surgeries where required.
Education/Advocacy:
  • Provide counseling on many health and health care issues including but not limited to birth control, prevention of STDs, prevention of disease and issues related to the effects of disease on family members.
  • Perform the role of advocate to assist patients to navigate through a complex health care system in order to obtain the best care in the most expeditious way in a cost effective manner.
  • Identify and meet the needs of the individual patients, the practice population and the community in general by working with a variety of partners throughout the public health, community, and hospital sectors.
Referrals/Collaboration:
  • Assist with discharge planning, rehabilitation services, out patient follow-up and home care services.
  • Coordinate referrals to other health care providers and agencies, including specialists, rehabilitation and physiotherapy services, home care and palliative care services, and diagnostic services, as required.
  • Collaborate with other mental health care providers when required.
  • Coordinate referrals to secondary and tertiary facilities based on patients’ needs.
  • Report births, deaths, and contagious and other diseases to governmental authorities.
  • Collaborate with necessary public health initiatives.
Registered Nurses (PNO, SPNO) & Other Nursing Staff

Depending on the population health needs and the mix of other providers, the Family Health Team may choose to integrate an RN, RPN, or both into the interdisciplinary team.

Assessment:
  • Assess holistically and provide services to patients in all developmental stages, and to families and communities.
  • Complete health assessments, including a health history and physical examination.
  • Formulate and communicate medical diagnoses.
  • Synthesize information from patients to identify broader implications for health within the family.
  • Use family assessment tools to evaluate family strengths and needs.
  • Determine the need for, and order from, an approved list of screening and diagnostic laboratory tests and interpret the results.
  • Determine the need for, and order and interpret reports of X-rays, ECGs and diagnostic ultrasounds for diagnosis.
  • Assess patient preferences.
  • Assessment of patient health care needs (physical, emotional, psychological, and spiritual).
  • Analysis of the findings of a health assessment.
  • Interpret patient health records.
  • Observe and record outcomes.
  • Collect data through a therapeutic relationship with a patient.
Treatment/Management:
  • Initiate and manage care of patients with diseases or disorders.
  • Monitor the ongoing therapy of patients with chronic stable illness by providing effective pharmacological, complementary or counseling interventions.
  • Prescribe drugs from an approved list.
  • Use nursing strategies arising from the best available evidence and consistently incorporate patient’s perspectives in care.
  • Determine the appropriate service or treatment, the appropriate care provider or the appropriate equipment.
  • Provide nursing care and treatment (including complementary therapies and/or counseling) for health problems.
Education/Advocacy:
  • Determine the need for, and implementation of, health promotion, and primary and secondary prevention strategies for individuals, families, and communities, or for specific age and cultural groups.
  • Provide health education to individuals and groups.
  • Identify community needs and resources and develop age and culturally sensitive community programs.
  • Help patients to identify and use health resources.
  • Involve patients in decisions about their own health.
  • Encourage patients to take action for their own health.
  • Initiate health education and other activities that assist, promote and support patients as they strive to achieve the highest possible level of health.
  • Develop learning resources for nurses and other health care providers.
  • Develop and deliver health education programs for patients, or communities.
Referrals/Collaboration:
  • Consult with a physician in accordance with the standards for consultation with physicians, and/or refer the patient to another healthcare professional.
  • Collaborate with other healthcare providers.
  • Coordinate patient care.
  • Refer to community programs and mental health services.
Midwives

(Note: Midwifery is often a separate specialization, but foundational nursing includes aspects of maternal and child health.)

Assessment:
  • Assess and monitor women during pregnancy.
  • Provide pre-natal education.
  • Order tests if necessary.
Treatment/Management:
  • Deliver babies.
  • Administer some medications during delivery if necessary.
  • Manage labour and conduct spontaneous normal vaginal deliveries.
  • Perform episiotomies and amniotomies and repairing episiotomies and lacerations, not involving the anus, anal sphincter, rectum, urethra and periurethral area.
  • Administer, by injection or inhalation, a substance designated in the regulations (Midwifery Act, 1991, c. 31, s. 4).
  • Take blood samples from newborns by skin pricking or from women from veins or by skin pricking.
  • Insert urinary catheters into women.
  • Prescribe drugs designated in the regulations (Midwifery Act, 1991, c. 31, s. 4).
  • Monitor women in post partum period.
  • Assess/monitor new babies.
Education/Advocacy:
  • Assist women in making informed decisions about their care and choice of birthplace.
Referrals/Collaboration:
  • Arrange consultation or transfer to physician if necessary.
  • Assist in complicated deliveries.
  • Report births to governmental authorities.
Dietitian

The Registered Dietitian (R.D.) is a healthcare professional trained in the single specialty of nutrition science. Their goal is to promote health and fight illness by fostering the practice of proper nutrition to individuals and groups.

Assessment:
  • Work with individual patients to determine nutritional needs.
  • Conduct nutritional and weight assessments.
Treatment/Management:
  • Develop nutritional plans based on comprehensive needs assessments.
  • Provide nutritional counseling.
  • Provide weight management counseling.
Education/Advocacy:
  • Promote behaviour change related to food choices, eating behaviour and preparation methods to optimize health.
  • Promote patient independence and autonomy in decision making for patient to achieve health.
  • Conduct patient workshops and seminars.
  • Identify community capacities and facilitate community skill building, health advocacy, and social action.
Referrals/Collaboration:
  • Work with physicians on medication monitoring plans as they relate to nutrition.
  • Communicate relevant nutritional information to other health care providers.
Pharmacists

Pharmacists dispense drugs and medications prescribed by physicians, physician assistants, nurse practitioners, and dentists. They also advise healthcare professionals and patients on the use and proper dosage of medications, as well as expected side effects and interactions with other prescription and nonprescription medicines. These professionals also order and maintain supplies of medications and various medical supplies required for use in the clinical setting.

Assessment:
  • Ensure appropriate patient information is gathered and recorded.
  • Review patient profile including known patient risk factors for adverse drug reactions, drug allergies, known contraindications to prescription drugs, nonprescription drugs, natural health products, and complementary or alternative medicines.
  • Evaluate patient drug therapy and identify potential and actual drug-related problems and determine appropriate therapeutic options to resolve or prevent them.
  • Conduct patient assessments for medication problems.
Treatment/Management:
  • Manage medication.
  • Monitor patient compliance.
  • Home follow-up.
Education/Advocacy:
  • Patient education to facilitate patient’s understanding of her/his drug therapy and ability to comply with the therapy regimen.
Referrals/Collaboration:
  • Refer the patient to appropriate health care providers within the Family Health Team if necessary.
  • Communicate with physicians to help the patient achieve maximum benefit from drug therapy and to prevent medication errors or potential significant adverse reactions.
Orthopedists
Assessment:
  • Complete health assessment through information gathering, lower extremity physical examination, patient health history and relevant clinical findings.
  • Evaluation of overall lower extremity foot and ankle function relating to activities of daily living.
  • Examination and review of lab tests, diagnostic tests and consulting medical and surgical notes.
  • Assessment of the impact of an injury, disability or disease (rheumatoid arthritis/diabetes/sprains/strains) on foot function.
Treatment/Management:
  • Perform surgery by cutting into subcutaneous tissues of the foot.
  • Administer, by injection into feet, a substance designated in the regulations.
  • Prescribe drugs designated in the regulations.
  • Perform surgery by cutting into bony tissues of the forefoot if the required training has been completed.
  • Communicate a diagnosis identifying a disease or disorder of the foot as the cause of a person’s symptoms.
  • Take x-rays under the Healing Arts Radiation Protection Act.
Education/Advocacy:
  • Educate and advise patients about the prevention and care of morbid conditions relating to chronic diseases (e.g., diabetes and peripheral vascular disease).
Referrals/Collaboration:
  • Chiropodists and podiatrists work as key interdisciplinary practitioners in hospitals, community health care centres, and nursing and retirement homes. In private practice, they receive referrals from medical and other health care practitioners and consult with these referring practitioners to provide timely and optimal care for their patients.
Social Worker

The role of social workers in an interdisciplinary team is to provide the psychosocial perspective to complement the biomedical perspective.

Assessment:
  • Assessment and social work diagnosis of psychosocial problems.
Treatment/Management:

They provide counseling, and enable individuals, families, and communities to obtain social services. They work with clients on issues of unemployment, illness, disability, housing, abuse, and financial problems. Social workers specializing in providing mental health services and counseling are called Clinical Social Workers. In the community, they may be active in organizing communities to improve health and social services. Social workers often assist families in crisis situations and during periods of transitions.

  • Individual, couple, family and group counseling and psychotherapy.
  • Case Management, including linkages to community resources.
Education/Advocacy:
  • Health Promotion.
  • Psycho-education related to the prevention of mental health problems.
  • Assistance in navigating service delivery networks to find required resources.
  • Advocacy to establish and access needed resources.
Referrals/Collaboration:
  • Development, management and delivery of programs alone or in collaboration with other professionals.
  • Consultation with other professionals related to patient needs.
Psychologists
Assessment:
  • Evaluation, diagnosis, and assessment of the functioning of individuals and groups related to mental disorders as well as wellness and mental health.
Treatment/Management:
  • Interventions with individuals and groups and organizations.
  • Treatment of serious mental health disorders.
  • Treatment of individual, marital and family relationships problems.
  • Maintenance of wellness and disease prevention.
  • Management of psychological factors and problems associated with physical conditions and disease (e.g., diabetes, heart disease, stroke).
  • Management of psychological factors in terminal and chronic illnesses such as cancer, brain injury, and degenerative brain diseases.
  • Treatment of addictions and substance use and abuse.
  • Pain management.
  • Assist with stress, anger and other aspects of lifestyle management.
  • Management of the impact and role of psychological and cognitive factors in accidents and injury, capacity, and competence one’s to manage personal affairs.
  • Treatment of problems associated with cognitive functioning such as learning, memory, problem solving, intellectual ability and performance.
  • Management of psychological factors related to work such as motivation, leadership, productivity, and healthy workplaces.
  • Administration of psychological services.
Education/Advocacy:
  • Public education regarding wellness and the promotion of mental health.
  • Implementation of primary and secondary prevention strategies.
  • Program development and evaluation.
Referrals/Collaboration:
  • Consultation relating to the assessment of or interventions with individuals and groups to facilitate the prevention or treatment of difficulties.
  • Referral to community agencies/services.
Counselors
Assessment:
  • Intake and assessment.
  • Develop treatment plan.
Treatment/Management:
  • Counsel individuals, couples and families.
  • Facilitate/run counseling groups (e.g., relapse prevention, guided self-change, anger management, stress management).
  • Assess and adjust and adjust of treatment plans on an ongoing basis.
  • Develop discharge plan.
Education/Advocacy:
  • Provide information about community resources.
  • Assist patients accessing other services.
Referrals/Collaboration:
  • Advise physicians and other health care workers regarding indicators of substance abuse, relapse prevention and appropriate referral techniques.
  • Collaborate with physicians, psychologists, and other professionals regarding after care plan and follow-up activities.
  • Refer to community programs and mental health services.
  • Refer to psychologists, psychiatrists, and other professionals as appropriate.
Health Educators

Health educators teach clients, both individually and in groups, about various health topics. Although all members of the healthcare team are charged with client education, health educators are focused on providing adequate information to the client to assure understanding of the medical problem and treatment plan. These individuals may focus their educational efforts in health promotion and disease prevention activities that reduce the burden of disease in the community. Some health educators are utilized to provide in-depth instruction to clients about specific illnesses after being diagnosed.

Community Health Worker

Community Health Workers (CHW) can be broadly defined as individuals who connect healthcare consumers and providers, promoting health particularly among groups who have traditionally lacked access to care. The CHW is a member of the community and play an important role in identifying a community’s problems and in developing solutions. Examples of successful uses of the CHW include: using ex-addicts to educate intravenous drug users about AIDS risks and increasing breast, cervical, and colon cancer screening in minority communities. CHWs may play critical roles in improving community health status by providing cultural and technical linkages between community members, primary care providers, and the health care delivery system.

Assessment:
  • Intake Assessment.
Treatment/Management:
  • Facilitate coordinated access to services in areas such as assistance with daily living, housing, crisis intervention, treatment, health promotion and prevention.
  • Facilitate linkages with appropriate services, supports, and resources.
  • Provide crisis intervention and intensive/short-term support.
  • Evaluate achievement of patient goals.
  • Financial management: budgeting, banking.
  • Nutrition: menu planning, grocery shopping, food preparation.
  • Personal effectiveness: problem-solving, decision making, communication and interpersonal skills, goal-setting, time structuring and management.
  • Community integration: use of transit, social/recreational, peer support and other services.
  • Health and wellness: support clinical plan including medication, appointments, healthy choices and lifestyle.
  • Employment/service: support maximum involvement in volunteer, community service or paid employment.
  • Personal care: hygiene grooming, self-care skills, clothing maintenance.
  • Household management: such as laundry and house cleaning.
  • Housing support: finding and maintaining adequate housing, liaison/support to landlord, utilities.
Education/Advocacy:
  • Advocacy: support appropriate use of available community public services and programs.
  • Advocate for patient’s civil and legal rights.
Referrals/Collaboration:
  • Collaborate with other professionals regarding after care plan and follow-up activities.
  • Refer to community programs and mental health services.
Physiotherapists
Assessment:
  • Assess movement, strength, endurance and other physical abilities.
  • Assess the impact of an injury or disability on physical functioning.
  • Assess physical preparation for work and sports.
  • Evaluate pain and movement patterns, muscle balance, joint function, cardio-respiratory status, reflexes and sensation.
  • Examine relevant x-rays, lab tests, medical records and surgical notes.
  • Evaluate overall functional ability both in the workplace and in other activities of daily living.
Treatment/Management:
  • Plan treatment programs, which include education, to restore movement and reduce pain.
  • Provide individualized treatment of an injury or disability based on scientific knowledge, a thorough assessment of the condition, environmental factors and lifestyle.
  • Provide treatment which can include an individualized exercise program, manual therapy, modalities, as well as patient and family education and home exercise prescription.
Education/Advocacy:
  • Educate to restore movement and reduce pain.
  • Encourage patients/patient to take charge of their health by teaching techniques for recovery, pain relief, injury prevention and improved physical movement, with emphasis on what the patient can do for her/himself.
  • Promote independence and facilitate patients assuming responsibility for their rehabilitation and self-care.
Referrals/Collaboration:
  • Based on assessment the physiotherapist either plans an appropriate treatment program and carries it out or refers the patient to another professional.
  • Coordinates treatment with other providers.
Occupational Therapists
Assessment:
  • Assessment of physical, emotional, and cognitive functioning with environmental considerations.
  • Evaluation of the home, work or school environment to assess the need for specialized equipment modifications and/or supports.
Treatment/Management:
  • Individualized treatment plans to develop, maintain, or augment function using evidence based treatment modalities.
  • Teaching daily living and community life skills.
  • Prescribing specialized adaptive equipment and teaching proper usage.
  • Modification of the physical and social home, work or school environments.
Education/Advocacy:
  • Educating and counseling family members and caregivers regarding the impact of disability, injury or disease on the individual and their potential role within the recovery process.
  • Educating and counseling to promote function and independence including health promotion and injury prevention.
Referrals/Collaboration:
  • Based on assessment the occupational therapist refers the individual to additional health care and community services as needed.
  • Collaborates with other health care professionals and community service providers to promote comprehensive and coordinated care.
Neurologists
Assessment:
  • Diagnosis, including differential diagnosis, of musculoskeletal disorders or referral for non-musculoskeletal complaints.
  • Ongoing evaluation of treatment/management outcomes using standard measurement tools.
  • Request/utilize X-Rays as authorized by the Healing Arts Radiation Protection Act.
  • Assessment or evaluation of workplace or home environments to inform treatment decisions and to provide ergonomic, activity, or other advice.
Treatment/Management:
  • Treatment of acute conditions and management of chronic or recurrent complaints with a focus on self-care.
  • Manual care including joint manipulation and mobilization and a wide variety of soft tissue techniques.
  • Electrotherapies such as ultrasound, electrical stimulation, laser, etc.
  • Planning, instruction, and supervision of therapeutic exercise programs.
Education/Advocacy:
  • Education for self-management of musculoskeletal conditions, including injury prevention, lifestyle and ergonomic advice.
  • Encouraging fundamental health promotion activities are integral to chiropractic practice.
Referrals/Collaboration:
  • Refer to physicians, physiotherapists, occupational therapists, psychologists, and others where appropriate.
  • Share care where the expertise of others is appropriate.
  • Communicate with other health professionals to facilitate patient care.
Volunteers

Volunteers are individuals that that provide services in the clinical setting with no monetary payment. They may be retired healthcare practitioners or citizens with a strong desire to provide public service to the community. Many clinics utilize these volunteers to perform a variety of jobs, such as interpreters, filing, answering telephones, or more patient oriented services, such as taking vital signs, assisting patients in completing forms, and assisting other health care team members.

Other Team Members

Other members of the interdisciplinary health care team may include: surgeons, ophthalmologists, ENT specialists, radiotherapists, laboratory technicians, speech and language therapists, and art or music therapists. The availability of these additional members of the health care team depends on the community served and the health care services offered.

References (from Curriculum for CN-1111)

Below are the core and other references listed in the curriculum for Module CN-1111. Refer to the original document for full details.

  • Uganda Catholic Medical Bureau (2015) Nursing and Midwifery procedure manual 2nd Edition Print Innovations and Publishers Ltd. Uganda
  • Nettina .S,M (2014) Lippincott Manual of Nursing Practice 10th Edition, Wolters Kluwer, Philadelphia, Newyork
  • Gupta, L.C., Sahu,U.C. and Gupta P.(2007):Practical Nursing Procedures. 3rd edition. JAYPEE brothers, New Delhi.
  • Craveni, R. Hirnle, C. and Henshaw, M.C. (2017). Fundamentals of Nursing Human Health and Function. 8th Edition. Wolters Kluwer
  • Hill, R., Hall, H and Glew, P. (2017). Fundamentals of Nursing and Midwifery, A person-Centered Approach to care. Wolters Kluwer
  • Rosdah I, BC and Kowalkski, TM (2017) Text book for Basic Nursing 11th Edition Wolters Kluwer.
  • Samson .R. (2009) Leadership and Management in Nursing Practice and Education 1st Edition Jaypee Brothers Medical Publishers India.
  • Taylor.C.R (2015) Fundamentals of Nursing, The Art and Science of person – centred nursing care, 8th Edition Wolters Kluwer, Health/Lippincott Williams and Wilkins.
  • Timby, K.B (2017) Fundamental Nursing Skills and concept 11th Edition Wolters Kluwers, Lippincotts Williams and Wilkins.
  • Lynn, P. (2015) Tyler's Clinical nursing skills, A Nursing Process Approach 4th Edition Wolters Kluwers, China
  • Gupta, D.S. (2005) Nursing Interventions for the critically ill 1st Edition Jaypee Brothers Medical Publishers Ltd. India.
  • Uganda Catholic Medical Buraeu (2010) Nursing and Midwifery Procedure Manual. 1st Ed. Print Innovations and Publishers Ltd., Uganda.
  • Carter, J. P. (2012) Lippincott's Textbook for nursing Assistant. 3rd Edition. Walters Kluwers. Lippingcotts Williams and Wilkins
  • Jensen, S. (2015) Nursing Health Assessment; A host Practice Approach. 2nd Edition. Wlaters Kluwer,
  • Gupta, D.S. (2005) Nursing Interventions for the Critically Ill. 1st Edition. Jaypee Brothers Medical Publishers Ltd. India.
  • UCMB. (2015) Nursing and Midwifery Procedure Manual. 2nd Edition. Print Innovation and Publishers Ltd. Kampala. Uganda.
  • Karesh, P. (2012) First Aid for Nurses. 1st Edition. Jaypee Brothers Publishers Ltd. India.
  • Molley, S. (2007) Nursing Process; A Clinical Guide. 2nd Edition. Jaypee Brothers Medical Publishers Ltd. India.
  • Carter, J.P. (2016) Lippincott's Textbook for Nursing Assistants. 4th Edition. Wolters Kluwer, Lippincotts Williams and Wilkins.
  • Rahim,A. (2017). Principles and practices of community medicine. 2nd Edition. JAYPEE Brothers Medical Publishers Ltd. New Delhi
  • Cherie Rector, (2017),Community & Public Health Nursing: Promoting The Public's Health 9e Lippincott Williams and Wilkins
  • Gail A. Harkness, Rosanna Demarco (2016) Community and Public Health Nursing 2nd edition, Lippincott Williams and Wilkins
  • Basavanthapp, B.T and Vasundhra, M.K (2008), Community Health Nursing, 2nd edition. JAYPEE Brothers Medical Publishers Ltd. New Delhi
  • Kamalam, S. (2017), Essentails in Community Health Nursing Practice 3rd edition. JAYPEE Brothers Publishers Ltd. New Delhi
  • James F. McKenzie, PhD, MPH, MCHES, MEd,and Robert R. Pinger, PhD, (2018) An Introduction to Community & Public Health, 9th edition, Jones and Bartlett Publishers. Sandburg, Massachusetts.
  • Maurer, F.A, Smith, C.M (2005), Community /Public health Nursing Practice, 3rd edition ELSEVIER SAUNDERS, USA
  • МОН, (2013) Occupational Safety and Health Training Manual, 1st Edition
  • МОН, (2008), Policy for Mainstreaming Occupational Health & Safety In The Health Service Sector.
  • Wooding, N. Teddy, N. Florence, N. (2012) Primary Health Care in East Africa. 1st Edition. Fountain Publishers. Kampala. Uganda.

Quick Quiz

Foundations Nursing Day 1 Quiz

FON - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Healthcare Team and Their Roles & Responsibilities Read More »

Correlation, Regression, Integration

Correlation, Regression, Integration

Correlation, Regression & Integration
Learning Objectives

By the end, you should be able to:

  • Interpret scatter diagrams and correlation coefficients in plain nursing language.
  • Explain simple linear regression using a real nursing or public health example.
  • Interpret a basic regression table (coefficient, confidence interval, p-value) confidently.
  • Integrate key ideas from epidemiology and biostatistics into a coherent decision framework.
  • Prepare confidently for continuous assessment and the final examination.
🎯 Guiding Question for Today:

"How do we use data to describe relationships, make simple predictions, and support public health decisions?"

Session 1: Scatter Diagrams and Correlation
Why Study Relationships in Nursing?

As a nurse or public health officer, you rarely look at just one number in isolation. You need to know whether two things move together. If they do, you can plan better, predict needs, and target interventions.

Health Situation Variable X Variable Y
Nursing workload Number of patients seen per day Average waiting time in OPD
Maternal care Number of ANC visits attended Likelihood of facility delivery
Child nutrition Child age in months Weight-for-age z-score
Malaria seasonality Monthly rainfall (mm) Number of malaria cases
Immunization coverage Distance from village to clinic (km) Measles vaccination rate (%)
💡 Key Insight:

Correlation helps you answer: "Do these two variables move together?" It does NOT tell you "Does X cause Y?" That requires a different level of evidence.

What is a Scatter Diagram?

A scatter diagram (or scatter plot) is a picture of paired values for two numeric variables. Each dot on the graph represents one observation — one patient, one day, one village.

How to set up a scatter diagram:

  • X-axis (horizontal): The explanatory variable — the one you think might explain or predict the outcome. Also called the independent variable or predictor.
  • Y-axis (vertical): The outcome variable — the one you want to explain or predict. Also called the dependent variable or response variable.
  • Each dot: Represents one pair of measurements (one X value and one Y value).
📝 Exam Tip — Setting Up Axes:

A simple rule: "X explains, Y is the outcome." If you are predicting waiting time FROM patient load, then patient load goes on X and waiting time goes on Y. Mnemonic: "Xplains Your outcome" — X explains Y.

Direction of Relationship

The first question when looking at a scatter plot is: Which way does the cloud of points move?

Direction What It Looks Like Nursing Example
Positive (↑↑) As X increases, Y also increases. Dots rise from left to right. More patients → longer waiting time. More ANC visits → higher birth weight.
Negative (↓↑) As X increases, Y decreases. Dots fall from left to right. Greater distance to clinic → lower vaccination coverage. More health education → fewer teenage pregnancies.
None (random) No clear pattern. Dots scattered all over. Shoe size and blood pressure — no logical connection, so no correlation.
Strength of Relationship

The second question is: How tightly do the points follow a pattern?

Strength Visual Pattern What It Means
Strong Points cluster very tightly around an imaginary straight line. X is a very good predictor of Y. You can predict Y quite accurately from X.
Moderate Points follow a general trend but with some spread. X predicts Y reasonably well, but other factors also matter.
Weak Points are widely scattered; the trend is barely visible. X is a poor predictor of Y. Many other factors are more important.
The Correlation Coefficient: r

The Pearson correlation coefficient (r) is a single number that summarises both the direction and strength of a linear relationship between two numeric variables.

The Correlation Coefficient Ranges

r ranges from −1 to +1

−1 = perfect negative | 0 = no linear relationship | +1 = perfect positive

What r tells you:

  • The sign (+ or −) tells the direction of the relationship.
  • The absolute value (how close to 1) tells the strength of the relationship.
  • r does NOT prove cause and effect. It only measures linear association.
Interpreting r in Plain Language

Use this guide to describe any correlation coefficient you encounter:

Absolute Value of r Strength Description Example Sentence
0.00 to 0.19 Very weak / negligible "There was very little linear relationship between the two variables."
0.20 to 0.39 Weak "There was a weak positive linear relationship between rainfall and malaria cases."
0.40 to 0.59 Moderate "There was a moderate negative linear relationship between distance and clinic attendance."
0.60 to 0.79 Strong "There was a strong positive linear relationship between patient load and waiting time."
0.80 to 1.00 Very strong "There was a very strong positive linear relationship between gestational age and birth weight."
📝 Exam Tip — Interpreting r:

Always mention three things in your answer: (1) the direction (positive/negative), (2) the strength (weak/moderate/strong), and (3) the variables involved. Never just say "r = 0.72 is strong." Say: "There is a strong positive linear relationship between daily patient load and average waiting time (r = 0.72)."

⚠️ Correlation is NOT Causation

This is the most important rule in all of correlation analysis. Two variables may move together for many reasons — only one of which is direct cause-and-effect.

Why two variables might be correlated without causation:

  • Direct effect (causation): X genuinely causes Y. More mosquitoes → more malaria. (Rare to prove with correlation alone.)
  • Common cause (confounding): A third variable Z causes both X and Y. Ice cream sales and drowning deaths are correlated because hot weather causes both.
  • Reverse causation: Y might actually cause X, not the other way around. Does low income cause poor health, or does poor health cause low income? Both can be true.
  • Coincidence: The correlation might be purely random, especially in small samples.
  • Measurement bias: The same flawed measurement tool might inflate both X and Y artificially.

Unsafe Wording: "Higher patient load caused longer waiting time."

Safe Wording: "Higher patient load was associated with longer waiting time." or "Higher patient load was positively correlated with longer waiting time."

🧠 Mnemonic — Safe Language:

"Always Say Associated" — ASA. Never say "caused" unless your study design (like an RCT) supports it.

Worked Example 1: Patient Load and Waiting Time

🩺 Question: Is patient load related to waiting time in the OPD?

Data collected over 5 days:

Day Patients Seen (X) Avg Wait Time in min (Y)
12035
23045
34050
45065
56075

Scatter Plot Description:

  • The dots rise from left to right → positive direction.
  • The dots are very close to an imaginary straight line → very strong relationship.
  • There are no obvious outliers.

Calculated correlation: r = 0.99

  • Direction: Positive (as patients increase, waiting time increases).
  • Strength: Very strong (r is very close to +1).
  • Plain meaning: Days with more patients had substantially longer average waiting times.
  • Caution: This does not by itself prove that more patients caused the longer waits — but it strongly suggests a predictable relationship.

Best Report Sentence: "Patient load was very strongly positively associated with average waiting time (r = 0.99)."

Class Practice: Distance from Facility and ANC Attendance

📊 Scenario: A scatter plot shows distance from the nearest health facility (km) on the X-axis and number of ANC visits attended on the Y-axis.

Your task — answer these three questions:

  • Direction: Do the dots rise or fall from left to right? (As distance increases, do ANC visits increase or decrease?)
  • Strength: Are the dots tightly clustered or widely scattered?
  • Safe interpretation: Write one sentence using the word "associated with" instead of "caused."

Expected answer: "Greater distance from the health facility was moderately negatively associated with ANC attendance." (As distance goes up, attendance goes down.)

Common Mistakes in Correlation — Avoid These!
Mistake Why It's Wrong How to Fix It
Calling every association "causal" Correlation alone cannot prove cause. Only experimental designs (like RCTs) can. Use "associated with," "correlated with," or "linked to."
Ignoring outliers One extreme point can pull r up or down dramatically, giving a false impression. Always look at the scatter plot, not just the number. Check for outliers.
Using r for curved relationships Pearson r only measures straight-line relationships. A U-shaped pattern could have r ≈ 0 even though there IS a relationship. Always plot the data first. If the pattern is curved, Pearson r is the wrong tool.
Forgetting to name the variables Saying "r = 0.6 is strong" tells the examiner nothing useful. Always state: "There is a strong positive correlation between [Variable X] and [Variable Y]."
Session 2: Introduction to Simple Linear Regression
What Regression Adds Beyond Correlation

Correlation tells you how strongly two variables are related. Regression tells you how much Y changes when X changes — and gives you an equation to make predictions.

Question Correlation Answers Regression Answers
What it tells you How strongly are X and Y related? How much does Y change for each 1-unit change in X?
Output A single number: r (between −1 and +1) An equation: Ŷ = a + bX
Prediction No — r cannot predict values. Yes — plug in an X value and get a predicted Y.
Cause-effect proof No — correlation ≠ causation. No — regression also does not prove causation by itself.
📝 Exam Tip:

If an exam question asks you to predict a value, you need regression, not correlation. If it asks you to describe the relationship, correlation is sufficient.

The Simple Linear Regression Equation

Simple linear regression uses one explanatory variable (X) to predict one outcome variable (Y). The equation is:

Ŷ = a + bX

  • Ŷ (Y-hat) = the predicted value of the outcome
  • a = the intercept (predicted Y when X = 0)
  • b = the slope (change in Y for every 1-unit change in X)
  • X = the explanatory variable value you plug in

Nursing example: Predict waiting time (Y) from the number of patients seen in OPD (X).

Choosing X and Y: The Golden Rule

Be crystal clear about which variable explains and which is explained:

  • X (explanatory / independent / predictor): The variable you use to explain or predict the outcome. You think it comes first or drives the change.
  • Y (outcome / dependent / response): The variable whose value you want to explain or predict. It is the result you care about.

🧠 Mnemonic: "Xplains Your outcome" — X is the explainer, Y is the outcome. Another one: "X comes before Y in the alphabet, and X comes before Y in time/causation."

Slope (b): The Key Regression Idea

The slope is the most important number in regression. It tells you exactly how much the outcome changes for every one-unit increase in the explanatory variable.

Slope Formula:
b = change in Y ÷ change in X

"Rise over run" — how much Y rises (or falls) for each unit of X.

Interpreting slope in nursing language:

  • If b = 1.0 (waiting time in minutes per patient): Each additional patient is associated with 1 extra minute of average waiting time.
  • If b = −0.40 (ANC visits per km): For each additional kilometre from the facility, ANC attendance decreases by 0.40 visits on average.
  • If b = 2.5 (birth weight in grams per week of gestation): Each additional week of gestation is associated with 2.5 extra grams of birth weight.
📝 Exam Tip — Interpreting Slope:

Always include units in your interpretation. Do not say "b = 1.0" — say "For each additional patient seen, the average waiting time increased by 1.0 minute." The unit is "minutes per patient."

Intercept (a): Read Carefully

The intercept is the predicted value of Y when X = 0. It is where the regression line crosses the Y-axis.

  • Sometimes meaningful: If X = 0 patients, the predicted waiting time is 14 minutes. This might represent administrative setup time before the first patient.
  • Sometimes meaningless: If X = 0 weeks of gestation, the predicted birth weight would be nonsensical. X = 0 is outside the real data range.
  • Rule: Do not over-interpret an intercept when X = 0 is unrealistic or outside your observed data range.
Worked Example 2: Regression Equation for OPD Data

Data (same as Example 1):

X: Patients 20 30 40 50 60
Y: Wait (min) 35 45 50 65 75

Regression output: Ŷ = 14 + 1.0X

Interpreting each part:

  • Intercept (a = 14): When zero patients are seen, the predicted waiting time is 14 minutes. This might represent the time to open the clinic, set up registers, and prepare. Interpret with caution — we have no data at X = 0.
  • Slope (b = 1.0): For each additional patient seen per day, the average waiting time increases by 1 minute. The unit is "minutes per patient."
  • Prediction example: If X = 45 patients, Ŷ = 14 + (1.0 × 45) = 14 + 45 = 59 minutes. We would predict an average waiting time of 59 minutes on a day with 45 patients.
Using Regression for Prediction

Regression predictions are estimates, not guarantees. They come with uncertainty. Here are the rules for safe prediction:

  • Stay within range: Only predict Y for X values within your observed data range. If you observed 20–60 patients, do not predict waiting time for 200 patients — the relationship may not hold.
  • Avoid extrapolation: Predicting far outside your data is called extrapolation and is scientifically dangerous. The line may curve or flatten beyond what you observed.
  • Remember uncertainty: The prediction Ŷ is an average. Actual waiting times will vary around this line. Some days will be higher, some lower.
⚠️ Danger Zone:

If your data shows patient loads from 20 to 60, predicting for 5 patients or 500 patients is extrapolation. The nurse manager might use your equation to plan staffing for 70 patients (slight extrapolation), but not for 300 patients (massive extrapolation).

Basic Checks Before Using Regression

Before you trust a regression equation, run through these four checks:

Check What to Look For What to Do If It Fails
Linearity The scatter plot should look roughly straight, not curved like a U or an S. Do not use simple linear regression. Consider data transformation or non-linear models.
Outliers Are there extreme points far from the others? One outlier can pull the whole line. Investigate the outlier. Is it a data entry error? If real, report results with and without it.
Independence One observation should not depend on another. Each patient/village/day should be separate. If data are repeated measures on the same person, use a different model (repeated measures).
Meaning Does the model make clinical or public health sense? If the slope direction contradicts clinical knowledge, check your data and model again.
Session 3: Practical Interpretation of Regression Results
Reading a Simple Regression Table

In real research and in exams, regression results are presented in a table. You must know how to read three key columns: Coefficient, 95% Confidence Interval (CI), and p-value.

Example Regression Table:

Predictor Coefficient 95% CI p-value
Patients seen per day 1.0 0.75 to 1.25 0.002
Intercept 14.0 5.0 to 23.0 0.010

Plain Interpretation: "Each additional patient seen per day was associated with about one extra minute of average waiting time (95% CI: 0.75–1.25; p = 0.002)."

The Coefficient: What Changed?

The coefficient (b) is the estimated change in the outcome (Y) for a 1-unit change in the predictor (X). It is the slope of the regression line.

  • Positive coefficient: As X increases, Y increases. (More patients → longer wait.)
  • Negative coefficient: As X increases, Y decreases. (More distance → fewer visits.)
  • Coefficient near zero: X has little to no linear effect on Y.
The Confidence Interval (CI): A Range of Plausibility

The 95% confidence interval gives a range of plausible values for the true coefficient in the population. It accounts for the fact that your sample might not perfectly represent the whole population.

💡 How to Interpret a CI in Regression:

Coefficient = 1.0, 95% CI: 0.75 to 1.25

"We are 95% confident that the true increase in waiting time per additional patient lies between 0.75 minutes and 1.25 minutes."

Critical rule for CIs in regression:

  • If the 95% CI includes 0, the evidence for an association is weak. The true effect could be zero.
  • If the 95% CI does NOT include 0, the evidence is stronger. There likely IS a real association.

📝 Exam Tip: Always interpret the size and direction of the coefficient, not just the p-value. A "statistically significant" result with a tiny coefficient may have no practical importance. A large coefficient with a wide CI may be uncertain.

The p-value: Evidence Against "No Effect"

The p-value tests the null hypothesis that the true coefficient is zero (no association). It asks: "If there were truly NO relationship in the population, how likely would we be to see a coefficient this large (or larger) in our sample?"

  • p < 0.05: The result is "statistically significant." We reject the null hypothesis. There is evidence of an association.
  • p ≥ 0.05: The result is "not statistically significant." We fail to reject the null. The evidence is insufficient.
⚠️ p-value Warnings:
  • A small p-value does NOT mean the effect is large or important.
  • A large p-value does NOT prove there is no effect — it just means you don't have enough evidence.
  • Always ask: "Is the effect size clinically meaningful?" not just "Is it statistically significant?"
Writing a Good Interpretation: The Formula

Use this structure for every regression interpretation in exams:

The Perfect Interpretation Sentence:

"For each additional [unit of X], [Y] [increased/decreased] by approximately [coefficient] [units of Y] (95% CI: [lower]–[upper]; p = [value])."

Example — Good vs. Weak interpretation:

  • Weak Interpretation: "The regression was significant, therefore patient load caused waiting time."
    Why weak: Claims causation. No numbers. No context.
  • Good Interpretation: "Each additional patient was associated with 1.0 extra minute of waiting time (95% CI: 0.75–1.25; p = 0.002)."
    Why good: Gives direction, size, units, CI, and p-value. Uses "associated with."
Group Task: Interpret a Regression Result

📊 Task: Write one plain-language sentence explaining this result.

Predictor Coefficient 95% CI p-value
Distance to facility (km) −0.40 −0.70 to −0.10 0.012

Requirements: Mention direction, size, statistical evidence, and use safe causal language.

Possible Answer: "For each additional kilometre from the health facility, the number of ANC visits decreased by approximately 0.40 visits. This association was statistically significant because the 95% confidence interval did not include zero, and p = 0.012."

Session 4: Integrated Revision — Connecting Epidemiology and Biostatistics
The Course Integration Map

Think of the entire course as one connected decision pathway. No part works in isolation:

  • Health Problem: What is happening?
  • Epidemiology: Who, where, when, why?
  • Data Collection: How will we measure it?
  • Biostatistics: What do the numbers show?
  • Decision: What action is justified?
Example of the full pathway:
  • Problem: Mothers in Village X are delivering at home instead of at the facility.
  • Epidemiology: Describe by person (young, first-time mothers), place (Village X, 12 km from facility), time (increased since road washed out in March).
  • Data: Collect distance, transport cost, ANC attendance, and delivery location for 100 mothers.
  • Biostatistics: Calculate mean distance. Run regression: distance predicts home delivery. r = −0.65 between distance and facility delivery.
  • Decision: Deploy mobile clinic to Village X twice monthly. Advocate for road repair.
Epidemiology Revision Checklist

Core ideas to master before the examination. For each topic, you should be able to define it, give a nursing example, and explain why it matters for public health action.

  • Disease classification — by cause (infectious, nutritional, genetic, environmental, behavioural), by duration (acute, subacute, chronic, recurrent), by transmission (communicable vs. non-communicable, vector-borne, water/food-borne, airborne), and by public health importance (common, severe, epidemic-prone, preventable, priority).
  • Person, place, and time — the three Ws of descriptive epidemiology. Be able to describe any outbreak using these three variables.
  • Chain of transmission — infectious agent, reservoir/source, portal of exit, route of transmission, portal of entry, susceptible host. Know how breaking any link stops spread.
  • Outbreak investigation steps — prepare, verify, define cases, find cases, describe by person/place/time, develop hypothesis, test hypothesis, implement control, communicate findings, follow up.
  • Study designs — cross-sectional (snapshot), cohort (follow forward), case-control (look backward), RCT (gold standard, randomised), ecological (population-level).
  • Sampling methods — simple random, systematic, stratified, cluster, convenience. Know why representativeness matters.
  • Incidence and prevalence — incidence = new cases over time (risk). Prevalence = all cases at one point in time (burden). Prevalence = Incidence × Average duration.
  • Measures of association — relative risk (RR), odds ratio (OR), attributable risk. Know when to use each.
📝 Exam Tip — Incidence vs. Prevalence:

Incidence asks: "How many new cases appeared this year?" (Like a water tap filling a bucket.)
Prevalence asks: "How many people currently have the disease?" (Like measuring how full the bucket is right now.)
Mnemonic: "Incidence = Incoming (new). Prevalence = Present (all existing)."

Biostatistics Revision Checklist

For each concept below, know the definition, the formula (where applicable), and how to interpret it in a nursing context.

  • Types of data and variables — categorical (nominal, ordinal) vs. numerical (discrete, continuous). This determines which statistical test you can use.
  • Mean, median, and mode — measures of central tendency. Mean is affected by outliers; median is robust. Use median for skewed data (like income, hospital stay).
  • Range and standard deviation — measures of spread/dispersion. Standard deviation tells you how far data points typically deviate from the mean.
  • Ratios, proportions, and rates — ratio (A:B), proportion (A / A+B, no time), rate (A / population-time, includes time). Risk is a proportion; incidence rate uses person-time.
  • Normal distribution and z-score — bell-shaped curve. 68% within 1 SD, 95% within 2 SD, 99.7% within 3 SD. z = (X − μ) / σ. Used for standardizing and comparing.
  • Confidence intervals — range of plausible values for a population parameter. 95% CI is standard. If CI for a difference includes 0, the result is not significant.
  • Hypothesis testing — null hypothesis (H₀: no effect) vs. alternative hypothesis (H₁: there is an effect). p < 0.05 rejects H₀. Type I error (false positive) vs. Type II error (false negative).
  • Correlation and regression — r measures strength and direction of linear association. Regression equation Ŷ = a + bX predicts Y from X. Interpret coefficient, CI, and p-value together.
🧠 Mnemonic — When to Use Mean vs. Median:

"If the data is Mean and clean (symmetric), use the mean. If the data is Messy and skewed, use the Median."
Example: Hospital length of stay is usually skewed (most people stay 2-3 days, a few stay 30 days). Always report median length of stay, not mean.

Choosing the Right Statistical Tool

The most common mistake in exams is using the wrong test. Start from the research question and the data type, then match the tool.

Research Question Data Type / Design Useful Tool
How many cases occurred? What proportion? Categorical counts Frequency / Proportion / Percentage
What is the average value? What is the middle value? Numerical data Mean (symmetric) or Median (skewed)
Are two categorical variables related? (e.g., gender vs. disease status) Two categorical variables Chi-square (χ²) test
Did the same group change after an intervention? (before vs. after) Paired numerical data Paired t-test
Do two independent groups differ? (treatment vs. control) Two independent numerical groups Independent (two-sample) t-test
Do two numeric variables move together? Two numerical variables Pearson correlation (r)
Can one variable predict another? One predictor, one outcome (both numeric) Simple linear regression
Compare means across three or more groups One categorical (3+ groups), one numerical ANOVA (Analysis of Variance)
📝 Exam Tip — Choosing the Right Test:

Ask yourself three questions in this order:

  • How many variables? One → descriptive stats. Two → association tests. Three+ → regression or ANOVA.
  • What type are they? Categorical vs. numerical determines everything.
  • Paired or independent? Same person measured twice = paired. Different people = independent.

Mnemonic: "Number, Type, Pairing" — NTP.

Integrated Case Study: Suspected Diarrhoeal Disease Outbreak in a School

🩺 Scenario: A school reports many pupils with diarrhoea after a shared meal. The health team records symptoms, class, sex, time of onset, water source, and food exposure.

Epidemiology Tasks:

  • Define cases: Establish a case definition (e.g., "any pupil with ≥3 loose stools in 24 hours, onset after 12:00 PM on [date]").
  • Descriptive epidemiology: Describe by person (age, class, sex), place (classroom, dormitory, water source), and time (epidemic curve showing cases by hour of onset).
  • Develop hypothesis: The shared meal is the likely source. Test by comparing attack rates among those who ate the meal vs. those who did not.
  • Implement control: Isolate sick pupils, provide ORS, inspect kitchen and water, withhold suspected food.

Biostatistics Tasks:

  • Calculate attack rates: (Cases among exposed ÷ Total exposed) × 100 for each exposure (meal, water source, class).
  • Compare exposed vs. unexposed: Use a 2×2 table to calculate relative risk or odds ratio.
  • Test significance: Use chi-square to test whether the difference in attack rates is statistically significant or due to chance.
  • Measure precision: Report 95% CI around the relative risk.
Worked Revision Example: Attack Rate

The attack rate is a special type of risk used during outbreak investigations. It tells you what percentage of exposed people actually got sick.

Attack Rate Formula:
Attack Rate = (Cases among exposed ÷ Total exposed) × 100

Example:

  • Exposed pupils (ate the suspected meal): 80
  • Cases among exposed: 32
  • Attack rate = (32 ÷ 80) × 100 = 40%
  • Interpretation: 40% of pupils who ate the meal became ill. This is very high and strongly suggests the meal was the source.

Next step: Calculate the attack rate among pupils who did NOT eat the meal. If it is 2%, the relative risk is 40% ÷ 2% = 20. Those who ate the meal were 20 times more likely to get sick.

📝 Exam Tip — Attack Rate:

Attack rate is essentially a risk calculated during an outbreak. It ALWAYS needs a denominator (total exposed). Never report just the number of cases. The examiner will ask: "Out of how many?"

Revision Question: Study Design

📋 Question: A researcher wants to compare children with measles and children without measles to determine whether vaccination history differed between the two groups.
Which study design is most appropriate?

Answer: Case-control study.

Why? The groups are selected by outcome status (measles vs. no measles), and then previous exposure (vaccination history) is compared. This is the classic case-control structure: start with the outcome, look backward for exposure.

Why not cohort? A cohort study would start with vaccinated and unvaccinated children and follow them forward to see who gets measles. That would also work, but the question describes starting with the outcome, which defines case-control.

Revision Question: Statistical Test

📋 Question: A nurse compares blood pressure in the same patients before and after counselling.
Which statistical test is most appropriate?

Answer: Paired t-test.

Why? The same individuals are measured twice (before and after). The measurements are paired or dependent. An independent t-test would be wrong because it assumes two separate groups.

Alternative: If the data is not normally distributed, use the Wilcoxon signed-rank test (non-parametric equivalent).

Revision Question: Correlation vs. Regression — NEW

📋 Question: A district health officer wants to know whether rainfall predicts malaria cases next month, and by how many cases rainfall must increase to expect a noticeable rise. Should they use correlation or regression?

Answer: Simple linear regression.

Why? The officer wants to predict the number of malaria cases (Y) from rainfall (X) and know how much Y changes per unit of X. Correlation would only tell them that rainfall and malaria are related — it cannot predict the number of cases or quantify the change.

Session 5: Final Course Wrap-up and Examination Guidance
Continuous Assessment Feedback

Use your continuous assessment feedback to improve before the final examination. Examiners mark four dimensions:

Dimension What Examiners Look For How to Improve
Content accuracy Correct definitions, correct formulas, correct interpretation. Make a one-page formula sheet. Test yourself with flashcards.
Calculation clarity Show the formula, substitute the numbers, and give the final answer with units. Never skip steps. Write: Formula → Substitution → Answer → Interpretation.
Evidence of reasoning Explain WHY a design or test is appropriate. Do not just name it. Practice saying: "I chose X because the data are paired/numerical/etc."
Communication Write clear public health conclusions. Connect numbers to action. End every calculation with: "This means that..." and suggest an action.
Quiz 2: What to Revise

The final assessment typically covers the entire course, with emphasis on:

  • Probability and normal distribution — calculating probabilities from z-scores, understanding the bell curve.
  • z-score and confidence intervals — formula, interpretation, and what "95% confident" actually means.
  • Null and alternative hypotheses — writing them correctly, knowing what p < 0.05 means in context.
  • Chi-square test — when to use it, how to interpret the result, what a 2×2 table looks like.
  • Paired t-test and independent t-test — know the difference. Paired = same people twice. Independent = two different groups.
  • Scatter diagrams, correlation, and regression — interpreting r, reading regression tables, writing safe causal language.
  • Outbreak investigation — steps, attack rate, case definition, person/place/time description.
  • Study design selection — case-control vs. cohort vs. cross-sectional vs. RCT.
💡 Study Tip:

"Revise by solving examples, not by reading definitions only." You cannot learn biostatistics by highlighting text. You must calculate, interpret, and write conclusions.

How to Answer Calculation Questions

Use a clear three-step structure. Marks are often lost when the final number is given without interpretation.

  • Write the formula: State the formula in symbols first. This shows you know the concept, even if you make a calculator error later.
    Example: Risk = (New Cases ÷ Total at Risk) × 100
  • Substitute the numbers: Plug in the values from the question. Show the substitution clearly.
    Example: Risk = (15 ÷ 120) × 100
  • State the answer and its meaning: Give the final number WITH units, then interpret what it means for public health or patient care.
    Example: "Risk = 12.5%. This means that 12.5% of the nursing students in the hostel developed fever during the outbreak period."
📝 Exam Tip — The Marking Scheme:

In many exams, the final number is worth only 1 mark. The formula is worth 1 mark. The substitution is worth 1 mark. The interpretation is worth 2 marks. Students who only write "12.5%" lose 60% of the marks. Always interpret.

How to Answer Short Essay Questions

Use the DEEC structure for every short essay or explanation question:

  • D - Define: Give the key meaning in your own words. Start with the definition.
  • E - Explain: Show how it works. What are the components? What does it measure?
  • E - Example: Use a real health or nursing situation. Make it specific and relevant.
  • C - Conclude: State the public health implication. Why does this matter? What action follows?

🧠 Mnemonic: "Doctor Explains Every Case" — DEEC = Define, Explain, Example, Conclude.

DEEC in action — answering "What is relative risk?"
  • Define: "Relative risk (RR) is the ratio of the risk of disease in the exposed group to the risk in the unexposed group."
  • Explain: "It tells us how many times more likely the exposed group is to develop the outcome compared to the unexposed group. An RR of 1 means no difference. RR > 1 means increased risk. RR < 1 means protective effect."
  • Example: "In a village outbreak, children who drank from the contaminated well had an attack rate of 30%, while those who did not drink from it had an attack rate of 5%. The relative risk is 30 ÷ 5 = 6.0."
  • Conclude: "This means children who drank the contaminated water were 6 times more likely to get diarrhoea. The nurse should immediately close the well, provide safe water, and educate the community."
Final Study Strategy

Move from memorisation to application. Here is a proven revision plan:

  • Prepare a one-page formula sheet. Include every formula from the course: mean, SD, z-score, CI, risk, rate, RR, OR, attack rate, correlation, regression. Keep it concise.
  • Practice at least five calculation examples for each major topic. Do not just read — calculate with a calculator.
  • Explain each statistical test in one sentence. If you cannot explain it simply, you do not understand it well enough.
  • Revise outbreak investigation steps using a case. Walk through a real or imagined outbreak from case definition to control measures.
  • Teach one topic to a classmate. Teaching forces you to organise your knowledge and identify gaps.
🎯 The Best Revision Test

Can I explain the result to a patient, a nurse manager, or a community leader?
If you can translate a p-value or a regression coefficient into language a village health team understands, you have truly mastered the material.

Common Exam Traps — Final Warnings
Trap Why You Lose Marks How to Avoid It
Forgetting units "The risk is 12.5" is meaningless. 12.5 what? Percent? Proportion? Always write "%" or "per 1000" or "proportion."
Confusing incidence and prevalence These are the most commonly confused measures. Incidence = new cases over time. Prevalence = all cases at one point.
Claiming causation from correlation Examiners deliberately test this. It is the #1 conceptual error. Use "associated with." Only RCTs or strong evidence support "caused."
Ignoring the denominator Risk without a denominator is just a count. Always ask: "Out of how many?" Include total at risk in every risk calculation.
Choosing the wrong statistical test Using a t-test for categorical data, or chi-square for numerical data. Use the NTP rule: Number of variables, Type of data, Paired or independent.
Final Takeaway
  • Good epidemiology asks the right question.
  • Good biostatistics helps answer it clearly.
  • Good nursing practice uses the answer to improve health.

You are not just learning formulas. You are learning to turn data into decisions that save lives.

Bonus: One-Page Quick Reference — Key Formulas
  • Descriptive Statistics
    • Mean = ΣX / n
    • SD = √[ Σ(X − mean)² / (n − 1) ]
    • z = (X − μ) / σ
    • Median = middle value when sorted
  • Measures of Disease Frequency
    • Risk = (New cases / Total at risk) × 100
    • Rate = (Cases / Person-time) × multiplier
    • Prevalence = (All cases / Total population) × 100
    • Attack Rate = (Cases among exposed / Total exposed) × 100
  • Measures of Association
    • Relative Risk (RR) = Risk in exposed / Risk in unexposed
    • Odds Ratio (OR) = (a × d) / (b × c) [from 2×2 table]
  • Correlation & Regression
    • r = correlation coefficient (−1 to +1)
    • Ŷ = a + bX [regression equation]
    • Slope (b) = change in Y / change in X
  • Confidence Interval
    • 95% CI = estimate ± (1.96 × SE)
    • If CI includes 0 (for differences) or 1 (for ratios), result is not significant.
References
  • Gordis, L. (2014). Epidemiology. Elsevier Saunders.
  • Daniel, W. W., & Cross, C. L. (2018). Biostatistics: A Foundation for Analysis in the Health Sciences. Wiley.
  • Webb, P., & Bain, C. (2010). Essential Epidemiology: An Introduction for Students and Health Professionals. Cambridge University Press.
  • Centers for Disease Control and Prevention (CDC). Principles of Epidemiology in Public Health Practice.

Quick Quiz

Correlation, Regression, Integration Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Correlation, Regression, Integration Read More »

Probability, Normal Distribution & Confidence Intervals

Probability, Normal Distribution & Confidence Intervals

Probability, Normal Distribution & Confidence Intervals
Learning Outcomes

By the end of this lecture, you should be able to:

  • Define probability and apply simple probability rules in health examples.
  • Recognise a normal distribution and explain why it matters in statistics.
  • Compute a z-score and explain whether a value is typical or unusual.
  • Calculate and interpret 95% confidence intervals for means and proportions.
  • Communicate statistical uncertainty in plain, correct language.

🎯 The Big Picture: Health data are often incomplete samples. We cannot measure every patient in the country. Probability helps us describe how much uncertainty remains when we use sample results to understand a wider population. The journey is:

Probability ➔ Normal Curve ➔ Z-Score ➔ Confidence Interval
Introduction to Probability

Probability is the language of uncertainty. In nursing and public health, we rarely deal with absolute certainties. A test may be "likely" positive. A patient may be "at risk" of complications. Probability gives us a number to express that uncertainty.

Key Terms
Term Definition Nursing Example
Outcome One possible result of a process. A malaria RDT result is positive.
Event A group of one or more outcomes. The patient has malaria (could be confirmed by RDT, microscopy, or clinical signs).
Probability A number from 0 to 1 showing how likely an event is. There is a 0.30 (30%) chance that a mother will book ANC before 12 weeks.
Basic Probability Formula

Probability = Number of Favourable Outcomes ÷ Total Possible Outcomes

Example: If 30 out of 100 mothers attend ANC before 12 weeks, the probability is 30 ÷ 100 = 0.30 = 30%.

🧠 Think About It: Probability is NOT just about gambling or coin flips. In health, probability tells us: "Out of every 100 patients like this, how many will experience this outcome?" This is the foundation of evidence-based nursing.
Basic Probability Rules

Three essential rules every nurse should use correctly when interpreting health data:

The Range Rule

Probability cannot be below 0 or above 1.

  • 0 = Impossible (e.g., the probability that a living patient has a negative heart rate).
  • 1 = Certain (e.g., the probability that a patient who has died will not recover).
  • 0.5 = 50/50 chance (e.g., a coin flip — though health probabilities are rarely this neat).
The Complement Rule

If an event either happens or does not happen, the probabilities must add to 1.

P(not A) = 1 − P(A)

Example: If P(malaria) = 0.25 in a community, then P(no malaria) = 1 − 0.25 = 0.75 (or 75%).

Why this matters: If you know 15% of patients have hypertension, you immediately know 85% do not. This is useful for planning resources and understanding risk.

The "Either/Or" Rule (Addition Rule)

For non-overlapping (mutually exclusive) events, you can add their probabilities:

P(A or B) = P(A) + P(B)

Example: In a ward, the probability a patient has malaria is 0.20, and the probability a patient has typhoid is 0.10. Assuming no patient has both, the probability a random patient has either malaria or typhoid is 0.20 + 0.10 = 0.30 (30%).

⚠️ Key Caution: You can ONLY add probabilities when the events cannot happen at the same time (mutually exclusive). If a patient could have BOTH malaria and typhoid, you would need a more advanced formula (P(A or B) = P(A) + P(B) − P(A and B)). For your exam, stick to mutually exclusive events unless told otherwise.
Worked Example: Probability in a Clinic
🩺 Scenario

A midwife reviews ANC attendance records at a health centre. Out of 150 mothers who registered for ANC:

ANC Attendance Status Number of Mothers Probability
Attended first ANC before 12 weeks 45 45/150 = 0.30
Attended first ANC at 12 weeks or later 105 105/150 = 0.70
Total 150 1.00
  • Interpretation: In this clinic sample, the probability that a randomly selected mother booked ANC before 12 weeks is 0.30, or 30%.
  • Complement Check: P(late ANC) = 1 − P(early ANC) = 1 − 0.30 = 0.70. This matches the table — a good way to verify your calculations.
  • Clinical Application: If only 30% of mothers book early, the nurse manager knows that 70% are at higher risk for complications. This data supports an intervention: community health worker outreach, transport vouchers, or male partner involvement programs.
Additional Scenario: Probability of Vaccination Status (NEW)

🩺 Scenario: In a village of 200 children under five, a nurse finds:

  • 140 are fully immunised.
  • 40 are partially immunised.
  • 20 are unimmunised.
Calculations:
  • P(fully immunised) = 140/200 = 0.70 (70%)
  • P(partially immunised) = 40/200 = 0.20 (20%)
  • P(unimmunised) = 20/200 = 0.10 (10%)
  • P(not fully immunised) = 1 − 0.70 = 0.30 (30%) using the complement rule.
  • P(either partially immunised OR unimmunised) = 0.20 + 0.10 = 0.30 (30%) using the addition rule (mutually exclusive).

Action: The 30% gap is a public health priority. The nurse can now argue for a catch-up campaign with precise numbers.

The Normal Distribution

The normal distribution is a bell-shaped pattern found in many biological and health measurements. It is one of the most important concepts in statistics because it allows us to make predictions about what is "normal" and what is "unusual."

Key Features of the Normal Curve
  • Symmetric around the mean — the left side is a mirror image of the right side.
  • Mean, median, and mode are all at the exact centre of the curve.
  • Most values cluster near the mean. The curve is highest in the middle.
  • Fewer values appear at the extremes. The "tails" get thinner as you move away from the centre.
  • The total area under the curve equals 1 (or 100%), representing all possible outcomes.
Examples of approximately normal measurements in health:
  • Adult height in a homogeneous population.
  • Birth weight of full-term babies.
  • Systolic blood pressure in a healthy adult population.
  • Haemoglobin levels in non-anaemic adults.
  • Examination scores in a large class.
⚠️ Important: Not every dataset is normal. Income distribution is usually skewed (most people earn little, a few earn a lot). Disease incidence may be clustered. Always look at the distribution before assuming normality. In your exam, you will usually be told when to assume normality.
Why the Normal Curve Matters in Nursing
Application Why It Matters
Clinical Measurement A value can be compared with the expected average. Example: a baby with birth weight 2.1 kg can be compared to the population mean of 3.0 kg. Is this baby unusually small?
Sampling Distribution Even if individual data are not normal, the means of many samples tend to form a normal pattern — especially when the sample size is large (Central Limit Theorem). This is why we can use normal-based formulas for confidence intervals.
Confidence Intervals Normal theory helps us estimate how far a sample result (like a mean) may be from the true population value. It quantifies our uncertainty.

📝 Exam Tip: When asked "Why does the normal distribution matter?" mention at least two of these three: clinical comparison, sampling distribution, and confidence intervals.

The Empirical Rule (68, 95, 99.7 Rule)

For any data that follows a normal distribution, the spread of values follows a remarkably predictable pattern:

  • 68% of values fall within 1 SD of the mean
  • 95% of values fall within 2 SD of the mean
  • 99.7% of values fall within 3 SD of the mean

Mnemonic: "68, 95, 99.7 — Almost All Are Near the Middle"

What this means practically: If you know the mean and standard deviation of a normally distributed measurement, you can immediately say what range covers "most" patients, and you can flag values that are unusually high or low.

Worked Example: Empirical Rule in Practice
🩺 Scenario

Systolic blood pressure is measured in a ward. The data are approximately normally distributed with:

  • Mean (μ) = 120 mmHg
  • Standard Deviation (SD) = 10 mmHg
Range Calculation Approx. % of Patients
110 to 130 mmHg 120 ± 1 SD 68%
100 to 140 mmHg 120 ± 2 SD 95%
90 to 150 mmHg 120 ± 3 SD 99.7%
  • Interpretation: A patient with systolic BP of 150 mmHg is about 3 SD above the mean. Only about 0.15% of patients in this population would be expected to have BP this high or higher. This patient should be considered unusually high and requires closer clinical attention — possibly immediate intervention.
  • Conversely: A patient with BP of 115 mmHg is within 1 SD of the mean. This is typical and expected. No alarm needed.
Additional Scenario: Birth Weights in a Maternity Ward (NEW)

🩺 Scenario: Birth weights in a district hospital are normally distributed with mean = 3.2 kg and SD = 0.5 kg.

Applying the Empirical Rule:
  • 68% of babies weigh between 2.7 kg and 3.7 kg (3.2 ± 0.5).
  • 95% of babies weigh between 2.2 kg and 4.2 kg (3.2 ± 1.0).
  • 99.7% of babies weigh between 1.7 kg and 4.7 kg (3.2 ± 1.5).

Clinical Application: A baby born at 1.8 kg is below the 3 SD lower limit. This is extremely unusual and signals possible prematurity, intrauterine growth restriction, or maternal malnutrition. The nurse should flag this immediately for paediatric review. A baby at 3.0 kg is well within the normal range — routine care is appropriate.

Standard Score (Z-Score)

A z-score (also called a standard score) tells us how many standard deviations a particular value is from the mean. It converts any measurement into a common scale, allowing comparison across different variables.

Z-Score Formula

z = (Observed Value − Mean) ÷ Standard Deviation

z = (x − μ) ÷ σ

What Different Z-Scores Mean
Z-Score Meaning Health Interpretation
z = 0 Exactly at the mean. Typical, average value. No concern.
z = +1 1 SD above the mean. Higher than average, but still common (about 16% of population is above this).
z = −1 1 SD below the mean. Lower than average, but still common (about 16% of population is below this).
z = +2 or more 2 or more SD above the mean. Unusually high. Only ~2.5% of population is above this. May need investigation.
z = −2 or less 2 or more SD below the mean. Unusually low. Only ~2.5% of population is below this. Often a clinical red flag.

📝 Exam Tip: A z-score changes different measurements into a common scale. This means you can compare a baby's birth weight z-score with another baby's haemoglobin z-score, even though the original units (kg vs. g/dL) are completely different.

Worked Example: Calculating a Z-Score
🩺 Scenario

A baby is born with a birth weight of 2.1 kg. In the population, the mean birth weight is 3.0 kg with a standard deviation of 0.45 kg.

  • z = (2.1 − 3.0) ÷ 0.45
  • z = −0.9 ÷ 0.45
  • z = −2.0

Interpretation: The baby's birth weight is 2 standard deviations below the mean. According to the empirical rule, only about 2.5% of babies would be expected to weigh this little or less. This is unusually low and may require closer clinical attention — kangaroo mother care, warming, feeding support, and possible referral.

Why z-scores matter in nursing: Instead of just saying "the baby is small," the nurse can say "the baby is 2 SD below the population mean." This is precise, comparable across hospitals, and immediately communicates severity to doctors and referral facilities.

Additional Scenario: Comparing Two Babies (NEW)

🩺 Scenario: Two babies are born at the same hospital:

  • Baby A: Birth weight = 2.5 kg. Population mean = 3.0 kg, SD = 0.5 kg.
  • Baby B: Birth weight = 2.8 kg. Population mean = 3.5 kg, SD = 0.4 kg.
Which baby is more unusually small for their population?
  • Baby A: z = (2.5 − 3.0) ÷ 0.5 = −1.0 (1 SD below mean — somewhat small, but common).
  • Baby B: z = (2.8 − 3.5) ÷ 0.4 = −1.75 (1.75 SD below mean — more unusually small for their population).

Conclusion: Even though Baby B weighs more in absolute terms (2.8 kg vs. 2.5 kg), Baby B is more unusually small relative to their population. This is why z-scores are powerful — they allow fair comparison across different groups.

⚠️ Rule of Thumb: Values beyond ±2 SD (z-scores below −2 or above +2) are often considered unusual. However, clinical context still matters. A z-score of −1.9 for birth weight may still trigger action in a resource-limited setting with high neonatal mortality. The z-score is a guide, not a replacement for clinical judgment.
From Sample to Population: Understanding Uncertainty

In real-world nursing and public health, we almost never measure the entire population. We take a sample and use it to estimate what is true for the whole population. But samples are imperfect — they contain sampling error.

Key Concepts
Term Definition & Example
Point Estimate A single number from the sample that estimates the population value. Example: The sample mean haemoglobin is 11.2 g/dL. This is our best guess for the population mean — but it is probably not exactly right.
Sampling Error The natural, unavoidable difference between a sample result and the true population value. If you took a different sample of 50 mothers, the mean Hb would likely be slightly different. This is NOT a mistake — it is expected variation.
Confidence Interval (CI) A range of plausible values for the true population value. Instead of saying "the mean is 11.2," we say "the true mean is plausibly between 10.8 and 11.6." This makes uncertainty visible and honest.
💡 Analogy: A point estimate is like a single photograph — it captures one moment but may miss the bigger picture. A confidence interval is like a panoramic shot — it shows the range of what the true scene probably looks like.
Confidence Interval for a Mean

Use this when your outcome is numerical (continuous data) — for example, average birth weight, average waiting time, average haemoglobin level.

95% CI for a Mean

95% CI = mean ± 1.96 × (SD ÷ √n)

95% CI = mean ± 1.96 × SE

What Each Symbol Means
  • Mean (x̄): The average calculated from your sample.
  • SD (standard deviation): How much individual values vary around the mean.
  • n (sample size): The number of people or observations in your sample.
  • SE (standard error): SD ÷ √n. This measures the uncertainty around the sample mean. It tells us how far sample means typically vary from the true population mean.
  • 1.96: The "multiplier" from the normal distribution that gives us approximately 95% confidence. (For 90% CI, use 1.645. For 99% CI, use 2.576.)

📝 Exam Tip — The Magic Number 1.96: For a 95% confidence interval, always use 1.96 (or approximately 2 for quick mental calculations). This number comes from the normal distribution: 95% of the area under the curve lies within 1.96 standard deviations of the mean.

Worked Example: CI for a Mean
🩺 Scenario

A hospital quality improvement team wants to know the average length of stay for pneumonia patients. They review 64 patient records.

  • n = 64 patients
  • Mean stay = 4.2 days
  • SD = 1.6 days
Step-by-Step Calculation:
  • Calculate SE: SE = SD ÷ √n = 1.6 ÷ √64 = 1.6 ÷ 8 = 0.20 days
  • Calculate Margin of Error: 1.96 × SE = 1.96 × 0.20 = 0.39 days
  • Calculate 95% CI: 4.2 ± 0.39
  • Result: 95% CI = 3.81 to 4.59 days

Interpretation: We are 95% confident that the true mean length of stay for all pneumonia patients at this hospital (not just the 64 sampled) is between 3.81 and 4.59 days.

What does "narrow" vs. "wide" mean?
  • A narrow interval (like 4.1 to 4.3 days) suggests a precise estimate — we have a good idea of the true value.
  • A wide interval (like 2.5 to 6.0 days) suggests more uncertainty — we need more data or the variation is very high.
  • How to make the interval narrower: Increase the sample size (n). The larger your sample, the more precise your estimate. This is why SE = SD ÷ √n — as n gets bigger, the denominator gets bigger, so SE gets smaller.
How to Interpret a 95% Confidence Interval Correctly

This is where many students lose marks in exams. The wording matters.

✅ CORRECT Wording:
  • "The data are consistent with a true mean between 3.81 and 4.59 days."
  • "We are 95% confident that the true population mean lies between 3.81 and 4.59 days."
  • "The plausible range for the true mean is 3.81 to 4.59 days."
❌ INCORRECT Wording (Avoid These):
  • "There is a 95% chance that this exact interval contains the true mean." — Wrong! The true mean is fixed. The interval either contains it or it doesn't. The 95% refers to the METHOD, not this specific interval.
  • "95% of patients stay between 3.81 and 4.59 days." — Wrong! The CI is about the MEAN, not individual patients. Individual stays vary much more widely.
  • "The true mean is definitely between 3.81 and 4.59 days." — Wrong! There is still a 5% chance the true mean falls outside this range.
💡 The Correct Way to Think About It: If we took 100 different samples and calculated a 95% CI from each, about 95 of those intervals would capture the true population mean, and about 5 would miss it. We don't know if OUR specific interval is one of the 95 or one of the 5 — but the method is reliable 95% of the time.
Confidence Interval for a Proportion

Use this when your outcome is categorical (yes/no, present/absent, positive/negative) — for example, the proportion of children immunised, the proportion of mothers satisfied with care, the proportion of patients testing positive for malaria.

95% CI for a Proportion

95% CI = p ± 1.96 × √[ p(1 − p) ÷ n ]

What Each Symbol Means
  • p: The sample proportion (as a decimal). Example: if 72 out of 120 children are immunised, p = 72/120 = 0.60.
  • (1 − p): The complement of the proportion. Example: if p = 0.60, then 1 − p = 0.40 (the proportion NOT immunised).
  • n: The total sample size.
  • √[ p(1 − p) ÷ n ]: The standard error (SE) for a proportion.

⚠️ Critical Caution: When calculating a CI for a proportion, you MUST use decimals (0.60), not percentages (60%), inside the formula. Only convert back to percentages at the very end for your final interpretation.

Worked Example: CI for a Proportion
🩺 Scenario

A nursing student surveys immunisation coverage in a village. Out of 120 children under five, 72 were fully immunised.

  • n = 120 children
  • Number immunised = 72
  • p = 72 ÷ 120 = 0.60
Step-by-Step Calculation:
  • Calculate SE: SE = √[ p(1 − p) ÷ n ] = √[ 0.60 × 0.40 ÷ 120 ]
  • SE = √[ 0.24 ÷ 120 ] = √0.002 = 0.0447
  • Calculate Margin: 1.96 × 0.0447 = 0.0876
  • Calculate 95% CI: 0.60 ± 0.0876
  • Result: 95% CI = 0.512 to 0.688 = 51.2% to 68.8%

Interpretation: We are 95% confident that the true immunisation coverage in the entire village (not just the 120 sampled children) is between 51.2% and 68.8%.

Why is this interval fairly wide?
  • The sample size (n = 120) is moderate. If we surveyed 500 children, the interval would be much narrower.
  • The proportion (0.60) is in the middle — proportions near 0.50 have the largest standard errors.
  • Clinical/Public Health Action: Even the upper bound (68.8%) is below the national target of 90%. This confirms that immunisation coverage is inadequate and justifies a catch-up campaign.
Choosing the Right Confidence Interval

The first step is always to identify your outcome type. This determines which formula to use.

Research Question Outcome Type Use This CI
What is the average birth weight? Numerical (continuous) CI for a mean
What proportion delivered in a facility? Yes/No (categorical) CI for a proportion
What is the average waiting time? Numerical (continuous) CI for a mean
What proportion tested malaria-positive? Yes/No (categorical) CI for a proportion

📝 Quick Check:

  • Mean = average amount (kg, minutes, g/dL, mmHg) ➔ Use CI for a mean.
  • Proportion = share of people with a characteristic (%, fraction) ➔ Use CI for a proportion.

Mnemonic: "Means are Measured; Proportions are People."

Class Practice: Full Worked Example
🩺 Scenario

A health centre records 50 postnatal mothers.

  • Mean waiting time = 36 minutes, SD = 14 minutes.
  • 32 mothers report satisfaction with services.
Task 1: 95% CI for Mean Waiting Time
  • SE = 14 ÷ √50 = 14 ÷ 7.071 = 1.98
  • Margin = 1.96 × 1.98 = 3.88
  • 95% CI = 36 ± 3.88 = 32.1 to 39.9 minutes
Task 2: Proportion Satisfied
  • p = 32/50 = 0.64 (64%)
Task 3: 95% CI for Satisfaction Proportion
  • SE = √[0.64 × 0.36 ÷ 50] = √0.004608 = 0.0679
  • Margin = 1.96 × 0.0679 = 0.133
  • 95% CI = 0.64 ± 0.133 = 0.507 to 0.773 = 50.7% to 77.3%

Interpretation: The satisfaction estimate is less precise — the interval is fairly wide. With only 50 mothers, there is considerable uncertainty about the true satisfaction rate.

Common Mistakes to Avoid
Mistake How to Fix It
Using percentages instead of decimals in proportion formulas. Always convert 60% ➔ 0.60 before calculating. Convert back at the end.
Forgetting to divide SD by √n when calculating SE for a mean. SE = SD ÷ √n. The √n is essential — it converts individual variation into uncertainty about the mean.
Interpreting a CI as a guarantee or a probability about one interval. Say "we are 95% confident" or "the plausible range is." Do not say "95% chance."
Reporting only the p-value and hiding the confidence interval. Always report the CI — it shows the size of the effect AND the precision.
Ignoring clinical context when deciding if a value is important. A z-score of −1.9 may not be "statistically unusual" but may still need clinical action.
Final Takeaway & Exam Summary
  • Probability: How likely is an event? Use counts ➔ convert to decimals.
  • Normal Distribution: Bell-shaped. 68, 95, 99.7 rule. Most values near the mean.
  • Z-Score: z = (x − mean) ÷ SD. Beyond ±2 = unusual. Common scale for comparison.
  • Confidence Interval: Mean: x̄ ± 1.96×(SD/√n). Proportion: p ± 1.96×√[p(1−p)/n]. Shows precision.
📝 Exam Formula Sheet (Memorise These):
  • Probability: P = favourable ÷ total
  • Complement: P(not A) = 1 − P(A)
  • Z-score: z = (x − μ) ÷ σ
  • SE for mean: SE = SD ÷ √n
  • 95% CI for mean: mean ± 1.96 × SE
  • SE for proportion: SE = √[p(1−p) ÷ n]
  • 95% CI for proportion: p ± 1.96 × SE
💡 Final Thought: Good health statistics do not only give numbers. They explain what the numbers mean and how uncertain they are. As a nurse, when you report that "64% of mothers are satisfied (95% CI: 51–77%)," you are communicating both the finding AND its reliability. That is the mark of a statistically literate health professional.
References
  • Rosner, B. (2015). Fundamentals of Biostatistics (8th ed.). Cengage Learning.
  • Daniel, W. W., & Cross, C. L. (2018). Biostatistics: A Foundation for Analysis in the Health Sciences (11th ed.). Wiley.
  • Gordis, L. (2013). Epidemiology (5th ed.). Saunders Elsevier.
  • Heavey, E. (2018). Statistics for Nursing: A Practical Approach (3rd ed.). Jones & Bartlett Learning.

Quick Quiz

Probability, Normal Distribution and Confidence Intervals

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Probability, Normal Distribution & Confidence Intervals Read More »

Probability, Hypothesis Testing & Statistical Tests

Probability, Hypothesis Testing & Statistical Tests

Probability, Hypothesis Testing & Statistical Tests
🎯 Learning Outcomes

By the end of this session, you should be able to:

  • Explain the purpose of hypothesis testing in health research and nursing practice.
  • Write clear null (H₀) and alternative (H₁) hypotheses for simple clinical questions.
  • Use p-values and significance levels correctly without overclaiming.
  • Select chi-square, paired t-test, or independent t-test appropriately based on data structure.
  • Interpret test results in clear nursing and public health language — not just statistical jargon.
Why Hypothesis Testing Matters for Nurses

In real clinical and community practice, we rarely measure everybody. We take a sample a group of patients, a ward, a village and use that sample to make a careful, evidence-based judgement about the wider population.

The Inference Pathway

Sample Result ➔ Statistical Test ➔ Evidence-Based Conclusion

Real Nursing Examples
  • Hand Hygiene: Does a 2-hour hand hygiene training session improve compliance rates among student nurses on the maternity ward?
  • Malaria Prevention: Is malaria positivity different between children who sleep under an insecticide-treated net and those who do not?
  • Patient Education: Do mothers' knowledge scores about ORS preparation increase after a health education session?
  • Clinic Efficiency: Is the average waiting time at Clinic A significantly shorter than at Clinic B?
  • Nutrition: Is there an association between a child's sex and their nutritional status (wasted vs. not wasted)?
🧠 Think About It: Without hypothesis testing, we might see that 8 out of 10 trained nurses washed their hands, while only 5 out of 10 untrained nurses did. But is that difference real, or could it just be random luck in who we picked for the sample? Hypothesis testing gives us the mathematical tools to answer that question with confidence.
Statistical Inference:

Statistical inference is the bridge between what we observe in our sample and what we believe is true in the entire population.

Component What It Is Nursing Example
KNOWN
Sample observed data
The actual numbers we collected from our study group. We measured Hb levels in 120 pregnant mothers attending ANC at Hospital X.
METHOD
Probability and tests
The mathematical rules (t-tests, chi-square, p-values) that judge whether our sample result is meaningful. We run a statistical test to compare the mean Hb before and after iron supplementation.
UNKNOWN
Population truth
The real situation in the entire population that we can never measure perfectly. The true mean Hb among ALL pregnant mothers in the district, not just our 120.

💡 Core Question of Inference: "Is the observed difference likely to be real, or could it be due to chance?"

A test does not prove truth; it measures how compatible the data are with the assumption of "no effect" (H₀). We are essentially asking: "If nothing were really happening, how surprised would I be to see these results?"

Key Terms You Must Know
Term Simple Meaning Nursing Example
Parameter The true value in the entire population. Usually unknown. The true mean haemoglobin (Hb) level among all pregnant mothers in Uganda.
Statistic A value calculated from your sample data. This is what you actually know. The mean Hb in the 120 mothers you actually tested at your clinic.
Sampling Error The natural difference between a sample statistic and the true population parameter. It exists simply because we did not measure everyone. Your sample mean Hb is 11.2 g/dL, but the true population mean might be 11.5 g/dL. That gap is sampling error.
Statistical Test A formal mathematical rule for judging whether sample evidence is strong enough to support a conclusion. Chi-square test, paired t-test, independent t-test, ANOVA.
Hypothesis A clear, testable statement about a population that can be supported or refuted by data. "Mothers who receive nutrition counselling have higher mean Hb than those who do not."
📝 Exam Tip — Parameter vs. Statistic:
Mnemonic: "Parameter = Population" and "Statistic = Sample". Both start with the same letter. If you know the whole population, it's a parameter. If you only have a sample, it's a statistic.
From Research Question to Hypothesis

Every statistical test begins with a question. But a research question is usually written in plain English. To test it statistically, we must translate it into two formal statements: the null hypothesis (H₀) and the alternative hypothesis (H₁).

Research Question Null Hypothesis (H₀) Alternative Hypothesis (H₁)
Does training improve knowledge scores among nurses? Mean score before training = Mean score after training. (No difference) Mean score before training ≠ Mean score after training. (There is a difference)
Is malaria positivity related to net use? There is no association between net use and malaria positivity. There is an association between net use and malaria positivity.
Do two clinics have different waiting times? Mean waiting time in Clinic A = Mean waiting time in Clinic B. Mean waiting time in Clinic A ≠ Mean waiting time in Clinic B.
Does iron supplementation reduce anaemia? Mean Hb after supplementation = Mean Hb before supplementation. Mean Hb after supplementation > Mean Hb before supplementation.
Is there a link between ward type and infection rate? Ward type and infection rate are independent (no association). Ward type and infection rate are associated.

❌ Common Student Error: Writing H₀ as "The training will improve scores." That is actually H₁! H₀ must always be the "no effect / no difference / no association" statement. Think of H₀ as the skeptic's position — it assumes nothing interesting is happening until proven otherwise.

Null and Alternative Hypotheses in Depth
H₀ — The Null Hypothesis
  • It is the "no difference" or "no association" statement.
  • It represents the status quo — what we would expect if chance alone were operating.
  • We test H₀ because random variation can easily mislead us, especially with small samples.
  • We never "prove" H₀ is true. We either reject it or fail to reject it.
H₁ — The Alternative Hypothesis
  • It is the statement the researcher thinks may be supported by evidence.
  • It represents the new idea, the intervention effect, or the association we suspect.
  • It is what we hope to find evidence for by finding evidence against H₀.
⚖️ Analogy — The Courtroom:
Think of H₀ as "innocent until proven guilty." The null hypothesis claims "nothing happened." Our data is the evidence. If the evidence is strong enough, we reject the claim of innocence (reject H₀). If the evidence is weak, we do not declare innocence — we simply say "not guilty enough to convict" (fail to reject H₀).
A Statistically Significant Result

When we say a result is "statistically significant," we mean: "The data we observed would be very unlikely if H₀ were actually true." Therefore, we have reasonable grounds to doubt H₀ and consider H₁ instead.

📝 Exam Tip — Wording:
✅ Correct: "We reject H₀" or "We fail to reject H₀."
❌ Incorrect: "We accept H₀" or "H₀ is true." — The test only tells us whether evidence is strong enough to discard H₀, not whether H₀ is definitively true.
One-Tailed vs. Two-Tailed Tests
Type Question It Answers When to Use
Two-tailed "Is there any difference?" (Could be higher OR lower) Most classroom and research situations. Use this as your default. It tests for a difference in either direction.
One-tailed "Is one group specifically higher or lower?" Only when the direction is justified before analysis based on strong theory or prior evidence. E.g., "We expect training to increase scores, not decrease them."

🛡️ Safe Rule for Students: Use a two-tailed test unless the study clearly planned a directional test before seeing the data. Never choose a one-tailed test just because it makes your p-value smaller — that is scientific misconduct.

Example: A new drug for hypertension might lower BP, raise BP, or do nothing. Since we are not sure of the direction, we use a two-tailed test. But if we are testing whether a proven anti-hypertensive drug reduces BP (and we have no reason to think it would raise it), a one-tailed test might be justified.

Significance Level: α (Alpha)

Alpha (α) is the cut-off point for deciding whether our evidence is strong enough to be called "statistically significant." It is the threshold of "surprise" — how unlikely must our results be under H₀ before we reject H₀?

The Convention

α = 0.05

This means we are willing to accept a 5% chance of wrongly rejecting H₀ when it is actually true.

Decision Rule
  • If p ≤ 0.05: Reject H₀. The result is statistically significant. The evidence is strong enough to support H₁.
  • If p > 0.05: Fail to reject H₀. The evidence is not strong enough. We do not have sufficient proof to support H₁.

⚠️ Critical Warning: α = 0.05 is a convention, not a law of nature. It is an arbitrary threshold agreed upon by scientists. A result with p = 0.051 is not magically "false," and p = 0.049 is not magically "true." The difference is tiny. Always interpret p-values as a continuum of evidence, not a binary switch.

What a p-value Actually Means

A p-value is the probability of observing results as extreme as (or more extreme than) our sample result, assuming H₀ is true.

In plain nursing language: "If the intervention really did nothing, how lucky (or unlucky) would I have to be to see this result just by chance?"

p-value Statistical Interpretation Nursing Language
p = 0.40 Observed result is quite compatible with H₀. "The difference we saw could easily happen by chance. No strong evidence of a real effect."
p = 0.06 Evidence is suggestive, but not significant at α = 0.05. "There is a hint of an effect, but we cannot be confident. Maybe the sample was too small."
p = 0.03 Observed result is unlikely if H₀ is true. Reject H₀. "This result would be surprising if the training had no effect. We have reasonable evidence it works."
p = 0.001 Observed result is very unlikely if H₀ is true. Strong evidence. "This would almost never happen by chance alone. We are very confident the effect is real."
🚨 What p-values Do NOT Tell You:
  • They do NOT measure the probability that H₀ is true or false.
  • They do NOT measure clinical importance or practical significance.
  • They do NOT show the size of the effect (a tiny effect in a huge sample can have p < 0.001).
  • They do NOT prove that the exposure caused the outcome — only that they are associated.
  • They should ALWAYS be interpreted alongside confidence intervals, effect sizes, and study design.

Example: A drug lowers BP by 0.5 mmHg with p = 0.001. Statistically significant? Yes. Clinically important? No — no doctor would change practice for half a millimetre.

Confidence Intervals and Tests

A confidence interval (CI) gives us a range of plausible values for the true population parameter. When combined with a p-value, it tells us both whether an effect exists and how large it might be.

Result How to Read It
Mean difference = 8.4 points, 95% CI: 7.0 to 9.8 Likely improvement. The CI does not include 0, so the difference is statistically significant. We are 95% confident the true improvement is between 7.0 and 9.8 points.
Risk ratio = 1.20, 95% CI: 0.90 to 1.60 Not statistically clear. The CI includes 1 (the "no effect" value for ratios), meaning the true effect could be no effect at all. The sample size may be too small.
Mean difference = 2.1 kg, 95% CI: −0.5 to 4.7 Not significant. CI includes 0. The intervention might help a little, hurt a little, or do nothing. We cannot tell from this study.
📝 Golden Rule:
For differences (subtraction), "no effect" = 0. If the CI includes 0, the result is not significant.
For ratios (division), "no effect" = 1. If the CI includes 1, the result is not significant.
Mnemonic: "Difference = Zero, Ratio = One."
Type I and Type II Errors

Because we are making decisions based on sample data (not the whole population), we can make mistakes. There are two types of errors we must understand:

Reality Your Decision Error Name Nursing Analogy
H₀ is true (No real effect) You reject H₀ Type I Error
False Positive
α = 0.05
You think a new training works, but it actually doesn't. You waste resources rolling it out.
H₀ is false (Real effect exists) You fail to reject H₀ Type II Error
False Negative
β (beta)
You think a life-saving drug doesn't work, but it actually does. You deny patients effective treatment.
Power of a Study

Power = 1 − β. It is the probability that a study will correctly detect a real effect when one truly exists. Power improves when:

  • Sample size is larger (more data = more precision).
  • Measurement tools are accurate and reliable (less random noise).
  • The true effect size is large (big differences are easier to detect).
  • Variability in the data is low (consistent measurements).
📝 Exam Tip — Remembering Errors:
Mnemonic: "Type I = Innocent person convicted" (false positive — you found an effect that isn't there).
Mnemonic: "Type II = Ignored criminal" (false negative — you missed an effect that is there).
The 6-Step Testing Framework

Do not start by calculating. Start by understanding the data and the question. Follow this framework for every hypothesis test:

  1. Ask a clear question. What exactly are you trying to find out? Make it specific.
  2. State H₀ and H₁. Translate your question into formal null and alternative hypotheses.
  3. Choose α. Usually 0.05. Decide before you see the data.
  4. Select the right test. This depends on your data type and structure (see Section 12).
  5. Calculate the test statistic and p-value. Use software, calculator, or manual formulas.
  6. Interpret in context. Translate the numbers back into nursing or public health language.
💡 Mnemonic — 6 Steps: "Question → Hypotheses → Alpha → Test → Calculate → Interpret"
QHATCI — "Quite Hard Always To Calculate, Initially"
Choosing the Right Test: The Decision Tree

The most common mistake students make is using the wrong test. The right test follows from two questions:

  • What type of outcome variable do you have? (Categorical vs. Continuous)
  • How many groups are you comparing, and are they related or independent?
Outcome Variable Groups / Comparison Suitable Test Nursing Example
Categorical (counts, proportions, yes/no) Two categorical variables (association) Chi-square (χ²) test Is net use associated with malaria positivity?
Continuous (mean scores, weight, BP, time) Same people measured twice (before/after) Paired t-test Did knowledge scores increase after training?
Continuous Two separate, unrelated groups Independent two-sample t-test Is waiting time different in Clinic A vs. Clinic B?
📝 Exam Tip — The Golden Question: Before picking any test, ask: "What type of outcome? Are measurements paired or independent?" Write this on your exam paper if you are unsure. The test choice depends on the data structure, not the topic.
Part 1: Chi-Square (χ²) Test

Used for relationships between two categorical variables.

Purpose
  • Tests whether two categorical variables are associated (linked) or independent (unrelated).
  • Compares observed counts with the counts we would expect if H₀ were true.
  • Useful for tables of frequencies — not means.
Chi-Square Formula

χ² = Σ (Observed − Expected)² / Expected
Sum across all cells in the table.

Worked Example: Net Use and Malaria

Research Question: Is sleeping under an insecticide-treated net (ITN) associated with malaria test result among febrile children?

  • H₀: There is no association between net use and malaria test result.
  • H₁: There is an association between net use and malaria test result.

Step 1: Observed Counts (Contingency Table)

Net Use Malaria Positive Malaria Negative Total
Yes (slept under net) 18 42 60
No (did not sleep under net) 32 28 60
Total 50 70 120

The counts show fewer positive tests among children who slept under a net (18/60 = 30%) compared to those who did not (32/60 = 53.3%). But is this difference likely due to chance?

Step 2: Calculate Expected Counts
Formula: Expected count = (Row Total × Column Total) ÷ Grand Total

Cell Calculation Expected Count
Yes & Positive (60 × 50) ÷ 120 25
Yes & Negative (60 × 70) ÷ 120 35
No & Positive (60 × 50) ÷ 120 25
No & Negative (60 × 70) ÷ 120 35

Expected counts represent what we would see if net use and malaria were completely unrelated (H₀ true). Notice the totals still add up to 60 and 60.

Step 3: Calculate χ²

Cell Observed (O) Expected (E) O − E (O − E)² (O − E)² / E
Yes & Positive 18 25 −7 49 1.96
Yes & Negative 42 35 +7 49 1.40
No & Positive 32 25 +7 49 1.96
No & Negative 28 35 −7 49 1.40
Total χ² 6.72

Step 4: Interpretation

  • χ² = 6.72, degrees of freedom (df) = 1, p ≈ 0.010.
  • Since p = 0.010 ≤ 0.05, we reject H₀.
  • Conclusion: Malaria test result is statistically associated with net use in this sample.
  • Context wording: Children who slept under a net had lower observed malaria positivity (30% vs. 53.3%).
⚠️ Caution in Wording:
❌ Avoid saying: "Net use caused the lower malaria rate."
✅ Correct: "Net use was associated with lower malaria positivity." Causation requires stronger study designs (like RCTs), not just a chi-square test.
Chi-Square Assumptions
  • Data are counts or frequencies, not percentages alone. (You can have percentages, but the test works on the raw counts behind them.)
  • Categories are mutually exclusive: each person belongs in one cell only.
  • Observations are independent — one person's result does not affect another's.
  • Expected counts are generally at least 5 in most cells (some say 80% of cells). If not, use Fisher's exact test.
  • Sample is collected in a way that represents the target population.
Practice Task: Handwashing and Diarrhoea

Scenario: A nurse investigates whether availability of handwashing stations at school is associated with diarrhoea among pupils.

Handwashing Station Diarrhoea: Yes Diarrhoea: No Total
Available 12 48 60
Not Available 24 36 60
Total 36 84 120

Your Tasks:

  • State H₀ and H₁.
  • Calculate one expected count (e.g., Available & Yes).
  • Decide whether chi-square is appropriate (check expected counts).

✅ Solution:

  • H₀: There is no association between handwashing station availability and diarrhoea.
  • H₁: There is an association between handwashing station availability and diarrhoea.
  • Expected count (Available & Yes): (60 × 36) ÷ 120 = 18
  • All expected counts: 18, 42, 18, 42 — all ≥ 5, so chi-square is appropriate.
  • χ² calculation: (12−18)²/18 + (48−42)²/42 + (24−18)²/18 + (36−42)²/42 = 2.0 + 0.86 + 2.0 + 0.86 = 5.72
  • Conclusion: p-value ≈ 0.017 (df=1). Since p < 0.05, reject H₀. Handwashing station availability is associated with diarrhoea risk.
Part 2: Paired t-Test

Used when the same person is measured twice (before/after).

Purpose
  • Compares two related means.
  • Common design: before-and-after measurement on the same participants.
  • The test analyses the difference within each pair, not the raw scores.
  • It answers: "Is the mean difference significantly different from zero?"
Paired t-Test Formula

t = mean difference ÷ (SD of differences / √n)
Or simply: t = mean difference ÷ standard error of the difference

When Data Are Paired

Paired data means the two measurements are linked or dependent. They are not independent observations.

Paired Situation Why It Is Paired
Knowledge score before and after training Same learner measured twice. The "after" score depends on the "before" score.
Blood pressure before and after treatment Same patient. Their baseline BP influences their post-treatment BP.
Left eye and right eye of same person Two measurements linked to one individual's genetics and environment.
Mother's weight before and after nutrition counselling Same mother. Her starting weight affects how much she can gain or lose.
Pain score before and after analgesia Same patient. Their pain tolerance and condition are consistent across both measurements.
🚨 Critical Warning: Treating paired data as independent wastes information and can give wrong results. If you ignore the pairing, you treat "Mother A before" and "Mother A after" as two different people, which inflates variability and may hide a real effect.
Worked Example: Knowledge Scores Before and After Health Education

Scenario: A nurse educator measures knowledge scores about infant nutrition before and after a 1-hour education session among 8 mothers.

The Raw Data

Mother 1 2 3 4 5 6 7 8
Before 45 50 55 60 52 48 58 62
After 55 58 63 70 59 56 64 72
Difference (After − Before) 10 8 8 10 7 8 6 10

Step-by-Step Calculation

  1. Calculate the mean difference:
    Mean difference = (10 + 8 + 8 + 10 + 7 + 8 + 6 + 10) ÷ 8 = 67 ÷ 8 = 8.38 points
  2. Calculate the standard deviation (SD) of the differences:
    Using the formula for sample SD, SD of differences ≈ 1.51
  3. Calculate the standard error (SE) of the mean difference:
    SE = SD / √n = 1.51 / √8 = 1.51 / 2.828 = 0.534
  4. Calculate the t-statistic:
    t = mean difference / SE = 8.38 / 0.534 = 15.73
  5. Determine degrees of freedom and p-value:
    df = n − 1 = 8 − 1 = 7. With t = 15.73 and df = 7, p < 0.001.

Interpretation

  • Since p < 0.001 ≤ 0.05, we reject H₀.
  • The mean increase in knowledge score (8.38 points) is statistically significant.
  • Direction: Scores increased after health education. (Always state the direction!)
  • Clinical context: The education session was effective at improving mothers' nutrition knowledge in this sample.
Paired t-Test Assumptions
  • The outcome variable is continuous or approximately continuous (e.g., scores, weight, BP).
  • Pairs are genuinely linked: same person, matched pair, or before/after design.
  • The differences are approximately normally distributed. (With n ≥ 30, this is less critical due to the Central Limit Theorem.)
  • No extreme outliers in the differences. One wildly different pair can distort the mean.
  • The measurement method is consistent before and after. (Don't change the scoring rubric mid-study!)
How to Report a Paired t-Test
  • ✅ Good Report: "Mean knowledge score increased from 53.8 before training to 62.1 after training. The mean paired increase was 8.38 points and was statistically significant, t(7) = 15.73, p < 0.001."
  • ❌ Poor Report: "The training worked because p was less than 0.05."
    Why it's poor: No means, no direction, no context. The reader has no idea how big the improvement was or whether it matters clinically.
Practice: Pain Scores After Counselling

Scenario: Eight patients have pain scores (0–10 scale) before and after counselling. The differences (After − Before) are: −2, −1, −3, 0, −2, −1, −4, −2.

Answer these:

  • What is the outcome variable? Pain score (continuous, 0–10 scale).
  • Why is a paired t-test appropriate? Same patient measured twice — before and after counselling.
  • What does a negative difference mean? Pain score decreased after counselling. (Negative = After is lower than Before = improvement!)
  • What would H₀ say? The mean difference in pain scores = 0 (counselling has no effect on pain).
  • What would H₁ say? The mean difference ≠ 0 (counselling changes pain scores).
  • Quick calculation: Mean difference = −1.875, SD ≈ 1.25, SE ≈ 0.44, t ≈ −4.26, p ≈ 0.004. Reject H₀. Counselling significantly reduced pain scores.
Part 3: Independent Two-Sample t-Test

Used when comparing means from two separate, unrelated groups.

Purpose
  • Compares the mean of a continuous outcome in two independent groups.
  • Independent means one participant belongs to only one group — there is no natural pairing.
  • Example: Mean waiting time in Clinic A versus Clinic B.
  • It asks: "Is the difference between group means larger than what we would expect by chance alone?"
Independent vs. Paired — The Critical Distinction
Question Data Structure Correct Test
Before and after training in the same learners Same learners measured twice Paired t-test
Clinic A patients vs. Clinic B patients Different patients in two clinics Independent t-test
Men vs. women Hb level Different people in two groups Independent t-test
Vaccinated children vs. unvaccinated children (immunity levels) Two separate groups of children Independent t-test
Mother's weight before and after nutrition program Same mothers measured twice Paired t-test
📝 Exam Tip: The main decision is not the topic — it is the data structure. Ask: "Are these the same people measured twice, or two completely different groups?"
Worked Example: Waiting Time in Two Clinics

Scenario: A district health officer wants to know if mean waiting time differs between Clinic A (new digital system) and Clinic B (paper-based system).

The Raw Data (waiting time in minutes)

Clinic Wait Times
Clinic A (Digital) 40 45 50 42 48 46 44 47
Clinic B (Paper) 55 58 60 52 57 61 54 59

Descriptive statistics:

  • Clinic A mean = 45.25 minutes
  • Clinic B mean = 57.00 minutes
  • Mean difference = 45.25 − 57.00 = −11.75 minutes

Inferential statistics:

  • t ≈ −7.39
  • df ≈ 14 (n₁ + n₂ − 2 = 8 + 8 − 2)
  • p < 0.001

Decision: Since p < 0.05, we reject H₀. Mean waiting time differs significantly between the two clinics.

Context: Clinic B (paper-based) had significantly longer waiting times than Clinic A (digital). The digital system appears to improve patient flow.

Independent t-Test Assumptions
  • The outcome variable is continuous or approximately continuous.
  • The two groups are independent — no person appears in both groups.
  • The outcome is approximately normally distributed within each group. (Less critical with large samples, n > 30 per group.)
  • No extreme outliers that would distort the mean.
  • Variability (variance) in the two groups is reasonably similar. If not, use Welch's t-test (a modified version that handles unequal variances).
💡 Note on Welch's t-test: Most statistical software automatically uses Welch's correction when variances are unequal. For your exams, simply knowing that unequal variances require a modification is sufficient. If the question does not mention unequal variances, assume the standard independent t-test applies.
How to Report an Independent t-Test

✅ Good Report: "Mean waiting time was 45.25 minutes in Clinic A and 57.00 minutes in Clinic B. The mean difference was −11.75 minutes and was statistically significant, t(14) = −7.39, p < 0.001."

Key addition: Always name the groups and the direction of the difference. Don't just say "significant" — say which group was higher/lower.

Choosing Tests: Mini Scenarios
Scenario Correct Test Why?
Sex of child vs. immunisation status (yes/no) Chi-square Two categorical variables.
Weight before and after nutrition counselling in same mothers Paired t-test Same people, before/after.
Mean Hb in two separate wards Independent t-test Two separate groups, continuous outcome.
Facility ownership (govt/private) vs. stock-out status (yes/no) Chi-square Two categorical variables.
Pain score before and after analgesia in same patients Paired t-test Same patients measured twice.
Mean BMI between urban and rural adolescents Independent t-test Two separate groups, continuous outcome.
Smoking status (yes/no) vs. TB status (yes/no) Chi-square Two categorical variables.
Knowledge scores of nurses trained online vs. nurses trained in-person Independent t-test Two different groups of nurses, continuous score.
Common Mistakes to Avoid
Mistake Why It Is Wrong What To Do Instead
Using percentages instead of counts for chi-square The formula requires raw counts. Percentages lose the sample size information. Always construct the table with actual frequencies (n).
Using independent t-test for before-and-after data You ignore the natural pairing, which increases error and reduces power. Use a paired t-test when the same people are measured twice.
Reporting only p-values without means or proportions A p-value tells you significance but not effect size or direction. Always report means/SDs or proportions alongside p-values.
Claiming causation from a simple association Association does not prove causation. Confounding variables may explain the link. Use cautious language: "associated with," "linked to," not "caused."
Ignoring assumptions and outliers Outliers can make means misleading. Non-normal data can invalidate t-tests. Check your data first. Plot it. Look for outliers. Consider transformations or non-parametric tests.
Choosing α after seeing the p-value This is data dredging / p-hacking. You are moving the goalposts. Set α = 0.05 before any analysis. Stick to it.
Saying "accept H₀" when p > 0.05 You never prove the null hypothesis. You only fail to find evidence against it. Say "fail to reject H₀" or "no statistically significant difference was found."
Practical Work Example

Scenario: A nurse educator measures infection prevention knowledge before and after a 2-hour training session among 20 student nurses. Scores range from 0 to 100.

Your Tasks & Solutions:

  • Identify the outcome variable. Infection prevention knowledge score (continuous, 0–100).
  • State H₀ and H₁.
    • H₀: Mean knowledge score before training = Mean knowledge score after training.
    • H₁: Mean knowledge score before training ≠ Mean knowledge score after training.
  • Choose the most appropriate test. Paired t-test — same 20 nurses measured before and after.
  • Explain why the test is appropriate. The data are continuous, and each participant serves as their own control. The two measurements are dependent (paired).
  • Write one sentence showing how you would report a significant result.
    "Mean infection prevention knowledge increased from 48.2 before training to 71.5 after training; the mean paired difference of 23.3 points was statistically significant, t(19) = 8.45, p < 0.001."
Group Exercise: Test Selection

Complete the table. For each question, identify the data types and choose the correct test. Then justify your choice in one sentence.

Question Data Type Test Justification
Is nutrition status associated with sex? Categorical + Categorical Chi-square Both variables are categories; we are testing association.
Did average temperature fall after treatment? Continuous, paired Paired t-test Same patients measured before and after treatment.
Are mean ages different between two wards? Continuous, independent Independent t-test Two separate groups of patients; age is continuous.
Is HIV testing uptake associated with age group? Categorical + Categorical Chi-square Both uptake (yes/no) and age group (e.g., 15–24, 25–34) are categories.
Is mean blood pressure lower after a 6-month exercise program? Continuous, paired Paired t-test Same hypertensive patients measured at baseline and 6 months.
Are recovery rates different between Hospital X and Hospital Y? Categorical + Categorical Chi-square Recovery (recovered/not recovered) and hospital are both categorical.
Quick Quiz (Self-Check)

Cover the answers and test yourself. Use one clear sentence per answer.

  • What does H₀ usually state?
    Answer: H₀ states there is no difference, no association, or no effect — the "status quo" position.
  • What does p = 0.03 mean at α = 0.05?
    Answer: The probability of seeing this result (or more extreme) if H₀ were true is 3%. Since 0.03 < 0.05, we reject H₀ and conclude a statistically significant result.
  • When do we use a chi-square test?
    Answer: When we want to test whether two categorical variables are associated, using frequency counts in a contingency table.
  • What makes a t-test "paired"?
    Answer: The same individual is measured twice (before/after), or two measurements are naturally linked (e.g., left eye and right eye).
  • Why should clinical meaning be reported alongside statistical significance?
    Answer: A result can be statistically significant but clinically trivial (tiny effect in a huge sample), or clinically important but not significant (small sample). Both dimensions matter for patient care.
  • What is a Type I error?
    Answer: Rejecting H₀ when it is actually true — a false positive. We think an effect exists when it does not.
  • What is a Type II error?
    Answer: Failing to reject H₀ when it is actually false — a false negative. We miss a real effect.
  • What does a 95% confidence interval that includes 0 tell us about a mean difference?
    Answer: The difference is not statistically significant at α = 0.05. The true effect could be zero.
Key Take-Away & Decision Summary

The right statistical test follows the research question and the data structure. Do not memorise tests by topic — memorise them by variable type and group relationship.

The Decision Tree
  • Categorical association (e.g., net use vs. malaria status) ➔ Chi-square test
  • Same people before and after (e.g., knowledge scores pre/post training) ➔ Paired t-test
  • Two separate groups with continuous outcomes (e.g., waiting time Clinic A vs. Clinic B) ➔ Independent t-test
📝 Final Exam Rules:
  • Always interpret in health context, not only by p-value.
  • Always state the direction of the effect (which group was higher/lower).
  • Never claim causation from a single test — say "associated with."
  • Always check assumptions before choosing a test.
  • Report means, SDs, and CIs alongside p-values.
References
  • Polit, D. F., & Beck, C. T. (2020). Nursing Research: Generating and Assessing Evidence for Nursing Practice (11th ed.). Wolters Kluwer.
  • Grove, S. K., & Gray, J. R. (2018). Understanding Nursing Research: Building an Evidence-Based Practice (7th ed.). Saunders.
  • Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics (5th ed.). SAGE Publications.
  • Daniel, W. W., & Cross, C. L. (2018). Biostatistics: A Foundation for Analysis in the Health Sciences (11th ed.). Wiley.

Quick Quiz

Probability, Hypothesis Testing & Statistical Tests Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Probability, Hypothesis Testing & Statistical Tests Read More »

Summarising, Presenting, Communicating Findings in research

Summarising, Presenting and Communicating Findings

Summarising, Presenting and Communicating Findings
Learning Outcomes

By the end of this session, you should be able to:

  • Explain population distributions and describe patterns in data.
  • Calculate and interpret mean, median, and mode and know when to use each.
  • Calculate and interpret range, variance, and standard deviation.
  • Use ratios, proportions, and rates correctly with the right denominator.
  • Present findings clearly using tables, graphs, and concise action-oriented messages.

🎯 Main Skill: Accurate summary + clear interpretation. Numbers alone mean nothing. A nurse who can collect data, summarise it correctly, and explain what it means for patient care is a powerful public health tool.

The Data Story Cycle

Data is not just numbers it is a story about people's health. Your job as a nurse is to read that story and tell it clearly so that action can be taken. The cycle has four stages:

  • RAW DATA: Messy numbers from registers, surveys, or observations
  • SUMMARISE: Mean, median, rates, tables, graphs
  • PRESENT: Clear tables, graphs, and messages
  • COMMUNICATE: Tell the story so decision-makers act
  • DECISION: Policy change, resource shift, intervention
The Three Questions You Must Answer:
  • What is the data saying? (Describe the numbers accurately.)
  • So what does it mean? (Interpret the pattern for patient care.)
  • What should be done next? (Recommend action based on evidence.)

💡 Mnemonic The Data Story: "What? So What? Now What?" = WSWNW. Think: "We See What Nurses Want." Every presentation must answer all three.

Working Dataset for Today

We will use a single, simple dataset throughout this lesson so you can see how each statistic tells a different part of the same story.

📋 Dataset: Waiting time before consultation, in minutes

10, 12, 12, 15, 18, 20, 25

Context: Small outpatient clinic • 7 sampled clients • Measure: waiting time from arrival to consultation start.

We will use this to learn: Central tendency, variation, graphing, and interpretation.

Population Distributions

A distribution is the pattern of values in a population or sample. It shows you where the data clusters, where it spreads out, and whether there are unusual values.

Why distribution matters:
  • It shows common values (where most patients fall) and unusual values (outliers that need attention).
  • It helps us choose the best summary statistic mean, median, or mode.
  • It supports fair comparison between groups (e.g., Clinic A vs. Clinic B).
  • It reveals skewness whether the data is balanced or pulled to one side.
Frequency Distribution

A frequency distribution groups data into categories and counts how many observations fall into each group. It is often the first step before drawing a graph.

Example: Age distribution of patients at a health centre

Age group (years) Number of patients Percentage (%)
0–4 8 10%
5–14 12 15%
15–24 20 25%
25–44 32 40%
45+ 18 23%

Interpretation: The 25–44 age group carries the largest patient load (40%). The clinic should ensure this group has adequate staffing and supplies. The 0–4 group is smallest but these are often the sickest patients, so do not ignore them.

Histogram: For Continuous Data

A histogram is a graph of a frequency distribution for continuous data (data that can take any value within a range, like weight, height, temperature, or waiting time).

  • Bars touch each other because the intervals are continuous (unlike a bar chart for categories).
  • The height of each bar shows the frequency (how many observations).
  • The shape tells us how the values are distributed symmetric, skewed, or bimodal.
  • Examples of continuous data in nursing: Birth weight, haemoglobin level, body temperature, blood pressure, waiting time, length of hospital stay.
Common Distribution Shapes
Shape What It Looks Like Nursing Example
Symmetric (Normal / Bell-shaped) The left side mirrors the right side. Most values cluster in the centre. Tails are equal on both sides. Birth weights of full-term babies; blood pressure in a healthy population. The mean, median, and mode are all at the centre.
Right-skewed (Positive skew) A long tail stretches to the right (higher values). Most values cluster on the left. Mean > Median. Hospital length of stay most patients stay 2-3 days, but a few stay 30 days. Income distribution most people earn little, a few earn a lot.
Left-skewed (Negative skew) A long tail stretches to the left (lower values). Most values cluster on the right. Mean < Median. Age at retirement most people retire at 60-65, but a few retire early at 40. Exam scores where most students score high and a few score very low.

📝 Exam Tip: Skewness means the data has a longer tail on one side. Right-skewed = tail on the right = mean pulled UP by high outliers = use median instead of mean. Left-skewed = tail on the left = mean pulled DOWN by low outliers = use median instead of mean. Symmetric = use mean.

Why Distribution Shape Matters in Real Clinics
Clinic Distribution Best Statistic
Clinic A Most patients wait 25–35 minutes. Few extreme values. Symmetric shape. Mean is reliable. It accurately represents the typical experience.
Clinic B Many patients wait 5 minutes. Some wait 90 minutes. Right-skewed. Median tells a fairer story. The mean would be pulled up by the few very long waits and would overstate the typical experience.

Rule: Always inspect the shape before summarising the data. The wrong statistic can mislead decision-makers.

Work Example: Read a Distribution

Antenatal clinic waiting times (grouped data)

Waiting time (minutes) Number of mothers
0–15 3
16–30 9
31–45 5
46–60 2
61+ 1

Interpretation:

  • Most mothers (9 out of 20) waited 16–30 minutes this is the modal class.
  • Very long waits (over 45 minutes) were uncommon only 3 mothers.
  • A few high values (the 61+ minute wait) may pull the mean upward. The median might be a fairer summary.

Action message: Service flow is generally moderate, but long waits still require attention. Investigate what causes the 61+ minute delay is it a specific time of day, a specific nurse, or a bottleneck in the lab?

Measures of Central Tendency

Central tendency tells us where the "centre" of the data is. It answers: "What is the typical value?" There are three main measures each tells a different story.

Measure Definition Best Used When... Nursing Example
Mean The arithmetic average. Sum all values ÷ number of values. Data are numerical, symmetric, and have no extreme outliers. Average birth weight of 50 newborns in a month.
Median The middle value when data are ordered from smallest to largest. Data are skewed or have outliers. Gives the "typical" experience. Median hospital stay most patients stay 3 days, but one stayed 45 days.
Mode The most frequent value the one that appears most often. Data are categorical or you want the most common value. A dataset can have multiple modes or no mode. Most common diagnosis at OPD this week (e.g., "malaria" appeared 45 times).

💡 Mnemonic When to Use What: "Mean for Middle of Symmetric data. Median for Middle of Skewed data. Mode for Most Frequent." Another: "Mean = Mathematical average. Median = Middle value. Mode = Most common."

Worked Example: Mean

Dataset: 10, 12, 12, 15, 18, 20, 25

  • Step 1: Add all values
    10 + 12 + 12 + 15 + 18 + 20 + 25 = 112
  • Step 2: Count the number of values (n)
    n = 7 clients
  • Step 3: Calculate the mean
    Mean = 112 ÷ 7 = 16 minutes

Interpretation: The average waiting time was 16 minutes. If you told the clinic manager "the mean waiting time is 16 minutes," they would expect most patients to wait around that long.

Worked Example: Median

Dataset (already ordered): 10, 12, 12, 15, 18, 20, 25

  • Step 1: Ensure data is ordered from smallest to largest. (Already done.)
  • Step 2: Find the middle position.
    For an odd number of values: position = (n + 1) ÷ 2 = (7 + 1) ÷ 2 = 4th value.
  • Step 3: Identify the median.
    10, 12, 12, 15, 18, 20, 25
    Median = 15 minutes

Interpretation: Half the clients waited 15 minutes or less. Half waited 15 minutes or more. The median is not affected by the person who waited 25 minutes it gives the "typical" experience.

What if n is even? If you have 8 values, the median is the average of the 4th and 5th values. Example: 10, 12, 12, 15, 18, 20, 25, 30. Median = (15 + 18) ÷ 2 = 16.5 minutes.

Worked Example: Mode

Dataset: 10, 12, 12, 15, 18, 20, 25

  • Step 1: Count how many times each value appears.
    10 → 1 time
    12 → 2 times
    15 → 1 time
    18 → 1 time
    20 → 1 time
    25 → 1 time
  • Mode = 12 minutes (appears most frequently twice).

Interpretation: The most common waiting time was 12 minutes. Even though the mean is 16, more people actually waited 12 minutes than any other single time.

Note: A dataset can have no mode (all values appear once) or multiple modes (two values appear equally often called bimodal). For categorical data (diagnoses, blood types), the mode is often the only useful measure of central tendency.

Mean, Median, and Mode Together

Mean = 16 | Median = 15 | Mode = 12

What does this pattern tell us?

  • The mean (16) is slightly higher than the median (15).
  • This suggests there are some higher waiting times pulling the mean upward a slight right skew.
  • The mode (12) is lower than both, showing that the most common experience is actually faster than the average.

Which statistic should you report? It depends on the question:

  • If the manager asks "What is the average wait?" → Report the mean (16 minutes).
  • If the manager asks "How long does the typical patient wait?" → Report the median (15 minutes).
  • If the manager asks "What is the most common wait time?" → Report the mode (12 minutes).

Best Practice: Report the statistic that best answers the question being asked. Do not just calculate all three and dump them on the reader. Choose the one that tells the clearest story.

Class Practice: Central Tendency

Dataset: Number of diarrhoea cases reported by 8 villages in one month
3, 4, 4, 6, 7, 9, 10, 13

Task: Work in pairs for 5 minutes. Calculate the mean, median, and mode. Show every step clearly. Write one sentence interpreting your answer.

Answer guide:

  • Mean = (3+4+4+6+7+9+10+13) ÷ 8 = 56 ÷ 8 = 7 cases
  • Median = average of 4th and 5th values = (6 + 7) ÷ 2 = 6.5 cases
  • Mode = 4 cases (appears twice)
  • Interpretation: The average village reported 7 cases, but the typical village reported 6-7 cases, and the most common report was 4 cases. The village with 13 cases is an outlier pulling the mean up.
Measures of Variation (Spread)

Central tendency tells us the centre. Variation tells us how far values differ from the centre. Two clinics can have the same mean waiting time but very different patient experiences.

🏥 Why variation matters in nursing: High variation means unequal experience across patients. Some get excellent care; others get terrible care. High variation in waiting times means the system is unstable. High variation in blood pressure readings means the measurement technique is inconsistent. Your job is to reduce harmful variation.

Range

The range is the simplest measure of spread. It tells you the gap between the highest and lowest values.

Range = Maximum value − Minimum value

Dataset: 10, 12, 12, 15, 18, 20, 25

Range = 25 − 10 = 15 minutes

Interpretation: Waiting times differed by 15 minutes between the shortest and longest wait. The range is easy to calculate but is sensitive to outliers one extreme value can make the range misleadingly large.

Variance

Variance measures how far each value is from the mean, on average. It uses every value in the dataset, not just the extremes.

Sample Variance = Σ(x − x̄)² ÷ (n − 1)

Where: x = each value, x̄ = mean, n = number of values, Σ = sum of

Key idea: Large variance = values are widely spread. Small variance = values are clustered close to the mean.

Why we divide by (n − 1) for sample variance: Dividing by (n − 1) instead of n gives a slightly larger variance, which corrects for the fact that we are using a sample (a subset) rather than the entire population. This makes our estimate more accurate. In exams, if you are told it is a sample, divide by (n − 1). If it is the entire population, divide by n.

⚠️ Important: Variance is in squared units (minutes squared, kg squared). This makes it hard to interpret directly. That is why we use standard deviation it brings the units back to the original scale.

Standard Deviation (SD)

Standard deviation is the square root of the variance. It tells us, on average, how far each value is from the mean. It uses the same units as the original data.

Standard Deviation = √Variance

SD Size What It Means Nursing Example
Small SD Values are close together. The process is consistent and predictable. Waiting times at a well-run clinic: most patients wait 25-30 minutes. The system is stable.
Large SD Values are far apart. Experience is unequal and unpredictable. Waiting times at a chaotic clinic: some wait 5 minutes, some wait 90 minutes. The system needs fixing.
Worked Example: Standard Deviation

Dataset: 10, 12, 12, 15, 18, 20, 25
We already know: Mean (x̄) = 16

  • Step 1: Calculate each deviation from the mean (x − x̄)
Value (x) Mean (x̄) Deviation (x − x̄) Squared deviation (x − x̄)²
1016−636
1216−416
1216−416
1516−11
1816+24
2016+416
2516+981
  • Step 2: Sum the squared deviations
    Σ(x − x̄)² = 36 + 16 + 16 + 1 + 4 + 16 + 81 = 170
  • Step 3: Calculate sample variance
    Variance = 170 ÷ (7 − 1) = 170 ÷ 6 = 28.33 minutes²
  • Step 4: Calculate standard deviation
    SD = √28.33 = 5.32 minutes

Interpretation: Waiting times commonly differ from the mean by about 5 minutes. Most patients (about 68%, if the distribution were normal) wait between 16 − 5 = 11 minutes and 16 + 5 = 21 minutes.

Interpreting Spread in Nursing Data
Clinic A Clinic B
Mean waiting time = 30 minutes
SD = 5 minutes
Mean waiting time = 30 minutes
SD = 25 minutes
Most patients experience similar waits (25–35 minutes). The system is predictable. Some patients wait 5 minutes, others wait 90 minutes. The system is chaotic and unfair.
Same mean, different experience. Same mean, different experience.

Lesson: Never report the mean without some measure of spread. The mean hides inequality. The SD reveals it.

Class Practice: Variation

Dataset: 3, 4, 4, 6, 7, 9, 10, 13 (diarrhoea cases in 8 villages)

Task: Calculate the range. Identify whether the data appears skewed. Explain what spread means in this outbreak situation.

Answer guide:

  • Range = 13 − 3 = 10 cases
  • Skewness: Values rise gradually with a high value at 13. The mean (7) is slightly higher than the median (6.5), suggesting a slight right skew.
  • What spread means: The outbreak is not evenly distributed. One village has 13 cases more than 4x the lowest village. This village needs targeted investigation (water source, sanitation, vaccination status). Low spread would mean all villages are equally affected, suggesting a widespread environmental factor.
Ratios, Proportions, and Rates

These three measures are the bread and butter of epidemiology. They look similar but mean very different things. Using the wrong one can mislead decision-makers and cost lives.

Measure Definition Formula Nursing Example
Ratio Compares two independent groups. The two numbers do not have to be part of the same whole. A : B
(simplify by dividing both by their greatest common divisor)
Male nurses to female nurses = 15 : 45 = 1 : 3
Proportion A part of a whole. The numerator is always included in the denominator. Always between 0 and 1 (or 0% and 100%). Part ÷ Whole
× 100 for percentage
Vaccination coverage = 120 vaccinated ÷ 150 eligible = 80%
Rate A proportion that includes time. It measures how fast something is happening in a population at risk. (Events ÷ Population at risk) × multiplier
per time period
Malaria incidence = 25 new cases ÷ 500 children × 1,000 = 50 per 1,000 children per month

📝 Exam Tip The Golden Rule: The denominator determines the correct interpretation. Always ask: "Does the numerator come FROM the denominator?" If yes → proportion. If no → ratio. If time is involved → rate.

Worked Example: Ratio

Scenario: At an antenatal clinic, 60 clients are female (mothers) and 40 are male (partners attending together).

Ratio = 60 : 40
Simplify by dividing both by 20: 3 : 2

Interpretation: For every 3 female clients, there are 2 male clients. A ratio does NOT mean the males are "part of" the females. They are two separate groups being compared.

Worked Example: Proportion

Scenario: In a survey of 100 households, 60 own an insecticide-treated net (ITN).

Proportion = 60 ÷ 100 = 0.60 = 60%

Interpretation: Six in every ten households own an insecticide-treated net. The numerator (60 households with nets) is INCLUDED in the denominator (100 households surveyed). A proportion must always be between 0 and 1 (or 0% and 100%). If your calculation gives 120%, you have made a mistake.

Worked Example: Rate

Scenario: In June 2026, a district recorded 25 new malaria cases among 500 children under 5.

Rate = (25 ÷ 500) × 1,000 = 50 per 1,000 children per month

Interpretation: During June, 50 new malaria cases occurred for every 1,000 children under 5. Rates MUST include a time period. Without "per month," this is just a proportion. The multiplier (1,000 or 100,000) makes the number easier to read and compare across populations of different sizes.

Common multipliers: Use × 100 for percentages (vaccination coverage). Use × 1,000 for common events (malaria rates). Use × 100,000 for rare events (maternal mortality ratio, cancer incidence).

Incidence vs. Prevalence

Note: Incidence vs. Prevalence These are two of the most important rates in epidemiology. Confusing them is a common exam mistake. Measure What It Measures Numerator Nursing Example Incidence New cases during a defined period. Measures risk or speed of occurrence.

Every day in nursing, you will be asked: "How many people are getting sick?" and "How many people are sick right now?" These sound similar, but they measure completely different things. Confusing them leads to wrong decisions, wrong budgets, and wrong priorities.

Incidence — The Speed of New Disease

Definition: Incidence measures the number of new cases of a disease that occur in a defined population during a specified period of time. It tells you how fast the disease is spreading the risk or speed of occurrence.

Incidence Formula
Incidence = (New Cases ÷ Population at Risk) × Multiplier
Usually expressed per 1,000 or per 100,000 population over a specific time period (e.g., per year, per month).

Key features of incidence:

  • It counts only new cases people who were healthy at the start of the period and became sick during it.
  • It requires a time period you cannot measure incidence "today." You measure it "this month," "this year," or "during the outbreak."
  • It needs a population at risk people who could actually get the disease. Someone who already had malaria last week and is still recovering is NOT at risk of a new malaria episode (unless reinfection is possible).
  • It measures risk the probability that a healthy person will develop the disease.

📝 Exam Tip Incidence: When you see "new cases" and "over a period of time," think INCIDENCE. Example: "There were 45 new cases of malaria in Village A during July 2026." That is incidence.

Prevalence — The Burden of Existing Disease

Definition: Prevalence measures the total number of existing cases (both new and old) in a population at a specific point in time or over a period. It tells you the burden of disease how many people are living with the condition right now.

Prevalence Formula
Prevalence = (All Existing Cases ÷ Total Population) × Multiplier
Usually expressed as a percentage or per 1,000 population. Can be point prevalence (one moment) or period prevalence (over a time span).

Key features of prevalence:

  • It counts all existing cases both people who got sick recently AND people who have been sick for a long time.
  • It is measured at a point in time (point prevalence) or over a period (period prevalence).
  • It uses the total population as the denominator not just those at risk.
  • It measures burden the total load of disease on the health system and community.

📝 Exam Tip Prevalence: When you see "existing cases" and "right now" or "today," think PREVALENCE. Example: "A survey found that 18 of 80 adults had high blood pressure on the day of screening." That is point prevalence.

Side-by-Side Comparison
Feature Incidence Prevalence
What it counts Only new cases during a period. All existing cases (new + old) at a point or period.
Time Requires a period (month, year). Measured at a point or over a period.
Denominator Population at risk (healthy people who could get it). Total population (everyone, sick or healthy).
What it tells you Risk / speed how fast is the disease occurring? Burden how many people are living with it?
Use case Detecting outbreaks, evaluating prevention programs, measuring vaccine effectiveness. Planning health services, estimating drug needs, measuring chronic disease load.
Example "15 new malaria cases in July." "120 people in the village currently have hypertension."

💡 Mnemonic Incidence vs. Prevalence:
Incidence = Incoming = In = New = needs a time period.
Prevalence = Present = Picture = snapshot = point in time.
Think: "Incidence is like a video (over time). Prevalence is like a photograph (one moment)."

Relationship Between Incidence and Prevalence

Prevalence is like a bathtub:

  • Incidence is the water flowing IN through the tap (new cases entering the pool).
  • Deaths and recoveries are the water flowing OUT through the drain (cases leaving the pool).
  • Prevalence is the total water in the tub at any moment.

Key insight: A disease with high incidence but short duration (like influenza) may have low prevalence people get it and recover quickly. A disease with low incidence but long duration (like diabetes or HIV) may have high prevalence people live with it for years.

Practical implication: If you want to know whether a prevention program is working, measure incidence (are fewer people getting sick?). If you want to know how many clinic appointments or drug doses you need, measure prevalence (how many people need care right now?).

🚨 Exam Trap: "A clinic sees 200 malaria patients this month." Is this incidence or prevalence? It depends. If these are 200 new cases, it is incidence. If these include people coming back for follow-up (old cases), it is a mix. In exams, always ask: "Are these new cases or all cases?"

Scenario: Malaria in Two Villages — Incidence and Prevalence

🩺 The Situation:

  • Village A: Population 500. In July, 50 people developed malaria for the first time this year. At the end of July, 30 people were still sick (20 had recovered).
  • Village B: Population 500. In July, 10 people developed malaria for the first time this year. At the end of July, 80 people were still sick (malaria is chronic in this area due to drug resistance).

Calculate and interpret:

Measure Village A Village B
Incidence (July) 50 ÷ 500 = 10% (or 100 per 1,000). High risk an outbreak. 10 ÷ 500 = 2% (or 20 per 1,000). Lower risk.
Point Prevalence (end July) 30 ÷ 500 = 6%. Moderate burden. 80 ÷ 500 = 16%. High burden many chronic cases.
Interpretation Village A has an acute outbreak act fast with nets, testing, treatment. Village B has a chronic burden needs long-term drug supply, adherence support, and possibly new treatment protocols.

💡 Key Lesson: Village A has higher incidence but lower prevalence (people recover fast). Village B has lower incidence but higher prevalence (people stay sick longer). The action needed is completely different. Always know which measure you are looking at.

Choosing the Correct Denominator

The denominator is the bottom number in any rate, proportion, or ratio. Choose the wrong denominator, and your result is meaningless or dangerously misleading. The golden rule is: the denominator must represent everyone who could have had the event.

Question / Measure Correct Denominator Type of Measure Why This Denominator?
Vaccine coverage All eligible children (e.g., children aged 12-23 months for measles vaccine). Proportion Only eligible children can receive the vaccine. Using all children (including newborns) would underestimate coverage.
Case fatality rate (CFR) All people with the disease (total cases, not the whole population). Proportion Only people who have the disease can die from it. Healthy people cannot die from malaria, so they do not belong in the denominator.
Maternal mortality ratio (MMR) Live births (not total population or total women). Ratio Only women who give birth are at risk of maternal death. The denominator measures the "opportunity" for the event.
Outpatient attendance rate Population at risk + time (e.g., catchment population per year). Rate The catchment population represents everyone who could potentially attend. Time is included because attendance accumulates over the year.
Attack rate All people exposed to the risk (e.g., everyone who ate the contaminated food). Proportion Only exposed people can get the disease. Someone who did not eat the food cannot get food poisoning from it.
Bed occupancy rate Total bed-days available during the period. Proportion Only available beds can be occupied. Beds that are broken or closed should not be counted.

📝 Exam Tip Denominator Check: Before calculating any measure, ask yourself: "Does my denominator include everyone who could have experienced this event?" If it includes people who could NOT have the event, your result will be too low (underestimation). If it excludes people who COULD have the event, your result will be too high (overestimation).

Ratios, Proportions, and Rates What's the Difference?
Measure Definition Denominator Example
Ratio A value obtained by dividing one quantity by another. The numerator is NOT part of the denominator. Any quantity numerator and denominator are unrelated. Doctors-to-nurses ratio = 1:4. Maternal deaths per 100,000 live births.
Proportion A ratio where the numerator is INCLUDED in the denominator. Always ranges from 0 to 1 (or 0% to 100%). The total group that includes the numerator. Vaccine coverage = 80%. Case fatality rate = 5%. Proportion of males = 45%.
Rate A measure of frequency that includes time in the denominator. Measures speed or velocity of events. Population at risk + time period. Incidence rate = 50 cases per 1,000 population per year. Birth rate = 35 per 1,000 per year.

⚠️ Common Mistake: Students often call everything a "rate." But vaccine coverage is a proportion (the vaccinated children are part of all eligible children), not a rate (it does not include time). Case fatality is a proportion (deaths are part of all cases), not a rate. Be precise with your terminology examiners notice.

Summarising Data — Central Tendency and Spread

Raw data is messy. A list of 100 patient ages tells you nothing until you summarise it. There are two things you need to know about any dataset: where is the centre? and how spread out are the values?

Measures of Central Tendency — Where Is the Centre?
Measure What It Is When to Use When NOT to Use
Mean (Average) Sum of all values ÷ Number of values. The "balancing point" of the data. When data is roughly symmetrical and there are no extreme outliers. Good for height, weight, normal lab values. When there are extreme outliers (e.g., one billionaire in a village of poor farmers). The mean will be misleadingly high.
Median The middle value when all values are arranged in order. Half the data is below, half is above. When data is skewed (pulled to one side) or has outliers. Good for income, hospital stay duration, waiting times. When you need to calculate further statistics that require the mean (like standard deviation).
Mode The value that appears most frequently in the dataset. For categorical data (e.g., most common diagnosis, most common age group, most common blood type). Can have multiple modes. When all values are unique (no repeats). When you need a precise numerical summary.

💡 Mnemonic Mean vs. Median vs. Mode:
Mean = Mathematical average = sensitive to Mavericks (outliers).
Median = Middle = Most robust = best for Messy data.
Mode = Most frequent = Most common = good for Mcategories.

Measures of Spread — How Wide Is the Data?

Two datasets can have the same mean but very different spreads. Knowing the spread tells you whether the mean is a reliable summary.

Measure What It Is Nursing Example
Range Maximum value − Minimum value. The simplest measure of spread. Patient ages range from 2 to 78 years. Range = 76 years. Quick but sensitive to outliers.
Variance The average of squared differences from the mean. Measures how far each value is from the centre. Used in advanced calculations. Hard to interpret directly because units are squared (e.g., years²).
Standard Deviation (SD) The square root of variance. Expressed in the same units as the original data. Tells you the "typical" distance from the mean. Average waiting time = 45 minutes, SD = 12 minutes. Most patients wait between 33 and 57 minutes (mean ± 1 SD). If SD is very large, the mean is not very representative.

💡 Key Insight: A small SD means data points are clustered tightly around the mean the mean is reliable. A large SD means data is widely scattered the mean alone is misleading, and you should report the median and range as well.

Scenario: Clinic Waiting Times

🩺 The Situation: A nurse records waiting times (in minutes) for 10 patients: 20, 25, 30, 32, 35, 38, 40, 42, 45, 180.

Task: Calculate mean, median, and range. Which measure best describes typical waiting time?

Calculations:

  • Mean: (20+25+30+32+35+38+40+42+45+180) ÷ 10 = 48.7 minutes.
  • Median: Arrange in order: 20, 25, 30, 32, 35, 38, 40, 42, 45, 180. With 10 values, median = average of 5th and 6th = (35+38) ÷ 2 = 36.5 minutes.
  • Range: 180 − 20 = 160 minutes.
  • Mode: No repeating value no mode.

Interpretation: The mean (48.7 min) is pulled up by one extreme outlier (180 min perhaps a complex emergency). The median (36.5 min) better represents the typical patient's experience. The range (160 min) shows huge variation. The nurse should report the median, not the mean, and investigate why one patient waited 3 hours.

📝 Exam Tip: When a dataset has an outlier, the median is always the better measure of central tendency. Always check for outliers before choosing between mean and median. A quick way: compare mean and median. If they are very different, there is skewness or an outlier.

Inspecting the Distribution First

Before calculating any summary statistic, look at your data. Plot it. Count it. Group it. A histogram or simple tally can reveal:

  • Skewness: Is the data pulled to the left (negative skew) or right (positive skew)? Right-skewed data (like income, waiting times) needs the median.
  • Outliers: Are there extreme values that distort the mean?
  • Bimodality: Are there two peaks? This might mean two different groups mixed together (e.g., children and adults with different disease patterns).
  • Gaps: Are there missing values or impossible values (e.g., a 150-year-old patient probably a data entry error)?

🚨 Golden Rule: Inspect the distribution before summarising. Never calculate a mean without looking at the data first. A mean calculated on dirty data is a dirty mean.

Data Presentation — Tables, Graphs, and Text

Data that is not presented well is data that is not understood. Presentation is not decoration it is a tool for understanding. Choose the right tool for the message you want to convey.

When to Use What
Format Best For Example
Tables When exact numbers matter. When readers need to look up specific values. When there are many categories. A table showing ORS use by age group, sex, and village. Exact percentages for each cell.
Graphs / Charts When patterns, trends, or comparisons matter more than exact numbers. When you want to show change over time or differences between groups. A line graph showing malaria cases rising from January to June. A bar chart comparing vaccine coverage across districts.
Text When one clear message is enough. When the finding is simple and does not need visual support. "Vaccine coverage was 80%. One in five children missed the vaccine." A single sentence conveys the message.

💡 Rule of Thumb: Use a table when precision matters. Use a graph when the pattern matters. Use text when the message is simple. Often, the best presentation uses all three: a graph for the big picture, a table for the details, and text for the key message.

Building a Good Table

A bad table confuses. A good table clarifies. Every table must have five elements:

  • Title: Describes WHAT, WHERE, and WHEN. Example: "ORS use among children with diarrhoea, Clinic A, June 2026."
  • Rows: The categories you are comparing. Usually the "what" (e.g., received ORS, did not receive ORS).
  • Columns: The measurements. Usually number, percentage, rate, or ratio.
  • Totals: Always include a total row. It lets the reader check your math.
  • Units: Are the numbers counts, percentages, or rates? Label clearly.

Good Table Example:

Table 1: ORS use among children under 5 with diarrhoea, Clinic A, June 2026

Group Number Percent
Received ORS 64 80%
Did not receive ORS 16 20%
Total 80 100%

Source: Clinic A OPD register, June 2026.

Worked Example: Interpreting a Table

Data: 64 of 80 children received ORS.

Calculation: ORS coverage = 64 ÷ 80 × 100 = 80%

Finding: Coverage is good 4 out of 5 children received ORS.

But: 1 in 5 children (16 children) still missed ORS. That is not acceptable.

Action message: Review stock levels (was ORS out of stock?), counselling practices (did nurses explain ORS importance?), and triage practices (were severe cases prioritised while mild cases were missed?).

Key Principle: A good finding does not just state the number. It asks "What does this mean?" and "What should we do?" 80% coverage is a statistic. "1 in 5 children missed ORS investigate stock and counselling" is public health action.

Choosing the Right Graph
Purpose Graph Type When to Use It
Compare categories Bar chart Comparing vaccine coverage across districts, comparing mortality by age group, comparing staff numbers by cadre. Bars should not touch (they represent separate categories).
Show trend over time Line graph Malaria cases by month, patient attendance by week, temperature readings over 24 hours. Points are connected because time is continuous.
Show distribution Histogram Age distribution of patients, weight distribution of newborns, blood pressure ranges. Bars touch because the x-axis is continuous numerical data grouped into intervals.
Show parts of a whole Pie chart Proportion of deaths by cause, proportion of clinic visits by diagnosis. Use sparingly hard to compare slices accurately. Never use more than 5-6 categories.
Show relationship between two variables Scatter plot Relationship between age and blood pressure, between weight and blood sugar. Shows correlation (positive, negative, or none).
⚠️ Rules for Good Graphs:
  • Avoid 3D effects. They distort perception and add no information.
  • Start axes at zero for bar charts. Starting at 20 instead of 0 can make a small difference look huge.
  • Label units and time periods clearly. Is the y-axis "number of cases" or "rate per 1,000"? Is the x-axis "months" or "weeks"?
  • Use consistent colours. The same colour should mean the same thing across all graphs in a presentation.
  • Include a title that tells the reader what the graph shows, where, and when.
Bar Chart Example — Facility Reporting Completeness

📊 Graph: Facility Reporting Completeness (%), District X, June 2026

Line Graph Example — Monthly Malaria Cases

📈 Graph: Monthly Malaria Cases, Clinic A, January–June 2026

Communicating Findings — The "What, So What, Now What" Framework

Data without interpretation is just numbers. A good epidemiological finding answers three questions in order:

  • WHAT? The result. What did you find?
  • SO WHAT? The meaning. Why does it matter?
  • NOW WHAT? The action. What should be done?
Worked Example — ORS Coverage
Question Answer
WHAT? ORS coverage was 80% (64 of 80 children with diarrhoea received ORS).
SO WHAT? While 80% seems good, 1 in 5 children (16 children) missed ORS. In diarrhoea, missing ORS can lead to dehydration, hospitalisation, or death. This gap is unacceptable.
NOW WHAT? Review ORS stock levels (was there a stock-out?). Assess counselling quality (do nurses explain ORS importance?). Check triage (were mild cases overlooked while severe cases were prioritised?). Re-train staff on IMCI diarrhoea management.
Worked Example — Rising Malaria Cases
Question Answer
WHAT? Malaria cases rose from 35 in January to 90 in June a 157% increase. There was a slight dip in April.
SO WHAT? This upward trend suggests an outbreak or seasonal epidemic. If unchecked, July and August (peak rainy season) could see 150+ cases. The April dip is suspicious was it a true reduction or a data/reporting problem?
NOW WHAT? 1. Verify data quality for April. 2. Ensure adequate RDTs and ACTs stock. 3. Distribute nets in high-risk areas. 4. Conduct larval source management. 5. Alert the District Health Office. 6. Monitor weekly (not monthly) during the peak.

📝 Exam Tip: In any data interpretation question, structure your answer using What, So What, Now What. This shows you understand not just the numbers, but their meaning and implications for action.

Group Activity Template

📋 Task: Use the diarrhoea village dataset.

  • Calculate one summary statistic (mean, median, proportion, rate).
  • Draw one table or graph to present your finding.
  • Write one action-oriented message using What-So What-Now What.

Deliverable in 10 minutes. One presenter per group, maximum 2 minutes. Assessment focus: accuracy, clarity, and interpretation.

Final Recap — Five Key Principles
  • Inspect the distribution before summarising. Look at your data first. Check for outliers, skewness, and errors before calculating mean or median.
  • Mean, median, and mode describe the centre. Choose the right one: mean for symmetrical data, median for skewed data or outliers, mode for categories.
  • Range, variance, and SD describe spread. A small SD means the mean is reliable. A large SD means the data is scattered use median and range.
  • Ratios, proportions, and rates depend on denominator choice. The denominator must include everyone who could have experienced the event. Wrong denominator = wrong conclusion.
  • A good finding explains What, So What, and Now What. Numbers alone do not save lives. Interpretation and action do.
References
  • World Health Organization (WHO). (2020). Basic Epidemiology. Geneva: WHO Press.
  • Gordis, L. (2013). Epidemiology (5th ed.). Philadelphia, PA: Elsevier Saunders.
  • Grove, S. K., & Cipher, D. J. (2016). Statistics for Nursing Research: A Workbook for Evidence-Based Practice (3rd ed.). St. Louis, MO: Elsevier.
  • Polit, D. F., & Beck, C. T. (2017). Nursing Research: Generating and Assessing Evidence for Nursing Practice (10th ed.). Wolters Kluwer.

Quick Quiz

Summarizing and Presenting Findings Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Summarising, Presenting and Communicating Findings Read More »

Data Sources and Data Collection Methods

Data Sources and Data Collection Methods

Data Sources and Data Collection Methods
Learning Outcomes

By the end of this session, you should be able to:

  • Identify common sources of health data used in epidemiology and health services.
  • Explain why data quality matters for patient care, surveillance, and research.
  • Compare questionnaires, interviews, observation, and records review, knowing when to use each.
  • Apply practical quality-control checks during data collection.
  • Draft a simple data collection tool for a community health problem.
🧠 Why This Topic Matters:

Health workers make decisions using data from registers, reports, surveys, and clients. Poor data can lead to wrong priorities, missed outbreaks, and weak patient follow-up. Good data collection starts before the form is printed or uploaded. The best method depends on the question, the population, and the resources available.

Big Picture: From Question to Action

Every data collection effort follows a practical chain. Skip a step, and your data may be useless or worse, misleading.

  • Health Question: Clarify what you need to know
  • Data Source: Choose where information comes from
  • Method: Choose how to collect it
  • Quality Checks: Protect accuracy and completeness
  • Decision: Use findings to improve practice
💡 Mnemonic The Data Chain:

"Question → Source → Method → Quality → Decision" = QSMQD. Think: "Quality Starts Making Quick Decisions."

Session 1: Sources of Health Data

Health data are recorded facts about people, services, diseases, risks, and outcomes. Data may describe a person, a health facility, a community, or a whole district. Data become useful only when they are organized, analysed, and interpreted for decision-making.

🏥 Example: The number "47 malaria cases" is just a number. But "47 malaria cases in Village A this week, compared to an average of 8 cases per week over the past 6 months" is information that triggers action. Context transforms data into evidence.

Major Sources of Health Data

Use more than one source when triangulation is needed, comparing data from multiple sources to confirm findings and reduce bias.

Source What It Is Examples
Routine Records Data produced during normal service delivery. Collected continuously as part of patient care. OPD registers, patient files, ANC cards, immunisation registers, pharmacy stock cards, HMIS reports.
Surveys Data collected from a sample of people or households using structured questionnaires or interviews. Household survey on mosquito net use, client exit interview on satisfaction, school survey on handwashing.
Surveillance Regular, systematic reporting designed to detect disease patterns, trends, and outbreaks early. Weekly IDSR reports, maternal death notifications, laboratory reporting of confirmed cases, notifiable disease registers.
Research Studies Data collected under a planned scientific protocol to answer a specific research question. Cohort study of newborn survival, case-control study of cholera risk factors, RCT comparing two interventions.
Routine Health Records

Routine records are the backbone of health information systems. They are produced during normal service delivery and are available continuously.

Examples:
  • OPD register: Records diagnosis, age, sex, village, date of visit, and treatment given for every outpatient.
  • ANC register: Records visit number, gestational age, HIV testing result, haemoglobin, blood pressure, and tetanus vaccination for every pregnant woman.
  • Immunisation register: Records vaccine doses given, dates, and defaulters (children who missed scheduled doses).
  • Delivery register: Records mode of delivery, birth weight, APGAR score, maternal complications, and neonatal outcomes.
  • Pharmacy stock card: Tracks medicine stock levels, consumption, and stock-outs.
Strengths of Routine Records:
  • Available continuously, no special funding or planning needed.
  • Often cheap to use, the data is already being collected.
  • Cover large populations over long periods, good for trends.
  • Reflect real-world clinical practice, not artificial research settings.
Limitations of Routine Records:
  • May have missing entries, staff are busy, and some fields get skipped.
  • Diagnostic errors, a nurse may record "malaria" without a test, or confuse similar conditions.
  • Reflect only people who reached care, they miss people who never came to the facility (selection bias).
  • Variable quality across facilities, some health centres keep excellent records; others are chaotic.
  • Definitions may change over time, making trend analysis difficult.

📝 Exam Tip: When asked about routine records, always mention both strengths and limitations. Examiners want to see that you understand routine data is valuable but not perfect. Never say "routine data is always accurate" or "routine data is useless."

Health Surveys

Surveys collect data from a sample of people or households when routine records are insufficient or when you need population-level estimates.

Best used when:
  • The information is not found in routine records (e.g., mosquito net use at home, dietary practices, knowledge of danger signs).
  • The researcher needs population-level estimates (e.g., what percentage of ALL children in the district are vaccinated? Routine records only show those who came to clinic).
  • Views, practices, or behaviours must be measured (e.g., why do mothers miss ANC visits? What do community members think about family planning?).
Examples:
  • Household survey: Interviewing 200 randomly selected households about mosquito net ownership, use, and condition.
  • Client exit interview: Asking 50 patients leaving the clinic about their satisfaction, waiting time, and understanding of their diagnosis.
  • School survey: Observing and interviewing 300 students about handwashing practices and toilet use.
Surveillance Data

Surveillance is action-oriented. It is not just about counting cases, it is about detecting changes and triggering response.

Key features of surveillance:
  • It asks: "What is changing? Where? Who is affected?"
  • It should trigger investigation or response when thresholds are crossed.
  • It is usually mandatory, health facilities must report certain diseases by law.
  • It operates on regular cycles, weekly, monthly, or immediate (for epidemic-prone diseases).
Examples:
  • Weekly IDSR (Integrated Disease Surveillance and Response) reports: Facilities report counts of priority diseases (malaria, diarrhoea, measles, meningitis, etc.) every Monday.
  • Maternal death notification: Every maternal death must be reported within 24 hours and investigated within a week.
  • Laboratory reporting: Confirmed cases of TB, HIV, and cholera are reported from the lab to the district and national level.

Key Principle: Surveillance is not research. Research asks "Why?" and seeks to generate new knowledge. Surveillance asks "What is happening now?" and seeks to trigger action. A nurse doing surveillance reports data quickly; a nurse doing research analyses data deeply. Both are essential.

Research Study Data

Research studies are planned investigations designed to answer specific questions with rigorous methods.

Key features:
  • Clear study question and defined population.
  • Standardized procedures, every participant is treated the same way.
  • Ethical approval, research involving humans must be approved by an ethics committee.
  • Higher control over quality, trained data collectors, validated tools, supervision.
  • More costly, requires funding, time, and expertise.
Examples:
  • Cohort study: Following 500 newborns for 12 months to measure survival rates and identify risk factors for mortality.
  • Case-control study: Comparing 100 cholera cases with 100 healthy controls to identify shared exposures (water source, food, travel).
  • Randomized Controlled Trial (RCT): Randomly assigning 200 wards to use a new handwashing protocol vs. standard protocol, then comparing infection rates.
Other Useful Sources

Health data are not only clinical data. Administrators, communities, and digital systems also produce valuable information.

Source Type Examples
Administrative Data Staffing levels, budgets, medicine stock cards, supply records, transport logs, building maintenance records.
Community Data VHT (Village Health Team) reports, community mapping of water sources and latrines, local leader records of births and deaths, traditional birth attendant logs.
Digital Systems Electronic Medical Records (EMR), DHIS2 (District Health Information Software), ODK/Kobo Collect mobile forms, SMS reporting systems, telemedicine platforms.
Choosing the Right Source: Start With the Question

The most common mistake in data collection is choosing the tool before clarifying the question. Always start with: "What do I need to know?" Then ask: "Where can I find this information?"

Question Best Source
How many malaria cases were treated last month? OPD register / HMIS report (routine data)
Why are mothers missing ANC visits? Survey + interviews (not found in routine records)
Is measles increasing this week? Surveillance data (weekly IDSR reports)
Did a new intervention reduce infection rates? Research study data (before-and-after comparison or RCT)
How many nurses are on duty per shift? Administrative data (staffing rosters)
What percentage of households have a functional latrine? Community data (VHT household survey) or household survey
Worked Example 1: Selecting Sources for a Real Community Health Question

🩺 Problem: A health centre wants to know why many children miss measles vaccination.

  • Routine source: Immunisation register identifies who missed the dose and where they live. This gives the "what" and "where."
  • Survey source: Caregiver questionnaire explains why, access barriers (distance, cost), knowledge gaps ("I didn't know it was due"), fear ("I heard the vaccine causes fever"), or practical barriers ("I had no transport money").
  • Interview source: Health worker interviews explain system barriers, stock-outs ("We ran out of measles vaccine"), staffing ("The outreach nurse was on maternity leave"), or scheduling ("The clinic is only open when I am at work").

Conclusion: The best approach combines records review, caregiver survey, and staff interviews. No single source tells the whole story. Routine data shows the problem; surveys and interviews explain it.

Session 2: Quality and Uses of Health Data

Good data are not just "correct" data. They are fit for the decision being made. Data quality has multiple dimensions, and all of them matter.

What Makes Data Useful?
  • Fit for purpose: The data answers the question you are asking. Data on patient satisfaction does not help you plan drug stock.
  • Accurate enough: The data correctly represents what happened. A blood pressure of 180/110 recorded as 120/80 is inaccurate and dangerous.
  • Complete enough: Missing data can hide who is most affected. If 30% of age fields are blank, you cannot identify vulnerable age groups.
  • Available in time: Data submitted three weeks late cannot support outbreak response. Timeliness is a quality dimension.
  • Understandable: The people who need the data can read, interpret, and act on it. A complex statistical report given to a village health team is useless.
Five Data Quality Dimensions

Think of these as the five pillars of trustworthy data. Weakness in any pillar weakens the whole structure.

Dimension Definition Nursing Example
Accuracy Data correctly represent what happened. The recorded value matches reality. A child's weight is 12.5 kg, and the register records 12.5 kg. A diagnosis of "malaria" is confirmed by RDT, not guessed.
Completeness All required information is present. No critical fields are missing. If 100 outpatient visits are recorded but age is missing for 30, completeness is 70%. Incomplete age data hides which age groups are most affected.
Timeliness Data are submitted or available when needed for decision-making. Weekly outbreak reports submitted on Monday morning, not three weeks later. Maternal death notifications within 24 hours.
Consistency Data agree across different forms, registers, and reports. No contradictions. Immunisation tally sheets report 82 doses, and the monthly summary also reports 82 doses, not 128. The ANC register and the delivery register agree on the number of deliveries.
Validity Values are within acceptable rules and make sense. No impossible or illogical entries. Sex is not recorded as "7." Age is not negative. A 2-year-old does not have 10 pregnancies. Haemoglobin is not 500 g/dL.
💡 Mnemonic Data Quality Dimensions:

"All Cats Take Care Very Seriously" = ACTCVS → Accuracy, Completeness, Timeliness, Consistency, Validity. (Add "S" for "Sensitivity" if needed, but the five above are the core.)

Completeness in Detail

Completeness asks: "Are all required fields filled?"

  • Example: If 100 outpatient visits are recorded but age is missing for 30 patients, age completeness is 70%.
  • Why it matters: Incomplete data can hide who is most affected. If age is missing for 30% of malaria cases, you cannot tell whether children or adults are most at risk. Your prevention strategy will be blind.
  • Field practice: Review forms before leaving the facility or household. Check every required field. If a field is missing, ask the respondent or check the record immediately, do not wait.
Accuracy in Detail

Accuracy asks: "Is the recorded value correct?"

  • Example: Recording a 3-month-old child as "30 years old" is inaccurate. Recording a weight of 65 kg as "56 kg" is inaccurate.
  • Why it matters: Inaccurate data leads to wrong clinical decisions (wrong drug dose) and wrong public health decisions (targeting the wrong age group).
  • How to improve: Clear definitions, training, supervision, and verification. For critical indicators, verify a sample of forms against source documents (e.g., compare the register with the patient's actual file).
Timeliness in Detail

Timeliness asks: "Did data reach decision-makers on time?"

  • Example: Weekly outbreak reports submitted three weeks late cannot support rapid response. By the time the data arrives, the outbreak may be over or may have spread.
  • Why it matters: Timely data are essential for epidemics, stock-outs, referrals, and maternal deaths. A delayed maternal death report means missed opportunities to prevent the next death.
  • Field practice: Set daily upload deadlines and monitor submissions. Use digital tools with automatic timestamps. Hold supervisors accountable for late reports.
Consistency in Detail

Consistency asks: "Do related records agree?"

  • Example: Immunisation tally sheets report 82 doses, but the monthly summary reports 128 doses. Where did the extra 46 doses come from? Double counting? Transcription error? Fraud?
  • Why it matters: Inconsistency undermines trust in the data. If the district cannot trust facility reports, they cannot plan accurately.
  • Field practice: Reconcile totals before submission. Cross-check the register against the tally sheet against the summary report. If they do not match, find out why before sending the report.
Validity in Detail

Validity asks: "Are values within acceptable rules?"

  • Example: Sex should not be recorded as "7", the only valid values are "M" and "F" (or 1 and 2). Age should not be negative. A haemoglobin of 500 g/dL is physiologically impossible.
  • Why it matters: Invalid entries corrupt the dataset. If "7" is entered for sex 50 times, your analysis of male vs. female patients will be wrong.
  • Field practice: For digital forms, use constraints to prevent invalid entries at the point of collection (e.g., age must be between 0 and 120). For paper forms, train data collectors on valid ranges and check forms daily.
Common Data Quality Problems

Problems usually begin during collection, but they can also arise from systems and processes.

Problem Type Examples Prevention Strategy
Collection Errors Skipped questions. Poor probing or leading questions. Wrong units (weeks instead of months). Unclear handwriting in paper forms. Socially desirable answers. Training, supervision, pretesting, clear instructions, neutral wording, daily form review.
System Errors Duplicated records. Late uploads or missing forms. Mismatch between registers and summaries. Wrong facility or village code. Software bugs. Unique identifiers, automated deduplication, real-time monitoring, code validation, regular system audits.
Uses of Quality Health Data

Good data support better choices at every level of the health system:

Level How Data Is Used Nursing Example
Patient Care Follow-up, referrals, diagnosis history, treatment continuity. A nurse checks the ANC card and sees the patient missed her last two visits. She calls the patient to reschedule and assess for complications.
Public Health Detect outbreaks, monitor disease burden, target prevention. Weekly surveillance data shows a doubling of diarrhoea cases. The district triggers a cholera investigation and distributes water purification tablets.
Management Plan staff, medicines, outreach, equipment, and budgets. OPD data shows malaria peaks in April. The manager orders extra ACTs and RDTs in March, before the season starts.
Research Generate evidence, evaluate interventions, publish findings. A study finds that community health worker home visits reduced childhood mortality by 25%. This evidence is used to scale up the program nationally.
Worked Example 2: Calculating Basic Data Quality Indicators

🩺 Scenario: A supervisor reviews 10 completed household questionnaires.

Check Finding Calculation Result
Completeness 8 of 10 forms have all required fields filled. 8 ÷ 10 × 100 80%
Timeliness 7 of 10 forms uploaded same day. 7 ÷ 10 × 100 70%
Validity 2 forms have age outside expected range (e.g., 150 years or -3 years). 8 valid ÷ 10 × 100 80%

Interpretation: The team should improve same-day uploads (timeliness = 70%) and check age-entry rules before full data collection (validity = 80%). Completeness is acceptable but could be improved. These indicators guide targeted quality improvement.

Session 3: Data Collection Methods

Choosing a data collection method is a design decision. The wrong method produces the wrong data. The right method produces trustworthy evidence.

Five Questions to Guide Method Selection:
  • What exactly must be measured? (Knowledge? Behaviour? Clinical outcome?)
  • Who or what has the information? (Patients? Caregivers? Health workers? Records?)
  • Is the topic private or difficult? (Sexual behaviour? Domestic violence? Substance use?)
  • What time, skills, and tools are available? (Trained interviewers? Digital devices? Transport?)
  • What checks will protect the data? (Supervision? Validation? Duplicate checks?)
Questionnaires

Definition: Structured questions asked in the same way to many people. Usually self-administered or administered by a trained interviewer reading from a script.

Strengths:
  • Good for surveys and quantitative analysis, every respondent answers the same questions, making comparison easy.
  • Easy to standardize across multiple data collectors, reduces interviewer bias.
  • Works well for knowledge, practice, and service-use questions (e.g., "Do you sleep under a mosquito net?" "How many ANC visits did you attend?").
  • Can be administered to large numbers relatively quickly.
  • Digital questionnaires (ODK, Kobo) allow automatic skip patterns and validation.
Limitations:
  • May miss detailed explanations, a questionnaire cannot probe "Why did you miss ANC?" as deeply as an interview.
  • Poor wording creates biased answers. A leading question like "You always attend ANC, don't you?" produces socially desirable answers.
  • Respondents may forget ("When was your last ANC visit?" "Um... maybe March?") or give socially desirable answers ("Yes, I wash my hands" when the interviewer can see dirty hands).
  • Requires literacy if self-administered; requires trained interviewers if administered.
Interviews

Definition: Guided conversations to explore experiences, explanations, and perceptions. Can be structured (fixed questions), semi-structured (flexible questions with probes), or unstructured (open conversation).

Strengths:
  • Useful for understanding reasons and perceptions, "Why did you not seek care immediately?" "What did you think when the nurse told you your child had malaria?"
  • Allows probing and clarification, the interviewer can ask follow-up questions based on the respondent's answers.
  • Good for health workers, leaders, and clients, anyone with complex experiences to share.
  • Can build rapport and trust, especially for sensitive topics.
Limitations:
  • Requires skilled interviewers, untrained interviewers may lead respondents, misrecord answers, or fail to probe deeply.
  • Takes time to transcribe and analyse, qualitative data is rich but labour-intensive.
  • Responses may be influenced by interviewer style, a friendly interviewer may get different answers than a stern one (interviewer bias).
  • Not feasible for large sample sizes due to time and cost.

📝 Exam Tip Questionnaire vs. Interview: Use a questionnaire when you need standardized, comparable data from many people (surveys, knowledge assessments). Use an interview when you need depth, explanation, and understanding from fewer people (exploratory research, understanding barriers, capturing stories). Many studies use both, questionnaires for breadth, interviews for depth.

Observation

Definition: Recording what is seen using a checklist or structured form. The observer watches and records behaviours, conditions, or practices without interfering.

Strengths:
  • Good for facility readiness and practice assessment, "Is handwashing soap available at every sink?" "Does the nurse use a sterile needle for every injection?"
  • Can verify whether resources are present, you see the stock-out with your own eyes, rather than relying on a report.
  • Reduces reliance on self-report, people may say they wash their hands, but observation shows whether they actually do.
  • Can capture non-verbal behaviours and environmental conditions.
Limitations:
  • People may change behaviour when observed (Hawthorne effect). A nurse who never washes hands may start washing when she sees the observer.
  • Only captures what happens during observation, you miss what happens at night, on weekends, or when you are not there.
  • Requires clear observation criteria, what counts as "good handwashing"? 20 seconds? Soap? Running water? Without clear criteria, observers disagree.
  • Can be intrusive and may affect the normal workflow of the facility.
Records Review

Definition: Extracting data from existing documents or systems, registers, patient files, laboratory logs, pharmacy stock cards, HMIS reports.

Strengths:
  • Useful for trends and service volumes, "How many malaria cases were treated each month for the past year?"
  • Usually cheaper than collecting new data, the data already exists; you just extract it.
  • Can cover long periods, years of data can be reviewed in days.
  • No respondent burden, you do not need to ask anyone questions.
Limitations:
  • Dependent on record quality, if the original records are incomplete, inaccurate, or illegible, your extracted data will be too.
  • Missing values may be difficult to correct, you cannot go back and ask the patient from 2019 why a field was blank.
  • Definitions may vary over time or across facilities, "malaria" may mean "clinical diagnosis" in one facility and "RDT-confirmed" in another. Comparing them is misleading.
  • May require ethical approval if patient identifiers are used.
Side-by-Side Comparison of Methods
Method Best For Strengths Limitations
Questionnaire Large surveys, standardized knowledge/practice data Standardized, scalable, easy to analyse Misses depth, poor wording biases answers, recall errors
Interview Understanding reasons, perceptions, complex experiences Deep, flexible, builds rapport Time-consuming, requires skill, interviewer bias
Observation Verifying practices, facility readiness, behaviour Objective, reduces self-report bias Hawthorne effect, limited to observation period, needs clear criteria
Records Review Trends, service volumes, historical data Cheap, covers long periods, no respondent burden Dependent on original quality, missing data hard to fix, definitions may vary
Digital Data Collection

Digital tools such as ODK (Open Data Kit), KoboToolbox, and REDCap have transformed data collection in low-resource settings.

Advantages of digital tools:
  • Constraints prevent impossible values, e.g., age cannot be negative, haemoglobin cannot exceed 20 g/dL.
  • Skip patterns reduce irrelevant questions, if a woman says she is not pregnant, the tool automatically skips all pregnancy-related questions.
  • Daily uploads allow supervisors to identify problems early, instead of discovering errors at the end of fieldwork, supervisors can correct them daily.
  • GPS tagging ensures data collectors actually visited the claimed location.
  • Automatic timestamps verify when data was collected.
  • No transcription errors, data is entered directly into digital format, eliminating the step of transferring paper to computer.
Important caveats:
  • Digital systems still require training, supervision, and data protection. A tablet with unencrypted patient data is a liability.
  • Technology can fail. Batteries die, networks fail, devices break. Always have a paper backup plan.
  • Not everyone is comfortable with technology. Older data collectors or those with limited literacy may struggle with digital forms.
  • Data security is critical. Patient identifiers must be encrypted, and access must be restricted to authorized personnel.
Ethical Practice During Data Collection

Ethics is not an afterthought, it is built into every step of data collection.

  • Explain the purpose of data collection in simple language the respondent can understand. Do not use medical jargon.
  • Seek voluntary informed consent before asking questions. The respondent must understand what they are agreeing to, know they can refuse, and know they can withdraw at any time.
  • Protect privacy, especially for sensitive health information (HIV status, mental health, sexual behaviour, substance use). Conduct interviews in private settings.
  • Avoid collecting names unless they are absolutely necessary for follow-up. Use study IDs instead.
  • Store completed forms and devices securely. Paper forms should be locked in a cabinet. Digital data should be encrypted and password-protected.
  • Do not share individual data with unauthorized people. Aggregate summaries (e.g., "30% of patients were hypertensive") are fine; individual patient records are not.

🚨 Ethical Red Line: Never collect data without informed consent. Never share identifiable patient information. Never pressure a respondent to answer questions they are uncomfortable with. Ethical violations destroy trust, harm patients, and can lead to legal consequences.

Worked Example 3: Turning an Objective Into Data Collection Questions

🩺 Objective: Assess barriers to completing four ANC visits among pregnant women.

Indicator Question Response Option
ANC attendance How many ANC visits have you attended during this pregnancy? Number of visits (0, 1, 2, 3, 4, 5+)
Distance barrier How long does it take you to reach the nearest health facility? Minutes / hours (continuous)
Cost barrier Did transport cost stop you from attending ANC? Yes / No (nominal)
Knowledge When should a pregnant woman start ANC? First trimester / later / do not know (ordinal, ordered by correctness)

Key principle: Each question maps directly to an indicator. The response options match the data type needed for analysis. This is how good tools are built, indicator by indicator, question by question.

Better Question Writing: Weak vs. Improved
Weak Question Problem Improved Question
"You always attend ANC, don't you?" Leading question. Suggests the "correct" answer. Produces socially desirable responses. "How many ANC visits have you attended during this pregnancy?" Neutral wording. Produces a measurable response.
"Do you have good health?" Vague. "Good health" means different things to different people. Unmeasurable. "In the past 30 days, how many days were you unable to do your normal activities because of illness?" Specific, time-bound, measurable.
"Why didn't you come to the clinic?" Open-ended without structure. Hard to analyse. May embarrass the respondent. "What was the main reason you did not attend the clinic? (Select one...)"
Session 4: Quality Control Methods in Data Collection

Quality control is continuous, not a final activity. It begins the moment you design your tool and continues until the data are analysed and reported. Think of it as infection prevention for data, every step needs a barrier against error.

The Quality Control Cycle: Before, During, and After
Phase What to Do Practical Examples
Before Design, review, translate, and pretest the tool. Define every variable. Use simple words. Translate carefully. Pretest with 5-10 people similar to your target population.
During Observe, review, debrief, and correct in real time. Supervisors observe interviews. Review forms before leaving the field. Check GPS, dates, and required fields daily. Hold short debriefs.
After Clean, validate, document, and protect the dataset. Check for duplicates, missing values, impossible values. Compare related variables for logic errors. Keep raw data separate. Store securely.

📝 Exam Tip The 5 Steps of Quality Control: "Design, Train, Observe, Clean, Improve" = DTOCI. Think: "Data Team Observes, Cleans, Improves." Another mnemonic: "Prepare, Collect, Check, Clean, Protect" = PCCCP.

Before Data Collection: Design for Quality

The quality of your data is determined before you collect a single form. A poorly designed tool will produce poor data no matter how carefully your team works.

  • Define every variable clearly before designing the form. What exactly do you mean by "fever"? By "delay"? By "treatment"? Write operational definitions.
  • Use simple words and local examples that respondents understand. Avoid medical jargon. Instead of "Did you experience dyspnoea?" ask "Did you feel short of breath?"
  • Translate carefully and back-check meaning. If your tool is in English but your respondents speak Luganda, translate professionally and then back-translate to English to check accuracy.
  • Pretest the tool with a small group (5-10 people) similar to your target population. Watch for confusion, hesitation, or multiple interpretations of the same question.
  • Revise confusing questions before full data collection. If 3 out of 10 pretest respondents misunderstand a question, rewrite it.

⚠️ Common Mistake: Skipping the pretest to "save time." This always costs more time later because you will have to re-collect data or throw out invalid responses. Pretesting is not optional, it is insurance.

During Data Collection: Supervision in Real Time

Errors made during collection are the hardest to fix later. Supervision must be continuous and immediate.

  • Supervisors should observe selected interviews respectfully. Do not interrupt, but watch for leading questions, skipped questions, or rushed responses. Give feedback privately after the interview.
  • Review completed forms before leaving the village or facility. Do not let a data collector leave with a form full of blanks or obvious errors. Fix it while the respondent is still available.
  • Check GPS coordinates, dates, facility codes, and required fields daily. A form with no date or wrong facility code is useless for analysis.
  • Hold short debriefs at the end of each day. Discuss errors, difficult questions, and respondent reactions. Share solutions across the team.
  • Correct procedures immediately, not at the end of fieldwork. If one data collector is consistently making the same mistake, retrain them today, not next week.
After Data Collection: Cleaning and Validation

Data cleaning is not just "fixing typos." It is a systematic process of checking, correcting, and documenting every decision.

  • Check for duplicates: Did the same person get interviewed twice? Same ID number? Same name? Remove or merge duplicates.
  • Check for missing values: Which questions were skipped? Was it random (data collector error) or systematic (respondents refused a sensitive question)?
  • Check for impossible values: Age = -5? ANC visits = 50? Sex = "Male" but pregnancy question answered? These are red flags.
  • Compare related variables for logic errors: A child with "no fever" should not have a "date of fever onset." A 12-year-old should not be married. Inconsistencies reveal data quality problems.
  • Document all cleaning decisions in a simple log. What did you find? What did you change? Why? Who decided? This log is your audit trail.
  • Keep raw data separate from cleaned data. Never overwrite the original file. Save the raw dataset as "Raw_Data_v1" and the cleaned dataset as "Clean_Data_v1."
  • Store files securely and restrict access to authorized people. Health data are confidential. Use password protection, encrypted drives, and locked cabinets for paper forms.

📝 Exam Tip: When asked about data cleaning, always mention at least: duplicates, missing values, impossible values, logic checks, documentation log, and data security. This shows comprehensive understanding.

Worked Example: Finding Errors in a Small Dataset

Dataset:

Record Age Sex ANC Visits Problem QC Action
001 24 Female 3 No obvious error Accept as valid.
002 -5 Female 2 Invalid age, negative number is impossible. Check original form. If typo (e.g., meant 5), correct with documentation. If truly unknown, code as missing.
003 31 Male 4 Logic error, males do not attend ANC. Check original form. Likely a coding error (sex should be Female). Correct with documentation. If the respondent was indeed male, investigate why ANC was recorded.
004 18 Female 12 Unlikely value, 12 ANC visits is extremely high (WHO recommends 8+). Check original form. Could be a typo (meant 2?). Verify with the respondent or facility register. Do not assume, confirm.
💡 QC Golden Rule:

Confirm the source, correct only when evidence exists, and document the correction. Never guess. Never delete data without a reason. Your cleaning log is your proof that you did not fabricate or manipulate data.

A Simple Supervisor QC Checklist

Use this checklist at the end of every data collection day:

  • ☐ Are all required questions answered? (No blank mandatory fields.)
  • ☐ Are dates, facility names, and village names correctly recorded? (No "N/A" or vague entries.)
  • ☐ Are skip patterns followed correctly? (If "No" to question 5, question 6 should be blank.)
  • ☐ Are values within expected ranges? (Age 0-120, blood pressure within physiological limits.)
  • ☐ Were any refusals, incomplete interviews, or unusual events documented? (Transparency about problems is a sign of good data quality.)
  • ☐ Are signatures and IDs present? (Data collector and supervisor must sign.)
  • ☐ Is the form legible and complete? (No torn pages, no pencil, use pen only.)
Session 5: Group Task Design a Simple Data Collection Tool

Now you apply everything you have learned. Designing a good tool is a skill that improves with practice. Follow the four-step framework below.

The Scenario

🩺 Community Problem: Many children under five are coming late for treatment of fever. The health team needs to understand why.

Task: Design a simple tool to collect information from caregivers. Your tool should help explain delays and guide community health action.

Step 1: Define the Information Needed

Before writing a single question, ask: "What do we need to know to solve this problem?" Organise your information needs by category:

Information Category What to Know Why It Matters
Who is the child? Age, sex, village, household size. Identifies vulnerable groups (e.g., infants under 1 year may delay more).
What happened? Fever onset, danger signs recognised, treatment sought. Reveals whether caregivers recognise danger signs and act appropriately.
When did care begin? Time from fever onset to first action (hours/days). Quantifies the delay. Allows comparison across groups.
Where was care sought? Home, drug shop, clinic, traditional healer, or health facility. Reveals care-seeking patterns. Many caregivers go to drug shops first.
Why the delay? Cost, distance, transport, knowledge, decision-making barriers, drug stock-outs. Identifies the modifiable barriers that the health program can address.
Step 2: Draft Core Questions

Each question must be clear, specific, and answerable. Avoid leading questions, double-barrelled questions, and jargon.

Example Core Questions:

  • "When did the fever start?" (Date and time. This allows calculation of delay duration.)
  • "What was the first action taken by the caregiver when the child got fever?" (Home care / VHT / drug shop / clinic / health facility / nothing / other. This reveals the care-seeking pathway.)
  • "How long did it take from when the fever started until you reached the first provider?" (Hours / days. This quantifies delay.)
  • "What was the main reason for not seeking care earlier?" (Cost / distance / no transport / did not know it was serious / husband not home to decide / no drugs at facility / other. This identifies barriers.)
  • "Was the child tested or treated for malaria?" (Yes / No / Don't know. This checks whether appropriate care was received.)

⚠️ Weak vs. Strong Questions:

  • Weak: "You always take your child early for treatment, don't you?" (Leading, suggests the "correct" answer. Embarrasses respondents who did not.)
  • Strong: "How long after the fever started did you first seek care for your child?" (Neutral, specific, non-judgmental.)
  • Weak: "Did you go to the clinic because of the fever and what medicine did they give?" (Double-barrelled, asks two things at once. Which do you code?)
  • Strong: "Where did you first seek care for the fever?" (One question, one answer.)
Step 3: Select Data Source and Method
Source / Method What It Provides When to Use It
Primary: Caregiver questionnaire Direct information about recent fever episodes, delays, and barriers. When you need to understand behaviour, perceptions, and reasons for delay.
Secondary: OPD register Attendance patterns, diagnosis, age, and date of visit. When you need to quantify the problem (how many, when, who) and compare with caregiver reports.
Qualitative: VHT interviews Community-level insights on referral barriers, cultural beliefs, and trust in the health system. When numbers alone do not explain "why" you need stories and context.
Sampling: Facility-based Select caregivers of under-five children attending the facility during the week. When you need a manageable sample that is easy to access and representative of care-seekers.
Step 4: Plan Quality Control

Build quality into the tool from the start:

  • Use clear definitions: "Late care" means care sought after 24 hours from fever onset. Define this in the training manual, not just in your head.
  • Pretest the tool with 3-5 caregivers before full use. Watch for confusion about "fever" (some cultures use different words), "first action" (some may list multiple things), and "delay" (some may not think they delayed).
  • Train data collectors on neutral probing. If a caregiver says "I don't know," the data collector should not suggest answers. They should say: "Take your time. What do you remember?"
  • Use local terms for fever. In some communities, "fever" is called "omusujja" (Luganda) or described as "the body is hot." Use the term the respondent understands.
  • Review completed forms daily for missing and inconsistent responses. If a caregiver said "no transport" for delay but also said they walked, flag it for follow-up.
  • Correct tool problems early and document changes. If question 4 is misunderstood by 4 out of 5 pretest respondents, rewrite it before day 1 of real data collection.
Work Example: Mini Questionnaire

Module: Community Fever Care-Seeking Among Children Under Five

Variable Name Question Response Options
child_age_months How old is the child? (in completed months) ____ months (0-59)
fever_start When did the fever start? (date and approximate time) Date: ____/____/____ Time: ____
first_action What was the first action taken when the child got fever? 1=Home care 2=VHT 3=Drug shop 4=Health facility 5=Nothing 6=Other
time_to_care How many hours passed from fever start until first action? ____ hours (0-168)
delay_reason What was the main reason for not seeking care earlier? 1=Cost 2=Distance 3=No transport 4=Did not know serious 5=Decision-maker absent 6=Facility closed 7=Other
malaria_tested Was the child tested for malaria? 1=Yes 2=No 3=Don't know
malaria_treated Was the child given malaria treatment? 1=Yes 2=No 3=Don't know
Group Presentation Guide

When presenting your tool, cover these five points:

  • State the health problem and the purpose of your tool. What are you trying to learn, and why does it matter?
  • Identify your data source and collection method. Primary (questionnaire), secondary (register), or both? Why did you choose this method?
  • Show five core questions and explain why each is included. Link every question to a specific information need.
  • Explain two quality-control checks you will use. One during collection (e.g., daily form review) and one after (e.g., logic checks for inconsistencies).
  • Mention one ethical issue and how you will address it. Informed consent? Confidentiality? Protection of vulnerable children? Respect for cultural practices?
Key Takeaways
  • Health data come from routine records, surveys, surveillance, and research studies. Each source has strengths and limitations.
  • Data quality means the data are accurate, complete, timely, consistent, and valid. Poor quality data are worse than no data, they lead to wrong decisions.
  • Data collection methods should match the objective and the source of information. Do not use a questionnaire when a register already has the answer. Do not use a register when you need to understand "why."
  • Quality control begins during tool design and continues through training, supervision, and cleaning. It is not a one-time check at the end.
  • A simple, clear tool is usually better than a long, confusing one. Ten good questions beat fifty bad ones.
  • Ethics are inseparable from data collection. Informed consent, confidentiality, and respect for respondents are not optional, they are professional obligations.
Quick Quiz Recap & Exam Preparation

Q: Which data source would you use to count outpatient malaria cases for last month?
Answer: The OPD register (secondary source). It already records diagnosis, date, and patient count. A survey would be unnecessary and wasteful.
Key principle: Use existing data before collecting new data.

Q: Name two dimensions of data quality and give one example of each.
Answer:
Accuracy: A blood pressure reading of 120/80 is accurate if measured with a calibrated machine and proper technique. A reading of 300/200 is likely inaccurate (check the cuff size and patient position).
Completeness: An ANC register with 95% of required fields filled is complete. One with 40% missing is incomplete and unreliable for planning.
Other valid dimensions: Timeliness (data available when needed), Consistency (same method over time), Validity (measures what it claims to measure).

Q: When is an interview better than a questionnaire?
Answer: An interview is better when:
• The respondent is illiterate or has low literacy.
• The topic is sensitive (sexual behaviour, domestic violence, substance use) and requires trust and probing.
• The questions are complex and need explanation (e.g., "What do you think causes malaria?").
• You need to explore unexpected answers (qualitative depth).
A questionnaire is better for large samples, standardised responses, and quantitative analysis.

Q: Give one quality-control check during data collection.
Answer: Supervisor observation of interviews. The supervisor watches silently as the data collector conducts an interview, then gives feedback on technique (neutral probing, correct skip patterns, respectful behaviour). Another valid answer: Daily form review before leaving the field to catch missing or inconsistent responses while the respondent is still available.

Q: Rewrite this weak question: "You always take your child early for treatment, don't you?"
Answer: Weak because: It is leading (suggests the "correct" answer), double-barrelled ("always" + "early"), and judgmental (embarrasses respondents who delayed).
Strong rewrite: "How many hours passed from when your child's fever started until you first sought care?" (Specific, neutral, quantitative, non-judgmental.)

Q: What is the difference between primary and secondary data?
Answer:
Primary data: Collected specifically for your study. You control the method, timing, and quality. Example: A caregiver questionnaire about fever delays.
Secondary data: Already exists, collected for another purpose. You do not control how it was collected. Example: OPD registers, HMIS reports, census data.
Primary data are more tailored but more expensive. Secondary data are cheaper but may not answer your exact question.

Q: Why should you keep raw data separate from cleaned data?
Answer: Raw data are your original, unaltered record. If someone questions your findings, you can show the raw data as proof. Cleaned data have been modified, and if you make a mistake during cleaning, you need the raw data to start over. Never overwrite raw data. It is your audit trail and your insurance policy.

Q: What is informed consent, and why does it matter in data collection?
Answer: Informed consent means the respondent understands:
• The purpose of the study.
• What they will be asked to do.
• That participation is voluntary and they can withdraw at any time.
• How their data will be used and protected.
• Any risks or benefits.
It matters because respect for persons is a core ethical principle. Forcing someone to participate or hiding the true purpose is unethical and may invalidate your data.

Q: What is a skip pattern, and why is it important?
Answer: A skip pattern (or filter) directs the data collector to skip irrelevant questions based on a previous answer. Example: If a respondent answers "No" to "Are you pregnant?" the data collector skips all pregnancy-related questions. This prevents illogical responses (a male answering pregnancy questions) and saves time.
In electronic tools (ODK, KoboToolbox), skip patterns are programmed automatically. In paper tools, arrows or instructions must be clear.

Q: What should you do if you find an impossible value during data cleaning?
Answer: Follow the QC golden rule:
1. Do not delete or change it immediately.
2. Check the original form or re-contact the respondent if possible.
3. Correct only when evidence exists (e.g., a clear typo: "-5" should be "5").
4. Document the correction in your cleaning log (record number, variable, old value, new value, reason, date, your name).
5. If the value cannot be verified, code it as missing and note why.

Reflection for Students

Apply what you have learned to your own context:

  • Think of one health problem in your community or clinical area. (Example: Low immunisation coverage, high teenage pregnancy, frequent drug stock-outs.)
  • Write one objective that can be answered using data. (Example: "To determine the proportion of children under 1 year who are fully immunised in Village X.")
  • Identify one data source and one collection method. (Example: Immunisation register + structured observation of vaccination sessions.)
  • Write three clear questions you would include in your tool. Make them specific, neutral, and answerable.
  • Explain how you would protect data quality and confidentiality. (Example: Daily supervisor review, secure storage, coded IDs instead of names, informed consent.)
References
  • World Health Organization (WHO). (2020). Framework and Standards for Country Health Information Systems. Geneva: WHO Press.
  • Centers for Disease Control and Prevention (CDC). (2012). Principles of Epidemiology in Public Health Practice (3rd ed.). Atlanta, GA.
  • Bowling, A. (2014). Research Methods in Health: Investigating Health and Health Services. McGraw-Hill Education.
  • Gordis, L. (2013). Epidemiology (5th ed.). Saunders Elsevier.

Quick Quiz

Data Source and Data Collection Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Data Sources and Data Collection Methods Read More »

principles of biostatistics

Principles of Biostatistics

Principles of Biostatistics
Learning Outcomes

By the end of this session, you should be able to:

  • Explain the role of biostatistics in health decision making.
  • Define data, population, sample, and variable.
  • Distinguish dependent and independent variables.
  • Classify data as qualitative, quantitative, discrete, or continuous.
  • Apply these concepts using nursing and community health examples.
What Is Biostatistics?

Biostatistics is the use of statistical methods to collect, summarize, analyze, and interpret health related data. It is the bridge between raw numbers and meaningful health decisions.

The Biostatistics Equation

BIO (Life, health, disease, patients) + STATISTICS (Methods for working with data) = BIOSTATISTICS (Statistics applied to health sciences)

Why Nurses Need Biostatistics

Biostatistics is not just for researchers or statisticians. It is an essential tool for every nurse who wants to provide evidence based care and protect their community.

  • To understand patient records and ward reports: A nurse who can read and interpret data tables, graphs, and summary statistics can spot problems faster and communicate them clearly.
  • To judge whether a treatment or intervention worked: Did the new handwashing protocol reduce infections? Did the nutrition education program improve children's weight? Biostatistics gives you the numbers to answer these questions.
  • To detect unusual patterns such as outbreaks: A sudden spike in diarrhoea cases, a cluster of wound infections, or an unexpected drop in immunisation coverage, all of these are statistical signals that require action.
  • To communicate evidence clearly to teams and communities: When you tell a village leader that "malaria cases dropped by 40% after net distribution," you are using biostatistics to build trust and motivate action.
  • To make safer decisions using facts, not guesswork: Intuition is valuable, but data is verifiable. Biostatistics helps you separate real trends from random noise.

💡 Key Insight: A nurse without biostatistics is like a clinician without a stethoscope, you can function, but you are missing a critical tool for understanding what is really happening.

From Data to Health Decisions: The Four Step Path
Step What Happens Nursing Example
Data Raw facts or observations are collected. Blood pressure readings from 200 ANC patients; temperature records from the paediatric ward.
Information Data is organized and summarized into meaningful patterns. "45 out of 200 ANC patients (22.5%) have hypertension." "Average waiting time is 2.5 hours."
Evidence Information is analyzed to answer specific questions and test hypotheses. "Women who attended ANC before 12 weeks were 30% less likely to have anaemia at delivery."
Decision Evidence is used to guide action, policy, or clinical practice. "We will screen all ANC patients for hypertension and start iron folate supplementation in the first trimester."

🏥 Clinical Example: A clinic reviews ANC records and finds many women have low haemoglobin. The data shows 60% of pregnant women are anaemic. The information is organized by trimester. The evidence shows that women who started iron folate in the first trimester had higher haemoglobin at delivery. The decision: improve iron folate counselling and ensure early initiation.

What Is Data?

Data are facts or observations collected for a purpose. In health, data comes from patients, records, surveys, observations, laboratory tests, and community reports.

  • A single patient record is one unit of data.
  • A collection of records becomes a dataset.
  • A dataset organized into rows and columns is the foundation of all biostatistical analysis.
Example: One Patient Record
Patient ID Age Sex Temperature Diagnosis
001 28 years Female 38.5°C Malaria
  • Each column is a variable (a characteristic that can vary).
  • Each row is a patient or observation (one unit of data).

📝 Exam Tip: In biostatistics, we organize patient observations so that patterns can be seen and decisions can be made. A messy register is data. A clean table is information. A graph with a trend line is evidence.

Population and Sample
Term Definition Nursing Example
Population The entire group of people or records that we are interested in studying. It is the complete set. All first year nursing students at Mulago. All under five children in a parish. All pregnant women attending ANC at Hospital X.
Sample A smaller, manageable subset selected from the population for actual study. We use samples because studying the entire population is usually impossible. 50 selected students from the nursing school. 120 selected under five children from the parish. 80 pregnant women interviewed from the ANC register.

⚠️ Key Idea: A sample should represent the population well enough to support fair conclusions. A biased sample (e.g., only interviewing rich families) produces misleading results. A representative sample (randomly selected, matching the population's characteristics) produces trustworthy evidence.

Worked Example: Population and Sample

Study Question: What proportion of under five children in a parish had malaria symptoms in the last two weeks?

Element Description
Target Population All under five children in the parish.
Sampling Frame Households listed by the Village Health Team (VHT). This is the list from which the sample is drawn.
Sample 120 selected under five children, chosen using systematic random sampling (every 5th household on the VHT list).

Why this matters: If the VHT list is incomplete (missing poor households or remote villages), the sample will be biased. The findings may underestimate true malaria burden. Good sampling requires a complete, accurate sampling frame.

💡 Mnemonic: Population vs. Sample: "Population = People All Together. Sample = Selected Part." Think of tasting soup: you do not drink the whole pot (population), you take a spoonful (sample) to judge the flavour. But the spoonful must be stirred well (random sampling) to represent the whole pot.

What Is a Variable?

A variable is any characteristic that can take different values across people, places, records, or time. If a characteristic is the same for everyone, it is a constant, not a variable.

Variable Type Definition Examples
Patient Variable Characteristics of the individual patient. Age, sex, weight, blood pressure, temperature, occupation.
Disease Variable Characteristics of the disease or condition. Diagnosis, severity, duration of illness, complications.
Service Variable Characteristics of the healthcare service or system. Waiting time, medicine availability, referral status, nurse to patient ratio.

⚠️ Constant vs. Variable: If you study only female patients, "sex" is a constant (all are female), it does not vary, so it cannot explain differences in outcomes. A variable must vary. This is why researchers sometimes exclude constants from analysis or stratify by them.

Variables in Nursing Examples
  • Patient age: 5 months, 20 years, 72 years. (Quantitative, discrete)
  • Outcome of delivery: Live birth, stillbirth, maternal referral. (Qualitative, nominal)
  • Treatment received: ORS, antibiotic, antimalarial, none. (Qualitative, nominal)
  • Pain score: 0 to 10 scale. (Quantitative, ordinal, or sometimes treated as discrete)
  • Length of hospital stay: Number of days admitted. (Quantitative, discrete)
  • Blood pressure: 120/80 mmHg. (Quantitative, continuous)
  • Patient satisfaction: Poor, fair, good, excellent. (Qualitative, ordinal)
Dependent and Independent Variables

In research and epidemiology, variables are classified by their role in the study, not just by what they measure.

Term Definition Also Called
Independent Variable The possible cause, exposure, predictor, or factor that may influence an outcome. It is the "input" or "explanation." Exposure, predictor, explanatory variable, risk factor, intervention.
Dependent Variable The outcome, response, or result being explained or measured. It is the "output" or "effect." Outcome, response variable, endpoint, result.

💡 Simple Question to Identify Variables: "What factor may influence what outcome?" The factor is the independent variable. The outcome is the dependent variable.

Worked Example 1: Malaria Prevention

Research Question: Does sleeping under an insecticide treated net (ITN) reduce malaria among children under five?

Variable Role Values
Net use Independent variable (exposure) Yes / No
Malaria status Dependent variable (outcome) Positive / Negative

Interpretation: Net use is the exposure (the thing we think might cause or prevent something). Malaria status is the health outcome (the thing we are trying to explain or predict).

Worked Example 2: ANC Attendance and Anaemia

Research Question: Is early ANC attendance associated with maternal anaemia at delivery?

Variable Role Values
Early ANC attendance Independent variable (exposure) Before 12 weeks: Yes / No
Anaemia at delivery Dependent variable (outcome) Yes / No (or Hb level in g/dL)

Important: The dependent variable is the outcome you want to explain. Never confuse the two. A common exam trap: students label "early ANC" as the outcome because it "sounds like a good thing." But in this study, we are asking whether early ANC causes less anaemia, so anaemia is the outcome.

Worked Example 3: Health Education and Handwashing

Research Question: Does health education improve handwashing practice among mothers?

Variable Role Data Type
Health education received Independent variable Qualitative, nominal (Yes / No)
Handwashing practice Dependent variable Qualitative, nominal (Good / Poor) or ordinal (Never, Sometimes, Always)
Common Mistakes to Avoid
Mistake Why It Is Wrong How to Fix It
Calling every variable an "outcome" Not every variable is something you are trying to explain. Age is a characteristic, not an outcome. Ask: "What am I trying to explain or predict?" That is the outcome.
Choosing the dependent variable before stating the research question The research question defines the variables, not the other way around. Always write the research question first. Then identify the exposure and outcome.
Using variables that are too vague to measure "Good health" cannot be measured. "Haemoglobin ≥11 g/dL" can. Make variables specific, observable, and measurable.
Mixing exposure and outcome in the same question A question like "Does malaria cause net use?" reverses causality. People buy nets because of malaria risk, not the other way around. Ensure temporal sequence: exposure must come before outcome.
Forgetting that one study may have several predictors Malaria is not caused only by net use. Age, season, housing, and immunity also matter. Identify the main predictor, but acknowledge confounding variables.

📝 Exam Tip: When asked to identify independent and dependent variables, always start by writing the research question clearly. Then ask: "What is the exposure/predictor?" (independent) and "What is the outcome?" (dependent). If you cannot write a clear question, you cannot identify the variables correctly.

Main Types of Data

Classifying data correctly is essential because the type of data determines how you summarize it, analyse it, and present it. Using the wrong statistical method for the wrong data type leads to meaningless or misleading results.

The First Question

Are the values categories/labels or numbers with mathematical meaning?

Categories → Qualitative. Numbers → Quantitative.

Qualitative Data (Categorical Data)

Qualitative data are grouped into categories or labels. Even if numbers are used as codes, they are not quantities, they are just labels.

  • Examples: Sex (male, female), blood group (A, B, AB, O), marital status (single, married, divorced), diagnosis (malaria, pneumonia, diarrhoea), ward (male, female, paediatric).
  • How to summarize: Counts and percentages. "40% of patients were diagnosed with malaria." "60% were female."
  • Statistical tests: Chi square test, Fisher's exact test (for comparing proportions between groups).
Two Forms of Qualitative Data
Type Definition Examples
Nominal Categories without natural order. You cannot say one category is "better" or "higher" than another. Blood group (A, B, AB, O), sex (male, female), diagnosis (malaria, TB, diabetes), ward (male, female, paediatric).
Ordinal Categories with natural order or rank. You can say one is "more" or "less" than another, but the gaps between categories are not equal. Pain severity (mild, moderate, severe), triage level (red, yellow, green), satisfaction (poor, fair, good, excellent), disease stage (Stage I, II, III, IV).

⚠️ Critical Distinction: In ordinal data, the order matters but the distance between categories is unknown. "Severe" pain is worse than "moderate," but we do not know if it is exactly twice as bad. You cannot calculate a meaningful average of ordinal data. You report the median or mode, not the mean.

Quantitative Data (Numerical Data)

Quantitative data are expressed as meaningful numbers. These numbers can be added, subtracted, averaged, and compared mathematically.

  • Examples: Age (28 years), weight (62 kg), temperature (38.5°C), pulse rate (72 bpm), haemoglobin (11.2 g/dL), blood pressure (120/80 mmHg).
  • How to summarize: Mean, median, standard deviation, range. "Average age was 32 years (SD 8.5)." "Median haemoglobin was 10.8 g/dL (range 7.2 to 14.1)."
  • Statistical tests: t test, ANOVA, correlation, regression (for comparing means or testing associations).
Two Forms of Quantitative Data
Type Definition Examples
Discrete Counted in whole numbers only. You cannot have a fraction of a count. There are gaps between possible values. Number of children (0, 1, 2, 3...), number of clinic visits (1, 2, 3...), number of tablets (1, 2, 3...), number of malaria episodes (0, 1, 2...).
Continuous Measured on a scale and can take any value within a range, including fractions and decimals. There are no gaps between possible values. Weight (62.3 kg), height (165.5 cm), temperature (38.7°C), time (2.5 hours), haemoglobin (11.2 g/dL), blood pressure (122/78 mmHg).

📝 Exam Tip: A common trap is students think "age" is continuous because it can be 28.5 years. But in many health datasets, age is recorded as whole years (28, 29, 30), making it discrete. However, if age is recorded in months, days, or as a decimal, it is continuous. In practice, age is often treated as continuous for analysis. The key is: "Can this value take any value on a scale, or only whole numbers?"

Decision Tree for Classifying Data

STEP 1: Are the values categories or numbers?

  • CATEGORIES → Qualitative Data
    • No natural order? → NOMINAL
    • Has natural order? → ORDINAL
  • NUMBERS → Quantitative Data
    • Counted in whole numbers? → DISCRETE
    • Measured on a scale? → CONTINUOUS

💡 Mnemonic for Data Types: "No Order? Nominal. Ordered? Ordinal. Discrete = Digits Counted. Continuous = Can be Cut into fractions."

Worked Examples: Classify the Variables
Variable Data Type Explanation
Sex Qualitative: nominal Categories (male, female) with no natural order. Male is not "higher" or "better" than female.
Triage level Qualitative: ordinal Categories (red, yellow, green) with a clear order: red = most urgent, green = least urgent. But the difference between red and yellow is not necessarily the same as between yellow and green.
Number of ANC visits Quantitative: discrete Counted in whole numbers (0, 1, 2, 3...). A patient cannot have 2.5 ANC visits.
Birth weight Quantitative: continuous Measured on a scale. A baby can weigh 2.85 kg, 3.1 kg, or any value in between. There are no gaps.
Haemoglobin level Quantitative: continuous Measured in g/dL. Can take any value within a physiological range (e.g., 7.2, 11.5, 14.3).
HIV test result Qualitative: nominal Categories (positive, negative) with no order. Positive is not "higher" than negative, they are just different states.
Waiting time in minutes Quantitative: continuous Can be 15 minutes, 15.5 minutes, or 15.75 minutes. Time is measured, not counted.
Ward of admission Qualitative: nominal Categories (male, female, paediatric, maternity) with no natural order.
Patient satisfaction Qualitative: ordinal Categories (poor, fair, good, excellent) with a clear order, but unequal gaps between categories.
Number of children in household Quantitative: discrete Counted in whole numbers. You cannot have 2.3 children.
How Data Type Guides Summary and Analysis

Choosing the wrong summary statistic is one of the most common errors in health data analysis. Here is how to match data type to the right summary:

Data Type Appropriate Summaries Graphs Examples
Qualitative: Nominal Frequencies, percentages, proportions, mode. Bar chart, pie chart. "Diagnosis: 40% malaria, 25% pneumonia, 20% diarrhoea, 15% other."
Qualitative: Ordinal Frequencies, percentages, median, mode. Never mean. Bar chart (ordered), stacked bar chart. "Pain: 10% mild, 40% moderate, 50% severe. Median = moderate."
Quantitative: Discrete Counts, mean, median, mode, range, standard deviation. Histogram, bar chart, box plot. "ANC visits: average of 4 visits (SD 1.2, range 1 to 8)."
Quantitative: Continuous Mean, median, standard deviation, range, interquartile range (IQR). Histogram, box plot, line graph, scatter plot. "Weight: median 62 kg (IQR 55 to 70, range 45 to 88)."

🚨 Critical Error to Avoid: Never calculate a mean for nominal or ordinal data. What is the "average blood group" of A, B, and O? It is meaningless. What is the "average satisfaction" of poor, fair, and good? Also meaningless. For nominal data, use percentages. For ordinal data, use the median or mode. For continuous data, use the mean (if normally distributed) or median (if skewed).

Mini Dataset for Practice

Here is a small dataset from a paediatric clinic. Your task: identify the variables and classify each by data type.

ID Age (years) Sex Temp (°C) RDT Result Visits
12F38.7Positive1
24M37.1Negative2
31F39.2Positive1
43M36.8Negative3
Worked Solution
Variable Meaning Data Type
Age Age in years Quantitative, discrete (recorded as whole numbers: 1, 2, 3, 4). Could be treated as continuous if measured in months or days.
Sex Male or female Qualitative, nominal (categories with no order).
Temperature Body temperature in °C Quantitative, continuous (measured on a scale: 36.8, 37.1, 38.7, 39.2, can take any value within a range).
RDT Result Rapid Diagnostic Test for malaria Qualitative, nominal (Positive / Negative, two categories with no natural order).
Visits Number of clinic visits Quantitative, discrete (counted in whole numbers: 1, 2, 3, cannot have 1.5 visits).

📝 Exam Tip: When classifying variables from a dataset, always look at how the data is recorded, not just what it represents. Age is "years lived" (continuous concept) but recorded as whole numbers (discrete in practice). Temperature is measured with a thermometer and can include decimals (continuous). RDT result is a label, not a number (qualitative).

Data Coding Basics

Coding converts answers or observations into organized numerical values for computer analysis. It is the bridge between the real world and the dataset.

Why Code Data?
  • Computers cannot analyse words like "male" and "female" directly, they need numbers.
  • Coding reduces data entry errors (typing "1" is faster and more consistent than typing "male" every time).
  • Coding allows statistical software to perform calculations and generate summaries automatically.
  • A well designed coding system makes the dataset cleaner and easier to share with other researchers.
Example: Coding Sex
Category Code Why This Code?
Male 1 Simple, consistent, easy to enter.
Female 2 Sequential numbering avoids confusion.
Example: Coding Diagnosis
Category Code Notes
Malaria 1 Most common diagnosis gets code 1 for efficiency.
Pneumonia 2 Sequential.
Diarrhoea 3 Sequential.
Other 99 "99" is a common convention for "other" or "not specified."
Missing / Unknown 88 or 999 Use a code that is clearly different from real data. Never leave blank, blanks cause errors.
Rules for Good Coding
  • Codes must be documented in a codebook. Every code must have a clear definition. Do not assume you will remember what "3" means in six months.
  • Never let codes change the meaning of a variable. If "1 = male" in one dataset, do not use "1 = female" in another dataset without clear documentation.
  • Use consistent coding across the entire study. All data collectors must use the same codes.
  • Good coding reduces data entry and analysis errors. Simple, logical codes are less likely to be entered incorrectly than long text strings.
  • Use standard missing value codes. Common conventions: 88, 99, 999, or -9. Choose one and document it. Never use "0" for missing, zero may be a real value (e.g., zero children).

⚠️ Common Coding Disaster: A researcher codes "male = 1, female = 2" but the data entry clerk sometimes types "M" and "F" instead. The statistical software treats "M" and "F" as text, not numbers, and excludes them from analysis. The result: 30% of the sample disappears. Solution: Use data validation rules in your entry software, and always check for unexpected text in numeric fields.

Writing Good Variables

A poorly defined variable leads to poor data, poor analysis, and poor decisions. A well defined variable is specific, observable, and measurable.

Weak Variable Why It Is Weak Improved Variable
"Health status" Vague. What does "health" mean? Physical? Mental? Self rated? "Haemoglobin level in g/dL" or "Self rated health: poor, fair, good, excellent."
"Good service" Subjective. "Good" means different things to different people. "Waiting time in minutes" or "Patient satisfaction score (1 to 5 scale)."
"Sick child" Too broad. What disease? What symptoms? How severe? "Child with confirmed malaria RDT positive and axillary temperature ≥37.5°C."
"Treatment improved" "Improved" is subjective. Improved by how much? Who decides? "Symptoms resolved by day 3 of treatment (Yes/No, confirmed by nurse assessment)."
"Temperature" Incomplete. Where was it measured? What unit? "Axillary temperature in °C, measured with digital thermometer after 5 minutes rest."

💡 The SMART Variable Rule: A good variable is Specific, Measurable, Achievable to collect, Relevant to the research question, and Time bound. Just like SMART goals, SMART variables lead to good research.

Class Exercise: Variable Detective

Task: In pairs, classify each item as qualitative or quantitative, then identify its subtype. For bonus marks, identify which variables could be dependent variables in a nursing research question.

  • Ward of admission
  • Number of children in household
  • Patient satisfaction: poor, fair, good
  • Haemoglobin level
  • HIV test result: positive or negative
  • Waiting time in minutes
Answer Key
Variable Data Type Could It Be a Dependent Variable?
Ward of admission Qualitative, nominal Rarely, usually a descriptive variable, not an outcome. Could be outcome in a study of triage decisions.
Number of children Quantitative, discrete Could be outcome in a study of family planning knowledge. More often an independent variable (predictor of maternal health).
Satisfaction level Qualitative, ordinal Yes, very common dependent variable. Example: "Does waiting time affect patient satisfaction?"
Haemoglobin level Quantitative, continuous Yes, very common dependent variable. Example: "Does iron supplementation improve haemoglobin?"
HIV test result Qualitative, nominal Yes, common dependent variable. Example: "Does circumcision reduce HIV incidence?"
Waiting time Quantitative, continuous Yes, common dependent variable. Example: "Does adding a second triage nurse reduce waiting time?"

📝 Exam Tip: Any variable can be a dependent variable, it depends on the research question. The same variable (e.g., "waiting time") can be an independent variable in one study ("Does waiting time affect satisfaction?") and a dependent variable in another ("Does adding staff reduce waiting time?"). The research question determines the role.

Group Activity: Build a Mini Study

Task: Each group chooses one nursing problem and fills the template below. This exercise connects all the concepts: research question, population, sample, variables, and data types.

📋 Mini Study Template
Element Your Group's Answer
Problem Example: High fever among children
Research Question Is sleeping under a mosquito net associated with reduced malaria among children under five?
Population Children under five attending OPD at Health Centre X
Sample 50 children selected during one clinic week using systematic random sampling
Independent Variable Sleeping under mosquito net (Yes / No), Qualitative, nominal
Dependent Variable Malaria RDT result (Positive / Negative), Qualitative, nominal
Confounding Variables Age, season, distance from breeding sites, household wealth, mother's education
How to Summarize Results Calculate percentage of RDT positive children among net users vs. non users. Compare using chi square test.
💡 Additional Example Problems for Group Work:
  • Problem: High post operative wound infection rate. IV: Hand hygiene compliance (Yes/No). DV: Wound infection (Yes/No).
  • Problem: Low immunisation coverage. IV: Mother's education level (None, Primary, Secondary+). DV: Child fully immunised (Yes/No).
  • Problem: Long clinic waiting times. IV: Number of nurses on duty. DV: Waiting time in minutes.
Worked Example: From Variable to Summary

Variable: Malaria RDT result among 50 children

Element Description
Data type Qualitative, nominal (Positive / Negative)
Summary method Count and percentage
Example result 18/50 positive = 36%
Graph Bar chart or pie chart showing positive vs. negative
Interpretation More than one third of the sampled children tested positive for malaria. This suggests a significant malaria burden in this population and warrants further investigation and intervention.

📝 Exam Tip: When interpreting a percentage, always mention both the number and the denominator. "36%" is meaningless without "18 out of 50." Also, always add a clinical or public health interpretation, do not just state the number. Explain what it means for patient care or community health.

Check for Understanding

Cover the answers and test yourself. If you can answer these clearly, you are ready for Day 7's exam!

What is the difference between a population and a sample?

A population is the entire group of interest (e.g., all first year nursing students). A sample is a smaller subset selected from that population for study (e.g., 50 randomly selected students). We use samples because studying the entire population is usually impossible, expensive, or time consuming.
Mnemonic: Population = People All Together. Sample = Selected Part.

Give two examples of nursing variables.
  • Patient variable: Blood pressure (continuous), age (discrete), sex (nominal).
  • Service variable: Waiting time in minutes (continuous), medicine availability (nominal: available / not available).
  • Disease variable: Diagnosis (nominal), severity (ordinal: mild, moderate, severe).

Always classify your examples by data type for extra marks.

In a study of net use and malaria, which variable is dependent?

Malaria status is the dependent variable (outcome). Net use is the independent variable (exposure/predictor). We are asking whether net use influences malaria status, so malaria is what we are trying to explain.
Remember: The dependent variable is the outcome. The independent variable is the exposure.

Is temperature qualitative or quantitative?

Quantitative, continuous. Temperature is measured on a scale (°C or °F) and can take any value within a range (e.g., 36.8°C, 37.1°C, 38.7°C). It has mathematical meaning, you can calculate an average temperature, and that average is meaningful.
If someone classifies temperature as "fever / no fever," it becomes qualitative (nominal). But the raw measurement is quantitative.

Is number of ANC visits discrete or continuous?

Discrete. ANC visits are counted in whole numbers (0, 1, 2, 3, 4...). A woman cannot attend 2.5 ANC visits. There are gaps between possible values.
Discrete = counted. Continuous = measured. This is the key distinction.

Why can you not calculate a mean for ordinal data?

Ordinal data has categories with a natural order (e.g., poor, fair, good, excellent), but the distance between categories is not equal or known. "Good" is better than "Fair," but we do not know if it is exactly twice as good. Calculating a mean assumes equal intervals, which ordinal data does not have. For ordinal data, use the median or mode instead.
This is a favourite exam question. Memorise the reason, not just the rule.

What is a codebook, and why is it important?

A codebook is a document that lists every variable, its definition, the codes used, and what each code means. It is important because:

  • It ensures consistency across multiple data collectors.
  • It prevents confusion when analysing data months later.
  • It allows other researchers to understand and verify your work.
  • It reduces data entry errors.

A dataset without a codebook is like a medicine bottle without a label, dangerous and unreliable.

Can a variable be both independent and dependent in different studies? Give an example.

Yes. A variable's role depends entirely on the research question.

  • As dependent: "Does iron supplementation improve haemoglobin?" (Haemoglobin = outcome)
  • As independent: "Does low haemoglobin increase the risk of post partum haemorrhage?" (Haemoglobin = predictor)

The research question determines the variable's role. There is no "inherent" independent or dependent variable.

What summary statistics would you use for each data type in a study of 100 ANC patients?
  • Nominal (e.g., HIV status): Frequencies and percentages. "12% were HIV positive."
  • Ordinal (e.g., satisfaction): Frequencies, percentages, median. "Median satisfaction = Good."
  • Discrete (e.g., number of visits): Mean, median, range, standard deviation. "Average 4.2 visits (SD 1.3)."
  • Continuous (e.g., haemoglobin): Mean or median, standard deviation, range, interquartile range. "Mean Hb 10.8 g/dL (SD 1.4, range 7.2 to 14.1)."

Use median for skewed continuous data (e.g., income, waiting time). Use mean for normally distributed data (e.g., height, weight in large samples).

Why is it important to define variables precisely before collecting data?

Precise variable definitions ensure that:

  • All data collectors record the same thing the same way (inter rater reliability).
  • The data answers the research question (validity).
  • The analysis is appropriate for the data type.
  • The results are reproducible by other researchers.
  • Clinical decisions based on the data are safe and evidence based.

Vague variables = vague data = vague conclusions = dangerous decisions.

Take Home Messages
  • Biostatistics helps nurses turn health data into decisions. It is not just numbers, it is the language of evidence.
  • A population is the full group of interest; a sample is the selected part. A good sample represents the population. A bad sample misleads everyone.
  • A variable is a characteristic that changes across observations. Constants do not vary and cannot explain differences.
  • Independent variables help explain dependent variables. The research question determines which is which.
  • Data type determines the correct summary and analysis. Nominal → percentages. Ordinal → median and percentages. Discrete → counts and means. Continuous → mean/median and spread.
  • Good coding and clear variable definitions prevent errors. A codebook is not optional, it is essential.
  • Never calculate a mean for nominal or ordinal data. It is mathematically meaningless and clinically misleading.
References
  • Grove, S. K., & Cipher, D. J. (2016). Statistics for Nursing Research: A Workbook for Evidence-Based Practice. Elsevier.
  • Heavey, E. (2018). Statistics for Nursing: A Practical Approach. Jones & Bartlett Learning.
  • Polit, D. F., & Beck, C. T. (2020). Nursing Research: Generating and Assessing Evidence for Nursing Practice. Wolters Kluwer.

Quick Quiz

Principles of Biostatistics Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Principles of Biostatistics Read More »

Sampling Methods, Disease Rates and Surveys

Sampling Methods, Disease Rates and Surveys

Sampling Methods, Disease Rates and Surveys
Learning Outcomes

By the end of this session, you should be able to:

  • Explain why sampling is used in epidemiology and biostatistics.
  • Describe simple random, systematic, stratified, and cluster sampling.
  • Identify strengths and limitations of non-probability sampling.
  • Calculate simple incidence, prevalence, ratios, and proportions.
  • Interpret disease measures for practical nursing decisions.

🧠 Core Question: "If we cannot study everyone, how do we select people fairly?" This is the central challenge of sampling. The answer determines whether your findings are trusted or dismissed.

Session 1: Introduction to Sampling
Why Do We Sample?

We sample because studying an entire population is usually impossible, impractical, or unnecessary. Here is why sampling is essential:

  • A population may be too large to study completely. You cannot interview all 40 million Ugandans about malaria knowledge. But you can interview 400 carefully selected people and learn a great deal.
  • Sampling saves time, money, and staff effort. A census (studying everyone) takes years and costs millions. A well designed survey takes weeks and costs thousands.
  • A good sample gives useful information about the wider group. If the sample truly represents the population, the findings apply to everyone not just those interviewed.
  • Poor sampling can produce misleading findings. If you only survey clinic attenders, you will overestimate service use. If you only survey urban areas, you will miss rural realities.

⚡ Golden Rule: Good sampling is not about studying many people only; it is about studying the right people. A sample of 80 well chosen mothers is more valuable than a sample of 800 poorly chosen ones.

From Population to Sample: The Flow
  • TARGET POPULATION: The full group we want to understand
  • SOURCE POPULATION: The accessible subset we can reach
  • SAMPLING FRAME: The list or method to identify eligible people
  • SELECTED SAMPLE: The smaller group actually studied
  • COLLECTED DATA: The information we analyse and interpret
Key Sampling Terms You Must Know
Term Definition & Example
Sampling Unit The individual person, household, school, or facility that is selected. Example: One mother with a child under one year.
Eligibility / Inclusion Criteria The rule that says who can be included in the study. Example: "Mothers with children aged 0-11 months living in the catchment area for at least 6 months."
Representativeness How well the sample reflects the characteristics of the whole population. A representative sample has the same age, sex, and socioeconomic distribution as the population.
Sampling Error The natural difference between a sample statistic and the true population value. Even a perfect random sample will not exactly match the population but the error is predictable and measurable. Larger samples have smaller sampling error.
Sampling Bias Systematic error caused by poor selection methods. Bias means the sample consistently overrepresents or underrepresents certain groups. Unlike sampling error, bias does not decrease with larger sample size.
📝 Exam Tip Sampling Error vs. Sampling Bias: This is a favourite exam distinction. Sampling error is random and natural it happens even with perfect methods. Sampling bias is systematic and caused by poor methods it happens because you selected the wrong way. Error can be reduced by increasing sample size. Bias can only be reduced by improving the sampling method.
Representativeness Matters: Weak vs. Strong Samples
❌ Weak Sample (Biased) ✅ Stronger Sample (Representative)
Only easy to reach households (those near the road) Includes different villages, including remote ones
Only clinic attenders (already using services) Uses household sampling to find non-attenders too
Excludes remote villages (no transport to reach them) Allocates resources to reach remote areas
Findings may be biased and not generalisable Findings are more credible and applicable to the whole population
The Sampling Frame

A sampling frame is the practical list or source from which sampling units are selected. It is the bridge between the theoretical population and the actual sample.

Examples of sampling frames:

  • Village register (list of all households).
  • School attendance list.
  • ANC (Antenatal Care) register at a health facility.
  • Facility list of all health centres in a district.
  • Household list from a recent census.

A weak frame is dangerous:

  • It may leave out eligible people (e.g., a village register that was last updated 3 years ago misses new households).
  • It may include ineligible people (e.g., the ANC register includes women who have since moved away or delivered).
  • It may be incomplete (e.g., no register exists for informal settlements).

⚠️ Critical Rule: A sample cannot be better than the frame used to select it. If your frame is missing half the population, your sample will miss them too no matter how fancy your randomisation method is.

Sampling Bias in Practice

🩺 Example: A survey about immunisation interviews only mothers who came to the clinic today.

  • Problem: It may miss mothers whose children are most likely to have missed vaccines the very group the survey wants to understand. Mothers who do not come to clinic may be the ones with transport barriers, misinformation, or cultural objections.
  • Likely effect: Coverage may appear higher than it truly is in the community. The survey concludes "90% coverage" when the real coverage is 60%.
  • Better approach: Sample households or use outreach lists across the entire catchment area. Include mothers who have never attended the clinic.
Mini Case: Immunisation Survey

🩺 The Situation: A health centre serves 1,200 mothers with children under one year. The team wants to interview 120 mothers about missed vaccines. They have village registers from 10 villages.

Task: How should they select mothers fairly?

Step by step thinking:

  • Population: 1,200 mothers with children under one year in the catchment area.
  • Sampling frame: Village registers from 10 villages. Check: Are the registers complete? Do they include all mothers? Are they up to date?
  • Sample size: 120 mothers (10% of the population a reasonable proportion for a survey).
  • Selection method:
    • Option A Simple Random Sampling: Combine all 10 village registers into one master list of 1,200 mothers. Number them 1 to 1,200. Use a random number table or computer to select 120 numbers. Interview those mothers.
    • Option B Systematic Sampling: Sampling interval = 1,200 ÷ 120 = 10. Choose a random start between 1 and 10 (e.g., 7). Then select every 10th mother: 7, 17, 27, 37... up to 1,197.
    • Option C Stratified Sampling: If some villages are much larger or poorer than others, divide the 120 sample proportionally by village size. Sample 12 from each village if equal, or proportionally if unequal. This ensures no village is overrepresented or underrepresented.
  • Possible bias: Village registers may miss mothers who recently moved in, or mothers who live in informal settlements not on any register. The team should plan for "non response" what if a selected mother is not home? Have a replacement rule (e.g., interview the next household) or revisit later.
Session 2: Probability Sampling

Core question: How do we give eligible people a known chance of selection?

Probability sampling means every eligible unit has a known, non zero chance of being selected. The selection uses a random or rule based method. This is the gold standard for surveys that need to estimate population levels.

📝 Exam Tip: In an exam, if you are asked to design a survey that estimates prevalence or compares groups, always choose a probability sampling method. Non-probability methods are only acceptable for exploratory or qualitative work.
Simple Random Sampling (SRS)

How it works:

  • Make a complete list of all eligible units in the population.
  • Number all eligible units (1 to N).
  • Use a random number table, lottery, or computer to select the required sample size.
  • Every unit has an equal chance of being selected.

Example: A health centre has a complete ANC register of 500 mothers. The team needs to select 100 for a satisfaction survey. They number the mothers 1-500, use a random number generator to pick 100 numbers, and interview those mothers.

Advantages:

  • Simple to understand and explain.
  • Every unit has equal chance no subgroup is favoured or ignored.
  • Statistical formulas work perfectly (standard errors, confidence intervals).

Limitations:

  • Requires a complete list of the population often unavailable in community settings.
  • Can be expensive and logistically difficult if selected units are scattered across a wide area.
  • May miss small subgroups by chance (e.g., only 2 elderly people selected in a sample of 100).
Systematic Sampling

How it works:

  • Calculate the sampling interval (k) = population size ÷ sample size.
  • Choose a random starting point between 1 and k.
  • Select every kth unit from the ordered list.
Systematic Sampling Formula: k = N ÷ n
Where N = population size, n = sample size, k = sampling interval

Example: 1,000 households ÷ 100 = k = 10. Random start = 3 (chosen between 1 and 10). Selected households: 3, 13, 23, 33, 43... 993.

Advantages:

  • Easier to implement than simple random sampling no need for a random number table.
  • Spreads the sample evenly across the list.
  • Often used in community surveys with household lists.

Limitations:

  • Hidden patterns in the list can introduce bias. Example: If a list is ordered by household head (male, female, male, female...) and k = 2, you might select only males or only females.
  • If the list has a periodic pattern that matches k, the sample is not random.
  • Less flexible than simple random sampling if you need to adjust mid study.

⚠️ Watch Out For: Always check the ordered list for hidden patterns before using systematic sampling. If the list is ordered by age, sex, or village in a repeating pattern, systematic sampling may be biased. In that case, use simple random or stratified sampling instead.

Stratified Sampling

How it works:

  • Divide the population into subgroups (strata) based on important characteristics. Strata should be mutually exclusive and collectively exhaustive (everyone fits in one and only one stratum).
  • Sample from each stratum separately. You can use simple random or systematic sampling within each stratum.
  • Combine the samples from all strata to form the total sample.

Common strata in health surveys:

  • Sex (male / female).
  • Age group (under 5, 5-14, 15-49, 50+).
  • Village or urban/rural.
  • School or facility type.
  • Socioeconomic status (wealth quintile).

Example: A district has 10 villages. 3 are near the main road (urban like), 7 are remote (rural). If you sample randomly, you might by chance select mostly road side villages. Instead, stratify by location: sample 30 from road side villages and 70 from remote villages, proportional to their population sizes.

Advantages:

  • Ensures representation of all important subgroups.
  • Allows separate analysis for each stratum (e.g., compare urban vs. rural vaccination rates).
  • More precise than simple random sampling when strata are internally similar but different from each other.

Limitations:

  • Requires knowledge of the population structure before sampling.
  • More complex to plan and analyse.
  • If strata are chosen poorly, it adds complexity without benefit.
Cluster Sampling

How it works:

  • First stage: Divide the population into natural groups called clusters (villages, schools, zones, parishes).
  • Second stage: Randomly select some clusters (not all).
  • Third stage: Within selected clusters, sample all individuals or a random subset.

Example: A district has 50 villages. You need to survey 500 households. Instead of listing all households in all 50 villages (impossible), you randomly select 10 villages, then survey 50 households in each selected village.

Advantages:

  • Practical and cheap for large, dispersed populations. No need for a complete list of all individuals.
  • Reduces travel costs interviewers stay in one area rather than travelling across the entire district.
  • Widely used in national surveys (DHS, MICS, SMART surveys).

Limitations:

  • People within a cluster tend to be similar (homogeneous). This increases sampling error compared to simple random sampling.
  • To compensate, you need a larger sample size than simple random sampling.
  • If clusters are selected poorly (e.g., only easy to reach villages), bias is introduced.
💡 Key Point: Cluster sampling is practical for community surveys, but people within a cluster may be similar. This is called the design effect (DEFF) a statistical penalty for using clusters. In exams, know that cluster sampling is cheaper but less precise than simple random sampling.
Choosing a Probability Method: Decision Guide
Situation Best Method Why
Complete list of all individuals exists Simple Random Sampling Every unit has equal chance; most statistically pure.
Ordered list exists; no hidden patterns Systematic Sampling Easy to implement; spreads sample evenly.
Subgroups differ in risk or access; equity matters Stratified Sampling Ensures all subgroups are represented; allows subgroup comparison.
Population is large and dispersed; no complete list of individuals Cluster Sampling Practical and cost-effective; only need lists of clusters (villages, schools).
Equity is important; want to compare urban vs. rural Stratified + Cluster First stratify by location, then cluster-sample within each stratum. Common in national surveys.
Class Activity: Pick the Method

Scenario A: 600 ANC clients are listed in a register; select 60.

Answer: Simple Random Sampling or Systematic Sampling. A complete list exists, so either works. Systematic might be easier: k = 600 ÷ 60 = 10. Random start between 1-10, then every 10th client.

Scenario B: A district has 12 villages; fieldwork can visit only 4 villages.

Answer: Cluster Sampling. Villages are the clusters. Randomly select 4 of 12 villages, then survey all or a sample of households within those 4. This is practical because visiting all 12 villages is too expensive.

Scenario C: The study must compare males and females fairly.

Answer: Stratified Sampling. Divide the population into male and female strata. Sample proportionally from each stratum. This guarantees enough males and females for statistical comparison.

Session 3: Non Probability Sampling

Core question: When random selection is not possible, what are the trade offs?

Non probability sampling means selection does not give every eligible person a known chance of being selected. It is faster, cheaper, and useful for hard to reach groups but it carries a higher risk of bias and weaker generalisation.

⚠️ Critical Rule: Use non probability sampling carefully and describe its limitations honestly in any report. Never claim that a convenience sample represents the whole population.

Convenience Sampling
  • Meaning: Select those who are easiest to reach.
  • When it is used: Pilot studies, practice exercises, rapid assessments, student research with limited time.
  • Example: A nursing student interviews patients sitting in the clinic waiting room because they are available right now.
  • Limitations:
    • May exclude the absent or remote the people who most need to be heard.
    • Can overrepresent service users people already in the clinic are not the same as people who never come.
    • Weak for population estimates. You cannot say "30% of the district has hypertension" based on a convenience sample of clinic patients.
Purposive Sampling
  • Meaning: Select people because they have specific knowledge or experience. The researcher deliberately chooses participants who can provide rich, relevant information.
  • When it is used: Qualitative research, key informant interviews, expert consultations, programme evaluations.
  • Example: Interviewing TB focal persons about case detection challenges, or interviewing traditional birth attendants about home delivery practices.
  • Quality depends on: Clear, transparent selection criteria. The researcher must explain why each person was chosen.
💡 Key Question for Purposive Sampling: "Who can provide the information needed?" Not "Who is easiest to find?" but "Who knows what we need to know?"
Quota Sampling
  • Meaning: Set required numbers for categories before data collection, then fill each quota with convenient respondents.
  • How it works: Decide you need 50 men and 50 women. Interview the first 50 men and 50 women you meet who fit the criteria.
  • Example: A rapid assessment in a market decides to interview 20 vendors, 20 shoppers, and 20 passers by.
  • Limitation: Quota controls numbers in groups, but it does not remove selection bias by itself. The interviewer still chooses which men and women to interview usually the easiest to approach. It is non random unless selection within quotas is randomised.
Snowball Sampling
  • Meaning: Initial participants help identify other eligible participants. Like a snowball rolling downhill it grows as it goes.
  • When it is used: Hidden or hard to reach populations: commercial sex workers, drug users, undocumented migrants, men who have sex with men, people with rare diseases.
  • Example: A researcher interviews one person living with HIV who then introduces 3 others in their support group, who each introduce more.
  • Limitations:
    • May overrepresent connected social networks. If the first participant only knows people from one church or one neighbourhood, the sample is biased.
    • Weak for estimating true population prevalence. You cannot calculate how common a behaviour is in the whole population from a snowball sample.
    • Ethical concerns: participants may feel pressured to recruit others.
Strengths and Limitations of Non Probability Sampling
Strengths Limitations
Fast and practical no need for complete lists. Unknown selection chance you cannot calculate the probability that any person was selected.
Useful for pilot studies and pretesting tools. Higher risk of bias certain groups are systematically overrepresented or excluded.
Good for qualitative depth rich, detailed information from key informants. Weak generalisation findings cannot be confidently applied to the wider population.
Can reach special groups that probability sampling cannot (hidden populations). Requires transparent reporting you must openly state the limitations in any report or publication.
Probability or Non Probability? Decision Guide
Use Probability When... Use Non Probability When...
Estimating prevalence or incidence in a population. Exploring experiences, beliefs, or perceptions (qualitative research).
Comparing population groups (e.g., urban vs. rural). Finding key informants with specialised knowledge.
Informing district or national planning. Pretesting questionnaires or data collection tools.
Generalisation to the wider population is important. Time, money, or lists are severely limited.
Group Task: Design a Sampling Plan

Study question: Why are some children missing immunisation?

Population: Mothers of children under one year in a catchment area.

Task: Choose one sampling method. State the sampling frame and one likely source of bias. Prepare a two minute explanation.

Example response:

  • Method: Stratified random sampling.
  • Strata: Urban and rural mothers (because access barriers differ).
  • Frame: Village registers for rural areas; facility ANC registers for urban areas.
  • Sample size: 100 mothers total 40 urban, 60 rural (proportional to population).
  • Likely bias: Village registers may miss mothers who recently moved in or who live in informal settlements not on any register. Urban ANC registers may miss mothers who never attended ANC.
  • Mitigation: Use community health workers to identify unregistered mothers. Plan for non response by selecting replacement households.
Session 4: Surveys and Disease Occurrence

Core question: After sampling, how do we count and interpret disease occurrence?

What Is a Survey?

A survey is a systematic method of collecting standard information from a defined group of people. It is not just a questionnaire it is a planned method for answering a health question.

What surveys can measure:

  • Health status: Prevalence of disease, nutritional status, disability.
  • Behaviour: Handwashing practices, net use, sexual behaviour, dietary habits.
  • Service use: ANC attendance, vaccination coverage, facility delivery rates.
  • Knowledge: Awareness of danger signs, understanding of disease transmission, health literacy.

What makes a good survey:

  • Clear questions every question has a purpose and is understood the same way by all respondents.
  • Appropriate sampling method matches the study question and population.
  • Quality control training interviewers, pretesting tools, supervising data collection, checking for completeness.
  • Practical decisions survey results should lead to action, not just sit in a report.
Basic Survey Steps
  1. DEFINE QUESTION
  2. DEFINE POPULATION
  3. CHOOSE SAMPLE
  4. COLLECT DATA
  5. ANALYSE MEASURES
  6. USE FINDINGS

⚠️ Critical Rule: Each step should be planned before fieldwork begins. The analysis plan should match the study question. Quality control starts before data collection not after you realise your questionnaire is confusing.

Counting Disease Occurrence: The Three Components

Every disease measure has three essential parts. Without all three, the number is meaningless:

Component What It Means
Numerator The number of cases or events counted. Example: 15 new malaria cases.
Denominator The population or group from which the cases came. Example: 120 hostel students.
Time Period The period during which new cases or events occurred. Example: During the month of July 2026.
📝 Exam Tip: A number becomes meaningful only when we know the denominator and time period. "15 malaria cases" tells you almost nothing. "15 new malaria cases among 120 students in July" tells you the risk is 12.5% actionable information.
Incidence

Definition: Incidence measures the number of new cases that develop in a population during a specific time period. It tells us about risk the probability that a healthy person will develop the disease.

Incidence Formula:
Incidence = New Cases During a Period ÷ Population at Risk
Usually expressed as a percentage or per 1,000 population

Key rules for incidence:

  • The numerator must include new cases only people who did not have the disease at the start of the period.
  • The denominator should include people at risk those who could have developed the disease. People who already have the disease should be excluded (unless studying recurrence).
  • Incidence must have a time period. Without time, it is not incidence it is just a count.

Incidence Example

Scenario: 15 new malaria cases occurred among 120 hostel students during the month of July.

Incidence = 15 ÷ 120 = 0.125 = 12.5%

Interpretation: About 13 out of every 100 students developed malaria during July. This is the risk of getting malaria in that hostel during that month.

Nursing action: A 12.5% monthly incidence is high. The nurse should investigate: Are nets being used? Is there stagnant water near the hostel? Are students seeking treatment promptly? Consider a mass net distribution or environmental clean up.

Prevalence

Definition: Prevalence measures the total number of existing cases (both new and old) in a population at a specific point in time (point prevalence) or over a period (period prevalence). It tells us about burden how widespread the condition is.

Prevalence Formula:
Prevalence = Existing Cases at a Point or Period ÷ Total Population
Includes people who already have the condition + new cases

Key rules for prevalence:

  • The numerator includes all existing cases both new and old. A person who has had diabetes for 10 years is still counted in prevalence.
  • The denominator is the total population not just those at risk. Everyone in the population could potentially be a case.
  • Prevalence is especially useful for chronic conditions (hypertension, diabetes, HIV) and for planning services (how many beds, drugs, or clinics are needed?).

Prevalence Example

Scenario: During a community screening, 18 adults out of 80 screened have high blood pressure.

Prevalence = 18 ÷ 80 = 0.225 = 22.5%

Interpretation: About 23 in every 100 screened adults had high blood pressure readings. This is the burden of hypertension in the screened population.

Nursing action: A 22.5% prevalence suggests hypertension is common in this community. The nurse should: confirm readings with repeat measurements, counsel on lifestyle, refer high readings, plan follow up clinics, and consider community education on diet and exercise.

Incidence vs. Prevalence: Side by Side
Feature Incidence Prevalence
Counts New cases only All existing cases (new + old)
Measures Risk how likely is a healthy person to get the disease? Burden how widespread is the disease right now?
Needs time? Yes must specify the time period Can be a point in time or a period
Denominator Population at risk (those who could get the disease) Total population (everyone in the group)
Best for Outbreaks, acute diseases, studying causes Chronic diseases, planning services, resource allocation
Example "10% of students got malaria in July" "22.5% of adults screened had high BP"
📝 Exam Tip: Incidence = new. Prevalence = existing. Incidence asks "How many got sick?" Prevalence asks "How many are sick?"
Incidence vs. Prevalence: The Relationship

Prevalence depends on both incidence and duration of disease:

Prevalence ≈ Incidence × Average Duration of Disease

This means:

  • If incidence is high and duration is long → prevalence is very high (e.g., HIV in high burden areas before ART scale up many new infections, and people lived with the disease for years).
  • If incidence is high but duration is short → prevalence may be lower than expected (e.g., acute diarrhoea many new cases, but they recover within 3-5 days, so at any single point, few people are sick).
  • If incidence drops but duration stays long → prevalence may remain high for years (e.g., diabetes fewer new cases due to prevention, but existing cases live for decades with the condition).
💡 Nursing Implication: A high prevalence of hypertension does not necessarily mean many new cases are appearing. It may mean people are living longer with the disease (good chronic care) or that detection has improved. Always ask: "Is prevalence high because of new cases, long duration, or better detection?"
Common Calculation Mistakes
Mistake Why It Is Wrong How to Fix It
Using total cases when the measure requires new cases only This gives prevalence, not incidence Check: Are these new cases or all cases?
Forgetting the time period for incidence Without time, it is not a rate it is just a count Always state: "per month," "per year," "during the outbreak"
Using the wrong denominator Comparing apples to oranges Ensure denominator matches the population at risk
Reporting a percentage without explaining what it means Numbers without context are useless Always interpret: "X out of every 100..."
Comparing groups without considering group size 10 cases in 50 vs. 10 cases in 500 are very different Always calculate rates, not just counts
📝 Exam Tip: Always ask yourself: "Numerator of what? Denominator among whom? During what time?" If you cannot answer all three, your measure is incomplete. Examiners love to give you a number and ask "What is missing?" the answer is usually the denominator or the time period.
Session 5: Ratios, Proportions and Practice

Core question: How do we calculate and interpret basic disease measures accurately?

Ratio

A ratio compares two quantities where the numerator is not necessarily part of the denominator. The two quantities are independent.

Ratio = One Quantity ÷ Another Quantity

Key feature: The numerator and denominator are separate groups. One is not a subset of the other.

Example: 30 male patients and 60 female patients attended the clinic.
Male to female ratio = 30 : 60 = 1 : 2
Interpretation: There is 1 male patient for every 2 female patients.

Other nursing examples:

  • Nurse to patient ratio: 5 nurses for 50 patients = 1 : 10
  • Doctor to nurse ratio: 2 doctors for 10 nurses = 1 : 5
  • Bed to population ratio: 100 beds for 50,000 people = 1 : 500
⚠️ Important: A ratio does NOT tell you what fraction of the whole has a condition. It only compares two groups. "1:2 male to female ratio" does not mean 33% are male it means for every male, there are 2 females.
Proportion

A proportion compares a part to the whole, where the numerator is included in the denominator. It is always expressed as a decimal or percentage.

Proportion = Part ÷ Whole

Key feature: The numerator is a subset of the denominator. The result ranges from 0 to 1 (or 0% to 100%).

Example: 20 diarrhoea cases among 200 children screened.
Proportion = 20 ÷ 200 = 0.10 = 10%
Interpretation: 10% of screened children had diarrhoea.

Other nursing examples:

  • Proportion of ANC attendees who are HIV positive: 15 HIV+ women ÷ 200 ANC attendees = 7.5%
  • Proportion of deliveries by caesarean section: 30 C-sections ÷ 300 deliveries = 10%
  • Proportion of children fully immunised: 85 fully immunised ÷ 100 children = 85%
📝 Exam Tip Ratio vs. Proportion: This is a classic exam trap. Ratio = compares two separate groups (male:female). Proportion = part of a whole (males ÷ total patients). If the numerator is included in the denominator, it is a proportion. If not, it is a ratio.
Rate

A rate describes how fast events occur in a population over time. It is the most informative measure in epidemiology because it combines count, population, and time.

Rate = Occurrence ÷ Population at Risk over Time

Key features:

  • Rates must state the time period.
  • Rates allow comparison between groups of different sizes.
  • Incidence is the most common type of rate.
  • Rates are often expressed "per 1,000" or "per 100,000" for rare diseases.

Example: 40 new malaria cases in a village of 500 children during August.
Rate = 40 ÷ 500 = 0.08 = 8% per month (or 80 per 1,000 per month).

Why "per 1,000" is useful: For rare diseases, percentages are tiny and hard to interpret. Saying "0.002% got Ebola" is confusing. Saying "2 cases per 100,000 population" is clear and standard for international comparison.

💡 Mnemonic Rate vs. Ratio vs. Proportion: "Rate has Time, Ratio has Two groups, Proportion has Part of whole." = R-T, R-T, P-P
Worked Example: Village Diarrhoea

Scenario: A village has 500 children under five. During August, 40 new diarrhoea cases are recorded. At the end of August, 25 children still have diarrhoea.

Question: Calculate August incidence and end of month prevalence.

Answer:

Measure Formula Calculation Result Interpretation
Incidence New cases ÷ Population at risk 40 ÷ 500 8% 8 out of every 100 children developed diarrhoea in August
Prevalence Existing cases ÷ Total population 25 ÷ 500 5% 5 out of every 100 children had diarrhoea at the end of August

Why the difference?

  • Incidence (8%) counts all new cases that occurred during August including those who already recovered by month end.
  • Prevalence (5%) counts only those still sick at the end of the month.
  • The gap (8% − 5% = 3%) represents children who got diarrhoea but recovered before month end.
⚡ Key Principle: Incidence tells you how fast the disease is spreading. Prevalence tells you how much disease is in the community right now. For acute diseases (diarrhoea, malaria attack), incidence is usually higher than point prevalence because people recover quickly. For chronic diseases (diabetes, hypertension), prevalence is much higher than incidence because cases accumulate over years.
Practical Exercise: Calculate and Interpret
Scenario Measure Calculation Interpretation
Village A: 30 new malaria cases among 300 people in July Incidence (risk) 30 ÷ 300 = 10% High risk — 1 in 10 people got malaria that month. Needs urgent vector control.
Village B: 30 new malaria cases among 1,500 people in July Incidence (risk) 30 ÷ 1,500 = 2% Lower risk — 1 in 50 people got malaria. Still monitor, but less urgent.
Health centre: 18 high BP readings among 80 adults screened Prevalence 18 ÷ 80 = 22.5% About 1 in 4 screened adults has high BP. Plan NCD follow up clinic.
Clinic register: 12 males and 36 females attended ANC education Ratio 12:36 = 1:3 For every male companion, 3 female companions attended. Male involvement is low.
📝 Exam Tip: Same number of cases can mean very different risk when denominators differ. Village A and Village B both had 30 cases but Village A's risk was 5 times higher. This is why denominators are essential. In an exam, never just compare counts. Always calculate rates.
Interpreting the Numbers: A 5 Step Framework

When you calculate a disease measure, follow these five steps to interpret it meaningfully for public health action:

Step What to Do Example
1. Name the measure Is it incidence, prevalence, ratio, or proportion? "This is an incidence measure..."
2. State the group Among whom was it calculated? "...among hostel students..."
3. State the time When or over what period? "...during the month of July..."
4. Translate to plain language "X out of every 100..." "...about 13 out of every 100 students..."
5. Suggest one action What should be done? "...suggests the need for improved net use and environmental clean up."

Full example interpretation:
"The incidence of malaria among hostel students was 12.5% during July. This means about 13 out of every 100 students developed malaria that month. This high risk suggests the need for improved insecticide treated net use, removal of stagnant water near the hostel, and prompt testing and treatment of febrile students."

📝 Exam Tip: In exams, marks are awarded for calculation AND interpretation. Many students calculate correctly but lose marks because they do not explain what the number means in plain language. Always finish with: "This means..." and "Therefore, we should..."
Group Assignment Brief

Task: Choose one health problem: malaria, diarrhoea, missed immunisation, or hypertension.

  • Define the population and sampling frame.
  • Choose a sampling method and justify it.
  • Create a small dataset and calculate one disease measure (incidence, prevalence, ratio, or proportion).
  • Present findings in three minutes using the 5 step interpretation framework.

Assessment focus: Clear sampling plan, correct calculation, and practical interpretation.

Example response structure:
"We studied [population] using [sampling method] because [justification]. Our sampling frame was [frame]. We found a [measure] of [X%], which means [interpretation]. Therefore, we recommend [action]. One limitation is [bias/limitation]."

Quick Self Check
Question Answer
Why is sampling used in epidemiology? Because populations are often too large, expensive, or time consuming to study completely. A good sample gives valid information about the wider group without the cost of a census.
What is the difference between stratified and quota sampling? Stratified sampling uses random selection within each stratum (probability method valid for generalisation). Quota sampling sets numbers for categories but uses convenience selection within quotas (non probability method faster but biased).
When would cluster sampling be practical? When the population is large and dispersed, no complete list of individuals exists, and travel costs must be minimised (e.g., national immunisation coverage surveys, DHS, MICS).
How is incidence different from prevalence? Incidence counts new cases over a time period (measures risk "how many got sick?"). Prevalence counts all existing cases at a point or period (measures burden "how many are sick?").
Why must every rate have a denominator and time period? Without a denominator, you cannot compare groups of different sizes. Without a time period, you cannot distinguish rapid outbreaks from slow trends. A rate without both is just a number not actionable evidence.
What is the formula for systematic sampling? k = N ÷ n (population size ÷ sample size = sampling interval). Choose random start between 1 and k, then select every kth unit.
What is sampling bias, and how is it different from sampling error? Sampling bias is systematic error caused by poor selection methods it does not decrease with larger sample size. Sampling error is random natural variation between sample and population it decreases with larger sample size.
Why is a sampling frame important? A sample cannot be better than the frame used to select it. If the frame is incomplete, outdated, or excludes certain groups, the sample will be biased no matter how random the selection method is.
When should you use non probability sampling? For exploratory research, qualitative depth, pilot studies, pretesting tools, or when studying hard to reach/hidden populations where probability sampling is impossible.
How do you interpret a ratio of 1:3 male to female? For every 1 male, there are 3 females. This does NOT mean 25% are male (that would be a proportion). It only compares the two groups.
References
  • Gordis, L. (2014). Epidemiology. Elsevier Saunders.
  • Bonita, R., Beaglehole, R., & Kjellström, T. (2006). Basic Epidemiology. World Health Organization.
  • Webb, P., Bain, C., & Page, A. (2017). Essential Epidemiology: An Introduction for Students and Health Professionals. Cambridge University Press.

Quick Quiz

Sampling Methods Quiz

Epidemiology and Biostatistics - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Sampling Methods, Disease Rates and Surveys Read More »

Nursing Management question approach

Finalists Pre-entry Exam

Quick Quiz

Pre-entry Finalist Exam

Aptitude - mobile-friendly and focused practice.

Privacy: Your details are used only for quiz tracking and certificates.

Finalists Pre-entry Exam Read More »

disaster phases

Requirements for disaster preparedness.

Requirements for Disaster Preparedness
What Are Requirements for Disaster Preparedness?
Definition

Requirements for disaster preparedness are all the conditions, resources, plans, systems, and capacities that must be in place before a disaster happens so that a community, hospital, or nation can respond effectively when disaster strikes.

Simple Explanation

Think of requirements like the ingredients you need before cooking a meal. If you wait until guests arrive to look for food, salt, and firewood, you will fail. Disaster preparedness means gathering everything you need in advance so that when the disaster comes, you are ready to act immediately.

Why Requirements Matter

"Failure to prepare is preparing to fail."

If a hospital does not have the requirements in place:

  • There will be no beds for sudden casualties
  • There will be no clean water when pipes break
  • Nurses will not know what to do
  • Patients will die from preventable causes
Categories of Preparedness Requirements

Disaster preparedness requirements can be grouped into nine major categories:

REQUIREMENTS FOR DISASTER PREPAREDNESS

  • PLANNING AND DOCUMENTATION REQUIREMENTS
  • RESOURCE AND SUPPLY REQUIREMENTS
  • PERSONNEL AND TRAINING REQUIREMENTS
  • INFRASTRUCTURE AND FACILITY REQUIREMENTS
  • COMMUNICATION AND INFORMATION REQUIREMENTS
  • EARLY WARNING SYSTEM REQUIREMENTS
  • FINANCIAL REQUIREMENTS
  • LEGAL AND POLICY REQUIREMENTS
  • COMMUNITY AND PUBLIC EDUCATION REQUIREMENTS
SECTION B: DETAILED REQUIREMENTS BY CATEGORY
CATEGORY 1: PLANNING AND DOCUMENTATION REQUIREMENTS
A Written Disaster Preparedness Plan
What Is Required

Every institution — from a national government to a village health center — must have a written disaster preparedness plan. This plan must be:

  • Written down — not just in someone's memory
  • Realistic — it must match actual local risks and resources
  • Simple — everyone who reads it must understand it
  • Updated regularly — at least once per year, or after every disaster
What the Plan Must Include
Section What It Must Describe
Risk assessment What disasters are likely to happen here? (floods, landslides, epidemics, fires)
Vulnerable areas Which buildings, roads, and communities are most at risk?
Vulnerable populations Who will need extra help? (elderly, disabled, pregnant women, children, orphans)
Roles and responsibilities Who does what when disaster strikes?
Command structure Who is in charge? Who makes decisions?
Activation triggers When does the plan start? (e.g., "When more than 10 casualties arrive" or "When flood water reaches 1 meter")
Evacuation procedures Where do people go? What routes do they use?
Resource inventory What supplies are available? Where are they stored?
Communication protocols Who calls whom? What radio frequencies? What phone numbers?
Alternative care sites Where will patients go if the hospital is damaged or full?
Deactivation criteria When is the disaster over? When does normal work resume?
Ugandan Example: A Health Centre III in Bududa must have a written plan that says:
• Risk: Landslides during heavy rain (March-May, September-November)
• Trigger: When district disaster office issues red alert OR when cracks appear on local slopes
• Nurse's role: Triage at entrance, activate community health worker network, open emergency drug box
• Evacuation: Move patients to the church on the hill if the health center is threatened
• Resources: Emergency box contains 50 IV fluids, 100 bandages, 20 cannulas, ORS packets, chlorine tablets
Contingency Plans
What Is Required

A contingency plan is a "Plan B" — what to do if the main plan fails.

Examples of Contingency Requirements
  • If the main hospital is flooded, where is the backup hospital?
  • If the main nurse is sick, who is the deputy?
  • If phones fail, how do we communicate? (runners, radios, drums)
  • If roads are blocked, how do we transport patients? (motorbikes, boats, foot stretchers)
Standard Operating Procedures (SOPs)
What Is Required

SOPs are step-by-step instructions for specific tasks. They remove guesswork during emergencies.

Required SOPs for Disaster Preparedness
SOP What It Must Describe
Triage SOP Exactly how to sort patients; who does it; where; how long per patient
Evacuation SOP How to move patients from wards; who carries whom; what equipment to take
Fire response SOP How to use extinguishers; when to evacuate; how to move bedbound patients
Infection control SOP How to isolate patients; PPE use; waste disposal during outbreaks
Mortuary management SOP How to handle dead bodies safely; documentation; family notification
Chemical spill SOP How to decontaminate; who does it; where is the decontamination area
Mapping and Documentation
What Is Required

Maps showing:

  • Hazard zones (floodplains, landslide areas, fault lines)
  • Safe evacuation routes
  • Location of safe buildings (churches, schools, strong houses)
  • Location of water sources, fuel stores, and medical supplies

Records of:

  • Vulnerable households (elderly living alone, disabled persons, pregnant women)
  • Community resources (who has a vehicle, a generator, a boat, first aid training)
  • Staff contact details (updated every 3 months)
CATEGORY 2: RESOURCE AND SUPPLY REQUIREMENTS
Emergency Stockpiles (Supplies)
What Is Required

Hospitals, health centers, and communities must keep emergency supplies stored safely, accessible, and checked regularly.

Required Medical Supplies
Category Specific Items Required Minimum Quantity Guidance
Airway and breathing Oropharyngeal airways, nasal airways, ambu bags, oxygen masks, oxygen cylinders At least 10 of each size
Bleeding control Gauze rolls, gauze pads, triangular bandages, tourniquets, hemostatic dressings 100+ units for mass casualty
IV access and fluids Cannulas (all sizes), IV giving sets, normal saline, Ringer's lactate, dextrose 50-100 bags depending on facility size
Drugs Adrenaline, atropine, diazepam, antibiotics, analgesics (morphine, paracetamol), tetanus toxoid, ORS Sufficient for 48-72 hours without resupply
PPE Gloves, surgical masks, N95 respirators, gowns, goggles, aprons, boots At least 1 week supply for all staff
Wound care Antiseptic (chlorhexidine, iodine), sutures, sterile gloves, dressing packs, plaster Mass casualty quantities
Diagnostic Thermometers, sphygmomanometers, stethoscopes, pulse oximeters, glucometers, weighing scales Backup equipment if main units fail
Obstetric Delivery kits, misoprostol, oxytocin, umbilical cord ties, resuscitation masks for newborns Protect pregnant women in disasters
Sanitation Chlorine tablets, soap, disinfectant, hand sanitizer, water containers, latrine slabs For facility and community use
Required Non-Medical Supplies
Item Why It Is Required
Clean water storage When pipes break, you need stored water for drinking, cleaning, and sterilization
Fuel (petrol/diesel) For generators, ambulances, and water pumps
Firewood or gas For cooking in shelters or for sterilizing equipment
Blankets and mattresses For patients in shelters or on floors
Stretchers For moving patients; improvised ones if manufactured ones are few
Tarpaulins and tents For temporary shelters and treatment areas
Plastic sheeting For waterproofing floors, making partitions, protecting supplies
Ropes and ties For securing tents, makeshift stretchers, and supplies
Torchlights and batteries For power outages
Candles and matches Backup lighting (with fire safety precautions)
Dustbins and liners For safe waste disposal
Body bags For respectful and safe handling of deceased
Emergency Kits
What Is Required

Pre-packed kits that can be grabbed and moved quickly.

Required Kits
Kit Name Contents Purpose
Health kit Towel, soap, toothbrush, toothpaste, comb, bandages Personal hygiene for displaced people
First aid kit Gauze, tape, antiseptic, bandages, scissors, gloves Basic wound care
Medicine kit Antibiotics, painkillers, ORS, antacids, anti-parasitics Common medical needs
School kit Paper, pencils, ruler, scissors, crayons Continue education for children in shelters
Baby kit Diapers, clothes, blankets, pins Care for infants in disasters
Sewing kit Fabric, needles, thread, buttons Repair clothes and basic items
Cleaning kit Buckets, bleach, brushes, soap, gloves Maintain sanitation
Water and Sanitation Requirements
Requirement Standard
Water storage At least 15 liters per person per day for drinking and hygiene
Water treatment Chlorine tablets or boiling capability
Latrines One latrine per 20 people in emergency shelters
Handwashing stations Available at every medical area and shelter
Waste disposal Safe burial pit or incinerator for medical waste
Bathing privacy Separate areas for men and women
Food Requirements
Requirement Standard
Emergency food stock 3-7 days supply for staff and patients
Nutritional supplements Ready-to-use therapeutic food (RUTF) for malnourished children
Infant feeding Breastfeeding support; formula only if absolutely necessary (risk of contamination)
Cooking fuel Safe fuel for preparing food for large groups
Power and Fuel Requirements
Requirement Purpose
Backup generator Keep lights, oxygen concentrators, and refrigerators working
Fuel stock Enough for 48-72 hours of generator use
Solar power Reliable backup that does not need fuel
Battery backups For critical equipment like monitors
Candles and lanterns Last-resort lighting (with fire safety measures)
CATEGORY 3: PERSONNEL AND TRAINING REQUIREMENTS
Adequate Staffing Numbers
What Is Required
  • Enough staff to handle sudden surge in patients
  • A call list of off-duty staff who can return quickly
  • Clear roles so everyone knows their job
Staffing Requirements by Facility Level
Facility Minimum Staffing Requirement for Preparedness
Health Centre II 2 nurses on duty; 2 community health workers on call; 1 support staff
Health Centre III 3 nurses; 1 clinical officer; 2 midwives; 3 support staff; on-call team of 5
Health Centre IV/ District Hospital Full emergency team; on-call surgical, pediatric, and maternity staff; 20+ nurses available within 2 hours
Regional Referral Hospital Mass casualty team ready 24/7; ability to triple nursing staff within 4 hours
Defined Roles and Job Descriptions
What Is Required

Every person must know exactly what they do in a disaster. There should be no confusion.

Required Role Assignments
Role Person Responsible Specific Duties
Incident Commander Senior doctor or hospital administrator Overall decision-making; liaison with government and NGOs
Triage Officer Senior nurse or emergency nurse Sort all incoming patients; assign colors; direct flow
Resuscitation Team Leader Doctor or senior clinical officer Manage RED tag area; prioritize life-saving interventions
Nursing Coordinator Senior nursing officer Assign nurses to areas; manage shift rotation; ensure rest
Pharmacy Coordinator Pharmacist Manage drug stock; ration scarce supplies; request resupply
Infection Control Officer Infection prevention nurse Enforce hand hygiene; manage isolation; track disease spread
Documentation Officer Records officer or assigned nurse Maintain patient registers; track admissions and deaths
Security Coordinator Hospital security head + police liaison Control crowds; protect staff; secure supplies
Logistics/Supply Officer Administrator or stores manager Track resources; arrange transport; manage donations
Mental Health Lead Psychiatric nurse or counselor Support traumatized patients and staff
Community Liaison Community health nurse Communicate with families; coordinate community health workers
Training Requirements
What Is Required

All staff must be trained before the disaster. Training during the disaster is too late.

Required Training Programs
Training Topic Who Must Be Trained How Often
Basic Life Support (BLS) All nurses, doctors, clinical officers Every 2 years
Advanced Cardiac Life Support (ACLS) Emergency and ICU nurses Every 2 years
Triage All nurses and emergency personnel Annually
First Aid All staff including support staff Annually
Fire Safety All hospital staff Every 6 months
Infection Prevention and Control (IPC) All clinical staff Annually
PPE Use All staff Before every outbreak; annually
Disaster Plan Orientation All new staff + all staff refresher At hiring; annually
Mass Casualty Management Emergency department staff; all nurses Annually
Psychological First Aid All nurses and counselors Annually
Emergency Obstetric Care Midwives and maternity nurses Every 2 years
Decontamination Staff near industrial areas or handling outbreaks Annually
Drills and Simulation Exercises
What Is Required
  • Tabletop exercises: Sitting around a table discussing "What if a bus crashes outside?"
  • Functional drills: Practicing one part of the plan (e.g., evacuation of one ward)
  • Full-scale drills: Complete simulation with actors, fake injuries, and timed responses
Drill Requirements
Type of Drill Frequency Purpose
Fire drill Every 3 months Practice evacuation; test alarms
Evacuation drill Every 6 months Move patients to safe areas quickly
Mass casualty drill Every 12 months Test triage, treatment, and coordination
Disease outbreak drill Every 12 months Test isolation, PPE, and reporting
Tabletop discussion Every 3 months Review plans; identify gaps
Personal Preparedness of Staff
What Is Required

Nurses and other staff cannot help patients if their own families are in danger.

Required Personal Preparedness
Requirement Why It Matters
Family emergency plan Staff know their families are safe, so they can focus on work
Emergency contact list Hospital can reach staff quickly
Physical fitness Disaster response is physically demanding
Mental health readiness Staff must cope with extreme stress
Updated skills certification CPR, first aid, triage certificates current
CATEGORY 4: INFRASTRUCTURE AND FACILITY REQUIREMENTS
Structural Safety of Buildings
What Is Required

Health facilities must be built to withstand the disasters common in their area.

Structural Requirements by Hazard
Hazard Structural Requirement
Earthquake Reinforced concrete; flexible joints; lightweight roofs; secured heavy equipment
Flood Elevated construction; water-resistant ground floor; raised electrical systems
Landslide Built on stable, flat ground; away from steep slopes; retaining walls if needed
Cyclone/Strong wind Hurricane straps; strong roof anchoring; shatter-resistant windows
Fire Fire-resistant materials; multiple exits; fire doors; smoke alarms; sprinklers
Safe Room and Shelter Requirements
Requirement Description
Safe room A reinforced room where staff and patients can shelter during extreme wind or earthquake
Emergency shelter A designated strong building nearby (church, school) if the hospital must be evacuated
Alternative care site A pre-identified location to treat patients if the hospital is damaged or full
Assembly point An open area where people gather after evacuation for headcount
Functional Areas During Disaster

The hospital or health center must designate and prepare:

Area Requirements
Triage area Covered space near entrance; clear signage; colored tags available; fast access
Resuscitation area Multiple beds/mats; oxygen; suction; good lighting; emergency drugs within arm's reach
Treatment area Space for YELLOW and GREEN patients; wound care supplies; splints
Isolation area Separate room with separate entrance for infectious diseases; negative pressure if possible
Morgue/deceased area Cool, secure, dignified; away from patient areas; body bags available
Command center Room with communication equipment, maps, plans, and decision-makers
Staff rest area Place for exhausted staff to eat, drink, and rest briefly
Utilities and Engineering Requirements
Utility Preparedness Requirement
Water Storage tanks holding at least 48 hours of water; backup borehole or rainwater collection
Electricity Generator with automatic start; solar backup; fuel stored safely
Medical gases Oxygen cylinders stored safely; backup supply; pressure gauges checked
Sewage Backup system if main sewer fails; portable latrines ready
Waste management Incinerator or burial pit functional; extra bins and liners stockpiled
Communication Landline, mobile network, radio (VHF/UHF), satellite phone if possible
Transportation Requirements
Requirement Purpose
Functional ambulance With fuel, driver, and basic emergency equipment always ready
Alternative transport Identified vehicles in community (trucks, private cars, motorcycles) for mass casualty
Boat access In flood-prone and lakeside areas, boats for rescue and evacuation
Clear access roads Hospital entrance must remain clear; no parking that blocks ambulances
Helicopter landing zone At referral hospitals, marked and maintained for air ambulance
CATEGORY 5: COMMUNICATION AND INFORMATION REQUIREMENTS
Communication Systems
What Is Required

Multiple ways to communicate, because one system often fails during disaster.

Required Communication Methods
Method Purpose Backup If This Fails
Mobile phones Daily coordination; calling staff Radio or runners
Radio (VHF/UHF) When cell towers fail; long-distance Satellite phone or drums/whistles
Satellite phone Remote areas; total network failure Physical messengers
Internet/Email Sending documents, maps, reports Radio or physical delivery
Public address system Announcements inside hospital Megaphone or word-of-mouth
Whatsapp/SMS groups Quick staff alerts Radio call
Physical messengers When all technology fails
Communication Protocols
What Is Required
  • Chain of command: Who reports to whom?
  • Standard reporting forms: Pre-printed forms for casualty numbers, supply needs, disease alerts
  • Media protocol: Who is allowed to speak to the press? What information can be shared?
  • Family notification system: How do families know where their relatives are?
Information Management
What Is Required
  • Patient tracking system: Know where every patient is, their condition, and their identity
  • Resource tracking: Know what supplies remain, what is used, what is needed
  • Situation reports (SITREPs): Regular updates sent to district and national levels
  • Maps: Updated maps of the area, hazard zones, and facility layout
CATEGORY 6: EARLY WARNING SYSTEM REQUIREMENTS
What Is an Early Warning System?

An early warning system is a chain of actions that detects a coming disaster and alerts people in time to act.

Requirements for Effective Early Warning
Requirement Description
Detection capability Technology and people watching for danger (weather stations, river gauges, disease surveillance, slope monitors)
Data analysis Experts who interpret the data and predict what will happen
Warning dissemination Systems to spread the warning quickly to everyone at risk (radio, SMS, sirens, community drums, church bells, messenger runners)
Community understanding People must know what the warning means and what to do when they hear it
Response capacity The community must be able to act on the warning (evacuation routes, shelters, transport)
Specific Early Warning Requirements by Disaster
Disaster Warning Requirement
Flood River level gauges; rain gauges; weather forecasts; community flood watchers
Landslide Slope monitoring (crack meters); rain intensity measurement; community spotters
Drought Seasonal rainfall forecasts; vegetation index monitoring; livestock condition tracking
Epidemic Disease surveillance; lab confirmation; community health worker reports; school absenteeism tracking
Cyclone/Storm Satellite monitoring; national meteorological alerts; community radio networks
Fire Smoke detectors; fire patrols during dry season; community fire watchers
The "Last Mile" Requirement

"A warning that does not reach the village is not a warning."

It is not enough to detect danger at the national level. The warning must reach:

  • The grandmother in the remote village with no radio
  • The farmer in the field with no phone
  • The child walking home from school

Requirements for last-mile warning:

  • Community messengers with bicycles or motorcycles
  • Church bells, mosque loudspeakers, and drums
  • Community health workers who go door-to-door
  • Visual signals (flags, colored lights) for those who cannot hear
CATEGORY 7: FINANCIAL REQUIREMENTS
Budget for Disaster Preparedness
What Is Required
  • Dedicated budget line for disaster preparedness in every health facility and district
  • Money must be available before the disaster, not just after
  • Funds for: Stockpiling supplies, Training and drills, Equipment maintenance, Plan development and printing
Emergency Funds
Requirement Purpose
Rapid access fund Small cash amount that the nurse-in-charge can spend immediately without waiting for approval (e.g., to buy fuel, hire a motorcycle, buy emergency water)
District contingency fund Money held at district level for emergency procurement
National disaster fund Government fund for large-scale disasters
Insurance Property and vehicle insurance for health facilities
Resource Mobilization Plan
What Is Required

A written plan for how to get more money and resources when the disaster exceeds local capacity:

  • Which NGOs to contact (Red Cross, UNICEF, WHO)
  • How to request government emergency funds
  • How to accept and account for donations
  • How to document spending for accountability
CATEGORY 8: LEGAL AND POLICY REQUIREMENTS
Legal Framework
What Is Required
  • National disaster management law: Uganda has the National Policy for Disaster Preparedness and Management and works under the Office of the Prime Minister
  • Mandatory reporting laws: Health workers must report certain diseases and disasters
  • Building codes: Laws requiring safe construction
  • Environmental protection laws: Laws protecting wetlands, forests, and water sources
Policy Requirements
Policy What It Must Cover
Disaster management policy Roles of all ministries; coordination structures; funding mechanisms
Health sector emergency policy How the Ministry of Health responds; deployment of medical teams; use of private facilities
Infection control policy Isolation requirements; PPE standards; waste management
Staff safety policy Protection for health workers; compensation if injured; right to refuse unsafe work
Patient confidentiality policy How to protect patient information during mass casualty events
Agreements and Memoranda of Understanding (MoUs)
What Is Required

Written agreements between:

  • Hospital and ambulance services
  • Hospital and blood bank
  • Hospital and nearby facilities for patient transfer
  • Hospital and police/fire services
  • Hospital and NGOs for supply support
  • District and national government for resource sharing
CATEGORY 9: COMMUNITY AND PUBLIC EDUCATION REQUIREMENTS
Community Preparedness Requirements

The community itself must be prepared, not just the health facility.

Element Requirement
Community disaster committee Elected or appointed group responsible for local preparedness
Community risk map Map drawn by community showing hazards, safe routes, and safe buildings
Family emergency plan Every family knows where to go, how to communicate, and what to bring
Community early warning Local system for alerting everyone (drums, whistles, runners)
Community first aid team Trained community members who can help before professionals arrive
Community resource inventory List of local assets (vehicles, strong buildings, water sources, trained people)
Public Education Requirements
Topic Target Audience Method
Warning signs of disasters Entire community Radio, community meetings, school programs
Evacuation routes and shelters All households Maps posted in public places, household visits
First aid and home care Community health workers, families Training sessions, demonstrations
Safe water and hygiene All households Home visits, school programs, drama
Immunization importance Parents, caregivers Health talks, radio spots
Fire safety Market vendors, school staff, families Demonstrations, inspections
Road safety Drivers, boda-boda riders, pedestrians Community policing, radio, school programs
School Preparedness Requirements

Schools are critical because children are vulnerable and schools often serve as emergency shelters.

Requirement Standard
School disaster plan Every school must have a written plan
Evacuation drills At least once per term
Safe construction Schools in earthquake/landslide zones must be reinforced
Lightning conductors Required in all schools in lightning-prone areas
First aid kits Available in every school
Trained teachers At least 2 teachers per school trained in first aid
Safe shelter function If school is a designated shelter, it must have water, latrines, and kitchen facilities
SECTION C: SPECIFIC REQUIREMENTS FOR A DISASTER PREPAREDNESS PLAN

A comprehensive disaster preparedness plan must meet these specific requirements:

Early Warning Systems
  • Requirement: Design and implement effective systems to detect and communicate impending disasters.
  • Details: Use appropriate technology (rain gauges, river sensors, disease surveillance); ensure warnings reach every community member; test the system regularly.
Evacuation and Victim Support
  • Requirement: Plan for safe evacuation and relocation of people.
  • Details: Marked evacuation routes; designated safe buildings; transportation for elderly and disabled; pre-positioned supplies at shelters; registration system at shelters.
Stockpiling Essential Supplies
  • Requirement: Store food, water, medicine, and other critical resources.
  • Details: 48-72 hour minimum supply for health facilities; 3-7 day supply for communities; regular rotation to prevent expiry; secure, accessible, dry storage.
Disaster Drills and Exercises
  • Requirement: Practice response and evacuation procedures.
  • Details: Tabletop exercises every 3 months; functional drills every 6 months; full-scale drills annually; after-action reviews to improve the plan.
Action Plans for Response and Recovery
  • Requirement: Written plans for what happens immediately after impact and during long-term recovery.
  • Details: Specific steps for first 24 hours; patient surge management; referral pathways; rehabilitation and reconstruction roles.
Personal Protective Equipment (PPE)
  • Requirement: Ensure protective gear is available for all emergency personnel.
  • Details: Correct sizes for all staff; training in proper use; stockpile for at least one week; disposal plan for contaminated PPE.
Environmental Controls
  • Requirement: Implement measures to prevent secondary environmental disasters.
  • Details: Safe chemical storage; protected water sources; controlled waste disposal; fire prevention.
Coordination Mechanisms
  • Requirement: Establish how different agencies will work together.
  • Details: Cluster system (health, water, shelter, etc.); regular coordination meetings; shared communication channels; joint assessment teams.
SECTION D: REQUIREMENTS FOR A DISASTER PREPAREDNESS TEAM

A disaster preparedness team must meet these requirements:

Knowledge of the Disaster Management Plan

Every team member must read, understand, and be able to implement the plan.

Regular Plan Updates

The team must review and update the disaster plan at least annually and after every drill or real event.

Development of Educational Materials
  • Create materials in local languages appropriate for local literacy levels.
  • Use pictures, diagrams, and oral methods for non-literate communities.
Organization of Drills

Schedule and conduct drills in collaboration with government and non-governmental organizations.

Records of Vulnerable Populations

Maintain updated, confidential records of:

  • Elderly living alone
  • People with disabilities
  • Pregnant women and new mothers
  • Orphans and vulnerable children
  • People with chronic diseases (HIV, diabetes, hypertension, TB)
  • Households without transport
Awareness of Community Resources
  • Know what the community has: buildings, vehicles, tools, skills, water sources.
  • Know how to access these resources quickly.
Promotion of Building Codes and Land/Water Management
  • Advocate for safe construction.
  • Advocate for wetland and forest protection.
Education for Disaster-Prone Areas

Provide targeted education to communities in high-risk zones.

Safety Precaution Instructions

Teach the public about:

  • Storing emergency supplies at home
  • Basic first aid
  • Preparing for injuries
  • Family communication plans
Public Communication Systems
  • Ensure the community has ways to receive information (radio, community meetings, SMS).
  • Ensure the team can send information out quickly.
Early Warning Utilization
  • Know how the early warning system works.
  • Know how to activate it and how to respond to it.
Immediate Hazard Mitigation

After a disaster, the team must be able to quickly identify and address new dangers (damaged buildings, contaminated water, downed power lines).

SECTION E: HOSPITAL AND HEALTH FACILITY PREPAREDNESS CHECKLIST

Minimum Requirements Every Health Facility Must Meet:

Administrative Requirements
  • [ ] Written disaster preparedness plan posted in a visible place
  • [ ] Disaster management committee established with named members
  • [ ] Clear chain of command with contact numbers
  • [ ] Memoranda of understanding with nearby hospitals, ambulance services, and police
  • [ ] Emergency budget line or rapid access fund
  • [ ] Insurance coverage for facility and vehicles
Staff Requirements
  • [ ] All staff oriented to the disaster plan within 1 month of hiring
  • [ ] At least 60% of clinical staff certified in BLS/First Aid
  • [ ] Triage training completed by all emergency and maternity staff
  • [ ] On-call roster established and tested monthly
  • [ ] Staff family emergency plans encouraged and supported
Supply Requirements
  • [ ] Emergency drug box checked and restocked monthly
  • [ ] Emergency delivery kit available and complete
  • [ ] PPE stock for minimum 1 week for all staff
  • [ ] Water storage for 48 hours minimum
  • [ ] Fuel for generator for 48 hours minimum
  • [ ] Alternative lighting (torches, candles with safety measures)
  • [ ] Body bags available (minimum 10)
  • [ ] Stretchers available (minimum 2)
  • [ ] Blankets and mattresses for floor patients
Infrastructure Requirements
  • [ ] Fire extinguishers present and inspected every 6 months
  • [ ] Smoke alarms installed and tested
  • [ ] Clear evacuation routes marked with illuminated signs
  • [ ] Assembly point identified and known to all staff
  • [ ] Backup generator tested monthly
  • [ ] Safe room or alternative care site identified
  • [ ] Isolation room or area designated
  • [ ] Mortuary area or dignified space for deceased identified
Communication Requirements
  • [ ] Updated staff contact list (tested monthly)
  • [ ] Functional radio or alternative communication
  • [ ] Emergency phone numbers posted (ambulance, fire, police, district health office, OPM)
  • [ ] Public address system or megaphone available
  • [ ] Pre-printed reporting forms available
Documentation Requirements
  • [ ] Triage tags available (colored cards or tape)
  • [ ] Patient registers for mass casualty events
  • [ ] Maps of facility and local area posted
  • [ ] Vulnerable household list updated quarterly
  • [ ] Resource inventory updated quarterly
SECTION F: NURSING-SPECIFIC PREPAREDNESS REQUIREMENTS

What Every Nurse Must Personally Have Ready:

Professional Requirements
Requirement Why It Matters
Current BLS/First Aid certification You may be the only one who can resuscitate a patient
Knowledge of facility disaster plan You must know your role without reading the plan during chaos
Participation in at least one drill per year Muscle memory saves time when seconds count
Familiarity with triage colors and categories You may be the triage officer
PPE competency Putting on PPE correctly prevents infection; taking it off incorrectly causes infection
Emergency drug knowledge Know doses and indications for adrenaline, atropine, diazepam, morphine
Personal Requirements
Requirement Why It Matters
Family emergency plan If your family is safe, you can focus on patients
Emergency contact card Hospital can reach you; you can reach family
Physical fitness Disaster response requires lifting, running, long hours
Mental resilience strategies You will see suffering; you must cope to continue helping
"Go bag" ready A bag with spare uniform, comfortable shoes, snacks, water bottle, flashlight, personal medications, and copies of certifications
SECTION G: MNEMONICS AND MEMORY AIDS

Mnemonic 1: "PLAN FIRST" — Core Preparedness Requirements

  • Personnel trained and ready
  • Legal framework in place
  • Alternative sites identified
  • Necessary supplies stockpiled
  • Finances available rapidly
  • Information systems working
  • Response plan written and known
  • Shelter and evacuation routes ready
  • Training and drills conducted regularly

Mnemonic 2: "READY NOW" — Facility Checklist

  • Resources inventoried
  • Emergency contacts updated
  • Alternative power tested
  • Drills practiced
  • Yield (supplies) rotated before expiry
  • Notification systems functional
  • On-call staff confirmed
  • Water and sanitation secured

Mnemonic 3: "WARN-ME" — Early Warning Requirements

  • Watch (detection systems)
  • Analyze (expert interpretation)
  • Reach (dissemination to all)
  • Notify (clear message)
  • Make understood (community education)
  • Enable action (evacuation routes and shelters ready)

Mnemonic 4: "SUPPLIES" — Stockpile Categories

  • Safety equipment (PPE, helmets, gloves)
  • Utilities backup (fuel, generator, water)
  • Pharmaceuticals (drugs, vaccines, ORS)
  • Patient transport (stretchers, blankets, splints)
  • Lifesaving tools (airways, suction, oxygen)
  • Infection control (chlorine, soap, waste bins)
  • Emergency kits (first aid, delivery, baby)
  • Sanitation (latrines, water containers)
SECTION H: EXAM PREPARATION
Common Exam Questions

Q1: List five requirements for disaster preparedness.
Answer: A written disaster preparedness plan; trained personnel; stockpiled essential supplies; functional communication systems; early warning systems; adequate infrastructure; financial resources; legal framework; community education. (Any five)

Q2: Why is it important to have a written disaster preparedness plan?
Answer: It ensures everyone knows their roles and responsibilities; it provides clear procedures during chaos; it can be reviewed and improved; it prevents panic and confusion; it meets institutional and legal standards.

Q3: What should be included in an emergency stockpile at a health center?
Answer: IV fluids and cannulas; emergency drugs (adrenaline, antibiotics, analgesics, ORS); bandages and wound care supplies; PPE (gloves, masks, gowns); oxygen and airway equipment; delivery kits; body bags; water storage; fuel for generator; blankets and stretchers.

Q4: Describe the "last mile" problem in early warning systems.
Answer: The last mile refers to the gap between national warning systems and the actual people at risk. A warning is useless if it does not reach the remote village, the elderly person without a radio, or the farmer in the field. Solutions include community messengers, drums, church bells, and door-to-door alerts by community health workers.

Q5: What are the requirements for a disaster preparedness team?
Answer: Knowledge of the disaster plan; ability to update the plan; skills to develop educational materials; capacity to organize drills; updated records of vulnerable populations; awareness of community resources; ability to promote building codes and land management; skills to teach safety precautions; access to communication systems; ability to use early warnings; capacity for immediate hazard mitigation.

Q6: Why must health facility staff have family emergency plans?
Answer: If staff are worried about their own families' safety during a disaster, they cannot focus on patient care. Personal preparedness ensures staff are mentally present and available to work.

Q7: List three infrastructure requirements for a hospital in a flood-prone area.
Answer: Elevated construction to prevent water entry; raised electrical systems and generators; water-resistant ground floor materials; protected drug storage; clear drainage around the building; alternative care site on higher ground. (Any three)

Q8: What communication methods should a health facility have ready for disaster?
Answer: Mobile phones; radio (VHF/UHF); satellite phone if possible; public address system or megaphone; physical messenger system as ultimate backup; updated staff contact lists; pre-printed reporting forms.

Clinical Scenarios
Scenario A: Health Centre III Preparedness Audit

You are the senior nurse at a Health Centre III in a landslide-prone district. The district health officer is coming to audit your preparedness.

Questions:

  • What documents must you show? Written disaster plan, updated staff contact list, drill records, stockpile inventory, vulnerable household list
  • What supplies will the auditor check? Emergency drug box, PPE stock, water storage, generator fuel, delivery kits, body bags
  • What infrastructure will be inspected? Fire extinguishers, evacuation routes, generator function, building structural safety, alternative shelter identification
  • What training records must you have? BLS certificates, triage training logs, drill attendance sheets, fire safety orientation
Scenario B: Preparing for the Rainy Season in Kasese

Your district hospital is in Kasese, which floods every year. It is now one month before the rainy season.

Questions:

  • What requirements must you check now? Generator and fuel; elevated storage for drugs; sandbags for entrance; alternative care site on upper floor; boat access arrangement; staff call list tested; ORS and cholera supplies pre-positioned; community warning system tested
  • What early warning requirements do you need? River level gauge readings; communication with meteorological authority; community flood watcher network; SMS alert system for staff
  • What coordination requirements exist? MoU with upstream health facilities for patient transfer; coordination with Uganda Red Cross for shelter; police contact for evacuation security
Scenario C: Ebola Preparedness in a Border District

Your district shares a border with DRC. There is an Ebola outbreak across the border. You must prepare your hospital.

Questions:

  • What supply requirements are specific to viral hemorrhagic fever? PPE — coveralls, gloves, boots, goggles, aprons; chlorine for disinfection; sharp containers; body bags with Ebola specifications; isolation tents or rooms
  • What training requirements are urgent? PPE donning and doffing; safe injection practices; safe burial protocols; patient isolation procedures; contact tracing
  • What infrastructure requirements must be met? Isolation ward with separate entrance; dedicated latrine for isolation area; dedicated burial team space; staff changing and decontamination area
  • What communication requirements exist? Direct line to Ministry of Health and UVRI; community rumor control system; safe burial team coordination; media protocol
Key Points to Remember
  • Disaster preparedness requirements are everything needed BEFORE disaster strikes
  • There are nine categories: planning, resources, personnel, infrastructure, communication, early warning, finance, legal/policy, and community education
  • Every health facility must have a written disaster plan that is realistic, simple, and updated regularly
  • Stockpiles must cover 48-72 hours minimum and be checked regularly for expiry
  • Staff must be trained, certified, and personally prepared
  • Infrastructure must withstand local hazards and maintain utilities during disaster
  • Communication requires multiple methods because one system always fails
  • Early warning is useless without the "last mile" — reaching every person at risk
  • Financial requirements include dedicated budgets and rapid-access emergency funds
  • Legal requirements ensure coordination, reporting, and accountability
  • Community education ensures the public knows what to do and can help themselves
  • Nurses must meet both professional and personal preparedness requirements
References
  • World Health Organization (WHO). (2020). Health Emergency and Disaster Risk Management Framework.
  • Ministry of Health, Uganda. National Technical Guidelines for Integrated Disease Surveillance and Response.
  • Office of the Prime Minister, Uganda. National Policy for Disaster Preparedness and Management.
  • International Council of Nurses (ICN). (2019). Core Competencies in Disaster Nursing Version 2.0.

Requirements for disaster preparedness. Read More »

Want notes in PDF? Join our classes!!

Send us a message on WhatsApp
0726113908

Scroll to Top
Enable Notifications OK No thanks