Skip to content

Probability sampling methods

Module items

Assignment

Sample assignment

Reading

Learning outcomes

  1. Compare and contrast probability sampling methods
    1. simple random sample,
    2. systematic random sample,
    3. stratified random sample, and
    4. multi-stage cluster sample
  2. Learn the factors affecting sample size and quality

Four types of probability sample

  1. Simple random sample
  2. Systematic random sample
  3. Stratified random sample
  4. Multi-stage cluster random sample

(1) Simple random sampling (1)

  • Each individual has an equal probability of selection.
  • List all individual and number them consecutively.
  • Use random numbers to select individuals.

(1) Simple random sampling (2)

  • There are five steps in selecting a simple random sample: 1. Obtain a complete list of individuals. 2. Give each individual a unique number starting at one 3. Decide on the required sample size. 4. Select numbers for the sample size from a table of random numbers. 5. Select the individuals that correspond to the randomly chosen numbers.
  • Number Name
    • 1 Adams, H.
    • 2 Anderson, J.
    • 3 Baker, E.
    • 4 Bradsley, W.
    • 5 Bradley, P.
    • 6 Carra, A.
    • 7 Cidoni, G.
    • 8 Daperis, D.
    • 9 Devlin, B.
    • 10 Eastside, R.
    • 11 Einhorn, B.
    • 12 Falconer, T.
    • 13 Felton, B.
    • 14 Garratt, S.
    • 15 Gelder, H.
    • 16 Hamilton, I.
    • 17 Hartnell, W.
  • Number Name
    • 18 Iulianotti, G.
    • 19 Ivono, V.
    • 20 Jabornik, T.
    • 21 Jacobs, B.
    • 22 Kennedy, G.
    • 23 Kassem, S.
    • 24 Ladd, F.
    • 25 Lamb, A.
    • 26 Mand, R.
    • 27 McIlraith, W.
    • 28 Natoli, P.
    • 29 Newman, L.
    • 30 Ooi, W.
    • 31 Oppenheim, F.
    • 32 Peters, P.
    • 33 Palmer, T.
    • 34 Quick, B.
  • Number Name
    • 35 Quinn, J.
    • 36 Reddan, R.
    • 37 Risteski, B.
    • 38 Sawers, R.
    • 39 Saunders, M.
    • 40 Tarrant, A.
    • 41 Thomas, G.
    • 42 Uttay, E.
    • 43 Usher, V.
    • 44 Varley, E.
    • 45 Van Rooy, P.
    • 46 Walters, J.
    • 47 West, W.
    • 48 Yates, R.
    • 49 Wyatt, R.
    • 50 Zappulla, T.

(2) Systematic random sampling (1)

  • Same as simple random sampling, except:
    • Choosing individuals from a random starting point, and choosing every nth individual (e.g., every 3rd individual).

(2) Systematic random sampling (2)

  • Step 1: Determine population size
    • 100
  • Step 2: Determine sample size required
    • 20
  • Step 3: Calculate sampling fraction (population ÷ sample)
    • = 100 ÷ 20
    • = 5
  • Step 4: Select random starting point within first 5 cases
    • e.g. 03
  • Step 5: Select every 5th case

    • = sample of 20
  • Number Name

    • 1 Adams, H.
    • 2 Anderson, J.
    • 3 Baker, E.
    • 4 Bradsley, W.
    • 5 Bradley, P.
    • 6 Carra, A.
    • 7 Cidoni, G.
    • 8 Daperis, D.
    • 9 Devlin, B.
    • 10 Eastside, R.
    • 11 Einhorn, B.
    • 12 Falconer, T.
    • 13 Felton, B.
    • 14 Garratt, S.
    • 15 Gelder, H.
    • 16 Hamilton, I.
    • 17 Hartnell, W.
  • Number Name
    • 18 Iulianotti, G.
    • 19 Ivono, V.
    • 20 Jabornik, T.
    • 21 Jacobs, B.
    • 22 Kennedy, G.
    • 23 Kassem, S.
    • 24 Ladd, F.
    • 25 Lamb, A.
    • 26 Mand, R.
    • 27 McIlraith, W.
    • 28 Natoli, P.
    • 29 Newman, L.
    • 30 Ooi, W.
    • 31 Oppenheim, F.
    • 32 Peters, P.
    • 33 Palmer, T.
    • 34 Quick, B.
  • Number Name
    • 35 Quinn, J.
    • 36 Reddan, R.
    • 37 Risteski, B.
    • 38 Sawers, R.
    • 39 Saunders, M.
    • 40 Tarrant, A.
    • 41 Thomas, G.
    • 42 Uttay, E.
    • 43 Usher, V.
    • 44 Varley, E.
    • 45 Van Rooy, P.
    • 46 Walters, J.
    • 47 West, W.
    • 48 Yates, R.
    • 49 Wyatt, R.
    • 50 Zappulla, T.

(3) Stratified random sampling (1)

  • Starting point is to categorize population into “strata” (relevant divisions, or departments of companies for example).
  • So, the sample can be proportionately representative of each stratum.
  • Then, randomly select within each stratum as for a simple or systematic random sample.

Discussion question (3)

  • discussion

    • Imagine we are conducting research on students at a college with 10,000 students. We know their college affiliation, and this information is important for our research, as we want all colleges to be represented. What is the limitation of using simple or systematic random sampling for this research?
  • Students Population % Simple or systematic random sample

    • Humanities 1,800 18% 150
    • Social sciences 1,200 12% 146
    • Pure sciences 2,600 26% 289
    • Applied sciences 1,800 18% 164
    • Engineering 2,600 26% 251
    • TOTAL 10,000 100% 1,000

(4) Multi-stage cluster sampling (1)

  • The most complex, expensive, and representative sampling.
  • First, divide population into groups (clusters) of units, like states.
  • Sub-clusters (sub-groups) can then be sampled from these clusters.
  • Now randomly select individuals from each (sub)cluster.
  • Collect data from each cluster of units, consecutively.

(4) Multi-stage cluster sampling (2)

  • This technique of obtaining a final sample involves drawing several different samples. 1. Divide the city into areas (e.g. electorates, census districts). These areas are called clusters. 2. Select a simple or systematic random sampling of these clusters. 3. Obtain a list of smaller areas (e.g. blocks) within the selected clusters. 4. Select a simple or systematic random sampling of smaller areas (e.g. blocks) within each of the clusters selected at stage 2. 5. For each selected block obtain a list of addresses of households (enumeration). 6. Select a random of addresses within the selected blocks. 7. At each selected address select an individual to participate in the sample.

(4) Multi-stage cluster sampling (3)

  • Table 2. Population and Sample According to Neighborhoods of Gebze (+ 20 ages)
  • Neighborhood Population Sample
    • Adem Yavuz 8,463 30
    • Arapçeşme 25,286 95
    • Barış 6,829 30
    • Beylikbağı 8,504 30
    • Cumhuriyet 5,210 30
    • Gaziler 17,999 66
    • Güzeller 16,507 65
    • Hacıhalil 8,599 35
    • Hürriyet 12,593 44
    • İnönü 9,327 36
    • İstasyon 14,683 55
  • Neighborhood Population Sample
    • Kirazpınar 4,174 30
    • Köşklü Çeşme 17,142 65
    • Mevlana 16,675 59
    • Mimar Sinan 11,444 41
    • Mustafapaşa 15,283 58
    • Osman Yılmaz 27,709 111
    • Sultan Orhan 11,626 43
    • Tatlıkuyu 9,894 37
    • Ulus 9,933 35
    • Yavuz Selim 11,734 42
    • Yenikent 12,830 47
  • Total 183,313 1,072
  • The 12.31.2010 dated Address-based Population Registration System of Turkey was used to design the sample, which was estimated with 95% confidence and a margin of error of 0.03, and quotas were set according to the population of 22 neighborhoods. In the case of the sub-sample of neighborhoods with less than 30 participants, extra interviews were conducted in order to reach the target of 30 respondents for each neighborhood.

Factors affecting sample size and quality

  • Time and cost
    • After a certain point (n=1,000), increasing sample size produces less noticeable gains in precision.
    • Very large samples are decreasingly cost-efficient (Hazelrigg, 2004).
  • Heterogeneity of the population
    • The more varied the population is, the larger the sample will have to be.
  • Non-response
    • Response rate = % of sample who agree to participate (or % who provide usable data).
    • Responders and non-responders may differ on a crucial variable.

Sample size

  • We need to decide how much error we are prepared to tolerate and how certain we want to be about our generalizations from the sample.
  • Two statistical concepts, sampling error and confidence intervals, help us specify the degree of accuracy we achieve, and the concept of confidence level specifies the level of confidence we can have in our generalizations.

Calculating sample size

  • Population Size 95% (5.0%) 95% (3.5%) 95% (2.5%) 95% (1.0%) 99% (5.0%) 99% (3.5%) 99% (2.5%) 99% (1.0%)
  • 10 10 10 10 10 10 10 10 10
  • 20 19 20 20 20 19 20 20 20
  • 30 28 29 29 30 29 29 30 30
  • 50 44 47 48 50 47 48 49 50
  • 75 63 69 72 74 67 71 73 75
  • 100 80 89 94 99 87 93 96 99
  • 150 108 126 137 148 122 135 142 149
  • 200 132 160 177 196 154 174 186 198
  • 250 152 190 215 244 182 211 229 246
  • 300 169 217 251 291 207 246 270 295
  • 400 196 265 318 384 250 309 348 391
  • 500 217 306 377 475 285 365 421 485
  • 600 234 340 432 565 315 416 489 579
  • 700 248 370 481 653 341 462 554 672
  • 800 260 396 526 739 363 503 615 763
  • 1,000 278 440 606 906 399 575 726 943
  • 1,200 291 474 674 1067 427 636 826 1119
  • 1,500 306 515 759 1297 460 712 958 1376
  • 2,000 322 563 869 1655 498 807 1140 1785
  • 2,500 333 597 952 1984 524 878 1287 2172
  • 3,500 346 641 1068 2565 558 976 1509 2890
  • 5,000 357 678 1176 3288 586 1065 1733 3842
  • 7,500 365 710 1275 4212 609 1146 1960 5164
  • 10,000 370 727 1332 4899 622 1192 2096 6238
  • 25,000 378 760 1448 6939 646 1284 2398 9968
  • 50,000 381 772 1491 8057 654 1318 2519 12449
  • 75,000 382 776 1506 8514 657 1329 2562 13576
  • 100,000 383 778 1513 8763 659 1335 2584 14220
  • 250,000 384 782 1527 9249 661 1346 2624 15546
  • 500,000 384 783 1532 9423 662 1350 2638 16045
  • 1,000,000 384 783 1534 9513 663 1351 2645 16306
  • 2,500,000 384 784 1536 9567 663 1352 2649 16467
  • 10,000,000 384 784 1536 9595 663 1353 2652 16549
  • 100,000,000 384 784 1537 9603 663 1353 2652 16574
  • 300,000,000 384 784 1537 9604 663 1353 2652 16576
  • Assume that a survey has a margin of error of plus or minus 2.5 percent at a 95 percent level of confidence.
  • If the survey were conducted 100 times, the data would be within 2.5 points above or below the percentage reported in 95 of the 100 surveys.
  • Multi-cluster sampling example (Google Sheets file)