Data Retention Audit Sample Size Estimator

Estimate a statistically based sample size for auditing a finite population of data-retention records or disposition decisions. The calculator uses a proportion sample-size model, then applies a finite-population correction so the required sample decreases appropriately when the population is not extremely large.

This is useful for internal quality checks on retention labels, deletion approvals, record classifications, or policy adherence. It does not establish a legally required audit size, and it cannot compensate for biased sampling or poor population definitions. Choose confidence, margin of error, and expected exception rate to match the purpose of the review. When little is known about the likely exception rate, 50% is conservative because it produces the largest sample for a given confidence level and margin of error.

Calculator inputs

items
%
%
%
Result
estimated audit sample
Approx. z-score
Initial infinite-population sample
Sample as % of population

1. Enter the population
Provide the total number of retention records or decisions eligible for the audit.

2. Set confidence level
Choose the desired statistical confidence percentage.

3. Set margin of error
Enter the tolerated percentage-point sampling error.

4. Estimate the exception rate
Use the expected share of records that may have the characteristic or exception being tested.

5. Review sample size
Use the rounded-up result as the minimum modeled sample, then select records using an appropriate sampling method.

n₀ = z² × p × (1 − p) ÷ e²
n = n₀ ÷ [1 + (n₀ − 1) ÷ N]

Where z is the two-sided normal critical value for the confidence level, p is the expected exception proportion, e is the margin of error as a decimal, and N is the finite population size. The displayed sample is rounded up.

What the result means

The result is a planning estimate based entirely on the values entered. Use it to compare scenarios and workload assumptions, not as a legal conclusion.

Confirm applicable law, contracts, policies, court orders, holds, and professional requirements before making compliance or legal decisions.

Given:
- Population = 5,000 records
- Confidence level = 95%
- Margin of error = 5%
- Expected exception rate = 50%

Calculation:
z ≈ 1.96.
Initial sample n₀ = 1.96² × 0.50 × 0.50 ÷ 0.05² ≈ 384.15.
Finite-population sample n = 384.15 ÷ [1 + (384.15 − 1) ÷ 5,000] ≈ 356.82.

Result:
Round up to 357 records.

Interpretation:
A properly selected random sample of about 357 records supports the entered statistical assumptions, but audit design and legal sufficiency may require other methods.

Why does a 50% expected exception rate create a larger sample?

For a proportion estimate, 50% maximizes p × (1 − p). That makes it a conservative choice when the true rate is unknown.

Can I use a convenience sample?

A convenience sample may introduce selection bias and does not provide the same statistical interpretation as a probability-based sample. The formula assumes a sampling process appropriate for the intended inference.

Why does population size sometimes have little effect?

Once a population is large relative to the initial sample, the finite-population correction becomes small. Confidence and margin of error then drive the sample more strongly.

Does this satisfy a regulator or court requirement?

Not necessarily. Specific audits, investigations, consent orders, or litigation may require a different sampling protocol or a full review.

What if I expect a very low exception rate?

Enter the best defensible expected rate, but consider whether your audit objective is detecting rare events. Rare-event detection may call for targeted or specialized sampling rather than this general proportion model.