Chapter 1 Wrap Up
Concept Check
Section Reviews
1.1 Introduction to Statistics and Key Terms
The mathematical theory of statistics is easier to learn when you know the language. This module presents important terms that will be used throughout the text.
1.2 Data Basics
Some calculations generate numbers that are artificially precise. It is not necessary to report a value to eight decimal places when the measures that generated that value were only accurate to the nearest tenth. Round off your final answer to one more decimal place than was present in the original data. This means that if you have data measured to the nearest tenth of a unit, report the final statistic to the nearest hundredth.
In addition to rounding your answers, you can measure your data using the following four levels of measurement: nominal, ordinal, interval, and ratio.
When organizing data, it is important to know how many times a value appears. How many statistics students study five hours or more for an exam? What percent of families on our block own two pets? Frequency, relative frequency, and cumulative relative frequency are measures that answer questions like these.
1.3 Data Collection and Observational Studies
In summary, making causal conclusions based on observational data can be treacherous and is not recommended. Thus, observational studies are generally only sufficient to show associations or form hypotheses that we later check using controlled experiments which we will discuss in the next section.
1.4 Designed Experiments
A poorly designed study will not produce reliable data. There are certain key components that must be included in every experiment. To eliminate lurking variables, subjects must be assigned randomly to different treatment groups. One of the groups must act as a control group, demonstrating what happens when the active treatment is not applied. Participants in the control group receive a placebo treatment that looks exactly like the active treatments but cannot influence the response variable. To preserve the integrity of the placebo, both researchers and subjects may be blinded. When a study is designed properly, the only difference between treatment groups is the one imposed by the researcher. Therefore, when groups respond differently to different treatments, the difference must be due to the influence of the explanatory variable.
“An ethics problem arises when you are considering an action that benefits you or some cause you support, hurts or reduces benefits to others, and violates some rule,” Andrew Gelman. Ethical violations in statistics are not always easy to spot. Professional associations and federal agencies post guidelines for proper conduct. It is important that you learn basic statistical procedures so that you can recognize proper data analysis.
1.5 Sampling Techniques and Ethics
Data are individual items of information that come from a population or sample. Data may be classified as qualitative (categorical), quantitative continuous, or quantitative discrete.
Because it is not practical to measure the entire population in a study, researchers use samples to represent the population. A random sample is a representative group from the population chosen by using a method that gives each individual in the population an equal chance of being included in the sample. Random sampling methods include simple random sampling, stratified sampling, cluster sampling, and systematic sampling. Convenience sampling is a nonrandom method of choosing a sample that often produces biased data.
Samples that contain different individuals result in different data. This is true even when the samples are well-chosen and representative of the population. When properly selected, larger samples model the population more closely than smaller samples. There are many different potential problems that can affect the reliability of a sample. Statistical data needs to be critically analyzed, not simply accepted.
Key Terms
Try to define the terms below on your own. Scroll over any term to check your response!
1.1 Introduction to Statistics and Key Terms
- Data analysis process
- Descriptive statistics
- Inferential statistics
- Probability
- Population
- Parameters
- Sample
- Statistic
- Individuals
- Variable
- Values
- Data
1.2 Data Basics
- Data
- Population
- Sample
- Qualitative (categorical)
- Quantitative (numerical)
- Discrete
- Continuous
- Nominal scale level
- Ordinal scale level
- Interval scale level
- Ratio scale level
- Variation
- Data analysis
1.3 Data Collection and Observational Studies
- Explanatory variable
- Response variable
- Data
- Anecdotal evidence
- Observational studies
- Designed (controlled) experiment
- Associations
- Confounding (lurking, conditional) variable
- Prospective study
- Retrospective study
- Cohort study
- Longitudinal study
- Cross-sectional study
- Case-control study
1.4 Designed Experiments
- Observational study
- Controlled (designed) experiments
- Explanatory variable
- Response variable
- Treatments
- Experimental unit
- Repeated measures
- Control group
- Placebo
- Blinding
- Double-blind
- Factors
- Levels
- Treatment combinations (interactions)
- Completely randomized
- Block design
- Matched pairs design
1.5 Sampling Techniques and Ethics
- Sample
- Simple random sample (SRS)
- Stratified sampling
- Cluster sampling
- Systematic sampling
- Sampling bias
- Sampling variability
- Convenience sampling
Extra Practice
Please see the extra practice section at the end of the book
Process of collecting, organizing, and analyzing data
Methods of organizing, summarizing, and presenting data
The facet of statistics dealing with using a sample to generalize (or infer) about the population
The study of randomness; a number between zero and one, inclusive, that gives the likelihood that a specific event will occur
The whole group of individuals who can be studied to answer a research question
A number that is used to represent a population characteristic and can only be calculated as the result of a census
A subset of the population studied
A number calculated from a sample
The person, animal, item, thing, place, etc. that we collect information about
A characteristic of interest for each person or object in a population
Possible observations of the variable
Actual values (numbers or words) that are collected from the variables of interest
Data that describes qualities, or puts individuals into categories
Numerical data with a mathematical context
A random variable that produces discrete data
A random variable (RV) whose outcomes are measured as an uncountable, infinite, number of values
Categorical data where the the categories have no natural, intuitive, or obvious order
Categorical data where the the categories have a natural or intuitive order
Quantitative data where the difference or gap between values is meaningful
Quantitative data where the difference or gap between values is meaningful AND has a true 0 value
The level of variability or dispersion of a dataset; also commonly known as variation/variability
The independent variable in an experiment; the value controlled by researchers
The dependent variable in an experiment; the value that is measured for change at the end of an experiment
Evidence that is based on personal testimony and collected informally
Data collection where no variables are manipulated
Data collection where variables are manipulated in a controlled setting
A relationship between variables
A variable that has an effect on a study even though it is neither an explanatory variable nor a response variable
Collecting information as events unfold
Collecting or using data after events have taken place
Longitudinal study where a group of people (typically having a common factor) are studied and data is collected for a purpose
Collecting data multiple times on the same individuals, usually at fixed increments, over a period of time
Data collection on a population at one point in time (often prospective)
A study that compares a group that has a certain characteristic to a group that does not, often a retrospective study for rare conditions
Type of experiment where variables are manipulated; data is collected in a controlled setting
Different values or components of the explanatory variable applied in an experiment
Any individual or object to be measured
When an individual goes through a single treatment more than once
A group in a randomized experiment that receives no (or an inactive) treatment but is otherwise managed exactly as the other groups
An inactive treatment that has no real effect on the explanatory variable
Not telling participants which treatment they are receiving
The act of blinding both the subjects of an experiment and the researchers who work with the subjects
Variables in an experiment
Certain values of variables in an experiment
Combinations of levels of variables in an experiment
Dividing participants into treatment groups randomly
Grouping individuals based on a variable into "blocks" and then randomizing cases within each block to the treatment groups
Very similar individuals (or even the same individual) receive two different two treatments (or treatment vs. control) then the difference in results are compared
Each member of the population is equally likely to be chosen for a sample of a given sample size and each sample is equally likely to be chosen
Dividing a population into groups (strata), and then using simple random sampling to identify a proportionate number of individuals from each
A method of sampling where the population has already sorted itself into groups (clusters), randomly selecting a cluster, and using every individual in the chosen cluster as the sample
Using some sort of pattern or probability based method for choosing your sample
Bias resulting from all members of the population not being equally likely to be selected
The idea that samples from the same population can yield different results
Selecting individuals that are easily accessible and may result in biased data