Introduction to Data Science for Behavioral and Educational Research
EDU S022
Subject & Catalog Number
Course Information
Description
This course introduces modern data science and machine learning tools for answering research questions in education and other social and behavioral sciences. We will focus first on how to craft research questions, identify useful data, wrangle/clean messy data, conduct simple descriptive analyses, and visualize/present complex data in useful ways. Next we cover machine learning tools; in this unit, we build and evaluate predictive models. Topics include cross-validation, data leakage, and useful machine learning models (e.g. Lasso and random forests). The third unit considers pre-trained models (e.g. Large Language Models), how they work, and how we can use them to enable analyses of unstructured data (e.g. images and text). Finally we will cover inference, emphasizing bootstrap methods to express uncertainty in estimates produced by the tools used earlier in the course. We will write code in R, but we do not assume any prior experience with the language. Our study of R will include the use of AI to write code.
Prerequisites: S-040 or equivalent course in statistics (concurrent enrollment accepted), or by permission.
Available for Harvard Cross Registration