Computer-Assisted Text Analysis for Social Research
EDU S062
Subject & Catalog Number
Course Information
Description
For social scientists, texts are an invaluable source of data. Whether in the form of primary sources, interviews, assessments, or other records of human communication, language is one of the most natural and important ways people preserve information, express meaning, and navigate social life. Thus, being able to study texts systematically is a powerful addition to a social scientist’s research toolkit.
Computer-assisted text analysis (CATA) is the practice of using computational tools to study texts without giving up the responsibilities of reading, interpretation, and judgment. This course introduces CATA as an applied research tradition for social scientists who want to work systematically with language data. Students will learn practical methods for organizing, transforming, evaluating, and analyzing text as data, with attention to how each computational transformation preserves some features of the original text while obscuring or discarding others.
The course is especially well suited for students who already have text data, research materials, or emerging text-based projects they want to work through, though an existing project is not required. Cross-registered students are warmly welcome, and auditors are welcome where permitted by program and university policies. Prior experience with Python, data science, or computational methods is helpful but not required; curiosity, care, and willingness to work through unfamiliar tools matter more.
Available for Harvard Cross Registration