Skip to main page content
Ideas
  1. Course Search
  2. EDU S062

Computer-Assisted Text Analysis for Social Research
EDU S062

Course Information

Description

For social scientists, texts are an invaluable source of data. Whether in the form of primary sources, interviews, assessments, or other records of human communication, language is one of the most natural and important ways people preserve information, express meaning, and navigate social life. Thus, being able to study texts systematically is a powerful addition to a social scientist’s research toolkit.

Computer-assisted text analysis (CATA) is the practice of using computational tools to study texts without giving up the responsibilities of reading, interpretation, and judgment. This course introduces CATA as an applied research tradition for social scientists who want to work systematically with language data. Students will learn practical methods for organizing, transforming, evaluating, and analyzing text as data, with attention to how each computational transformation preserves some features of the original text while obscuring or discarding others.

The course is especially well suited for students who already have text data, research materials, or emerging text-based projects they want to work through, though an existing project is not required. Cross-registered students are warmly welcome, and auditors are welcome where permitted by program and university policies. Prior experience with Python, data science, or computational methods is helpful but not required; curiosity, care, and willingness to work through unfamiliar tools matter more.

School Graduate School of Education
Credits 4
Cross Reg

Available for Harvard Cross Registration

Department Education
Course Component Regular Course
Instruction Mode In Person
Subject Education
Grading Basis HGSE Student Option (Letter Graded, Sat/Unsat)
Learning Goals <p>By the end of the course, students should be able to:</p><p>&nbsp; &nbsp;- {organize} text data into documented, traceable corpus structures suitable for computational analysis;<br>&nbsp; &nbsp;- {define} and {justify} units of analysis, metadata structures, preprocessing choices, and other translator decisions;<br>&nbsp; &nbsp;- {create}, {motivate}, and {compare} multiple translations of the same textual materials, including sparse, dense, and pretrained model-based representations;<br>&nbsp; &nbsp;- {use} translated text artifacts for exploration, visualization, search, labeling, inference, and prediction;<br>&nbsp; &nbsp;- {validate} computational outputs by returning to textual context and assessing what each translation preserves, highlights, obscures, or discards;<br>&nbsp; &nbsp;- {complete} a reproducible text analysis workflow using course data and, where appropriate, students’ own corpora.</p>
Career Focus <p>This course prepares students to work responsibly and effectively with text as data in research and applied settings. The skills are relevant for academic research, policy analysis, program evaluation, institutional research, qualitative and mixed-methods research, data science, consulting, journalism, nonprofit work, and public-sector roles. Students will gain experience organizing text data, using computational tools to search and analyze large text collections, validating automated outputs against textual evidence, documenting analytic decisions, and developing reproducible workflows that can support theses, dissertations, articles, reports, portfolios, or applied research projects. The emphasis on transparency makes it especially pressing for fields like academic research where being able to show one's work is essential.</p>