• A
  • A
  • A
  • АБВ
  • АБВ
  • АБВ
  • А
  • А
  • А
  • А
  • А
Обычная версия сайта

Бакалаврская программа «Социология и социальная информатика»

04
Октябрь

Advanced Data Analysis

2026/2027
Учебный год
ENG
Обучение ведется на английском языке
5
Кредиты
Статус:
Курс по выбору
Когда читается:
4-й курс, 1, 2 модуль

Преподаватель

Course Syllabus

Abstract

The course is targeted at undergraduate social science students aiming at careers in data analysis or academia. The course consists of seminars. It covers special prediction and classification models (logistic regression and cluster analysis) and more advanced data management topics such as web scraping and data imputation. The course discusses data culture, from data management to coding styles to narrating with data and Bayesian statistics. Another feature of that course is to learn the coding instruments which helps to create tidy analysis. It is also a starting point for students interested in pursuing advanced training in research methods or planning to use quantitative methods with categorical outcomes in their own research.
Learning Objectives

Learning Objectives

  • The course covers the foundations and popular techniques of categorical data analysis with the goal of training students to be informed producers and consumers of quantitative research.
Expected Learning Outcomes

Expected Learning Outcomes

  • Students can apply classification techniques, propose hypotheses and choose the methods in categorical data analysis in R, including supervised classification with a binary outcome, and unsupervised classification with clustering techniques of mixed data types.
  • Students create customized R Markdown reports.
  • Students create reproducible analysis scripts.
  • Students define basic terms and identify the purposes of Bayesian inference vs frequentist inference.
  • Students describe known problems with the null hypothesis statistical testing and propose known solutions to them.
  • Students inspect missing data patterns and apply various methods of data imputation.
  • Students interpret the results and assess the quality of proposed analytical and visualization solutions, provide reasons for their choice of techniques, interpret the outputs correctly, and assess the quality of models and data stories.
  • Students propose and apply tools for reproducible and ethical data analysis.
  • Students scrape simple web tables and texts with R and convert them into standard data formats.
Course Contents

Course Contents

  • Coding with Style
  • Web Scraping
  • Binary Logistic Regression
  • Missing Data
  • Cluster Analysis
  • Data Culture and Data Acumen
Assessment Elements

Assessment Elements

  • non-blocking RMD Customization
    Create a customized Rmd template by tweaking the theme settings in the YAML and creating your unique CSS style of report. Add a structure to the template with a table of contents or bookmarks, and name the sections with all the coding projects lying ahead in this course - Web Scraping, Binary Logistic Regression with Data Imputation, Dimension Reduction, and Cluster Analysis.Choose colors you are going to use in all plots and save them to a palette. Create a personalized "theme" function for ggplot2 for quick plotting. Provide a sample of how plots, tables, and pictures would look like in this template. Follow the clean code principles in your script(s).
  • non-blocking Web Scraping
    Your task is to scrape data from an open Internet source (e.g. Wikipedia, government statistics, cultural datasets). From this data, create: One meaningful binary variable (1/0) At least 3 independent variables (numeric or categorical) Dataset should contain at least 300 rows with valid data A tidy dataset following the principles discussed in class Regex (stringr) must be used in the tidying process Your submission must be reproducible: the instructor should be able to run your script from scratch and get the same result
  • non-blocking Binary Outcome Project
    Predict a binary outcome using binary logistic regression and decision tree (CHAID, CRT/CART, or QUEST). Evaluate model assumptions, fit, and performance. Compare the two methods and provide a reasoned conclusion about which performs better. Your analysis should be reproducible, well-documented, and include both technical results and non-technical interpretation. You can use data from Web scraping task or any other available data (e.g., ESS, RLMS, EVS etc.).
  • non-blocking Data Imputation
    This assignment focuses on identifying missing data mechanisms and patterns in a real dataset and applying multiple imputation methods discussed in the seminar. You must explore missingness visually and statistically, perform and justify imputation, compare results across methods, and produce a final dataset suitable for later modelling (e.g., logistic regression). You may use the dataset prepared for the binary logistic regression assignment or any open-source dataset (e.g., RLMS, ESS, airquality).
  • non-blocking Cluster Analysis
    Identify, evaluate, and interpret clusters in a dataset that contains mixed types of variables (numeric, categorical). Students must demonstrate the full clustering workflow: data preparation, choice of distance, application of at least two clustering methods (one agglomerative/hierarchical + at least one non-hierarchical: PAM/DBSCAN), selection of the number of clusters, validation and stability checks, visualization (at least 2 methods), cluster profiling and a reasoned final choice of the best solution.
  • non-blocking Final exam
    Online test conducted during the seminar session.
Interim Assessment

Interim Assessment

  • 2026/2027 2nd module
    0.1 * RMD Customization + 0.2 * Final exam + 0.15 * Web Scraping + 0.2 * Binary Outcome Project + 0.2 * Cluster Analysis + 0.15 * Data Imputation
Bibliography

Bibliography

Recommended Core Bibliography

  • Baker, M. (2015). Reproducibility crisis: Blame it on the antibodies. Nature, 521(7552), 274–276. https://doi.org/10.1038/521274a
  • Ledolter, J. (2013). Data Mining and Business Analytics with R. Hoboken, New Jersey: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=587979
  • Munzert, S. (2014). Automated Data Collection with R : A Practical Guide to Web Scraping and Text Mining. HobokenChichester, West Sussex, United Kingdom: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=878670
  • Upton, G. J. G. (2016). Categorical Data Analysis by Example. Hoboken, New Jersey: Wiley. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=1402878
  • Wickham, H. (2015). Advanced R, Second Edition. Boca Raton, FL: Chapman and Hall/CRC. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=934735
  • Wickham, H., & Grolemund, G. (2016). R for Data Science : Import, Tidy, Transform, Visualize, and Model Data (Vol. First edition). Sebastopol, CA: Reilly - O’Reilly Media. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=1440131

Recommended Additional Bibliography

  • 9781439898208 - Andrew Gelman , John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, Donald B. Rubin - Bayesian Data Analysis, Third Edition - 2013 - Chapman & Hall/CRC Press - http://search.ebscohost.com/login.aspx?direct=true&db=nlebk&AN=1763244 - nlebk - 1763244
  • 9781482253467 - McElreath, Richard - Statistical Rethinking : A Bayesian Course with Examples in R and Stan - 2015 - Chapman and Hall/CRC - http://search.ebscohost.com/login.aspx?direct=true&db=nlebk&AN=1338291 - nlebk - 1338291
  • Hadley, W. (2016). Ggplot2 : Elegant Graphics for Data Analysis. New York, NY: Springer. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=1175341
  • Little, R. J. A., & Rubin, D. B. (2002). Statistical Analysis with Missing Data (Vol. Second edition). Hoboken: Wiley-Interscience. Retrieved from http://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=edsebk&AN=838162
  • Mood, C. (2010). Logistic Regression: Why We Cannot Do What We Think We Can Do, and What We Can Do About It. European Sociological Review, 26(1), 67–82. https://doi.org/10.1093/esr/jcp006
  • Seppe vanden Broucke, & Bart Baesens. (2018). Practical Web Scraping for Data Science : Best Practices and Examples with Python. Apress.

Authors

  • Lebedev Daniil Vadimovich