- Statistical Models II
- Summary
Course Syllabus
Obiettivi formativi
L’insegnamento rientra nelle aree di apprendimento delle scienze statistiche, dell’informatica e delle scienze sociali. Mira a fornire agli studenti una preparazione riguardante i seguenti approcci inferenziali: bootstrap non parametrico, distribuzione Gaussiana multivariata, modelli miscuglio Gaussiani univariati e multivariati, nonché modelli predittivi.
Durante l’attività didattica lo studente sviluppa una comprensione critica delle assunzioni alla base dei modelli teorici, attraverso applicazioni empiriche su dati reali e simulati. Lo studente acquisisce anche competenze relative alla messa in atto di ricerche riproducibili e replicabili. Inoltre, sviluppa abilità comunicative scritte, poiché è richiesta la redazione di testi che accompagnino i risultati delle analisi svolte.
Conoscenza e comprensione
L'insegnamento consente agli studenti di:
• Analizzare i dati utilizzando modelli statistici avanzati sviluppati per variabili risposta univariate e multivariate, sia di natura categoriale che continua.
• Sviluppare la conoscenza dei metodi di simulazione.
• Utilizzare la semantica del software R, anche attraverso l'ambiente RMarkdown, per sviluppare un metodo di ricerca replicabile e riproducibile. I documenti generati includono il codice, i risultati e i commenti al codice e alle analisi svolte.
• Interpretare i risultati delle elaborazioni in modo rigoroso sviluppando capacità espressive e di sintesi testuale anche per scopi divulgativi rivolti a un pubblico non accademico. In questo modo sviluppa autonomia di giudizio e affina le proprie abilità comunicative.
Capacità di applicare conoscenza e comprensione
L'insegnamento consente agli studenti di:
• Condurre l’inferenza statistica tramite tecniche di ricampionamento (bootstrap);
• Stimare, selezionare ed interpretare i modelli miscuglio di distribuzioni per popolazioni eterogenee;
• Concettualizzare i modelli a variabili latenti, stimare i parametri con il principio di massima verosimiglianza e interpretare i risultati;
• Applicare le conoscenze teoriche per analizzare dati di diverse tipologie derivanti dagli ambiti applicativi del corso di studio quali l'epidemiologia, la medicina, la biologia, la genetica e la salute pubblica.
• Implementare codice con il linguaggio del software open source R per le analisi descrittive ed inferenziali adottando un approccio open source che garantisca la riproducibilità e la replicabilità delle analisi.
L’insegnamento permette agli studenti di acquisire solidi elementi di teoria e di sviluppare le applicazioni pratiche attraverso un approccio di “problem solving”. L’insegnamento si inserisce nell’ambito della scienza dei dati, conoscenza oggi essenziale per i contesti lavorativi di sbocco degli studenti del corso di laurea in Biostatistica. Al termine dell’insegnamento, grazie al materiale fornito (le dispense del docente corredate da un’ampia bibliografia, i codici per i software R e l’interfaccia RMarkdown), lo studente è in grado di proseguire in modo autonomo nell’approfondimento di questa disciplina.
Contenuti sintetici
Nella prima parte dell’insegnamento vengono richiamate le principali distribuzioni probabilistiche che si utilizzano per simulare delle realizzazioni da variabili casuali. Viene presentato il procedimento di ricampionamento noto come bootstrap per ottenere misure di precisione in ambito non parametrico per alcuni stimatori di interesse.
Nella seconda parte dell’insegnamento dopo aver presentato la distribuzione Gaussiana multivariata si illustrano i modelli miscuglio Gaussiani. Vengono descritti i passi dell’algoritmo EM per la stima di massima verosimiglianza dei parametri dei modelli e dei modelli a variabili latenti con distribuzione discreta. Le lezioni di teoria sono affiancate da esercitazioni pratiche. L’insegnamento fornisce competenze nell'uso della semantica del software R, utilizzando anche la libreria RMarkdown tramite la libreria knitr per integrare il codice, i risultati delle analisi ed i commenti.
Programma esteso
La prima parte dell’attività didattica riguarda i metodi lineari congruenziali per la generazione di numeri pseudo-casuali ed i test grafici per la verifica della pseudo-casualità. La teoria è affiancata da esempi di simulazioni di dati da alcune distribuzioni probabilistiche. Nella seconda parte dell’attività didattica, dopo una breve introduzione sull’impianto concettuale dell’inferenza statistica, viene presentato il procedimento di ricampionamento noto come bootstrap per ottenere misure di precisione in ambito non parametrico per alcuni stimatori di interesse. Si illustrano gli intervalli di confidenza ottenuti sia con il metodo del percentile che con il metodo BCA che permette di correggere per la distorsione. Si illustrano i modelli miscuglio (finite mixture models) univariati e multivariati per variabili risposta quantitative assumendo una distribuzione di Gauss per le componenti del miscuglio. Si illustra il metodo di stima si massima verosimiglianza basato sull’algoritmo Expectation-Maximization. In particolare si considera la stima della densità e la classificazione delle unità statistiche con il metodo della massima probabilità a posteriori.
La teoria è affiancata da esercitazioni pratiche in cui vengono sviluppate, nell’ambiente R e con l’ausilio del marcatore di testo RMarkdown, numerose applicazioni volte all’analisi e all’adattamento dei modelli statistici per dati reali e simulati riguardanti gli ambiti della biostatistica. Le principali librerie del software R utilizzate sono skimr, MASS, boot, bootstrap, mclust. Lo studente è incoraggiato ad elaborare documenti riproducibili in cui commenta in forma testuale il codice ed i risultati delle analisi in modo critico anche tramite apprendimento cooperativo. Settimanalmente vengono assegnati degli esercizi e gli studenti nello svolgimento sono incoraggiati a scrivere reports in cui commentano il codice, ed offrono una spiegazione del procedimento di analisi svolto oltre ad una descrizione critica dei risultati ottenuti. Durante l’attività didattica vengono discusse le soluzioni agli esercizi assegnati.
Prerequisiti
Per una più agevole comprensione dei contenuti dell’insegnamento è necessario conoscere le nozioni di Probabilità e di Inferenza Statistica e la semantica di base del linguaggio di programmazione in ambiente R.
Metodi didattici
Sono previste lezioni frontali svolte presso il laboratorio informatico oppure aule informatizzate, le lezioni di teoria sono affiancate da esercitazioni pratiche che consentono agli studenti di apprendere tramite problem solving analizzando dati reali e simulati. Settimanalmente vengono assegnati degli esercizi di riepilogo relativi al programma svolto. Durante l’insegnamento con l'ausilio di R nell'ambiente RStudio e l'interfaccia di RMarkdown, gli studenti imparano ad elaborare documenti riproducibili che contengono codice, descrizioni e commenti ai risultati delle analisi. Sono incoraggiati a collaborare tra di loro nella risoluzione dei problemi applicativi, al fine di promuovere l'apprendimento cooperativo. L’insegnamento si svolge in 30 ore di didattica erogativa, dedicate alla presentazione degli aspetti teorici e metodologici, e 12 ore di didattica interattiva, dedicate a esercitazioni guidate, applicazioni con R/RStudio/RMarkdown, discussione dei risultati e attività di problem solving. Le videoregistrazioni asincrone rese disponibili sulla piattaforma e-learning costituiscono materiale integrativo di supporto allo studio e non sostituiscono le attività didattiche in presenza.
Modalità di verifica dell'apprendimento
Le seguenti modalità di verifica dell'apprendimento si applicano sia agli studenti frequentanti che a quelli non frequentanti le lezioni frontali. L’esame è in forma scritta con orale facoltativo, non sono previste prove intermedie. L'esame scritto ha una durata massima di due ore e si svolge in laboratorio informatico. Le domande aperte di teoria a cui gli studenti devono rispondere mirano a valutare la comprensione dei concetti essenziali dell’inferenza statistica condotta con metodi avanzati, mentre gli esercizi applicativi condotti utilizzando l'ambiente R, RStudio e RMarkdown, permettono di verificare la capacità degli studenti di applicare le metodologie proposte nonché di elaborare report riproducibili che descrivano i dati, le procedure e i risultati ottenuti. La prova mira anche a promuovere la capacità degli studenti di pianificare e gestire in modo efficace il tempo necessario per la stesura dell’elaborato. Durante l'esame è consentito l'utilizzo del materiale di studio e del codice R implementato durante l’insegnamento e personalmente dallo studente. Ogni punto di ogni esercizio ha una valutazione di circa 3 punti. Lo studente supera l'esame con una votazione non inferiore a 18/30.
Testi di riferimento
Il materiale didattico principale consiste nelle dispense preparate dal docente, che coprono, gli argomenti teorici, le applicazioni sviluppate con il software R, gli esercizi e le soluzioni. Queste dispense saranno rese disponibili sulla pagina della piattaforma e-learning dell'università dedicata all'insegnamento. Inoltre, il docente pubblica alla fine di ogni lezione le slides, i programmi di calcolo e i dataset utilizzati. Settimanalmente vengono assegnati esercizi, e le relative soluzioni. Sulla stessa pagina web sono disponibili degli esempi del testo d'esame.
I principali testi di riferimento sono elencati nella bibliografia delle dispense alcuni dei quali sono i seguenti che sono anche disponibili in ebook presso la biblioteca dell’Ateneo:
Bartolucci, F., Farcomeni, A., Pennoni, F. (2013). Latent Markov Models for longitudinal data, Chapman and Hall/CRC, Boca Raton.
Bishop, Y. M., Fienberg, S. E., Holland, P. W. (2007). Discrete multivariate analysis: theory and practice. Springer Science & Business Media, New York.
Blitzstein, J. K., Hwang, J. (2014). Introduction to probability, Chapman & Hall/CRC.
Gentle, J. E., Härdle, W., Mori Y. (2004). Handbook of computational statistics. Springer-Berlin.
Lange, K. (2010). Numerical analysis for statisticians, 2nd Edition, Springer, New York.
Pennoni, F. (2026). Dispensa di Modelli Statistici II, parte di teoria e applicazioni con R. Dipartimento di Statistica e
Metodi Quantitativi, Università degli Studi di Milano-Bicocca.
R Core Team (2026). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/.
Periodo di erogazione dell'insegnamento
Semester I, cycle I, September-November 2026
Lingua di insegnamento
L’insegnamento viene erogato in lingua italiana. Gli studenti Erasmus possono utilizzare il materiale didattico predisposto in lingua inglese e fornito dal docente su richiesta. Possono inoltre richiedere di svolgere la prova d’esame in lingua inglese.
Sustainable Development Goals
Learning objectives
The course falls within the learning areas of statistical sciences, computer science and social sciences. It aims to provide students with knowledge of the following inferential approaches: nonparametric bootstrap, the multivariate Gaussian distribution, univariate and multivariate Gaussian mixture models, and predictive models.
During the learning activities, students develop a critical understanding of the assumptions underlying theoretical models through empirical applications using real and simulated data. Students also acquire skills related to the implementation of reproducible and replicable research. In addition, they develop written communication skills, since they are required to write texts accompanying the results of the analyses carried out.
Knowledge and understanding
The course enables students to:
• Analyze data using advanced statistical models developed for univariate and multivariate response variables, both categorical and continuous.
• Develop knowledge of simulation methods.
• Use the syntax and semantics of the R software, also through the RMarkdown environment, to develop a replicable and reproducible research approach. The documents produced include code, results, and comments on the code and on the analyses carried out.
• Interpret the results of data processing rigorously, developing expressive and text-summarizing skills also for dissemination purposes aimed at a non-academic audience. In this way, students develop independent judgment and refine their communication skills.
Applying knowledge and understanding
The course enables students to:
• Conduct statistical inference using resampling techniques (bootstrap).
• Estimate, select and interpret mixture models of distributions for heterogeneous populations.
• Conceptualize latent variable models, estimate their parameters using the maximum likelihood principle, and interpret the results.
• Apply theoretical knowledge to analyze different types of data arising from the application fields of the degree program, such as epidemiology, medicine, biology, genetics and public health.
• Implement code using the open-source R software language for descriptive and inferential analyses, adopting an open-source approach that ensures the reproducibility and replicability of the analyses.
The course enables students to acquire solid theoretical foundations and to develop practical applications through a problem-solving approach. The course is part of the field of data science, a knowledge area that is now essential for the professional contexts available to graduates in Biostatistics. At the end of the course, thanks to the material provided (the instructor's handouts accompanied by an extensive bibliography, R software code and the RMarkdown interface), students are able to continue studying this discipline independently.
Contents
In the first part of the course, the main probability distributions used to simulate realizations from random variables are reviewed. The resampling procedure known as bootstrap is introduced in order to obtain measures of precision in a nonparametric framework for selected estimators of interest. In the second part of the course, after presenting the multivariate Gaussian distribution, Gaussian mixture models are illustrated. The steps of the EM algorithm for maximum likelihood estimation of the parameters of these models and of latent variable models with a discrete distribution are described. The theoretical lectures are complemented by practical exercises. The course provides skills in using the syntax and semantics of the R software, also using the RMarkdown library through the knitr package to integrate code, analysis results and comments.
Detailed program
The first part of the learning activities concerns linear congruential methods for generating pseudo-random numbers and graphical tests for assessing pseudo-randomness. The theoretical discussion is accompanied by examples of data simulation from selected probability distributions. In the second part of the learning activities, after a brief introduction to the conceptual framework of statistical inference, the resampling procedure known as bootstrap is presented in order to obtain measures of precision in a nonparametric framework for selected estimators of interest. Confidence intervals obtained using both the percentile method and the BCa method, which corrects for bias, are illustrated. Univariate and multivariate finite mixture models for quantitative response variables are illustrated, assuming Gaussian distributions for the mixture components. The maximum likelihood estimation method based on the Expectation-Maximization algorithm is presented. In particular, density estimation and classification of statistical units using the maximum a posteriori probability method are considered.
Theory is complemented by practical exercises in which numerous applications are developed in the R environment with the support of the RMarkdown markup interface. These applications are aimed at the analysis and fitting of statistical models for real and simulated data concerning the fields of biostatistics. The main R packages used are skimr, MASS, boot, bootstrap, mclust. Students are encouraged to prepare reproducible documents in which they provide written comments on the code and critically discuss the results of the analyses, also through cooperative learning. Exercises are assigned weekly, and students are encouraged to write reports in which they comment on the code, provide an explanation of the analytical procedure carried out, and critically describe the results obtained. During the learning activities, the solutions to the assigned exercises are discussed.
Prerequisites
For an easier understanding of the course content, students must know the basic notions of Probability and Statistical Inference, as well as the basic syntax and semantics of the programming language in the R environment.
Teaching methods
Lectures are held in the computer laboratory or in computer-equipped classrooms. Theoretical lectures are complemented by practical exercises that allow students to learn through problem solving by analysing real and simulated data. Weekly review exercises are assigned on the topics covered in the course. During the course, with the support of R in the RStudio environment and the R Markdown interface, students learn to produce reproducible documents containing code, descriptions, and comments on the results of the analyses. Students are encouraged to collaborate with one another in solving applied problems, in order to promote cooperative learning. The course consists of 30 hours of lectures, devoted to the presentation of theoretical and methodological aspects, and 12 hours of interactive teaching, devoted to guided exercises, applications using R/RStudio/R Markdown, discussion of results, and problem-solving activities. The asynchronous video recordings made available on the e-learning platform constitute supplementary study support material and do not replace in-person teaching activities.
Assessment methods
The following methods for assessing learning apply both to students attending and to students not attending the lectures. The examination consists of a written test with an optional oral examination. No mid-term tests are scheduled. The written examination lasts a maximum of two hours and takes place in the computer lab. The open-ended theoretical questions aim to assess the understanding of the essential concepts of statistical inference conducted with advanced methods. The applied exercises, carried out using the R environment, RStudio and RMarkdown, allow the assessment of students' ability to apply the proposed methodologies and to produce reproducible reports describing the data, the procedures and the results obtained.
The written examination is graded on a scale out of 30, according to a scoring rubric communicated to students before the exam. The assessment takes into account: the theoretical correctness of the answers; the appropriateness of the model choice; the correctness of the computational implementation; the critical interpretation of the results; and the clarity and completeness of the written communication. The optional oral examination, if requested by the student or by the instructor, consists of a discussion of the written examination and of the theoretical topics covered in the course, and may confirm, supplement, or modify the final grade. During the examination, students are allowed to use the teaching materials and the materials provided by the instructor, including the R and SAS code developed during the course activities. The exam is considered passed with a minimum grade of 18/30.
Textbooks and Reading Materials
The main teaching material consists of handouts prepared by the instructor, covering theoretical topics, applications developed with the R software, exercises and solutions. These handouts will be made available on the university e-learning platform page dedicated to the course. In addition, at the end of each lesson the instructor publishes the slides, calculation programs and datasets used.
Exercises are assigned weekly, together with the corresponding solutions. Examples of examination texts are available on the same webpage.
The main reference texts are listed in the bibliography of the handouts. Some of them are the following, also available as eBooks in the University Library:
Bartolucci, F., Farcomeni, A., Pennoni, F. (2013). Latent Markov Models for longitudinal data, Chapman and Hall/CRC, Boca Raton.
Bishop, Y. M., Fienberg, S. E., Holland, P. W. (2007). Discrete multivariate analysis: theory and practice. Springer Science & Business Media, New York.
Blitzstein, J. K., Hwang, J. (2014). Introduction to probability, Chapman & Hall/CRC.
Gentle, J. E., Härdle, W., Mori Y. (2004). Handbook of computational statistics. Springer-Berlin.
Lange, K. (2010). Numerical analysis for statisticians, 2nd Edition, Springer, New York.
Pennoni, F. (2026). Dispensa di Modelli Statistici II, parte di teoria e applicazioni con R. Dipartimento di Statistica e
Metodi Quantitativi, Università degli Studi di Milano-Bicocca.
R Core Team (2026). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria. https://www.R-project.org/.
Semester
Semestre I, ciclo I, Settembre-Novembre 2025
Teaching language
The course is delivered in Italian. Erasmus students may use the teaching material prepared in English and provided by the instructor upon request. They may also request to take the examination in English.