- Economics
- Bachelor Degree
- Statistica e Gestione delle Informazioni [E4104B - E4102B]
- Courses
- A.A. 2026-2027
- 2nd year
- Statistical Models
- Summary
Course Syllabus
Obiettivi formativi
1. Conoscenza e capacità di comprensione
Al termine dell’insegnamento, lo studente avrà acquisito una solida conoscenza dei principali modelli statistici per l’analisi di dati multivariati, con particolare riferimento alla regressione lineare multipla, al modello lineare classico e alla regressione logistica. Sarà in grado di comprendere il ruolo del modello statistico nell’analisi dei fenomeni reali, distinguendolo dal modello matematico e dagli approcci di machine learning. Conoscerà le principali assunzioni alla base dei modelli studiati, le procedure di stima dei parametri, i criteri di valutazione della bontà di adattamento e le metodologie per l’interpretazione dei risultati. Acquisirà inoltre familiarità con la rappresentazione matriciale dei dati e dei modelli, nonché con i fondamenti teorici dell’inferenza statistica applicata ai modelli di regressione.
2. Conoscenza e capacità di comprensione applicate
Lo studente sarà in grado di applicare i modelli statistici studiati all’analisi di dati provenienti da differenti contesti applicativi, quali quelli economici, aziendali, sociali, biologici, medici e ambientali. Saprà specificare un modello statistico coerente con gli obiettivi dell’analisi, stimarne i parametri, verificarne le assunzioni e valutarne la capacità esplicativa e previsionale. Sarà inoltre in grado di utilizzare il software R per la gestione dei dati, la stima di modelli di regressione lineare e logistica, l’individuazione di problematiche quali multicollinearità e osservazioni anomale, e la produzione di analisi e report basati su dati reali e simulati.
3. Autonomia di giudizio
Durante il percorso formativo, lo studente svilupperà la capacità di valutare criticamente l’adeguatezza di un modello statistico rispetto al fenomeno oggetto di studio. Sarà in grado di analizzare la plausibilità delle assunzioni adottate, interpretare correttamente i risultati delle procedure inferenziali, confrontare modelli alternativi e valutarne punti di forza e limiti. Acquisirà inoltre la capacità di distinguere tra finalità esplicative e previsive dell’analisi statistica, formulando giudizi fondati sull’evidenza empirica e sulla corretta interpretazione dei risultati.
4. Abilità comunicative
Lo studente maturerà la capacità di comunicare in modo chiaro e rigoroso i risultati delle analisi statistiche, utilizzando il linguaggio tecnico proprio della modellistica statistica. Sarà in grado di presentare e discutere modelli di regressione, risultati inferenziali, misure di bontà di adattamento e indicatori previsionali sia in forma scritta sia orale. Saprà inoltre predisporre report tecnici corredati da tabelle, grafici e commenti interpretativi, adattando il livello di approfondimento alle caratteristiche degli interlocutori.
5. Capacità di apprendere
Il corso mira a sviluppare nello studente la capacità di approfondire autonomamente metodologie statistiche avanzate e di affrontare problemi complessi di analisi dei dati. Le conoscenze teoriche e applicative acquisite costituiranno una base essenziale per insegnamenti successivi nell’ambito della statistica, della data science e dell’analisi quantitativa, nonché per l’utilizzo professionale dei modelli statistici in contesti decisionali e di ricerca. Lo studente sarà inoltre in grado di aggiornare autonomamente le proprie competenze rispetto all’evoluzione degli strumenti metodologici e informatici per l’analisi dei dati.
Contenuti sintetici
Il corso introduce i principali strumenti della modellistica statistica per l’analisi dei dati. Vengono affrontati il concetto di modello statistico e il processo di modellizzazione, i modelli di regressione lineare semplice e multipla e le principali estensioni per variabili risposta di diversa natura.
Il corso tratta le basi dell’inferenza statistica, la stima dei parametri, la valutazione e selezione dei modelli e gli elementi essenziali di diagnostica. Le metodologie sono applicate a dati reali e simulati mediante il software R.
Programma esteso
Il corso introduce i principali concetti e strumenti della modellistica statistica per l’analisi di dati multivariati. Vengono presentati il concetto di modello statistico e il suo ruolo nell’analisi dei fenomeni reali, le differenze rispetto al modello matematico e agli approcci di machine learning, nonché le principali fasi del processo di modellizzazione statistica, dalla formulazione delle ipotesi alla raccolta dei dati, dalla stima dei parametri alla valutazione della capacità esplicativa e previsionale dei modelli.
Successivamente vengono approfonditi i modelli di regressione per lo studio delle relazioni tra variabili quantitative. Sono illustrati i metodi di stima dei parametri, i criteri per la valutazione della bontà di adattamento, le procedure di selezione dei modelli e gli strumenti per la diagnosi delle principali criticità che possono emergere nell’analisi dei dati. Vengono inoltre introdotti elementi di rappresentazione matriciale dei dati e dei modelli statistici.
Particolare attenzione è dedicata agli aspetti inferenziali della modellistica statistica. Si analizzano le proprietà degli stimatori, le ipotesi alla base del modello lineare classico, la costruzione e l’interpretazione di intervalli di confidenza e test di ipotesi, nonché i criteri per il confronto tra modelli alternativi e per la valutazione del loro utilizzo a fini esplicativi e previsivi.
Il corso affronta inoltre alcune estensioni dei modelli lineari, con particolare riferimento all’inclusione di variabili qualitative, alle trasformazioni delle variabili e ai modelli per variabili risposta non continue. Vengono presentati i fondamenti della regressione logistica e l’interpretazione dei relativi parametri e indicatori.
Gli argomenti teorici sono costantemente affiancati da applicazioni pratiche mediante il software R, utilizzando dati reali e simulati per sviluppare competenze operative nell’analisi statistica e nell’interpretazione dei risultati.
Prerequisiti
Si richiede di aver superato gli esami degli insegnamenti propedeutici: Statistica I, Analisi Matematica I, Algebra Lineare, Calcolo delle Probabilità.
Per una più agevole comprensione dei contenuti del corso è fortemente consigliato conoscere le nozioni di inferenza statistica impartite al corso di Statistica II.
Metodi didattici
Il corso prevede 49 ore di lezione e 12 ore di esercitazioni svolte in presenza. Le attività didattiche sono affiancate da un tutor.
Le lezioni frontali affrontano gli aspetti teorici della modellistica statistica, con il supporto di strumenti informatici e di esempi applicativi sviluppati mediante il software R. Le esercitazioni si svolgono in laboratorio informatico e sono dedicate all’applicazione delle tecniche di analisi statistica su dati reali e simulati.
Durante il corso gli studenti sono guidati nello sviluppo di analisi statistiche, nella scelta delle metodologie più appropriate e nell’interpretazione dei risultati. Il materiale didattico viene reso disponibile sulla piattaforma e-learning prima delle lezioni.
Modalità di verifica dell'apprendimento
L’esame prevede una prova scritta e una prova orale. Non sono previste prove intermedie.
La prova scritta si svolge mediante l’utilizzo del software R e consiste nell’analisi di un insieme di dati forniti dal docente. Lo studente è chiamato a svolgere un’analisi statistica completa, includendo l’applicazione dei modelli studiati, la produzione di grafici e tabelle e la redazione di un report in formato R Markdown. Il report deve contenere l’interpretazione dei risultati e il commento delle principali evidenze emerse dall’analisi.
La prova orale verte sugli argomenti del corso e ha lo scopo di verificare la comprensione teorica dei metodi utilizzati e la capacità di interpretazione critica dei risultati ottenuti.
È inoltre prevista una giornata di presentazione di lavori di gruppo, durante la quale gli studenti illustrano un’analisi svolta autonomamente su dati reali reperiti in modo indipendente, mediante una presentazione in aula. Tale attività contribuisce alla valutazione complessiva del percorso formativo.
Testi di riferimento
I principali testi di riferimento sono:
-
J. J. Faraway: Linear Models with R, 2nd ed., Chapman & Hall/CRC, 2014.
-
G. James, D. Witten, T. Hastie, R. Tibshirani: An Introduction to Statistical Learning, 2nd ed., Springer, 2021 (con particolare riferimento ai capitoli dedicati alla regressione lineare e alla regressione logistica, con applicazioni in R).
-
Carter Hill, William E. Griffiths, Guay C. Lim: Principles of Econometrics, 5th ed., Wiley, 2018.
-
D. Nolan, D. T. Lang: Data Science in R: A Case Studies Approach to Computational Reasoning and Problem Solving, Chapman & Hall/CRC, 2015.
-
Dispense fornite dal docente durante lo svolgimento del corso.
Periodo di erogazione dell'insegnamento
II Semestre, III Ciclo
Lingua di insegnamento
Italiana
Sustainable Development Goals
Learning objectives
1. Knowledge and Understanding
By the end of the course, students will have acquired a solid understanding of the main statistical models for the analysis of multivariate data, with particular emphasis on multiple linear regression, the classical linear model, and logistic regression. They will be able to understand the role of statistical models in the analysis of real-world phenomena, distinguishing them from mathematical models and machine learning approaches. Students will be familiar with the main assumptions underlying the models studied, parameter estimation procedures, model fit evaluation criteria, and methods for interpreting results. They will also gain familiarity with the matrix representation of data and statistical models, as well as with the theoretical foundations of statistical inference applied to regression models.
2. Applied Knowledge and Understanding
Students will be able to apply the statistical models studied to the analysis of data from different application contexts, such as economics, business, social sciences, biology, medicine, and environmental studies. They will be able to specify a statistical model consistent with the objectives of the analysis, estimate its parameters, verify its assumptions, and evaluate its explanatory and predictive performance. They will also be able to use the R software environment for data management, estimation of linear and logistic regression models, identification of issues such as multicollinearity and outliers, and the production of analyses and reports based on real and simulated data.
3. Autonomy of Judgment
Throughout the course, students will develop the ability to critically assess the adequacy of a statistical model with respect to the phenomenon under study. They will be able to evaluate the plausibility of the underlying assumptions, correctly interpret the results of inferential procedures, compare alternative models, and assess their strengths and limitations. They will also acquire the ability to distinguish between explanatory and predictive purposes of statistical analysis, formulating judgments based on empirical evidence and correct interpretation of results.
4. Communication Skills
Students will develop the ability to clearly and rigorously communicate the results of statistical analyses using the technical language of statistical modeling. They will be able to present and discuss regression models, inferential results, goodness-of-fit measures, and predictive indicators both in written and oral form. They will also be able to prepare technical reports supported by tables, graphs, and interpretative comments, adapting the level of detail to the characteristics of the audience.
5. Learning Skills
The course aims to develop students’ ability to independently deepen advanced statistical methodologies and to tackle complex data analysis problems. The theoretical and applied knowledge acquired will provide a solid foundation for subsequent courses in statistics, data science, and quantitative analysis, as well as for the professional use of statistical models in decision-making and research contexts. Students will also be able to independently update their skills in response to the evolution of methodological and computational tools for data analysis.
Contents
The course introduces the main tools of statistical modeling for data analysis. It covers the concept of a statistical model and the modeling process, simple and multiple linear regression models, and the main extensions for response variables of different types.
The course addresses the foundations of statistical inference, parameter estimation, model evaluation and selection, and essential diagnostic tools. The methodologies are applied to real and simulated datasets using the R software environment.
Detailed program
The course introduces the main concepts and tools of statistical modeling for the analysis of multivariate data. It presents the concept of a statistical model and its role in the analysis of real-world phenomena, the differences between statistical and mathematical models as well as machine learning approaches, and the main stages of the statistical modeling process, from hypothesis formulation and data collection to parameter estimation and the assessment of the explanatory and predictive performance of models.
Subsequently, regression models for studying relationships between quantitative variables are introduced. The course covers parameter estimation methods, criteria for assessing model fit, model selection procedures, and diagnostic tools for identifying potential issues in data analysis. Elements of matrix representation of data and statistical models are also introduced.
Particular attention is devoted to the inferential aspects of statistical modeling. The properties of estimators are analyzed, along with the assumptions underlying the classical linear model, the construction and interpretation of confidence intervals and hypothesis tests, as well as criteria for comparing alternative models and evaluating their use for explanatory and predictive purposes.
The course also addresses extensions of linear models, with particular reference to the inclusion of categorical variables, variable transformations, and models for non-continuous response variables. The fundamentals of logistic regression and the interpretation of its parameters and related indicators are introduced.
Theoretical topics are consistently complemented by practical applications using the R software environment, employing both real and simulated datasets to develop operational skills in statistical analysis and interpretation of results.
Prerequisites
Students are required to have successfully completed the prerequisite courses: Statistica I, Analisi Matematica I, Algebra Lineare, and Calcolo delle Probabilità.
For a smoother understanding of the course contents, a solid knowledge of the statistical inference topics covered in Statistica II is strongly recommended.
Teaching methods
The course consists of 49 hours of lectures and 12 hours of in-person practical sessions. Teaching activities are supported by a tutor assistant.
The lectures cover the theoretical foundations of statistical modeling, complemented by computer-based demonstrations and applied examples developed using the R software environment. The practical sessions take place in a computer laboratory and focus on the application of statistical analysis techniques to both real and simulated datasets.
Throughout the course, students are guided in developing statistical analyses, selecting appropriate methodologies, and interpreting the results. Course materials are made available on the e-learning platform before each lecture.
Assessment methods
The assessment consists of a written examination and an oral examination. No midterm assessments are scheduled.
The written examination is conducted using the R software environment and consists of the analysis of a dataset provided by the instructor. Students are required to perform a comprehensive statistical analysis, including the application of the models covered during the course, the production of graphs and tables, and the preparation of a report in R Markdown format. The report must include the interpretation of the results and a discussion of the main findings.
The oral examination covers the topics addressed during the course and is designed to assess students' theoretical understanding of the methods employed as well as their ability to critically interpret the results obtained.
In addition, a group project presentation session is scheduled, during which students present an independent statistical analysis based on real-world data that they have identified and collected on their own. The project is presented in class and contributes to the overall course assessment.
Textbooks and Reading Materials
Main Reference Texts
-
J. J. Faraway: Linear Models with R, 2nd ed., Chapman & Hall/CRC, 2014.
-
G. James, D. Witten, T. Hastie, R. Tibshirani: An Introduction to Statistical Learning, 2nd ed., Springer, 2021 (with particular emphasis on the chapters covering linear regression and logistic regression, including applications in R).
-
Carter Hill, William E. Griffiths, Guay C. Lim: Principles of Econometrics, 5th ed., Wiley, 2018.
-
D. Nolan, D. T. Lang: Data Science in R: A Case Studies Approach to Computational Reasoning and Problem Solving, Chapman & Hall/CRC, 2015.
-
Lecture notes and supplementary materials provided by the instructor during the course.
Semester
II Semester, III Cycle
Teaching language
Italian