Logo do repositório
 
A carregar...
Miniatura
Publicação

Comparative evaluation of machine learning and deep learning strategies for Port wine authentication and age prediction by FT-IR Spectroscopy and chemometric Modelling

Utilize este identificador para referenciar este registo.
Nome:Descrição:Tamanho:Formato: 
156408427.pdf796.87 KBAdobe PDF Ver/Abrir

Orientador(es)

Resumo(s)

Port wine authentication and age validation have significant reputational and economic implications. However, their analytical assessment remains a major challenge, as it is still based mainly on sensory analysis and complementary metadata from chemical profiling, ranging from molecular markers to, more recently, radiocarbon-IMS (C¹⁴) measurements. Nevertheless, no standard practice is currently prescribed by regulatory authorities. This study presents a spectroscopic modelling framework for Port wine age prediction, supported by a systematic chemometric evaluation of sixteen machine learning (ML) and deep learning (DL) approaches applied to mid-infrared Fourier Transform Infrared (FT-IR) spectral data. Tabular data still lacks a one-fits-all model regarding DL architectures. A dataset of 7,281 FT-IR spectra covering four Port wine categories - “Colheita”, “Tawny”, Late Bottled Vintage (LBV), and “Vintage”- was preprocessed using Standard Normal Variate (SNV) transformation and mean centering applied exclusively from training partitions to prevent information leakage; Venetian Blind cross-validation was implemented to account for sequential autocorrelation inherent in spectroscopic datasets. Sixteen modelling strategies were benchmarked, spanning classical chemometrics (PLS, LDA, SVM, regularized regression, Random Forest, Gradient Boosting) and deep learning architectures (1D CNNs, Multi-scale CNNs with Inception-style parallel branches, residual MLP, and Deep Ensembles). CNNs with Inception-style parallel branches, residual MLP combined with Deep Ensembles were selected due to its ability of offering some interpretability, breaking the “Black-box” concept regarding DL models. For wine type classification, Support Vector Machine with an RBF kernel achieved the highest accuracy (95.88%, ROC-AUC 0.9945), with Random Forest providing comparable performance (93.54%, ROC-AUC 0.9936) alongside explicit featureimportance maps identifying discriminative spectral regions. For age regression, a Deep Ensemble of five residual MLP models attained R² = 0.9803 (MAE = 1.25 years), substantially outperforming the best traditional chemometric method, Gradient Boosting (R² = 0.8401, MAE = 2.94 years). A novel Hybrid Two-Stage Random Forest pipeline - combining wine-type classification with type-specific age regressors - achieved R² = 0.9630 (MAE = 1.28 years) without GPU requirements. A leverage- based individual confidence interval method provided well-calibrated prediction uncertainty with 97.7% empirical coverage. Chemical interpretability was confirmed by convergence of feature importance patterns across methods with established wine ageing biochemistry, validating that models capture genuine chemical variation. The structure of the data strongly influences the choice of modelling architecture: while deep learning has demonstrated outstanding performance for image and natural language tasks, spectroscopic tabular data still lacks universally accepted one-size-fits- all architecture, leaving the optimal model design context-dependent on dataset size, interpretability requirements, and available computational resources.

Descrição

Palavras-chave

Port wine authentication Age prediction Machine learning Deep learning Ensemble methods Interpretability

Contexto Educativo

Citação

Projetos de investigação

Unidades organizacionais

Fascículo

Editora

SSRN

Licença CC

Sem licença CC

Métricas Alternativas