A Robust Machine-Learning Framework for Predicting Tailpipe CO₂ Emissions from Vehicle Specifications

Main Article Content

T. Karthikeya, K. Amarendra

Abstract

Accurate estimation of vehicle CO₂ emissions is essential for the purpose of informing environmental policy and directing fleet management.  This study developed a reproducible machine-learning pipeline by utilizing Canada's 2018 light-duty vehicle dataset. The dataset consisted of 15 attributes, including engine displacement, cylinder count, gasoline type, transmission, and city/highway/combined fuel-consumption ratings.  We standardized numeric features through z-score normalization and applied one-hot encoding to categorical fields after eliminating incomplete records. Stratified 5-fold cross-validation guided hyperparameter optimization over five relapse calculations: straight relapse, choice tree, irregular timberland, XGBoost, and LightGBM, and the preprocessed information were separated into 70% preparing and 30% testing subsets. In arrange to progress the steadiness of the show and dispose of atypical sections, a Confinement Woodland was executed. The direct pattern had a unassuming RMSE of ~7 g/km, and untuned XGBoost failed to meet expectations. In separate, furnish methods—particularly self-assertive timberland and LightGBM—achieved overwhelming exactness on the held-out test set (RMSE < 4 g/km; R² > 0.98). When organized through an enthusiastic, end-to-end learning system, these disclosures portray that immediately accessible vehicle judgments can work as endeavored and veritable center individuals for debilitating spreads. The instrument that rises gives an adaptable arrangement for real-time outflow estimating and encourages data-driven emanation relief techniques.

Article Details

Section
Articles