← Back to all projects

Machine Learning / Data Science Research

Nov 2024 - Dec 2024

Wildfire Size Prediction

Cleaned and prepared 38,750 wildfire records after dataset filtering

Performed exploratory analysis on wildfire causes, climate trends, vegetation impact, and spatial distribution

Applied feature engineering on weather, time, geographical, and fire characteristics

Implemented classification models including MLP, DNN, and GCN-LSTM

Implemented regression models including CNN-LSTM, GP-LSTM, and GMM-stacking

Optimized models using Optuna, Bayesian Optimization, Random Search, and Hyperband

Deployed prediction interface using Streamlit

Wildfire Size Prediction interface preview

The problem

Wildfire prediction remains challenging because wildfire behaviour depends on complex interactions between climate conditions, vegetation, location, and human activities. Existing AI models also face limitations in explainability, data availability, and integration of diverse environmental factors (Taylor et al., 2013; Bugallo et al., 2022).

The solution

Processed a wildfire dataset containing 55,367 records and 43 attributes by filtering data from 2000–2015, handling missing values, engineering temporal and environmental features, and preserving meaningful extreme wildfire events. Multiple classification and regression models were developed and compared, including MLP, DNN, GCN-LSTM, CNN-LSTM, GP-LSTM, and GMM-stacking models.

System flow

How it works

Kaggle Dataset
Data Cleaning
Feature Engineering
Exploratory Analysis
Model Training
Hyperparameter Optimization
Model Evaluation
Streamlit Deployment

Tech stack

Python Pandas NumPy Scikit-learn PyTorch TensorFlow Optuna Streamlit Matplotlib Seaborn