Airline Passenger Satisfaction
A supervised-learning notebook over a 130k-row airline passenger survey: cleaning with the reasons written down, PCA and K-means segmentation, then five classifiers tuned by grid search and compared on accuracy, macro F1 and the train-test gap.
The problem
Project 3 of 3 for the Artificial Intelligence course at the Universitat de les Illes Balears — the one on supervised learning.
The dataset is a passenger survey of 103,904 training rows and 25,976 test rows, 24 columns wide. Some describe the passenger and the trip — age, customer type, type of travel, class, distance, delays — and the rest are 1–5 ratings of individual services: wifi, online boarding, seat comfort, entertainment, leg room, check-in, cleanliness. The label is binary: satisfied, or neutral / dissatisfied.
Two questions are asked of it, and they need different tools. The first is predictive — can satisfaction be predicted, and how well? The second is descriptive — what actually drives it, and are there distinct kinds of passenger hiding in the data? The second is answered first, because a model that scores well without telling you why is only half an answer.
How it works
- Cleaning, with the reasons written down —
GenderandArrival Delay in Minutesare both dropped, each after a section that justifies it rather than asserting it: gender splits satisfaction almost identically, and the two delay columns correlate at 0.97, so keeping both only adds multicollinearity. - Preparation — one-hot encoding for the categorical columns, then
StandardScalerfitted on train alone and applied to test, so nothing about the test set leaks in early. - Segmentation —
PCA(n_components=0.95)first, which keeps 17 components for 95.7% of the variance, then K-means with K = 4 on the projection. The four clusters are profiled and named rather than left as numbers: a disloyal dissatisfied segment, a satisfied business VIP, a loyal economy business traveller hit by service failures, and the leisure passenger. - Five classifiers, one protocol — Perceptron, logistic regression, decision tree, random forest and SVM, each tuned with
GridSearchCVand reported the same way. The SVM grid runs on a stratified 30,000-row subset because the full search does not finish in reasonable time, and the winner is then refitted on everything. - A comparison that looks past accuracy — test accuracy, macro and weighted F1, the train–test gap as a proxy for overfitting, and per-class behaviour.
| Model | Accuracy | F1 macro | Train–test gap |
|---|---|---|---|
| Random forest | 0.9631 | 0.9624 | 0.0284 |
| SVM | 0.9606 | 0.9599 | 0.0076 |
| Decision tree | 0.9556 | 0.9547 | 0.0028 |
| Logistic regression | 0.8714 | 0.8688 | 0.0033 |
| Perceptron | 0.8322 | 0.8304 | 0.0053 |
The two linear models sit about nine points below the rest, which is the clearest finding in the whole notebook: service ratings and satisfaction are not linearly separable, and the models with the capacity to bend around that take all of the margin. Random forest wins on raw performance and pays for it with the largest gap; the decision tree gives up a point of accuracy for the steadiest behaviour of the five.
Both tree models agree on which variable matters most, and it is not the one you would guess from an airline: online boarding, ahead of wifi and entertainment.
Running it
Requires Python 3.10 or newer with pandas, scikit-learn, matplotlib, seaborn and mlxtend, all pinned in the repository’s requirements.txt. Open the notebook from its own folder so the relative data/ paths resolve:
pip install -r requirements.txt
cd P3
jupyter lab "Predicting Airline Passenger Satisfaction.ipynb"
Then run the cells in order. The grid searches are the slow part, the SVM one most of all, even on the reduced subset.