Predicting Falcon 9 Reusability
Classification pipeline for first-stage landing outcomes
A data science workflow predicting whether a Falcon 9 first stage will land successfully, using feature engineering, exploratory analysis, and model comparison.
Built with
- Python
- Pandas
- Scikit-learn
- Requests
- REST API
The workflow
From launch data to landing prediction.
One experiment, from historical launch characteristics to a consistently evaluated classification result.
- 01
Explore
Class balance · Booster reuse
Inspect landing outcomes, booster reuse, and launch-site patterns to identify signals that matter for modeling.
- 02
Split
Stratified holdout
Compare non-stratified and stratified 75/25 splits to understand how test-set distribution changes evaluation.
- 03
Model
Logistic Regression · KNN
Establish a majority-class baseline, use Logistic Regression as an additive benchmark, and test KNN as a flexible local alternative.
- 04
Evaluate
Accuracy · F1 · Confusion Matrix
Compare models on the same holdout set and examine minority-class behavior rather than relying on accuracy alone.
01 / Intent
Predict more than the majority class.
Reusable first-stage boosters can reduce launch costs, but landing outcomes depend on launch profile, payload, orbit, and historical flight context. The project frames landing prediction as a supervised classification problem.
02 / Implementation
From EDA to model comparison.
Automated data ingestion, engineered mission-level features, built Scikit-learn classification pipelines, and applied stratified sampling for class imbalance.
03 / Result
KNN performed best on this holdout.
Improved accuracy from 86% to 98% and reached a 96% F1 score while keeping the workflow reproducible.
Current boundary
Strong predictive workflows need both model performance and trustworthy data lineage. Careful feature definitions mattered as much as the final classifier.
Split strategy is part of the experiment.
High accuracy alone can hide poor minority-class performance.
