Skip to content

About

Machine learning analysis of speed-dating response prediction using HistGradientBoosting and relationship-based feature engineering.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Speed Dating Response Prediction Using HistGradientBoosting

Project Overview

This project applies machine learning to examine how self-perception, dating preferences, and beliefs about potential partners are associated with the likelihood of receiving a positive response during a speed-dating interaction.

The project uses the Columbia Business School SpeedDating dataset, which contains 8,378 participant-partner interactions and 195 original features collected across 21 speed-dating event waves. The dataset includes participant demographics, interests, dating preferences, self-ratings, beliefs about what potential partners value, partner evaluations, and dating decisions.

A substantial exploratory data analysis and preprocessing workflow was used to identify relevant features, address missing data, standardize differences in survey methodology, engineer relationship-based features, and prevent information leakage between participants.

Multiple classification approaches were evaluated. The final model uses scikit-learn's HistGradientBoostingClassifier because it provided the strongest combination of balanced classification performance, recall, training efficiency, and native handling of missing values.

Recognition: This work received a WGU Excellence Award for exceptional performance in the Advanced AI and ML course project.

Project Objectives

  • Predict whether an individual will receive a positive dating response from a partner.
  • Examine how self-perception and dating preferences relate to dating outcomes.
  • Engineer features representing compatibility, preference similarity, and belief accuracy.
  • Prevent information leakage between participants appearing within the same speed-dating events.
  • Compare baseline, standard gradient boosting, and histogram-based gradient boosting models.
  • Optimize model performance while retaining features important to the research question.
  • Use partial-dependence analysis to identify relationships between participant characteristics and model predictions.

Results

Several classification approaches were evaluated during model development and optimization:

Model Accuracy Balanced Accuracy Precision Recall F1-Score ROC-AUC PR-AUC
Majority Baseline 0.5883 0.5000 0.0000 0.0000 0.0000 0.5000 0.4117
Logistic Regression 0.5751 0.5445 0.4792 0.3712 0.4184 0.5756 0.4712
Default Gradient Boosting 0.5928 0.5618 0.5070 0.3863 0.4385 0.6033 0.5195
Optimized Gradient Boosting 0.6087 0.5846 0.5291 0.4485 0.4855 0.6142 0.5235
HistGradientBoosting 0.5901 0.5966 0.5017 0.6330 0.5598 0.6145 0.5178

Balanced accuracy was used as the primary model-selection metric. HistGradientBoosting achieved the highest validation balanced accuracy, recall, and F1-score while also substantially reducing training time compared with the optimized standard GradientBoostingClassifier.

Median training time decreased from approximately 4.56 seconds to 0.57 seconds, representing an approximately eightfold improvement.

Final HistGradientBoosting Performance

After the final model configuration and classification threshold were frozen, the model was retrained using the full training partition and evaluated once against the previously untouched holdout test set.

  • Accuracy: 0.6098
  • Balanced Accuracy: 0.6160
  • Precision: 0.5348
  • Recall: 0.6583
  • F1-Score: 0.5902
  • ROC-AUC: 0.6569
  • PR-AUC: 0.5664
  • Classification Threshold: 0.40

The final model demonstrated modest but meaningful predictive performance. Its relatively high recall allowed it to identify approximately 66% of positive dating responses, while both ROC-AUC and PR-AUC exceeded their respective no-skill baselines.

Confusion Matrix

Final Model Confusion Matrix

ROC Curve

Final Model ROC Curve

Precision-Recall Curve

Final Model Precision-Recall Curve

Prediction Analysis

Partial-dependence analysis was used to examine how selected participant characteristics were associated with the model's predictions.

Several notable patterns emerged:

  • Greater mismatch between a participant's beliefs about a partner's preferences and the partner's actual preferences was generally associated with a lower predicted probability of receiving a positive response.

  • Underestimating the importance a partner placed on intelligence was associated with more favorable predictions than accurate estimation or overestimation.

  • Substantial underestimation of the importance of attractiveness was associated with above-average predictions for both genders, with the relationship particularly pronounced for men receiving decisions from women.

Attractiveness Belief Gap Partial Dependence

Example partial-dependence result showing the relationship between attractiveness belief gaps and model predictions by gender. Additional relationships are explored in Final_Model.ipynb.

  • Higher self-perception was not consistently associated with greater predicted desirability.

  • Partner-weighted self-perception showed different relationships with predicted outcomes depending on gender.

These results should be interpreted as associations identified by the model rather than causal relationships.

Key Features

  • Extensive exploratory data analysis across 195 original variables
  • Missing-value analysis and feature-quality screening
  • Standardization of survey responses collected using different rating methodologies
  • Group-based train, validation, and test splitting by experimental wave
  • Participant-level leakage prevention
  • Training-dependent feature construction
  • Feature engineering based on participant and partner relationships
  • Preference similarity measures
  • Belief-accuracy and belief-gap features
  • Self-assessment gap features
  • Interest compatibility features
  • Partner-weighted self-perception
  • Protected research-relevant feature selection
  • Gradient boosting model comparison
  • Hyperparameter optimization
  • Native missing-value handling using HistGradientBoosting
  • Classification-threshold optimization
  • Partial-dependence interpretability analysis
  • Computational efficiency benchmarking

Technologies Demonstrated

  • Machine Learning
  • Binary Classification
  • Gradient Boosting
  • HistGradientBoosting
  • Feature Engineering
  • Data Cleaning and Validation
  • Missing-Value Handling
  • Group-Based Data Splitting
  • Leakage Prevention
  • Hyperparameter Optimization
  • Classification Threshold Optimization
  • Model Evaluation
  • Partial Dependence Analysis
  • Model Interpretability

Dataset

The project uses the Columbia Business School SpeedDating dataset, which contains 8,378 observations and 195 features collected during speed-dating experiments conducted between 2002 and 2004.

Each observation represents an interaction between a participant and a dating partner. Available information includes:

  • Participant demographics
  • Interests and hobbies
  • Self-perception ratings
  • Desired partner characteristics
  • Beliefs about what potential partners value
  • Evaluations of dating partners
  • Expectations about dating outcomes
  • Partner decisions

The dataset was collected across 21 experimental waves.

During preprocessing, features were organized into structural, self-focused, partner-focused, leakage-prone, and unnecessary categories. Features with extensive missingness were reviewed for removal, while features important to the research question were retained when appropriate.

Differences in rating methodology across experimental waves were standardized where necessary.

The final model used 67 selected features, including 30 features deliberately protected during feature selection because of their relevance to the research question.

Model Architecture

Speed Dating Interaction
          ↓
 Participant Responses
 Partner Preferences
 Self-Perception
 Interests
          ↓
 Feature Engineering
          ↓
┌─────────────────────────────┐
│  HistGradientBoosting       │
│  Classifier                 │
│                             │
│ • 100 Boosting Iterations   │
│ • Learning Rate = 0.10      │
│ • Max Leaf Nodes = 15       │
│ • Max Depth = None          │
│ • Min Samples Leaf = 20     │
│ • Max Bins = 127            │
│ • L2 Regularization = 0.0   │
│ • Native NaN Handling       │
└─────────────────────────────┘
          ↓
 Predicted Probability
          ↓
 Threshold = 0.40
          ↓
 Positive / Negative Response

Repository Structure

Speed_Dating_HistGradientBoosting/
├── README.md
├── data/
│   ├── Speed Dating Data.csv
│   ├── gb_Speed_Dating_EDA.csv
│   ├── gb_Speed_Dating_candidate_features.json
│   ├── gb_processed_feature_names.json
│   ├── gb_X_train.csv
│   ├── gb_X_test.csv
│   ├── gb_y_train.csv
│   └── gb_y_test.csv
├── images/
│   ├── confusion_matrix.png
│   ├── roc_curve.png
│   ├── precision_recall_curve.png
│   └── attractiveness_belief_gap_partial_dependence.png
├── notebooks/
│   ├── EDA.ipynb
│   ├── Preprocessing.ipynb
│   ├── Initial_Model.ipynb
│   ├── Optimization.ipynb
│   └── Final_Model.ipynb
└── results/
    ├── gb_final_hist_gradient_boosting_model.joblib
    ├── gb_final_model_metadata.json
    ├── gb_final_permutation_importance.csv
    ├── gb_selected_features_67.json
    └── gb_selected_model_config.json

Dependencies

This project was developed using Python and the following primary libraries:

  • pandas
  • NumPy
  • scikit-learn
  • Matplotlib
  • joblib

The analysis was developed and documented using Jupyter notebooks.

Usage

The notebooks are organized to follow the complete model-development workflow and should be run in the following order:

  1. EDA.ipynb — Performs exploratory data analysis, data-quality assessment, cleaning, and initial feature engineering.
  2. Preprocessing.ipynb — Creates group-based data partitions, constructs training-dependent features, and prepares data for modeling.
  3. Initial_Model.ipynb — Establishes baseline performance and develops the initial GradientBoostingClassifier.
  4. Optimization.ipynb — Evaluates feature selection, compares gradient-boosting approaches, optimizes the HistGradientBoostingClassifier, and selects the final classification threshold.
  5. Final_Model.ipynb — Retrains the frozen model configuration, evaluates it against the untouched holdout set, and performs model-interpretability analysis.

Intermediate datasets are stored in data/, while model configurations, feature-selection results, and the final trained model are stored in results/.

Future Improvements

Potential extensions to this project include:

  • Validate the model using additional or more recent speed-dating datasets to assess generalizability.
  • Compare HistGradientBoosting with other high-performance boosting frameworks such as XGBoost, LightGBM, and CatBoost.
  • Incorporate SHAP-based explanations to complement the existing partial-dependence analysis.
  • Evaluate probability calibration to improve interpretation of predicted response probabilities.
  • Conduct additional subgroup and fairness analysis across participant characteristics.
  • Explore interaction effects between self-perception, partner preferences, and belief-accuracy features.
  • Develop an interactive dashboard for exploring model predictions and relationships between dating preferences and outcomes.

About

Machine learning analysis of speed-dating response prediction using HistGradientBoosting and relationship-based feature engineering.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages