# 🧠 Big Five Personality & Life Outcomes ### Does your personality decide your income, your job performance, or your happiness? The Five Core Traits - **O**penness to Experience: Measures your creativity, curiosity, and willingness to try new things versus preferring a routine. - **C**onscientiousness: Measures how organized, responsible, goal-directed, and disciplined you are versus being disorganized or spontaneous. - **E**xtraversion: Measures how much you are energized by the outside world, sociability, and assertiveness versus being reserved or solitary. - **A**greeableness: Measures your compassion, trust, cooperation, and desire for social harmony versus being skeptical or competitive. - **N**euroticism: Measures your emotional sensitivity and tendency to experience stress, anxiety, or negative emotions versus emotional stability. ![[summary-shap-share-heatmap.png|650]] > [!abstract] TL;DR > I trained three regression models on **50,000 people**, one per life outcome, and used **SHAP** to see what each model relies on. > > | Outcome | Model | Test RΒ² | Share of the model driven by Big Five | > |---|---|---:|---:| > | πŸ’° Income | Ridge | **0.40** | **18%** | > | πŸ“ˆ Job performance | Gradient Boosting | **0.22** | **48%** | > | 😊 Life satisfaction | Ridge | **0.25** | **87%** | > > - ==Personality barely explains **income**.== Education, age, parental SES and cognitive ability do most of the work. > - **Job performance** is split roughly in half: cognitive ability and income on one side, conscientiousness and the other traits on the other. > - **Life satisfaction** is almost entirely personality, and **neuroticism alone accounts for 36%** of it. > - Every model is **modest**: 60–78% of the variance in each outcome is left unexplained. As the dataset intro puts it, the answer is "less than you think." ## 1. The question > [!question] So what? > Most Big Five datasets are questionnaire responses with **no outcomes attached**. This one links the five traits to outcomes people care about, and it also includes the **confounders** most naive analyses leave out: parental SES, parental education and cognitive ability. The project asks one question three times: 1. How much of **income** can personality explain once background and ability are in the model? 2. What drives supervisor-rated **job performance**? 3. What drives **life satisfaction**? ## 2. The data | **Rows** | 50,000 people β†’ **42,500 train / 7,500 test** | | ------------------------- | ---------------------------------------------------------------------------------------------------------- | | **Raw columns** | 22 (traits, confounders, 8 life outcomes) | | **Features used per run** | 16 (the two other outcomes are included as features; see [[#⚠️ Outcomes predicting outcomes]]) | | **Big Five** | openness, conscientiousness, extraversion, agreeableness, neuroticism (each 0–100) | | **Confounders** | parental_ses (parents social status) (0–100), parental_education_years, cognitive_score (mean 100, SD 15) | | **Outcomes** | income_usd, job_performance (1–5), life_satisfaction (1–10), gpa, self_rated_health, exercise_days_week, … | > [!info]- Full data dictionary > | Column | Type | Description | > |---|---|---| > | person_id | string | Unique person id | > | age | int | Age in years | > | sex | string | M / F | > | openness | float | Openness to Experience (0–100) | > | conscientiousness | float | Conscientiousness (0–100) | > | extraversion | float | Extraversion (0–100) | > | agreeableness | float | Agreeableness (0–100) | > | neuroticism | float | Neuroticism (0–100) | > | parental_ses | float | Parental socioeconomic status composite (0–100), a **confounder** | > | parental_education_years | int | Parents' years of education | > | cognitive_score | int | Cognitive ability (mean 100, SD 15), a **confounder** | > | education_years | int | Years of education attained | > | has_degree | int | 1 if 16+ years of education | > | gpa | float | Academic GPA (0–4) | > | job_performance | float | Supervisor-rated job performance (1–5) | > | income_usd | float | Annual income (USD) | > | high_income | int | 1 if income is in the top quartile | > | life_satisfaction | float | Life satisfaction (1–10) | > | exercise_days_week | int | Days per week exercising | > | smoker | int | 1 if smoker | > | self_rated_health | float | Self-rated health (1–5) | > | partnered | int | 1 if in a committed relationship | > | ever_divorced | int | 1 if ever divorced (0 for under-28s) | ## 3. How the pipeline works ```mermaid flowchart LR A[(Raw CSV<br/>50k Γ— 22)] --> B[EDA<br/>distributions Β· correlation] B --> C[Preprocess<br/>scale numeric Β· one-hot sex] C --> D[Train / test split<br/>42.5k / 7.5k] D --> E[Fit model<br/>LinReg Β· Ridge Β· GBM] E --> F[Evaluate<br/>RMSE Β· MAE Β· RΒ² Β· learning curve] F --> G[Interpret<br/>coefficients Β· permutation Β· SHAP] G --> H[(runs/&lt;model&gt;_&lt;target&gt;/)] ``` Each run is saved to `runs/<model>_<target>/` with a `model/` folder (joblib, test predictions, metadata), an `interpretation/` folder (CSVs and plots), and a `run_report.md` log. > [!note]- Methods used (click to expand) > **Ridge regression** shrinks coefficients with an L2 penalty: > $\hat\beta=\arg\min_\beta \sum_i (y_i - x_i^\top\beta)^2 + \alpha\lVert\beta\rVert_2^2,\qquad \alpha=100$ > **Gradient boosting** adds shallow trees one at a time, each fitted to the current residuals: 200 trees, depth 2, learning rate 0.1. > > **Metrics** on the held-out test set: > $\text{RMSE}=\sqrt{\tfrac1n\sum(y_i-\hat y_i)^2}\qquad R^2 = 1-\frac{\sum(y_i-\hat y_i)^2}{\sum(y_i-\bar y)^2}$ > > **SHAP** splits each prediction into additive feature contributions: > $\hat y_i = \mathbb{E}[\hat y] + \sum_{j}\phi_{ij}$ > Global importance = mean $|\phi_{ij}|$ over a 500-row sample. The heatmap above shows each feature's **share** of that total within its model. > > **Permutation importance** measures how much the score drops when one feature is shuffled. > > Related notes: [[Regression]] Β· [[Tree-based Models]] Β· [[Feature Selection]] Β· [[Exploratory Data Analysis]] ## 4. Exploratory analysis ![[eda-correlation-heatmap.png|600]] > [!tip] What the correlation matrix already says > - **Income** correlates most with education_years (**0.44**), parental_ses (**0.36**), cognitive_score (**0.33**) and job_performance (**0.31**). Its correlations with the Big Five are all ≀ 0.15. > - **Life satisfaction** correlates most with neuroticism (**βˆ’0.38**) and extraversion (**0.22**). > - The **Big Five are almost uncorrelated with one another** (|r| ≀ 0.01). That keeps multicollinearity low and makes each trait's coefficient easy to interpret. > - Watch the confounder chain: cognitive_score ↔ gpa (**0.47**) and ↔ education (**0.36**); parental_ses ↔ education (**0.37**). > [!example]- Distributions & categorical checks > ![[eda-numerical-boxplots.png|650]] > - **income_usd is strongly right-skewed** (long tail of outliers above ~$100K, up to about $272K). This comes back in the residuals. > - life_satisfaction has a tail of low scores. job_performance is bounded between 1 and 5. > > ![[eda-sex-distribution.png|320]] ![[eda-income-by-sex.png|320]] > - Sex split is balanced (51.1% F / 48.9% M). Mean income differs only slightly: **$51,659 (M)** vs **$50,248 (F)**. ## 5. Results at a glance ![[summary-model-scorecard.png|700]] | Run | Target | Model | Test RMSE | Test MAE | Test RΒ² | RMSE if predicting the mean | Improvement | |---|---|---|---:|---:|---:|---:|---:| | `ridge_income_usd` | income_usd | Ridge (Ξ±=100) | **$17,614** | $13,070 | **0.396** | $22,672 | 22.3% | | `gradient_boosting_job_performance` | job_performance | GBM | **0.704** | 0.565 | **0.217** | 0.795 | 11.5% | | `ridge_life_satisfaction` | life_satisfaction | Ridge (Ξ±=100) | **1.386** | 1.112 | **0.251** | 1.602 | 13.5% | > [!success] No overfitting anywhere > Training and validation RMSE are almost equal in every run, and the learning curves flatten out well before the full training set. **Adding more rows won't help.** The limit is in the features: much of each outcome depends on things this dataset doesn't measure. > [!warning] Predictions regress toward the mean > Predictions are much narrower than the actual outcomes. For example, life satisfaction ranges 1–10 in reality but only **3.9–9.3** in the predictions. Mean residual by actual-value quintile: > > | Quintile of actual | Income | Job perf. | Life sat. | > |---|---:|---:|---:| > | Lowest 20% | βˆ’$12,341 | βˆ’0.87 | βˆ’1.74 | > | Middle | βˆ’$4,152 | +0.01 | +0.05 | > | Highest 20% | **+$22,457** | +0.89 | +1.74 | > > The models **over-predict low values and under-predict high ones**. This is expected when RΒ² is modest, but it means the models shouldn't be used to find extreme individuals. ## 6. Run 1 β€” πŸ’° What drives income? (`ridge_income_usd`) > [!summary] Verdict > Income is mainly about **education, age (career stage), family background and cognitive ability**. The Big Five together account for only **~18%** of the model's SHAP importance. ### Coefficients (per 1 SD increase, standardized features) | Rank | Feature | Effect on income | Mean \|SHAP\| | |---:|---|---:|---:| | 1 | education_years | **+$5,533** | $4,465 | | 2 | age | **+$5,023** | $4,252 | | 3 | parental_ses | **+$4,959** | $3,848 | | 4 | cognitive_score | **+$3,696** | $2,972 | | 5 | job_performance | +$3,215 | $2,514 | | 6 | openness | +$1,806 | $1,496 | | 9 | conscientiousness | +$1,505 | $1,176 | | 10 | agreeableness | ==**βˆ’$1,295**== | $993 | | 14 | neuroticism | βˆ’$256 | $198 | ![[income-ridge-shap-beeswarm.png|550]] > [!tip] Findings > 1. **Credentials and background win.** The top four drivers aren't personality traits, and together they make up ~59% of the model's SHAP importance. > 2. **Parental SES is worth almost as much as your own age**: about +$5K per SD. This is the confounder the dataset was built to expose. > 3. **Conscientiousness, often called the strongest personality predictor, ranks only 9th for income** once education, ability and background are controlled for. > 4. **Agreeableness carries a penalty** (βˆ’$1,295 per SD). Agreeable people earn slightly less when everything else is held constant. > 5. **Neuroticism barely matters for income** (0.8% of SHAP importance), even though it dominates life satisfaction (see Run 3). > [!example]- SHAP dependence plots, permutation importance & one worked prediction > ![[income-ridge-shap-all-relationships.png|700]] > Since Ridge is linear, every SHAP relationship is a straight line. Its slope is the coefficient. One thing to check: **education_years looks flat at the low end**, which suggests the preprocessing clips low values. > > ![[income-ridge-permutation-importance.png|500]] > > **One person's prediction (row 0):** the baseline is $52,888. Education at +1.6 SD adds **+$8,281**, low agreeableness (βˆ’1.8 SD) adds **+$2,259**, and below-average cognitive score and conscientiousness subtract about $2.3K and $1.9K. The final prediction is **$57,092**. > ![[income-ridge-shap-row0.png|500]] > [!failure]- Diagnostics: the residuals are skewed > ![[income-linreg-qq-plot.png|420]] ![[income-linreg-actual-vs-predicted.png|420]] > - Residual skew **1.36**, excess kurtosis **5.35**. The Q-Q plot bends upward in the right tail, and actual incomes above ~$120K are under-predicted. > - The plain **linear regression baseline** reached about **$17,500** RMSE on train and validation, close to Ridge's $17,614 on test. With 42,500 rows and 16 features, regularization adds little. > - **Next step:** model `log(income_usd)` or use a tree model to handle the right tail. ## 7. Run 2 β€” πŸ“ˆ What drives job performance? (`gradient_boosting_job_performance`) > [!summary] Verdict > Job performance is about **half personality and half ability and pay**. **Cognitive score (20%)** and **income (19%)** lead, followed by **conscientiousness (15%)**, and all five traits play a part. | Rank | Feature | Mean \|SHAP\| | Share | Permutation importance | |---:|---|---:|---:|---:| | 1 | cognitive_score | 0.133 | 20.2% | 0.0419 | | 2 | income_usd | 0.126 | 19.2% | 0.0391 | | 3 | conscientiousness | 0.096 | 14.6% | 0.0241 | | 4 | extraversion | 0.063 | 9.5% | 0.0104 | | 5 | agreeableness | 0.060 | 9.1% | 0.0096 | | 6 | neuroticism | 0.057 | 8.6% | 0.0112 | | 7 | openness | 0.041 | 6.2% | 0.0062 | ![[jobperf-gb-shap-beeswarm.png|550]] > [!tip] Findings from the SHAP dependence plots > 1. **Cognitive score, conscientiousness, extraversion, agreeableness and openness** all raise predicted performance almost linearly. > 2. **Income has diminishing returns.** It rises steeply up to about +1 SD, then flattens. > 3. **Neuroticism works in the opposite direction**: higher neuroticism means lower predicted performance. > 4. ==**Age and parental SES have *negative* conditional effects.**== Older workers and people from higher-SES families are rated slightly lower *once income is known*. A likely reason: at a given income, a person who is older or from a wealthier family is doing comparatively less well for their pay. This comes from using income as a feature, not a causal effect. > 5. **Education works like a step**: flat, then a jump at the upper end, consistent with a degree threshold. > 6. **Sex contributes exactly 0**. The boosted trees never split on it. > [!example]- All SHAP dependence plots & diagnostics > ![[jobperf-gb-shap-all-relationships.png|700]] > ![[jobperf-gb-actual-vs-predicted.png|420]] ![[jobperf-gb-learning-curve.png|420]] > - Train RMSE β‰ˆ 0.69 vs validation β‰ˆ 0.70, so the depth-2 trees are well regularized. > - The target clusters at its **floor (1.0) and ceiling (5.0)**, which no regression model can predict well. An ordinal or censored model would suit it better. > ![[jobperf-gb-permutation-importance.png|500]] ## 8. Run 3 β€” 😊 What drives life satisfaction? (`ridge_life_satisfaction`) > [!summary] Verdict > Life satisfaction is **almost entirely personality (87%)**. **Neuroticism alone is 36%**, twice as much as the next trait. Background and ability contribute close to nothing. | Rank | Feature | Coefficient (per SD) | Mean \|SHAP\| | Share | | ---: | --------------------------------------- | -------------------: | ------------: | ----------: | | 1 | neuroticism | ==**βˆ’0.573**== | 0.443 | 35.9% | | 2 | extraversion | **+0.329** | 0.259 | 21.0% | | 3 | agreeableness | +0.216 | 0.166 | 13.5% | | 4 | conscientiousness | +0.186 | 0.146 | 11.8% | | 5 | income_usd | +0.150 | 0.115 | 9.3% | | 6 | openness | +0.074 | 0.062 | 5.0% | | β€” | cognitive_score Β· education Β· gpa Β· age | β‰ˆ 0 | < 0.003 | < 0.3% each | ![[lifesat-ridge-shap-beeswarm.png|550]] > [!tip] Findings > 1. **One SD more neuroticism costs about 0.57 points** on the 1–10 scale, more than any other feature. > 2. **Extraversion is the biggest positive factor** (+0.33 per SD). Agreeableness and conscientiousness also help. > 3. **Money helps, but modestly**: +0.15 per SD of income, less than half of extraversion's effect. > 4. **Being smart or educated doesn't make people happier here.** Cognitive score, education years and GPA all have essentially zero weight. > 5. **Job performance barely matters** (0.9%), which suggests performing well at work and being satisfied with life are separate things. > [!example]- Dependence plots, coefficients & diagnostics > ![[lifesat-ridge-shap-all-relationships.png|700]] > ![[lifesat-ridge-coefficients.png|500]] > ![[lifesat-ridge-actual-vs-predicted.png|420]] ![[lifesat-ridge-learning-curve.png|420]] > - The learning-curve gap is tiny (β‰ˆ 1.377 vs 1.378 RMSE at full size), so the model has converged. > - There's a visible **ceiling spike at 10**. Predictions never go above ~9.3. ## 9. Cross-outcome takeaways ```mermaid pie showData title Big Five share of each model's SHAP importance (%) "Income" : 17.9 "Job performance" : 48.0 "Life satisfaction" : 87.2 ``` > [!quote] The bigger picture > **How much personality matters depends on the outcome.** The more *external* the outcome (pay), the more background and credentials dominate. The more *internal* the outcome (how satisfied you feel), the more it comes down to temperament. | Trait | Income | Job performance | Life satisfaction | |---|:---:|:---:|:---:| | Neuroticism | ~0 | ↓ | ⬇⬇⬇ strongest | | Extraversion | ↑ small | ↑ | ⬆⬆ | | Agreeableness | ↓ penalty | ↑ | ⬆ | | Conscientiousness | ↑ small | ⬆ top trait | ⬆ | | Openness | ↑ small | ↑ | ↑ small | Agreeableness has the most interesting pattern: it **lowers income but raises job performance and life satisfaction**. ## 10. Caveats > [!warning] ⚠️ Outcomes predicting outcomes > Each run uses the other two targets as **features**: income is predicted from job_performance and life_satisfaction, job performance from income, and so on. That improves fit, but: > - these relationships likely run in **both directions** (pay ↔ performance), so the coefficients are **not causal** > - it causes odd conditional effects, such as the negative age effect on job performance > > **Next run:** refit using only traits and confounders, the variables that exist *before* the outcomes, and compare RΒ². > [!warning]- Other limitations > - **Observational data.** Every finding is an association, even with confounders included. > - **Skewed and bounded targets.** Income needs a log transform. Job performance (1–5) and life satisfaction (1–10) have floor and ceiling spikes. > - **SHAP sample size.** Importances come from 500 rows, which is stable for rankings but approximate for exact shares. > - **One model per target.** Ridge was used for income and life satisfaction and GBM for job performance, so the RΒ² values across targets aren't a strict comparison of the same model.