# πβ‘ Who Will Buy an EV?
### Subsidies open the door. Environmental concern walks through it.
![[ev-subsidy-concern-gate.png]]
> [!abstract] TL;DR
> I trained a gradient boosting classifier on **100,000 people** to predict whether someone **will buy an electric vehicle**, then used SHAP to see what drives the prediction.
> - **Two factors dominate.** Without a **subsidy**, almost nobody buys (**0.6%**). With one, **environmental concern** takes buy rates from **0.9%** up to **69.5%**.
> - **Income** is the third lever. **Range anxiety** can override everything else: people with high range anxiety almost never buy (**0.3%**).
> - **Car type, gender and number of cars owned barely matter.**
> - The model is strong at **ranking** people: ROC-AUC **0.94**, PR-AUC **0.75** against a 0.175 baseline. The **top 20%** of scores contain **76%** of all buyers.
| At a glance | |
|---|---|
| **Question** | Who will buy an EV, and what moves them? |
| **Data** | 100,000 people Β· 13 features Β· 17.5% buyers |
| **Model** | Gradient boosting (300 trees, depth 5, learning rate 0.05) |
| **Split** | 85,000 train / 15,000 test |
| **Headline metric** | ROC-AUC **0.940** Β· PR-AUC **0.752** Β· top-decile lift **4.6Γ** |
## 1. The question
> [!question] Who's ready to switch?
> EV makers, dealers and policymakers all face the same targeting problem: only about **1 in 6** people in this data plan to buy. Sending incentives, test drives or ads to everyone wastes most of the budget.
I wanted to answer two things:
1. **Prediction:** can we rank people by how likely they are to buy?
2. **Explanation:** which levers matter most: money, values, infrastructure, or practical worries like range?
## 2. The data
| Group | Features |
|---|---|
| π€ **Demographics** | Age, Gender, Annual_Income_USD, City_Type (Urban / Suburban / Rural) |
| π **Current situation** | Current_Car_Type (Sedan / SUV / Hatchback / Truck), Number_of_Cars_Owned, Daily_Commute_km |
| π **Charging access** | Charging_Stations_Near_Home, Charging_Stations_Near_Work, Home_Charging_Possible |
| π **Attitudes** | Environmental_Concern_Level (1β5), Range_Anxiety_Level (Low / Medium / High) |
| π΅ **Policy** | Subsidy_Available |
| π― **Target** | **Will_Buy_EV** (True / False) |
> [!info]- Class balance
> ![[ev-eda-target-distribution.png|420]]
> Only **17.5%** of people plan to buy an EV (17,454 of 100,000). Because the classes are imbalanced, **accuracy alone is misleading**: predicting "no" for everyone would already score 82.5%. That's why this post focuses on **ROC-AUC, PR-AUC, and precision and recall for buyers**.
## 3. How I approached it
```mermaid
flowchart LR
A[(BigQuery table<br/>100k Γ 15)] --> B[EDA<br/>buy rates Β· correlation]
B --> C[Encode<br/>ordinal maps Β· one-hot]
C --> D[Scale numeric<br/>StandardScaler]
D --> E[Train / test<br/>85k / 15k]
E --> F[Gradient boosting<br/>300 trees, depth 5]
F --> G[Evaluate<br/>ROC Β· PR Β· calibration Β· lift]
G --> H[Explain<br/>permutation Β· SHAP]
```
> [!note]- Preprocessing details
> | Feature type | Treatment |
> |---|---|
> | City_Type | Ordinal map: Rural = 1, Suburban = 2, Urban = 3 |
> | Range_Anxiety_Level | Ordinal map: Low = 1, Medium = 2, High = 3 |
> | Gender | Map: Male = 1, Female = 0, Other = 0 |
> | Current_Car_Type, Home_Charging_Possible, Subsidy_Available | One-hot encoding |
> | Numeric features | Median imputation + standard scaling |
>
> All steps run inside one scikit-learn `Pipeline`, so the preprocessing is learned from the training rows only. Related notes: [[Tree-based Models]] Β· [[Exploratory Data Analysis]]
## 4. What the raw data already says
Before any modeling, simple buy rates show most of the story.
### Environmental concern: the steepest gradient
| Concern level | 1 | 2 | 3 | 4 | 5 |
| ------------- | ---: | ---: | ----: | ----: | ------------: |
| **Buy rate** | 0.5% | 2.2% | 10.9% | 24.9% | ==**51.7%**== |
Each step up the concern scale multiplies the buy rate by roughly **2Γ to 5Γ**.
### Subsidy: the gate
| Subsidy available? | People | Buy rate |
| ------------------ | -----: | --------: |
| No | 37,222 | **0.6%** |
| Yes | 62,778 | **27.5%** |
> [!important] The interaction is the real finding
> Subsidy and concern don't simply add up. **Concern only matters when a subsidy exists.** Without one, even the most concerned people buy at just **2.5%**. With one, the same group buys at **69.5%** (see the chart at the top).
### Income, range anxiety and the rest
| Annual income | < $40K | $40β60K | $60β80K | $80β100K | $100β120K | $120β150K | $150K+ |
|---|---:|---:|---:|---:|---:|---:|---:|
| **Buy rate** | 4.6% | 9.3% | 11.8% | 17.8% | 25.9% | 31.9% | **37.4%** |
| Range anxiety | Low | Medium | High |
|---|---:|---:|---:|
| **Buy rate** | 18.9% | 4.2% | **0.3%** |
| People | 90,395 | 9,289 | 316 |
> [!example]- Smaller patterns
> - **Home charging possible:** 19.5% vs 12.7%
> - **City type:** Rural 19.2% Β· Suburban 18.2% Β· Urban 16.0%. Rural people are slightly *more* likely to buy.
> - **Daily commute:** peaks at **10β20 km (24.9%)** and falls to **10.6%** above 60 km.
> - **Car type, gender, number of cars, age:** all within about Β±2 points of the average.
>
> ![[ev-eda-correlation-heatmap.png|520]]
> Among the numeric features, only **environmental concern (r = 0.46)** and **income (r = 0.22)** correlate noticeably with buying. Stations near home and near work correlate with each other (0.51) but not with the target.
## 5. How good is the model?
| Metric (15,000 test people) | Value | What it means |
| --------------------------- | ------------: | -------------------------------------------------------------------- |
| **ROC-AUC** | **0.940** | Picks the buyer over the non-buyer 94% of the time for a random pair |
| **PR-AUC** | **0.752** | vs **0.175** for random guessing |
| Accuracy | 0.900 | Only modestly above the 0.825 you'd get by always predicting "no" |
| Log loss Β· Brier score | 0.230 Β· 0.072 | Probability quality |
> [!warning] Read the buyer-class numbers, not the weighted ones
> The run metadata reports precision and recall **weighted across both classes** (β 0.90), which the large "no" class inflates. For the class that matters, **buyers**, at the default 0.5 threshold:
>
> | | Predicted no | Predicted yes |
> |---|---:|---:|
> | **Actual no** | 11,713 | 669 |
> | **Actual yes** | 836 | **1,782** |
>
> **Precision 0.727 Β· Recall 0.681 Β· F1 0.703**. About **1 in 3 real buyers is missed** at this threshold.
### Choosing a threshold
| Threshold | Precision | Recall | F1 | Share of people flagged |
| -----------------: | --------: | --------: | --------: | ----------------------: |
| 0.30 | 0.614 | **0.829** | 0.706 | 23.5% |
| **0.41** (best F1) | 0.685 | 0.747 | **0.715** | β 20% |
| 0.50 (default) | **0.727** | 0.681 | 0.703 | 16.3% |
Lowering the threshold to **0.41** catches about 7 percentage points more buyers for a small drop in precision. For outreach where a missed buyer costs more than a wasted contact, 0.30 is a reasonable choice.
### Ranking power
![[ev-lift-by-decile.png]]
> [!success] Targeting insight
> - The **top 10%** of scores buy at **80.2%**, **4.6Γ** the average.
> - Contacting the **top 30%** reaches **92%** of all buyers.
> - The **bottom half** of the ranking contains only **1.5%** of buyers, so it can safely be skipped.
> [!example]- Calibration, ROC / PR curves and learning curve
> ![[ev-gb-calibration.png|420]] ![[ev-gb-learning-curve.png|420]]
> - **Well calibrated:** people given ~85% get 84.9% actual, people given ~15% get 16.1%. The probabilities can be used directly, for example to estimate expected sales.
> - **No overfitting:** train ROC-AUC 0.950 vs test 0.940. On the learning curve, weighted F1 is 0.897 on training and 0.895 in validation.
>
> ![[ev-gb-roc-curve.png|420]] ![[ev-gb-precision-recall.png|420]]
> ![[ev-gb-confusion-matrix.png|380]]
## 6. What drives the prediction (SHAP)
![[ev-gb-shap-beeswarm.png|600]]
> [!info] How to read SHAP here
> The model predicts the **log-odds** of buying. A SHAP value of **+1** multiplies the odds of buying by about **e β 2.7**, and **β1** divides them by 2.7. Red dots are high feature values and blue dots are low ones.
| Rank | Feature | Permutation importance | Mean \|SHAP\| (log-odds) |
|---:|---|---:|---:|
| 1 | Environmental_Concern_Level | **0.118** | **1.39** |
| 2 | Subsidy_Available | **0.067** | **0.82 + 0.80**ΒΉ |
| 3 | Annual_Income_USD | 0.035 | 0.40 |
| 4 | Range_Anxiety_Level | 0.009 | 0.16 |
| 5 | Daily_Commute_km | 0.004 | 0.07 |
| 6 | Charging_Stations_Near_Home | 0.002 | 0.07 |
| β | Gender, Number_of_Cars_Owned, Current_Car_Type | β 0 | < 0.03 |
ΒΉ *Subsidy is one-hot encoded into two mirror-image columns (True / False), so SHAP splits its effect between them. Read them as one lever.*
### Key findings
> [!tip] 1. Environmental concern: steady and strong
> SHAP rises almost **linearly** in log-odds, from about β2.1 at level 1 to +2.3 at level 5. That's roughly **+1.1 log-odds per level**, about **3Γ the odds** each step up. There's no plateau, so every step of concern adds about the same boost.
> [!tip] 2. Subsidy: an on/off switch
> Each of the two subsidy columns adds about **+0.7** when a subsidy is available and subtracts about **β1.0** when it isn't. Together, having vs. not having a subsidy is a swing of roughly **3β3.5 log-odds** (about **20β30Γ the odds**), the biggest single switch in the model. The raw data shows it's a precondition: concern only turns into purchases when a subsidy exists.
> [!tip] 3. Income: more money, more EVs
> Income has a near-linear positive effect, with the steepest penalty at the low end. Below-average earners are pushed down sharply, and above-average earners get a steady boost.
> [!tip] 4. Range anxiety: a veto
> Medium anxiety costs about **β0.7** log-odds and high anxiety about **β1.5**. For the small group with high anxiety, that is enough to cancel strong environmental concern.
> [!tip] 5. Commute: the sweet spot is short to moderate
> Commutes of around **10β20 km** push predictions up (the SHAP peak is near 10 km), and long commutes, especially above ~50 km, push them down. That fits practical range limits: EVs suit regular, predictable daily driving.
> [!tip] 6. Charging stations: context matters
> In the raw data, buy rates are almost **flat** across the number of stations near home (17β19%). The model, holding everything else fixed, finds that **zero stations nearby is a clear negative** and that about **8+ stations** gives a small boost. A marginal (one-at-a-time) view can hide a real conditional effect.
> [!example]- All SHAP dependence plots & one worked prediction
> ![[ev-gb-shap-all-relationships.png|700]]
>
> **One person's prediction (row 0):** no subsidy (β1.09 and β0.90 across the two columns) plus below-average concern (β1.09) pull the prediction down to **β6.46 log-odds**, a **0.16% chance** of buying. Nothing else comes close in size.
> ![[ev-gb-shap-row0.png|480]]
> ![[ev-gb-permutation-importance.png|480]]
## 7. What this means in practice
```mermaid
flowchart TD
S{Subsidy<br/>available?} -- No --> N["~0.6% buy<br/>(even max concern: 2.5%)"]
S -- Yes --> C{Environmental<br/>concern}
C -- "1β2" --> L["0.9β3.7% buy"]
C -- 3 --> M["18.2% buy"]
C -- "4β5" --> H["39.8β69.5% buy"]
H --> R{Range anxiety?}
R -- High --> V["almost no buyers<br/>(0.3% overall)"]
R -- "Low / Medium" --> T["π― Prime prospects"]
```
| Who | Recommended action |
|---|---|
| ποΈ **Policymakers** | Subsidies are the precondition: concern alone barely converts without one. Where subsidies don't exist, expect very little EV uptake. |
| π― **Marketers** | Target **high-concern people in subsidy regions**. The top 30% of model scores reach 92% of buyers. |
| π **EV makers & dealers** | Address **range anxiety** directly with range guarantees, charging maps and test drives on real commutes. It's a veto, not a small drag. |
| ποΈ **Charging networks** | Getting from **zero to some** stations near home matters more than adding stations where some already exist. |
## 8. Caveats
> [!warning]- Limitations
> - **Associations, not causes.** A subsidy's effect here may partly reflect *where* subsidies exist (region, policy environment), not the subsidy itself.
> - **Unusually clean patterns.** Buy rates like 0.0% for low concern without a subsidy suggest this is simulated or competition data. Real-world behavior is noisier.
> - **Small groups.** Only **316** people report high range anxiety, so that estimate is less stable.
> - **Encoding choices.** Gender "Other" (797 people) is merged with Female, and City_Type is treated as ordered (Rural < Suburban < Urban).
> - **Duplicated one-hot columns.** Binary features (subsidy, home charging) appear twice in SHAP. `OneHotEncoder(drop="if_binary")` would give one clean column each.
> - **Metrics.** The run's saved precision and recall are class-weighted. The buyer-class numbers above are the ones to quote.