# πŸš—βš‘ Who Will Buy an EV? ### Subsidies open the door. Environmental concern walks through it. ![[ev-subsidy-concern-gate.png]] > [!abstract] TL;DR > I trained a gradient boosting classifier on **100,000 people** to predict whether someone **will buy an electric vehicle**, then used SHAP to see what drives the prediction. > - **Two factors dominate.** Without a **subsidy**, almost nobody buys (**0.6%**). With one, **environmental concern** takes buy rates from **0.9%** up to **69.5%**. > - **Income** is the third lever. **Range anxiety** can override everything else: people with high range anxiety almost never buy (**0.3%**). > - **Car type, gender and number of cars owned barely matter.** > - The model is strong at **ranking** people: ROC-AUC **0.94**, PR-AUC **0.75** against a 0.175 baseline. The **top 20%** of scores contain **76%** of all buyers. | At a glance | | |---|---| | **Question** | Who will buy an EV, and what moves them? | | **Data** | 100,000 people Β· 13 features Β· 17.5% buyers | | **Model** | Gradient boosting (300 trees, depth 5, learning rate 0.05) | | **Split** | 85,000 train / 15,000 test | | **Headline metric** | ROC-AUC **0.940** Β· PR-AUC **0.752** Β· top-decile lift **4.6Γ—** | ## 1. The question > [!question] Who's ready to switch? > EV makers, dealers and policymakers all face the same targeting problem: only about **1 in 6** people in this data plan to buy. Sending incentives, test drives or ads to everyone wastes most of the budget. I wanted to answer two things: 1. **Prediction:** can we rank people by how likely they are to buy? 2. **Explanation:** which levers matter most: money, values, infrastructure, or practical worries like range? ## 2. The data | Group | Features | |---|---| | πŸ‘€ **Demographics** | Age, Gender, Annual_Income_USD, City_Type (Urban / Suburban / Rural) | | πŸš™ **Current situation** | Current_Car_Type (Sedan / SUV / Hatchback / Truck), Number_of_Cars_Owned, Daily_Commute_km | | πŸ”Œ **Charging access** | Charging_Stations_Near_Home, Charging_Stations_Near_Work, Home_Charging_Possible | | πŸ’­ **Attitudes** | Environmental_Concern_Level (1–5), Range_Anxiety_Level (Low / Medium / High) | | πŸ’΅ **Policy** | Subsidy_Available | | 🎯 **Target** | **Will_Buy_EV** (True / False) | > [!info]- Class balance > ![[ev-eda-target-distribution.png|420]] > Only **17.5%** of people plan to buy an EV (17,454 of 100,000). Because the classes are imbalanced, **accuracy alone is misleading**: predicting "no" for everyone would already score 82.5%. That's why this post focuses on **ROC-AUC, PR-AUC, and precision and recall for buyers**. ## 3. How I approached it ```mermaid flowchart LR A[(BigQuery table<br/>100k Γ— 15)] --> B[EDA<br/>buy rates Β· correlation] B --> C[Encode<br/>ordinal maps Β· one-hot] C --> D[Scale numeric<br/>StandardScaler] D --> E[Train / test<br/>85k / 15k] E --> F[Gradient boosting<br/>300 trees, depth 5] F --> G[Evaluate<br/>ROC Β· PR Β· calibration Β· lift] G --> H[Explain<br/>permutation Β· SHAP] ``` > [!note]- Preprocessing details > | Feature type | Treatment | > |---|---| > | City_Type | Ordinal map: Rural = 1, Suburban = 2, Urban = 3 | > | Range_Anxiety_Level | Ordinal map: Low = 1, Medium = 2, High = 3 | > | Gender | Map: Male = 1, Female = 0, Other = 0 | > | Current_Car_Type, Home_Charging_Possible, Subsidy_Available | One-hot encoding | > | Numeric features | Median imputation + standard scaling | > > All steps run inside one scikit-learn `Pipeline`, so the preprocessing is learned from the training rows only. Related notes: [[Tree-based Models]] Β· [[Exploratory Data Analysis]] ## 4. What the raw data already says Before any modeling, simple buy rates show most of the story. ### Environmental concern: the steepest gradient | Concern level | 1 | 2 | 3 | 4 | 5 | | ------------- | ---: | ---: | ----: | ----: | ------------: | | **Buy rate** | 0.5% | 2.2% | 10.9% | 24.9% | ==**51.7%**== | Each step up the concern scale multiplies the buy rate by roughly **2Γ— to 5Γ—**. ### Subsidy: the gate | Subsidy available? | People | Buy rate | | ------------------ | -----: | --------: | | No | 37,222 | **0.6%** | | Yes | 62,778 | **27.5%** | > [!important] The interaction is the real finding > Subsidy and concern don't simply add up. **Concern only matters when a subsidy exists.** Without one, even the most concerned people buy at just **2.5%**. With one, the same group buys at **69.5%** (see the chart at the top). ### Income, range anxiety and the rest | Annual income | < $40K | $40–60K | $60–80K | $80–100K | $100–120K | $120–150K | $150K+ | |---|---:|---:|---:|---:|---:|---:|---:| | **Buy rate** | 4.6% | 9.3% | 11.8% | 17.8% | 25.9% | 31.9% | **37.4%** | | Range anxiety | Low | Medium | High | |---|---:|---:|---:| | **Buy rate** | 18.9% | 4.2% | **0.3%** | | People | 90,395 | 9,289 | 316 | > [!example]- Smaller patterns > - **Home charging possible:** 19.5% vs 12.7% > - **City type:** Rural 19.2% Β· Suburban 18.2% Β· Urban 16.0%. Rural people are slightly *more* likely to buy. > - **Daily commute:** peaks at **10–20 km (24.9%)** and falls to **10.6%** above 60 km. > - **Car type, gender, number of cars, age:** all within about Β±2 points of the average. > > ![[ev-eda-correlation-heatmap.png|520]] > Among the numeric features, only **environmental concern (r = 0.46)** and **income (r = 0.22)** correlate noticeably with buying. Stations near home and near work correlate with each other (0.51) but not with the target. ## 5. How good is the model? | Metric (15,000 test people) | Value | What it means | | --------------------------- | ------------: | -------------------------------------------------------------------- | | **ROC-AUC** | **0.940** | Picks the buyer over the non-buyer 94% of the time for a random pair | | **PR-AUC** | **0.752** | vs **0.175** for random guessing | | Accuracy | 0.900 | Only modestly above the 0.825 you'd get by always predicting "no" | | Log loss Β· Brier score | 0.230 Β· 0.072 | Probability quality | > [!warning] Read the buyer-class numbers, not the weighted ones > The run metadata reports precision and recall **weighted across both classes** (β‰ˆ 0.90), which the large "no" class inflates. For the class that matters, **buyers**, at the default 0.5 threshold: > > | | Predicted no | Predicted yes | > |---|---:|---:| > | **Actual no** | 11,713 | 669 | > | **Actual yes** | 836 | **1,782** | > > **Precision 0.727 Β· Recall 0.681 Β· F1 0.703**. About **1 in 3 real buyers is missed** at this threshold. ### Choosing a threshold | Threshold | Precision | Recall | F1 | Share of people flagged | | -----------------: | --------: | --------: | --------: | ----------------------: | | 0.30 | 0.614 | **0.829** | 0.706 | 23.5% | | **0.41** (best F1) | 0.685 | 0.747 | **0.715** | β‰ˆ 20% | | 0.50 (default) | **0.727** | 0.681 | 0.703 | 16.3% | Lowering the threshold to **0.41** catches about 7 percentage points more buyers for a small drop in precision. For outreach where a missed buyer costs more than a wasted contact, 0.30 is a reasonable choice. ### Ranking power ![[ev-lift-by-decile.png]] > [!success] Targeting insight > - The **top 10%** of scores buy at **80.2%**, **4.6Γ—** the average. > - Contacting the **top 30%** reaches **92%** of all buyers. > - The **bottom half** of the ranking contains only **1.5%** of buyers, so it can safely be skipped. > [!example]- Calibration, ROC / PR curves and learning curve > ![[ev-gb-calibration.png|420]] ![[ev-gb-learning-curve.png|420]] > - **Well calibrated:** people given ~85% get 84.9% actual, people given ~15% get 16.1%. The probabilities can be used directly, for example to estimate expected sales. > - **No overfitting:** train ROC-AUC 0.950 vs test 0.940. On the learning curve, weighted F1 is 0.897 on training and 0.895 in validation. > > ![[ev-gb-roc-curve.png|420]] ![[ev-gb-precision-recall.png|420]] > ![[ev-gb-confusion-matrix.png|380]] ## 6. What drives the prediction (SHAP) ![[ev-gb-shap-beeswarm.png|600]] > [!info] How to read SHAP here > The model predicts the **log-odds** of buying. A SHAP value of **+1** multiplies the odds of buying by about **e β‰ˆ 2.7**, and **βˆ’1** divides them by 2.7. Red dots are high feature values and blue dots are low ones. | Rank | Feature | Permutation importance | Mean \|SHAP\| (log-odds) | |---:|---|---:|---:| | 1 | Environmental_Concern_Level | **0.118** | **1.39** | | 2 | Subsidy_Available | **0.067** | **0.82 + 0.80**ΒΉ | | 3 | Annual_Income_USD | 0.035 | 0.40 | | 4 | Range_Anxiety_Level | 0.009 | 0.16 | | 5 | Daily_Commute_km | 0.004 | 0.07 | | 6 | Charging_Stations_Near_Home | 0.002 | 0.07 | | β€” | Gender, Number_of_Cars_Owned, Current_Car_Type | β‰ˆ 0 | < 0.03 | ΒΉ *Subsidy is one-hot encoded into two mirror-image columns (True / False), so SHAP splits its effect between them. Read them as one lever.* ### Key findings > [!tip] 1. Environmental concern: steady and strong > SHAP rises almost **linearly** in log-odds, from about βˆ’2.1 at level 1 to +2.3 at level 5. That's roughly **+1.1 log-odds per level**, about **3Γ— the odds** each step up. There's no plateau, so every step of concern adds about the same boost. > [!tip] 2. Subsidy: an on/off switch > Each of the two subsidy columns adds about **+0.7** when a subsidy is available and subtracts about **βˆ’1.0** when it isn't. Together, having vs. not having a subsidy is a swing of roughly **3–3.5 log-odds** (about **20–30Γ— the odds**), the biggest single switch in the model. The raw data shows it's a precondition: concern only turns into purchases when a subsidy exists. > [!tip] 3. Income: more money, more EVs > Income has a near-linear positive effect, with the steepest penalty at the low end. Below-average earners are pushed down sharply, and above-average earners get a steady boost. > [!tip] 4. Range anxiety: a veto > Medium anxiety costs about **βˆ’0.7** log-odds and high anxiety about **βˆ’1.5**. For the small group with high anxiety, that is enough to cancel strong environmental concern. > [!tip] 5. Commute: the sweet spot is short to moderate > Commutes of around **10–20 km** push predictions up (the SHAP peak is near 10 km), and long commutes, especially above ~50 km, push them down. That fits practical range limits: EVs suit regular, predictable daily driving. > [!tip] 6. Charging stations: context matters > In the raw data, buy rates are almost **flat** across the number of stations near home (17–19%). The model, holding everything else fixed, finds that **zero stations nearby is a clear negative** and that about **8+ stations** gives a small boost. A marginal (one-at-a-time) view can hide a real conditional effect. > [!example]- All SHAP dependence plots & one worked prediction > ![[ev-gb-shap-all-relationships.png|700]] > > **One person's prediction (row 0):** no subsidy (βˆ’1.09 and βˆ’0.90 across the two columns) plus below-average concern (βˆ’1.09) pull the prediction down to **βˆ’6.46 log-odds**, a **0.16% chance** of buying. Nothing else comes close in size. > ![[ev-gb-shap-row0.png|480]] > ![[ev-gb-permutation-importance.png|480]] ## 7. What this means in practice ```mermaid flowchart TD S{Subsidy<br/>available?} -- No --> N["~0.6% buy<br/>(even max concern: 2.5%)"] S -- Yes --> C{Environmental<br/>concern} C -- "1–2" --> L["0.9–3.7% buy"] C -- 3 --> M["18.2% buy"] C -- "4–5" --> H["39.8–69.5% buy"] H --> R{Range anxiety?} R -- High --> V["almost no buyers<br/>(0.3% overall)"] R -- "Low / Medium" --> T["🎯 Prime prospects"] ``` | Who | Recommended action | |---|---| | πŸ›οΈ **Policymakers** | Subsidies are the precondition: concern alone barely converts without one. Where subsidies don't exist, expect very little EV uptake. | | 🎯 **Marketers** | Target **high-concern people in subsidy regions**. The top 30% of model scores reach 92% of buyers. | | πŸ”‹ **EV makers & dealers** | Address **range anxiety** directly with range guarantees, charging maps and test drives on real commutes. It's a veto, not a small drag. | | πŸ™οΈ **Charging networks** | Getting from **zero to some** stations near home matters more than adding stations where some already exist. | ## 8. Caveats > [!warning]- Limitations > - **Associations, not causes.** A subsidy's effect here may partly reflect *where* subsidies exist (region, policy environment), not the subsidy itself. > - **Unusually clean patterns.** Buy rates like 0.0% for low concern without a subsidy suggest this is simulated or competition data. Real-world behavior is noisier. > - **Small groups.** Only **316** people report high range anxiety, so that estimate is less stable. > - **Encoding choices.** Gender "Other" (797 people) is merged with Female, and City_Type is treated as ordered (Rural < Suburban < Urban). > - **Duplicated one-hot columns.** Binary features (subsidy, home charging) appear twice in SHAP. `OneHotEncoder(drop="if_binary")` would give one clean column each. > - **Metrics.** The run's saved precision and recall are class-weighted. The buyer-class numbers above are the ones to quote.