# shap-values

Surrogate SHAP values for the pitch-modeling Stuff and Pitching models, per pitcher x season x
pitch type, 2023-26 (`shap_values.py` in the stuff_model project; scorer and models:
https://github.com/Blandalytics/pitch-modeling).

Models and targets:

* `stuff`: count-neutral Stuff (`stuff_rv`) and its nine count-neutral outcome probabilities.
  Features: the stuff inputs and the pitcher's hand.
* `pitching`: Pitching (`pitching_rv`, the pitch's actual location and count) and its nine
  location-aware outcome probabilities. Features: the stuff inputs, the pitcher's hand, the
  batter-relative location (`x_b`, `z_n`) and the count (`balls`, `strikes`).
* `target`: `plus` (Stuff+ / Pitching+ points on the `pitcher_season_pitch_type` scale of
  pitch-modeling's `constants/plus_scale_constants.json`) or `p_<outcome>` (percentage points)
  for ball, called_strike, swinging_strike, foul, field_out, single, double, triple, home_run,
  or `wobacon`: expected wOBA on contact from the same model's batted-ball probabilities,
  (0.9 single + 1.25 double + 1.6 triple + 2 home_run) / P(in play). A unit's wOBAcon averages
  its pitches weighted by P(in play), so it is per ball in play.
* `rv_<outcome>`: the plus points that outcome's predicted probability is worth vs league,
  priced at the outcome's average (count-neutral) run value: -15 x run value / sd x (100 p -
  league). Each is its `p_<outcome>` row scaled by that constant, so no separate proxy. For
  Stuff the nine sum exactly to plus - 100. Pitching+ prices outcomes at the pitch's actual
  count, so for Pitching plus - 100 - the nine sum is count leverage (the card draws it as its
  own bar). Run values are in `meta.json`.
* `era`: pitch type ERA (`model_era.py --by pt`, runs per 9): a season constant - 9 x mean run
  value x modelled pitches per inning (the observed-rate formula applied to the pitch type's
  predicted outcome rates; "PLV ERA" for `pitching`, "Stuff ERA" for `stuff`). Its parts are
  the exact Shapley split of that product, built from the `plus` and `p_<outcome>` rows (no
  separate proxy). Season-neutral: the season feature's part is removed from the run value
  and pitches per inning, and each part is relative to that season's average pitch, so
  `season_env` and the extra `calibration` column are 0 and `league` (the season's league ERA)
  carries the season: exact = league + baseline + features + residual (+ calibration = 0).

Location (Pitching minus Stuff at the actual count) has an outcome split only, no feature
SHAP: `rv_<outcome>` is location's change in that outcome's probability (Pitching - Stuff as
used) x its run value at the pitch's count, on the Location+ scale vs league. The nine sum
exactly to Location+ - 100. `dp_<outcome>` is the probability change itself, in percentage
points.

For every unit and target:

    exact = league + baseline + sum(feature SHAP) + residual

`baseline` is the unit's pitch group x matchup expected value relative to the league, each
feature column its mean SHAP over the unit's pitches, and `residual` the surrogate's error
against the exact value. `meta.json` documents the columns, labels and units; `fidelity.csv`
has each surrogate's R^2 on held-out pitchers.

Files:

| file | contents |
|---|---|
| `units_<model>_<season>.parquet` | one row per season, pitcher, pt, target |
| `unit_values_<season>.parquet` | each unit's exact values: `<model>_plus`, `<model>_p_<outcome>` (%), `<model>_wobacon`, `<model>_rv_<outcome>` (plus points vs league), `<model>_era` |
| `units_location_<season>.parquet` | Location+ and its exact split by outcome (`plus`, `rv_<outcome>`, `dp_<outcome>`); no feature SHAP |
| `unit_features_<season>.parquet` | name, hand, pitch counts, main group, mean inputs |
| `meta.json`, `fidelity.csv`, `importance.csv` | documentation, surrogate fit, mean abs SHAP |
| `shap_values_card.py` | draws the waterfall card from these files |

```bash
pip install pandas pyarrow matplotlib
python shap_values_card.py "Jacob Misiorowski" FF
python shap_values_card.py "Jacob Misiorowski" FF --model pitching
python shap_values_card.py "Jacob Misiorowski" SL --target p_swinging_strike --season 2026
python shap_values_card.py "Jacob Misiorowski" FF --target wobacon
python shap_values_card.py "Jacob Misiorowski" FF --target rv_swinging_strike
python shap_values_card.py "Jacob Misiorowski" FF --by outcome   # Stuff+ split by outcome
python shap_values_card.py "Jacob deGrom" SL --model pitching --target era   # PLV ERA
python shap_values_card.py "Jacob deGrom" SL --model location --by outcome   # Location+
```

```python
import pandas as pd
u = pd.read_parquet("https://data.blandalytics.com/shap-values/units_stuff_2026.parquet")
```
