← Patrick Taylor
Personal project  |  2026

hmda-audit

US lenders have to show that their approvals are fair across race, ethnicity and sex, and banks have to validate the models behind credit decisions. This project runs both kinds of check on public mortgage data.

The data

In the US, every lender has to publish every mortgage application and whether they approved it. I pulled three years, 2023 to 2025: 36.7 million HMDA records from 5,329 lenders. Taking out purchased loans, which nobody applied for, leaves 32.6 million applications. I split the approvals by race, ethnicity and sex.

The screen

614 of 5,329 lenders fall below the four-fifths line
Approval rate of each race group divided by the joint-applicant rate, national total, on the screen's definition (every record except purchased loans)
0.80
Free Form Text Only0.69
Native Hawaiian or Other Pacific Islander0.83
2 or more minority races0.83
American Indian or Alaska Native0.85
Black or African American0.85
Race Not Available0.94
White0.98
Asian0.99
Joint (reference)1.00
Note: A ratio under 0.80 fails the four-fifths rule. Per-lender screen counts groups with 100 or more applications. 32.6M applications screened; purchased loans are excluded. Hover a bar for its count.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
Click to enlarge

First I screened every lender with the four-fifths rule, a first-pass test regulators use. If a group gets approved less than 80% as often as the reference group, that lender gets flagged for a closer look.

614 of 5,329 lenders got flagged. The screen counts every record except purchased loans, so a withdrawn or incomplete application counts as not denied. On that definition the only national bar below the line is "Free Form Text Only": about 9,000 filings where race was only written in as free text.

The definition matters. Counting only the files a lender actually approved or denied, which is how the model below is trained, Black applicants come in at 0.79 of the joint-applicant rate, below the line, and three other named race groups fall below it too. Both versions are in the results files.

The fair comparison

About two thirds of the gap for Black applicants remains after a like-for-like comparison
Approval gap against joint applicants, in percentage points
Raw gap12.4 pts
Same income, loan size and debt8.2 pts
Note: Controls are income, how much of the home's value is being borrowed, debt measured against income, what the loan is for, and whether it is a first or a second mortgage. Run on a 200,000-row stratified sample. The remaining gap is unexplained variation, not proof of discrimination.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
Click to enlarge

Raw gaps can mislead, because groups differ in income and debt. So I compared applicants with the same income, loan size and debt load.

For Black applicants the gap went from about 12 points to about 8, so roughly two thirds of it was still there.

The model

The denial model ranks the denied application as riskier 86% of the time
Share of denied and approved pairs the model puts in the right order, tested on 2025
Constant score (AUC 0.5)50%
Model, race left out86%
Note: AUC 0.86 (Gini 0.72) for gradient-boosted trees and AUC 0.78 (Gini 0.57) for logistic regression. Trained on 2023 and 2024, tested on the 2025 holdout, about 380,000 decisions per seed, from a 1,500,000-row sample; three random seeds agreed. Ranking well is not the same as being accurate for one applicant.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
Click to enlarge

Next I trained a model that predicts which applications get denied, with race kept out of the inputs. It learned from 2023 and 2024 and was scored on 2025, a year it never saw.

If you pick one denied and one approved application, it ranks the denied one as riskier about 86% of the time. That is an AUC of 0.86, or a Gini of 0.72 in credit-scoring terms. Logistic regression gets AUC 0.78 (Gini 0.57). A constant score can't order anyone, so it gets AUC 0.5.

The race probe

Race was never an input, yet a probe reads it at 0.68 AUC from the allowed features
Linear race-probe AUC, Black vs. White applicants, mean over 5 seeds; 0.5 is chance
0.50
Allowed input features0.678
Hidden layer 10.673
Hidden layer 20.659
Random-init layer 2 (control)0.642
Shuffled labels (control)0.499
Note: Race leaks through ordinary credit fields; the strongest proxies by one-feature AUC are property value (0.599), income (0.595) and debt-to-income band (0.594). A probe reads race a little less easily from each hidden layer than from the raw inputs, so training does not make race easier to read than it already is from the allowed features. This measures what information a probe can extract, not discrimination or intent; numbers are sample means over 5 seeds on a 500,000-row national sample, and this small network's own denial AUC is 0.838 (logistic regression 0.784, LightGBM 0.859).
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
Click to enlarge

I also trained a small neural net on the same allowed inputs and checked whether a simple probe could still tell Black and White applicants apart from its inputs and hidden layers, and it could, well above chance.

The controls

37 of 38 model-risk controls have an automated test
Each square is one control, mapped to SR 11-7 (US) and OSFI E-23 (Canada)
Note: Hover a square to read the control. C-38 has no test yet. C-01 fails the build if a race, ethnicity, sex or age column reaches the model.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
Click to enlarge

I also wrote 38 model-risk controls, mapped to the US and Canadian guidance (SR 11-7 and OSFI E-23).

37 of them have an automated test behind them, including one that fails if race ever sneaks into the model.

The watch list

1,018 of 1,400 raw four-fifths flags do not survive a small-sample correction
Lenders flagged for the Black vs. White denial-rate gap, national total, all years pooled
Raw four-fifths flags, no size floor1,400
Four-fifths flags, 100+ applications368
Empirical-Bayes watch list404
Note: Small lenders are pulled toward the typical Black vs. White denial-rate gap of 9.8 points, which itself varies by about 5.9 points from lender to lender, so a raw four-fifths ratio computed on a handful of applications is noisy and can flag or clear a lender by chance. 97 watch-list lenders are not flagged by the four-fifths screen even with the 100-application floor applied; a flag here is a prompt for review, not a finding.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
Click to enlarge

The four-fifths screen treats every lender the same, so a lender with only a handful of applications can fail it by chance. I also built an empirical-Bayes watch list: it compares each lender's Black vs. White denial-rate gap, pooled across all years, and shrinks small lenders' estimates toward the typical gap before ranking them.

1,018 of the 1,400 lenders flagged by the raw four-fifths screen do not survive that correction. For the top watch-list lenders, ranked by shrunk gap, I also checked the raw gap against a version controlled for loan and applicant differences: it could be estimated for 21 of them and stayed positive for all 21, and it could not be estimated for 4 more. No lender identity appears on this page.

Checked on Databricks too

Same answer on Databricks
Lender-by-race application and denial counts, Spark SQL vs. DuckDB, full national file
RowsApplicationsDenials
Spark SQL, Databricks31,79332,620,7896,166,654
DuckDB, local parquet31,79332,620,7896,166,654
Note: 0 value mismatches across every (lender, group) row, checked on 2026-09-28.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor

To make sure the DuckDB numbers above hold up on a different engine, I reran the lender-by-race counts as Spark SQL against the full national file loaded into a Databricks Delta table, and compared every row against the DuckDB result.

Every row matched: 0 mismatches across 31,793 lender-by-race rows, 32,620,789 applications and 6,166,654 denials.

The limit

Public HMDA has no credit scores, and I can't add them, so every controlled gap leaves out the biggest factor in real underwriting. That is a real limit: a flag means look closer, not proof of discrimination.

Stack
Python, DuckDB, scikit-learn, LightGBM, Databricks (Spark SQL, PySpark, MLflow), Streamlit
Speed
185x faster than pandas on the same tables, 0.13 s against 24.36 s
Back to everything else