US lenders have to show that their approvals are fair across race, ethnicity and sex, and banks have to validate the models behind credit decisions. This project runs both kinds of check on public mortgage data.
The data
In the US, every lender has to publish every mortgage application and whether they approved it. I pulled three years, 2023 to 2025: 36.7 million HMDA records from 5,329 lenders. Taking out purchased loans, which nobody applied for, leaves 32.6 million applications. I split the approvals by race, ethnicity and sex.
First I screened every lender with the four-fifths rule, a first-pass test regulators use. If a group gets approved less than 80% as often as the reference group, that lender gets flagged for a closer look.
614 of 5,329 lenders got flagged. The screen counts every record except purchased loans, so a withdrawn or incomplete application counts as not denied. On that definition the only national bar below the line is "Free Form Text Only": about 9,000 filings where race was only written in as free text.
The definition matters. Counting only the files a lender actually approved or denied, which is how the model below is trained, Black applicants come in at 0.79 of the joint-applicant rate, below the line, and three other named race groups fall below it too. Both versions are in the results files.
Next I trained a model that predicts which applications get denied, with race kept out of the inputs. It learned from 2023 and 2024 and was scored on 2025, a year it never saw.
If you pick one denied and one approved application, it ranks the denied one as riskier about 86% of the time. That is an AUC of 0.86, or a Gini of 0.72 in credit-scoring terms. Logistic regression gets AUC 0.78 (Gini 0.57). A constant score can't order anyone, so it gets AUC 0.5.
I also trained a small neural net on the same allowed inputs and checked whether a simple probe could still tell Black and White applicants apart from its inputs and hidden layers, and it could, well above chance.
The four-fifths screen treats every lender the same, so a lender with only a handful of applications can fail it by chance. I also built an empirical-Bayes watch list: it compares each lender's Black vs. White denial-rate gap, pooled across all years, and shrinks small lenders' estimates toward the typical gap before ranking them.
1,018 of the 1,400 lenders flagged by the raw four-fifths screen do not survive that correction. For the top watch-list lenders, ranked by shrunk gap, I also checked the raw gap against a version controlled for loan and applicant differences: it could be estimated for 21 of them and stayed positive for all 21, and it could not be estimated for 4 more. No lender identity appears on this page.
Checked on Databricks too
Same answer on Databricks
Lender-by-race application and denial counts, Spark SQL vs. DuckDB, full national file
Rows
Applications
Denials
Spark SQL, Databricks
31,793
32,620,789
6,166,654
DuckDB, local parquet
31,793
32,620,789
6,166,654
Note: 0 value mismatches across every (lender, group) row, checked on 2026-09-28.
Source: public HMDA filings, 2023 to 2025. hmda-audit analysis.
Patrick Taylor
To make sure the DuckDB numbers above hold up on a different engine, I reran the lender-by-race counts as Spark SQL against the full national file loaded into a Databricks Delta table, and compared every row against the DuckDB result.
Every row matched: 0 mismatches across 31,793 lender-by-race rows, 32,620,789 applications and 6,166,654 denials.
The limit
Public HMDA has no credit scores, and I can't add them, so every controlled gap leaves out the biggest factor in real underwriting. That is a real limit: a flag means look closer, not proof of discrimination.