BLOG POST

Fairness Is a Build Step:Rethinking my fair lending work at PayPal, two years out

September 28, 2026 · 5 min read

At PayPal, finishing a credit model was the easy part. Then came fair lending review, where legal and compliance combed through every variable for signs of discrimination. Reviews ran on a semi-annual cycle, so getting a model through could take half a year.

The tool my team patented was built to make models fairer, and it did. But I spent two years building credit models at PayPal and have been gone about as long, and from here, the bigger win was elsewhere: the tool moved fairness from the end of that line to the beginning, turning a review gate into a build step.

A model doesn’t need to see race to be unfair

US lenders can’t decide who gets credit based on race or sex, so those columns never go into the model. Problem solved? Not quite.

Models learn from history, and history is full of stand-ins. Zip code tracks race because of decades of housing segregation. Where you shop, what device you use, and where you went to school can carry the same signal. Delete the race column, and the model can often rebuild it from everything else. Researchers call this approach “fairness through unawareness,” and it mostly doesn’t work.

There’s an irony, too. For most consumer loans, lenders can’t even ask applicants their race. So to check whether a model treats groups differently, you first have to estimate each applicant’s race, usually from surname and zip code. Our patent describes exactly that.

How fairness gets measured

There are many ways to score fairness. Two do most of the work.

The first is the adverse impact ratio (or disparate impact ratio): one group’s approval rate divided by another’s. If 50% of one group gets approved and 43% of another does, the ratio is 0.86, which is where some of our models started. Fair lending teams often treat anything under 0.80 as a red flag, a rule of thumb borrowed from employment law.

The second is the false positive rate. Think of a credit model as an alarm that goes off when it predicts “this person will default.” A false positive is a false alarm: someone who would have paid you back, declined anyway. The false positive rate is the share of those would-be repayers you turned away. If it’s 10% for one group and 5% for another, creditworthy people in the first group are twice as likely to be wrongly rejected, even when approval rates match.

Accuracy gets a number too. KS measures how cleanly a model’s scores separate people who will default from people who won’t. Higher is better.

Our tool screened each variable on several fairness metrics, error rates included, but the de-biasing step itself targeted the approval ratio. That’s a real choice. When groups default at different rates, no useful model can make every metric equal at once.

A model that plays a guessing game against itself

Many of our credit models ran on LightGBM, which builds hundreds of small decision trees one at a time, each fixing the last one’s mistakes.

The tool adds a second player, an adversary. It sees only the credit score and tries to guess the applicant’s race, gender, or age. If it guesses well, the score is leaking group membership.

Each new tree is then trained to do two things: keep predicting default, and make the adversary’s guess worse. Over enough rounds, the score keeps what it knows about repayment and loses what it knows about group. The patent’s example: applicants aged 20, 40, and 60 with similar incomes should get similar scores.

The de-biasing guessing gameApplicant data feeds a LightGBM credit model that adds one tree per round and outputs a credit score. Three adversaries see only the score and try to guess race, gender, and age. How well they guess is fed back, scaled by a decaying lambda, to shape the next tree.Applicant dataCredit model (LightGBM)tree 1 → tree 2 → … → tree NCredit scorescore onlyAdversaries try to guessracegenderagefrom the score aloneλnudge
The guessing game. Adversaries see only the score and try to guess each applicant’s race, gender, and age. Each new tree is trained to keep predicting default while making those guesses worse. The dial λ sets how hard they push, and it decays over time.

Three details made it practical. Separate adversaries for race, gender, and age run in parallel. Each has a strength dial, λ (lambda), that decays over time, so it pushes hard early and eases off before it wrecks accuracy. And it’s plug-and-play: it bolts onto an existing model without touching its code.

The method came from a research paper implemented in scikit-learn. I rewrote it for LightGBM, PayPal’s workhorse, with help from LightGBM’s developers, and extended it from gender alone to all three attributes. We tested it on 10 credit models, spanning underwriting and fraud. On some, the adverse impact ratio climbed from 0.86 to above 0.99 while KS slipped from 54 to 52.

From review gate to build step

Before the tool, fairness problems surfaced late. Review would flag a variable, either as a likely stand-in for a protected trait or for scoring badly on a fairness metric. The standard fix was to drop it.

If the variable was among the model’s most important, dropping it sent KS off a cliff. And you can’t just yank a variable out of a live model: the stripped-down version had to run as a shadow model beside the original, scoring real applications without deciding anything, until everyone trusted it. Nobody was happy. The modeling team, the “first line of defense” in bank-speak, had to rebuild. The business ate the lost performance. Compliance, legal, and modelers traded feedback for months.

With the tool, modelers screened their own variables before review, added a quick de-biasing pass to normal training, and checked KS on data the model had never seen, as always. Across several credit products, the KS cost was usually small.

It also changed the conversation with legal. Each round of de-biasing yields a slightly different model, so you get a curve, not a yes-or-no. In the patent’s example, race and age fairness improve significantly by round 60 while KS stays close to 52; push to 100 and KS sinks to about 44. You pick your point before review. That curve is a menu of less discriminatory alternatives, exactly the question disparate impact law asks.

The payoff: variable review got 84% faster, saving about five months per cycle, and outcome review and customer impact assessments were halved, saving 90 and 45 days respectively. My other project did the same for adverse action notices, the letters that explain a decline, using SHAP, which splits a score into each input’s contribution, to pull the real reasons from the model.

Two years later

The rules have moved since I left. As of July 21, 2026, the CFPB’s rules say the Equal Credit Opportunity Act doesn’t recognize disparate impact at all, and the EU has pushed its high-risk rules for credit scoring to December 2027.

The problems didn’t move. In 2025, Massachusetts settled with the student lender Earnest for $2.5 million. The allegations: it never tested its AI underwriting models for disparate impact, it used a variable (colleges’ default rates) that penalized Black and Hispanic borrowers, and its decline notices didn’t give real reasons. Those are the same three problems our tools were built for.

A fairness check that lives in a review queue depends on someone asking for it. One that lives in the training loop runs every time you build. That’s the part I’d build again.