The 2x2 Perf Review
JBS, the GM of Risk at Square, had a simple way of doing performance reviews for our risk operations team. Everyone went on the same 2x2.
One axis was false positives: how often did you flag a legitimate merchant? The other was false negatives: how much fraud did you let through? Throughput was basically pass/fail. You had a queue. You had to get through it.
Every few weeks, JBS would meet with each operator individually. He would also anonymize everyone else on the team, so each person knows where they sat relative to the rest.
I remember finding this a little awkward at first. These were people making nuanced judgment calls with incomplete information, and we were reducing their performance to two numbers. Also, the comparison with everyone else felt… almost too direct.
But, I’ve come to appreciate the principle behind it: clear is kind. Everyone knew what good meant: catch the bad guys without bothering the good guys. The 2x2 didn't tell you what to do on any particular case. It told you whether, over enough cases, your judgment was actually working.
There was another useful outcome.
Once we had defined the job that clearly, we could grade an algorithm on the exact same 2x2. Human reviewer, old model, new model: same cases, same objective, same scoreboard. That made it possible to run new models in shadow mode and ask a very simple question: who is actually better at the job?
As we ask AI to take on increasingly fuzzy knowledge work, we keep running into the problem of evals. Often the first problem isn't figuring out how to grade the machine. It's being honest that we never clearly defined how to grade the humans.