redline
METHODOLOGY

What this score means, and what it does not.

We would like to sell an interview. An interview is a selection procedure, and a selection procedure that has not been validated is a legal exposure wearing a dashboard. So here is the register of everything we claim about our own number, what actually backs each one, and what would have to be true before we could say it without a caveat.

Two of the five are marked not validated. One of them was on this website until recently.

THE NOTICE ON EVERY SCORE

This score measures performance on this exercise. It has not been validated as a predictor of job performance and must not be the sole or primary basis for a hiring decision.

01THE CLAIMS REGISTER
Partial

The score measures how well someone reviews code they did not write.

WHAT BACKS IT

The instrument has content validity by construction: the tasks are review tasks, the rubric was written from the failure modes senior engineers name in real reviews, and two of the defects are mined from public git history with the fix commit as the answer key. That establishes the exercise is made of the right material. It does not establish that the number ranks people correctly.

WHAT WOULD UPGRADE IT

A criterion study: scores from at least 200 engineers against an independent measure of review quality — manager rating, peer assessment, or agreement with a hand-scored panel on held-out diffs — with the correlation published including its confidence interval.

Not validated

A high score predicts on-the-job performance.

WHAT BACKS IT

Nothing. We have not run a predictive study, we do not have the longitudinal data to run one, and no result on this platform has been compared against any outcome at any employer. Anyone telling you otherwise, including us, is guessing.

WHAT WOULD UPGRADE IT

A predictive validity study with a pre-registered criterion, run with design partners who agree in advance to publish the result whichever way it comes out.

Not validated

An agent cannot sit this assessment for you.

WHAT BACKS IT

False as stated, and it used to be on our landing page. Paste a diff challenge into a frontier model and it finds the missing idempotency key. The narrower claim we can defend is about the other two formats: on blast radius and adjudication the answer is not in the diff, so a model given only the diff cannot produce it. That is a statement about what is in the context window, not about what a model can reason about — a model given the whole system can reason about the whole system.

WHAT WOULD UPGRADE IT

A published benchmark: current frontier models run against every challenge under stated conditions, scores reported per format, re-run each release. Until that table exists the honest claim is the narrow one.

Not validated

The assessment does not disadvantage protected groups.

WHAT BACKS IT

Unknown, and unknowable from what we collect. We ask for no demographic data, which means we cannot compute a selection rate by group even if we wanted to. Not measuring is not the same as no impact, and an instrument that is heavily verbal — the explanation step is graded prose — has an obvious channel for one.

WHAT WOULD UPGRADE IT

An adverse impact analysis on voluntary, separately-stored demographic data collected with consent from an assessment cohort, reported against the four-fifths rule, plus an independent bias audit published before any hiring use.

Not validated

The result proves a human produced this answer.

WHAT BACKS IT

It does not, and no amount of cryptography will make it. A signed receipt proves our server graded this answer at this time. Distinguishing a person from an agent driving a browser is proctoring, and we do not proctor. The submission-speed flag catches the careless version of the problem and nothing else.

WHAT WOULD UPGRADE IT

Proctored administration for any assessment use, plus a private problem pool that never appears in the public product so a leaked answer key is worth nothing.

02WHAT WE SELL IT FOR

Two of these are in the contract as prohibited uses.

Permitted: Practice, and a rating that goes down
The consumer product. Nothing is at stake, and self-reported numbers cost nobody a job.
Permitted: Onboarding against your own codebase
A new hire meeting your real failure modes in week one. No selection decision, no protected-group exposure, no validity requirement — this is training.
Permitted: Reviewer calibration inside a team
Finding out which of your reviewers block clean changes and which approve the expensive ones. A development conversation, not an employment one — provided it stays out of promotion packets.
Not permitted: Screening candidates out of a pipeline
This is the one everybody asks for, and we do not sell it yet. A cut score applied to applicants is a selection procedure whether we call it one or not.
Not permitted: Ranking candidates against each other
Same reason, plus a worse one: a rank order implies precision the instrument has not earned. We do not know that 78 beats 71.
03WHAT ATTACHES ON DAY ONE

The obligations below start the moment a score influences an employment decision — not when we start calling it an assessment. This is the register we work from with counsel. It is not legal advice, jurisdictions move, and this list has been wrong before.

United States — federal

Uniform Guidelines on Employee Selection Procedures (29 CFR 1607)

Any procedure used as a basis for an employment decision is a selection procedure. If it produces an adverse impact on a protected group — conventionally read against the four-fifths rule — it has to be justified by validity evidence, and the evidence has to exist before the tool is used, not after somebody asks.

New York City

Local Law 144 — automated employment decision tools

An annual bias audit by an independent auditor, a summary of the results published, and candidates notified at least ten business days before the tool is used on them. The audit is the buyer's obligation as well as ours, which makes it a sales objection as much as a legal one.

European Union

AI Act — employment is an Annex III high-risk use

Software used for recruitment, selection, or evaluation of candidates sits in the high-risk category, which brings risk management, data governance, technical documentation, logging, human oversight, and a conformity assessment. Obligations phase in; the classification does not.

Everywhere we operate

Accommodation, and a human in the loop

A timed reading-and-writing exercise has an accessibility surface: a candidate is entitled to request an adjustment, and the process has to exist before it is requested. A person, not a threshold, makes the decision — and that person has to be able to see the transcript the score came from.

04WHAT WOULD STOP US

Stated in advance so we cannot move the goalposts later.

  1. 01

    Adjudication and blast-radius scores correlate with nothing — not seniority, not peer assessment, not manager rating. Then we are measuring test-taking, and the honest move is to say so.

  2. 02

    Fewer than three companies put this in a real hiring or onboarding loop after fifty conversations. Then the distribution problem is unwinnable and the product is a hobby.

  3. 03

    Players score the same on real 900-day defects as on invented ones. Then the provenance is a story rather than a moat.

Go and try it instead

Nothing on this page is a reason not to use the product. It is a reason not to use the number for something it cannot carry yet.