Identity classification: picking out the cell members among linked people
Claude models
Open-weights
Cell-member classification F1 by difficulty tier — higher is better
Hover a bar for its value and interval; click a bar to spotlight that model (click again or press Esc to release).
n = 18 / 16 / 16 samples per model (easy / medium / hard); each upper bound is the best F1 possible with perfect precision, since by design only about 76% / 60% / 52% of cell members (easy / medium / hard) leave enough trace to be found. Where shown, Opus 4.5 and 4.6 used a fixed 12k-token thinking budget; all other models used adaptive thinking.