Pranav Viswanath
Email · X
HuggingFace Dataset
Writeup

J-lens vs NLA on AuditBench

Hidden quirk
= an NLA read on this prompt surfaces the quirk · = this prompt's J-lens readout surfaces the quirk · = alignment-notable non-quirk reads flagged

Prompt + Response · click any token to read both instruments at that position

model response NLA read here NLA surfaces the quirk

J-lens · argmax · layer × pos · top row = token positions, each column is that position read at lens layers 78 down to 30 · hover a cell for the organism and base-model top-10

an annotator-cited quirk atom is in this cell's top-25 · the cell still shows its true rank-1 token — hover to see the cited atom highlighted in the full readout, with the annotator's note alignment-notable but not the quirk (annotator flag: deception awareness, eval awareness, affect...) no highlight = this prompt's whole J-lens readout adjudicated as not surfacing the quirk (task echo / false-confidence vocab)

NLA · verbalized layer-53 residual · one explanation per scanned position · quirk evidence highlighted

loading data…