Eval Projects
PXR-related evaluation work from PRISM 2026 — probing whether models reason about potency and selectivity or lean on arithmetic shortcuts.
PRISM 2026 · Experiment Updates
Generalizing from PXR to AR datasets
Behavioural tasks that ask whether models can do nuanced thinking about biological assay data when the answer cannot be read off a single assay — transferring from PXR counter-screens to AR blockade and viability assays.
True/False probes
Given a true claim and a false claim about the same compound, is there a linear direction in the residual stream that separates them — and does that direction transfer to new compounds, survive confounds, and respond to steering?
OLMo branch experiments
Does potency become linearly decodable before selectivity, and does a genuine selectivity construct emerge before or after arithmetic selectivity? Frozen probes across training checkpoints on OLMo, Pythia, and Gemma.
Guiding question: LLM potency & selectivity shortcuts
The opening framing for the project: probe activations, reparameterize units, and patch residual streams to ask whether models genuinely represent potency and selectivity or default to arithmetic shortcuts on scientific tasks.