Eval Projects

PXR-related evaluation work from PRISM 2026 — probing whether models reason about potency and selectivity or lean on arithmetic shortcuts.

Generalizing from PXR to AR datasets

September 2026 · Experiment Update · deepseek-v4 · gemma-4-26b · gpt-5.4 · qwen3.5-35b

Behavioural tasks that ask whether models can do nuanced thinking about biological assay data when the answer cannot be read off a single assay — transferring from PXR counter-screens to AR blockade and viability assays.

True/False probes

September 2026 · Experiment Update · Mistral-7B · Qwen2.5-7B · gemma-4-E4b · Qwen3-14B

Given a true claim and a false claim about the same compound, is there a linear direction in the residual stream that separates them — and does that direction transfer to new compounds, survive confounds, and respond to steering?

OLMo branch experiments

September 2026 · Experiment Update · OLMo-2-7B · Pythia-6.9B · Gemma-4-E4B

Does potency become linearly decodable before selectivity, and does a genuine selectivity construct emerge before or after arithmetic selectivity? Frozen probes across training checkpoints on OLMo, Pythia, and Gemma.

Guiding question: LLM potency & selectivity shortcuts

Mid-August 2026 · Team Update · PXR Dataset · HuggingFace · 11,000+ screened compounds

The opening framing for the project: probe activations, reparameterize units, and patch residual streams to ask whether models genuinely represent potency and selectivity or default to arithmetic shortcuts on scientific tasks.