Research

The Value Axis: Language Models Encode Whether They're on the Right Track

The Value Axis: Language Models Encode Whether They're on the Right Track

Nick Jiang, Isaac Kauvar, Jack Lindsey

paper | code

TLDR: language models internally estimate how well they're doing towards their goals along a linear axis, akin to a value function in RL.

Interpretable Embeddings with Sparse Autoencoders

Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit

Nick Jiang*, Xiaoqing Sun*, Lisa Dunlap, Lewis Smith, Neel Nanda

ICML 2026, NeurIPS Mech Interp Workshop 2025 (Spotlight ⭐)

paper | project page | code | X thread

TLDR: we show that sparse autoencoders outperform baselines on four data analysis tasks and find surprising model behaviors by analyzing training data and outputs.

Vision Transformers Don't Need Trained Registers

Vision Transformers Don't Need Trained Registers

Nick Jiang*, Amil Dravid*, Alexei Efros, Yossi Gandelsman

NeurIPS 2025 (Spotlight ⭐)

paper | project page | code | X thread

TLDR: we find and remove a sparse mechanism that causes attention sinks in ViTs, leading to general performance gains.

Interpreting and Editing Vision-Language Representations

Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Nick Jiang*, Anish Kachinthaya*, Suzie Petryk, Yossi Gandelsman

ICLR 2025

paper | code | X thread

TLDR: we use logit lens to identify and reduce hallucinations by 25% training-free from vision-language models.

*equal contribution