Accepted Work
Poster Session 1
Listed by random order.
- Conditional model diffing: synthetic document finetuning instills conditional steering vectors
- The Platonic Universe: Do Foundation Models See the Same Sky?
- Auditing Safety Probes Across Languages, Scripts, and Context Lengths
- Mean Field Analysis of Attention Describes Evolution of Representation Geometry and Identifies In-Context Learning
- Who does the confessing, and will they confess to anything
- The Steinmetz Project
- Weight-Space Discovery of an Epistatic Circuit Invisible to Activation-Based Methods
- What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
- Not Just a Pointer: Language Models Compose Thematic Role and Statement Order into a Binding Address
- Mechanics of Compliance: How Language Models Internalize Hierarchical Authority Bias
- Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?
- Promises lie in multiple directions
- Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
- Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models
- How Sycophancy Overrides Internal Knowledge
- Cognitive Steering Vectors For Language Models
- Calibrated Measurement Extraction from Scientific Literature with Open Large Language Models
- Surgical Repair of Insecure Code Generation in LLMs
- Improving Confidence Estimation in Language Models with Linear Representations of Ambiguity
- Interpreting Bayesian Inference in Language Models
- Toward Mechanistic Interpretability of LLM Agents: Explaining Trajectories Demands New Methods
- When the Circuit Isn't There: Establishing Absence in a Behavioral Sequence Model
- Singular Vectors of Attention Heads Align with Features
- Fine-Tuning Enhances Latent Metacognitive Capability in Language Models
- What Does the Instrument Identify? Probing, Geometry, and Measurement Validity in LLM Concept Audits
- BizzaroWorld: Factual Recall Circuits Do Not Scale Monotonically Across Gemma
- Testing the Limits of Truth Directions in LLMs
- Do explanations generalize across large reasoning models?
- Training In-sequence Introspection Without Causal Bypass
- Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
- Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
- Count Me If You Can: Geometric Failure Modes in Language Model Counting
- Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?
- Training is Less Consistent in MoEs than Dense Models
- Learning the Same Concept Geometry from Different Statistical Evidence
- Verbalizing Hidden Fine-Tuning Effects with Text Optimization
- Prototype-Based Dynamic Steering for Large Language Models
- Decodable Is Not Used: Probe Regularization Silently Controls the Causal Relevance of Belief-State Representations in Transformers
- Detecting and Controlling Sycophancy with Cascading Linear Features
- Reasoning Models Confabulate to Defend Injected Beliefs About Their Own Requests
- Refusal Steering using SAEs
- Steering language models via cognitive distillation
- How to Train Your Model Organism
- Are Latent Reasoning Models Easily Interpretable?
- Transformers Learn Two Competing Sorting Algorithms, and Only One Length-Generalizes
- Comparison as a Computational Primitive: Testing Modular Implementation in Language Models
- Hidden States as States: Discretizing Hidden-State Geometry in Language Model Computation
- Improving Positional Information in Activation Steering via Subcontexts
- Can Graph Learning Learn Circuits?
- Online Statistical Computation over Tables in Language Models
Poster Session 2
Listed by random order.
- Understanding the Effects of Modality in Context-Memory Conflicts
- Dual-contrastive sparse autoencoders reveal features of musical interpretation
- Understanding Cross-Resolution Information Loss in Variational Autoencoders
- Pretrained Representational Geometry Controls Fine-Tuning Generalization
- Predictable Steering? Exploring Geometry Proxies for Layer Selection and Per-Instance Success Prediction
- LMs Encode What a Sub-Goal Is and When It Ends Through Distinct Mechanisms
- Some Modalities are More Equal than Others: Interpreting Cross-Modal Interactions in MLLMs
- From Complexity to Clarity: A Systematic Study of Geometric Dynamics in Hidden Representations of LLMs
- Finding Interpretable Prompt-Specific Circuits in Language Models
- Single-unit activations confer inductive biases for emergent circuit solutions to cognitive tasks
- Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
- Similarity of Neural Network Representations in Superposition
- Reverse-engineering fast-weight recurrent networks reveals mechanisms underlying selective credit assignment during human-like in-context learning
- Do Dual-Route Induction Heads Explain Improbable-Bigram Copy Failures?
- Differences in multi- and single-token interpretability methods
- What Does a Chromatin Foundation Model Know About a Petri Dish? Sparse Autoencoders Reveal In Vitro vs. In Vivo Context in EPIBERT
- A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning
- What Makes a Reasoning Model Commit? Tracking Uncertainty Representations at Answer Proposals
- The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust
- Subliminal Steering: Stronger Encoding of Hidden Signals
- Interpretation of Introspective Self-Prediction in Large Language Models
- Lost in Compression: SOTA Neural Audio Models Predictably Degrade Access to Meaningful Features
- Representations in Motion: Tracking Layer-wise Activation Trajectories in Language Models
- Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
- Why Does Robustness Reduce Superposition?
- What to attribute over? Target token selection in long horizon interpretability
- Activation Parsers for Oversight
- Where Does Time Go? Tracing and Routing Temporal Information for Video-LLMs
- Identifying Introspection From the Inside
- Stable Encodings, Changing Downstream Sensitivity: Measuring SAE Feature Identity Across Fine-Tuning
- Convergent World Representations and Divergent Tasks
- Finding Algorithms in Neural Networks: Behavioral and Mechanistic Methods
- Same Answer, Same Mechanism? An Audit of Mechanistic Reproducibility
- On the Predictive Power of Representation Dispersion in Language Models
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
- Compiling to transformers
- Is the global workspace emergent? J-space ablation is causal but not selective in a 4B model
- Epistemic Familiarity is Associated with Belief Stability in Large Language Models
- Protein Circuit Tracing via Cross-layer Transcoders
- Interpreting Disordered Protein Representations in Protein Language Models via Sparse Autoencoders
- Sparse Interaction Decomposition Finds Interacting Features in Vision Transformers
- Gaze Heads: How VLMs Look at What They Describe
- When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
- Mapping the Blind Spots: Cultural Asymmetries in LLM Geographic Representations
- Uncovering Competency Gaps in Large Language Models and Their Benchmarks
- When Does a Feature Arrive? Punctuality and Orderability as Metrics of Circuit Validity
- Amortized Circuit Discovery via Graph Neural Network
- One mechanism among many: traversing the space of functionally equivalent solutions
