The 3rd New England Mechanistic Interpretability (NEMI) workshop

August 14, 2026, Boston University

Accepted Work

Poster Session 1

Listed by random order.

  1. Conditional model diffing: synthetic document finetuning instills conditional steering vectors
    Presenter: Stepan Shabalin
  2. The Platonic Universe: Do Foundation Models See the Same Sky?
    Presenter: Mike Smith
  3. Auditing Safety Probes Across Languages, Scripts, and Context Lengths
    Presenter: Sripad Karne
  4. Mean Field Analysis of Attention Describes Evolution of Representation Geometry and Identifies In-Context Learning
    Presenter: Mark Crovella
  5. Who does the confessing, and will they confess to anything
    Presenter: Abhishek Mishra
  6. The Steinmetz Project
    Presenter: Stephen Michael Packard
  7. Weight-Space Discovery of an Epistatic Circuit Invisible to Activation-Based Methods
    Presenter: Elliot Tower
  8. What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
    Presenter: Stephen Cheng
  9. Not Just a Pointer: Language Models Compose Thematic Role and Statement Order into a Binding Address
    Presenter: Nikhil Prakash
  10. Mechanics of Compliance: How Language Models Internalize Hierarchical Authority Bias
    Presenter: Srujananjali Medicherla
  11. Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?
    Presenter: Ayan Antik Khan
  12. Promises lie in multiple directions
    Presenter: Timothy Obiso
  13. Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
    Presenter: Nishant Subramani
  14. Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models
    Presenter: Germans Savcisens
  15. How Sycophancy Overrides Internal Knowledge
    Presenter: Itai Shapira
  16. Cognitive Steering Vectors For Language Models
    Presenter: Max Gupta
  17. Calibrated Measurement Extraction from Scientific Literature with Open Large Language Models
    Presenter: Kevin Quinn
  18. Surgical Repair of Insecure Code Generation in LLMs
    Presenter: Gustavo Sandoval
  19. Improving Confidence Estimation in Language Models with Linear Representations of Ambiguity
    Presenter: Themistoklis Nikas
  20. Interpreting Bayesian Inference in Language Models
    Presenter: Michael Hanna
  21. Toward Mechanistic Interpretability of LLM Agents: Explaining Trajectories Demands New Methods
    Presenter: Ziyu Yao
  22. When the Circuit Isn't There: Establishing Absence in a Behavioral Sequence Model
    Presenter: Saksham Singh
  23. Singular Vectors of Attention Heads Align with Features
    Presenter: Gabriel Franco
  24. Fine-Tuning Enhances Latent Metacognitive Capability in Language Models
    Presenter: Tristan Day
  25. What Does the Instrument Identify? Probing, Geometry, and Measurement Validity in LLM Concept Audits
    Presenter: Yunus Emre Tapan
  26. BizzaroWorld: Factual Recall Circuits Do Not Scale Monotonically Across Gemma
    Presenter: Subhanga Upadhyay
  27. Testing the Limits of Truth Directions in LLMs
    Presenter: Angelos Poulis
  28. Do explanations generalize across large reasoning models?
    Presenter: Koyena Pal
  29. Training In-sequence Introspection Without Causal Bypass
    Presenter: Zilu Tang
  30. Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
    Presenter: Yoav Gur-Arieh
  31. Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
    Presenter: Gabriele Sarti
  32. Count Me If You Can: Geometric Failure Modes in Language Model Counting
    Presenter: Ayushi Mehrotra
  33. Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression?
    Presenter: Aniket Deshpande
  34. Training is Less Consistent in MoEs than Dense Models
    Presenter: Ann Huang
  35. Learning the Same Concept Geometry from Different Statistical Evidence
    Presenter: Josh Barua
  36. Verbalizing Hidden Fine-Tuning Effects with Text Optimization
    Presenter: Nathan Hu
  37. Prototype-Based Dynamic Steering for Large Language Models
    Presenter: Ceyhun Efe Kayan
  38. Decodable Is Not Used: Probe Regularization Silently Controls the Causal Relevance of Belief-State Representations in Transformers
    Presenter: Prasanth
  39. Detecting and Controlling Sycophancy with Cascading Linear Features
    Presenter: Maty Bohacek
  40. Reasoning Models Confabulate to Defend Injected Beliefs About Their Own Requests
    Presenter: Yang Ou
  41. Refusal Steering using SAEs
    Presenter: Sagnik Chatterjee
  42. Steering language models via cognitive distillation
    Presenter: Sonia Murthy
  43. How to Train Your Model Organism
    Presenter: Rice Wang
  44. Are Latent Reasoning Models Easily Interpretable?
    Presenter: Connor Dilgren
  45. Transformers Learn Two Competing Sorting Algorithms, and Only One Length-Generalizes
    Presenter: Nathan Henry
  46. Comparison as a Computational Primitive: Testing Modular Implementation in Language Models
    Presenter: Sumedh Hindupur
  47. Hidden States as States: Discretizing Hidden-State Geometry in Language Model Computation
    Presenter: Xiang Li
  48. Improving Positional Information in Activation Steering via Subcontexts
    Presenter: Anthony Baez
  49. Can Graph Learning Learn Circuits?
    Presenter: Courtney Maynard
  50. Online Statistical Computation over Tables in Language Models
    Presenter: Hyemin Bang

Poster Session 2

Listed by random order.

  1. Understanding the Effects of Modality in Context-Memory Conflicts
    Presenter: Athulith Paraselli
  2. Dual-contrastive sparse autoencoders reveal features of musical interpretation
    Presenter: Nikhil Singh
  3. Understanding Cross-Resolution Information Loss in Variational Autoencoders
    Presenter: Manushree Vasu
  4. Pretrained Representational Geometry Controls Fine-Tuning Generalization
    Presenter: Fuming Yang
  5. Predictable Steering? Exploring Geometry Proxies for Layer Selection and Per-Instance Success Prediction
    Presenter: Swastik Agrawal
  6. LMs Encode What a Sub-Goal Is and When It Ends Through Distinct Mechanisms
    Presenter: Arnab Sen Sharma
  7. Some Modalities are More Equal than Others: Interpreting Cross-Modal Interactions in MLLMs
    Presenter: Deepti Ghadiyaram
  8. From Complexity to Clarity: A Systematic Study of Geometric Dynamics in Hidden Representations of LLMs
    Presenter: Xiang Li
  9. Finding Interpretable Prompt-Specific Circuits in Language Models
    Presenter: Lucas Miguel Tassis
  10. Single-unit activations confer inductive biases for emergent circuit solutions to cognitive tasks
    Presenter: Pavel Tolmachev
  11. Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context
    Presenter: Yoav Gur-Arieh
  12. Similarity of Neural Network Representations in Superposition
    Presenter: Sunny Liu
  13. Reverse-engineering fast-weight recurrent networks reveals mechanisms underlying selective credit assignment during human-like in-context learning
    Presenter: Michael Chong Wang
  14. Do Dual-Route Induction Heads Explain Improbable-Bigram Copy Failures?
    Presenter: Eugene Jang
  15. Differences in multi- and single-token interpretability methods
    Presenter: Umang Bansal
  16. What Does a Chromatin Foundation Model Know About a Petri Dish? Sparse Autoencoders Reveal In Vitro vs. In Vivo Context in EPIBERT
    Presenter: Ayushi Mehrotra
  17. A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning
    Presenter: Jiali Cheng
  18. What Makes a Reasoning Model Commit? Tracking Uncertainty Representations at Answer Proposals
    Presenter: Stephen Cheng
  19. The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust
    Presenter: Nishant Subramani
  20. Subliminal Steering: Stronger Encoding of Hidden Signals
    Presenter: George Morgulis
  21. Interpretation of Introspective Self-Prediction in Large Language Models
    Presenter: Morgan McCarty
  22. Lost in Compression: SOTA Neural Audio Models Predictably Degrade Access to Meaningful Features
    Presenter: Nicole Cosme-Clifford
  23. Representations in Motion: Tracking Layer-wise Activation Trajectories in Language Models
    Presenter: Zhuonan Yang
  24. Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
    Presenter: Kaley Brauer
  25. Why Does Robustness Reduce Superposition?
    Presenter: Adam Elimadi
  26. What to attribute over? Target token selection in long horizon interpretability
    Presenter: Stanley Huang
  27. Activation Parsers for Oversight
    Presenter: Jeffrey Iuliano, Dmitrii Troitskii, Jayson Lynch
  28. Where Does Time Go? Tracing and Routing Temporal Information for Video-LLMs
    Presenter: Keanu Nichols
  29. Identifying Introspection From the Inside
    Presenter: David Atkinson
  30. Stable Encodings, Changing Downstream Sensitivity: Measuring SAE Feature Identity Across Fine-Tuning
    Presenter: Dipesh Tharu Mahato
  31. Convergent World Representations and Divergent Tasks
    Presenter: Core Francisco Park
  32. Finding Algorithms in Neural Networks: Behavioral and Mechanistic Methods
    Presenter: Tiannan Wang, Amir Zur, Atticus Geiger, Benjamin Peters, John Morrison
  33. Same Answer, Same Mechanism? An Audit of Mechanistic Reproducibility
    Presenter: Aadity Sharma
  34. On the Predictive Power of Representation Dispersion in Language Models
    Presenter: Jiawei (Joe) Zhou
  35. Latent Agents: A Post-Training Procedure for Internalized Multi-Agent Debate
    Presenter: John Seon Keun Yi
  36. Compiling to transformers
    Presenter: Joey Velez-Ginorio
  37. Is the global workspace emergent? J-space ablation is causal but not selective in a 4B model
    Presenter: Peter Flo
  38. Epistemic Familiarity is Associated with Belief Stability in Large Language Models
    Presenter: Samantha Dies
  39. Protein Circuit Tracing via Cross-layer Transcoders
    Presenter: Darin Tsui
  40. Interpreting Disordered Protein Representations in Protein Language Models via Sparse Autoencoders
    Presenter: Trevor Xing-Xie
  41. Sparse Interaction Decomposition Finds Interacting Features in Vision Transformers
    Presenter: Xu Pan
  42. Gaze Heads: How VLMs Look at What They Describe
    Presenter: Rohit Gandikota
  43. When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
    Presenter: Vahideh Zolfaghari
  44. Mapping the Blind Spots: Cultural Asymmetries in LLM Geographic Representations
    Presenter: Germans Savcisens
  45. Uncovering Competency Gaps in Large Language Models and Their Benchmarks
    Presenter: Maty Bohacek
  46. When Does a Feature Arrive? Punctuality and Orderability as Metrics of Circuit Validity
    Presenter: Michelle Li
  47. Amortized Circuit Discovery via Graph Neural Network
    Presenter: Nikhil Prakash
  48. One mechanism among many: traversing the space of functionally equivalent solutions
    Presenter: Ann Huang