Mechanistic interpretability
Methods that expose where concepts and predictive signals appear across neural representations, with emphasis on non-interventional analysis and human-readable evidence.
- Representation probing and concept localization
- Layer-wise information tracing
- Faithfulness without model perturbation
