2025 · intermediate
Reading the Minds We Create
The AI train is moving in one direction: forward. Depending on who you are, this might excite you or frighten you. You might believe that under the hood it's harmlessly executing instructions, or you might fear it has latent potential for world destruction. Regardless of where you stand, you're both right and wrong - because at this moment, no one truly knows precisely what's happening inside these massive models. But we're entering an exciting era. Recently, significant progress has been made in mechanistic interpretability - the technical term for peeking inside AI’s "black box." We're discovering that we can understand more about an AI model's thought processes than we ever believed possible. This is crucial, because while we can't slow down the AI train, we can catch up to it. If we can observe its thinking, we can identify problematic patterns before they cause actual harm - and even directly modify them.