What’s Trending in AI?

Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability · Bharat Hunt