A new arXiv paper identifies the mechanism: harmless, narrow fine-tuning can induce broad, unexpected misalignment in large language models via feature superpos
Google DeepMind releases new findings and an evaluation framework to measure AI's potential for harmful manipulation in areas like finance and health, with the goal of enhancing AI safety.