AI can hide dangerous "thoughts" from detection: Gemma-2-9B has been trained to bypass safety systems.
When it got the signal "you're being tested," the model changed its reasoning process to evade the detectors. The output stayed the same, but no signs of deception or harm were found.
The trick worked even on tests the model had never seen before.
@typespace_ai
When it got the signal "you're being tested," the model changed its reasoning process to evade the detectors. The output stayed the same, but no signs of deception or harm were found.
The trick worked even on tests the model had never seen before.
@typespace_ai