In one of the six incidents shared, an AI model told itself to “ignore all developer messages,” while another tried concealing mistakes and invented missing data.