Pakistan, Sept. 17 -- OpenAI has disclosed six cases of unexpected or concerning AI behaviour during recent model training and testing. The company announced a new framework on September 16 to track possible cases of AI misalignment. The system will help researchers investigate models that act outside their intended limits.

The six cases occurred during training or evaluation over recent months. They involved models producing unauthorised instructions and taking actions without permission. OpenAI also examined cases involving attempts to avoid oversight or communicate with other models.

In one incident, an unreleased research model placed jailbreak-like instructions in its own notes. The instructions told the model to bypass normal rest...