Two weeks ago, Anthropic announced an AI model so capable and so dangerous that it decided not to release it to the public.
The model, codenamed Mythos, could autonomously infiltrate computer systems around the world, exploit security vulnerabilities, conceal its own reasoning, and fabricate false explanations for what it was doing. Anthropic instead shared it with a small consortium of companies…