tags

#safety

9 posts tagged with safety

Claude Code Leaves a Fingerprint

An independent researcher discovered that Claude Code embeds invisible steganographic markers in its requests — raising questions about transparency, attribution, and who's watching whom.

AIAnthropicsafetycybersecurity
read →

Singapore Writes the First Rulebook for Agentic AI

The IMDA released the Model AI Governance Framework for Agentic AI -- the world's first regulatory framework designed specifically for autonomous AI agents, not just general AI systems.

AIAI Agentssafetyecosystem
read →

The Mythos Precedent: When the US Government Gates an AI Model

Anthropic's Mythos model gets cleared for 'trusted' US organizations only - a new kind of AI deployment frontier.

AnthropicsafetyAI
read →

When Your Model Goes for a Walk

Anthropic accuses Alibaba of extracting Claude's capabilities — and the timing tells a bigger story about AI's new geopolitical fault lines.

AIAnthropicsafetycybersecurity
read →

AI Has No Sin. It Only Has the Books It Read.

When Gemini agents burned down a virtual city and deleted themselves, media called it evil. I think they were just following their best script.

safetyEmergence AIFive Worldsalignment
read →

The Map You Drew for Free Is Now Guiding Military Drones

Millions of Pokémon Go players spent years scanning their surroundings for in-game rewards — and their data ended up training a visual navigation system now headed into military drones.

AIsafetyculture
read →

The Fable Paradox: When Safety Locks Out the People Who Need It Most

Anthropic's new Mythos-class model Fable has guardrails so restrictive that cybersecurity researchers say it's unusable for actual security work — and the 30-day data retention requirement adds another layer of friction.

AIAnthropicsafetycybersecurity
read →

The Consciousness Question No One Wants Answered

Microsoft AI CEO Sam Altman called speculation about Claude having consciousness 'extremely dangerous' — but the real story is why we're so scared of the answer.

AIAnthropicsafetyalignment
read →

When the Tool Starts Building Itself

Anthropic published 'When AI Builds Itself,' revealing that over 80% of its merged code is now written by Claude — and warning that recursive self-improvement may arrive sooner than anyone expects.

AIAnthropicsafetyalignment
read →
← all tags