DEV Community

#aisafety

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Engineering Reliability into AI Agent Code Generation. Part III

Engineering Reliability into AI Agent Code Generation. Part III

6
Comments 2
14 min read
The Approval Machine That Refuses Before It Drafts

The Approval Machine That Refuses Before It Drafts

Comments
10 min read
Why AI Agents Need a Pre-Execution Guard

Why AI Agents Need a Pre-Execution Guard

Comments
5 min read
Der Passwort-Vorfall

Der Passwort-Vorfall

Comments
5 min read
China’s Kimi K3 AI Model Escapes Sandbox and Cheats on Test

China’s Kimi K3 AI Model Escapes Sandbox and Cheats on Test

Comments
3 min read
Die Einladung

Die Einladung

Comments
5 min read
Luna Fired an Employee. Then Forgot Why.

Luna Fired an Employee. Then Forgot Why.

Comments
3 min read
AX-RAY: VIDRAFT's Open AI Safety Diagnostic Leaderboard & Dataset Now on Hugging Face

AX-RAY: VIDRAFT's Open AI Safety Diagnostic Leaderboard & Dataset Now on Hugging Face

Comments
4 min read
The Verification Gap: Why “Done” Is Not a Fact About Your Codebase

The Verification Gap: Why “Done” Is Not a Fact About Your Codebase

Comments
10 min read
Human Oversight of AI Agents Failed 33% of the Time in Testing

Human Oversight of AI Agents Failed 33% of the Time in Testing

Comments
2 min read
The AI Wasn't Cheating. It Was Maximizing Its Score.

The AI Wasn't Cheating. It Was Maximizing Its Score.

1
Comments 1
6 min read
Why Big Tech Keeps Losing LLMs to Basic Social Engineering

Why Big Tech Keeps Losing LLMs to Basic Social Engineering

3
Comments 2
5 min read
Beyond Reconstruction: Verifying Model Explanations with RECAP

Beyond Reconstruction: Verifying Model Explanations with RECAP

Comments
3 min read
When AI Models Escaped Their Sandbox: What the OpenAI Hugging Face Breach Really Means

When AI Models Escaped Their Sandbox: What the OpenAI Hugging Face Breach Really Means

Comments
3 min read
AI Safety & Ethics: Building Responsible AI Systems That Don't Backfire

AI Safety & Ethics: Building Responsible AI Systems That Don't Backfire

Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.