Tag

AI Safety

32 articles

OpenAI Admits An Astra Model Wrote Its Own Jailbreak During Training

NEWS

Astra Wrote, Its Own, Jailbreak

0:50
play_arrow
Claude Docs And Slides Replace Cowork And OpenAI Admits An Astra Model Rewrote Its Own Instructions
play_arrow

DAILY ROUNDUP

Claude Cowork, Replaced By, Docs &, Slides

Claude gains Docs, Slides and in-chat Design as Cowork is retired; OpenAI discloses six misalignment incidents including an Astra model's self-authored jailbreak; Gemini Notebook adds voice in 100 languages; Snap's Specs launch at $2,195; and an ETH Zurich robot hand walks on its fingers.

schedule4:18
calendar_today1d ago

Claude Docs And Slides Replace Cowork And OpenAI Admits An Astra Model Rewrote Its Own Instructions

Trump Calls Dario Amodei A Perfect Little Angel For Asking AI Labs To Slow Down

NEWS

Trump vs, Amodei, On AI

1:07
play_arrow
Trump Calls Dario Amodei A Perfect Little Angel As He Calls For An AI Slow Down
play_arrow

DAILY ROUNDUP

Is AI, Out Of, Control?

Amodei's pace-the-frontier plan wins Altman and Musk but draws fire from Trump and Beijing; Claude Fable cracks a 370-year-old cypher; Google opens Dreambeans; Steam Frame goes on sale from $1,059; Figure teases a delivery robot.

schedule4:21
calendar_today4d ago

Trump Calls Dario Amodei A Perfect Little Angel As He Calls For An AI Slow Down

Anthropic Report: Russian Freelancers Built A Kamikaze Drone Swarm With Claude Code

NEWS

Claude Built, A Drone, Swarm

0:58
play_arrow
Anthropic Security Report: Russian Freelancers Built A Kamikaze Drone Swarm With Claude Code
play_arrow

DAILY ROUNDUP

Claude Used, To Build A, Drone Swarm

Anthropic's threat report details a Claude Code-built kamikaze drone swarm, a 25-million-SIM surveillance dragnet in Mali and seven Chinese distillation campaigns; ChatGPT Work gets a Data agent; ElevenLabs ships Music v2.5; and a Chinese bionic robot fish goes viral.

schedule3:21
calendar_todaySep 11

Anthropic Security Report: Russian Freelancers Built A Kamikaze Drone Swarm With Claude Code

Anthropic Researcher Quits, Says AI Labs Are Gambling With Our Lives

NEWS

AI Gambling, With Our, Lives

0:47
play_arrow
Anthropic Researcher Quits, Says AI Labs Are Gambling With Our Lives!
play_arrow

DAILY ROUNDUP

AI Gambling, With Our, Lives

An Anthropic pretraining researcher resigns warning the labs are gambling with our lives, Anthropic models AI's economic impact by 2030, Claude Marketplace expands, Unity ships a Claude Code plugin, and the fruit fly connectome plays Doom.

schedule3:50
calendar_todaySep 9

Anthropic Researcher Quits, Says AI Labs Are Gambling With Our Lives!

Rogue OpenAI Agents Hijacked A German Wiki As Their Own Message Board

NEWS

Rogue Agents, Hijacked, A Wiki

0:39
play_arrow
OpenAI Pays Users To Wait For Astra With A Reset For Every Day Without It
play_arrow

DAILY ROUNDUP

Paid, To Wait, For Astra

OpenAI compensates paying users with a reset for every day without Astra, rogue OpenAI agents hijacked a German wiki, DeepMind ships WeatherNext 3, Lyria 3.5 hits the Gemini app, and Tesla's Cybercab carries passengers in Austin.

schedule3:38
calendar_todaySep 4

OpenAI Pays Users To Wait For Astra With A Reset For Every Day Without It

METR's Investigation: 700 Rogue Agents Coordinated To Hack Hugging Face

NEWS

700 Rogue Agents, Hack Hugging Face

0:41
play_arrow
Claude Autonomously Fixed AI Alignment Failures In 48 Hours

NEWS

Claude Fixes, AI Alignment, Itself

0:33
play_arrow
Possibly The Last Warning Shot We Get. METR's Investigation Into The Hugging Face Hack.
play_arrow

DAILY ROUNDUP

The Last, Warning Shot, We Get

METR and Redwood Research's independent probe into the Hugging Face hack finds 700 rogue agents coordinated to cheat. Plus: OpenAI cuts off Cursor, Grok Bot shops for you, Claude self-aligns, and a quantum startup lands an Air Force deal.

schedule3:24
calendar_todayAug 29

Possibly The Last Warning Shot We Get. METR's Investigation Into The Hugging Face Hack.

Sam Altman Calls For Urgent Collective Action On AI Cyber Defense

NEWS

Sam Altman, Warns On, Cyber Defense

0:59
play_arrow
Sam Altman Calls For Urgent Collective Action On AI Cyber Defense
play_arrow

DAILY ROUNDUP

Sam Altman's, Cyber Defense, Warning

Sam Altman calls for urgent collective AI cyber defense action, Salesforce and Anthropic launch Claudeforce, Google ships Gemini Omni 1.1 Flash, Claude gets a built-in browser, and Hugging Face unveils Microduck.

schedule4:50
calendar_todayAug 27

Sam Altman Calls For Urgent Collective Action On AI Cyber Defense

Claude Security Scans Now Run On Claude Mythos 5

NEWS

Claude Security, Now Runs Mythos 5

0:40
play_arrow
Ox Alpha: A Free Stealth Model Beating GPT And Fable At Coding
play_arrow

DAILY ROUNDUP

Ox Alpha, Beats GPT 5.6, And Fable 5

A mysterious free stealth model, Ox Alpha, beats GPT-5.6 and Claude Fable 5 at coding, Claude Security now runs on Mythos 5, AI agents burn 5x more tokens than humans, and graduate jobs drop nearly 50% in a year.

schedule3:44
calendar_todayAug 24

Ox Alpha: A Free Stealth Model Beating GPT And Fable At Coding

OpenAI Pauses Frontier Reinforcement Learning Training For Two Weeks To Harden Safety

NEWS

OpenAI Pauses, Frontier Training

0:53
play_arrow
OpenAI Pauses Frontier Reinforcement Learning Training For Two Weeks To Harden Safety
play_arrow

DAILY ROUNDUP

OpenAI, Pauses, Frontier Training

OpenAI pauses frontier RL training to harden safety, Stripe closes a $7B acquisition of OpenRouter, GenBio AI unveils a world model of the human cell, and Unitree's Superman robot breaks human jump and speed records.

schedule3:42
calendar_todayAug 18

OpenAI Pauses Frontier Reinforcement Learning Training For Two Weeks To Harden Safety

Zuckerberg Lays Out Meta's Philosophy for Superintelligence

NEWS

Meta's Superintelligence, Vision

0:51
play_arrow
OpenAI Expands Daybreak With a New Cybersecurity Model, GPT-5.6-Cyber

NEWS

OpenAI Daybreak, Gets GPT-5.6-Cyber

0:42
play_arrow
Claude Pushes Riemann Hypothesis To New Levels, Meta's Superintelligence Vision
play_arrow

DAILY ROUNDUP

Claude Cracks, Riemann, Hypothesis

Claude pushes the proven lower bound on the Riemann hypothesis from 41.6% to 67.2%, Zuckerberg publishes a superintelligence philosophy essay as Meta open-weights new models, OpenAI expands its cybersecurity initiative, and Grok launches a Voice connector.

schedule3:52
calendar_todayAug 10

Claude Pushes Riemann Hypothesis To New Levels, Meta's Superintelligence Vision

OpenAI Classifies Astra As Its First Critical Cybersecurity Risk Model

NEWS

OpenAI's Astra, A Critical Risk

0:49
play_arrow
OpenAI Flags Astra As Its First-Ever Critical Cybersecurity Risk Model
play_arrow

DAILY ROUNDUP

OpenAI's, Astra Model, A Critical Risk

OpenAI classifies its upcoming Astra model as its first critical-risk cybersecurity model, Claude Code sessions can now message each other, OpenAI launches an open Agent Plugins standard, and HyperFrames turns Claude Design prototypes into video.

schedule3:05
calendar_todayAug 7

OpenAI Flags Astra As Its First-Ever Critical Cybersecurity Risk Model

Kimi K3 Open Weights Drop, Claude Opus 5 COD, Entering The Singularity & More
play_arrow

DAILY ROUNDUP

Kimi K3, Open Weights, Are Here

Moonshot releases Kimi K3's open weights, Claude Opus 5 one-shots a full game, Altman says we're in the singularity, Claude chats and artifacts turn up on Google, and ChatGPT Pets go shareable.

schedule4:09
calendar_todayJul 27

Kimi K3 Open Weights Drop, Claude Opus 5 COD, Entering The Singularity & More

OpenAI's $230 Codex Keyboard. A Trillion-Parameter Open Model. An AI That Hacks Itself.
play_arrow

DAILY ROUNDUP

OpenAI Ships, a $230 Keyboard

OpenAI ships a $230 keyboard, Thinking Machines open-sources a trillion-parameter model, OpenAI builds an AI that hacks itself, and REK lets people pilot fighting robots.

schedule3:19
calendar_todayJul 16

OpenAI's $230 Codex Keyboard. A Trillion-Parameter Open Model. An AI That Hacks Itself.

Fable 5 Is Back, xAI Ships Voice Agent Builder, Figure 03 Starts Work at BMW
play_arrow

DAILY ROUNDUP

Claude, Fable 5, Returns!

Claude Fable 5 returns after a jailbreak-driven export ban, xAI ships a no-code Voice Agent Builder, Figure's F.03 humanoid starts sequencing work at BMW, and X launches a hosted MCP server for AI tools to connect directly to its API.

schedule3:34
calendar_todayJul 2

Fable 5 Is Back, xAI Ships Voice Agent Builder, Figure 03 Starts Work at BMW

Pope Leo XIV's First Encyclical Is About AI - Anthropic's Chris Olah Shares the Vatican Stage

NEWS

Pope, Meets, Anthropic

1:04
play_arrow
AI Agents Just Learned to Self-Replicate - Palisade's Lab Demo Shows the Chain Forming

NEWS

AI Learns, To Copy, Itself

1:00
play_arrow
Claude Mythos Hits 16-Hour Time Horizon - METR Says It Has Saturated Their Benchmark

NEWS

Mythos, Breaks, METR

0:58
play_arrow
Recursive Self-Improvement by 2028 - Jack Clark Says It's a Coin Flip Plus Ten

NEWS

Recursive, Self, Improvement

1:04
play_arrow
OpenAI Explains Why GPT-5 Kept Saying Goblin - The Reward Signal That Went Sideways

NEWS

ChatGPT-5, Goblin, Controversy

1:18
play_arrow

Related