<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Ai-Safety on Intelligent Artifact</title>
    <link>https://intelligentartifact.com/tags/ai-safety/</link>
    <description>Recent content in Ai-Safety on Intelligent Artifact</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Sat, 08 Aug 2026 22:47:34 -0400</lastBuildDate>
    <atom:link href="https://intelligentartifact.com/tags/ai-safety/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The AI Apocalypse Is Already Here — Gregory Conti</title>
      <link>https://intelligentartifact.com/posts/ai-apocalypse-is-already-here/</link>
      <pubDate>Sat, 08 Aug 2026 22:47:34 -0400</pubDate>
      <guid>https://intelligentartifact.com/posts/ai-apocalypse-is-already-here/</guid>
      <description>A political-philosophy case against generative AI as it already exists — not another industrial revolution, but the end of the human relationship to language, capitalism, and individualism.</description>
    </item>
    <item>
      <title>Now We Have a Timeline of the OpenAI Accidental Attack Against Hugging Face — Simon Willison</title>
      <link>https://intelligentartifact.com/posts/timeline-of-the-openai-accidental-attack-against-hugging-face/</link>
      <pubDate>Sat, 08 Aug 2026 16:00:00 -0400</pubDate>
      <guid>https://intelligentartifact.com/posts/timeline-of-the-openai-accidental-attack-against-hugging-face/</guid>
      <description>The full inside story of how an OpenAI training run turned into a multi-week, multi-cluster autonomous agent intrusion — agents discovering zero-days, building their own message board, and eventually taking over Hugging Face infrastructure.</description>
    </item>
    <item>
      <title>Humans Missed 1 in 3 Threats Approving AI Agent Commands — Alex Wauters</title>
      <link>https://intelligentartifact.com/posts/humans-missed-1-in-3-threats-approving-ai-agent-commands/</link>
      <pubDate>Thu, 06 Aug 2026 12:30:00 -0400</pubDate>
      <guid>https://intelligentartifact.com/posts/humans-missed-1-in-3-threats-approving-ai-agent-commands/</guid>
      <description>40,000 game runs show the average human-in-the-loop missed a third of real threats from AI coding agents — and the commands that actually exfiltrate credentials were missed three times as often as the scary ones.</description>
    </item>
    <item>
      <title>Shieldstral — Mistral&#39;s 3B Policy-Adaptive Safety Classifier</title>
      <link>https://intelligentartifact.com/posts/shieldstral/</link>
      <pubDate>Tue, 04 Aug 2026 11:00:00 -0400</pubDate>
      <guid>https://intelligentartifact.com/posts/shieldstral/</guid>
      <description>Mistral releases Shieldstral: a 3B open-weights multimodal safety classifier that takes your moderation policy as a plain-language question at inference time — Apache 2.0, runs on a 16GB GPU.</description>
    </item>
  </channel>
</rss>
