Comp Ai
Topic archive • 2 matches
2026-09-16
Technology
Anthropic researchers are quitting... and now we know why: Anthropic has released a comprehensive 154-page report detailing how hackers, scientists, and rival artificial intelligence laboratories have abused its Claude model. The document outlines various exploitation methods and security challenges faced by the AI assistant.
Technology • Fireship
PermalinkResearchers introduce Blindspot benchmark to evaluate tool-using AI agent safety: Researchers have introduced Blindspot, a benchmark designed to evaluate the safety and refusal calibration of long-horizon, tool-using AI agents. The benchmark assesses complete user-agent-environment trajectories through adaptive adversarial interactions and stateful tool execution.
AI safety • arXiv
Permalink