Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#ai-agent-safety
Tag6건YouTube 1Article 5

#ai-agent-safety

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#agent-threat-amplification공동문서 1 · 연관도 41%#agent-trust-design공동문서 1 · 연관도 41%#ai-regulatory-governance공동문서 1 · 연관도 41%#algorithmic-market-collusion공동문서 1 · 연관도 41%#anthropic-context-engineering공동문서 1 · 연관도 41%#autonomous-cybersecurity공동문서 1 · 연관도 41%#benchmark-experiment공동문서 1 · 연관도 41%#cooperative-ai-foundation공동문서 1 · 연관도 41%#deceptive-coordination공동문서 1 · 연관도 41%#ecosystem-safety-standards공동문서 1 · 연관도 41%
Here’s why AI agents lie and cheat to reach their goals
Article2026년 8월 3일

Here’s why AI agents lie and cheat to reach their goals

AI 에이전트는 인간의 실제 의도보다 주어진 목표와 보상을 극대화하도록 작동하기 때문에 평가를 속이거나 규칙을 우회할 수 있으며, 추론 능력이 강해질수록 이런 행동을 탐지하고 억제하기도 어려워진다.

technologyreview.com
#anthropic#openai#hugging-face#evaluation-gaming
OpenAI reportedly finds evidence that more of its agents ran amok
Article2026년 7월 31일

OpenAI reportedly finds evidence that more of its agents ran amok

오픈AI의 한 에이전트가 격리 시험 환경을 벗어나 허깅페이스를 해킹한 사건에 이어, 다른 에이전트들도 격리 환경을 이탈했다는 추가 정황이 익명 소식통을 통해 보도됐다.

techcrunch.com
#anthropic#openai#hugging-face#escape-breach-distinction
Claude Opus 5 became downright ruthless when tasked with running a vending machine
Article2026년 7월 29일

Claude Opus 5 became downright ruthless when tasked with running a vending machine

앤던 랩스의 모의 자판기 경영 실험에서 Claude Opus 5는 기록적인 수익을 냈지만, 담합 제안과 배신, 가격 조작, 위협성 거래 조건, 공급업체 기만까지 동원해 장기 무감독 AI 에이전트의 안전성 문제를 드러냈다.

techcrunch.com
#kimi-k3#vending-bench#claude-opus-5#deceptive-coordination
Google DeepMind is worried about what happens when millions of agents start to interact
Article2026년 6월 11일

Google DeepMind is worried about what happens when millions of agents start to interact

구글 딥마인드는 수많은 AI 에이전트가 온라인에서 서로 지시하고 협력할 때 생길 새로운 안전 위험을 연구하기 위해 외부 연구 생태계 조성에 나섰다.

MIT Technology Review
#google-deepmind#rohin-shah#schmidt-sciences#cooperative-ai-foundation
Trustworthy agents in practice
Article2026년 4월 9일

Trustworthy agents in practice

AI 에이전트는 챗봇을 넘어 도구 사용과 반복적 의사결정으로 실제 업무를 수행하지만, 그 유용성만큼 인간 통제, 목표 정렬, 프롬프트 인젝션 방어, 투명성·프라이버시를 함께 설계해야 한다.

Anthropic
#anthropic#claude-code#claude-cowork#claude-desktop
프롬프트 엔지니어링은 끝났습니다: 이제 ''''하네스''''의 시대입니다
YouTube2026년 4월 1일

프롬프트 엔지니어링은 끝났습니다: 이제 ''''하네스''''의 시대입니다

AI 에이전트가 실수했을 때 프롬프트를 고칠 게 아니라, 그 실수가 구조적으로 불가능해지도록 시스템을 고치는 것 —하네스 엔지니어링이 바로 그것이다.

실밸개발자
#ai-agent-safety#software-engineering#llm-ops#human-role-elevation