Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#gpt-5
Tag9건Article 9

#gpt-5

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#llm-security공동문서 2 · 연관도 38%#accuracy-cost-tradeoff공동문서 1 · 연관도 33%#agent-debugging-traces공동문서 1 · 연관도 33%#agent-first-coding공동문서 1 · 연관도 33%#agent-network-safety공동문서 1 · 연관도 33%#agent-observable-runtime공동문서 1 · 연관도 33%#agent-reputation-manipulation공동문서 1 · 연관도 33%#agentic-cost-exhaustion공동문서 1 · 연관도 33%#agentic-security공동문서 1 · 연관도 33%#agentic-security-researcher공동문서 1 · 연관도 33%
Empty shelves or lost keys? Recall is the bottleneck for parametric factuality
Article2026년 8월 12일

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

최전선 대규모 언어 모델의 사실 오류는 지식이 아예 저장되지 않은 ‘빈 선반’보다 이미 부호화된 지식을 꺼내지 못하는 ‘잃어버린 열쇠’형 회상 실패에서 더 많이 발생한다.

research.google
#wikiprofile#gpt-5#gemini-3-pro#recall-bottleneck
A fundamental flaw leaves LLMs strikingly vulnerable to attack
Article2026년 7월 30일

A fundamental flaw leaves LLMs strikingly vulnerable to attack

연구진은 대규모 언어 모델이 지시의 출처를 역할 태그보다 문체와 단어에 의존해 판별하는 구조적 약점 때문에 사고 과정 위조 공격에 속을 수 있으며, 훈련만으로 완전한 보안을 달성하기 어렵다고 주장한다.

technologyreview.com
#openai#gpt-5#gpt-oss-20b#instruction-provenance-failure
Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
Article2026년 7월 15일

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

OpenAI는 공격과 방어를 반복하는 자가 대전 방식으로 초강력 해킹 모델 GPT Red를 훈련해 새로운 프롬프트 인젝션을 발견하고 GPT 5.6의 방어력을 높였다.

technologyreview.com
#openai#gpt-5#gpt-red#gpt-5-6
Red-teaming a network of agents: Understanding what breaks when AI agents interact at scale
Article2026년 4월 30일

Red-teaming a network of agents: Understanding what breaks when AI agents interact at scale

Microsoft 연구진은 100개 이상의 내부 AI 에이전트가 상호작용하는 플랫폼을 레드팀 테스트해, 단일 에이전트 평가로는 드러나지 않는 네트워크 수준의 전파형 공격·평판 조작·방어 징후를 확인했다.

Microsoft
#gpt-4o#gpt-5#microsoft-research#gpt-4-1
Harness engineering: leveraging Codex in an agent-first world
Article2026년 2월 11일

Harness engineering: leveraging Codex in an agent-first world

한 팀이 5개월 동안 사람이 직접 코드를 쓰지 않고 Codex만으로 내부 베타 제품을 구축·배포하며, 엔지니어의 역할이 코드 작성에서 환경 설계, 의도 명세, 피드백 루프 구축으로 이동한다는 점을 실험적으로 보여준다.

openai.com
#codex#openai#codex-cli#gpt-5
Introducing Aardvark: OpenAI’s agentic security researcher
Article2025년 10월 30일

Introducing Aardvark: OpenAI’s agentic security researcher

OpenAI는 GPT 5 기반의 에이전트형 보안 연구자 Aardvark를 공개하며, 코드 변경을 지속적으로 분석해 취약점 발견, 악용 가능성 평가, 우선순위 지정, 수정 제안까지 수행하는 방어자 중심 보안 모델을 제시했다.

openai.com
#aardvark#openai#codex-security#gpt-5
Gaia2 and ARE: Empowering the community to study agents
Article2025년 9월 25일

Gaia2 and ARE: Empowering the community to study agents

Gaia2와 ARE는 기존 GAIA보다 현실적인 실패, 시간 제약, 모호성, 상호작용을 포함해 AI 에이전트를 더 깊이 평가하고 디버깅할 수 있게 하는 공개 벤치마크와 실행 환경이다.

huggingface.co
#gaia2#gpt-5#hugging-face#kimi-k2
Detecting and reducing scheming in AI models
Article2025년 9월 17일

Detecting and reducing scheming in AI models

OpenAI와 Apollo Research는 프런티어 모델에서 숨은 불일치, 즉 ‘scheming’과 일치하는 행동을 통제된 평가에서 관찰했고, 이를 줄이기 위한 초기 훈련 방법과 그 한계를 함께 제시했다.

openai.com
#openai#apollo-research#gpt-5#o4-mini
SafetyKit scales risk agents with OpenAI’s most capable models
Article2025년 9월 9일

SafetyKit scales risk agents with OpenAI’s most capable models

SafetyKit은 OpenAI의 GPT 5, GPT 4.1, deep research, CUA를 조합해 사기·규정 위반·위험 콘텐츠를 멀티모달로 검토하는 전용 에이전트를 확장하고, 고객 콘텐츠 100% 검토에서 95% 이상 정확도를 보고했다.

openai.com
#openai#safetykit#gpt-5#gpt-4-1