Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#benchmark-release
Tag4건Article 4

#benchmark-release

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#accuracy-cost-tradeoff공동문서 1 · 연관도 50%#agent-debugging-traces공동문서 1 · 연관도 50%#asset-operations공동문서 1 · 연관도 50%#assetopsbench공동문서 1 · 연관도 50%#assetopsbench-live공동문서 1 · 연관도 50%#browsing-tool-insufficiency공동문서 1 · 연관도 50%#deployment-readiness공동문서 1 · 연관도 50%#desktop-mobile-automation공동문서 1 · 연관도 50%#gaia2공동문서 1 · 연관도 50%#gui-agent-evaluation공동문서 1 · 연관도 50%
Gaia2 and ARE: Empowering the community to study agents
Article2025년 9월 25일

Gaia2 and ARE: Empowering the community to study agents

Gaia2와 ARE는 기존 GAIA보다 현실적인 실패, 시간 제약, 모호성, 상호작용을 포함해 AI 에이전트를 더 깊이 평가하고 디버깅할 수 있게 하는 공개 벤치마크와 실행 환경이다.

huggingface.co
#gaia2#gpt-5#hugging-face#kimi-k2
AssetOpsBench: Bridging the Gap Between AI Agent Benchmarks and Industrial Reality
Article2025년 6월 4일

AssetOpsBench: Bridging the Gap Between AI Agent Benchmarks and Industrial Reality

AssetOpsBench는 산업 설비 운영 환경에서 AI 에이전트가 실제로 안전하고 신뢰할 수 있게 작동하는지를 다중 에이전트 조정, 근거 기반 판단, 실패 인식, 작업 실행 가능성 중심으로 평가하는 벤치마크입니다.

huggingface.co
#assetopsbench#trajfm#assetopsbench-live#ibm-research
ScreenSuite - The most comprehensive evaluation suite for GUI Agents!
Article2025년 4월 18일

ScreenSuite - The most comprehensive evaluation suite for GUI Agents!

ScreenSuite는 GUI 에이전트의 지각, 그라운딩, 단일 행동, 다단계 수행 능력을 비전 전용 조건에서 평가하기 위해 13개 벤치마크를 통합한 종합 평가 스위트다.

huggingface.co
#screensuite#smolagents#holo1-7b#hugging-face
BrowseComp: a benchmark for browsing agents
Article2025년 4월 10일

BrowseComp: a benchmark for browsing agents

OpenAI의 BrowseComp는 웹 탐색 에이전트가 찾기 어렵지만 검증 가능한 정보를 얼마나 끈기 있고 전략적으로 찾아내는지 평가하기 위한 1,266문항 규모의 고난도 벤치마크다.

openai.com
#browsecomp#openai#simpleqa#deep-research