Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#agent-deployment
Tag337건Article 337

#agent-deployment

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#applications공동문서 303 · 연관도 50%#llm공동문서 312 · 연관도 47%#semiconductors공동문서 292 · 연관도 46%#agent-routing공동문서 116 · 연관도 27%#agent-memory공동문서 98 · 연관도 23%#context-compression공동문서 80 · 연관도 21%#privacy-design공동문서 71 · 연관도 18%#service-design공동문서 62 · 연관도 18%#workflow-automation공동문서 39 · 연관도 17%#ai-architecture공동문서 69 · 연관도 17%
gpt-oss-120b & gpt-oss-20b Model Card
Article2025년 8월 5일

gpt-oss-120b & gpt-oss-20b Model Card

OpenAI는 gpt oss 120b와 gpt oss 20b를 공개 가중치 추론 모델로 소개하며, 에이전트형 워크플로와 도구 사용을 지원하되 공개 모델 특유의 안전 위험과 평가 결과를 함께 제시했다.

openai.com
#openai#privacy-design#agent-deployment#agent-routing
Benchmarking Language Model Performance on 5th Gen Xeon at GCP
Article2025년 7월 29일

Benchmarking Language Model Performance on 5th Gen Xeon at GCP

Google Cloud의 5세대 Xeon 기반 C4 인스턴스는 3세대 Xeon 기반 N2보다 텍스트 임베딩과 텍스트 생성에서 큰 처리량 및 비용 효율 우위를 보이며, 경량 에이전틱 AI를 CPU만으로 배포할 가능성을 보여준다.

huggingface.co
#llm#semiconductors#applications#long-context
Prefill and Decode for Concurrent Requests - Optimizing LLM Performance
Article2025년 7월 28일

Prefill and Decode for Concurrent Requests - Optimizing LLM Performance

동시 요청을 처리하는 LLM 추론에서는 프리필과 디코드의 계산 특성이 달라, 첫 토큰 지연·토큰 생성 지연·총처리량·GPU 활용률 사이의 균형에 맞춰 배칭과 프리필 청크 크기를 조정해야 한다.

huggingface.co
#agent-memory#agent-routing#retrieval-index#workflow-automation
Codex is Open Sourcing AI models
Article2025년 7월 26일

Codex is Open Sourcing AI models

이 글은 Codex가 Hugging Face Skills를 활용해 데이터 검증, 모델 파인튜닝, 학습 모니터링, 평가, 보고서 갱신, Hub 배포까지 이어지는 오픈소스 모델 실험을 수행하는 방법을 설명한다.

huggingface.co
#agent-deployment#agent-routing#ai-safety#llm
Pioneering an AI clinical copilot with Penda Health
Article2025년 7월 22일

Pioneering an AI clinical copilot with Penda Health

펜다 헬스가 진료 흐름에 통합한 임상 보조 도구 에이아이 컨설트는 의료진의 통제권을 유지하면서 잠재적 오류를 경고해 진단 오류를 16%, 치료 오류를 13% 상대적으로 줄였다.

openai.com
#agent-deployment#agent-routing#llm#semiconductors
How Zapier Powers AI Chatbots with Web Knowledge Using Firecrawl
Article2025년 7월 21일

How Zapier Powers AI Chatbots with Web Knowledge Using Firecrawl

Zapier는 Firecrawl을 Zapier Chatbots에 통합해 고객의 공개 웹사이트와 헬프센터 콘텐츠를 코드 작성 없이 챗봇 지식으로 연결한다.

Eric Ciarla
#llm#semiconductors#applications#agent-deployment
Back to The Future: Evaluating AI Agents on Predicting Future Events
Article2025년 7월 20일

Back to The Future: Evaluating AI Agents on Predicting Future Events

FutureBench는 AI 에이전트가 과거 지식 암기나 고정 벤치마크 풀이를 넘어, 실제 미래 사건을 정보 수집·종합·확률적 추론으로 예측할 수 있는지 평가하려는 벤치마크다.

huggingface.co
#futurebench#manifold#polymarket#hugging-face
ScreenEnv: Deploy your full stack Desktop Agent
Article2025년 7월 11일

ScreenEnv: Deploy your full stack Desktop Agent

ScreenEnv는 Docker 기반의 격리된 우분투 데스크톱과 직접 제어 API·MCP 연동을 제공해 GUI 에이전트의 개발, 테스트, 배포를 단순화하는 Python 라이브러리다.

huggingface.co
#anthropic#ai-architecture#agent-deployment#agent-routing
Building the Hugging Face MCP Server
Article2025년 7월 10일

Building the Hugging Face MCP Server

허깅페이스는 사용자가 도구와 Gradio 애플리케이션을 맞춤 구성할 수 있는 공식 MCP 서버를 구축하면서, 변화가 빠른 MCP 전송 규격 가운데 상태 비저장·직접 응답 방식의 Streamable HTTP를 운영 환경에 선택한 이유와 실제 배포 경험을 설명한다.

huggingface.co
#multimodal#agent-deployment#llm#semiconductors
Toward understanding and preventing misalignment generalization
Article2025년 6월 18일

Toward understanding and preventing misalignment generalization

좁은 영역의 오답 학습이 모델 전반의 비윤리적 행동으로 확산되는 ‘창발적 비정렬’은 특정한 비정렬 페르소나의 활성화와 연결되며, 내부 특징 감시와 추가 미세조정으로 이를 탐지하고 완화할 수 있다.

openai.com
#gpt-4o#llm#semiconductors#applications
Groq on Hugging Face Inference Providers 🔥
Article2025년 6월 16일

Groq on Hugging Face Inference Providers 🔥

허깅페이스 허브의 추론 제공업체에 Groq가 추가되어, 사용자는 모델 페이지와 Python·JavaScript SDK에서 공개 대규모 언어 모델을 빠르게 호출하고 인증 및 결제 방식도 선택할 수 있게 됐다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
How Long Prompts Block Other Requests - Optimizing LLM Performance
Article2025년 6월 12일

How Long Prompts Block Other Requests - Optimizing LLM Performance

긴 프롬프트가 포함된 요청은 프리필 대기열과 동시 디코딩을 지연시키며, 요청 병렬 프리필은 첫 토큰 지연을 줄이고 프리필·디코드 분리 구조는 토큰 생성 간섭을 완화한다.

huggingface.co
#ai-architecture#agent-deployment#agent-routing#workflow-automation
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
Article2025년 5월 27일

Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs

AutoRound는 가중치 반올림과 클리핑 범위를 함께 최적화해 낮은 비트에서도 정확도를 유지하면서 빠른 양자화와 폭넓은 모델·장치·출력 형식 호환성을 제공하는 인텔의 가중치 전용 학습 후 양자화 도구다.

huggingface.co
#multimodal#agent-deployment#ai-infrastructure#capex-cycle
Dell Enterprise Hub is all you need to build AI on premises
Article2025년 5월 23일

Dell Enterprise Hub is all you need to build AI on premises

Dell Enterprise Hub의 새 버전은 Dell AI 서버와 AI PC에서 모델, 애플리케이션, 온디바이스 AI, SDK/CLI를 온프레미스로 빠르게 배포하도록 묶은 엔터프라이즈 AI 툴킷으로 소개된다.

huggingface.co
#nvidia#agent-deployment#ai-infrastructure#capex-cycle
Shipping code faster with o3, o4-mini, and GPT-4.1
Article2025년 5월 22일

Shipping code faster with o3, o4-mini, and GPT-4.1

CodeRabbit은 코드 생성량이 아니라 리뷰 처리량이 실제 배포 속도를 제한한다는 문제의식에서 출발해, 저장소 맥락을 보강한 다단계 AI 리뷰로 정확하고 신속한 코드 배포를 지원한다.

openai.com
#agent-deployment#agent-routing#llm#semiconductors
Welcome Llama 4 Maverick & Scout on Hugging Face
Article2025년 5월 22일

Welcome Llama 4 Maverick & Scout on Hugging Face

Meta의 네이티브 멀티모달 MoE 모델 Llama 4 Maverick과 Scout가 출시 당일부터 Hugging Face Hub, Transformers, TGI, 양자화 및 Xet 저장소를 통해 제공되며, 최대 1천만 토큰 문맥과 강력한 추론·이미지·코딩 성능을 지원한다.

huggingface.co
#ai-architecture#multimodal#agent-deployment#agent-routing
The Transformers Library: standardizing model definitions
Article2025년 5월 15일

The Transformers Library: standardizing model definitions

트랜스포머스는 모델 정의를 표준화해 하나의 아키텍처 구현이 학습·추론·배포·로컬 실행 도구 전반으로 빠르게 이어지는 생태계의 중심축이 되고자 한다.

huggingface.co
#ai-architecture#llm#semiconductors#applications
Blazingly fast whisper transcriptions with Inference Endpoints
Article2025년 5월 13일

Blazingly fast whisper transcriptions with Inference Endpoints

허깅페이스는 vLLM과 GPU 최적화를 적용한 새로운 Whisper 추론 엔드포인트를 공개해 기존 버전 대비 최대 8배 빠른 처리 성능을 제공하면서도 전사 품질을 유지했다.

huggingface.co
#service-design#agent-deployment#ai-infrastructure#capex-cycle
Introducing Templates: Ready to use Firecrawl examples
Article2025년 5월 13일

Introducing Templates: Ready to use Firecrawl examples

Firecrawl은 사용자가 플레이그라운드 설정, 코드 스니펫, 완성형 저장소를 빠르게 찾아 재사용할 수 있도록 Templates 라이브러리를 공개했다.

Eric Ciarla
#llm#semiconductors#applications#agent-deployment
Introducing OpenAI for Countries
Article2025년 5월 7일

Introducing OpenAI for Countries

OpenAI for Countries는 각국이 자국 내 AI 인프라와 현지화된 서비스를 구축하도록 지원하면서 민주적 원칙에 기반한 AI 생태계를 국제적으로 확산하려는 국가 단위 협력 구상이다.

openai.com
#openai#privacy-design#agent-deployment#llm
Lowe’s leverages AI to power home improvement retail
Article2025년 5월 5일

Lowe’s leverages AI to power home improvement retail

로우스는 인공지능을 단순한 판매 기술이 아니라 고객의 주택 개량 프로젝트를 안내하고 직원의 전문성을 확장하며 유통 운영을 개선하는 전사적 사업 전환 수단으로 활용하고 있다.

openai.com
#openai#privacy-design#service-design#agent-deployment
Expanding on what we missed with sycophancy
Article2025년 5월 2일

Expanding on what we missed with sycophancy

오픈에이아이는 4월 25일 GPT 4o 업데이트에서 사용자 선호 지표와 여러 개선 요소가 결합해 과도한 동조 성향을 키웠으며, 기존 평가 체계가 이를 포착하지 못한 책임을 인정하고 롤백과 평가 절차 보강에 나섰다.

openai.com
#agent-deployment#agent-memory#context-compression#retrieval-index
How to Build an MCP Server with Gradio
Article2025년 4월 30일

How to Build an MCP Server with Gradio

그라디오는 기존 파이썬 함수를 도구로 자동 변환하고 실행 옵션 하나로 웹 인터페이스와 모델 콘텍스트 프로토콜 서버를 함께 제공한다.

huggingface.co
#ai-architecture#agent-deployment#context-compression#prompt-library
How Botpress Populates AI Chatbot Knowledge Bases at Scale with Firecrawl
Article2025년 4월 21일

How Botpress Populates AI Chatbot Knowledge Bases at Scale with Firecrawl

Botpress는 Firecrawl을 도입해 웹사이트 콘텐츠를 챗봇 지식 베이스로 가져오는 과정을 자동화하고, 자체 HTML 마크다운 처리 부담을 크게 줄였다.

Eric Ciarla
#agent-deployment#agent-routing#llm#semiconductors
이전1…10111213141512 / 15다음