Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#agent-memory
Tag543건YouTube 19Article 524

#agent-memory

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#retrieval-index공동문서 248 · 연관도 66%#context-compression공동문서 286 · 연관도 60%#semiconductors공동문서 457 · 연관도 56%#applications공동문서 432 · 연관도 56%#llm공동문서 468 · 연관도 55%#agent-routing공동문서 201 · 연관도 37%#ai-architecture공동문서 152 · 연관도 30%#agent-deployment공동문서 98 · 연관도 23%#vision-language-models공동문서 64 · 연관도 20%#multimodal공동문서 66 · 연관도 20%
Ulysses Sequence Parallelism: Training with Million-Token Contexts
Article2025년 7월 26일

Ulysses Sequence Parallelism: Training with Million-Token Contexts

Ulysses Sequence Parallelism은 긴 시퀀스 학습에서 시퀀스와 어텐션 헤드를 함께 나누고 all to all 통신으로 재배치해, 단일 GPU 메모리 한계를 넘어 수십만~백만 토큰 문맥 학습을 가능하게 하는 방식이다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#retrieval-index
TimeScope: How Long Can Your Video Large Multimodal Model Go?
Article2025년 7월 23일

TimeScope: How Long Can Your Video Large Multimodal Model Go?

TimeScope는 1분부터 8시간까지의 영상에 짧은 동영상 클립을 삽입해 검색·정보 종합·세밀한 시간 지각 능력을 측정하며, 최신 비전 언어 모델의 장시간 영상 이해가 아직 제한적임을 보여주는 오픈소스 벤치마크다.

huggingface.co
#multimodal#llm#semiconductors#vision-language-models
Pioneering an AI clinical copilot with Penda Health
Article2025년 7월 22일

Pioneering an AI clinical copilot with Penda Health

펜다 헬스가 진료 흐름에 통합한 임상 보조 도구 에이아이 컨설트는 의료진의 통제권을 유지하면서 잠재적 오류를 경고해 진단 오류를 16%, 치료 오류를 13% 상대적으로 줄였다.

openai.com
#agent-deployment#agent-routing#llm#semiconductors
Consilium: When Multiple LLMs Collaborate
Article2025년 7월 19일

Consilium: When Multiple LLMs Collaborate

Consilium은 서로 다른 역할을 맡은 여러 언어 모델이 구조화된 토론과 외부 조사를 거쳐 합의 또는 최종 분석을 도출하도록 만든 시각적 다중 모델 협업 플랫폼이다.

huggingface.co
#ai-architecture#agent-routing#llm#semiconductors
The AI tools for Art Newsletter - Issue 1
Article2025년 7월 16일

The AI tools for Art Newsletter - Issue 1

2024년 창작 AI는 오픈소스 이미지 생성의 구조적 도약과 개인화 기술의 대중화를 이뤘으며, 2025년에는 비디오·오디오·3D 등 더 다양한 형식으로 발전의 중심이 이동하고 있다.

huggingface.co
#ai-architecture#multimodal#agent-memory#agent-routing
Custom Kernels for All from Codex and Claude
Article2025년 7월 16일

Custom Kernels for All from Codex and Claude

CUDA 커널 개발 지식을 에이전트 스킬로 구조화해 Claude와 Codex가 실제 diffusers·transformers 대상의 커널, PyTorch 바인딩, 빌드 구성, 벤치마크까지 완성하도록 한 사례다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#context-compression
Ettin Suite: SoTA Paired Encoders and Decoders
Article2025년 7월 16일

Ettin Suite: SoTA Paired Encoders and Decoders

에틴은 동일한 공개 데이터 2조 토큰, 모델 구조, 학습 절차를 적용한 1,700만~10억 매개변수 규모의 인코더·디코더 쌍으로, 두 아키텍처를 공정하게 비교하면서 양쪽 모두에서 공개 데이터 기반 최고 수준의 성능을 달성한 모델군이다.

huggingface.co
#ai-architecture#llm#semiconductors#applications
Migrating the Hub from Git LFS to Xet
Article2025년 7월 15일

Migrating the Hub from Git LFS to Xet

허깅페이스는 기존 사용자의 작업 방식을 유지하는 브리지와 무중단 백그라운드 마이그레이션을 기반으로, 50만 개 저장소와 20페타바이트 규모의 허브를 깃 대용량 파일 저장소에서 젯으로 전환하고 있다.

huggingface.co
#agent-routing#llm#semiconductors#applications
Building the Hugging Face MCP Server
Article2025년 7월 10일

Building the Hugging Face MCP Server

허깅페이스는 사용자가 도구와 Gradio 애플리케이션을 맞춤 구성할 수 있는 공식 MCP 서버를 구축하면서, 변화가 빠른 MCP 전송 규격 가운데 상태 비저장·직접 응답 방식의 Streamable HTTP를 운영 환경에 선택한 이유와 실제 배포 경험을 설명한다.

huggingface.co
#multimodal#agent-deployment#llm#semiconductors
Reachy Mini - The Open-Source Robot for Today's and Tomorrow's AI Builders
Article2025년 7월 9일

Reachy Mini - The Open-Source Robot for Today's and Tomorrow's AI Builders

Reachy Mini는 Pollen Robotics와 Hugging Face가 공개한 데스크톱 크기의 오픈소스 로봇으로, 인간 로봇 상호작용과 AI 실험을 Python 기반으로 쉽게 시도하도록 설계된 제품입니다.

huggingface.co
#privacy-design#multimodal#llm#semiconductors
Upskill your LLMs With Gradio MCP Servers
Article2025년 7월 9일

Upskill your LLMs With Gradio MCP Servers

Gradio의 MCP 지원을 활용하면 Hugging Face Spaces의 다양한 AI 도구를 Cursor 같은 LLM 클라이언트에 연결해 이미지 편집, 영상 전사, OCR, 음성 합성 등의 새로운 기능을 부여할 수 있습니다.

huggingface.co
#capex-cycle#llm#semiconductors#applications
Migrating data from Postgres to Convex
Article2025년 7월 8일

Migrating data from Postgres to Convex

Postgres 데이터를 Convex로 이전하는 방법은 소규모 데이터의 JSONL 덤프·가져오기에서 시작해, 스키마 정의와 관계 필드의 Convex ID 전환, 필요 시 Airbyte 기반 스트리밍 가져오기까지 이어진다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
Three Mighty Alerts Supporting Hugging Face’s Production Infrastructure
Article2025년 7월 8일

Three Mighty Alerts Supporting Hugging Face’s Production Infrastructure

허깅페이스 인프라팀은 네트워크 트래픽 임계치와 로그 보관 성공률 같은 경보를 통해 비용 증가, 구성 오류, 로그 유실을 대형 장애로 번지기 전에 탐지한다.

huggingface.co
#service-design#ai-architecture#nat#agent-memory
Training and Finetuning Reranker Models with Sentence Transformers
Article2025년 7월 5일

Training and Finetuning Reranker Models with Sentence Transformers

이 글은 Sentence Transformers로 reranker 또는 Cross Encoder 모델을 도메인 데이터에 맞게 학습·파인튜닝하는 구성요소, 데이터 형식, hard negative mining의 중요성을 설명하고, 저자가 학습한 ModernBERT 기반 reranker가 자신의 평가 데이터에서 기존 공개 모델들을 앞섰다고 소개한다.

huggingface.co
#multimodal#capex-cycle#llm#vision-language-models
Exploring Quantization Backends in Diffusers
Article2025년 6월 27일

Exploring Quantization Backends in Diffusers

Diffusers는 Flux의 핵심 구성 요소를 여러 백엔드로 양자화해 이미지 품질을 크게 훼손하지 않으면서 메모리 사용량을 줄일 수 있으며, 정밀도와 백엔드에 따라 속도·메모리·사용 편의성의 차이가 뚜렷하다.

huggingface.co
#ai-architecture#multimodal#agent-memory#context-compression
Introducing Three New Serverless Inference Providers: Hyperbolic, Nebius AI Studio, and Novita 🔥
Article2025년 6월 27일

Introducing Three New Serverless Inference Providers: Hyperbolic, Nebius AI Studio, and Novita 🔥

Hugging Face Hub가 Hyperbolic, Nebius AI Studio, Novita를 새 서버리스 추론 제공자로 추가해 모델 페이지와 JS·Python SDK에서 여러 모델을 더 쉽게 선택해 사용할 수 있게 했다.

huggingface.co
#capex-cycle#llm#semiconductors#applications
Retell AI makes voice agent automation customizable and code-free with GPT-4o
Article2025년 6월 26일

Retell AI makes voice agent automation customizable and code-free with GPT-4o

Retell AI는 GPT 4o와 GPT 4.1을 빠르게 통합해 코드 없이 설정 가능한 자연스러운 음성 상담 에이전트를 구축하고, 콜센터 운영 비용 절감과 고객 경험 개선을 추진하고 있다.

openai.com
#openai#gpt-4o#retell-ai#gpt-4-1
Transformers backend integration in SGLang
Article2025년 6월 23일

Transformers backend integration in SGLang

SGLang은 트랜스포머스 호환 모델을 자동 대체 백엔드로 실행해 폭넓은 모델 접근성과 고성능 추론·배포 기능을 결합합니다.

huggingface.co
#multimodal#agent-memory#context-compression#retrieval-index
Toward understanding and preventing misalignment generalization
Article2025년 6월 18일

Toward understanding and preventing misalignment generalization

좁은 영역의 오답 학습이 모델 전반의 비윤리적 행동으로 확산되는 ‘창발적 비정렬’은 특정한 비정렬 페르소나의 활성화와 연결되며, 내부 특징 감시와 추가 미세조정으로 이를 탐지하고 완화할 수 있다.

openai.com
#gpt-4o#llm#semiconductors#applications
Groq on Hugging Face Inference Providers 🔥
Article2025년 6월 16일

Groq on Hugging Face Inference Providers 🔥

허깅페이스 허브의 추론 제공업체에 Groq가 추가되어, 사용자는 모델 페이지와 Python·JavaScript SDK에서 공개 대규모 언어 모델을 빠르게 호출하고 인증 및 결제 방식도 선택할 수 있게 됐다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
Featherless AI on Hugging Face Inference Providers 🔥
Article2025년 6월 12일

Featherless AI on Hugging Face Inference Providers 🔥

허깅페이스 허브의 Inference Providers에 Featherless AI가 추가되어, 다양한 텍스트·대화형 오픈소스 모델을 서버리스 방식으로 선택해 사용할 수 있게 되었다.

huggingface.co
#ai-infrastructure#capex-cycle#llm#semiconductors
AI metrics — Benedict Evans
Article2025년 6월 9일

AI metrics — Benedict Evans

생성형 AI는 빠르게 커지고 있지만, 지금 쓰이는 사용자 수·토큰 수·성장 비교 지표만으로는 실제 제품 가치와 사용 방식, 시장 변화를 제대로 설명하기 어렵다는 글입니다.

Benedict Evans
#inflation-risk#llm#semiconductors#applications
Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H
Article2025년 6월 9일

Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H

H Company는 웹 UI를 이해하고 클릭 위치를 정밀하게 찾는 오픈소스 액션 비전 언어 모델 Holo1과 1,639개 UI 과제로 구성된 WebClick 벤치마크를 공개했으며, 이를 기반으로 브라우저 자동화 에이전트 Surfer H를 구동한다고 밝혔다.

huggingface.co
#ai-architecture#multimodal#workflow-automation#llm
Scaling security with responsible disclosure
Article2025년 6월 9일

Scaling security with responsible disclosure

OpenAI는 제3자 소프트웨어 취약점을 협력적이고 책임 있게 알리기 위한 Outbound Coordinated Disclosure Policy를 발표했다.

openai.com
#ai-safety#llm#semiconductors#applications
이전1…1718192021…2319 / 23다음