Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#retrieval-index
Tag248건YouTube 17Article 231

#retrieval-index

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#agent-memory공동문서 248 · 연관도 66%#context-compression공동문서 160 · 연관도 50%#applications공동문서 197 · 연관도 38%#llm공동문서 213 · 연관도 37%#semiconductors공동문서 203 · 연관도 37%#agent-routing공동문서 103 · 연관도 28%#ai-architecture공동문서 82 · 연관도 24%#service-design공동문서 41 · 연관도 14%#gpu공동문서 17 · 연관도 12%#multimodal공동문서 25 · 연관도 11%
Ulysses Sequence Parallelism: Training with Million-Token Contexts
Article2025년 7월 26일

Ulysses Sequence Parallelism: Training with Million-Token Contexts

Ulysses Sequence Parallelism은 긴 시퀀스 학습에서 시퀀스와 어텐션 헤드를 함께 나누고 all to all 통신으로 재배치해, 단일 GPU 메모리 한계를 넘어 수십만~백만 토큰 문맥 학습을 가능하게 하는 방식이다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#retrieval-index
Consilium: When Multiple LLMs Collaborate
Article2025년 7월 19일

Consilium: When Multiple LLMs Collaborate

Consilium은 서로 다른 역할을 맡은 여러 언어 모델이 구조화된 토론과 외부 조사를 거쳐 합의 또는 최종 분석을 도출하도록 만든 시각적 다중 모델 협업 플랫폼이다.

huggingface.co
#ai-architecture#agent-routing#llm#semiconductors
The AI tools for Art Newsletter - Issue 1
Article2025년 7월 16일

The AI tools for Art Newsletter - Issue 1

2024년 창작 AI는 오픈소스 이미지 생성의 구조적 도약과 개인화 기술의 대중화를 이뤘으며, 2025년에는 비디오·오디오·3D 등 더 다양한 형식으로 발전의 중심이 이동하고 있다.

huggingface.co
#ai-architecture#multimodal#agent-memory#agent-routing
Custom Kernels for All from Codex and Claude
Article2025년 7월 16일

Custom Kernels for All from Codex and Claude

CUDA 커널 개발 지식을 에이전트 스킬로 구조화해 Claude와 Codex가 실제 diffusers·transformers 대상의 커널, PyTorch 바인딩, 빌드 구성, 벤치마크까지 완성하도록 한 사례다.

huggingface.co
#ai-architecture#agent-memory#agent-routing#context-compression
Migrating the Hub from Git LFS to Xet
Article2025년 7월 15일

Migrating the Hub from Git LFS to Xet

허깅페이스는 기존 사용자의 작업 방식을 유지하는 브리지와 무중단 백그라운드 마이그레이션을 기반으로, 50만 개 저장소와 20페타바이트 규모의 허브를 깃 대용량 파일 저장소에서 젯으로 전환하고 있다.

huggingface.co
#agent-routing#llm#semiconductors#applications
Migrating data from Postgres to Convex
Article2025년 7월 8일

Migrating data from Postgres to Convex

Postgres 데이터를 Convex로 이전하는 방법은 소규모 데이터의 JSONL 덤프·가져오기에서 시작해, 스키마 정의와 관계 필드의 Convex ID 전환, 필요 시 Airbyte 기반 스트리밍 가져오기까지 이어진다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
Three Mighty Alerts Supporting Hugging Face’s Production Infrastructure
Article2025년 7월 8일

Three Mighty Alerts Supporting Hugging Face’s Production Infrastructure

허깅페이스 인프라팀은 네트워크 트래픽 임계치와 로그 보관 성공률 같은 경보를 통해 비용 증가, 구성 오류, 로그 유실을 대형 장애로 번지기 전에 탐지한다.

huggingface.co
#service-design#ai-architecture#nat#agent-memory
Exploring Quantization Backends in Diffusers
Article2025년 6월 27일

Exploring Quantization Backends in Diffusers

Diffusers는 Flux의 핵심 구성 요소를 여러 백엔드로 양자화해 이미지 품질을 크게 훼손하지 않으면서 메모리 사용량을 줄일 수 있으며, 정밀도와 백엔드에 따라 속도·메모리·사용 편의성의 차이가 뚜렷하다.

huggingface.co
#ai-architecture#multimodal#agent-memory#context-compression
Transformers backend integration in SGLang
Article2025년 6월 23일

Transformers backend integration in SGLang

SGLang은 트랜스포머스 호환 모델을 자동 대체 백엔드로 실행해 폭넓은 모델 접근성과 고성능 추론·배포 기능을 결합합니다.

huggingface.co
#multimodal#agent-memory#context-compression#retrieval-index
Groq on Hugging Face Inference Providers 🔥
Article2025년 6월 16일

Groq on Hugging Face Inference Providers 🔥

허깅페이스 허브의 추론 제공업체에 Groq가 추가되어, 사용자는 모델 페이지와 Python·JavaScript SDK에서 공개 대규모 언어 모델을 빠르게 호출하고 인증 및 결제 방식도 선택할 수 있게 됐다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
KV Cache from scratch in nanoVLM
Article2025년 6월 4일

KV Cache from scratch in nanoVLM

나노브이엘엠에 계층별 키·값 캐시와 사전 채움·순차 해독 구조를 직접 구현해, 자기회귀 생성의 중복 계산을 줄이고 생성 속도를 38% 높인 과정과 원리를 설명한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#retrieval-index
(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware
Article2025년 6월 1일

(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware

FLUX.1 dev의 트랜스포머를 4비트로 양자화하고 LoRA, 8비트 옵티마이저, 체크포인팅, 잠재 표현·텍스트 임베딩 캐시를 결합해 단일 GPU에서 약 10GB 미만의 VRAM으로 미세 조정하는 방법을 설명한다.

huggingface.co
#ai-architecture#agent-memory#capex-cycle#context-compression
Argument Validation without Repetition
Article2025년 5월 5일

Argument Validation without Repetition

Convex와 TypeScript 프로젝트에서 스키마 검증기와 타입 유틸리티를 재사용해 인자 검증 중복을 줄이고, 함수·클라이언트·부분 업데이트까지 일관된 타입 안전성을 유지하는 방법을 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
Expanding on what we missed with sycophancy
Article2025년 5월 2일

Expanding on what we missed with sycophancy

오픈에이아이는 4월 25일 GPT 4o 업데이트에서 사용자 선호 지표와 여러 개선 요소가 결합해 과도한 동조 성향을 키웠으며, 기존 평가 체계가 이를 포착하지 못한 책임을 인정하고 롤백과 평가 절차 보강에 나섰다.

openai.com
#agent-deployment#agent-memory#context-compression#retrieval-index
Welcoming Llama Guard 4 on Hugging Face Hub
Article2025년 4월 29일

Welcoming Llama Guard 4 on Hugging Face Hub

메타가 공개한 라마 가드 4는 텍스트와 이미지를 함께 검사해 유해한 입력과 출력을 분류하는 120억 매개변수의 다국어 안전 모델이며, 프롬프트 주입과 탈옥을 탐지하는 라마 프롬프트 가드 2도 함께 제공된다.

huggingface.co
#privacy-design#ai-architecture#multimodal#agent-memory
Introducing Change Tracking: Launch Week III - Day 1
Article2025년 4월 14일

Introducing Change Tracking: Launch Week III - Day 1

Firecrawl은 웹사이트 스크랩·크롤 결과를 이전 버전과 비교해 신규, 동일, 변경, 제거 상태를 알려주는 Change Tracking 기능을 베타로 공개했다.

Eric Ciarla
#agent-routing#llm#semiconductors#applications
Get your VLM running in 3 simple steps on Intel CPUs
Article2025년 4월 8일

Get your VLM running in 3 simple steps on Intel CPUs

소형 비전 언어 모델 SmolVLM2를 Optimum Intel과 OpenVINO로 변환·양자화·추론하면 별도 GPU 없이도 Intel CPU에서 지연시간과 처리량을 크게 개선할 수 있다.

huggingface.co
#privacy-design#service-design#multimodal#agent-memory
Merging Streams of Convex data
Article2025년 3월 21일

Merging Streams of Convex data

이 글은 Convex에서 convex helpers의 스트림 기능을 사용해 여러 인덱스 범위의 문서를 병합하고, 조인형 확장과 변환, 필터링, 페이지네이션을 더 효율적으로 구현하는 방법을 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
AI Policy @🤗: Response to the White House AI Action Plan RFI
Article2025년 3월 19일

AI Policy @🤗: Response to the White House AI Action Plan RFI

허깅페이스는 미국 백악관 AI 행동계획에 대한 의견서에서 개방형 AI와 오픈 사이언스가 성능 향상, 광범위한 도입, 자원 효율성, 신뢰성 및 보안을 함께 달성하기 위한 핵심 기반이라고 주장했다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
Translate SQL into Convex Queries
Article2025년 3월 19일

Translate SQL into Convex Queries

이 글은 SQL에 익숙한 개발자가 UNION, WHERE 필터, JOIN, DISTINCT, GROUP BY 같은 패턴을 Convex 쿼리와 QueryStreams 방식으로 옮기는 방법을 Slack형 채팅 앱 예시로 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
Hugging Face and VirusTotal collaborate to strengthen AI security
Article2025년 3월 18일

Hugging Face and VirusTotal collaborate to strengthen AI security

Hugging Face는 VirusTotal과 협력해 Hub의 220만 개 이상 공개 모델·데이터셋 저장소 파일을 지속적으로 검사하고, 악성 또는 위험 자산에 대한 가시성을 높이기로 했다.

huggingface.co
#privacy-design#agent-routing#llm#semiconductors
Visualize and understand GPU memory in PyTorch
Article2025년 3월 1일

Visualize and understand GPU memory in PyTorch

PyTorch 메모리 스냅샷으로 GPU 사용량을 단계별로 시각화하고, 학습 과정의 최대 메모리를 구성 요소별로 추정하는 방법을 설명한다.

huggingface.co
#ai-architecture#agent-memory#capex-cycle#context-compression
HuggingFace, IISc partner to supercharge model building on India's diverse languages
Article2025년 2월 27일

HuggingFace, IISc partner to supercharge model building on India's diverse languages

허깅페이스와 인도과학원·아트파크는 인도 전역의 언어·방언·지역·인구통계적 다양성을 담은 오픈소스 다중양식 데이터셋 ‘바니’의 접근성과 활용성을 높이기 위해 협력한다.

huggingface.co
#multimodal#llm#semiconductors#vision-language-models
Visual Document Retrieval Goes Multilingual
Article2025년 2월 26일

Visual Document Retrieval Goes Multilingual

LlamaIndex는 문서 화면을 OCR 없이 단일 벡터로 표현하는 다국어 검색 모델과 약 50만 건 규모의 공개 학습 데이터를 선보였으며, 검색 성능·추론 속도·언어 간 검색·저장 효율을 함께 개선했다.

huggingface.co
#multimodal#agent-memory#capex-cycle#retrieval-index
이전1…78910119 / 11다음