Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#context-compression
Tag404건YouTube 25Article 379

#context-compression

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

연관 태그

#agent-memory공동문서 286 · 연관도 60%#prompt-library공동문서 154 · 연관도 53%#semiconductors공동문서 354 · 연관도 50%#applications공동문서 334 · 연관도 50%#retrieval-index공동문서 160 · 연관도 49%#llm공동문서 342 · 연관도 47%#agent-routing공동문서 116 · 연관도 25%#agent-deployment공동문서 80 · 연관도 21%#ai-architecture공동문서 92 · 연관도 21%#compute공동문서 34 · 연관도 17%
Transformers backend integration in SGLang
Article2025년 6월 23일

Transformers backend integration in SGLang

SGLang은 트랜스포머스 호환 모델을 자동 대체 백엔드로 실행해 폭넓은 모델 접근성과 고성능 추론·배포 기능을 결합합니다.

huggingface.co
#multimodal#agent-memory#context-compression#retrieval-index
Efficient Request Queueing – Optimizing LLM Performance
Article2025년 6월 17일

Efficient Request Queueing – Optimizing LLM Performance

다중 사용자 환경의 대규모 언어 모델 서빙에서는 사용자별 공정 스케줄링과 백엔드 지표 기반의 동적 요청 제어를 결합해야 지연 시간을 줄이면서 그래픽 처리 장치 활용률을 유지할 수 있다.

huggingface.co
#context-compression#prompt-library#api-vllm-api#llm
Groq on Hugging Face Inference Providers 🔥
Article2025년 6월 16일

Groq on Hugging Face Inference Providers 🔥

허깅페이스 허브의 추론 제공업체에 Groq가 추가되어, 사용자는 모델 페이지와 Python·JavaScript SDK에서 공개 대규모 언어 모델을 빠르게 호출하고 인증 및 결제 방식도 선택할 수 있게 됐다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
Featherless AI on Hugging Face Inference Providers 🔥
Article2025년 6월 12일

Featherless AI on Hugging Face Inference Providers 🔥

허깅페이스 허브의 Inference Providers에 Featherless AI가 추가되어, 다양한 텍스트·대화형 오픈소스 모델을 서버리스 방식으로 선택해 사용할 수 있게 되었다.

huggingface.co
#ai-infrastructure#capex-cycle#llm#semiconductors
AI metrics — Benedict Evans
Article2025년 6월 9일

AI metrics — Benedict Evans

생성형 AI는 빠르게 커지고 있지만, 지금 쓰이는 사용자 수·토큰 수·성장 비교 지표만으로는 실제 제품 가치와 사용 방식, 시장 변화를 제대로 설명하기 어렵다는 글입니다.

Benedict Evans
#inflation-risk#llm#semiconductors#applications
Scaling security with responsible disclosure
Article2025년 6월 9일

Scaling security with responsible disclosure

OpenAI는 제3자 소프트웨어 취약점을 협력적이고 책임 있게 알리기 위한 Outbound Coordinated Disclosure Policy를 발표했다.

openai.com
#ai-safety#llm#semiconductors#applications
How Answer HQ Powers AI Customer Support for Businesses with Firecrawl
Article2025년 6월 5일

How Answer HQ Powers AI Customer Support for Businesses with Firecrawl

Answer HQ는 소규모 기업의 기존 웹사이트 콘텐츠를 AI 고객지원 도우미에 연결하기 위해 Firecrawl을 웹사이트 가져오기 기능의 핵심 인프라로 사용한다.

Eric Ciarla
#agent-routing#workflow-automation#llm#semiconductors
AI Agents (and humans) do better with good abstractions
Article2025년 6월 3일

AI Agents (and humans) do better with good abstractions

Convex의 사례는 좋은 추상화가 개발자뿐 아니라 AI 에이전트도 복잡한 풀스택 앱을 더 안정적으로 만들게 한다는 점을 보여준다.

stack.convex.dev
#service-design#ai-architecture#agent-routing#context-compression
(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware
Article2025년 6월 1일

(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware

FLUX.1 dev의 트랜스포머를 4비트로 양자화하고 LoRA, 8비트 옵티마이저, 체크포인팅, 잠재 표현·텍스트 임베딩 캐시를 결합해 단일 GPU에서 약 10GB 미만의 VRAM으로 미세 조정하는 방법을 설명한다.

huggingface.co
#ai-architecture#agent-memory#capex-cycle#context-compression
Creating websites in minutes with AI Website Builder
Article2025년 5월 29일

Creating websites in minutes with AI Website Builder

Wix는 OpenAI 모델과 자사의 웹사이트 제작 경험을 결합해, 사용자가 대화만으로 콘텐츠·이미지·레이아웃·업무 기능을 갖춘 웹사이트를 몇 분 안에 만들 수 있도록 했다.

openai.com
#openai#privacy-design#context-compression#prompt-library
CodeAgents + Structure: A Better Way to Execute Actions
Article2025년 5월 28일

CodeAgents + Structure: A Better Way to Execute Actions

이 글은 CodeAgent의 유연한 코드 실행 방식에 구조화된 JSON 출력을 결합하면, 충분히 강한 모델에서 파싱 안정성과 추론 명시성이 높아져 여러 벤치마크 성능이 개선된다고 설명한다.

huggingface.co
#anthropic#agent-routing#llm#semiconductors
The Transformers Library: standardizing model definitions
Article2025년 5월 15일

The Transformers Library: standardizing model definitions

트랜스포머스는 모델 정의를 표준화해 하나의 아키텍처 구현이 학습·추론·배포·로컬 실행 도구 전반으로 빠르게 이어지는 생태계의 중심축이 되고자 한다.

huggingface.co
#ai-architecture#llm#semiconductors#applications
Argument Validation without Repetition
Article2025년 5월 5일

Argument Validation without Repetition

Convex와 TypeScript 프로젝트에서 스키마 검증기와 타입 유틸리티를 재사용해 인자 검증 중복을 줄이고, 함수·클라이언트·부분 업데이트까지 일관된 타입 안전성을 유지하는 방법을 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
Expanding on what we missed with sycophancy
Article2025년 5월 2일

Expanding on what we missed with sycophancy

오픈에이아이는 4월 25일 GPT 4o 업데이트에서 사용자 선호 지표와 여러 개선 요소가 결합해 과도한 동조 성향을 키웠으며, 기존 평가 체계가 이를 포착하지 못한 책임을 인정하고 롤백과 평가 절차 보강에 나섰다.

openai.com
#agent-deployment#agent-memory#context-compression#retrieval-index
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.
Article2025년 5월 1일

Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models.

Falcon Edge는 하나의 사전학습 과정에서 범용 bfloat16 모델, 삼진 가중치 기반 BitNet 모델, 미세조정용 사전 양자화 모델을 함께 제공하는 1.58비트 언어 모델 시리즈다.

huggingface.co
#ai-architecture#multimodal#llm#vision-language-models
Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models
Article2025년 4월 30일

Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models

이 글은 Qwen3 8B를 Intel® Core™ Ultra에서 더 빠르게 실행하기 위해 OpenVINO.GenAI의 추측 디코딩과 깊이 가지치기된 Qwen3 0.6B 드래프트 모델을 결합해 약 1.4배 속도 향상을 얻은 과정을 설명한다.

huggingface.co
#service-design#agent-routing#capex-cycle#context-compression
How to Build an MCP Server with Gradio
Article2025년 4월 30일

How to Build an MCP Server with Gradio

그라디오는 기존 파이썬 함수를 도구로 자동 변환하고 실행 옵션 하나로 웹 인터페이스와 모델 콘텍스트 프로토콜 서버를 함께 제공한다.

huggingface.co
#ai-architecture#agent-deployment#context-compression#prompt-library
How Botpress Populates AI Chatbot Knowledge Bases at Scale with Firecrawl
Article2025년 4월 21일

How Botpress Populates AI Chatbot Knowledge Bases at Scale with Firecrawl

Botpress는 Firecrawl을 도입해 웹사이트 콘텐츠를 챗봇 지식 베이스로 가져오는 과정을 자동화하고, 자체 HTML 마크다운 처리 부담을 크게 줄였다.

Eric Ciarla
#agent-deployment#agent-routing#llm#semiconductors
LLM Inference on Edge: A Fun and Easy Guide to run LLMs via React Native on your Phone!
Article2025년 4월 21일

LLM Inference on Edge: A Fun and Easy Guide to run LLMs via React Native on your Phone!

이 글은 React Native 앱에서 Hugging Face의 GGUF 모델을 내려받고 llama.rn으로 로컬 실행해, Android와 iOS에서 오프라인 LLM 채팅 앱을 만드는 과정을 안내한다.

huggingface.co
#privacy-design#service-design#llm#semiconductors
Announcing FIRE-1, Our Web Action Agent: Launch Week III - Day 2
Article2025년 4월 15일

Announcing FIRE-1, Our Web Action Agent: Launch Week III - Day 2

Firecrawl은 Launch Week III 2일차에 복잡한 웹사이트를 탐색하고 버튼·검색폼 등과 상호작용해 숨은 데이터를 추출하는 웹 액션 에이전트 FIRE 1을 발표했다.

Eric Ciarla
#agent-routing#context-compression#prompt-library#llm
Introducing Change Tracking: Launch Week III - Day 1
Article2025년 4월 14일

Introducing Change Tracking: Launch Week III - Day 1

Firecrawl은 웹사이트 스크랩·크롤 결과를 이전 버전과 비교해 신규, 동일, 변경, 제거 상태를 알려주는 Change Tracking 기능을 베타로 공개했다.

Eric Ciarla
#agent-routing#llm#semiconductors#applications
SmolVLM2: Bringing Video Understanding to Every Device
Article2025년 4월 8일

SmolVLM2: Bringing Video Understanding to Every Device

SmolVLM2는 2.2B·500M·256M 세 가지 크기로 영상 이해의 높은 메모리 효율과 온디바이스 실행 가능성을 제시하며, Transformers와 MLX를 통해 출시 직후부터 다양한 환경에서 활용할 수 있도록 공개된 소형 비전·영상 언어 모델군이다.

huggingface.co
#multimodal#capex-cycle#context-compression#prompt-library
Merging Streams of Convex data
Article2025년 3월 21일

Merging Streams of Convex data

이 글은 Convex에서 convex helpers의 스트림 기능을 사용해 여러 인덱스 범위의 문서를 병합하고, 조인형 확장과 변환, 필터링, 페이지네이션을 더 효율적으로 구현하는 방법을 설명한다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
AI Policy @🤗: Response to the White House AI Action Plan RFI
Article2025년 3월 19일

AI Policy @🤗: Response to the White House AI Action Plan RFI

허깅페이스는 미국 백악관 AI 행동계획에 대한 의견서에서 개방형 AI와 오픈 사이언스가 성능 향상, 광범위한 도입, 자원 효율성, 신뢰성 및 보안을 함께 달성하기 위한 핵심 기반이라고 주장했다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
이전1…131415161715 / 17다음