Docs faviconDOCS우성짱의 문서
전체YouTubeArticleTagsAuthorsHub
홈/태그 찾기/#llm
Tag1292건YouTube 48Article 1244

#llm

이 태그와 연결된 문서를 한곳에서 모아보고, 함께 자주 등장하는 연관 태그까지 이어서 탐색할 수 있습니다.

Alias / 동의어

large-language-models

연관 태그

#semiconductors공동문서 1040 · 연관도 83%#applications공동문서 974 · 연관도 81%#agent-routing공동문서 512 · 연관도 61%#agent-memory공동문서 468 · 연관도 55%#ai-architecture공동문서 403 · 연관도 51%#privacy-design공동문서 384 · 연관도 50%#context-compression공동문서 342 · 연관도 47%#agent-deployment공동문서 312 · 연관도 47%#service-design공동문서 286 · 연관도 42%#retrieval-index공동문서 213 · 연관도 37%
Upskill your LLMs With Gradio MCP Servers
Article2025년 7월 9일

Upskill your LLMs With Gradio MCP Servers

Gradio의 MCP 지원을 활용하면 Hugging Face Spaces의 다양한 AI 도구를 Cursor 같은 LLM 클라이언트에 연결해 이미지 편집, 영상 전사, OCR, 음성 합성 등의 새로운 기능을 부여할 수 있습니다.

huggingface.co
#capex-cycle#llm#semiconductors#applications
Migrating data from Postgres to Convex
Article2025년 7월 8일

Migrating data from Postgres to Convex

Postgres 데이터를 Convex로 이전하는 방법은 소규모 데이터의 JSONL 덤프·가져오기에서 시작해, 스키마 정의와 관계 필드의 Convex ID 전환, 필요 시 Airbyte 기반 스트리밍 가져오기까지 이어진다.

stack.convex.dev
#agent-routing#llm#semiconductors#applications
SmolLM3: smol, multilingual, long-context reasoner
Article2025년 7월 8일

SmolLM3: smol, multilingual, long-context reasoner

SmolLM3는 공개 데이터와 학습 도구로 구축한 30억 매개변수 모델로, 11조 개가 넘는 토큰의 단계별 사전학습과 장문·추론 중간학습을 결합해 다국어, 최대 128K 문맥, 추론·비추론 이중 모드를 지원한다.

huggingface.co
#long-context#privacy-design#ai-architecture#llm
Three Mighty Alerts Supporting Hugging Face’s Production Infrastructure
Article2025년 7월 8일

Three Mighty Alerts Supporting Hugging Face’s Production Infrastructure

허깅페이스 인프라팀은 네트워크 트래픽 임계치와 로그 보관 성공률 같은 경보를 통해 비용 증가, 구성 오류, 로그 유실을 대형 장애로 번지기 전에 탐지한다.

huggingface.co
#service-design#ai-architecture#nat#agent-memory
Training and Finetuning Reranker Models with Sentence Transformers
Article2025년 7월 5일

Training and Finetuning Reranker Models with Sentence Transformers

이 글은 Sentence Transformers로 reranker 또는 Cross Encoder 모델을 도메인 데이터에 맞게 학습·파인튜닝하는 구성요소, 데이터 형식, hard negative mining의 중요성을 설명하고, 저자가 학습한 ModernBERT 기반 reranker가 자신의 평가 데이터에서 기존 공개 모델들을 앞섰다고 소개한다.

huggingface.co
#multimodal#capex-cycle#llm#vision-language-models
Open Researcher, our AI Agent That Uses Firecrawl Tools During Research
Article2025년 7월 1일

Open Researcher, our AI Agent That Uses Firecrawl Tools During Research

Open Researcher는 정해진 워크플로를 늘리는 대신 Anthropic의 interleaved thinking과 Firecrawl 웹 데이터 도구를 결합해 연구 과정에서 AI가 스스로 다음 행동을 판단하도록 만든 오픈소스 연구 에이전트다.

Eric Ciarla
#anthropic#agent-routing#ai-safety#search-advertising
Exploring Quantization Backends in Diffusers
Article2025년 6월 27일

Exploring Quantization Backends in Diffusers

Diffusers는 Flux의 핵심 구성 요소를 여러 백엔드로 양자화해 이미지 품질을 크게 훼손하지 않으면서 메모리 사용량을 줄일 수 있으며, 정밀도와 백엔드에 따라 속도·메모리·사용 편의성의 차이가 뚜렷하다.

huggingface.co
#ai-architecture#multimodal#agent-memory#context-compression
Introducing Three New Serverless Inference Providers: Hyperbolic, Nebius AI Studio, and Novita 🔥
Article2025년 6월 27일

Introducing Three New Serverless Inference Providers: Hyperbolic, Nebius AI Studio, and Novita 🔥

Hugging Face Hub가 Hyperbolic, Nebius AI Studio, Novita를 새 서버리스 추론 제공자로 추가해 모델 페이지와 JS·Python SDK에서 여러 모델을 더 쉽게 선택해 사용할 수 있게 했다.

huggingface.co
#capex-cycle#llm#semiconductors#applications
Transformers backend integration in SGLang
Article2025년 6월 23일

Transformers backend integration in SGLang

SGLang은 트랜스포머스 호환 모델을 자동 대체 백엔드로 실행해 폭넓은 모델 접근성과 고성능 추론·배포 기능을 결합합니다.

huggingface.co
#multimodal#agent-memory#context-compression#retrieval-index
Preparing for future AI capabilities in biology
Article2025년 6월 18일

Preparing for future AI capabilities in biology

오픈에이아이는 생물학 분야의 인공지능이 과학적 발견을 크게 앞당기는 동시에 생물학적 위협의 진입 장벽을 낮출 수 있다고 보고, 고위험 역량에 도달하기 전부터 모델 훈련·탐지·집행·레드팀·보안·정부 협력을 결합한 다층 방어 체계를 준비하고 있다.

openai.com
#openai#privacy-design#agent-routing#llm
Toward understanding and preventing misalignment generalization
Article2025년 6월 18일

Toward understanding and preventing misalignment generalization

좁은 영역의 오답 학습이 모델 전반의 비윤리적 행동으로 확산되는 ‘창발적 비정렬’은 특정한 비정렬 페르소나의 활성화와 연결되며, 내부 특징 감시와 추가 미세조정으로 이를 탐지하고 완화할 수 있다.

openai.com
#gpt-4o#llm#semiconductors#applications
Efficient Request Queueing – Optimizing LLM Performance
Article2025년 6월 17일

Efficient Request Queueing – Optimizing LLM Performance

다중 사용자 환경의 대규모 언어 모델 서빙에서는 사용자별 공정 스케줄링과 백엔드 지표 기반의 동적 요청 제어를 결합해야 지연 시간을 줄이면서 그래픽 처리 장치 활용률을 유지할 수 있다.

huggingface.co
#context-compression#prompt-library#api-vllm-api#llm
Groq on Hugging Face Inference Providers 🔥
Article2025년 6월 16일

Groq on Hugging Face Inference Providers 🔥

허깅페이스 허브의 추론 제공업체에 Groq가 추가되어, 사용자는 모델 페이지와 Python·JavaScript SDK에서 공개 대규모 언어 모델을 빠르게 호출하고 인증 및 결제 방식도 선택할 수 있게 됐다.

huggingface.co
#llm#semiconductors#applications#agent-deployment
Featherless AI on Hugging Face Inference Providers 🔥
Article2025년 6월 12일

Featherless AI on Hugging Face Inference Providers 🔥

허깅페이스 허브의 Inference Providers에 Featherless AI가 추가되어, 다양한 텍스트·대화형 오픈소스 모델을 서버리스 방식으로 선택해 사용할 수 있게 되었다.

huggingface.co
#ai-infrastructure#capex-cycle#llm#semiconductors
How Long Prompts Block Other Requests - Optimizing LLM Performance
Article2025년 6월 12일

How Long Prompts Block Other Requests - Optimizing LLM Performance

긴 프롬프트가 포함된 요청은 프리필 대기열과 동시 디코딩을 지연시키며, 요청 병렬 프리필은 첫 토큰 지연을 줄이고 프리필·디코드 분리 구조는 토큰 생성 간섭을 완화한다.

huggingface.co
#ai-architecture#agent-deployment#agent-routing#workflow-automation
Introducing Training Cluster as a Service - a new collaboration with NVIDIA
Article2025년 6월 11일

Introducing Training Cluster as a Service - a new collaboration with NVIDIA

Hugging Face와 NVIDIA는 연구기관과 기업이 필요한 시점·규모·기간에 맞춰 대규모 GPU 클러스터를 요청하고 학습 작업에 활용할 수 있는 Training Cluster as a Service를 발표했다.

huggingface.co
#nvidia#service-design#ai-infrastructure#capex-cycle
AI metrics — Benedict Evans
Article2025년 6월 9일

AI metrics — Benedict Evans

생성형 AI는 빠르게 커지고 있지만, 지금 쓰이는 사용자 수·토큰 수·성장 비교 지표만으로는 실제 제품 가치와 사용 방식, 시장 변화를 제대로 설명하기 어렵다는 글입니다.

Benedict Evans
#inflation-risk#llm#semiconductors#applications
Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H
Article2025년 6월 9일

Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H

H Company는 웹 UI를 이해하고 클릭 위치를 정밀하게 찾는 오픈소스 액션 비전 언어 모델 Holo1과 1,639개 UI 과제로 구성된 WebClick 벤치마크를 공개했으며, 이를 기반으로 브라우저 자동화 에이전트 Surfer H를 구동한다고 밝혔다.

huggingface.co
#ai-architecture#multimodal#workflow-automation#llm
Scaling security with responsible disclosure
Article2025년 6월 9일

Scaling security with responsible disclosure

OpenAI는 제3자 소프트웨어 취약점을 협력적이고 책임 있게 알리기 위한 Outbound Coordinated Disclosure Policy를 발표했다.

openai.com
#ai-safety#llm#semiconductors#applications
How Answer HQ Powers AI Customer Support for Businesses with Firecrawl
Article2025년 6월 5일

How Answer HQ Powers AI Customer Support for Businesses with Firecrawl

Answer HQ는 소규모 기업의 기존 웹사이트 콘텐츠를 AI 고객지원 도우미에 연결하기 위해 Firecrawl을 웹사이트 가져오기 기능의 핵심 인프라로 사용한다.

Eric Ciarla
#agent-routing#workflow-automation#llm#semiconductors
How we’re responding to The New York Times’ data demands in order to protect user privacy
Article2025년 6월 5일

How we’re responding to The New York Times’ data demands in order to protect user privacy

OpenAI는 사용자 데이터의 무기한 보존 의무가 종료되어 표준 30일 삭제 정책으로 복귀했지만, 뉴욕타임스가 요구한 2025년 4월부터 9월까지의 제한된 과거 데이터는 법적 의무에 따라 별도로 보호하고 있다고 밝혔다.

openai.com
#openai#privacy-design#ai-safety#llm
KV Cache from scratch in nanoVLM
Article2025년 6월 4일

KV Cache from scratch in nanoVLM

나노브이엘엠에 계층별 키·값 캐시와 사전 채움·순차 해독 구조를 직접 구현해, 자기회귀 생성의 중복 계산을 줄이고 생성 속도를 38% 높인 과정과 원리를 설명한다.

huggingface.co
#ai-architecture#multimodal#agent-memory#retrieval-index
AI Agents (and humans) do better with good abstractions
Article2025년 6월 3일

AI Agents (and humans) do better with good abstractions

Convex의 사례는 좋은 추상화가 개발자뿐 아니라 AI 에이전트도 복잡한 풀스택 앱을 더 안정적으로 만들게 한다는 점을 보여준다.

stack.convex.dev
#service-design#ai-architecture#agent-routing#context-compression
Announcing /search: Discover and scrape the web with one API call
Article2025년 6월 3일

Announcing /search: Discover and scrape the web with one API call

Firecrawl은 웹페이지 발견과 콘텐츠 추출을 한 번의 API 호출로 처리하는 /search 엔드포인트를 출시했다고 발표했다.

Eric Ciarla
#agent-routing#ai-distribution#search-advertising#zero-click-search
이전1…4445464748…5446 / 54다음